İçindekiler (11)
MathReader & LaTeX Grabber is a specialized, privacy-centric browser extension engineered to bridge the critical gap between web-rendered scientific literature and local digital document pipelines. Built on the Manifest V3 standard, the tool operates entirely within the client runtime without relaying document state or user interactions to external telemetry endpoints.
The extension targets a persistent friction point in the computational sciences: the non-destructive extraction, normalization, and semantic archiving of mathematically dense web content (such as preprints, forum threads, reference encyclopedias, and computational notebooks) into portable, publication-grade targets including Tokenized CommonMark (Markdown), Standalone Styled HTML, and Vector-Preserved Print PDF.
The foundational principle of the extension is semantic idempotency: mathematically complex DOM trees must retain their precise typographic, structural, and algebraic integrity across formats, solving the persistent problem of corrupted formulas, missing graphical dependencies, and invasive navigational boilerplate.
1. Core Functional Taxonomy: How the Extension Operates
A systematic audit of the implementation reveals five primary subsystems that cooperate across the browser runtime:
A. Universal Mathematical Detection Engine
The extension deploys an ambient DOM inspection routine matching rendered nodes across every major mathematical web layout system:
KaTeXrendering trees (.katex,.katex-display)MathJaxversions 2 and 3 output envelopes (mjx-container,.MathJax,.MathJax_Display)Temmlstructures (.temml)- W3C Native
<math>(MathML) subtrees LaTeXMLoutputs (.ltx_Math,.ltx_equation)- MediaWiki math extensions (
.mwe-math-element, along with high-DPI inline/display SVG and fallback image representations)
Upon cursor registration over any recognized mathematical component, the extension mounts a low-overhead, absolute-positioned floating UI node (.lyc-overlay). This provides zero-friction clipboard extraction with immediate visual feedback (Kopyalandı ✓) and is fully accessible via a global keyboard shortcut (Alt+C).
B. Recursive Abstract Syntax Tree (AST) MathML-to-LaTeX Transpiler
A significant challenge in web scraping occurs when platforms strip raw TeX annotations from rendered output, leaving only structural MathML. The internal mathmlToLatex() transpiler performs recursive node analysis directly against the browser DOM:
- Rational Fractions and Radicals: Correctly balances numerator and denominator boundaries using
\frac{}{}, alongside arbitrary root indexes via\sqrt[]{}. - Multi-tier Scripting: Resolves complex subscript, superscript, and simultaneous limit placements via
msub,msup,msubsup,munder,mover, andmunderover. - Accent and Vector Mapping: Accurately translates overhead operators into formal TeX modifiers including
\vec{},\hat{},\bar{},\dot{},\ddot{}, and generalized\overset{}{}/\underset{}{}constructions. - Tabular and Linear Systems: Converts
<mtable>,<mtr>, and<mtd>configurations into canonical\begin{matrix} ... \end{matrix}LaTeX tabular syntax. - Exhaustive Symbol and Blackboard Normalization: Maps hundreds of Unicode entities to LaTeX commands, including the entire Greek alphabet (lower and upper case), calculus primitives, contour integrals, and blackboard-bold sets (
\mathbb{R},\mathbb{C},\mathbb{N},\mathbb{Z},\mathbb{Q}).
C. Heuristic Content Isolation & Platform-Specific Adapters
Unlike naive scrapers that capture redundant document hierarchies, the extension evaluates document topology through automated scoring:
- StackExchange / StackOverflow Family Adapter: Isolates the root question body, visually marks the verified solution with green emphasis (
Accepted Answer ✓), and iterates subsequent responses demarcated by structured horizontal dividers. - Wikipedia Academic Citation Normalizer: Strips interactive editing hooks (
[edit]) and diagnostic cleanup tags, reconstructs citation anchors into clean inline suprascripts, and standardizes reference bibliographies within dedicatedol.wiki-referenceslists. - Interactive Computational Notebook Support: Directly isolates notebook cells and Markdown blocks within
#notebook-containerand.jp-Notebookruntimes. - Algorithmic Scorer Fallback: Calculates paragraph volume, filters out high link-density ratios (rejecting promotional link farms where anchor tags exceed 40% of the body), and applies positive scoring coefficients to mathematical density.
- User Selection Fallback Mode: Whenever an active text selection exists, the scoring heuristic yields priority, isolating only the user-designated DOM fragment while preserving math and image resolution.
D. Parallel Base64 In-Memory Graphic Resolution
To permanently resolve the "broken image" failure mode inherent to offline documents, the runtime processes images before serializing the DOM:
// Dual-layer graphic resolution pipeline:
// 1. Concurrent network fetch with Cache-Control: force-cache and no-referrer
// 2. Offscreen HTML5 Canvas pixel extraction fallback for authenticated memory states
Candidate image paths are evaluated across srcset arrays, high-resolution data-src, and <picture><source> tags. Resolved streams are converted into immutable data:image/png;base64 strings, ensuring exported files render completely offline without depending on external hosting.
E. Zero-Tab Sandboxed Silent Iframe Printing
To generate clean PDF documents without exposing extension internal paths (such as chrome-extension://), the script creates a sandboxed, zero-dimension iframe (position: fixed; left: -9999px; width: 0; height: 0;). The sanitized document, complete with injected KaTeX typography and print pagination rules (page-break-inside: avoid), is injected into the frame to trigger window.print() directly before teardown.
2. Specific Critical Problems Solved by the Architecture
Problem 1: Destructive Copy-Paste of Formulas
Standard browser selection turns mathematical expressions into disjointed plain text strings, stripping denominators, subscripts, and fractional layout. This tool provides deterministic extraction directly to LaTeX with customizable delimitations:
- Raw TeX:
\lim_{x \to 0} \frac{\sin x}{x} = 1 - Inline Math:
$\lim_{x \to 0} \frac{\sin x}{x} = 1$ - Display Block:
$$\n\lim_{x \to 0} \frac{\sin x}{x} = 1\n$$
Problem 2: Print Degradation and Web Clutter
Standard web printing includes sidebars, floating sticky banners, tracking pixels, and author biographies. The extension enforces a strict semantic whitelist (ALLOWED_TAGS), completely removing non-whitelisted elements, unverified style sheets, and tracking attributes to produce a clean, book-like layout.
Problem 3: Markdown Pipeline Incompatibilities
When capturing documentation for personal knowledge bases (e.g., Obsidian, Logseq, or GitHub-flavored Markdown), formula formatting often gets corrupted by Markdown parsers. The extension utilizes a Tokenized Markdown Engine: formulas are extracted and swapped with temporary identifiers (___MATH_TOKEN_N___), HTML tags are mapped to CommonMark primitives, and mathematical delimiters are restored at the final stage.
3. Security Architecture & Engineering Quality
The extension implements a hardened security posture:
- Zero External Data Transmission: Operates entirely locally; user data never leaves the browser context.
- Proactive XSS Neutralization: Strips executable attributes (including
onclick,onerror, andsrcdoc) across all processed DOM nodes. - Strict Isolation: Hyperlinks are validated against safe protocols (
http:,https:,mailto:) and forced to userel="noopener noreferrer".
Henüz yorum yapılmamış. İlk yorumu siz yapın!