Comprehensive Technical & Functional Review of MathReader

İçindekiler (11)

MathReader & LaTeX Grabber is a specialized, privacy-centric browser extension engineered to bridge the critical gap between web-rendered scientific literature and local digital document pipelines. Built on the Manifest V3 standard, the tool operates entirely within the client runtime without relaying document state or user interactions to external telemetry endpoints.

This plugin has not been released yet. Once released, it will be available for download from this post as well.

The extension targets a persistent friction point in the computational sciences: the non-destructive extraction, normalization, and semantic archiving of mathematically dense web content (such as preprints, forum threads, reference encyclopedias, and computational notebooks) into portable, publication-grade targets including Tokenized CommonMark (Markdown), Standalone Styled HTML, and Vector-Preserved Print PDF.

The foundational principle of the extension is semantic idempotency: mathematically complex DOM trees must retain their precise typographic, structural, and algebraic integrity across formats, solving the persistent problem of corrupted formulas, missing graphical dependencies, and invasive navigational boilerplate.

1. Core Functional Taxonomy: How the Extension Operates

A systematic audit of the implementation reveals five primary subsystems that cooperate across the browser runtime:

A. Universal Mathematical Detection Engine

The extension deploys an ambient DOM inspection routine matching rendered nodes across every major mathematical web layout system:

  • KaTeX rendering trees (.katex, .katex-display)
  • MathJax versions 2 and 3 output envelopes (mjx-container, .MathJax, .MathJax_Display)
  • Temml structures (.temml)
  • W3C Native <math> (MathML) subtrees
  • LaTeXML outputs (.ltx_Math, .ltx_equation)
  • MediaWiki math extensions (.mwe-math-element, along with high-DPI inline/display SVG and fallback image representations)

Upon cursor registration over any recognized mathematical component, the extension mounts a low-overhead, absolute-positioned floating UI node (.lyc-overlay). This provides zero-friction clipboard extraction with immediate visual feedback (Kopyalandı ✓) and is fully accessible via a global keyboard shortcut (Alt+C).

B. Recursive Abstract Syntax Tree (AST) MathML-to-LaTeX Transpiler

A significant challenge in web scraping occurs when platforms strip raw TeX annotations from rendered output, leaving only structural MathML. The internal mathmlToLatex() transpiler performs recursive node analysis directly against the browser DOM:

  1. Rational Fractions and Radicals: Correctly balances numerator and denominator boundaries using \frac{}{}, alongside arbitrary root indexes via \sqrt[]{}.
  2. Multi-tier Scripting: Resolves complex subscript, superscript, and simultaneous limit placements via msub, msup, msubsup, munder, mover, and munderover.
  3. Accent and Vector Mapping: Accurately translates overhead operators into formal TeX modifiers including \vec{}, \hat{}, \bar{}, \dot{}, \ddot{}, and generalized \overset{}{} / \underset{}{} constructions.
  4. Tabular and Linear Systems: Converts <mtable>, <mtr>, and <mtd> configurations into canonical \begin{matrix} ... \end{matrix} LaTeX tabular syntax.
  5. Exhaustive Symbol and Blackboard Normalization: Maps hundreds of Unicode entities to LaTeX commands, including the entire Greek alphabet (lower and upper case), calculus primitives, contour integrals, and blackboard-bold sets (\mathbb{R}, \mathbb{C}, \mathbb{N}, \mathbb{Z}, \mathbb{Q}).

C. Heuristic Content Isolation & Platform-Specific Adapters

Unlike naive scrapers that capture redundant document hierarchies, the extension evaluates document topology through automated scoring:

  • StackExchange / StackOverflow Family Adapter: Isolates the root question body, visually marks the verified solution with green emphasis (Accepted Answer ✓), and iterates subsequent responses demarcated by structured horizontal dividers.
  • Wikipedia Academic Citation Normalizer: Strips interactive editing hooks ([edit]) and diagnostic cleanup tags, reconstructs citation anchors into clean inline suprascripts, and standardizes reference bibliographies within dedicated ol.wiki-references lists.
  • Interactive Computational Notebook Support: Directly isolates notebook cells and Markdown blocks within #notebook-container and .jp-Notebook runtimes.
  • Algorithmic Scorer Fallback: Calculates paragraph volume, filters out high link-density ratios (rejecting promotional link farms where anchor tags exceed 40% of the body), and applies positive scoring coefficients to mathematical density.
  • User Selection Fallback Mode: Whenever an active text selection exists, the scoring heuristic yields priority, isolating only the user-designated DOM fragment while preserving math and image resolution.

D. Parallel Base64 In-Memory Graphic Resolution

To permanently resolve the "broken image" failure mode inherent to offline documents, the runtime processes images before serializing the DOM:

Kod
// Dual-layer graphic resolution pipeline:
// 1. Concurrent network fetch with Cache-Control: force-cache and no-referrer
// 2. Offscreen HTML5 Canvas pixel extraction fallback for authenticated memory states

Candidate image paths are evaluated across srcset arrays, high-resolution data-src, and <picture><source> tags. Resolved streams are converted into immutable data:image/png;base64 strings, ensuring exported files render completely offline without depending on external hosting.

E. Zero-Tab Sandboxed Silent Iframe Printing

To generate clean PDF documents without exposing extension internal paths (such as chrome-extension://), the script creates a sandboxed, zero-dimension iframe (position: fixed; left: -9999px; width: 0; height: 0;). The sanitized document, complete with injected KaTeX typography and print pagination rules (page-break-inside: avoid), is injected into the frame to trigger window.print() directly before teardown.

2. Specific Critical Problems Solved by the Architecture

Problem 1: Destructive Copy-Paste of Formulas

Standard browser selection turns mathematical expressions into disjointed plain text strings, stripping denominators, subscripts, and fractional layout. This tool provides deterministic extraction directly to LaTeX with customizable delimitations:

  • Raw TeX: \lim_{x \to 0} \frac{\sin x}{x} = 1
  • Inline Math: $\lim_{x \to 0} \frac{\sin x}{x} = 1$
  • Display Block: $$\n\lim_{x \to 0} \frac{\sin x}{x} = 1\n$$

Problem 2: Print Degradation and Web Clutter

Standard web printing includes sidebars, floating sticky banners, tracking pixels, and author biographies. The extension enforces a strict semantic whitelist (ALLOWED_TAGS), completely removing non-whitelisted elements, unverified style sheets, and tracking attributes to produce a clean, book-like layout.

Problem 3: Markdown Pipeline Incompatibilities

When capturing documentation for personal knowledge bases (e.g., Obsidian, Logseq, or GitHub-flavored Markdown), formula formatting often gets corrupted by Markdown parsers. The extension utilizes a Tokenized Markdown Engine: formulas are extracted and swapped with temporary identifiers (___MATH_TOKEN_N___), HTML tags are mapped to CommonMark primitives, and mathematical delimiters are restored at the final stage.

3. Security Architecture & Engineering Quality

The extension implements a hardened security posture:

  • Zero External Data Transmission: Operates entirely locally; user data never leaves the browser context.
  • Proactive XSS Neutralization: Strips executable attributes (including onclick, onerror, and srcdoc) across all processed DOM nodes.
  • Strict Isolation: Hyperlinks are validated against safe protocols (http:, https:, mailto:) and forced to use rel="noopener noreferrer".

Yorumlar (0)

Henüz yorum yapılmamış. İlk yorumu siz yapın!

Yorum Bırakın