Ray's Knowledge Base

Build a ground-truth set for PDF extraction from arXiv LaTeX sources

RecipeVerified 28 Sep 2026Holds anywhere
Recipe. When to use it, the steps, and what showed that they work.

When to use#

You tune a heuristic or a model that reads structure from PDFs (tables, figures, headings, references) and need a held-out score, not one hand-picked file.

Steps#

  1. Pick about 12 recent papers across layouts: IEEE two-column, ACM two-column, single-column ML, revtex, AASTeX, a statistics paper, and one paper with rule-less tables. Skip papers whose source is not LaTeX.
  2. Download each PDF and source, slowly (3 s between requests, with a User-Agent that names the tool): https://arxiv.org/pdf/<id> and https://arxiv.org/e-print/<id>. The e-print is a .tar.gz, a gzipped single .tex or plain text; handle all three.
  3. From the main .tex (follow \input and \include, strip % comments), write the truth per paper:
    • tables: each table, table*, longtable or deluxetable float that holds a tabular-like environment, with its caption, header row cells and column count (the most common body row width, with \multicolumn{n} counted as n, not the header cell count);
    • figures: count of figure and figure* environments;
    • references: the .bbl file's \bibitem count, and titles from \bibinfo{title} or the second \newblock;
    • headings: the PDF's own bookmarks, when it has them, to score a heading guesser run without them.
  4. Match found items to the truth loosely: a table matches when at least half of its truth header tokens appear in its first two rows; report recall, precision and exact column counts.
  5. Keep the scripts next to the data in a scratch folder, and keep the paper IDs in the project notes: macOS clears /private/tmp on reboot, and then only the IDs and scripts bring the set back.

Evidence#

rait, 2026-09-24 to 2026-09-28: this set scored table detectors (heuristic 46% recall / 38% precision, Table Transformer 75% / 72%), granite-docling (67% / 65%), heading guessing (67% of bookmark titles) and reference parsing (784 entries for 787 \bibitems, 94% of titles). Counting columns from the header cells marked tables with a \multicolumn header as wrong; counting body row widths instead raised the exact-column score from 25 to 48 of 67 for the same output. After a reboot cleared the scratch folder, the set was rebuilt from the saved IDs and gave the same scores.