Build a ground-truth set for PDF extraction from arXiv LaTeX sources
RecipeVerified 28 Sep 2026Holds anywhere
Recipe. When to use it, the steps, and what showed that they work.
When to use#
You tune a heuristic or a model that reads structure from PDFs (tables, figures, headings, references) and need a held-out score, not one hand-picked file.
Steps#
- Pick about 12 recent papers across layouts: IEEE two-column, ACM two-column, single-column ML, revtex, AASTeX, a statistics paper, and one paper with rule-less tables. Skip papers whose source is not LaTeX.
- Download each PDF and source, slowly (3 s between requests, with a User-Agent that names the tool):
https://arxiv.org/pdf/<id>andhttps://arxiv.org/e-print/<id>. The e-print is a.tar.gz, a gzipped single.texor plain text; handle all three. - From the main
.tex(follow\inputand\include, strip%comments), write the truth per paper:- tables: each
table,table*,longtableordeluxetablefloat that holds a tabular-like environment, with its caption, header row cells and column count (the most common body row width, with\multicolumn{n}counted as n, not the header cell count); - figures: count of
figureandfigure*environments; - references: the
.bblfile's\bibitemcount, and titles from\bibinfo{title}or the second\newblock; - headings: the PDF's own bookmarks, when it has them, to score a heading guesser run without them.
- tables: each
- Match found items to the truth loosely: a table matches when at least half of its truth header tokens appear in its first two rows; report recall, precision and exact column counts.
- Keep the scripts next to the data in a scratch folder, and keep the paper IDs in the project notes: macOS clears
/private/tmpon reboot, and then only the IDs and scripts bring the set back.
Evidence#
rait, 2026-09-24 to 2026-09-28: this set scored table detectors (heuristic 46% recall / 38% precision, Table Transformer 75% / 72%), granite-docling (67% / 65%), heading guessing (67% of bookmark titles) and reference parsing (784 entries for 787 \bibitems, 94% of titles). Counting columns from the header cells marked tables with a \multicolumn header as wrong; counting body row widths instead raised the exact-column score from 25 to 48 of 67 for the same output. After a reboot cleared the scratch folder, the set was rebuilt from the saved IDs and gave the same scores.