Ray's Knowledge Base

rait's held-out PDF evaluation set: 12 arXiv papers and the scores to beat

FactVerified 28 Sep 2026Holds project: rait
Fact. A statement and the evidence for it.

Statement#

rait's PDF features are scored on 12 arXiv papers, built as in the arXiv ground-truth recipe. The set is not in the repository; rebuild it from these IDs:

2003.13350 ml        2006.13453 revtex   2008.11881 ieee    2012.00825 ieee
2108.05316 ieee      2109.03670 ml       2109.06716 norules 2203.03754 revtex
2209.02334 acm       2209.04554 acm      2209.05563 econ    2210.08587 aastex

Scores at commit 81ff1a0 (2026-09-27), to compare against after a change:

  • pdf tables with both models: 87 tables, 93 found, 67 matched: 77% recall, 72% precision, exact column count 51 of 67. Without the structure model: 75% / 72%, 42 of 65. Heuristic only: 39% / 41%.
  • Figure PDFs from the user's own paper (45 files): the models find no tables; the heuristic finds 1, which is a table image.
  • pdf figures: 83 figures on the set, where the LaTeX sources have about 85 figure environments plus one paper's figure macros; all 8 figures of the user's paper are right.
  • pdf outline --headings against each paper's own bookmarks (9 papers): 67% of bookmark titles found.
  • pdf references against 11 .bbl files: 784 entries for 787 \bibitems (9 papers exact), 94% of checkable titles; the user's paper gives exactly its 62 entries.

A table matches when at least half of its truth header tokens appear in its header or first row. Exact column counts use the most common body row width with \multicolumn{n} counted as n.

Evidence#

Measured in the session of 2026-09-24 to 2026-09-28 with the release binary and scripts in a scratch folder (eval.py, outline_eval.py, refs_eval.py, figs.py). After a reboot cleared the scratch folder, the set was rebuilt from these IDs and the table score was unchanged (75% / 72%).