Ray's Knowledge Base

rait finds PDF tables with Table Transformer in Rust, and word positions fill the cells

DecisionVerified 28 Sep 2026Holds project: rait
Decision. The context, the choice, why, and the options that were turned down.

Context#

rait pdf tables and rait pdf text must find tables in arbitrary PDFs, not only in one paper. A word-gap heuristic alone overfit: on a held-out set of 12 arXiv papers (87 tables) it scored 46% recall and 38% precision, and it called 29 of 45 figure PDFs tables. The binary must stay self-contained and pure Rust where possible.

Decision#

  • Table Transformer detection (microsoft/table-transformer-detection, MIT, revision 2357cbe) finds table regions on each page rendered to 800 px.
  • Table Transformer structure v1.1 (microsoft/table-transformer-structure-recognition-v1.1-all, revision 7587a7e) finds rows, columns, the header and spanning cells in each crop (800 px, at most 200 dpi).
  • The PDF's own words (pdftotext -bbox-layout) fill the grid; sideways tables and regions without a grid fall back to word gaps inside the region.
  • Both models run in crates/rait-vision: candle 0.9.2 for safetensors loading, with own kernels (batch norm folded into convolutions, tiled im2col with sgemm from Accelerate on macOS and the gemm crate elsewhere, fused softmax and layernorm).
  • rait models fetch downloads the weights with curl into the model cache and checks their SHA-256; they are never embedded. Without them, the word-gap heuristic runs, and every result names its detector.
  • A table that continues on the next page under the same header row, with no Table N caption above it, is joined and gets end_page.

Why#

On the held-out set, detection plus word geometry scored 75% recall and 72% precision; with the structure model 77% and 72%, and exact column counts rose from 42 of 65 to 51 of 67. The models found no tables in the 45 figure PDFs. The Rust port matches ONNX Runtime and transformers within 2.5e-5 (logits) and 5e-6 (structure) and runs a 14-page paper in about 0.7 s.

Rejected options#

  • granite-docling-258M: 67% recall and 65% precision on the same set, 7.4 s per page even on the GPU (MLX), and it turned sideways tables into pictures. Kept as a possible opt-in mode for scanned PDFs.
  • tract and ONNX Runtime (ort): the user asked for a pure-Rust path without them; tract also needed with_ignore_value_info to load the int8 model.
  • Laya as a yes/no table verifier on the text of each region: zero-shot it kept 40% of real regions and 34% of false ones; text alone does not separate a table from math.
  • Splitting a model column at a vertical gap between words: the gaps in wrongly merged columns were about one space wide, and it split real columns.