Ray's Knowledge Base

Code that assumes the dftracer schema breaks other record schemas silently

PitfallVerified 27 Sep 2026Holds project: dftracer-utils
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.

Symptom#

With the genesis schema (records carry only run, and a run dictionary is built from RUN lines), select/query on resolved.run.* in the docs example returned empty columns, with no error.

Cause#

The View's ride-along index builds passed the built-in dftracer dictionaries (FH, HH, SH, PR) instead of the dictionaries of the file's record schema, so the run dictionary was never built. The dftracer lookup tables are hard-coded in several places in the reader and the View (about 5 before stage 6b; resolved.fpath and resolved.hostname were the only resolved columns).

Fix#

  • Take dictionaries, roles and fields from the plan's record schema (plan_record_schema(plan).dictionaries), never from the dftracer built-ins, in every build and read path: the View scan, the executor (2 places), the indexed export, the column-type harvest and the server (TraceIndex::record_schema()).
  • When you add a record schema, test a full user workflow on it (index, then select and query the new resolved columns), not only unit tests of the schema.

Evidence#

  • Stage 8 (genesis-schema, commit 9bde805f): the docs example showed empty resolved.run.* until view_scan.cpp, view_executor.cpp and view_export_index.cpp used the plan schema's dictionaries. The docs guide was also changed to index first, because TraceViewer builds only the members tier.