Typed record decode got 35% slower until per-leaf work left the hot loop
PitfallVerified 27 Sep 2026Holds project: dftracer-utils
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.
Symptom#
Reading a file with a user record schema (typed fields with roles) cost about 35% more CPU than reading the same file as generic. Smaller slowdowns showed in the same change: about 5% CPU on generic files, the evaluator, and group_by.
Cause#
Work that depends only on the plan or the schema ran per record or per JSON leaf:
- a lookup of the declared field for every leaf path;
- decoding role paths (time, duration, entity) for plans that never use them;
- passing a record schema with no fields, so the decoder took the typed path for nothing;
- in the evaluator, a general comparison before the common string case; in
group_by, a combined window check and line probes inside the loop.
Fix#
- Align
path_fieldswith the projection, so the decoder gets the declared field of each projected path by index, with no lookup per leaf. - Project role paths only for
plan.timedplans. - Pass a null schema when the schema has no fields.
- Test the string fast path first in the evaluator.
- Split
in_windowintoin_role_windowand hoistprobe_linesout of the loop ingroup_by.
After each change, rerun the A/B benchmark (index_bench, and a CPU comparison of the same file read as the user schema and as generic).
Evidence#
- Stage 7b (
record-schema, commiteb2bb89c): the user schema went from +35% to +3.4% CPU againstgeneric, with the same wall time; generic JSON lines were 543 ms against 544 ms CPU; dftracer traces were within 2%. - The remaining 3.4% is the cost of converting values to their declared types.