Ray's Knowledge Base

Typed record decode got 35% slower until per-leaf work left the hot loop

PitfallVerified 27 Sep 2026Holds project: dftracer-utils
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.

Symptom#

Reading a file with a user record schema (typed fields with roles) cost about 35% more CPU than reading the same file as generic. Smaller slowdowns showed in the same change: about 5% CPU on generic files, the evaluator, and group_by.

Cause#

Work that depends only on the plan or the schema ran per record or per JSON leaf:

  • a lookup of the declared field for every leaf path;
  • decoding role paths (time, duration, entity) for plans that never use them;
  • passing a record schema with no fields, so the decoder took the typed path for nothing;
  • in the evaluator, a general comparison before the common string case; in group_by, a combined window check and line probes inside the loop.

Fix#

  • Align path_fields with the projection, so the decoder gets the declared field of each projected path by index, with no lookup per leaf.
  • Project role paths only for plan.timed plans.
  • Pass a null schema when the schema has no fields.
  • Test the string fast path first in the evaluator.
  • Split in_window into in_role_window and hoist probe_lines out of the loop in group_by.

After each change, rerun the A/B benchmark (index_bench, and a CPU comparison of the same file read as the user schema and as generic).

Evidence#

  • Stage 7b (record-schema, commit eb2bb89c): the user schema went from +35% to +3.4% CPU against generic, with the same wall time; generic JSON lines were 543 ms against 544 ms CPU; dftracer traces were within 2%.
  • The remaining 3.4% is the cost of converting values to their declared types.