A byte pre-filter on JSON lines must not assume how the writer spelled the JSON
PitfallVerified 27 Sep 2026Holds anywhere
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.
Symptom#
A filter that first scans raw JSON lines for a text needle and parses only the matching lines drops lines that do match the query. In dftracer-utils:
- A needle
"field":"value"missed lines written as"field": "value", the default spacing of Pythonjson.dumps. - A needle for a dotted path,
"args.name":"t37", never matched, because the text of the line is"args":{"name":"t37"}. Soname == "thread_name"style tests on dotted fields returned no lines.
Cause#
The needle encoded one spelling of the key and value. Real writers differ in spacing, key order, nesting and escaping, and a dotted query path is not text in the line.
Fix#
- Build needles from values only: a quoted string, or the digits of an integer. Never from a key, a
key:valuepair or a dotted path. - Make needles only from values that no writer would escape. For any other value, do not filter; parse the line.
- Use the needles only as a necessary condition: a line without the needle cannot match, and a line with it is still parsed and checked by the real evaluator.
- Test the pre-filter against the same query with the pre-filter off, on randomized data, and require equal results.
Evidence#
- dftracer-utils stage 6a (
raw-prefilter) replaced the reader's old pre-filter with one needle builder for the View and the reader, and added a fuzz test againstuse_index=False. - On the laghos trace,
name == "thread_name"returned 0 records before and all 16 after. That result needed two fixes: this needle, and a separate reader bug where a nestedargs.nameoverwrote the top-levelnameinside the reader's filter.