The aggregation tier holds no file id, so one file's rows cannot be replaced
FactVerified 27 Sep 2026Holds project: dftracer-utils
Fact. A statement and the evidence for it.
Statement#
The pre-aggregation tier (the aggregation and system_metrics column families, the dftracer.agg extension from stage 10a) stores rows as RocksDB merge operands. The row key holds a shard, the map type and the interned cat, name, pid, tid, hhash, fhash, the time bucket and extra keys. It holds no file id. The merge operator adds every file's operands into the same rows.
Consequences:
- You cannot remove or replace the share of one file. Any rebuild of an aggregated file, for a changed source, a new checkpoint size, a new schema or a forced rebuild, must clear the whole tier first and aggregate every file again. Otherwise the file's old rows stay, and the next aggregation adds its new rows on top, so its events are counted twice.
- Before stage 10a, the build avoided this by deleting the whole index root on any changed source or interval change. Stage 10a clears only the tier (
agg::tier::clear) and aggregates every dftracer file again in an aggregation-only pass. - A reader may use the tier only when every file of the view has a current
dftracer.aggmanifest entry. Checking only the file registry is not enough: a file indexed without aggregation is in the registry but not in the tier.
Evidence#
serialize_agg_key_intoinsrc/dftracer/utils/index/schemas/dft/agg/aggregation_serialization.cppwrites the key fields listed above and no file id.- Before stage 10a,
resolve_and_build.cppranfs::remove_all(root)whenneeds_augmentation || stale_detected, with the comment that a merged aggregation polluted by a changed source cannot be refined in place. - Stage 10a test
index/test_agg_extension: after a source change the tier is cleared and the count over two files equals a fresh index (75 events); a view over a file without an entry reads the traces (50 events, where the tier holds only 30).