duql replaces dictionaries, resolved columns and the metadata phase with queries
Context#
Stage 12 of the extensible-index plan needed dictionaries (resolved.run.*) and metadata for path-decoded formats such as the new genesis long format. Two designs were tried and rejected on 2026-09-27: a schema-wide kind role (one field that names each record's kind), then a per-dictionary {path, value} row selector.
Decision#
- A query language, duql, carries these features. The design is
docs/plans/2026-09-27-duql.md(gitignored, like all ofdocs/plans/). - A source (profile) is a set of named row sets:
data(the default),all(every record) and one per lookup (runs,files). Lookups are queries:let,in (from ...), the arrowrun -> runs.app,lookup ... into. - No feature adds a schema-wide attribute; each feature is expressed as a query.
- Regex: Highway literal tests and a literal prefilter plus PCRE2 now; Vectorscan later for yes/no matching and the multi-pattern prefilter.
- Stage order: 12a1 grammar, 12a2 semantics, 12b pipeline, 12c window/expand/pivot, 12d lookups, 12e sources, 12f genesis long format, 12g builder, 12h plugins, 12i Vectorscan.
Grammar decisions from stage 12a1 (duql-grammar, 2026-09-27):
- Two trees: the duql syntax tree (
src/dftracer/utils/duql/syntax.h) and today'squery::QueryNode, joined by a lowering step. Constructs the engine cannot evaluate lower to an error naming the stage that adds them. - Stage words (
where,sample, ...) are keywords only at a stage start; a query that starts with one is a stage, so such a field is written in backticks there. Expression keywordsand or not in like ilike between is escape null true falseare reserved. -is always an operator; insort, a leading-is the direction and takes a full expression.pivottakes its key below theinlevel.sourcemembers are separated by;. The legacy"text" in fieldis chosen by lookahead (a string,[not] in, then a path orany(path)).- The Lark copy
scripts/duql.larkand the C++ parser read the same samples,tests/duql/grammar_samples.duql.
Why#
The user: "if everytime we have new feature and we add kind attribute then it is not schemaless per se"; the goal is jq-like querying of any JSON shape, with stability, performance, correctness and future-proofing ahead of a quick feature. Research on KQL, PRQL, GROQ, Malloy, PartiQL, LINQ and others showed that a lookup can be a plain named sub-query whose key set the engine pushes down.
Rejected options#
kindrole on the schema: a new schema attribute per feature.- Per-dictionary
{path, value}selector: the user must know the selecting value; cannot select by shape. - SQL as the main language: rigid; may come later as a second front end over the same tree.
Evidence#
- The genesis long-format writer is kept uncommitted in the working tree; the engine edits for the
kindrole were reverted.