diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index fcf1cc501..610a5e5ea 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -43,7 +43,7 @@ primitive" to "add a fusion rule". | D-WFL-EXPR | **Mask algebra is globally NON-MATERIALIZING by default.** A mask EXPRESSION denotes membership; it does not imply a bitmap exists. Materialization happens only at an explicit TERMINAL, when the membership set is requested as a carrier. Stronger than "folds are zero-copy" because folding, masking, ternlog, gating, projection and reduction all join ONE algebra — the expression stays unevaluated as population state all the way to a low-entropy terminal | Queued | three concepts kept distinct in every plan: MASKING (operation) · MASK EXPRESSION (composition) · MATERIALIZED MASK (bitmap). Falsified if a plan cannot express a multi-operand masking chain that emits no membership bits | | D-WFL-MASKOP | ⊘ **`MaskOp` must not semantically mean "produce a Scratch mask" — it must mean CONTRIBUTE TO A MASK EXPRESSION.** Scratch is one physical LOWERING, never the semantics. Read-verified: `Terminal::Keep{mask}` (`ir.rs:184`, *"the final mask itself stays in `mask`… nothing is reduced"*) IS the materialization election, but `MaskOp::And{a,b,dst}` is `dst = a & b` — every op is an assignment, so the ops destroy at level N−1 the choice the terminals encode at level N, and `exec.rs:566` then forces every slot to `words_for(n_rows)` | Queued | **This is the deepest correction in the arc and it precedes W0/W1.** Falsified if changing `MaskOp` semantics does not remove the need for per-op fixes | | D-WFL-SEAMB″ | ⊘ `Pred::Range → Scratch` is a **SYMPTOM, not the disease** — every earlier framing (performance complaint · T1 conformance failure · fold-law violation · absent decision) was chasing one op. Fixing `Range` alone leaves `And`/`Or`/`Xor`/`AndNot`/`Ternlog` all writing full planes | Queued | the fix is judged at the execution MODEL, not at one variant | -| D-WFL-FUSE | **Half the "missing primitives" dissolve into lowering rules.** The fused `popcount(a & b)` of `D-WFL-T1-FUSED` is not a bespoke instruction — it is what a fuser emits for `MaskExpr → Terminal::Count`. `fuse.rs` already collapses a Boolean tree into one ternlog; what it does NOT do is fuse across the **op → terminal** boundary, which is exactly the boundary D-WFL-MASKOP moves | In PR for the single-op case — `Program::fused_ternlog` folds one resident 2/3-input op → Count/Any; multi-op trees (fuse.rs → one ternlog → this fold) still OPEN | re-audit every "missing op" against this before minting any. A primitive that a fusion rule could emit is not a primitive | +| D-WFL-FUSE | **Half the "missing primitives" dissolve into lowering rules.** The fused `popcount(a & b)` of `D-WFL-T1-FUSED` is not a bespoke instruction — it is what a fuser emits for `MaskExpr → Terminal::Count`. `fuse.rs` already collapses a Boolean tree into one ternlog; what it does NOT do is fuse across the **op → terminal** boundary, which is exactly the boundary D-WFL-MASKOP moves | **Shipped** — single-op #1270, multi-op chains ≤3 planes #1272 (>3 planes stays tiled, boundary measured); was: In PR for the single-op case — `Program::fused_ternlog` folds one resident 2/3-input op → Count/Any; multi-op trees (fuse.rs → one ternlog → this fold) still OPEN | re-audit every "missing op" against this before minting any. A primitive that a fusion rule could emit is not a primitive | **Why it took a day, so it is not repeated:** the design already encoded the distinction (`Keep` vs `Count`; `WideFieldMask` as a field-PARTICIPATION @@ -64,8 +64,8 @@ materialization, and that is the ideal Layer-0 operation, not a compromise. |---|---|---|---| | D-WFL-AXIS | **Three independent axes, not one binary:** OPERATORS (fold · mask/ternlog · project · rotate · neighbour · reduce) × CARRIERS (canonical lane · range · descriptor · resident mask · cached mask) × MATERIALIZATION CHOICE (fused vs materialized bitmap). The BBB question is not *fold or mask?* but **is this membership relation transient algebra, or has it been PROMOTED to a mask carrier?** | Queued | every plan must record the promotion, not the operator choice. Falsified if a plan can promote a membership relation to a carrier without that appearing in it | | D-WFL-ENTROPY | **Representation entropy should follow ANSWER entropy.** *"Do these two million-row regions intersect?"* ≈ 1 bit; building 125 KB of mask to find it is the obscenity — and 125 KB is Seam B's MEASURED number at N=1M, not rhetoric. *"How many overlap?"* = 32–64 bits; also fused. *"Give me the overlap, six thoughts will manipulate it"* justifies the bitmap, which is then low-entropy **relative to its future workload** | Queued | the ratio `materialized bytes : answer bytes` reported per operation (§6's `R_info`), with the downstream workload named whenever it exceeds 1 | -| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | In PR — W2b-A shipped (fused `Range ∩ plane → Count/Any`, 0 derived words; `entries/2026-09-23-terminal-elects-materialization-range-plane.md`); W2b-B reuse burden OPEN | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs | -| D-WFL-EXTENT | Absolute execution extent: `execute_extent(program, planes, foreign, scratch, out, lo..hi)`; the extent is an OUTER restriction in `Planes`' absolute row coordinates (never a rebasing); `execute_into` = extent `0..n_rows` | In PR — shipped with split/merge, non-rebasing and structural gates; `entries/2026-09-23-absolute-execution-extent.md` | whole == merge of any partition in any order; extent starting at 65/129 reads absolute rows (worker-local reading shown to differ); a 1-row extent over 1M rows visits one word. OQ-5 untouched | +| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | **Shipped (W2b-A, #1268)**; W2b-B reuse burden still OPEN; was: In PR — W2b-A shipped (fused `Range ∩ plane → Count/Any`, 0 derived words; `entries/2026-09-23-terminal-elects-materialization-range-plane.md`); W2b-B reuse burden OPEN | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs | +| D-WFL-EXTENT | Absolute execution extent: `execute_extent(program, planes, foreign, scratch, out, lo..hi)`; the extent is an OUTER restriction in `Planes`' absolute row coordinates (never a rebasing); `execute_into` = extent `0..n_rows` | **Shipped (#1269)**; was: In PR — shipped with split/merge, non-rebasing and structural gates; `entries/2026-09-23-absolute-execution-extent.md` | whole == merge of any partition in any order; extent starting at 65/129 reads absolute rows (worker-local reading shown to differ); a 1-row extent over 1M rows visits one word. OQ-5 untouched | **Shortest form:** fold the datasets, mask the folds, materialize only when the mask itself is worth keeping. @@ -82,7 +82,7 @@ defines what a fold IS — never what the machine may do. | D-WFL-SIBLING | **The rule governs the TRANSITION, not the bytes:** crossing from fold-native to mask-native execution must be deliberate and visible at the T2 planning membrane. Once MASK is elected, behaving like a mask engine (AND → TERNLOG → shift → cache) is legitimate. Forbidden only: the planner believes it is folding, a helper silently allocates `words_for(N)`, and nobody made the decision | Queued | the BBB question must be answerable for every plan: *who elected the mask, on what basis?* Falsified if a plan can become mask-native without an election appearing in it | | D-WFL-SEAMB′ | ⊘ **restates Seam B more precisely than every earlier framing** (performance complaint · T1 conformance failure · fold-law violation — all circling this). The defect in `Pred::Range` is NOT that it writes a mask. It is that the planner can neither elect nor decline: there is exactly ONE path, so **the choice does not exist**. Seam B is an ABSENT DECISION, not a present mask | Queued | fixed when both paths exist and the plan records which was taken — not when the mask disappears | | D-WFL-W2b″ | ⊘ **supersedes D-WFL-W2b′'s "must not write".** W2b demonstrates BOTH legal paths over identical semantics: FOLD-NATIVE (`Range ∩ resident → Count/Any`, no second mask) and MASK-NATIVE (`→ a bounded/cached mask` because a consumer reuses it). Pipeline vs materialize | Queued | the two arms differentially checked against each other AND the oracle — identical row sets, identical Count/Any. The earlier "zero derived buffers, asserted by counter" gate now scopes to the FOLD arm only | -| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | In PR — plane∩plane closed as the `AND2` table of the generalized `mask_ternlog_popcount`/`_any` (ndarray #322; no new ISA primitive — `U64x8` composition suffices); `entries/2026-09-23-ternlog-count-any-fold.md` | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking | +| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | **Shipped** (ndarray #322 + #1270); was: In PR — plane∩plane closed as the `AND2` table of the generalized `mask_ternlog_popcount`/`_any` (ndarray #322; no new ISA primitive — `U64x8` composition suffices); `entries/2026-09-23-ternlog-count-any-fold.md` | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking | | D-WFL-ELECT | the election rule, static first: `terminal Count → FOLD`; `one AND then Count → probably FOLD`; `reuse_count > 1 → consider MASK`; `shared cached result → MASK`; `Wabe frontier reused → maybe MASK`; `~11 ns cached mask → almost certainly MASK`. DuckDB-style dynamic costing later | Queued | static rules must be inspectable in the plan. Falsified if the rule set fires the same way on every program (it would carry no information — cf. the can-it-stay-silent twin) | ## D-WFL — the cache scoping (2026-09-19): frozen is fine, marching is the disaster @@ -118,7 +118,7 @@ newly written derived buffer. | D-id | scope | status | gate | |---|---|---|---| | D-WFL-W2b′ | **respec: bounded composition must not WRITE the intersection.** `WRONG: Range × resident mask → write a bounded mask → Count/Any`. `RIGHT: peek only the intersecting resident words → AND in registers → Count/Any`. The moment the bounded mask is written the program has crossed into RECONSTRUCT — legitimately perhaps, but it is no longer a fold and must be named | Queued | zero derived buffers allocated or written between the bound and the terminal, asserted by counter. A bounded-mask write fails the wave even at 12 words | -| D-WFL-T1-FUSED | the clean case for the anti-zoo rule licensing a NEW T1 primitive: a fused `popcount(a[i] & b[i])` accumulated over a word span. It cannot be expressed by the existing algebra without an intermediate buffer, so it exposes a genuinely new zero-copy operation rather than a convenience | In PR — composition falsifier HELD at the ISA level (`U64x8` ternlog/popcnt everywhere); the slice loop landed generalized over all 256 tables (ndarray #322); differential gate met | differential vs `mask_and` + `popcount_batch_u64` over the same span; identical answer, zero intermediate bytes. Falsified if composition already achieves it without a buffer | +| D-WFL-T1-FUSED | the clean case for the anti-zoo rule licensing a NEW T1 primitive: a fused `popcount(a[i] & b[i])` accumulated over a word span. It cannot be expressed by the existing algebra without an intermediate buffer, so it exposes a genuinely new zero-copy operation rather than a convenience | **Shipped** (ndarray #322 + #1270); was: In PR — composition falsifier HELD at the ISA level (`U64x8` ternlog/popcnt everywhere); the slice loop landed generalized over all 256 tables (ndarray #322); differential gate met | differential vs `mask_and` + `popcount_batch_u64` over the same span; identical answer, zero intermediate bytes. Falsified if composition already achieves it without a buffer | ## D-WFL-L0 — the foundational ruling, to land BEFORE any W0/W1 code (2026-09-19) @@ -136,7 +136,7 @@ corrects two errors of my own from earlier the same day. | D-id | scope | status | gate | |---|---|---|---| -| D-WFL-L0 | the six clauses above, landed in `.claude/plans/waben-fold-execution-loop-v1.md` §1 | **In PR (#1251)** | none — a ruling, not a measurement. Its falsifier is textual: if a future session can quote `membrane-tiers.md` to argue a full mask is a legal fold output, clause 6 failed to draw the line | +| D-WFL-L0 | the six clauses above, landed in `.claude/plans/waben-fold-execution-loop-v1.md` §1 | **Shipped (#1251)** | none — a ruling, not a measurement. Its falsifier is textual: if a future session can quote `membrane-tiers.md` to argue a full mask is a legal fold output, clause 6 failed to draw the line | **Wave mapping this makes self-evident:** W1 ADDRESS integrity · W2 compact carrier preservation · W3 ROTATE · W4 spatial/local operators · W5 metacognitive @@ -170,7 +170,7 @@ text this arc had already written. Append-only: the earlier sections stand. | D-WFL-W0 | **Two attestations, not one.** ⊘ `(key, ordinal)` is NOT a fix: with `ordinal` = post-sort position, `K→A,K→B` and `K→B,K→A` both digest as `(K,0),(K,1)`. Split instead — `OrderedLaneWitness` keeps `key_digest = H(K0,K1,…)`; a separate `RowDomain` carries `row_order_digest = H(ID0,ID1,…)` over a STABLE row identity (NodeGuid sequence or writer source ordinals), never the semantic key. Executor requires both. Duplicate keys stay legal per `ordered_lane.rs:194` | Queued | a swap of two rows behind one key must fail while key_digest, lens, version and n_rows are all unchanged. Deliverable is a DECISION: the smallest stable row identity that already exists where lane and planes are assembled — minting a new one is the failure mode | | D-WFL-W1-AUTH | **W0 says WHAT is attested; W1 must say WHO MAY MINT IT.** Without it the defect moves up one level — `Program.row_order_digest == Planes.row_order_digest` is still metadata vs metadata. PREFERRED: `AttestedPlanes<'a>` borrowed from the sealed row image (keys + mask planes + value lanes + row-identity sequence), typed constructor, no public assembly — unattested planes UNREPRESENTABLE. Second best: recompute `H(NodeGuid_0…n)` once at the attachment boundary, cached against the immutable borrow. REJECTED: a digest field on a freely-constructed `Planes` | Queued | THE acceptance case: identical keys / version / lens / n_rows / key_digest / `RowDomain` **copied by the caller**, but row identities A↔B swapped and one value plane swapped to match ⇒ execution MUST refuse before the `Range` is consumed. Red before the fix, green only when the actual executed plane ordering is attested | | D-WFL-W1′ | ⊘ **Carrying a `RowDomain` is not verifying one.** Comparing program-domain against `planes.domain` compares two metadata copies; a caller can permute a lane with every label intact, so a metadata check PASSES the permuted-planes falsifier. Three enforcement shapes: derive the digest from lane contents; carry the seal's permutation; or make `Planes` constructible only from a sealed lane (unattested planes unrepresentable) | Queued | the permuted-planes case is load-bearing precisely because a metadata-only check passes it | -| D-WFL-W2a′ | scratch necessity must be DERIVED from the validated program shape (`requires_scratch() == false` for the range-terminal shape), never an ad-hoc caller flag — a flag is a claim, a derived predicate is a proof | Queued | scratch words required = 0, mask words written = 0, answer = `hi-lo`. A zero-mask program still demanding `words_for(N)` fails the wave | +| D-WFL-W2a′ | scratch necessity must be DERIVED from the validated program shape (`requires_scratch() == false` for the range-terminal shape), never an ad-hoc caller flag — a flag is a claim, a derived predicate is a proof | **Shipped (#1268)** — `Program::requires_scratch()` is derived from `fused_terminal`/`fused_ternlog`, never a caller flag | scratch words required = 0, mask words written = 0, answer = `hi-lo`. A zero-mask program still demanding `words_for(N)` fails the wave | | D-WFL-W5′ | ⊘ the `w_slot` = thought-track hypothesis is close to REFUTED: `AttentionMaskSoA::touch` (`attention_mask.rs:83-88`) finds by `mailbox_id` alone and overwrites `w_slot`, so a second track at one node ERASES the first. Do not write an acceptance criterion assuming two tracks visible at one NodeGuid | Queued | either a separate mailbox identity per track with that mapping carried explicitly, or the identification is wrong | | D-WFL-W5″ | the publication decision must test the DELTA's own emptiness, not `Terminal::RangeAny`. RangeAny is endpoint arithmetic for one un-`under`ed `Pred::Range`; `delta = A_{t+1} \ A_t` is an arbitrary, possibly fragmented mask, so a non-empty PREFIX reports "publish" on every converged step | Queued | a fixed-point step (delta empty, prefix non-empty) must publish nothing | | D-WFL-W6.0 | **ATTEND vs EPISTEMIC FIRE.** `AlphaOverlay` is attention memory, not epistemic truth. Ruling: *a non-empty Boolean delta is sufficient for an ATTENTION effect and never sufficient for an EPISTEMIC effect.* Attention may move focus and the alpha trace; epistemic effects require provenance + evidence identity before touching TruthU8/NARS | Queued | the floor that unblocks W6 without a full cognitive theory. Falsified by any path where a bare intersection revises truth | @@ -191,7 +191,7 @@ consumer.** | D-id | scope | status | gate / falsifier | |---|---|---|---| | D-WFL-W1 | row identity + order attestation: `RowDomain {version, lens, permutation-or-row-identity, n_rows}` on `Planes`, checked in `validate`. Binds to the PERMUTATION — NOT key uniqueness (`ordered_lane.rs:194`: equal keys are indistinguishable, so refusing duplicates is a semantic regression that also proves nothing). Also the precondition for replay (`E-REPLAY-CAN-BE-CHEAPER-THAN-STORAGE-1`) | Queued | wrong-version / wrong-lens / permuted-planes (same rows, same length, same key digest, different associated order), each disable-verified RED first. Falsified if a permuted-planes program still returns the oracle's answer | -| D-WFL-W2a | range-native terminals, ZERO scratch: `Count = hi-lo`, `Any = lo!=hi`, gated on the program being exactly one un-`under`ed `Pred::Range`. Requires `execute()`'s unconditional `scratch.words == words_for(n_rows)` check to become conditional on the program needing planes | Queued | N from 1K to 100M, same width, varying absolute position — **mask words written must be 0**. Falsified if any mask word is written, or cost moves with N or position | +| D-WFL-W2a | range-native terminals, ZERO scratch: `Count = hi-lo`, `Any = lo!=hi`, gated on the program being exactly one un-`under`ed `Pred::Range`. Requires `execute()`'s unconditional `scratch.words == words_for(n_rows)` check to become conditional on the program needing planes | **Shipped (#1268)** — `fused_terminal` un-gated `Pred::Range`: Count = hi−lo, Any = lo2 key columns) +**Supersedes:** the W-D line under "What is still open" in `entries/2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md` and `entries/2026-09-23-quack-having-sym-sum-presence-mask.md`. + +## What landed +- **ndarray #324:** `GroupKeyAddr::Pair { hi, lo, stride }`, a third address for the ONE keyed-reduction walker (no new walker). The group is `hi * stride + lo`, computed in u64 (cannot overflow). A minor key `lo >= stride` names no group and is dropped, which is the family's zero-fallback. Five public `masked_group_*_pair` kernels, each reusing its Resident sibling's fold closure unchanged. +- **mask-risc:** `GroupKey::Pair`. Four `GroupReduce` folds: Count, MinI32, MaxI32, SumSymI32. The scalar reference computes the composite independently. +- **quack:** `GroupAddr::Pair` lowers to `GroupKey::Pair`. `lower_group_avg` refuses a Pair key with `LowerError::GroupAvgPairKey` rather than answer with half a fraction. + +No composite key lane is materialised at any layer. + +## Gates +- DuckDB oracle (1.5.5), `GROUP BY (cost_center, status)`, 24 groups encoded as `cost_center*3 + status`: COUNT, MIN, and a sparse MAX in which 15 groups read NULL. The expected values were regenerated by `oracle.py`; all 32 existing rows regenerated byte-identical. +- mask-risc: Pair equals a precomputed composite `Lane` under the same executor, and equals `reference_execute`, for all four folds. Test size is n = 1000, spanning more than one tile. A `lo == stride` row that would alias `(1, 0)`'s slot is dropped. +- **Disable runs** (committed first; each went red, then was restored): + - ndarray: deleting the stride guard, transposing `hi`/`lo`, and u32 wrapping arithmetic. + - mask-risc exec: `hi`/`lo` swapped. + - quack: lowering with `stride + 1`. + - reference: the stride guard dropped. +- Correction made along the way: the ndarray stride test was **vacuous** in its first version. Its dropped rows also fell past the group universe, so the universe check dropped them too. It was rewritten so each dropped row would otherwise alias a real slot. + +## Open +- AVG and full-range (coalescing) SUM over a Pair key. `masked_group_sum_i32_pair` exists in ndarray, but no mask-risc terminal routes a Pair key to it (`GroupSumI32` / `GroupSumViaI32` take a single lane). Refused by name today. +- More than two key columns. Nothing is built for it; a nested Pair or an N-ary address would be its own decision. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 27d6f62d0..d64164dc1 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,16 +25,18 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -154 entries, 2026-08-06 .. 2026-09-23. +156 entries, 2026-08-06 .. 2026-09-23. | date | entry id | finding | file | |---|---|---|---| | 2026-09-23 | `ternlog-count-any-fold` | | [2026-09-23-ternlog-count-any-fold.md](2026-09-23-ternlog-count-any-fold.md) | | 2026-09-23 | `terminal-elects-materialization-range-plane` | | [2026-09-23-terminal-elects-materialization-range-plane.md](2026-09-23-terminal-elects-materialization-range-plane.md) | | 2026-09-23 | `report-plan-zero-copy-pivot-docir-convergence` | ReportPlan lowers into Quack; pivot = view over shared Arc; reports become OGAR ObjectSlot sources via grid_of (OGAR #307) | [2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md](2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md) | +| 2026-09-23 | `quack-w-d-multi-key-group-by` | | [2026-09-23-quack-w-d-multi-key-group-by.md](2026-09-23-quack-w-d-multi-key-group-by.md) | | 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) | | 2026-09-23 | `program-collapse-boolean-chains` | | [2026-09-23-program-collapse-boolean-chains.md](2026-09-23-program-collapse-boolean-chains.md) | | 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) | +| 2026-09-23 | `collapse-probe-v4-and-dispatch-split` | | [2026-09-23-collapse-probe-v4-and-dispatch-split.md](2026-09-23-collapse-probe-v4-and-dispatch-split.md) | | 2026-09-23 | `absolute-execution-extent` | | [2026-09-23-absolute-execution-extent.md](2026-09-23-absolute-execution-extent.md) | | 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) | | 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) | diff --git a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs index 79cf92ef6..cc0af61ad 100644 --- a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs +++ b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs @@ -1,6 +1,6 @@ //! Multi-op Boolean chain → Count/Any at N = 1M rows: collapsed fold vs tiled. //! -//! The same raw op sequence, three physical endings: +//! The same raw op sequence, four physical endings: //! //! - FOLD: `Program::fused_ternlog` interprets the chain symbolically and //! lowers it onto ONE `mask_ternlog_{popcount,any}` pass; no slot is carved @@ -11,22 +11,43 @@ //! terminal reads the last one back. Same answer, today's scratch path. //! - KEEP: the chain with a `Keep` terminal into an `Out::Mask`, then //! `popcount_batch_u64` / `mask_any` over the kept words. +//! - BULK: the same op sequence evaluated ONCE per op over the extent's +//! whole touched-word span, into preallocated whole-population scratch +//! buffers (allocated once per chain, outside every timed closure) — the +//! same word WRITES the tiled arm pays, but with no per-tile interpreter +//! dispatch: one `ndarray::simd` facade call per op, full span wide, +//! instead of one call per op per tile. Two derived columns report the +//! MEASURED GAPS between these paths, not isolated costs: +//! `wr_ns = bulk_ns - fold_ns` is the gap between one fused pass that +//! reads at most three planes and writes nothing, and one pass PER OP that +//! reads and writes whole-span buffers; it mixes write cost with the +//! pass-count and read difference. `disp_ns = tiled_ns - bulk_ns` is the +//! gap between the same ops run per 8-word tile and run once over the +//! whole span; it mixes per-call overhead with batch-size effects (cache +//! residency, loop setup). Attributing either gap to one cause would need +//! matched-work controls this probe does not have. //! -//! Reported per chain × extent: median ns per arm and the derived words the -//! tiled arm writes (`ops × touched words`; the fold writes 0). A four-plane -//! chain is measured on the tiled path only: it is outside the algebra (one -//! ternlog addresses three inputs), and its row records where the collapse -//! stops. Every result is checked against a bit-serial scalar oracle. +//! Reported per chain × extent: median ns per arm (fold/bulk/tiled/keep), +//! the derived words the tiled arm writes (`ops × touched words`; the fold +//! writes 0), and the two BULK-derived columns above. A four-plane chain is +//! measured on the bulk and tiled paths only: it is outside the fold algebra +//! (one ternlog addresses three inputs), and its row records where the +//! collapse stops. Every result is checked against a bit-serial scalar +//! oracle. //! //! `cargo run --release -p lance-graph-mask-risc --example program_collapse_probe` +use std::ops::Range; use std::time::Instant; use lance_graph_mask_risc::exec::{execute_extent, Scratch}; +use lance_graph_mask_risc::ternlog_dispatch::ternlog_dispatch; use lance_graph_mask_risc::{ touched_words, Foreign, MaskOp, Operand, Out, Planes, Program, Terminal, Value, FUSED_SLOT_CAP, }; -use ndarray::simd::{mask_any, popcount_batch_u64}; +use ndarray::simd::{ + mask_and, mask_andnot, mask_any, mask_not, mask_or, mask_xor, popcount_batch_u64, +}; fn lcg(seed: &mut u64) -> u64 { *seed = seed @@ -92,6 +113,141 @@ fn retarget(ops: &[MaskOp], from: u16, to: u16) -> Vec { .collect() } +/// The scratch slot an op writes — every [`MaskOp`] variant has one. +fn op_dst(op: &MaskOp) -> u16 { + match *op { + MaskOp::Pred { dst, .. } + | MaskOp::And { dst, .. } + | MaskOp::Or { dst, .. } + | MaskOp::Xor { dst, .. } + | MaskOp::AndNot { dst, .. } + | MaskOp::Not { dst, .. } + | MaskOp::Ternlog { dst, .. } + | MaskOp::Gather { dst, .. } => dst, + } +} + +/// The bits of word `w` that fall inside `[lo, hi)` (absolute rows) — the +/// BULK arm's own copy of `exec::edge_mask` (private to that module), used +/// to restrict the terminal's first and last touched word to the extent +/// exactly as `exec::run_fused`/`run_fused_ternlog` do. +fn edge_mask(w: usize, lo: usize, hi: usize) -> u64 { + let base = w * 64; + let from = lo.saturating_sub(base).min(64); + let to = (hi - base).min(64); + let upper = if to == 64 { u64::MAX } else { (1u64 << to) - 1 }; + upper & (u64::MAX << from) +} + +/// Every BULK scratch buffer except the one an op is writing — this arm's +/// analogue of `exec::Slots`, over per-slot `Vec` buffers rather than a +/// flat tiled arena. +struct BulkSlots<'s> { + /// Buffers `[0, hole)`. + left: &'s [Vec], + /// Buffers `(hole, bufs.len())`. + right: &'s [Vec], + hole: usize, +} + +impl BulkSlots<'_> { + fn get(&self, i: usize, n: usize) -> &[u64] { + debug_assert!( + i != self.hole, + "bulk arm: slot {i} read while it is this op's write target" + ); + if i < self.hole { + &self.left[i][..n] + } else { + &self.right[i - self.hole - 1][..n] + } + } +} + +/// Evaluate `ops` ONCE per op over `span` (a word range, e.g. from +/// [`touched_words`]) into `bufs`, one preallocated whole-population buffer +/// per scratch slot the chain uses — no per-tile dispatch, one +/// `ndarray::simd` facade call per op over the whole span. Leaf operands +/// (`Operand::Plane`) read `leaves[i][span]` directly, no copy. `bufs` must +/// have one entry per distinct scratch slot `ops` addresses, each at least +/// `span.len()` words long; nothing is allocated here. +/// +/// No edge masking happens here: bits outside the extent `[lo, hi)` but +/// inside `span`'s first/last word are left as whatever the ops computed +/// them to be. That is sound because those positions are masked OUT by +/// [`bulk_terminal`] purely by WORD POSITION, so their VALUE never reaches +/// the reported result — the same reasoning `exec::run_fused_ternlog` relies +/// on for an odd ternlog's tail bits. +fn bulk_eval(ops: &[MaskOp], leaves: &[&[u64]; 4], bufs: &mut [Vec], span: Range) { + let n = span.len(); + for op in ops { + let dst = usize::from(op_dst(op)); + let (left, rest) = bufs.split_at_mut(dst); + let (mid, right) = rest.split_at_mut(1); + let slots = BulkSlots { + left: &*left, + right: &*right, + hole: dst, + }; + let rd = |o: Operand| -> &[u64] { + match o { + Operand::Plane(i) => &leaves[usize::from(i)][span.clone()], + Operand::Scratch(i) => slots.get(usize::from(i), n), + } + }; + let d = &mut mid[0][..n]; + match *op { + MaskOp::And { a, b, .. } => mask_and(rd(a), rd(b), d), + MaskOp::Or { a, b, .. } => mask_or(rd(a), rd(b), d), + MaskOp::Xor { a, b, .. } => mask_xor(rd(a), rd(b), d), + MaskOp::AndNot { a, b, .. } => mask_andnot(rd(a), rd(b), d), + // `n * 64` as the tail-clear bound disables `mask_not`'s own + // clipping (it is a no-op exactly at a word boundary, which `n * + // 64` always is) — any population-edge clipping is `bulk_eval`'s + // caller's job via `bulk_terminal`, not an interior op's. + MaskOp::Not { a, .. } => mask_not(rd(a), n * 64, d), + MaskOp::Ternlog { imm, a, b, c, .. } => ternlog_dispatch(imm, rd(a), rd(b), rd(c), d), + other => panic!("bulk arm: unsupported op {other:?}"), + } + } +} + +/// Fold a BULK terminal slot (`buf`, the `span.len()`-word result of +/// [`bulk_eval`], `span` starting at absolute word `w0`) into `Count` or +/// `Any` over `[lo, hi)`, restricting the first and last word with +/// [`edge_mask`] exactly as `exec::run_fused`/`run_fused_ternlog` do — the +/// interior words are read as they are, since [`touched_words`] guarantees +/// every word strictly between the first and last is wholly inside the +/// extent. +fn bulk_terminal(buf: &[u64], w0: usize, lo: usize, hi: usize, want_count: bool) -> Value { + let n = buf.len(); + if n == 0 { + return if want_count { + Value::Count(0) + } else { + Value::Bool(false) + }; + } + if n == 1 { + let w = buf[0] & edge_mask(w0, lo, hi); + return if want_count { + Value::Count(w.count_ones() as usize) + } else { + Value::Bool(w != 0) + }; + } + let head = [buf[0] & edge_mask(w0, lo, hi)]; + let tail = [buf[n - 1] & edge_mask(w0 + n - 1, lo, hi)]; + let interior = &buf[1..n - 1]; + if want_count { + let c = + popcount_batch_u64(&head) + popcount_batch_u64(interior) + popcount_batch_u64(&tail); + Value::Count(c as usize) + } else { + Value::Bool(mask_any(&head) || mask_any(interior) || mask_any(&tail)) + } +} + fn main() { let n = 1usize << 20; let words = n.div_ceil(64); @@ -207,8 +363,20 @@ fn main() { let extents = [("1%", mid, mid + n / 100), ("whole", 0, n)]; let cap = FUSED_SLOT_CAP as u16; println!( - "{:>17} {:>5} {:>5} {:>3} {:>10} {:>10} {:>10} {:>6} {:>6} {:>9}", - "chain", "ext", "term", "ops", "fold_ns", "tiled_ns", "keep_ns", "t/f", "k/f", "tiled_wr" + "{:>17} {:>5} {:>5} {:>3} {:>10} {:>10} {:>10} {:>10} {:>6} {:>6} {:>9} {:>10} {:>10}", + "chain", + "ext", + "term", + "ops", + "fold_ns", + "bulk_ns", + "tiled_ns", + "keep_ns", + "t/f", + "k/f", + "tiled_wr", + "wr_ns", + "disp_ns" ); for (name, ops, last, oracle) in chains { let term = |mask| (Terminal::Count { mask }, Terminal::Any { mask }); @@ -232,6 +400,12 @@ fn main() { let mut ts = Scratch::for_program(&tiled[0], n).expect("scratch"); let mut ks = Scratch::for_program(&keep, n).expect("scratch"); let mut out = vec![0u64; words]; + // One buffer per distinct scratch slot `ops` addresses, each sized + // to the LARGEST span either extent below touches (`words`, the + // "whole" extent's span) — allocated once per chain, reused across + // both extents and both terminals by slicing to the live span. + let max_slot = ops.iter().map(op_dst).max().unwrap_or(0); + let mut bufs: Vec> = (0..=max_slot).map(|_| vec![0u64; words]).collect(); for (ename, lo, hi) in extents { let want = (lo..hi) .filter(|&r| oracle(bit(&pa, r), bit(&pb, r), bit(&pc, r), bit(&pd, r))) @@ -260,6 +434,16 @@ fn main() { } else { (f64::NAN, expect) }; + let (bns, bv) = median(reps, || { + bulk_eval(&ops, &masks, &mut bufs, span.clone()); + bulk_terminal( + &bufs[usize::from(last)][..span.len()], + span.start, + lo, + hi, + k == 0, + ) + }); let (tns, tv) = median(reps, || { execute_extent( &tiled[k], @@ -289,10 +473,16 @@ fn main() { } }); assert_eq!(fv, expect, "fold {name} {ename} {tname}"); + assert_eq!(bv, expect, "bulk {name} {ename} {tname}"); assert_eq!(tv, expect, "tiled {name} {ename} {tname}"); assert_eq!(kv, expect, "keep {name} {ename} {tname}"); + // NaN propagates automatically: when `fns` is NaN (the 4p + // chain, uncollapsible) `wr_ns` is NaN too, matching the + // fold_ns column's own NaN-means-not-collapsible convention. + let wr_ns = bns - fns; + let disp_ns = tns - bns; println!( - "{name:>17} {ename:>5} {tname:>5} {:>3} {fns:>10.0} {tns:>10.0} {kns:>10.0} {:>6.2} {:>6.2} {:>9}", + "{name:>17} {ename:>5} {tname:>5} {:>3} {fns:>10.0} {bns:>10.0} {tns:>10.0} {kns:>10.0} {:>6.2} {:>6.2} {:>9} {wr_ns:>10.0} {disp_ns:>10.0}", ops.len(), tns / fns, kns / fns, @@ -303,6 +493,8 @@ fn main() { } println!( "(n = {n}, {words} population words; tiled_wr = derived words the tiled path \ - writes, ops × touched words; the fold writes 0; fold_ns NaN = not collapsible)" + writes, ops × touched words; the fold writes 0; fold_ns NaN = not collapsible; \ + wr_ns = bulk_ns - fold_ns and disp_ns = tiled_ns - bulk_ns are measured \ + gaps between paths, not isolated write or dispatch costs)" ); } diff --git a/crates/lance-graph-mask-risc/src/exec.rs b/crates/lance-graph-mask-risc/src/exec.rs index 4d2fda6a4..f608507cc 100644 --- a/crates/lance-graph-mask-risc/src/exec.rs +++ b/crates/lance-graph-mask-risc/src/exec.rs @@ -28,9 +28,11 @@ use ndarray::simd::{ le_i32_to_mask, le_i32_to_mask_under, lt_i32_to_mask, lt_i32_to_mask_under, mask_all, mask_and, mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32, mask_not, mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_xor, - mask_xor_assign, masked_group_count_u32, masked_group_count_u32_via, masked_group_max_i32, - masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_via, masked_group_sum_i32, - masked_group_sum_i32_via, masked_group_sum_sym_i32, masked_group_sum_sym_i32_via, + mask_xor_assign, masked_group_count_u32, masked_group_count_u32_pair, + masked_group_count_u32_via, masked_group_max_i32, masked_group_max_i32_pair, + masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_pair, + masked_group_min_i32_via, masked_group_sum_i32, masked_group_sum_i32_via, + masked_group_sum_sym_i32, masked_group_sum_sym_i32_pair, masked_group_sum_sym_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32, masked_sum_i32, ne_i32_to_mask, ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, popcount_batch_u64, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, ternary_match_u64_to_mask, @@ -1015,7 +1017,8 @@ pub fn execute_extent( }; return Ok(run_fused(f, planes)); } - // The Boolean-membership fold: a single 2/3-input op over resident planes, + // The Boolean-membership fold: a chain of Boolean ops over at most three + // resident planes, collapsed symbolically to one ternlog table (#1272) and // folded by Count/Any — also no slot, no membership bit written. if let Some(f) = program.fused_ternlog() { let mut written = [0u64; FUSED_SLOT_CAP.div_ceil(64)]; @@ -1397,6 +1400,45 @@ pub fn execute_extent( o, ) } + (GroupKey::Pair { hi, lo, stride }, GroupFold::Count) => { + masked_group_count_u32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::MinI32(v)) => { + masked_group_min_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::MaxI32(v)) => { + masked_group_max_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::SumSymI32(v)) => { + masked_group_sum_sym_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } } } } diff --git a/crates/lance-graph-mask-risc/src/ir.rs b/crates/lance-graph-mask-risc/src/ir.rs index a892bafdf..775ae3eb8 100644 --- a/crates/lance-graph-mask-risc/src/ir.rs +++ b/crates/lance-graph-mask-risc/src/ir.rs @@ -384,6 +384,15 @@ pub enum GroupKey { /// `U32` lane of this table, `key` a `U32` lane of the foreign table. /// The two hops are fused; no remapped key lane is materialised. Via { fk: u16, key: u16 }, + /// The group of row `i` is `lanes[hi][i] * stride + lanes[lo][i]` — the + /// composite address of a two-column `GROUP BY`. `hi` and `lo` are both + /// `U32` lanes of this table. Fused: no composite key lane is ever + /// materialised. Zero fallback: a minor key `lanes[lo][i] >= stride` + /// names no group and drops the row, exactly like a resolved key past + /// the group universe — same contract as + /// [`ndarray::simd::masked_group_count_u32_pair`] and its + /// min/max/sum siblings. + Pair { hi: u16, lo: u16, stride: u32 }, } /// What a [`Terminal::GroupReduce`] folds into each group's slot. diff --git a/crates/lance-graph-mask-risc/src/reference.rs b/crates/lance-graph-mask-risc/src/reference.rs index 3738ccc7c..9035f9b40 100644 --- a/crates/lance-graph-mask-risc/src/reference.rs +++ b/crates/lance-graph-mask-risc/src/reference.rs @@ -531,6 +531,10 @@ pub(crate) fn validate( check_lane(planes, fk, LaneKind::U32)?; check_foreign_lane(foreign, key, LaneKind::U32)?; } + GroupKey::Pair { hi, lo, .. } => { + check_lane(planes, hi, LaneKind::U32)?; + check_lane(planes, lo, LaneKind::U32)?; + } } match fold { GroupFold::Count => {} @@ -921,7 +925,7 @@ pub fn reference_execute_into( Some(LaneRef::U32(v)) => &v[..], _ => &[][..], }, - GroupKey::Lane(_) => &[][..], + GroupKey::Lane(_) | GroupKey::Pair { .. } => &[][..], }; for r in survivors(mask) { let k = match key { @@ -933,6 +937,21 @@ pub fn reference_execute_into( } remap[idx] as usize } + GroupKey::Pair { hi, lo, stride } => { + let lo_v = u32_at(planes, lo, r); + if lo_v >= stride { + continue; + } + // hi and lo are both u32, stride is u32: widen to + // u64 first so the multiply-add cannot overflow, + // matching ndarray::simd's GroupKeyAddr::Pair. + let hi_v = u32_at(planes, hi, r); + let composite = u64::from(hi_v) * u64::from(stride) + u64::from(lo_v); + match usize::try_from(composite) { + Ok(k) => k, + Err(_) => continue, + } + } }; if k >= o.len() { continue; diff --git a/crates/lance-graph-mask-risc/tests/foreign.rs b/crates/lance-graph-mask-risc/tests/foreign.rs index 6c6bcb518..776c7c163 100644 --- a/crates/lance-graph-mask-risc/tests/foreign.rs +++ b/crates/lance-graph-mask-risc/tests/foreign.rs @@ -1223,6 +1223,9 @@ fn sym_sum_agrees_with_full_range_sum_plus_count() { key, val: 2, }), + // The loop iterates Lane and Via only: a Pair key has no + // full-range group-sum terminal to compare the sym sum against. + GroupKey::Pair { .. } => unreachable!("loop yields Lane and Via only"), }; let count = run(Terminal::GroupReduce { mask: S0, @@ -1337,3 +1340,226 @@ fn group_reduce_refuses_wrong_lanes_and_a_missing_sink() { }) )); } + +/// FAILS IF: `GroupKey::Pair { hi, lo, stride }` (the fused two-column +/// `GROUP BY` address, `group = hi[i] * stride + lo[i]`) disagrees with an +/// EQUIVALENT `GroupKey::Lane` run over a precomputed composite lane, for +/// any of the four folds, OR either disagrees with the independent +/// row-at-a-time oracle. Multi-tile (`n` past `64 * TILE_WORDS`) so the +/// tiled `_pair` kernels are actually exercised, not just their single-tile +/// remainder. +#[test] +fn pair_key_group_reduce_matches_a_precomputed_composite_lane() { + const STRIDE: u32 = 3; + const HI_RANGE: u32 = 5; + const GROUPS: usize = (HI_RANGE * STRIDE) as usize; // 15 + // 64 * TILE_WORDS (8) = 512; this must exceed it to force multiple tiles. + const N: usize = 1000; + + let mut s = 0xF00D_BEEFu64; + let hi: Vec = (0..N) + .map(|_| (lcg(&mut s) % u64::from(HI_RANGE)) as u32) + .collect(); + let lo: Vec = (0..N) + .map(|_| (lcg(&mut s) % u64::from(STRIDE)) as u32) + .collect(); + // Precomputed composite lane, independently reproducing the pair + // formula so it can stand in for `GroupKey::Lane`. + let comp: Vec = hi.iter().zip(&lo).map(|(&h, &l)| h * STRIDE + l).collect(); + let value: Vec = (0..N).map(|_| (lcg(&mut s) % 4000) as i32 - 2000).collect(); + // A third lane purely to drive the filter predicate below. + let flag: Vec = (0..N).map(|_| (lcg(&mut s) % 3) as u32).collect(); + + let lanes = [ + LaneRef::U32(&hi), // 0 + LaneRef::U32(&lo), // 1 + LaneRef::U32(&comp), // 2 + LaneRef::I32(&value), // 3 + LaneRef::U32(&flag), // 4 + ]; + let masks: [&[u64]; 0] = []; + let planes = Planes { + n_rows: N, + masks: &masks, + lanes: &lanes, + }; + let words = tile_words_for(N); + const { + assert!( + N > 64 * lance_graph_mask_risc::exec::TILE_WORDS, + "must span >1 tile" + ) + }; + + for fold in [ + GroupFold::Count, + GroupFold::MinI32(3), + GroupFold::MaxI32(3), + GroupFold::SumSymI32(3), + ] { + let mk_program = |key: GroupKey| { + Program::new( + vec![MaskOp::Pred { + pred: Pred::NeU32 { lane: 4, v: 1 }, + under: None, + dst: 0, + }], + Terminal::GroupReduce { + mask: S0, + key, + fold, + }, + ) + }; + let run = |key: GroupKey| { + let p = mk_program(key); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + let mut out = vec![7i64; GROUPS]; + let got = execute_into( + &p, + &planes, + &Foreign::NONE, + &mut scratch, + Out::I64(&mut out), + ) + .expect("runs"); + (got, out) + }; + + let (pair_val, pair_out) = run(GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }); + let (lane_val, lane_out) = run(GroupKey::Lane(2)); + assert_eq!(pair_val, Value::GroupReduced, "{fold:?}"); + assert_eq!(pair_val, lane_val, "{fold:?}: value"); + assert_eq!( + pair_out, lane_out, + "{fold:?}: Pair must match the composite Lane exactly" + ); + + // Both must also agree with the independent oracle, for the Pair key. + let p_pair = mk_program(GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }); + let mut want_out = vec![-7i64; GROUPS]; + let want = + reference_execute_into(&p_pair, &planes, &Foreign::NONE, Out::I64(&mut want_out)) + .expect("oracle runs"); + assert_eq!(pair_val, want, "{fold:?}: oracle value"); + assert_eq!(pair_out, want_out, "{fold:?}: oracle sink"); + } + + // Anti-vacuity, computed independently of the executor/oracle under + // test (straight from the raw fixture lanes): at least 8 of the 15 + // composite groups must actually be reached by a selected row, or this + // test would pass just as well against an implementation that filled + // every slot identically. + let non_empty = pair_group_non_empty_count(&hi, &lo, &flag, STRIDE, GROUPS); + assert!( + non_empty >= 8, + "need >= 8 non-empty groups, got {non_empty}" + ); +} + +/// Independent count (NOT the executor or the oracle) of how many of the +/// `groups` composite slots see at least one row that both clears the +/// `flag != 1` filter and has an in-range minor key — the anti-vacuity +/// guard above, computed without depending on the code under test. +fn pair_group_non_empty_count( + hi: &[u32], + lo: &[u32], + flag: &[u32], + stride: u32, + groups: usize, +) -> usize { + let mut seen = vec![false; groups]; + for i in 0..hi.len() { + if flag[i] == 1 { + continue; + } + if lo[i] >= stride { + continue; + } + let g = (hi[i] * stride + lo[i]) as usize; + if g < groups { + seen[g] = true; + } + } + seen.iter().filter(|&&b| b).count() +} + +/// FAILS IF: a row whose minor key equals `stride` (or exceeds it) is +/// silently counted into the group its bits would alias to, instead of +/// being dropped — the `GroupKey::Pair` zero-fallback contract +/// (`lo[i] >= stride` names no group). +#[test] +fn pair_key_drops_a_minor_key_at_stride() { + const STRIDE: u32 = 3; + const GROUPS: usize = 8; + // row: hi, lo, and whether its minor key is in-range + // 0: (0, 0) -> comp 0, valid + // 1: (0, 3) -> lo == stride, DROPPED (would alias comp 3 if not guarded) + // 2: (1, 0) -> comp 3, valid -- the genuine occupant of group 3 + // 3: (1, 1) -> comp 4, valid + // 4: (2, 5) -> lo > stride, DROPPED + let hi: Vec = vec![0, 0, 1, 1, 2]; + let lo: Vec = vec![0, 3, 0, 1, 5]; + let n = hi.len(); + let lanes = [LaneRef::U32(&hi), LaneRef::U32(&lo)]; + let all_bits = [0b1_1111u64]; + let masks: [&[u64]; 1] = [&all_bits]; + let planes = Planes { + n_rows: n, + masks: &masks, + lanes: &lanes, + }; + let p = Program::new( + vec![], + Terminal::GroupReduce { + mask: Operand::Plane(0), + key: GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }, + fold: GroupFold::Count, + }, + ); + let words = tile_words_for(n); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + let mut got_out = vec![9i64; GROUPS]; + let got = execute_into( + &p, + &planes, + &Foreign::NONE, + &mut scratch, + Out::I64(&mut got_out), + ) + .expect("runs"); + let mut want_out = vec![-9i64; GROUPS]; + let want = reference_execute_into(&p, &planes, &Foreign::NONE, Out::I64(&mut want_out)) + .expect("oracle runs"); + assert_eq!(got, Value::GroupReduced); + assert_eq!(got, want); + assert_eq!(got_out, want_out); + // Only the genuine (1, 0) row (row 2) lands in group 3; the aliased, + // out-of-range row 1 must not also be counted there. + assert_eq!( + got_out[3], 1, + "group (1,0)=3 must count only the genuine row, not the row whose lo==stride" + ); + // Exactly 3 of the 5 rows have an in-range minor key (rows 0, 2, 3). + assert_eq!( + got_out.iter().sum::(), + 3, + "rows with lo >= stride must never be counted anywhere" + ); +} diff --git a/crates/lance-graph-quack/src/lib.rs b/crates/lance-graph-quack/src/lib.rs index 9db93881f..68c072492 100644 --- a/crates/lance-graph-quack/src/lib.rs +++ b/crates/lance-graph-quack/src/lib.rs @@ -768,6 +768,12 @@ pub enum LowerError { /// the executor refuses a zero-length sink — so the plan would be /// unexecutable. Refused here, where the caller can see why. EmptyGroupUniverse, + /// `GROUP BY (hi, lo) AVG(val)`: an AVG over a two-column key. There is + /// no full-range pair `GROUP SUM` terminal yet — the count half + /// ([`GroupAgg::Count`] over [`GroupAddr::Pair`]) would lower, the sum + /// half has nowhere to go — so [`lower_group_avg`] refuses rather than + /// answer with only half the fraction. + GroupAvgPairKey, } impl core::fmt::Display for LowerError { @@ -788,6 +794,12 @@ impl core::fmt::Display for LowerError { LowerError::EmptyGroupUniverse => { write!(f, "a grouped plan needs at least one group (groups == 0)") } + LowerError::GroupAvgPairKey => { + write!( + f, + "AVG over a two-column (Pair) group key has no SUM terminal yet" + ) + } } } } @@ -920,6 +932,19 @@ pub enum GroupAddr { /// The group key column on the FOREIGN table. key: ForeignLane, }, + /// `hi[row] * stride + lo[row]` — a two-column `GROUP BY` over a + /// composite address, both `hi` and `lo` `u32` columns of this table. + /// Fused: no composite key column is ever materialised. Same + /// zero-fallback as [`GroupKey::Pair`]: a minor key `lo[row] >= stride` + /// names no group and drops the row. + Pair { + /// The major key column (`u32`). + hi: Col, + /// The minor key column (`u32`), valid only in `0..stride`. + lo: Col, + /// The minor key's cardinality: `group = hi * stride + lo`. + stride: u32, + }, } /// What an [`Agg::GroupReduce`] folds per group. @@ -1225,6 +1250,7 @@ pub fn lower_group_avg( let sum_agg = match key { GroupAddr::Local(k) => Agg::GroupSumI32 { key: k, val }, GroupAddr::Via { fk, key } => Agg::GroupSumViaI32 { fk, key, val }, + GroupAddr::Pair { .. } => return Err(LowerError::GroupAvgPairKey), }; Ok(GroupAvgPlan { sum: lower(&Query { @@ -2018,6 +2044,11 @@ fn terminal_of(agg: Agg, mask: Operand) -> Terminal { fk: fk.0, key: key.0, }, + GroupAddr::Pair { hi, lo, stride } => GroupKey::Pair { + hi: hi.0, + lo: lo.0, + stride, + }, }, fold: agg.fold(), }, diff --git a/crates/lance-graph-quack/tests/duckdb/cases.tsv b/crates/lance-graph-quack/tests/duckdb/cases.tsv index c5b0a30b4..07592745d 100644 --- a/crates/lance-graph-quack/tests/duckdb/cases.tsv +++ b/crates/lance-graph-quack/tests/duckdb/cases.tsv @@ -18,6 +18,9 @@ join_group_sum_country SELECT p.country, SUM(l.amount) FROM line l JOIN partner group_min_cc WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MIN(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=1 GROUP BY g.cost_center ORDER BY g.cost_center 0:-4962;1:-4983;2:-4965;3:-4997;4:-4904;5:-4997;6:-4941;7:-4937 group_max_cc WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=1 GROUP BY g.cost_center ORDER BY g.cost_center 0:19951;1:19872;2:19980;3:19977;4:19967;5:19967;6:19998;7:19848 group_max_cc_sparse WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=2 AND l.qty>45 AND l.cost_center<3 GROUP BY g.cost_center ORDER BY g.cost_center 0:17719;1:14535;2:13253;3:NULL;4:NULL;5:NULL;6:NULL;7:NULL +group2_count_cc_st WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, COUNT(l.rid) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>10 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:91;1:280;2:49;3:83;4:297;5:47;6:87;7:291;8:42;9:79;10:288;11:44;12:68;13:266;14:41;15:73;16:319;17:40;18:82;19:285;20:45;21:105;22:255;23:35 +group2_min_cc_st WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, MIN(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>10 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:-4881;1:-4962;2:-3625;3:-4950;4:-4983;5:-4426;6:-4053;7:-4965;8:-4827;9:-4866;10:-4984;11:-3734;12:-4655;13:-4904;14:-3438;15:-4384;16:-4997;17:-4534;18:-4800;19:-4941;20:-4629;21:-4772;22:-4937;23:-954 +group2_max_cc_st_sparse WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>45 AND l.cost_center<3 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:17771;1:19951;2:17719;3:19586;4:19458;5:14535;6:15928;7:19393;8:13253;9:NULL;10:NULL;11:NULL;12:NULL;13:NULL;14:NULL;15:NULL;16:NULL;17:NULL;18:NULL;19:NULL;20:NULL;21:NULL;22:NULL;23:NULL join_group_count_country WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS country) SELECT g.country, COUNT(l.rid) FROM g LEFT JOIN (line l JOIN partner p ON p.rid=l.partner_id AND l.status=1) ON p.country=g.country GROUP BY g.country ORDER BY g.country 0:562;1:227;2:714;3:173;4:283;5:208;6:347;7:311 join_group_min_country WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS country) SELECT g.country, MIN(l.amount) FROM g LEFT JOIN (line l JOIN partner p ON p.rid=l.partner_id AND l.status=1) ON p.country=g.country GROUP BY g.country ORDER BY g.country 0:-4984;1:-4690;2:-4983;3:-4902;4:-4903;5:-4997;6:-4966;7:-4997 sum_case_posted_cc3 SELECT SUM(CASE WHEN cost_center=3 THEN amount ELSE 0 END) FROM line WHERE status=1 2446433 diff --git a/crates/lance-graph-quack/tests/duckdb/oracle.py b/crates/lance-graph-quack/tests/duckdb/oracle.py index 7db146852..df13ae3b4 100644 --- a/crates/lance-graph-quack/tests/duckdb/oracle.py +++ b/crates/lance-graph-quack/tests/duckdb/oracle.py @@ -47,6 +47,9 @@ "having_sparse_count_cc", "having_sparse_sum_lt_cc", "join_having_count_country", + "group2_count_cc_st", + "group2_min_cc_st", + "group2_max_cc_st_sparse", } BOOLEAN = {"exists_neg"} ROWS = {"rows_proj"} diff --git a/crates/lance-graph-quack/tests/duckdb_differential.rs b/crates/lance-graph-quack/tests/duckdb_differential.rs index 1ea9124bd..11fdabbbf 100644 --- a/crates/lance-graph-quack/tests/duckdb_differential.rs +++ b/crates/lance-graph-quack/tests/duckdb_differential.rs @@ -772,6 +772,97 @@ fn group_max_cc_sparse() { assert_case(&cases, "group_max_cc_sparse", &actual); } +/// `COUNT(*) GROUP BY (cost_center, status)` — the two-column `GROUP BY`, +/// `GroupAddr::Pair { hi: COST_CENTER, lo: STATUS, stride: 3 }` (status is +/// `0..3`, so `stride=3` is the minor key's real cardinality). The SQL +/// encodes the SAME composite key the Rust side folds: +/// `cost_center * 3 + status`. +#[test] +fn group2_count_cc_st() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_count_cc_st", + &planes, + &Foreign::NONE, + Filter::cmp(QTY, Cmp::GtI32(10)), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::Count, + 24, + ); + print_metric("group2_count_cc_st", &m); + assert_case(&cases, "group2_count_cc_st", &actual); +} + +/// `MIN(amount) GROUP BY (cost_center, status)` — same `Pair` key as +/// [`group2_count_cc_st`], a different fold. +#[test] +fn group2_min_cc_st() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_min_cc_st", + &planes, + &Foreign::NONE, + Filter::cmp(QTY, Cmp::GtI32(10)), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::MinI32(AMOUNT), + 24, + ); + print_metric("group2_min_cc_st", &m); + assert_case(&cases, "group2_min_cc_st", &actual); +} + +/// `MAX(amount) GROUP BY (cost_center, status)` over a filter tight enough +/// (`qty>45 AND cost_center<3`) that whole `cost_center` bands are empty — +/// the two-column sibling of [`group_max_cc_sparse`], exercising the SQL +/// `NULL` encoding of an empty group under the `Pair` key. +#[test] +fn group2_max_cc_st_sparse() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_max_cc_st_sparse", + &planes, + &Foreign::NONE, + Filter::and([ + Filter::cmp(QTY, Cmp::GtI32(45)), + Filter::or([ + Filter::cmp(COST_CENTER, Cmp::EqU32(0)), + Filter::cmp(COST_CENTER, Cmp::EqU32(1)), + Filter::cmp(COST_CENTER, Cmp::EqU32(2)), + ]), + ]), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::MaxI32(AMOUNT), + 24, + ); + assert!( + actual.contains(":NULL"), + "fixture must leave some group empty or this case tests nothing: {actual}" + ); + print_metric("group2_max_cc_st_sparse", &m); + assert_case(&cases, "group2_max_cc_st_sparse", &actual); +} + #[test] fn group_sum_cc() { let cases = load_cases();