From 0293250374e9f5b9025f9da7008e3aa64d937df0 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 19:44:01 +0000 Subject: [PATCH 1/5] board: reconcile D-WFL status drift after #1251/#1268/#1269/#1270/#1272 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nine STATUS_BOARD rows still read "In PR" or "Queued" for work that has merged. Each flip was checked against origin/main before editing: - D-WFL-FUSE Shipped: single-op #1270, multi-op chains ≤3 planes #1272 - D-WFL-W2b‴ Shipped W2b-A #1268; W2b-B reuse burden stays OPEN - D-WFL-EXTENT Shipped #1269 (`execute_extent`) - D-WFL-T1-FUSED, T1-FUSED′ Shipped (ndarray #322 + #1270) - D-WFL-L0 Shipped #1251 - D-WFL-W2a Shipped #1268: un-gated Pred::Range, Count = hi−lo, Any = lo Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- .claude/board/STATUS_BOARD.md | 18 +++++++++--------- crates/lance-graph-mask-risc/src/exec.rs | 3 ++- 2 files changed, 11 insertions(+), 10 deletions(-) diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index fcf1cc501..610a5e5ea 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -43,7 +43,7 @@ primitive" to "add a fusion rule". | D-WFL-EXPR | **Mask algebra is globally NON-MATERIALIZING by default.** A mask EXPRESSION denotes membership; it does not imply a bitmap exists. Materialization happens only at an explicit TERMINAL, when the membership set is requested as a carrier. Stronger than "folds are zero-copy" because folding, masking, ternlog, gating, projection and reduction all join ONE algebra — the expression stays unevaluated as population state all the way to a low-entropy terminal | Queued | three concepts kept distinct in every plan: MASKING (operation) · MASK EXPRESSION (composition) · MATERIALIZED MASK (bitmap). Falsified if a plan cannot express a multi-operand masking chain that emits no membership bits | | D-WFL-MASKOP | ⊘ **`MaskOp` must not semantically mean "produce a Scratch mask" — it must mean CONTRIBUTE TO A MASK EXPRESSION.** Scratch is one physical LOWERING, never the semantics. Read-verified: `Terminal::Keep{mask}` (`ir.rs:184`, *"the final mask itself stays in `mask`… nothing is reduced"*) IS the materialization election, but `MaskOp::And{a,b,dst}` is `dst = a & b` — every op is an assignment, so the ops destroy at level N−1 the choice the terminals encode at level N, and `exec.rs:566` then forces every slot to `words_for(n_rows)` | Queued | **This is the deepest correction in the arc and it precedes W0/W1.** Falsified if changing `MaskOp` semantics does not remove the need for per-op fixes | | D-WFL-SEAMB″ | ⊘ `Pred::Range → Scratch` is a **SYMPTOM, not the disease** — every earlier framing (performance complaint · T1 conformance failure · fold-law violation · absent decision) was chasing one op. Fixing `Range` alone leaves `And`/`Or`/`Xor`/`AndNot`/`Ternlog` all writing full planes | Queued | the fix is judged at the execution MODEL, not at one variant | -| D-WFL-FUSE | **Half the "missing primitives" dissolve into lowering rules.** The fused `popcount(a & b)` of `D-WFL-T1-FUSED` is not a bespoke instruction — it is what a fuser emits for `MaskExpr → Terminal::Count`. `fuse.rs` already collapses a Boolean tree into one ternlog; what it does NOT do is fuse across the **op → terminal** boundary, which is exactly the boundary D-WFL-MASKOP moves | In PR for the single-op case — `Program::fused_ternlog` folds one resident 2/3-input op → Count/Any; multi-op trees (fuse.rs → one ternlog → this fold) still OPEN | re-audit every "missing op" against this before minting any. A primitive that a fusion rule could emit is not a primitive | +| D-WFL-FUSE | **Half the "missing primitives" dissolve into lowering rules.** The fused `popcount(a & b)` of `D-WFL-T1-FUSED` is not a bespoke instruction — it is what a fuser emits for `MaskExpr → Terminal::Count`. `fuse.rs` already collapses a Boolean tree into one ternlog; what it does NOT do is fuse across the **op → terminal** boundary, which is exactly the boundary D-WFL-MASKOP moves | **Shipped** — single-op #1270, multi-op chains ≤3 planes #1272 (>3 planes stays tiled, boundary measured); was: In PR for the single-op case — `Program::fused_ternlog` folds one resident 2/3-input op → Count/Any; multi-op trees (fuse.rs → one ternlog → this fold) still OPEN | re-audit every "missing op" against this before minting any. A primitive that a fusion rule could emit is not a primitive | **Why it took a day, so it is not repeated:** the design already encoded the distinction (`Keep` vs `Count`; `WideFieldMask` as a field-PARTICIPATION @@ -64,8 +64,8 @@ materialization, and that is the ideal Layer-0 operation, not a compromise. |---|---|---|---| | D-WFL-AXIS | **Three independent axes, not one binary:** OPERATORS (fold · mask/ternlog · project · rotate · neighbour · reduce) × CARRIERS (canonical lane · range · descriptor · resident mask · cached mask) × MATERIALIZATION CHOICE (fused vs materialized bitmap). The BBB question is not *fold or mask?* but **is this membership relation transient algebra, or has it been PROMOTED to a mask carrier?** | Queued | every plan must record the promotion, not the operator choice. Falsified if a plan can promote a membership relation to a carrier without that appearing in it | | D-WFL-ENTROPY | **Representation entropy should follow ANSWER entropy.** *"Do these two million-row regions intersect?"* ≈ 1 bit; building 125 KB of mask to find it is the obscenity — and 125 KB is Seam B's MEASURED number at N=1M, not rhetoric. *"How many overlap?"* = 32–64 bits; also fused. *"Give me the overlap, six thoughts will manipulate it"* justifies the bitmap, which is then low-entropy **relative to its future workload** | Queued | the ratio `materialized bytes : answer bytes` reported per operation (§6's `R_info`), with the downstream workload named whenever it exceeds 1 | -| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | In PR — W2b-A shipped (fused `Range ∩ plane → Count/Any`, 0 derived words; `entries/2026-09-23-terminal-elects-materialization-range-plane.md`); W2b-B reuse burden OPEN | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs | -| D-WFL-EXTENT | Absolute execution extent: `execute_extent(program, planes, foreign, scratch, out, lo..hi)`; the extent is an OUTER restriction in `Planes`' absolute row coordinates (never a rebasing); `execute_into` = extent `0..n_rows` | In PR — shipped with split/merge, non-rebasing and structural gates; `entries/2026-09-23-absolute-execution-extent.md` | whole == merge of any partition in any order; extent starting at 65/129 reads absolute rows (worker-local reading shown to differ); a 1-row extent over 1M rows visits one word. OQ-5 untouched | +| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | **Shipped (W2b-A, #1268)**; W2b-B reuse burden still OPEN; was: In PR — W2b-A shipped (fused `Range ∩ plane → Count/Any`, 0 derived words; `entries/2026-09-23-terminal-elects-materialization-range-plane.md`); W2b-B reuse burden OPEN | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs | +| D-WFL-EXTENT | Absolute execution extent: `execute_extent(program, planes, foreign, scratch, out, lo..hi)`; the extent is an OUTER restriction in `Planes`' absolute row coordinates (never a rebasing); `execute_into` = extent `0..n_rows` | **Shipped (#1269)**; was: In PR — shipped with split/merge, non-rebasing and structural gates; `entries/2026-09-23-absolute-execution-extent.md` | whole == merge of any partition in any order; extent starting at 65/129 reads absolute rows (worker-local reading shown to differ); a 1-row extent over 1M rows visits one word. OQ-5 untouched | **Shortest form:** fold the datasets, mask the folds, materialize only when the mask itself is worth keeping. @@ -82,7 +82,7 @@ defines what a fold IS — never what the machine may do. | D-WFL-SIBLING | **The rule governs the TRANSITION, not the bytes:** crossing from fold-native to mask-native execution must be deliberate and visible at the T2 planning membrane. Once MASK is elected, behaving like a mask engine (AND → TERNLOG → shift → cache) is legitimate. Forbidden only: the planner believes it is folding, a helper silently allocates `words_for(N)`, and nobody made the decision | Queued | the BBB question must be answerable for every plan: *who elected the mask, on what basis?* Falsified if a plan can become mask-native without an election appearing in it | | D-WFL-SEAMB′ | ⊘ **restates Seam B more precisely than every earlier framing** (performance complaint · T1 conformance failure · fold-law violation — all circling this). The defect in `Pred::Range` is NOT that it writes a mask. It is that the planner can neither elect nor decline: there is exactly ONE path, so **the choice does not exist**. Seam B is an ABSENT DECISION, not a present mask | Queued | fixed when both paths exist and the plan records which was taken — not when the mask disappears | | D-WFL-W2b″ | ⊘ **supersedes D-WFL-W2b′'s "must not write".** W2b demonstrates BOTH legal paths over identical semantics: FOLD-NATIVE (`Range ∩ resident → Count/Any`, no second mask) and MASK-NATIVE (`→ a bounded/cached mask` because a consumer reuses it). Pipeline vs materialize | Queued | the two arms differentially checked against each other AND the oracle — identical row sets, identical Count/Any. The earlier "zero derived buffers, asserted by counter" gate now scopes to the FOLD arm only | -| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | In PR — plane∩plane closed as the `AND2` table of the generalized `mask_ternlog_popcount`/`_any` (ndarray #322; no new ISA primitive — `U64x8` composition suffices); `entries/2026-09-23-ternlog-count-any-fold.md` | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking | +| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | **Shipped** (ndarray #322 + #1270); was: In PR — plane∩plane closed as the `AND2` table of the generalized `mask_ternlog_popcount`/`_any` (ndarray #322; no new ISA primitive — `U64x8` composition suffices); `entries/2026-09-23-ternlog-count-any-fold.md` | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking | | D-WFL-ELECT | the election rule, static first: `terminal Count → FOLD`; `one AND then Count → probably FOLD`; `reuse_count > 1 → consider MASK`; `shared cached result → MASK`; `Wabe frontier reused → maybe MASK`; `~11 ns cached mask → almost certainly MASK`. DuckDB-style dynamic costing later | Queued | static rules must be inspectable in the plan. Falsified if the rule set fires the same way on every program (it would carry no information — cf. the can-it-stay-silent twin) | ## D-WFL — the cache scoping (2026-09-19): frozen is fine, marching is the disaster @@ -118,7 +118,7 @@ newly written derived buffer. | D-id | scope | status | gate | |---|---|---|---| | D-WFL-W2b′ | **respec: bounded composition must not WRITE the intersection.** `WRONG: Range × resident mask → write a bounded mask → Count/Any`. `RIGHT: peek only the intersecting resident words → AND in registers → Count/Any`. The moment the bounded mask is written the program has crossed into RECONSTRUCT — legitimately perhaps, but it is no longer a fold and must be named | Queued | zero derived buffers allocated or written between the bound and the terminal, asserted by counter. A bounded-mask write fails the wave even at 12 words | -| D-WFL-T1-FUSED | the clean case for the anti-zoo rule licensing a NEW T1 primitive: a fused `popcount(a[i] & b[i])` accumulated over a word span. It cannot be expressed by the existing algebra without an intermediate buffer, so it exposes a genuinely new zero-copy operation rather than a convenience | In PR — composition falsifier HELD at the ISA level (`U64x8` ternlog/popcnt everywhere); the slice loop landed generalized over all 256 tables (ndarray #322); differential gate met | differential vs `mask_and` + `popcount_batch_u64` over the same span; identical answer, zero intermediate bytes. Falsified if composition already achieves it without a buffer | +| D-WFL-T1-FUSED | the clean case for the anti-zoo rule licensing a NEW T1 primitive: a fused `popcount(a[i] & b[i])` accumulated over a word span. It cannot be expressed by the existing algebra without an intermediate buffer, so it exposes a genuinely new zero-copy operation rather than a convenience | **Shipped** (ndarray #322 + #1270); was: In PR — composition falsifier HELD at the ISA level (`U64x8` ternlog/popcnt everywhere); the slice loop landed generalized over all 256 tables (ndarray #322); differential gate met | differential vs `mask_and` + `popcount_batch_u64` over the same span; identical answer, zero intermediate bytes. Falsified if composition already achieves it without a buffer | ## D-WFL-L0 — the foundational ruling, to land BEFORE any W0/W1 code (2026-09-19) @@ -136,7 +136,7 @@ corrects two errors of my own from earlier the same day. | D-id | scope | status | gate | |---|---|---|---| -| D-WFL-L0 | the six clauses above, landed in `.claude/plans/waben-fold-execution-loop-v1.md` §1 | **In PR (#1251)** | none — a ruling, not a measurement. Its falsifier is textual: if a future session can quote `membrane-tiers.md` to argue a full mask is a legal fold output, clause 6 failed to draw the line | +| D-WFL-L0 | the six clauses above, landed in `.claude/plans/waben-fold-execution-loop-v1.md` §1 | **Shipped (#1251)** | none — a ruling, not a measurement. Its falsifier is textual: if a future session can quote `membrane-tiers.md` to argue a full mask is a legal fold output, clause 6 failed to draw the line | **Wave mapping this makes self-evident:** W1 ADDRESS integrity · W2 compact carrier preservation · W3 ROTATE · W4 spatial/local operators · W5 metacognitive @@ -170,7 +170,7 @@ text this arc had already written. Append-only: the earlier sections stand. | D-WFL-W0 | **Two attestations, not one.** ⊘ `(key, ordinal)` is NOT a fix: with `ordinal` = post-sort position, `K→A,K→B` and `K→B,K→A` both digest as `(K,0),(K,1)`. Split instead — `OrderedLaneWitness` keeps `key_digest = H(K0,K1,…)`; a separate `RowDomain` carries `row_order_digest = H(ID0,ID1,…)` over a STABLE row identity (NodeGuid sequence or writer source ordinals), never the semantic key. Executor requires both. Duplicate keys stay legal per `ordered_lane.rs:194` | Queued | a swap of two rows behind one key must fail while key_digest, lens, version and n_rows are all unchanged. Deliverable is a DECISION: the smallest stable row identity that already exists where lane and planes are assembled — minting a new one is the failure mode | | D-WFL-W1-AUTH | **W0 says WHAT is attested; W1 must say WHO MAY MINT IT.** Without it the defect moves up one level — `Program.row_order_digest == Planes.row_order_digest` is still metadata vs metadata. PREFERRED: `AttestedPlanes<'a>` borrowed from the sealed row image (keys + mask planes + value lanes + row-identity sequence), typed constructor, no public assembly — unattested planes UNREPRESENTABLE. Second best: recompute `H(NodeGuid_0…n)` once at the attachment boundary, cached against the immutable borrow. REJECTED: a digest field on a freely-constructed `Planes` | Queued | THE acceptance case: identical keys / version / lens / n_rows / key_digest / `RowDomain` **copied by the caller**, but row identities A↔B swapped and one value plane swapped to match ⇒ execution MUST refuse before the `Range` is consumed. Red before the fix, green only when the actual executed plane ordering is attested | | D-WFL-W1′ | ⊘ **Carrying a `RowDomain` is not verifying one.** Comparing program-domain against `planes.domain` compares two metadata copies; a caller can permute a lane with every label intact, so a metadata check PASSES the permuted-planes falsifier. Three enforcement shapes: derive the digest from lane contents; carry the seal's permutation; or make `Planes` constructible only from a sealed lane (unattested planes unrepresentable) | Queued | the permuted-planes case is load-bearing precisely because a metadata-only check passes it | -| D-WFL-W2a′ | scratch necessity must be DERIVED from the validated program shape (`requires_scratch() == false` for the range-terminal shape), never an ad-hoc caller flag — a flag is a claim, a derived predicate is a proof | Queued | scratch words required = 0, mask words written = 0, answer = `hi-lo`. A zero-mask program still demanding `words_for(N)` fails the wave | +| D-WFL-W2a′ | scratch necessity must be DERIVED from the validated program shape (`requires_scratch() == false` for the range-terminal shape), never an ad-hoc caller flag — a flag is a claim, a derived predicate is a proof | **Shipped (#1268)** — `Program::requires_scratch()` is derived from `fused_terminal`/`fused_ternlog`, never a caller flag | scratch words required = 0, mask words written = 0, answer = `hi-lo`. A zero-mask program still demanding `words_for(N)` fails the wave | | D-WFL-W5′ | ⊘ the `w_slot` = thought-track hypothesis is close to REFUTED: `AttentionMaskSoA::touch` (`attention_mask.rs:83-88`) finds by `mailbox_id` alone and overwrites `w_slot`, so a second track at one node ERASES the first. Do not write an acceptance criterion assuming two tracks visible at one NodeGuid | Queued | either a separate mailbox identity per track with that mapping carried explicitly, or the identification is wrong | | D-WFL-W5″ | the publication decision must test the DELTA's own emptiness, not `Terminal::RangeAny`. RangeAny is endpoint arithmetic for one un-`under`ed `Pred::Range`; `delta = A_{t+1} \ A_t` is an arbitrary, possibly fragmented mask, so a non-empty PREFIX reports "publish" on every converged step | Queued | a fixed-point step (delta empty, prefix non-empty) must publish nothing | | D-WFL-W6.0 | **ATTEND vs EPISTEMIC FIRE.** `AlphaOverlay` is attention memory, not epistemic truth. Ruling: *a non-empty Boolean delta is sufficient for an ATTENTION effect and never sufficient for an EPISTEMIC effect.* Attention may move focus and the alpha trace; epistemic effects require provenance + evidence identity before touching TruthU8/NARS | Queued | the floor that unblocks W6 without a full cognitive theory. Falsified by any path where a bare intersection revises truth | @@ -191,7 +191,7 @@ consumer.** | D-id | scope | status | gate / falsifier | |---|---|---|---| | D-WFL-W1 | row identity + order attestation: `RowDomain {version, lens, permutation-or-row-identity, n_rows}` on `Planes`, checked in `validate`. Binds to the PERMUTATION — NOT key uniqueness (`ordered_lane.rs:194`: equal keys are indistinguishable, so refusing duplicates is a semantic regression that also proves nothing). Also the precondition for replay (`E-REPLAY-CAN-BE-CHEAPER-THAN-STORAGE-1`) | Queued | wrong-version / wrong-lens / permuted-planes (same rows, same length, same key digest, different associated order), each disable-verified RED first. Falsified if a permuted-planes program still returns the oracle's answer | -| D-WFL-W2a | range-native terminals, ZERO scratch: `Count = hi-lo`, `Any = lo!=hi`, gated on the program being exactly one un-`under`ed `Pred::Range`. Requires `execute()`'s unconditional `scratch.words == words_for(n_rows)` check to become conditional on the program needing planes | Queued | N from 1K to 100M, same width, varying absolute position — **mask words written must be 0**. Falsified if any mask word is written, or cost moves with N or position | +| D-WFL-W2a | range-native terminals, ZERO scratch: `Count = hi-lo`, `Any = lo!=hi`, gated on the program being exactly one un-`under`ed `Pred::Range`. Requires `execute()`'s unconditional `scratch.words == words_for(n_rows)` check to become conditional on the program needing planes | **Shipped (#1268)** — `fused_terminal` un-gated `Pred::Range`: Count = hi−lo, Any = lo Date: Wed, 23 Sep 2026 20:02:36 +0000 Subject: [PATCH 2/5] probe: BULK arm splits the tiled cost into writes vs per-tile dispatch; v3 and v4 measured program_collapse_probe gains a fourth arm. BULK evaluates each op once over the whole touched span, into preallocated buffers, with the same ndarray::simd kernels the executor calls per tile. It writes the same derived words as the tiled path but makes one facade call per op, not one per op per tile. Two derived columns: wr_ns = bulk - fold (the writes) and disp_ns = tiled - bulk (per-tile overhead). Every arm is asserted equal to the bit-serial oracle. Measured at N = 1M, whole population: - disp_ns is 90-94 % of tiled_ns at both v3 and v4 (~30 ns per op per 8-word tile). The writes cost at most ~19 us. Most of #1272's 10-26x is skipping the tile interpreter; write elimination is the smaller share. - At v4 every collapsible 3-plane Count folds in ~6.2 us regardless of its truth table (v3: 8.6-17.3 us). The tiled path does not move between tiers. - At 1 % extents BULK beats the fold, which pays a fixed per-call cost (validate + re-running the symbolic recognizer). Entry: .claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md. Tile size is deliberately NOT changed; it is bound to the scratch-size contract, and the entry records it as OPEN. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- ...23-collapse-probe-v4-and-dispatch-split.md | 41 ++++ .claude/board/entries/README.md | 3 +- .../examples/program_collapse_probe.rs | 212 +++++++++++++++++- 3 files changed, 244 insertions(+), 12 deletions(-) create mode 100644 .claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md diff --git a/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md b/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md new file mode 100644 index 000000000..716bea7a0 --- /dev/null +++ b/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md @@ -0,0 +1,41 @@ +# 2026-09-23 — The tiled path is dispatch-bound, not write-bound; AVX-512 flattens the fold + +**Status:** MEASURED · OPEN (per-tile call overhead; the fold's fixed per-call cost) +**D-ids:** D-WFL-FUSE (measurement follow-up to #1272). Corrects the attribution in `entries/2026-09-23-program-collapse-boolean-chains.md`, which says the tiled gap is "not only write elimination" but does not split it. + +## What was added +`examples/program_collapse_probe.rs` has a fourth arm, **BULK**. It evaluates the same chain once per op over the extent's whole touched span, into whole-population buffers preallocated outside the timing. It uses the same `ndarray::simd` kernels the executor uses per tile (`mask_and/or/xor/andnot/not`, and `ternlog_dispatch` for runtime immediates). So BULK writes every derived word the tiled path writes, but with one facade call per op instead of one per op per tile. + +Two derived columns come from it: +- `wr_ns = bulk − fold`: the cost of the writes themselves. +- `disp_ns = tiled − bulk`: the per-tile overhead on top of the same writes. + +This is a first-order split: BULK still pays one call per op, so `disp_ns` is per-TILE overhead, not all interpretation cost. Every arm is asserted equal to the bit-serial oracle at both tiers. + +Command: `cargo run --release -p lance-graph-mask-risc --example program_collapse_probe`. v4 was built with `RUSTFLAGS="-C target-cpu=x86-64-v4"` into a throwaway target dir. N = 1 048 576 rows (16 384 words), median ns. + +## Whole population, Count + +| chain | tier | fold | bulk | tiled | wr (bulk−fold) | disp (tiled−bulk) | +|---|---|---|---|---|---|---| +| `(a&b)\|!c` | v3 | 8.6 µs | 17.5 µs | 220 µs | 8.8 µs | 202 µs | +| `(a&b)\|!c` | v4 | 6.2 µs | 19.6 µs | 280 µs | 13.4 µs | 260 µs | +| `((a^b)&!c)\|(a&c)` | v3 | 15.5 µs | 20.6 µs | 362 µs | 5.1 µs | 341 µs | +| `((a^b)&!c)\|(a&c)` | v4 | 6.2 µs | 25.5 µs | 354 µs | 19.3 µs | 329 µs | +| `maj(a,b,c)^a` | v3 | 17.3 µs | 16.1 µs | 167 µs | −1.2 µs | 151 µs | +| `maj(a,b,c)^a` | v4 | 6.2 µs | 16.5 µs | 183 µs | 10.3 µs | 167 µs | +| `a&d` (empty by data) | v3 | 6.5 µs | 8.3 µs | 113 µs | 1.8 µs | 105 µs | +| `a&d` (empty by data) | v4 | 4.7 µs | 9.8 µs | 120 µs | 5.1 µs | 110 µs | +| 4-plane (not collapsible) | v3 | — | 28.2 µs | 296 µs | — | 267 µs | +| 4-plane (not collapsible) | v4 | — | 23.8 µs | 274 µs | — | 251 µs | + +## Readings +- **The tiled path's cost is ~90 % per-tile overhead.** `disp_ns` accounts for 90–94 % of `tiled_ns` on every whole-population row, at both tiers. `TILE_WORDS = 8` (one 512-bit vector per facade call, `exec.rs`) means 2 048 tiles per op, and the gap works out to roughly 30 ns per op per tile. The derived-word writes cost at most ~19 µs. **So most of #1272's 10–26× comes from skipping the tile interpreter, not from eliminating writes.** Write elimination is real, but it is the smaller share. +- **AVX-512 flattens the fold.** At v4 every collapsible 3-plane Count costs ~6.2 µs whatever its truth table: native VPTERNLOG is one instruction per word for any immediate. At v3 the same folds cost 8.6–17.3 µs, because AVX2 composes each table from its own instruction sequence. The tiled path does not move between tiers, since it is dispatch-bound. At v4 the fold beats BULK by 2.1–4.1×. At v3 it ranges from 0.9× (`maj^a`, where the fold is slower) to 2.0×. +- **At small extents BULK beats the fold.** At 1 % of the population (165–660 words), BULK takes 95–280 ns against the fold's 190–285 ns. The fold pays a fixed per-call cost: `validate` plus re-running the symbolic recognizer (`fused_ternlog`) on every execute. That cost is only amortized at large extents. +- `wr_ns` is negative on several 1 % rows. That is per-call overhead dominating a few hundred words of work, not a negative write cost. + +## Open +- Per-tile call overhead on the tiled path (~30 ns per op per tile). Tile size is bound to the scratch-size contract (`slots × TILE_WORDS`, independent of `n_rows`), so it is **not** changed here. A larger tile trades scratch footprint for fewer calls. That is a decision for whoever owns the scratch contract, informed by these numbers. OPEN; nothing built. +- The fold's fixed per-call cost. Caching the recognized `FusedTernlog` on `Program` would remove the re-interpretation per execute. Not measured separately; OPEN. +- Only x86-64 was measured. NEON was not. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 27d6f62d0..d81b8cb1e 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,7 +25,7 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -154 entries, 2026-08-06 .. 2026-09-23. +155 entries, 2026-08-06 .. 2026-09-23. | date | entry id | finding | file | |---|---|---|---| @@ -35,6 +35,7 @@ exactly one of them, never both. | 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) | | 2026-09-23 | `program-collapse-boolean-chains` | | [2026-09-23-program-collapse-boolean-chains.md](2026-09-23-program-collapse-boolean-chains.md) | | 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) | +| 2026-09-23 | `collapse-probe-v4-and-dispatch-split` | | [2026-09-23-collapse-probe-v4-and-dispatch-split.md](2026-09-23-collapse-probe-v4-and-dispatch-split.md) | | 2026-09-23 | `absolute-execution-extent` | | [2026-09-23-absolute-execution-extent.md](2026-09-23-absolute-execution-extent.md) | | 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) | | 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) | diff --git a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs index 79cf92ef6..b3eca6e6e 100644 --- a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs +++ b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs @@ -1,6 +1,6 @@ //! Multi-op Boolean chain → Count/Any at N = 1M rows: collapsed fold vs tiled. //! -//! The same raw op sequence, three physical endings: +//! The same raw op sequence, four physical endings: //! //! - FOLD: `Program::fused_ternlog` interprets the chain symbolically and //! lowers it onto ONE `mask_ternlog_{popcount,any}` pass; no slot is carved @@ -11,22 +11,41 @@ //! terminal reads the last one back. Same answer, today's scratch path. //! - KEEP: the chain with a `Keep` terminal into an `Out::Mask`, then //! `popcount_batch_u64` / `mask_any` over the kept words. +//! - BULK: the same op sequence evaluated ONCE per op over the extent's +//! whole touched-word span, into preallocated whole-population scratch +//! buffers (allocated once per chain, outside every timed closure) — the +//! same word WRITES the tiled arm pays, but with no per-tile interpreter +//! dispatch: one `ndarray::simd` facade call per op, full span wide, +//! instead of one call per op per tile. It splits the tiled/fold gap into +//! two additive pieces: `wr_ns = bulk_ns - fold_ns` is the cost of writing +//! derived words at all (fold writes none), and `disp_ns = tiled_ns - +//! bulk_ns` is the cost of the per-tile interpreter loop on top of those +//! same writes. This is a FIRST-ORDER decomposition, not an exact one: +//! bulk still pays one facade call per op (i.e. one round of +//! interpretation), so `disp_ns` is the per-TILE overhead layered on top +//! of that single call, not the cost of interpretation itself. //! -//! Reported per chain × extent: median ns per arm and the derived words the -//! tiled arm writes (`ops × touched words`; the fold writes 0). A four-plane -//! chain is measured on the tiled path only: it is outside the algebra (one -//! ternlog addresses three inputs), and its row records where the collapse -//! stops. Every result is checked against a bit-serial scalar oracle. +//! Reported per chain × extent: median ns per arm (fold/bulk/tiled/keep), +//! the derived words the tiled arm writes (`ops × touched words`; the fold +//! writes 0), and the two BULK-derived columns above. A four-plane chain is +//! measured on the bulk and tiled paths only: it is outside the fold algebra +//! (one ternlog addresses three inputs), and its row records where the +//! collapse stops. Every result is checked against a bit-serial scalar +//! oracle. //! //! `cargo run --release -p lance-graph-mask-risc --example program_collapse_probe` +use std::ops::Range; use std::time::Instant; use lance_graph_mask_risc::exec::{execute_extent, Scratch}; +use lance_graph_mask_risc::ternlog_dispatch::ternlog_dispatch; use lance_graph_mask_risc::{ touched_words, Foreign, MaskOp, Operand, Out, Planes, Program, Terminal, Value, FUSED_SLOT_CAP, }; -use ndarray::simd::{mask_any, popcount_batch_u64}; +use ndarray::simd::{ + mask_and, mask_andnot, mask_any, mask_not, mask_or, mask_xor, popcount_batch_u64, +}; fn lcg(seed: &mut u64) -> u64 { *seed = seed @@ -92,6 +111,141 @@ fn retarget(ops: &[MaskOp], from: u16, to: u16) -> Vec { .collect() } +/// The scratch slot an op writes — every [`MaskOp`] variant has one. +fn op_dst(op: &MaskOp) -> u16 { + match *op { + MaskOp::Pred { dst, .. } + | MaskOp::And { dst, .. } + | MaskOp::Or { dst, .. } + | MaskOp::Xor { dst, .. } + | MaskOp::AndNot { dst, .. } + | MaskOp::Not { dst, .. } + | MaskOp::Ternlog { dst, .. } + | MaskOp::Gather { dst, .. } => dst, + } +} + +/// The bits of word `w` that fall inside `[lo, hi)` (absolute rows) — the +/// BULK arm's own copy of `exec::edge_mask` (private to that module), used +/// to restrict the terminal's first and last touched word to the extent +/// exactly as `exec::run_fused`/`run_fused_ternlog` do. +fn edge_mask(w: usize, lo: usize, hi: usize) -> u64 { + let base = w * 64; + let from = lo.saturating_sub(base).min(64); + let to = (hi - base).min(64); + let upper = if to == 64 { u64::MAX } else { (1u64 << to) - 1 }; + upper & (u64::MAX << from) +} + +/// Every BULK scratch buffer except the one an op is writing — this arm's +/// analogue of `exec::Slots`, over per-slot `Vec` buffers rather than a +/// flat tiled arena. +struct BulkSlots<'s> { + /// Buffers `[0, hole)`. + left: &'s [Vec], + /// Buffers `(hole, bufs.len())`. + right: &'s [Vec], + hole: usize, +} + +impl BulkSlots<'_> { + fn get(&self, i: usize, n: usize) -> &[u64] { + debug_assert!( + i != self.hole, + "bulk arm: slot {i} read while it is this op's write target" + ); + if i < self.hole { + &self.left[i][..n] + } else { + &self.right[i - self.hole - 1][..n] + } + } +} + +/// Evaluate `ops` ONCE per op over `span` (a word range, e.g. from +/// [`touched_words`]) into `bufs`, one preallocated whole-population buffer +/// per scratch slot the chain uses — no per-tile dispatch, one +/// `ndarray::simd` facade call per op over the whole span. Leaf operands +/// (`Operand::Plane`) read `leaves[i][span]` directly, no copy. `bufs` must +/// have one entry per distinct scratch slot `ops` addresses, each at least +/// `span.len()` words long; nothing is allocated here. +/// +/// No edge masking happens here: bits outside the extent `[lo, hi)` but +/// inside `span`'s first/last word are left as whatever the ops computed +/// them to be. That is sound because those positions are masked OUT by +/// [`bulk_terminal`] purely by WORD POSITION, so their VALUE never reaches +/// the reported result — the same reasoning `exec::run_fused_ternlog` relies +/// on for an odd ternlog's tail bits. +fn bulk_eval(ops: &[MaskOp], leaves: &[&[u64]; 4], bufs: &mut [Vec], span: Range) { + let n = span.len(); + for op in ops { + let dst = usize::from(op_dst(op)); + let (left, rest) = bufs.split_at_mut(dst); + let (mid, right) = rest.split_at_mut(1); + let slots = BulkSlots { + left: &*left, + right: &*right, + hole: dst, + }; + let rd = |o: Operand| -> &[u64] { + match o { + Operand::Plane(i) => &leaves[usize::from(i)][span.clone()], + Operand::Scratch(i) => slots.get(usize::from(i), n), + } + }; + let d = &mut mid[0][..n]; + match *op { + MaskOp::And { a, b, .. } => mask_and(rd(a), rd(b), d), + MaskOp::Or { a, b, .. } => mask_or(rd(a), rd(b), d), + MaskOp::Xor { a, b, .. } => mask_xor(rd(a), rd(b), d), + MaskOp::AndNot { a, b, .. } => mask_andnot(rd(a), rd(b), d), + // `n * 64` as the tail-clear bound disables `mask_not`'s own + // clipping (it is a no-op exactly at a word boundary, which `n * + // 64` always is) — any population-edge clipping is `bulk_eval`'s + // caller's job via `bulk_terminal`, not an interior op's. + MaskOp::Not { a, .. } => mask_not(rd(a), n * 64, d), + MaskOp::Ternlog { imm, a, b, c, .. } => ternlog_dispatch(imm, rd(a), rd(b), rd(c), d), + other => panic!("bulk arm: unsupported op {other:?}"), + } + } +} + +/// Fold a BULK terminal slot (`buf`, the `span.len()`-word result of +/// [`bulk_eval`], `span` starting at absolute word `w0`) into `Count` or +/// `Any` over `[lo, hi)`, restricting the first and last word with +/// [`edge_mask`] exactly as `exec::run_fused`/`run_fused_ternlog` do — the +/// interior words are read as they are, since [`touched_words`] guarantees +/// every word strictly between the first and last is wholly inside the +/// extent. +fn bulk_terminal(buf: &[u64], w0: usize, lo: usize, hi: usize, want_count: bool) -> Value { + let n = buf.len(); + if n == 0 { + return if want_count { + Value::Count(0) + } else { + Value::Bool(false) + }; + } + if n == 1 { + let w = buf[0] & edge_mask(w0, lo, hi); + return if want_count { + Value::Count(w.count_ones() as usize) + } else { + Value::Bool(w != 0) + }; + } + let head = [buf[0] & edge_mask(w0, lo, hi)]; + let tail = [buf[n - 1] & edge_mask(w0 + n - 1, lo, hi)]; + let interior = &buf[1..n - 1]; + if want_count { + let c = + popcount_batch_u64(&head) + popcount_batch_u64(interior) + popcount_batch_u64(&tail); + Value::Count(c as usize) + } else { + Value::Bool(mask_any(&head) || mask_any(interior) || mask_any(&tail)) + } +} + fn main() { let n = 1usize << 20; let words = n.div_ceil(64); @@ -207,8 +361,20 @@ fn main() { let extents = [("1%", mid, mid + n / 100), ("whole", 0, n)]; let cap = FUSED_SLOT_CAP as u16; println!( - "{:>17} {:>5} {:>5} {:>3} {:>10} {:>10} {:>10} {:>6} {:>6} {:>9}", - "chain", "ext", "term", "ops", "fold_ns", "tiled_ns", "keep_ns", "t/f", "k/f", "tiled_wr" + "{:>17} {:>5} {:>5} {:>3} {:>10} {:>10} {:>10} {:>10} {:>6} {:>6} {:>9} {:>10} {:>10}", + "chain", + "ext", + "term", + "ops", + "fold_ns", + "bulk_ns", + "tiled_ns", + "keep_ns", + "t/f", + "k/f", + "tiled_wr", + "wr_ns", + "disp_ns" ); for (name, ops, last, oracle) in chains { let term = |mask| (Terminal::Count { mask }, Terminal::Any { mask }); @@ -232,6 +398,12 @@ fn main() { let mut ts = Scratch::for_program(&tiled[0], n).expect("scratch"); let mut ks = Scratch::for_program(&keep, n).expect("scratch"); let mut out = vec![0u64; words]; + // One buffer per distinct scratch slot `ops` addresses, each sized + // to the LARGEST span either extent below touches (`words`, the + // "whole" extent's span) — allocated once per chain, reused across + // both extents and both terminals by slicing to the live span. + let max_slot = ops.iter().map(op_dst).max().unwrap_or(0); + let mut bufs: Vec> = (0..=max_slot).map(|_| vec![0u64; words]).collect(); for (ename, lo, hi) in extents { let want = (lo..hi) .filter(|&r| oracle(bit(&pa, r), bit(&pb, r), bit(&pc, r), bit(&pd, r))) @@ -260,6 +432,16 @@ fn main() { } else { (f64::NAN, expect) }; + let (bns, bv) = median(reps, || { + bulk_eval(&ops, &masks, &mut bufs, span.clone()); + bulk_terminal( + &bufs[usize::from(last)][..span.len()], + span.start, + lo, + hi, + k == 0, + ) + }); let (tns, tv) = median(reps, || { execute_extent( &tiled[k], @@ -289,10 +471,16 @@ fn main() { } }); assert_eq!(fv, expect, "fold {name} {ename} {tname}"); + assert_eq!(bv, expect, "bulk {name} {ename} {tname}"); assert_eq!(tv, expect, "tiled {name} {ename} {tname}"); assert_eq!(kv, expect, "keep {name} {ename} {tname}"); + // NaN propagates automatically: when `fns` is NaN (the 4p + // chain, uncollapsible) `wr_ns` is NaN too, matching the + // fold_ns column's own NaN-means-not-collapsible convention. + let wr_ns = bns - fns; + let disp_ns = tns - bns; println!( - "{name:>17} {ename:>5} {tname:>5} {:>3} {fns:>10.0} {tns:>10.0} {kns:>10.0} {:>6.2} {:>6.2} {:>9}", + "{name:>17} {ename:>5} {tname:>5} {:>3} {fns:>10.0} {bns:>10.0} {tns:>10.0} {kns:>10.0} {:>6.2} {:>6.2} {:>9} {wr_ns:>10.0} {disp_ns:>10.0}", ops.len(), tns / fns, kns / fns, @@ -303,6 +491,8 @@ fn main() { } println!( "(n = {n}, {words} population words; tiled_wr = derived words the tiled path \ - writes, ops × touched words; the fold writes 0; fold_ns NaN = not collapsible)" + writes, ops × touched words; the fold writes 0; fold_ns NaN = not collapsible; \ + wr_ns = bulk_ns - fold_ns (cost of writing derived words); \ + disp_ns = tiled_ns - bulk_ns (per-tile interpreter overhead on top of the same writes))" ); } From 869ba24d3dc9a4619268a67099fccdd89a18b3a3 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 20:08:07 +0000 Subject: [PATCH 3/5] =?UTF-8?q?W-D:=20two-column=20GROUP=20BY=20=E2=80=94?= =?UTF-8?q?=20GroupKey::Pair=20(mask-risc)=20+=20GroupAddr::Pair=20(quack)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Parity item W-D. A two-column `GROUP BY hi, lo` lowers to ONE Terminal::GroupReduce keyed by the composite `hi * stride + lo`, and delegates to ndarray #324's `masked_group_*_pair` kernels. No composite key lane is materialised anywhere. mask-risc - `GroupKey::Pair { hi, lo, stride }`. Both lanes are validated like `Lane` (U32, in range). Four GroupReduce arms: Count / MinI32 / MaxI32 / SumSymI32. Partial-extent refusal and sink seeding match on the TERMINAL, so Pair inherits both unchanged. - The scalar reference computes the composite independently (u64 widen; drop when `lo >= stride` or when the composite is past the universe). It never calls the kernel it checks. quack - `GroupAddr::Pair { hi, lo, stride }` lowers to `GroupKey::Pair`. - `lower_group_avg` refuses a Pair key with the new `LowerError::GroupAvgPairKey` (`LowerError` is already non_exhaustive): no full-range pair GROUP SUM terminal exists, so AVG would have only half its fraction. Tests - mask-risc `foreign.rs`: - Pair equals a precomputed composite Lane under the same executor, and equals `reference_execute`, for all four folds (n = 1000, i.e. more than one tile; >= 8 non-empty groups, counted independently). - A `lo == stride` row that would alias `(1, 0)`'s slot is dropped. - An existing exhaustive match gains an `unreachable!` arm for Pair. The loop there yields Lane/Via only, and Pair has no full-range sum to compare against. - quack DuckDB differential: three new cases, GROUP BY (cost_center, status), 24 groups via key = cost_center*3 + status: COUNT, MIN, and a sparse MAX with 15 NULL groups. Expected values were regenerated by `oracle.py` against DuckDB 1.5.5. The three ids were added to its GROUPED set; all 32 existing rows regenerated byte-identical. Depends on ndarray #324 (CI resolves ndarray via the local path dep). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- crates/lance-graph-mask-risc/src/exec.rs | 47 +++- crates/lance-graph-mask-risc/src/ir.rs | 9 + crates/lance-graph-mask-risc/src/reference.rs | 21 +- crates/lance-graph-mask-risc/tests/foreign.rs | 226 ++++++++++++++++++ crates/lance-graph-quack/src/lib.rs | 31 +++ .../lance-graph-quack/tests/duckdb/cases.tsv | 3 + .../lance-graph-quack/tests/duckdb/oracle.py | 3 + .../tests/duckdb_differential.rs | 91 +++++++ 8 files changed, 427 insertions(+), 4 deletions(-) diff --git a/crates/lance-graph-mask-risc/src/exec.rs b/crates/lance-graph-mask-risc/src/exec.rs index 3aae26d5a..f608507cc 100644 --- a/crates/lance-graph-mask-risc/src/exec.rs +++ b/crates/lance-graph-mask-risc/src/exec.rs @@ -28,9 +28,11 @@ use ndarray::simd::{ le_i32_to_mask, le_i32_to_mask_under, lt_i32_to_mask, lt_i32_to_mask_under, mask_all, mask_and, mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32, mask_not, mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_xor, - mask_xor_assign, masked_group_count_u32, masked_group_count_u32_via, masked_group_max_i32, - masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_via, masked_group_sum_i32, - masked_group_sum_i32_via, masked_group_sum_sym_i32, masked_group_sum_sym_i32_via, + mask_xor_assign, masked_group_count_u32, masked_group_count_u32_pair, + masked_group_count_u32_via, masked_group_max_i32, masked_group_max_i32_pair, + masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_pair, + masked_group_min_i32_via, masked_group_sum_i32, masked_group_sum_i32_via, + masked_group_sum_sym_i32, masked_group_sum_sym_i32_pair, masked_group_sum_sym_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32, masked_sum_i32, ne_i32_to_mask, ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, popcount_batch_u64, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, ternary_match_u64_to_mask, @@ -1398,6 +1400,45 @@ pub fn execute_extent( o, ) } + (GroupKey::Pair { hi, lo, stride }, GroupFold::Count) => { + masked_group_count_u32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::MinI32(v)) => { + masked_group_min_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::MaxI32(v)) => { + masked_group_max_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } + (GroupKey::Pair { hi, lo, stride }, GroupFold::SumSymI32(v)) => { + masked_group_sum_sym_i32_pair( + m, + lane_u32(planes, hi, t), + lane_u32(planes, lo, t), + stride, + lane_i32(planes, v, t), + o, + ) + } } } } diff --git a/crates/lance-graph-mask-risc/src/ir.rs b/crates/lance-graph-mask-risc/src/ir.rs index a892bafdf..775ae3eb8 100644 --- a/crates/lance-graph-mask-risc/src/ir.rs +++ b/crates/lance-graph-mask-risc/src/ir.rs @@ -384,6 +384,15 @@ pub enum GroupKey { /// `U32` lane of this table, `key` a `U32` lane of the foreign table. /// The two hops are fused; no remapped key lane is materialised. Via { fk: u16, key: u16 }, + /// The group of row `i` is `lanes[hi][i] * stride + lanes[lo][i]` — the + /// composite address of a two-column `GROUP BY`. `hi` and `lo` are both + /// `U32` lanes of this table. Fused: no composite key lane is ever + /// materialised. Zero fallback: a minor key `lanes[lo][i] >= stride` + /// names no group and drops the row, exactly like a resolved key past + /// the group universe — same contract as + /// [`ndarray::simd::masked_group_count_u32_pair`] and its + /// min/max/sum siblings. + Pair { hi: u16, lo: u16, stride: u32 }, } /// What a [`Terminal::GroupReduce`] folds into each group's slot. diff --git a/crates/lance-graph-mask-risc/src/reference.rs b/crates/lance-graph-mask-risc/src/reference.rs index 3738ccc7c..9035f9b40 100644 --- a/crates/lance-graph-mask-risc/src/reference.rs +++ b/crates/lance-graph-mask-risc/src/reference.rs @@ -531,6 +531,10 @@ pub(crate) fn validate( check_lane(planes, fk, LaneKind::U32)?; check_foreign_lane(foreign, key, LaneKind::U32)?; } + GroupKey::Pair { hi, lo, .. } => { + check_lane(planes, hi, LaneKind::U32)?; + check_lane(planes, lo, LaneKind::U32)?; + } } match fold { GroupFold::Count => {} @@ -921,7 +925,7 @@ pub fn reference_execute_into( Some(LaneRef::U32(v)) => &v[..], _ => &[][..], }, - GroupKey::Lane(_) => &[][..], + GroupKey::Lane(_) | GroupKey::Pair { .. } => &[][..], }; for r in survivors(mask) { let k = match key { @@ -933,6 +937,21 @@ pub fn reference_execute_into( } remap[idx] as usize } + GroupKey::Pair { hi, lo, stride } => { + let lo_v = u32_at(planes, lo, r); + if lo_v >= stride { + continue; + } + // hi and lo are both u32, stride is u32: widen to + // u64 first so the multiply-add cannot overflow, + // matching ndarray::simd's GroupKeyAddr::Pair. + let hi_v = u32_at(planes, hi, r); + let composite = u64::from(hi_v) * u64::from(stride) + u64::from(lo_v); + match usize::try_from(composite) { + Ok(k) => k, + Err(_) => continue, + } + } }; if k >= o.len() { continue; diff --git a/crates/lance-graph-mask-risc/tests/foreign.rs b/crates/lance-graph-mask-risc/tests/foreign.rs index 6c6bcb518..776c7c163 100644 --- a/crates/lance-graph-mask-risc/tests/foreign.rs +++ b/crates/lance-graph-mask-risc/tests/foreign.rs @@ -1223,6 +1223,9 @@ fn sym_sum_agrees_with_full_range_sum_plus_count() { key, val: 2, }), + // The loop iterates Lane and Via only: a Pair key has no + // full-range group-sum terminal to compare the sym sum against. + GroupKey::Pair { .. } => unreachable!("loop yields Lane and Via only"), }; let count = run(Terminal::GroupReduce { mask: S0, @@ -1337,3 +1340,226 @@ fn group_reduce_refuses_wrong_lanes_and_a_missing_sink() { }) )); } + +/// FAILS IF: `GroupKey::Pair { hi, lo, stride }` (the fused two-column +/// `GROUP BY` address, `group = hi[i] * stride + lo[i]`) disagrees with an +/// EQUIVALENT `GroupKey::Lane` run over a precomputed composite lane, for +/// any of the four folds, OR either disagrees with the independent +/// row-at-a-time oracle. Multi-tile (`n` past `64 * TILE_WORDS`) so the +/// tiled `_pair` kernels are actually exercised, not just their single-tile +/// remainder. +#[test] +fn pair_key_group_reduce_matches_a_precomputed_composite_lane() { + const STRIDE: u32 = 3; + const HI_RANGE: u32 = 5; + const GROUPS: usize = (HI_RANGE * STRIDE) as usize; // 15 + // 64 * TILE_WORDS (8) = 512; this must exceed it to force multiple tiles. + const N: usize = 1000; + + let mut s = 0xF00D_BEEFu64; + let hi: Vec = (0..N) + .map(|_| (lcg(&mut s) % u64::from(HI_RANGE)) as u32) + .collect(); + let lo: Vec = (0..N) + .map(|_| (lcg(&mut s) % u64::from(STRIDE)) as u32) + .collect(); + // Precomputed composite lane, independently reproducing the pair + // formula so it can stand in for `GroupKey::Lane`. + let comp: Vec = hi.iter().zip(&lo).map(|(&h, &l)| h * STRIDE + l).collect(); + let value: Vec = (0..N).map(|_| (lcg(&mut s) % 4000) as i32 - 2000).collect(); + // A third lane purely to drive the filter predicate below. + let flag: Vec = (0..N).map(|_| (lcg(&mut s) % 3) as u32).collect(); + + let lanes = [ + LaneRef::U32(&hi), // 0 + LaneRef::U32(&lo), // 1 + LaneRef::U32(&comp), // 2 + LaneRef::I32(&value), // 3 + LaneRef::U32(&flag), // 4 + ]; + let masks: [&[u64]; 0] = []; + let planes = Planes { + n_rows: N, + masks: &masks, + lanes: &lanes, + }; + let words = tile_words_for(N); + const { + assert!( + N > 64 * lance_graph_mask_risc::exec::TILE_WORDS, + "must span >1 tile" + ) + }; + + for fold in [ + GroupFold::Count, + GroupFold::MinI32(3), + GroupFold::MaxI32(3), + GroupFold::SumSymI32(3), + ] { + let mk_program = |key: GroupKey| { + Program::new( + vec![MaskOp::Pred { + pred: Pred::NeU32 { lane: 4, v: 1 }, + under: None, + dst: 0, + }], + Terminal::GroupReduce { + mask: S0, + key, + fold, + }, + ) + }; + let run = |key: GroupKey| { + let p = mk_program(key); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + let mut out = vec![7i64; GROUPS]; + let got = execute_into( + &p, + &planes, + &Foreign::NONE, + &mut scratch, + Out::I64(&mut out), + ) + .expect("runs"); + (got, out) + }; + + let (pair_val, pair_out) = run(GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }); + let (lane_val, lane_out) = run(GroupKey::Lane(2)); + assert_eq!(pair_val, Value::GroupReduced, "{fold:?}"); + assert_eq!(pair_val, lane_val, "{fold:?}: value"); + assert_eq!( + pair_out, lane_out, + "{fold:?}: Pair must match the composite Lane exactly" + ); + + // Both must also agree with the independent oracle, for the Pair key. + let p_pair = mk_program(GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }); + let mut want_out = vec![-7i64; GROUPS]; + let want = + reference_execute_into(&p_pair, &planes, &Foreign::NONE, Out::I64(&mut want_out)) + .expect("oracle runs"); + assert_eq!(pair_val, want, "{fold:?}: oracle value"); + assert_eq!(pair_out, want_out, "{fold:?}: oracle sink"); + } + + // Anti-vacuity, computed independently of the executor/oracle under + // test (straight from the raw fixture lanes): at least 8 of the 15 + // composite groups must actually be reached by a selected row, or this + // test would pass just as well against an implementation that filled + // every slot identically. + let non_empty = pair_group_non_empty_count(&hi, &lo, &flag, STRIDE, GROUPS); + assert!( + non_empty >= 8, + "need >= 8 non-empty groups, got {non_empty}" + ); +} + +/// Independent count (NOT the executor or the oracle) of how many of the +/// `groups` composite slots see at least one row that both clears the +/// `flag != 1` filter and has an in-range minor key — the anti-vacuity +/// guard above, computed without depending on the code under test. +fn pair_group_non_empty_count( + hi: &[u32], + lo: &[u32], + flag: &[u32], + stride: u32, + groups: usize, +) -> usize { + let mut seen = vec![false; groups]; + for i in 0..hi.len() { + if flag[i] == 1 { + continue; + } + if lo[i] >= stride { + continue; + } + let g = (hi[i] * stride + lo[i]) as usize; + if g < groups { + seen[g] = true; + } + } + seen.iter().filter(|&&b| b).count() +} + +/// FAILS IF: a row whose minor key equals `stride` (or exceeds it) is +/// silently counted into the group its bits would alias to, instead of +/// being dropped — the `GroupKey::Pair` zero-fallback contract +/// (`lo[i] >= stride` names no group). +#[test] +fn pair_key_drops_a_minor_key_at_stride() { + const STRIDE: u32 = 3; + const GROUPS: usize = 8; + // row: hi, lo, and whether its minor key is in-range + // 0: (0, 0) -> comp 0, valid + // 1: (0, 3) -> lo == stride, DROPPED (would alias comp 3 if not guarded) + // 2: (1, 0) -> comp 3, valid -- the genuine occupant of group 3 + // 3: (1, 1) -> comp 4, valid + // 4: (2, 5) -> lo > stride, DROPPED + let hi: Vec = vec![0, 0, 1, 1, 2]; + let lo: Vec = vec![0, 3, 0, 1, 5]; + let n = hi.len(); + let lanes = [LaneRef::U32(&hi), LaneRef::U32(&lo)]; + let all_bits = [0b1_1111u64]; + let masks: [&[u64]; 1] = [&all_bits]; + let planes = Planes { + n_rows: n, + masks: &masks, + lanes: &lanes, + }; + let p = Program::new( + vec![], + Terminal::GroupReduce { + mask: Operand::Plane(0), + key: GroupKey::Pair { + hi: 0, + lo: 1, + stride: STRIDE, + }, + fold: GroupFold::Count, + }, + ); + let words = tile_words_for(n); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + let mut got_out = vec![9i64; GROUPS]; + let got = execute_into( + &p, + &planes, + &Foreign::NONE, + &mut scratch, + Out::I64(&mut got_out), + ) + .expect("runs"); + let mut want_out = vec![-9i64; GROUPS]; + let want = reference_execute_into(&p, &planes, &Foreign::NONE, Out::I64(&mut want_out)) + .expect("oracle runs"); + assert_eq!(got, Value::GroupReduced); + assert_eq!(got, want); + assert_eq!(got_out, want_out); + // Only the genuine (1, 0) row (row 2) lands in group 3; the aliased, + // out-of-range row 1 must not also be counted there. + assert_eq!( + got_out[3], 1, + "group (1,0)=3 must count only the genuine row, not the row whose lo==stride" + ); + // Exactly 3 of the 5 rows have an in-range minor key (rows 0, 2, 3). + assert_eq!( + got_out.iter().sum::(), + 3, + "rows with lo >= stride must never be counted anywhere" + ); +} diff --git a/crates/lance-graph-quack/src/lib.rs b/crates/lance-graph-quack/src/lib.rs index 9db93881f..68c072492 100644 --- a/crates/lance-graph-quack/src/lib.rs +++ b/crates/lance-graph-quack/src/lib.rs @@ -768,6 +768,12 @@ pub enum LowerError { /// the executor refuses a zero-length sink — so the plan would be /// unexecutable. Refused here, where the caller can see why. EmptyGroupUniverse, + /// `GROUP BY (hi, lo) AVG(val)`: an AVG over a two-column key. There is + /// no full-range pair `GROUP SUM` terminal yet — the count half + /// ([`GroupAgg::Count`] over [`GroupAddr::Pair`]) would lower, the sum + /// half has nowhere to go — so [`lower_group_avg`] refuses rather than + /// answer with only half the fraction. + GroupAvgPairKey, } impl core::fmt::Display for LowerError { @@ -788,6 +794,12 @@ impl core::fmt::Display for LowerError { LowerError::EmptyGroupUniverse => { write!(f, "a grouped plan needs at least one group (groups == 0)") } + LowerError::GroupAvgPairKey => { + write!( + f, + "AVG over a two-column (Pair) group key has no SUM terminal yet" + ) + } } } } @@ -920,6 +932,19 @@ pub enum GroupAddr { /// The group key column on the FOREIGN table. key: ForeignLane, }, + /// `hi[row] * stride + lo[row]` — a two-column `GROUP BY` over a + /// composite address, both `hi` and `lo` `u32` columns of this table. + /// Fused: no composite key column is ever materialised. Same + /// zero-fallback as [`GroupKey::Pair`]: a minor key `lo[row] >= stride` + /// names no group and drops the row. + Pair { + /// The major key column (`u32`). + hi: Col, + /// The minor key column (`u32`), valid only in `0..stride`. + lo: Col, + /// The minor key's cardinality: `group = hi * stride + lo`. + stride: u32, + }, } /// What an [`Agg::GroupReduce`] folds per group. @@ -1225,6 +1250,7 @@ pub fn lower_group_avg( let sum_agg = match key { GroupAddr::Local(k) => Agg::GroupSumI32 { key: k, val }, GroupAddr::Via { fk, key } => Agg::GroupSumViaI32 { fk, key, val }, + GroupAddr::Pair { .. } => return Err(LowerError::GroupAvgPairKey), }; Ok(GroupAvgPlan { sum: lower(&Query { @@ -2018,6 +2044,11 @@ fn terminal_of(agg: Agg, mask: Operand) -> Terminal { fk: fk.0, key: key.0, }, + GroupAddr::Pair { hi, lo, stride } => GroupKey::Pair { + hi: hi.0, + lo: lo.0, + stride, + }, }, fold: agg.fold(), }, diff --git a/crates/lance-graph-quack/tests/duckdb/cases.tsv b/crates/lance-graph-quack/tests/duckdb/cases.tsv index c5b0a30b4..07592745d 100644 --- a/crates/lance-graph-quack/tests/duckdb/cases.tsv +++ b/crates/lance-graph-quack/tests/duckdb/cases.tsv @@ -18,6 +18,9 @@ join_group_sum_country SELECT p.country, SUM(l.amount) FROM line l JOIN partner group_min_cc WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MIN(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=1 GROUP BY g.cost_center ORDER BY g.cost_center 0:-4962;1:-4983;2:-4965;3:-4997;4:-4904;5:-4997;6:-4941;7:-4937 group_max_cc WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=1 GROUP BY g.cost_center ORDER BY g.cost_center 0:19951;1:19872;2:19980;3:19977;4:19967;5:19967;6:19998;7:19848 group_max_cc_sparse WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center) SELECT g.cost_center, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=2 AND l.qty>45 AND l.cost_center<3 GROUP BY g.cost_center ORDER BY g.cost_center 0:17719;1:14535;2:13253;3:NULL;4:NULL;5:NULL;6:NULL;7:NULL +group2_count_cc_st WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, COUNT(l.rid) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>10 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:91;1:280;2:49;3:83;4:297;5:47;6:87;7:291;8:42;9:79;10:288;11:44;12:68;13:266;14:41;15:73;16:319;17:40;18:82;19:285;20:45;21:105;22:255;23:35 +group2_min_cc_st WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, MIN(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>10 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:-4881;1:-4962;2:-3625;3:-4950;4:-4983;5:-4426;6:-4053;7:-4965;8:-4827;9:-4866;10:-4984;11:-3734;12:-4655;13:-4904;14:-3438;15:-4384;16:-4997;17:-4534;18:-4800;19:-4941;20:-4629;21:-4772;22:-4937;23:-954 +group2_max_cc_st_sparse WITH gc AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS cost_center), gs AS (SELECT UNNEST(GENERATE_SERIES(0,2)) AS status), g AS (SELECT gc.cost_center, gs.status FROM gc CROSS JOIN gs) SELECT g.cost_center*3+g.status, MAX(l.amount) FROM g LEFT JOIN line l ON l.cost_center=g.cost_center AND l.status=g.status AND l.qty>45 AND l.cost_center<3 GROUP BY g.cost_center, g.status ORDER BY g.cost_center*3+g.status 0:17771;1:19951;2:17719;3:19586;4:19458;5:14535;6:15928;7:19393;8:13253;9:NULL;10:NULL;11:NULL;12:NULL;13:NULL;14:NULL;15:NULL;16:NULL;17:NULL;18:NULL;19:NULL;20:NULL;21:NULL;22:NULL;23:NULL join_group_count_country WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS country) SELECT g.country, COUNT(l.rid) FROM g LEFT JOIN (line l JOIN partner p ON p.rid=l.partner_id AND l.status=1) ON p.country=g.country GROUP BY g.country ORDER BY g.country 0:562;1:227;2:714;3:173;4:283;5:208;6:347;7:311 join_group_min_country WITH g AS (SELECT UNNEST(GENERATE_SERIES(0,7)) AS country) SELECT g.country, MIN(l.amount) FROM g LEFT JOIN (line l JOIN partner p ON p.rid=l.partner_id AND l.status=1) ON p.country=g.country GROUP BY g.country ORDER BY g.country 0:-4984;1:-4690;2:-4983;3:-4902;4:-4903;5:-4997;6:-4966;7:-4997 sum_case_posted_cc3 SELECT SUM(CASE WHEN cost_center=3 THEN amount ELSE 0 END) FROM line WHERE status=1 2446433 diff --git a/crates/lance-graph-quack/tests/duckdb/oracle.py b/crates/lance-graph-quack/tests/duckdb/oracle.py index 7db146852..df13ae3b4 100644 --- a/crates/lance-graph-quack/tests/duckdb/oracle.py +++ b/crates/lance-graph-quack/tests/duckdb/oracle.py @@ -47,6 +47,9 @@ "having_sparse_count_cc", "having_sparse_sum_lt_cc", "join_having_count_country", + "group2_count_cc_st", + "group2_min_cc_st", + "group2_max_cc_st_sparse", } BOOLEAN = {"exists_neg"} ROWS = {"rows_proj"} diff --git a/crates/lance-graph-quack/tests/duckdb_differential.rs b/crates/lance-graph-quack/tests/duckdb_differential.rs index 1ea9124bd..11fdabbbf 100644 --- a/crates/lance-graph-quack/tests/duckdb_differential.rs +++ b/crates/lance-graph-quack/tests/duckdb_differential.rs @@ -772,6 +772,97 @@ fn group_max_cc_sparse() { assert_case(&cases, "group_max_cc_sparse", &actual); } +/// `COUNT(*) GROUP BY (cost_center, status)` — the two-column `GROUP BY`, +/// `GroupAddr::Pair { hi: COST_CENTER, lo: STATUS, stride: 3 }` (status is +/// `0..3`, so `stride=3` is the minor key's real cardinality). The SQL +/// encodes the SAME composite key the Rust side folds: +/// `cost_center * 3 + status`. +#[test] +fn group2_count_cc_st() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_count_cc_st", + &planes, + &Foreign::NONE, + Filter::cmp(QTY, Cmp::GtI32(10)), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::Count, + 24, + ); + print_metric("group2_count_cc_st", &m); + assert_case(&cases, "group2_count_cc_st", &actual); +} + +/// `MIN(amount) GROUP BY (cost_center, status)` — same `Pair` key as +/// [`group2_count_cc_st`], a different fold. +#[test] +fn group2_min_cc_st() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_min_cc_st", + &planes, + &Foreign::NONE, + Filter::cmp(QTY, Cmp::GtI32(10)), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::MinI32(AMOUNT), + 24, + ); + print_metric("group2_min_cc_st", &m); + assert_case(&cases, "group2_min_cc_st", &actual); +} + +/// `MAX(amount) GROUP BY (cost_center, status)` over a filter tight enough +/// (`qty>45 AND cost_center<3`) that whole `cost_center` bands are empty — +/// the two-column sibling of [`group_max_cc_sparse`], exercising the SQL +/// `NULL` encoding of an empty group under the `Pair` key. +#[test] +fn group2_max_cc_st_sparse() { + let cases = load_cases(); + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let planes = lanes.planes(); + let (actual, m) = run_group_reduce( + "group2_max_cc_st_sparse", + &planes, + &Foreign::NONE, + Filter::and([ + Filter::cmp(QTY, Cmp::GtI32(45)), + Filter::or([ + Filter::cmp(COST_CENTER, Cmp::EqU32(0)), + Filter::cmp(COST_CENTER, Cmp::EqU32(1)), + Filter::cmp(COST_CENTER, Cmp::EqU32(2)), + ]), + ]), + GroupAddr::Pair { + hi: COST_CENTER, + lo: STATUS, + stride: 3, + }, + GroupAgg::MaxI32(AMOUNT), + 24, + ); + assert!( + actual.contains(":NULL"), + "fixture must leave some group empty or this case tests nothing: {actual}" + ); + print_metric("group2_max_cc_st_sparse", &m); + assert_case(&cases, "group2_max_cc_st_sparse", &actual); +} + #[test] fn group_sum_cc() { let cases = load_cases(); From 69b88b705f89b2f0a6bd4e7bcd9401d016b34249 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 20:09:31 +0000 Subject: [PATCH 4/5] =?UTF-8?q?board:=20W-D=20entry=20=E2=80=94=20two-colu?= =?UTF-8?q?mn=20GROUP=20BY=20as=20one=20composite-addressed=20GroupReduce?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records the ndarray #324 + mask-risc + quack landing, the DuckDB gate, the disable runs (including the vacuous first ndarray stride test that was caught and rewritten), and the open items: AVG / full-range SUM over a Pair key, and more than two key columns. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- ...2026-09-23-quack-w-d-multi-key-group-by.md | 25 +++++++++++++++++++ .claude/board/entries/README.md | 3 ++- 2 files changed, 27 insertions(+), 1 deletion(-) create mode 100644 .claude/board/entries/2026-09-23-quack-w-d-multi-key-group-by.md diff --git a/.claude/board/entries/2026-09-23-quack-w-d-multi-key-group-by.md b/.claude/board/entries/2026-09-23-quack-w-d-multi-key-group-by.md new file mode 100644 index 000000000..cd45382d2 --- /dev/null +++ b/.claude/board/entries/2026-09-23-quack-w-d-multi-key-group-by.md @@ -0,0 +1,25 @@ +# 2026-09-23 — W-D: two-column GROUP BY as one composite-addressed GroupReduce + +**Status:** MEASURED (DuckDB differential green; disable-verified) · OPEN (AVG / full-range SUM over a Pair key; >2 key columns) +**Supersedes:** the W-D line under "What is still open" in `entries/2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md` and `entries/2026-09-23-quack-having-sym-sum-presence-mask.md`. + +## What landed +- **ndarray #324:** `GroupKeyAddr::Pair { hi, lo, stride }`, a third address for the ONE keyed-reduction walker (no new walker). The group is `hi * stride + lo`, computed in u64 (cannot overflow). A minor key `lo >= stride` names no group and is dropped, which is the family's zero-fallback. Five public `masked_group_*_pair` kernels, each reusing its Resident sibling's fold closure unchanged. +- **mask-risc:** `GroupKey::Pair`. Four `GroupReduce` folds: Count, MinI32, MaxI32, SumSymI32. The scalar reference computes the composite independently. +- **quack:** `GroupAddr::Pair` lowers to `GroupKey::Pair`. `lower_group_avg` refuses a Pair key with `LowerError::GroupAvgPairKey` rather than answer with half a fraction. + +No composite key lane is materialised at any layer. + +## Gates +- DuckDB oracle (1.5.5), `GROUP BY (cost_center, status)`, 24 groups encoded as `cost_center*3 + status`: COUNT, MIN, and a sparse MAX in which 15 groups read NULL. The expected values were regenerated by `oracle.py`; all 32 existing rows regenerated byte-identical. +- mask-risc: Pair equals a precomputed composite `Lane` under the same executor, and equals `reference_execute`, for all four folds. Test size is n = 1000, spanning more than one tile. A `lo == stride` row that would alias `(1, 0)`'s slot is dropped. +- **Disable runs** (committed first; each went red, then was restored): + - ndarray: deleting the stride guard, transposing `hi`/`lo`, and u32 wrapping arithmetic. + - mask-risc exec: `hi`/`lo` swapped. + - quack: lowering with `stride + 1`. + - reference: the stride guard dropped. +- Correction made along the way: the ndarray stride test was **vacuous** in its first version. Its dropped rows also fell past the group universe, so the universe check dropped them too. It was rewritten so each dropped row would otherwise alias a real slot. + +## Open +- AVG and full-range (coalescing) SUM over a Pair key. `masked_group_sum_i32_pair` exists in ndarray, but no mask-risc terminal routes a Pair key to it (`GroupSumI32` / `GroupSumViaI32` take a single lane). Refused by name today. +- More than two key columns. Nothing is built for it; a nested Pair or an N-ary address would be its own decision. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index d81b8cb1e..d64164dc1 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,13 +25,14 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -155 entries, 2026-08-06 .. 2026-09-23. +156 entries, 2026-08-06 .. 2026-09-23. | date | entry id | finding | file | |---|---|---|---| | 2026-09-23 | `ternlog-count-any-fold` | | [2026-09-23-ternlog-count-any-fold.md](2026-09-23-ternlog-count-any-fold.md) | | 2026-09-23 | `terminal-elects-materialization-range-plane` | | [2026-09-23-terminal-elects-materialization-range-plane.md](2026-09-23-terminal-elects-materialization-range-plane.md) | | 2026-09-23 | `report-plan-zero-copy-pivot-docir-convergence` | ReportPlan lowers into Quack; pivot = view over shared Arc; reports become OGAR ObjectSlot sources via grid_of (OGAR #307) | [2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md](2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md) | +| 2026-09-23 | `quack-w-d-multi-key-group-by` | | [2026-09-23-quack-w-d-multi-key-group-by.md](2026-09-23-quack-w-d-multi-key-group-by.md) | | 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) | | 2026-09-23 | `program-collapse-boolean-chains` | | [2026-09-23-program-collapse-boolean-chains.md](2026-09-23-program-collapse-boolean-chains.md) | | 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) | From 0088f5996c349f68605d2a5cec753be53be19858 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 20:34:06 +0000 Subject: [PATCH 5/5] probe: report wr_ns / disp_ns as path gaps, and correct the per-tile figure to 33-54 ns Addresses two CodeRabbit findings on #1275, both verified against the data. - The "~30 ns per op per tile" figure was an arithmetic error. The v3/v4 whole-population rows give disp_ns / (ops x 2048 tiles) = 33-54 ns. The Readings line and the Open line now carry that range. - wr_ns and disp_ns were described as isolated costs (the writes; per-tile dispatch). They are gaps between execution paths: - bulk - fold also carries the pass-count/read difference (k passes vs one fused pass); - tiled - bulk also carries batch-size effects. The probe doc, its printed footer and the entry now say so, and add that isolating either cost would need matched-work controls. The entry title and the v4 reading no longer call the tiled path "dispatch-bound". No measurement changed; only what the numbers are claimed to isolate. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- ...23-collapse-probe-v4-and-dispatch-split.md | 18 +++++++-------- .../examples/program_collapse_probe.rs | 22 ++++++++++--------- 2 files changed, 21 insertions(+), 19 deletions(-) diff --git a/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md b/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md index 716bea7a0..06676461e 100644 --- a/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md +++ b/.claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md @@ -1,16 +1,16 @@ -# 2026-09-23 — The tiled path is dispatch-bound, not write-bound; AVX-512 flattens the fold +# 2026-09-23 — Most of the tiled gap is per-tile, not the writes; AVX-512 flattens the fold -**Status:** MEASURED · OPEN (per-tile call overhead; the fold's fixed per-call cost) +**Status:** MEASURED · OPEN (the per-tile gap's decomposition; the fold's fixed per-call cost) **D-ids:** D-WFL-FUSE (measurement follow-up to #1272). Corrects the attribution in `entries/2026-09-23-program-collapse-boolean-chains.md`, which says the tiled gap is "not only write elimination" but does not split it. ## What was added `examples/program_collapse_probe.rs` has a fourth arm, **BULK**. It evaluates the same chain once per op over the extent's whole touched span, into whole-population buffers preallocated outside the timing. It uses the same `ndarray::simd` kernels the executor uses per tile (`mask_and/or/xor/andnot/not`, and `ternlog_dispatch` for runtime immediates). So BULK writes every derived word the tiled path writes, but with one facade call per op instead of one per op per tile. -Two derived columns come from it: -- `wr_ns = bulk − fold`: the cost of the writes themselves. -- `disp_ns = tiled − bulk`: the per-tile overhead on top of the same writes. +Two derived columns come from it. Both are **measured gaps between paths, not isolated costs**: +- `wr_ns = bulk − fold`: the gap between one fused pass (reads at most three planes, writes nothing) and one pass per op over whole-span buffers. It mixes the write cost with the pass-count and read difference. +- `disp_ns = tiled − bulk`: the gap between the same ops run per 8-word tile and run once over the whole span. It mixes per-call overhead with batch-size effects (cache residency, loop setup). -This is a first-order split: BULK still pays one call per op, so `disp_ns` is per-TILE overhead, not all interpretation cost. Every arm is asserted equal to the bit-serial oracle at both tiers. +Attributing either gap to a single cause would need matched-work controls, which this probe does not have. Every arm is asserted equal to the bit-serial oracle at both tiers. Command: `cargo run --release -p lance-graph-mask-risc --example program_collapse_probe`. v4 was built with `RUSTFLAGS="-C target-cpu=x86-64-v4"` into a throwaway target dir. N = 1 048 576 rows (16 384 words), median ns. @@ -30,12 +30,12 @@ Command: `cargo run --release -p lance-graph-mask-risc --example program_collaps | 4-plane (not collapsible) | v4 | — | 23.8 µs | 274 µs | — | 251 µs | ## Readings -- **The tiled path's cost is ~90 % per-tile overhead.** `disp_ns` accounts for 90–94 % of `tiled_ns` on every whole-population row, at both tiers. `TILE_WORDS = 8` (one 512-bit vector per facade call, `exec.rs`) means 2 048 tiles per op, and the gap works out to roughly 30 ns per op per tile. The derived-word writes cost at most ~19 µs. **So most of #1272's 10–26× comes from skipping the tile interpreter, not from eliminating writes.** Write elimination is real, but it is the smaller share. -- **AVX-512 flattens the fold.** At v4 every collapsible 3-plane Count costs ~6.2 µs whatever its truth table: native VPTERNLOG is one instruction per word for any immediate. At v3 the same folds cost 8.6–17.3 µs, because AVX2 composes each table from its own instruction sequence. The tiled path does not move between tiers, since it is dispatch-bound. At v4 the fold beats BULK by 2.1–4.1×. At v3 it ranges from 0.9× (`maj^a`, where the fold is slower) to 2.0×. +- **90–94 % of the tiled path is the tiled−bulk gap.** `disp_ns` is 90–94 % of `tiled_ns` on every whole-population row, at both tiers. `TILE_WORDS = 8` (one 512-bit vector per facade call, `exec.rs`) means 2 048 tiles per op, so the gap works out to **33–54 ns per op per tile** across the ten rows. The bulk−fold gap is at most ~19 µs. **So most of #1272's 10–26× is the per-tile execution path, not the derived-word writes.** This is a gap between two paths; how it divides between call overhead and batch-size effects is not separated here. +- **AVX-512 flattens the fold.** At v4 every collapsible 3-plane Count costs ~6.2 µs whatever its truth table: native VPTERNLOG is one instruction per word for any immediate. At v3 the same folds cost 8.6–17.3 µs, because AVX2 composes each table from its own instruction sequence. The tiled path does not move between tiers; it is dominated by the per-tile gap, which the wider vector does not shrink. At v4 the fold beats BULK by 2.1–4.1×. At v3 it ranges from 0.9× (`maj^a`, where the fold is slower) to 2.0×. - **At small extents BULK beats the fold.** At 1 % of the population (165–660 words), BULK takes 95–280 ns against the fold's 190–285 ns. The fold pays a fixed per-call cost: `validate` plus re-running the symbolic recognizer (`fused_ternlog`) on every execute. That cost is only amortized at large extents. - `wr_ns` is negative on several 1 % rows. That is per-call overhead dominating a few hundred words of work, not a negative write cost. ## Open -- Per-tile call overhead on the tiled path (~30 ns per op per tile). Tile size is bound to the scratch-size contract (`slots × TILE_WORDS`, independent of `n_rows`), so it is **not** changed here. A larger tile trades scratch footprint for fewer calls. That is a decision for whoever owns the scratch contract, informed by these numbers. OPEN; nothing built. +- The per-tile gap on the tiled path (33–54 ns per op per tile, split between call overhead and batch-size effects not yet separated). Tile size is bound to the scratch-size contract (`slots × TILE_WORDS`, independent of `n_rows`), so it is **not** changed here. A larger tile trades scratch footprint for fewer calls. That is a decision for whoever owns the scratch contract, informed by these numbers. OPEN; nothing built. - The fold's fixed per-call cost. Caching the recognized `FusedTernlog` on `Program` would remove the re-interpretation per execute. Not measured separately; OPEN. - Only x86-64 was measured. NEON was not. diff --git a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs index b3eca6e6e..cc0af61ad 100644 --- a/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs +++ b/crates/lance-graph-mask-risc/examples/program_collapse_probe.rs @@ -16,14 +16,16 @@ //! buffers (allocated once per chain, outside every timed closure) — the //! same word WRITES the tiled arm pays, but with no per-tile interpreter //! dispatch: one `ndarray::simd` facade call per op, full span wide, -//! instead of one call per op per tile. It splits the tiled/fold gap into -//! two additive pieces: `wr_ns = bulk_ns - fold_ns` is the cost of writing -//! derived words at all (fold writes none), and `disp_ns = tiled_ns - -//! bulk_ns` is the cost of the per-tile interpreter loop on top of those -//! same writes. This is a FIRST-ORDER decomposition, not an exact one: -//! bulk still pays one facade call per op (i.e. one round of -//! interpretation), so `disp_ns` is the per-TILE overhead layered on top -//! of that single call, not the cost of interpretation itself. +//! instead of one call per op per tile. Two derived columns report the +//! MEASURED GAPS between these paths, not isolated costs: +//! `wr_ns = bulk_ns - fold_ns` is the gap between one fused pass that +//! reads at most three planes and writes nothing, and one pass PER OP that +//! reads and writes whole-span buffers; it mixes write cost with the +//! pass-count and read difference. `disp_ns = tiled_ns - bulk_ns` is the +//! gap between the same ops run per 8-word tile and run once over the +//! whole span; it mixes per-call overhead with batch-size effects (cache +//! residency, loop setup). Attributing either gap to one cause would need +//! matched-work controls this probe does not have. //! //! Reported per chain × extent: median ns per arm (fold/bulk/tiled/keep), //! the derived words the tiled arm writes (`ops × touched words`; the fold @@ -492,7 +494,7 @@ fn main() { println!( "(n = {n}, {words} population words; tiled_wr = derived words the tiled path \ writes, ops × touched words; the fold writes 0; fold_ns NaN = not collapsible; \ - wr_ns = bulk_ns - fold_ns (cost of writing derived words); \ - disp_ns = tiled_ns - bulk_ns (per-tile interpreter overhead on top of the same writes))" + wr_ns = bulk_ns - fold_ns and disp_ns = tiled_ns - bulk_ns are measured \ + gaps between paths, not isolated write or dispatch costs)" ); }