Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 9 additions & 9 deletions .claude/board/STATUS_BOARD.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# 2026-09-23 — Most of the tiled gap is per-tile, not the writes; AVX-512 flattens the fold

**Status:** MEASURED · OPEN (the per-tile gap's decomposition; the fold's fixed per-call cost)
**D-ids:** D-WFL-FUSE (measurement follow-up to #1272). Corrects the attribution in `entries/2026-09-23-program-collapse-boolean-chains.md`, which says the tiled gap is "not only write elimination" but does not split it.

## What was added
`examples/program_collapse_probe.rs` has a fourth arm, **BULK**. It evaluates the same chain once per op over the extent's whole touched span, into whole-population buffers preallocated outside the timing. It uses the same `ndarray::simd` kernels the executor uses per tile (`mask_and/or/xor/andnot/not`, and `ternlog_dispatch` for runtime immediates). So BULK writes every derived word the tiled path writes, but with one facade call per op instead of one per op per tile.

Two derived columns come from it. Both are **measured gaps between paths, not isolated costs**:
- `wr_ns = bulk − fold`: the gap between one fused pass (reads at most three planes, writes nothing) and one pass per op over whole-span buffers. It mixes the write cost with the pass-count and read difference.
- `disp_ns = tiled − bulk`: the gap between the same ops run per 8-word tile and run once over the whole span. It mixes per-call overhead with batch-size effects (cache residency, loop setup).

Attributing either gap to a single cause would need matched-work controls, which this probe does not have. Every arm is asserted equal to the bit-serial oracle at both tiers.

Command: `cargo run --release -p lance-graph-mask-risc --example program_collapse_probe`. v4 was built with `RUSTFLAGS="-C target-cpu=x86-64-v4"` into a throwaway target dir. N = 1 048 576 rows (16 384 words), median ns.

## Whole population, Count

| chain | tier | fold | bulk | tiled | wr (bulk−fold) | disp (tiled−bulk) |
|---|---|---|---|---|---|---|
| `(a&b)\|!c` | v3 | 8.6 µs | 17.5 µs | 220 µs | 8.8 µs | 202 µs |
| `(a&b)\|!c` | v4 | 6.2 µs | 19.6 µs | 280 µs | 13.4 µs | 260 µs |
| `((a^b)&!c)\|(a&c)` | v3 | 15.5 µs | 20.6 µs | 362 µs | 5.1 µs | 341 µs |
| `((a^b)&!c)\|(a&c)` | v4 | 6.2 µs | 25.5 µs | 354 µs | 19.3 µs | 329 µs |
| `maj(a,b,c)^a` | v3 | 17.3 µs | 16.1 µs | 167 µs | −1.2 µs | 151 µs |
| `maj(a,b,c)^a` | v4 | 6.2 µs | 16.5 µs | 183 µs | 10.3 µs | 167 µs |
| `a&d` (empty by data) | v3 | 6.5 µs | 8.3 µs | 113 µs | 1.8 µs | 105 µs |
| `a&d` (empty by data) | v4 | 4.7 µs | 9.8 µs | 120 µs | 5.1 µs | 110 µs |
| 4-plane (not collapsible) | v3 | — | 28.2 µs | 296 µs | — | 267 µs |
| 4-plane (not collapsible) | v4 | — | 23.8 µs | 274 µs | — | 251 µs |

## Readings
- **90–94 % of the tiled path is the tiled−bulk gap.** `disp_ns` is 90–94 % of `tiled_ns` on every whole-population row, at both tiers. `TILE_WORDS = 8` (one 512-bit vector per facade call, `exec.rs`) means 2 048 tiles per op, so the gap works out to **33–54 ns per op per tile** across the ten rows. The bulk−fold gap is at most ~19 µs. **So most of #1272's 10–26× is the per-tile execution path, not the derived-word writes.** This is a gap between two paths; how it divides between call overhead and batch-size effects is not separated here.
- **AVX-512 flattens the fold.** At v4 every collapsible 3-plane Count costs ~6.2 µs whatever its truth table: native VPTERNLOG is one instruction per word for any immediate. At v3 the same folds cost 8.6–17.3 µs, because AVX2 composes each table from its own instruction sequence. The tiled path does not move between tiers; it is dominated by the per-tile gap, which the wider vector does not shrink. At v4 the fold beats BULK by 2.1–4.1×. At v3 it ranges from 0.9× (`maj^a`, where the fold is slower) to 2.0×.
- **At small extents BULK beats the fold.** At 1 % of the population (165–660 words), BULK takes 95–280 ns against the fold's 190–285 ns. The fold pays a fixed per-call cost: `validate` plus re-running the symbolic recognizer (`fused_ternlog`) on every execute. That cost is only amortized at large extents.
- `wr_ns` is negative on several 1 % rows. That is per-call overhead dominating a few hundred words of work, not a negative write cost.

## Open
- The per-tile gap on the tiled path (33–54 ns per op per tile, split between call overhead and batch-size effects not yet separated). Tile size is bound to the scratch-size contract (`slots × TILE_WORDS`, independent of `n_rows`), so it is **not** changed here. A larger tile trades scratch footprint for fewer calls. That is a decision for whoever owns the scratch contract, informed by these numbers. OPEN; nothing built.
- The fold's fixed per-call cost. Caching the recognized `FusedTernlog` on `Program` would remove the re-interpretation per execute. Not measured separately; OPEN.
- Only x86-64 was measured. NEON was not.
25 changes: 25 additions & 0 deletions .claude/board/entries/2026-09-23-quack-w-d-multi-key-group-by.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# 2026-09-23 — W-D: two-column GROUP BY as one composite-addressed GroupReduce

**Status:** MEASURED (DuckDB differential green; disable-verified) · OPEN (AVG / full-range SUM over a Pair key; >2 key columns)
**Supersedes:** the W-D line under "What is still open" in `entries/2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md` and `entries/2026-09-23-quack-having-sym-sum-presence-mask.md`.

## What landed
- **ndarray #324:** `GroupKeyAddr::Pair { hi, lo, stride }`, a third address for the ONE keyed-reduction walker (no new walker). The group is `hi * stride + lo`, computed in u64 (cannot overflow). A minor key `lo >= stride` names no group and is dropped, which is the family's zero-fallback. Five public `masked_group_*_pair` kernels, each reusing its Resident sibling's fold closure unchanged.
- **mask-risc:** `GroupKey::Pair`. Four `GroupReduce` folds: Count, MinI32, MaxI32, SumSymI32. The scalar reference computes the composite independently.
- **quack:** `GroupAddr::Pair` lowers to `GroupKey::Pair`. `lower_group_avg` refuses a Pair key with `LowerError::GroupAvgPairKey` rather than answer with half a fraction.

No composite key lane is materialised at any layer.

## Gates
- DuckDB oracle (1.5.5), `GROUP BY (cost_center, status)`, 24 groups encoded as `cost_center*3 + status`: COUNT, MIN, and a sparse MAX in which 15 groups read NULL. The expected values were regenerated by `oracle.py`; all 32 existing rows regenerated byte-identical.
- mask-risc: Pair equals a precomputed composite `Lane` under the same executor, and equals `reference_execute`, for all four folds. Test size is n = 1000, spanning more than one tile. A `lo == stride` row that would alias `(1, 0)`'s slot is dropped.
- **Disable runs** (committed first; each went red, then was restored):
- ndarray: deleting the stride guard, transposing `hi`/`lo`, and u32 wrapping arithmetic.
- mask-risc exec: `hi`/`lo` swapped.
- quack: lowering with `stride + 1`.
- reference: the stride guard dropped.
- Correction made along the way: the ndarray stride test was **vacuous** in its first version. Its dropped rows also fell past the group universe, so the universe check dropped them too. It was rewritten so each dropped row would otherwise alias a real slot.

## Open
- AVG and full-range (coalescing) SUM over a Pair key. `masked_group_sum_i32_pair` exists in ndarray, but no mask-risc terminal routes a Pair key to it (`GroupSumI32` / `GroupSumViaI32` take a single lane). Refused by name today.
- More than two key columns. Nothing is built for it; a nested Pair or an N-ary address would be its own decision.
4 changes: 3 additions & 1 deletion .claude/board/entries/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,16 +25,18 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately
opposite directions; the stranding this convention prevents shows up in
exactly one of them, never both.

154 entries, 2026-08-06 .. 2026-09-23.
156 entries, 2026-08-06 .. 2026-09-23.

| date | entry id | finding | file |
|---|---|---|---|
| 2026-09-23 | `ternlog-count-any-fold` | | [2026-09-23-ternlog-count-any-fold.md](2026-09-23-ternlog-count-any-fold.md) |
| 2026-09-23 | `terminal-elects-materialization-range-plane` | | [2026-09-23-terminal-elects-materialization-range-plane.md](2026-09-23-terminal-elects-materialization-range-plane.md) |
| 2026-09-23 | `report-plan-zero-copy-pivot-docir-convergence` | ReportPlan lowers into Quack; pivot = view over shared Arc<CellSpace>; reports become OGAR ObjectSlot sources via grid_of (OGAR #307) | [2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md](2026-09-23-report-plan-zero-copy-pivot-docir-convergence.md) |
| 2026-09-23 | `quack-w-d-multi-key-group-by` | | [2026-09-23-quack-w-d-multi-key-group-by.md](2026-09-23-quack-w-d-multi-key-group-by.md) |
| 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) |
| 2026-09-23 | `program-collapse-boolean-chains` | | [2026-09-23-program-collapse-boolean-chains.md](2026-09-23-program-collapse-boolean-chains.md) |
| 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) |
| 2026-09-23 | `collapse-probe-v4-and-dispatch-split` | | [2026-09-23-collapse-probe-v4-and-dispatch-split.md](2026-09-23-collapse-probe-v4-and-dispatch-split.md) |
| 2026-09-23 | `absolute-execution-extent` | | [2026-09-23-absolute-execution-extent.md](2026-09-23-absolute-execution-extent.md) |
| 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) |
| 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) |
Expand Down
Loading
Loading