Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude/board/STATUS_BOARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ materialization, and that is the ideal Layer-0 operation, not a compromise.
|---|---|---|---|
| D-WFL-AXIS | **Three independent axes, not one binary:** OPERATORS (fold · mask/ternlog · project · rotate · neighbour · reduce) × CARRIERS (canonical lane · range · descriptor · resident mask · cached mask) × MATERIALIZATION CHOICE (fused vs materialized bitmap). The BBB question is not *fold or mask?* but **is this membership relation transient algebra, or has it been PROMOTED to a mask carrier?** | Queued | every plan must record the promotion, not the operator choice. Falsified if a plan can promote a membership relation to a carrier without that appearing in it |
| D-WFL-ENTROPY | **Representation entropy should follow ANSWER entropy.** *"Do these two million-row regions intersect?"* ≈ 1 bit; building 125 KB of mask to find it is the obscenity — and 125 KB is Seam B's MEASURED number at N=1M, not rhetoric. *"How many overlap?"* = 32–64 bits; also fused. *"Give me the overlap, six thoughts will manipulate it"* justifies the bitmap, which is then low-entropy **relative to its future workload** | Queued | the ratio `materialized bytes : answer bytes` reported per operation (§6's `R_info`), with the downstream workload named whenever it exceeds 1 |
| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | Queued | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs |
| D-WFL-W2b‴ | ⊘ supersedes W2b″'s "fold arm vs mask arm" — **both arms mask.** W2b-A: `Range × resident → FUSED masking → Count/Any`, no result mask. W2b-B: `→ masking → a MATERIALIZED bounded mask`. Same masking semantics, different result carrier | In PR — W2b-A shipped (fused `Range ∩ plane → Count/Any`, 0 derived words; `entries/2026-09-23-terminal-elects-materialization-range-plane.md`); W2b-B reuse burden OPEN | differential across arms and against the oracle. **W2b-B carries a burden W2b-A does not: it must NAME and MEASURE the downstream reuse justifying the carrier — a materialization with no demonstrated consumer FAILS the arm.** That is what deliberate promotion costs |

**Shortest form:** fold the datasets, mask the folds, materialize only when the
mask itself is worth keeping.
Expand All @@ -69,7 +69,7 @@ defines what a fold IS — never what the machine may do.
| D-WFL-SIBLING | **The rule governs the TRANSITION, not the bytes:** crossing from fold-native to mask-native execution must be deliberate and visible at the T2 planning membrane. Once MASK is elected, behaving like a mask engine (AND → TERNLOG → shift → cache) is legitimate. Forbidden only: the planner believes it is folding, a helper silently allocates `words_for(N)`, and nobody made the decision | Queued | the BBB question must be answerable for every plan: *who elected the mask, on what basis?* Falsified if a plan can become mask-native without an election appearing in it |
| D-WFL-SEAMB′ | ⊘ **restates Seam B more precisely than every earlier framing** (performance complaint · T1 conformance failure · fold-law violation — all circling this). The defect in `Pred::Range` is NOT that it writes a mask. It is that the planner can neither elect nor decline: there is exactly ONE path, so **the choice does not exist**. Seam B is an ABSENT DECISION, not a present mask | Queued | fixed when both paths exist and the plan records which was taken — not when the mask disappears |
| D-WFL-W2b″ | ⊘ **supersedes D-WFL-W2b′'s "must not write".** W2b demonstrates BOTH legal paths over identical semantics: FOLD-NATIVE (`Range ∩ resident → Count/Any`, no second mask) and MASK-NATIVE (`→ a bounded/cached mask` because a consumer reuses it). Pipeline vs materialize | Queued | the two arms differentially checked against each other AND the oracle — identical row sets, identical Count/Any. The earlier "zero derived buffers, asserted by counter" gate now scopes to the FOLD arm only |
| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | Queued | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking |
| D-WFL-T1-FUSED′ | ⊘ upgrade from optimization to **enabler**: without a fused `popcount(a & b)` over a span there is no intermediate-buffer-free path, so **the fold-native arm does not exist at all**. The primitive CREATES the choice — which is exactly why Seam B had no decision in it | Refuted for `Range ∩ plane` (existing `popcount_batch_u64`/`mask_any` over the borrowed span suffice); plane∩plane `and_popcount` still OPEN | unchanged differential gate vs `mask_and` + `popcount_batch_u64`; the framing change raises its priority from nice-to-have to W2b-blocking |
| D-WFL-ELECT | the election rule, static first: `terminal Count → FOLD`; `one AND then Count → probably FOLD`; `reuse_count > 1 → consider MASK`; `shared cached result → MASK`; `Wabe frontier reused → maybe MASK`; `~11 ns cached mask → almost certainly MASK`. DuckDB-style dynamic costing later | Queued | static rules must be inspectable in the plan. Falsified if the rule set fires the same way on every program (it would carry no information — cf. the can-it-stay-silent twin) |

## D-WFL — the cache scoping (2026-09-19): frozen is fine, marching is the disaster
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# 2026-09-23 — The terminal elects materialization: fused `Range ∩ plane → Count/Any`

**Status:** MEASURED (flagship W2b-A) · OPEN (W2b-B reuse burden; plane∩plane and
compare→Count fusion; execution extent)
**D-ids:** D-WFL-W2b‴ (arm A shipped), D-WFL-T1-FUSED′ (refuted for this shape),
D-WFL-MASKOP / D-WFL-EXPR / D-WFL-SEAMB″ (corrected, still queued), D-WFL-FUSE (confirmed)

## What was actually wrong (re-audited on main 33df0710)
- **Stale:** D-WFL-MASKOP's "`exec.rs:566` forces every slot to `words_for(n_rows)`" and the
plan's "a `Pred::Range` at 1M rows writes 125 KB". Since #1266, scratch is tiled
(`TILE_WORDS = 8`, `tile_words_for`), so a slot is at most 8 words.
- **Still true, and stricter:** `Range ∩ resident plane → Count` wrote derived membership
on every tile. Quack lowers to `Pred{Range, under: gate}`. The executor then ran
`mask_set_range` and `mask_and_assign` into the tile slot, and `Count` read it back with
`popcount_batch_u64`. The work was also not bounded by the touched span: every tile of the
population was walked, however narrow the range.

## What changed
In `crates/lance-graph-mask-risc`:
- `Program::fused_terminal()` recognises `[Pred{Range, under: None | Plane}]` followed by
`Count` or `Any` of that slot. `Program::requires_scratch()` is DERIVED from it, and
`touched_words(lo, hi)` is the one spelling of the span.
- `execute_into` folds a fused program straight from the resident plane's touched words
plus two register-masked edge words. Validation stays total, using a local bookkeeping word.
- `Scratch::for_program` and `over_for_program` carve **zero** slots for such a program.
- `Keep` is never fused: it is the explicit election of a bitmap.

No second evaluator, no new IR and no ndarray change.

## What materialization disappeared, and what remains deliberate
- **Gone:** on the fold arm, derived words written = 0 and scratch slots carved = 0. Gated
by a poisoned caller arena that must stay all-`u64::MAX`; a twin test proves the same
probe sees `Keep`'s carve.
- **Allocation:** `Scratch::for_program` allocates 0 bytes for a fold, counted per thread.
- **Deliberate:** `Keep` still writes the demanded `Out::Mask`.

## D-WFL-T1-FUSED′ refuted for this shape
No new primitive was needed. A range's interior mask is all ones, so the relation is the
plane's own words over the span. The existing `popcount_batch_u64` and `mask_any` are
enough, called on the borrowed slice and on two one-word register temporaries. This
confirms D-WFL-FUSE: it was a lowering rule, not a primitive. The general slice∩slice
`popcount(a & b)` still lacks a buffer-free primitive (see the classification below).

## Measurements
`cargo run --release -p lance-graph-mask-risc --example range_fused_probe`. The two arms
agree on every case (asserted). Median ns:

| N | range | plane | materialized | fused | derived words written, mat / fused | plane words read, mat / fused |
|---|---|---|---|---|---|---|
| 4,096 | tiny | dense | 653 | 54 | 128 / 0 | 64 / 1 |
| 65,536 | 25% | dense | 9,002 | 119 | 2,048 / 0 | 1,024 / 257 |
| 1,048,576 | tiny | dense | 138,620 | 56 | 32,768 / 0 | 16,384 / 1 |
| 1,048,576 | 25% | dense | 141,688 | 1,072 | 32,768 / 0 | 16,384 / 4,097 |
| 1,048,576 | whole | dense | 148,396 | 4,435 | 32,768 / 0 | 16,384 / 16,384 |

- The fused latency is flat in N for a tiny range (54 → 56 ns from 4K to 1M rows) and
scales with the touched span, not the population.
- The word counts are derived from the executor's contract, not from instrumentation:
- Tiled path: two slots written over every tile.
- Fused path: reads exactly `touched_words`.
- The fused-arm zero is the one the gate test enforces.
- Sparse-scattered rows at 1M showed run-to-run noise, up to 297 µs on the materialized arm.

## Falsifiers (`tests/fused_terminal.rs`, disable-verified)
| gate | disable | result |
|---|---|---|
| fold == Keep→popcount/any == scalar oracle, over 4 row counts × 4 plane shapes × the named edges (empty `[65,65)` / `[0,0)`, single row, aligned/unaligned, inside one word, across a word, across tiles, sub-64 tail, whole) | tail edge word not counted | red: `count one n=1024 [60,70)` |
| poisoned arena untouched + zero slots carved | always carve the declared slots | red |
| `for_program` allocates 0 bytes for a fold | same | red |
| fusion recognised at all | `fused_terminal` → `None` | 4 of 7 red |

The allocation gate first flaked (120 stray bytes): the counter was process-wide while the
harness runs tests in parallel. It is now per-thread.

## Classification of the remaining ops (WORKING-MODEL, not measured)
| op → scalar terminal | class |
|---|---|
| `Range`, bare or gated by a resident plane | **fused** (this entry) |
| `Not(a)` → Count | fusion rule: `n − popcount(a)`; no primitive |
| `And` / `AndNot(plane, plane)` → Count | needs a T1 primitive: slice-slice `and_popcount` (ndarray has `U64x8::xor_popcount` as the precedent, but no AND form) |
| `Or` → Count | fusion rule once `and_popcount` exists: `\|a\| + \|b\| − \|a∧b\|` |
| `Xor` → Count | needs a u64-slice XOR popcount (the fused `hamming_distance_raw` exists on bytes only) |
| `Ternlog` → Count | needs `ternlog_popcount`, which generalises the three above |
| lane compares (`Gt`/`Eq`/… `_to_mask_under`) → Count | needs compare-count primitives; today a tile-local write remains |
| `Gather`, `Keep`, `ScatterOr` | deliberate materialization |

The desired direction is FEWER concepts: one `ternlog_popcount` would subsume `and`, `or`,
`xor` and `andnot` counting. Not built here; the flagship needed none of it.

## Still open
- **W2b-B**: the Keep arm is tested, but its downstream-reuse burden (name and measure the
consumer that justifies the carrier) is not met.
- **Execution extent** (PR C): `execute_into` has no ranged entry point.
- **D-WFL-MASKOP**: `MaskOp` still reads as an assignment. Only the terminal-side lowering
moved.
3 changes: 2 additions & 1 deletion .claude/board/entries/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,10 +25,11 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately
opposite directions; the stranding this convention prevents shows up in
exactly one of them, never both.

149 entries, 2026-08-06 .. 2026-09-23.
150 entries, 2026-08-06 .. 2026-09-23.

| date | entry id | finding | file |
|---|---|---|---|
| 2026-09-23 | `terminal-elects-materialization-range-plane` | | [2026-09-23-terminal-elects-materialization-range-plane.md](2026-09-23-terminal-elects-materialization-range-plane.md) |
| 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) |
| 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) |
| 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) |
Expand Down
145 changes: 145 additions & 0 deletions crates/lance-graph-mask-risc/examples/range_fused_probe.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
//! Measure `Range[lo, hi) ∩ resident plane → Count` two ways over the SAME
//! membership relation:
//!
//! - MATERIALIZED: `Pred::Range → Scratch(0)`, `And(Scratch(0), Plane(0)) →
//! Scratch(1)`, `Count(Scratch(1))` — the ops write derived membership one
//! tile at a time before the terminal counts it.
//! - FUSED: `Pred::Range` gated by the plane, `Count` — the executor folds the
//! plane's touched words and two register-masked edge words, writing none.
//!
//! Reported separately, never collapsed into one "speedup": latency (median
//! of repeats), derived membership words written, and plane words read. The
//! word counts are DERIVED from the executor's contract (the tiled path
//! writes every tile of every slot; the fused path reads `touched_words`), and
//! the fused count is cross-checked against the materialized one each run.
//!
//! `cargo run --release -p lance-graph-mask-risc --example range_fused_probe`

use std::time::Instant;

use lance_graph_mask_risc::exec::{execute_into, Scratch};
use lance_graph_mask_risc::{
touched_words, Foreign, MaskOp, Operand, Out, Planes, Pred, Program, Terminal, Value,
};

fn lcg(seed: &mut u64) -> u64 {
*seed = seed
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
*seed >> 11
}

fn plane(n: usize, set: impl Fn(usize) -> bool) -> Vec<u64> {
let mut w = vec![0u64; n.div_ceil(64)];
for r in 0..n {
if set(r) {
w[r / 64] |= 1u64 << (r % 64);
}
}
w
}

fn median_ns(mut f: impl FnMut() -> Value, reps: usize) -> (f64, Value) {
let mut times = Vec::with_capacity(reps);
let mut last = Value::Count(0);
for _ in 0..reps {
let t = Instant::now();
last = std::hint::black_box(f());
times.push(t.elapsed().as_nanos() as f64);
}
times.sort_by(|a, b| a.total_cmp(b));
(times[reps / 2], last)
}

fn main() {
println!(
"{:>8} {:>10} {:>18} {:>12} {:>12} {:>10} {:>10} {:>9} {:>9}",
"N", "range", "plane", "mat_ns", "fused_ns", "mat_wr", "fused_wr", "mat_rd", "fused_rd"
);
let mut seed = 0xfeed_u64;
for n in [4_096usize, 65_536, 1_048_576] {
let scattered: Vec<bool> = (0..n).map(|_| lcg(&mut seed).is_multiple_of(97)).collect();
let shapes: [(&str, Vec<u64>); 3] = [
("dense", plane(n, |r| r % 4 != 0)),
("sparse-clustered", plane(n, |r| (r / 4096) % 50 == 7)),
("sparse-scattered", plane(n, |r| scattered[r])),
];
let nn = n as u32;
let ranges: [(&str, u32, u32); 5] = [
("tiny", nn / 2, nn / 2 + 3),
("one-tile", 1000.min(nn - 600), 1000.min(nn - 600) + 512),
("1%", nn / 3, nn / 3 + nn / 100),
("25%", nn / 5, nn / 5 + nn / 4),
("whole", 0, nn),
];
let words = n.div_ceil(64);
for (pname, p) in &shapes {
let masks: [&[u64]; 1] = [p];
let planes = Planes {
n_rows: n,
masks: &masks,
lanes: &[],
};
for (rname, lo, hi) in ranges {
let materialized = Program::new(
vec![
MaskOp::Pred {
pred: Pred::Range { lo, hi },
under: None,
dst: 0,
},
MaskOp::And {
a: Operand::Scratch(0),
b: Operand::Plane(0),
dst: 1,
},
],
Terminal::Count {
mask: Operand::Scratch(1),
},
);
let fused = Program::new(
vec![MaskOp::Pred {
pred: Pred::Range { lo, hi },
under: Some(Operand::Plane(0)),
dst: 0,
}],
Terminal::Count {
mask: Operand::Scratch(0),
},
);
assert!(materialized.fused_terminal().is_none());
assert!(fused.fused_terminal().is_some());
let mut ms = Scratch::for_program(&materialized, n).expect("scratch");
let mut fs = Scratch::for_program(&fused, n).expect("scratch");
let reps = if n > 100_000 { 31 } else { 201 };
let (mat_ns, mv) = median_ns(
|| {
execute_into(&materialized, &planes, &Foreign::NONE, &mut ms, Out::None)
.expect("materialized")
},
reps,
);
let (fused_ns, fv) = median_ns(
|| {
execute_into(&fused, &planes, &Foreign::NONE, &mut fs, Out::None)
.expect("fused")
},
reps,
);
assert_eq!(mv, fv, "the two arms disagree: {pname} {rname} N={n}");
// Materialized: the range slot and the And slot are each
// written over every tile; the plane is read once over the
// whole population. Fused: nothing written; the plane is read
// over the touched span only.
let (mat_wr, mat_rd) = (2 * words, words);
let fused_rd = touched_words(lo, hi).len();
println!(
"{n:>8} {rname:>10} {pname:>18} {mat_ns:>12.0} {fused_ns:>12.0} \
{mat_wr:>10} {:>10} {mat_rd:>9} {fused_rd:>9}",
0
);
}
}
}
}
Loading
Loading