From b65afda25cddbaf33ef33867344ae0d580e2df89 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 11:05:20 +0000 Subject: [PATCH 1/2] board: CubeCL/LLVM boundary DECISION, regrade of the CubeCL audit, pinned _sym merge law - New entry 2026-09-23-cubecl-llvm-boundary-and-audit-regrade: the fold law, the masking substrate and R2IL are ours and never refactored toward CubeCL or LLVM. CubeCL is lab-only inspiration; LLVM only as a parallel compiler arm fed the same folded program. The chat-only CubeCL audit is regraded section by section; its tier picture (CubeCL IR -> LLVM beside T0, CubeCL types as the seam) is rewritten, and its proposed CubeCL PR is dropped. Its scheduler findings become five laws for our future scheduler. The shader-driver / stockfish-rs NNUE delta-dispatch direction is recorded as OPEN. - TD-SYM-SUM-MERGE-IS-NOT-ADDITION-1 gains the pinned merge law (bottom as identity, the row bound over the TOTAL rows of merged partials, #1263's wrapping law preserved). A law, not code: nothing is built until partial aggregation enters the execution path. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- .claude/board/SUPERSESSION-INDEX.md | 6 +- .claude/board/TECH_DEBT.md | 16 +++ ...-cubecl-llvm-boundary-and-audit-regrade.md | 98 +++++++++++++++++++ .claude/board/entries/README.md | 3 +- 4 files changed, 119 insertions(+), 4 deletions(-) create mode 100644 .claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md diff --git a/.claude/board/SUPERSESSION-INDEX.md b/.claude/board/SUPERSESSION-INDEX.md index 04224f36f..e4199db6b 100644 --- a/.claude/board/SUPERSESSION-INDEX.md +++ b/.claude/board/SUPERSESSION-INDEX.md @@ -102,13 +102,13 @@ a licence to act on it. | **RESCOPE** | `soa-migration-diff-resolution-2026-06-13` | `BindSpace`, `CollapseGateEmission`, `GateDecision`, `MergeMode` … | — | 1/5 | | **RESCOPE** | `cognitive-substrate-convergence-v1` | `BindSpace`, `CollapseGateEmission`, `GateDecision`, `MergeMode` … | PROPOSAL (sprint-10 architectural decisions | 3/13 | | **RESCOPE** | `cognitive-substrate-convergence-v2` | `BindSpace`, `CollapseGateEmission`, `GateDecision`, `MergeMode` … | ACTIVE — sprint-11 Phase A/B COMPLETE (pendi | 4/15 | -| **RESCOPE** | `bindspace-singleton-to-mailbox-soa-v1` | `BindSpace`, `CollapseGateEmission`, `ResonanceDto`, `ThinkingStyle` | CONJECTURE / design (migration spec). NOT ye | 2/19 | +| **RESCOPE** | `bindspace-singleton-to-mailbox-soa-v1` | `BindSpace`, `CollapseGateEmission`, `ResonanceDto`, `ThinkingStyle` | CONJECTURE / design (migration spec). NOT ye | 3/19 | | **RESCOPE** | `callcenter-membrane-v1` | `BindSpace`, `GateDecision`, `MergeMode`, `ThinkingStyle` | Active | 0/0 | -| **RESCOPE** | `causaledge64-mailbox-rename-soa-v1` | `BindSpace`, `GateDecision`, `MergeMode`, `ThinkingStyle` | Active (draft, 2026-05-14) | 0/10 | +| **RESCOPE** | `causaledge64-mailbox-rename-soa-v1` | `BindSpace`, `GateDecision`, `MergeMode`, `ThinkingStyle` | Active (draft, 2026-05-14) | 1/10 | | **RESCOPE** | `integrated-cognitive-planner-v1` | `BindSpace`, `GateDecision`, `ResonanceDto`, `dispatch_busdto` | — | 0/3 | | **RESCOPE** | `palantir-parity-cascade-v2` | `BindSpace`, `MergeMode`, `ResonanceDto`, `ThinkingStyle` | plan, not implementation. | 0/17 | | **RESCOPE** | `temporal-markov-and-style-classes-v1` | `BindSpace`, `MergeMode`, `StepMask`, `ThinkingStyle` | ACTIVE (operator-ratified 2026-07-10: "other | 1/19 | -| **RESCOPE** | `unified-soa-convergence-v1` | `BindSpace`, `CollapseGateEmission`, `ResonanceDto`, `ThinkingStyle` | PROPOSAL / integration plan. Design-spec onl | 3/25 | +| **RESCOPE** | `unified-soa-convergence-v1` | `BindSpace`, `CollapseGateEmission`, `ResonanceDto`, `ThinkingStyle` | PROPOSAL / integration plan. Design-spec onl | 4/25 | | **RESCOPE** | `alpha-reason-witness-shader-field-archaeology-pass-1` | `BindSpace`, `MergeMode`, `ResonanceDto` | SOURCE AUDIT / PLAN ONLY. No production wiri | 0/1 | | **RESCOPE** | `bindspace-mailbox-soa-dependency-map-v1` | `BindSpace`, `dispatch_busdto`, `persist_cycle` | MAP / preflight. No source wired yet. Read-b | 0/2 | | **RESCOPE** | `bindspace-mailbox-soa-w3-w4a-impl-v1` | `BindSpace`, `dispatch_busdto`, `persist_cycle` | v2 — 5-consolidation + 3-brutal-critic pass | 0/1 | diff --git a/.claude/board/TECH_DEBT.md b/.claude/board/TECH_DEBT.md index b4aea9521..2122da041 100644 --- a/.claude/board/TECH_DEBT.md +++ b/.claude/board/TECH_DEBT.md @@ -35,6 +35,22 @@ full-range SUM + COUNT over the same data and would go red on a wrong merge — but only once the merge path runs under it. It does not exist yet, so the test cannot fire on it today. +**Pinned law (2026-09-23, operator) — a law, not code.** With `⊥ = SYM_EMPTY_I64`: + +```text +⊥ ⊕ x = x x ⊕ ⊥ = x ⊥ ⊕ ⊥ = ⊥ +x ⊕ y = x.wrapping_add(y) (x, y ≠ ⊥) +``` + +- It is a commutative monoid with identity `⊥` only while no present value equals `⊥`. So + `GROUP_SUM_SYM_MAX_ROWS` bounds the **total rows across all merged partials**, never each + partial: two in-bound partials can merge into an out-of-bound total. +- It preserves #1263's row-first wrapping law: wrapping addition is associative and + commutative mod 2⁶⁴, so merging present partials equals the row-first fold over all their + rows. A merge must never introduce a non-wrapping step. +- No merge function and no test are written until partial aggregation actually enters the + execution path. See also `.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md`. + Cross-ref: `.claude/board/entries/2026-09-23-quack-having-sym-sum-presence-mask.md` (the `_sym` decision and the presence-mask boundary); ndarray #321; lance-graph #1266. diff --git a/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md b/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md new file mode 100644 index 000000000..a23692c43 --- /dev/null +++ b/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md @@ -0,0 +1,98 @@ +# 2026-09-23 — CubeCL / LLVM boundary, and the regrade of the CubeCL scheduler audit + +**Status:** DECISION (operator) · OPEN (shader-driver delta dispatch; the future scheduler) +**Supersedes:** the chat-only "CubeCL fold-before-steal audit" (never on the board). Its +proposed next step, a CubeCL dispatcher PR, is dropped. + +## DECISION +- **Ours, never refactored toward CubeCL or LLVM:** the T0 fold-before-dispatch law; the + masking substrate (bit planes, ternlog, the keyed-reduction family, `ndarray::simd` as the + only polyfill); R2IL as the IR. A representation we invented is never reshaped to match an + external one, however familiar the external one is. +- **CubeCL is lab-only:** a comparison point for `cognitive-shader-driver`, to inspire shader + and scheduling ideas. The CubeCL lab fork PR stays lab hygiene only. +- **The scheduler may borrow ideas** (rayon vs CubeCL, bare-metal folding, Mississippi Queen + board economics, hexagon) and rebuild them as OUR scheduler, which sits ABOVE the fold. +- **LLVM only as a PARALLEL compiler arm**, fed the same folded program. JIT codegen (compile + per query at run time) and folding into an explicit precompiled polyfill are different + philosophies; neither reshapes the other. + +**SCOPE:** every lowering, scheduler, shader-driver and backend decision in lance-graph and +ndarray. **BASIS:** operator direction, 2026-09-23. **REVISIT WHEN:** a measured result from +the parallel arm exists — and even then any argument for changing the IR, the fold law or the +polyfill goes to the operator, never straight into code. + +## Regrade of the old audit + +| Old section | Verdict | +|---|---| +| §1 #1260 residue (failed fold retires its slots; unobservable today) | KEEP | +| §2 #1261 residue; `GroupLowering::{Folded, Forest}` as the scheduler seam | KEEP. Still only quack's own tests call `lower_group_by_auto` (`crates/lance-graph-quack/src/lib.rs`). | +| §3 CubeCL CPU scheduler mechanics | KEEP as lab notes; reframed as lessons below | +| §4 8→64 overflow workers, suspected cross-stream deadlock | REGRADE to anti-patterns; fixing a lab reference is not our work | +| §5 tier split drawing "T0.5 CubeCL IR → LLVM O3" beside T0, with CubeCL's kernel types as the seam | **REWRITE** (below) — this was the boundary violation | +| §6 invariant ownership | lance-graph half KEEP; the CubeCL half becomes laws for OUR scheduler | +| §7 four falsifiers written in CubeCL | RE-TARGET at our scheduler, when it exists | +| "Next PR: CubeCL" | DROP | + +## Corrected tier picture + +``` +T1 semantics ─fold (T2)─► folded R2IL program ─► OUR dispatch unit ─► OUR scheduler + │ (program + requirements + extent + affinity) + └─► [parallel lab arm] LLVM/CubeCL compile of the SAME + folded program → result + timing, compared only +``` +T0 realization is `mask-risc` / `ndarray::simd` — one path, not a fork. The parallel arm +answers only: do results agree (an oracle); what are JIT compile+run vs fold+dispatch +latencies (never measured here); where does fused codegen beat the fold (a finding for us +to answer our own way). + +## Lessons for OUR scheduler +From reading CubeCL's CPU runtime (not built or run here; its CPU backend needs native LLVM): +1. **Extent is a first-class field of the dispatch unit.** CubeCL buries its stealable unit + (a cube range) inside compiled code with no ABI slot for it. mask-risc already exposes it: + the unit is `(program, tile range)`. +2. **A waiting task must not stall its worker.** CubeCL's 4-slot parked queue does exactly + that, and is the likely route to the suspected cross-stream deadlock. +3. **Special pools stay isolated.** CubeCL's barrier-overflow workers leak into ordinary work + permanently. +4. **Split by extent, merge by the fold's own algebra** — see + `TD-SYM-SUM-MERGE-IS-NOT-ADDITION-1` for the `_sym` SUM merge law. + +rayon splits recursively over a range (extent-native); CubeCL launches a fixed geometry. Our +fold turns most work into one small program over K slots, so the default is **inline**; +splitting is for the `Forest` residue or very large tile ranges only. + +Mississippi Queen / hexagon stays **[H]**: `.claude/plans/spog-alpha-channel-v1.md` §4 already +says *topology chooses neighborhood, masks choose admissibility, BLAS chooses magnitude*. For +scheduling that suggests an MQ-style budget pricing which tiles to dispatch next. Not +buildable: the MQ source docs are still missing (`.claude/board/LATEST_STATE.md`, "no MQ docs"). + +**Laws the future scheduler must carry:** exactly-once; barrier affinity; pool isolation; +blocked-doesn't-block; extent-merge-by-fold-algebra. The scheduler itself remains gated on +OQ-5 (rayon vendor vs `std::thread::scope`, `.claude/plans/causaledge64-mailbox-rename-soa-v1.md`, +`STATUS_BOARD.md` row D-CE64-MB-1-impl). + +## The residue is already native (operator framing) +The HAVING work shows the architecture already doing this "T0.5-ish" job in its own +vocabulary. The residue is not a task object or a generic IR. It is **group slot + reserved +empty code + exact fold kernel + presence normalization at exit**: tiny, typed, algebraic, +with materialization delayed to the boundary. The twin path (`_sym` SUM vs full-range SUM + +COUNT) is a built-in semantic witness, not debt: it catches `_sym` degrading to an ordinary +sum, and a genuine `i64::MIN` sum being mistaken for emptiness. + +## `cognitive-shader-driver` and the NNUE reference — OPEN +- The driver has no scheduler: one synchronous `run` per cycle + (`crates/cognitive-shader-driver/src/driver.rs`), recomputing its pre-pass and cascade on + every call. +- The NNUE lesson is the delta. stockfish-rs `HalfKaAccumulator::apply_move` updates 2 + columns on a quiet move vs 32 on a refresh, byte-exact by wrapping + (`stockfish-rs/src/eval/incremental.rs`, tests + `incremental_halfka_matches_full_refresh_every_ply` and + `incremental_column_count_is_bounded`). +- Caveat, measured: stockfish-rs's own search does not use it yet — `evaluate()` calls + `Accumulator::refresh` every time (`src/eval/network_eval.rs`), and `search.rs` calls that + `evaluate`. The reference is proven in isolation, not wired. +- Direction: **delta-driven dispatch** — touch only rows whose inputs changed; the fold law + applied along time. CubeCL contributes nothing here. No code change now. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 7408a85aa..c0579dff9 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,11 +25,12 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -148 entries, 2026-08-06 .. 2026-09-23. +149 entries, 2026-08-06 .. 2026-09-23. | date | entry id | finding | file | |---|---|---|---| | 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) | +| 2026-09-23 | `cubecl-llvm-boundary-and-audit-regrade` | | [2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md](2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md) | | 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) | | 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) | | 2026-09-22 | `E-CATS-FOLD-DOES-NOT-RETAIN-A-POPULATION-BITMAP-1` | CATS aggregate lowers to one tiled grouped terminal; bitmap realization is a requested boundary sink | [2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md](2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md) | From e8e797e455cc94562d6d850d2b1bd669432d195e Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 23 Sep 2026 11:58:58 +0000 Subject: [PATCH 2/2] board: correct three overclaims in the CubeCL/LLVM boundary entry (#1267 review) - Tier diagram: both arms receive the mask-risc Program that GroupLowering::Folded carries, not the upstream R2IL; an R2IL-consuming arm needs its own equivalence-tested R2IL->Program boundary. - Scheduler lesson 1: mask-risc has no ranged entry point yet (execute_into walks whole Planes 0..n_rows); marked REQUIRED WORK. - Twin witness: it checks the presence invariant, it does not exercise a present i64::MIN sum. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG --- ...-cubecl-llvm-boundary-and-audit-regrade.md | 24 ++++++++++++------- 1 file changed, 16 insertions(+), 8 deletions(-) diff --git a/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md b/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md index a23692c43..91c8c0455 100644 --- a/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md +++ b/.claude/board/entries/2026-09-23-cubecl-llvm-boundary-and-audit-regrade.md @@ -38,11 +38,15 @@ polyfill goes to the operator, never straight into code. ## Corrected tier picture ``` -T1 semantics ─fold (T2)─► folded R2IL program ─► OUR dispatch unit ─► OUR scheduler - │ (program + requirements + extent + affinity) - └─► [parallel lab arm] LLVM/CubeCL compile of the SAME - folded program → result + timing, compared only +T1 semantics ─fold (T2)─► R2IL/loco ─lower─► mask-risc Program ─► OUR dispatch unit ─► OUR scheduler + │ (program + requirements + extent + affinity) + └─► [parallel lab arm] LLVM/CubeCL compile of the SAME + mask-risc Program → result + timing, compared only ``` +The artifact both arms receive is the `lance_graph_mask_risc::Program` that +`GroupLowering::Folded` carries — not the upstream R2IL. An arm that ever consumes R2IL +instead first needs its own R2IL→Program boundary, equivalence-tested against the native +lowering; otherwise it is not a same-program oracle. T0 realization is `mask-risc` / `ndarray::simd` — one path, not a fork. The parallel arm answers only: do results agree (an oracle); what are JIT compile+run vs fold+dispatch latencies (never measured here); where does fused codegen beat the fold (a finding for us @@ -50,9 +54,12 @@ to answer our own way). ## Lessons for OUR scheduler From reading CubeCL's CPU runtime (not built or run here; its CPU backend needs native LLVM): -1. **Extent is a first-class field of the dispatch unit.** CubeCL buries its stealable unit - (a cube range) inside compiled code with no ABI slot for it. mask-risc already exposes it: - the unit is `(program, tile range)`. +1. **Extent must become a first-class field of the dispatch unit.** CubeCL buries its + stealable unit (a cube range) inside compiled code with no ABI slot for it. mask-risc has + none yet either: `execute_into` takes a whole `Planes` and walks every word from 0 to + `n_rows`. **REQUIRED WORK before any split:** a ranged entry point over + `(program, tile range)`, including rebasing every mask and lane for ranges that do not + start on a word boundary. 2. **A waiting task must not stall its worker.** CubeCL's 4-slot parked queue does exactly that, and is the likely route to the suspected cross-stream deadlock. 3. **Special pools stay isolated.** CubeCL's barrier-overflow workers leak into ordinary work @@ -80,7 +87,8 @@ vocabulary. The residue is not a task object or a generic IR. It is **group slot empty code + exact fold kernel + presence normalization at exit**: tiny, typed, algebraic, with materialization delayed to the boundary. The twin path (`_sym` SUM vs full-range SUM + COUNT) is a built-in semantic witness, not debt: it catches `_sym` degrading to an ordinary -sum, and a genuine `i64::MIN` sum being mistaken for emptiness. +sum and checks the presence invariant that would expose a genuine `i64::MIN` sum being +mistaken for emptiness. ## `cognitive-shader-driver` and the NNUE reference — OPEN - The driver has no scheduler: one synchronous `run` per cycle