Skip to content

mask-risc: collapse ≤3-leaf Boolean op chains onto the ternlog Count/Any fold - #1272

Merged
AdaWorldAPI merged 4 commits into
mainfrom
claude/fold-distillation-pr-wave-s57uj7
Sep 23, 2026
Merged

AdaWorldAPI merged 4 commits into
mainfrom
claude/fold-distillation-pr-wave-s57uj7

Conversation

@AdaWorldAPI

Copy link
Copy Markdown
Owner

Summary

Before this PR, the Count/Any fold from #1270 only handled a single Boolean op. Now any raw op chain over at most three resident planes gets the same treatment: Program::fused_ternlog() interprets the whole chain symbolically.

  • How it works: each plane is a leaf truth table (0xF0 / 0xCC / 0xAA). Each op combines tables using its own Boolean law; Ternlog applies its immediate bitwise to the three input tables. Each op writes the resulting table to its destination slot, so slot reuse is modelled exactly.
  • What the fold receives: the terminal slot's table becomes the fold immediate.
  • No allocation: recognition uses a fixed on-stack array of FUSED_SLOT_CAP slots, the same bound the validation bitmap uses.

Stays on the tiled path, unchanged:

  • a fourth distinct plane;
  • a Pred (comparison leaf) or Gather anywhere in the chain;
  • a scratch slot read before it is written;
  • a slot at or above the cap;
  • any terminal other than Count/Any (Keep remains the explicit election of a bitmap).

Scope: lane predicates are deliberately not attempted. Chains with more than three leaves are measured at their boundary only. No Rayon, OQ-5 or scheduler changes.

Gates

  • End to end: tests/program_collapse.rs checks that fuse_program output reaches the fold.
  • Differential: raw chains are checked against the tiled Keep path across extents, and against reference_execute over the whole population. The chains are a catalogue of 7 (slot reuse in place, a derived-input ternlog, a dead op, double complement, a & !a) plus 400 random chains, with an anti-vacuity check of at least 20 distinct counts.
  • Zero materialization: a poisoned-scratch gate plus a counting allocator, each with a Keep twin as a control.
  • Recogniser test updated: in tests/fused_ternlog.rs it now admits exactly the collapsible chains. Two exec.rs fixtures that needed an unfused program now use Terminal::All, because Not and chains fold now.
  • Disable runs: committed first, confirmed red, then restored to green. Each of these fails inside program_collapse.rs:
    • an interpreter limited to single ops;
    • a transposed table-apply;
    • slot reuse ignored;
    • a fourth leaf aliased to leaf 0.
  • Checks: full lance-graph-mask-risc and lance-graph-quack suites pass; clippy -D warnings and fmt are clean.

Measurement

cargo run --release -p lance-graph-mask-risc --example program_collapse_probe. Setup:

  • Tier: x86-64-v3 (the AVX2 polyfill, the repo default). v4 was not measured.
  • Size: N = 1M rows.
  • Values: median, oracle-checked.
chain ops fold Count tiled Count derived words written by the tiled path
(a&b)|!c 3 8.6 µs 224 µs 49 152
((a^b)&!c)|(a&c) 4 16.2 µs 352 µs 65 536
maj(a,b,c)^a 2 17.4 µs 166 µs 32 768
a&d, empty by data (Any: 4.4 vs 111 µs) 1 6.5 µs 120 µs 16 384
(a&b)|(c&d), 4 planes 3 not collapsed 283 µs 49 152

How to read the numbers:

  • The fold writes 0 derived words.
  • Any on non-empty data: the gap is dominated by the fold exiting at its first block. The data-empty row isolates the scan itself: 25× for Any, 18× for Count.
  • The gap is not all write elimination: the tiled path also pays per-tile interpreter dispatch, and this probe does not separate the two costs.
  • An algebraically constant chain (e.g. a&b&!a → table 0x00) monomorphizes to a fold that never reads the planes. That is why the empty row is empty by data.

Board entry: .claude/board/entries/2026-09-23-program-collapse-boolean-chains.md

🤖 Generated with Claude Code

https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG

…Any fold

fused_ternlog() now interprets the whole op sequence symbolically as
8-bit truth tables over at most three resident planes; the terminal
slot's table becomes the fold immediate. Chains, slot reuse, Not, and
Ternlog over derived inputs collapse; a fourth plane, Pred/Gather, or a
read-before-write scratch stay on the tiled path.

Differential tests against the Keep path and reference_execute across
extents; poisoned-scratch + zero-allocation gate for collapsed chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
…entry

program_collapse_probe: fold vs tiled vs Keep at N=1M (x86-64-v3), incl.
a data-empty chain that isolates the scan and a 4-plane chain that marks
where the collapse stops.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: 44ebbf2f-962d-4cad-9ced-74d7ee005f12

📥 Commits

Reviewing files that changed from the base of the PR and between d90600d and c47698f.

📒 Files selected for processing (7)
  • .claude/board/entries/2026-09-23-program-collapse-boolean-chains.md
  • .claude/board/entries/README.md
  • crates/lance-graph-mask-risc/examples/program_collapse_probe.rs
  • crates/lance-graph-mask-risc/src/exec.rs
  • crates/lance-graph-mask-risc/src/ir.rs
  • crates/lance-graph-mask-risc/tests/fused_ternlog.rs
  • crates/lance-graph-mask-risc/tests/program_collapse.rs
 _________________________________________________
< Because, even your code needs a second opinion. >
 -------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_a21e4a07-c99c-44da-b2a1-4675707486e6)

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review September 23, 2026 18:27
@AdaWorldAPI
AdaWorldAPI merged commit 9b6acb5 into main Sep 23, 2026
9 of 10 checks passed
AdaWorldAPI pushed a commit that referenced this pull request Sep 23, 2026
Nine STATUS_BOARD rows still read "In PR" or "Queued" for work that has
merged. Each flip was checked against origin/main before editing:

- D-WFL-FUSE      Shipped: single-op #1270, multi-op chains ≤3 planes #1272
- D-WFL-W2b‴      Shipped W2b-A #1268; W2b-B reuse burden stays OPEN
- D-WFL-EXTENT    Shipped #1269 (`execute_extent`)
- D-WFL-T1-FUSED, T1-FUSED′  Shipped (ndarray #322 + #1270)
- D-WFL-L0        Shipped #1251
- D-WFL-W2a       Shipped #1268: un-gated Pred::Range, Count = hi−lo,
                  Any = lo<hi, zero scratch (`run_fused`)
- D-WFL-W2a′      Shipped #1268: `requires_scratch()` derived from the
                  fused lowerings, never a caller flag
- D-WFL-2         Partially shipped: range Count/Any landed as a lowering,
                  the `BoundedMask` window is not built

Status cells only; every row's description is unchanged, and the old
status is kept after "was:" where it carried detail. The row count stays
at 2203 lines.

Also corrects one exec.rs comment that still described the Boolean fold
as "a single 2/3-input op"; since #1272 it collapses whole chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
AdaWorldAPI pushed a commit that referenced this pull request Sep 23, 2026
…h; v3 and v4 measured

program_collapse_probe gains a fourth arm. BULK evaluates each op once over
the whole touched span, into preallocated buffers, with the same
ndarray::simd kernels the executor calls per tile. It writes the same
derived words as the tiled path but makes one facade call per op, not one
per op per tile. Two derived columns: wr_ns = bulk - fold (the writes) and
disp_ns = tiled - bulk (per-tile overhead). Every arm is asserted equal to
the bit-serial oracle.

Measured at N = 1M, whole population:
- disp_ns is 90-94 % of tiled_ns at both v3 and v4 (~30 ns per op per
  8-word tile). The writes cost at most ~19 us. Most of #1272's 10-26x is
  skipping the tile interpreter; write elimination is the smaller share.
- At v4 every collapsible 3-plane Count folds in ~6.2 us regardless of its
  truth table (v3: 8.6-17.3 us). The tiled path does not move between
  tiers.
- At 1 % extents BULK beats the fold, which pays a fixed per-call cost
  (validate + re-running the symbolic recognizer).

Entry: .claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md.
Tile size is deliberately NOT changed; it is bound to the scratch-size
contract, and the entry records it as OPEN.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants