mask-risc: collapse ≤3-leaf Boolean op chains onto the ternlog Count/Any fold - #1272
Merged
Merged
Conversation
…Any fold fused_ternlog() now interprets the whole op sequence symbolically as 8-bit truth tables over at most three resident planes; the terminal slot's table becomes the fold immediate. Chains, slot reuse, Not, and Ternlog over derived inputs collapse; a fourth plane, Pred/Gather, or a read-before-write scratch stay on the tiled path. Differential tests against the Keep path and reference_execute across extents; poisoned-scratch + zero-allocation gate for collapsed chains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
…entry program_collapse_probe: fold vs tiled vs Keep at N=1M (x86-64-v3), incl. a data-empty chain that isolates the scan and a 4-plane chain that marks where the collapse stops. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches📝 Generate docstrings
Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_a21e4a07-c99c-44da-b2a1-4675707486e6) |
AdaWorldAPI
marked this pull request as ready for review
September 23, 2026 18:27
…ion-pr-wave-s57uj7
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
AdaWorldAPI
pushed a commit
that referenced
this pull request
Sep 23, 2026
Nine STATUS_BOARD rows still read "In PR" or "Queued" for work that has merged. Each flip was checked against origin/main before editing: - D-WFL-FUSE Shipped: single-op #1270, multi-op chains ≤3 planes #1272 - D-WFL-W2b‴ Shipped W2b-A #1268; W2b-B reuse burden stays OPEN - D-WFL-EXTENT Shipped #1269 (`execute_extent`) - D-WFL-T1-FUSED, T1-FUSED′ Shipped (ndarray #322 + #1270) - D-WFL-L0 Shipped #1251 - D-WFL-W2a Shipped #1268: un-gated Pred::Range, Count = hi−lo, Any = lo<hi, zero scratch (`run_fused`) - D-WFL-W2a′ Shipped #1268: `requires_scratch()` derived from the fused lowerings, never a caller flag - D-WFL-2 Partially shipped: range Count/Any landed as a lowering, the `BoundedMask` window is not built Status cells only; every row's description is unchanged, and the old status is kept after "was:" where it carried detail. The row count stays at 2203 lines. Also corrects one exec.rs comment that still described the Boolean fold as "a single 2/3-input op"; since #1272 it collapses whole chains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
AdaWorldAPI
pushed a commit
that referenced
this pull request
Sep 23, 2026
…h; v3 and v4 measured program_collapse_probe gains a fourth arm. BULK evaluates each op once over the whole touched span, into preallocated buffers, with the same ndarray::simd kernels the executor calls per tile. It writes the same derived words as the tiled path but makes one facade call per op, not one per op per tile. Two derived columns: wr_ns = bulk - fold (the writes) and disp_ns = tiled - bulk (per-tile overhead). Every arm is asserted equal to the bit-serial oracle. Measured at N = 1M, whole population: - disp_ns is 90-94 % of tiled_ns at both v3 and v4 (~30 ns per op per 8-word tile). The writes cost at most ~19 us. Most of #1272's 10-26x is skipping the tile interpreter; write elimination is the smaller share. - At v4 every collapsible 3-plane Count folds in ~6.2 us regardless of its truth table (v3: 8.6-17.3 us). The tiled path does not move between tiers. - At 1 % extents BULK beats the fold, which pays a fixed per-call cost (validate + re-running the symbolic recognizer). Entry: .claude/board/entries/2026-09-23-collapse-probe-v4-and-dispatch-split.md. Tile size is deliberately NOT changed; it is bound to the scratch-size contract, and the entry records it as OPEN. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Before this PR, the Count/Any fold from #1270 only handled a single Boolean op. Now any raw op chain over at most three resident planes gets the same treatment:
Program::fused_ternlog()interprets the whole chain symbolically.0xF0/0xCC/0xAA). Each op combines tables using its own Boolean law;Ternlogapplies its immediate bitwise to the three input tables. Each op writes the resulting table to its destination slot, so slot reuse is modelled exactly.FUSED_SLOT_CAPslots, the same bound the validation bitmap uses.Stays on the tiled path, unchanged:
Pred(comparison leaf) orGatheranywhere in the chain;Keepremains the explicit election of a bitmap).Scope: lane predicates are deliberately not attempted. Chains with more than three leaves are measured at their boundary only. No Rayon, OQ-5 or scheduler changes.
Gates
tests/program_collapse.rschecks thatfuse_programoutput reaches the fold.Keeppath across extents, and againstreference_executeover the whole population. The chains are a catalogue of 7 (slot reuse in place, a derived-input ternlog, a dead op, double complement,a & !a) plus 400 random chains, with an anti-vacuity check of at least 20 distinct counts.Keeptwin as a control.tests/fused_ternlog.rsit now admits exactly the collapsible chains. Twoexec.rsfixtures that needed an unfused program now useTerminal::All, becauseNotand chains fold now.program_collapse.rs:lance-graph-mask-riscandlance-graph-quacksuites pass; clippy-D warningsand fmt are clean.Measurement
cargo run --release -p lance-graph-mask-risc --example program_collapse_probe. Setup:x86-64-v3(the AVX2 polyfill, the repo default). v4 was not measured.(a&b)|!c((a^b)&!c)|(a&c)maj(a,b,c)^aa&d, empty by data (Any: 4.4 vs 111 µs)(a&b)|(c&d), 4 planesHow to read the numbers:
a&b&!a→ table0x00) monomorphizes to a fold that never reads the planes. That is why the empty row is empty by data.Board entry:
.claude/board/entries/2026-09-23-program-collapse-boolean-chains.md🤖 Generated with Claude Code
https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG