Skip to content

Commit db10f03

Browse files
committed
board: PR #79 arc entry (post-merge sha 9cb63e9)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DCfrD5y19cvFc4AoyydXYv
1 parent 9cb63e9 commit db10f03

1 file changed

Lines changed: 40 additions & 0 deletions

File tree

‎.claude/board/PR_ARC_INVENTORY.md‎

Lines changed: 40 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,46 @@
88
> anti-pattern the imported board rules name. Backfilled below in one
99
> pass rather than left stale; PR #4 onward gets its entry at merge time.
1010
11+
## PR #79 — `hop_cached_vs_gather`: the M1b tile pays off on hop two; the scatter walk is the access-shape question (opened 2026-09-16, merged `9cb63e9`, head `c10029b`)
12+
13+
**Added.** `native/lgj-abi/examples/hop_cached_vs_gather.rs` — `lgj_hop`'s
14+
selection three ways at 65 536 rows × 32 facets, every arm asserted
15+
bit-identical on every frontier before timing: `recompute` (the shipped
16+
body: per facet two contiguous `eq_u32` passes + `ternlog<AND3>` + scatter),
17+
`cached` (`sel_f = class_f ∧ struct_f` built once per store generation —
18+
32 × 8 KiB = 256 KiB, the §9 M1b tile keyed `(generation, edge_classid)` —
19+
then one `mask_and` per facet + scatter), `gather` (the scalar row walk on
20+
AosRows, the access-shape baseline, not a candidate). `LATEST_STATE.md`
21+
entry with the full table. No ABI symbol, no Java change.
22+
23+
**Measured** (Xeon 2.10 GHz, `avx512f=true`, release, median of 7, two
24+
runs, every output `black_box`ed inside its timed closure): cache build
25+
843–930 µs; cached **14.4–15.0 / 31.8–33.0 / 96–100 / 159–196 /
26+
903–1 056 µs** across rand7 / rand655 / classid / hop2 / all, against
27+
recompute 728–1 779 µs — break-even **0.9–1.3 hops** at every frontier.
28+
29+
**Locked.** *A store generation's predicates do not change between hops;
30+
re-deriving them per hop is 64 contiguous 256 KiB passes for nothing.* The
31+
scatter is `O(N/64 + selected bits)` per facet (it walks every word of
32+
`selected` before it can know which are empty), not O(frontier) —
33+
CodeRabbit's correction, carried into the doc and the board.
34+
35+
**Deferred.** Invalidation cost (a write to any classid/hi32 lane drops all
36+
32 masks — the registry already does this wholesale for `cached_carving`);
37+
the per-`edge_classid` multiplication of the 256 KiB; and the operator's
38+
actual question — inside the random-access scatter walk, decide a visited
39+
row's 32 facets with one zmm `mask_cmpeq_epi32` per 4 facets (8 loads) or
40+
one xmm compare per facet, instead of the scalar 64-load / 64-branch loop
41+
(`gather_xmm` / `gather_zmm_row`, named, not built).
42+
43+
**Review.** Codex P1 (timed closures returned `()`, outputs unread — LLVM
44+
could drop the scatter stores): fixed, re-measured, every number held.
45+
CodeRabbit: the O(frontier) claim, fixed; docstring-coverage warning
46+
(40 %) on private helpers, not addressed. Bugbot: usage limit, no run.
47+
48+
**Confidence.** HIGH on the numbers (two runs, ranges banked). The
49+
access-granularity arms are unmeasured; the door is named, not opened.
50+
1151
## PR #77 — first CI lint gate: fmt + clippy + rust-test, Rust pinned to 1.98.1 (opened 2026-09-05, head `6d4b1a2`)
1252

1353
**Added.** `.github/workflows/lint.yml` (three jobs: `format`, `clippy`,

0 commit comments

Comments
 (0)