|
8 | 8 | > anti-pattern the imported board rules name. Backfilled below in one |
9 | 9 | > pass rather than left stale; PR #4 onward gets its entry at merge time. |
10 | 10 |
|
| 11 | +## PR #79 — `hop_cached_vs_gather`: the M1b tile pays off on hop two; the scatter walk is the access-shape question (opened 2026-09-16, merged `9cb63e9`, head `c10029b`) |
| 12 | + |
| 13 | +**Added.** `native/lgj-abi/examples/hop_cached_vs_gather.rs` — `lgj_hop`'s |
| 14 | +selection three ways at 65 536 rows × 32 facets, every arm asserted |
| 15 | +bit-identical on every frontier before timing: `recompute` (the shipped |
| 16 | +body: per facet two contiguous `eq_u32` passes + `ternlog<AND3>` + scatter), |
| 17 | +`cached` (`sel_f = class_f ∧ struct_f` built once per store generation — |
| 18 | +32 × 8 KiB = 256 KiB, the §9 M1b tile keyed `(generation, edge_classid)` — |
| 19 | +then one `mask_and` per facet + scatter), `gather` (the scalar row walk on |
| 20 | +AosRows, the access-shape baseline, not a candidate). `LATEST_STATE.md` |
| 21 | +entry with the full table. No ABI symbol, no Java change. |
| 22 | + |
| 23 | +**Measured** (Xeon 2.10 GHz, `avx512f=true`, release, median of 7, two |
| 24 | +runs, every output `black_box`ed inside its timed closure): cache build |
| 25 | +843–930 µs; cached **14.4–15.0 / 31.8–33.0 / 96–100 / 159–196 / |
| 26 | +903–1 056 µs** across rand7 / rand655 / classid / hop2 / all, against |
| 27 | +recompute 728–1 779 µs — break-even **0.9–1.3 hops** at every frontier. |
| 28 | + |
| 29 | +**Locked.** *A store generation's predicates do not change between hops; |
| 30 | +re-deriving them per hop is 64 contiguous 256 KiB passes for nothing.* The |
| 31 | +scatter is `O(N/64 + selected bits)` per facet (it walks every word of |
| 32 | +`selected` before it can know which are empty), not O(frontier) — |
| 33 | +CodeRabbit's correction, carried into the doc and the board. |
| 34 | + |
| 35 | +**Deferred.** Invalidation cost (a write to any classid/hi32 lane drops all |
| 36 | +32 masks — the registry already does this wholesale for `cached_carving`); |
| 37 | +the per-`edge_classid` multiplication of the 256 KiB; and the operator's |
| 38 | +actual question — inside the random-access scatter walk, decide a visited |
| 39 | +row's 32 facets with one zmm `mask_cmpeq_epi32` per 4 facets (8 loads) or |
| 40 | +one xmm compare per facet, instead of the scalar 64-load / 64-branch loop |
| 41 | +(`gather_xmm` / `gather_zmm_row`, named, not built). |
| 42 | + |
| 43 | +**Review.** Codex P1 (timed closures returned `()`, outputs unread — LLVM |
| 44 | +could drop the scatter stores): fixed, re-measured, every number held. |
| 45 | +CodeRabbit: the O(frontier) claim, fixed; docstring-coverage warning |
| 46 | +(40 %) on private helpers, not addressed. Bugbot: usage limit, no run. |
| 47 | + |
| 48 | +**Confidence.** HIGH on the numbers (two runs, ranges banked). The |
| 49 | +access-granularity arms are unmeasured; the door is named, not opened. |
| 50 | + |
11 | 51 | ## PR #77 — first CI lint gate: fmt + clippy + rust-test, Rust pinned to 1.98.1 (opened 2026-09-05, head `6d4b1a2`) |
12 | 52 |
|
13 | 53 | **Added.** `.github/workflows/lint.yml` (three jobs: `format`, `clippy`, |
|
0 commit comments