Skip to content

Commit 8334bb7

Browse files
committed
simd: group_walk as a (group,row) visitor; add PowerSums power-sum fold
The private keyed-reduction walker no longer takes an i64 sink: it yields (group, row) and each fold owns its destination. The 15 existing folds were migrated mechanically; no public signature changes. New: PowerSums { n: u64, sum: i64, sum_sq: u128 } (#[repr(C)], 32 bytes, layout pinned at compile time) and masked_group_power_sums_i32 with _via / _pair, the exact degree-0/1/2 power sums per group in one pass. A scratch benchmark chose the shapes: the visitor costs nothing for the existing folds, and the record beats three separate lanes at large K. The u128 square sum is load-bearing: a u64-truncation disable run turns the four exactness tests red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ho2JosrXCnZPbB7RFssrse
1 parent b9f8e4e commit 8334bb7

3 files changed

Lines changed: 483 additions & 39 deletions

File tree

‎.claude/blackboard.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,11 @@
1+
## 2026-09-30 — `group_walk` becomes a (group,row) visitor; `PowerSums` + `masked_group_power_sums_i32{,_via,_pair}`
2+
3+
- **Walker:** the private `group_walk` now takes `groups: usize` and `fold: FnMut(group, row)`; it never sees a sink. Its `i64` slot was an accident of the first folds, not part of what it does. All 15 call sites were migrated mechanically (`|k, i| out[k] = …`, passing `out.len()`); no public signature changed.
4+
- **New:** `#[repr(C)] PowerSums { n: u64, sum: i64, sum_sq: u128 }`, the degree-0/1/2 power sums, exact integers, deliberately free of statistical vocabulary (meaning is jc's job upstream). Three kernels share one fold. Compile-time layout pin: size 32, offsets 0/8/16, align == `align_of::<u128>()` (16 on x86_64/aarch64, 8 on some cross targets; pin `align(16)` explicitly if it ever crosses an ABI), plus the arithmetic Σx row bound `(2^32-1)·2^31 ≤ i64::MAX`.
5+
- **Why this shape (scratch benchmark before the change: 72 cells of resident/via/pair keys × mask density 1/50/100 % × K 1…65 536 × realistic/extreme values, median of 15 reps, two runs):** the existing sum through the new visitor ran at 0.994× the old walker; a slot-generic walker compiles to the same 79-instruction stream as the visitor with a record sink, so the visitor dominates at zero cost. A record (AoS) beats three lanes (SoA) by 1.37× (median) at K = 65 536 and up to 2.5× on pair keys: one cache line per row instead of three, and 2 stack reloads in the hot loop instead of 6. Three separate passes cost 2.17× (median). Unexplained, reproducible in both runs: SoA beats AoS by ~35 % on dense resident/via keys at K = 16. Recorded, not chased; the decision does not depend on it.
6+
- **Evidence:** 9 new tests against an independent per-row i128/u128 oracle (resident/via at 7 lengths × sparse/half/dense masks; pair with both drops; an unreached group stays zero; agreement with the count/sum folds; accumulate; an extreme fixture whose exact Σx² > `u64::MAX`, asserted from the oracle first; 2^20 × `i32::MIN` exact). **Disable run** (Σx² truncated to u64 in the fold): exit 101, the 4 exactness tests red, the 5 others green. Even the ordinary randomized fixture overflows u64, so the width is needed on realistic data too. Full `cargo test --lib` 2516 passed; group doctests 17/17; clippy `-D warnings` and fmt clean.
7+
- **Next (not done):** mask-risc `GroupFold::PowerSumsI32` + `Out::PowerSums`, reusing `Value::GroupReduced` (no new Value variant), with an oracle arm and differential tests; only after that is green does R2IL get a byte. No real-crate codegen witness has been run for the walker yet; the no-regression figure comes from the scratch copy.
8+
19
## 2026-09-25 (2) — encryption: Argon2 KDF + envelope behind a default-on `kdf` feature
210

311
- `crates/encryption`: `argon2` is optional; `kdf = ["dep:argon2"]`, `default = ["kdf"]`,

‎src/simd.rs‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -838,6 +838,9 @@ pub use crate::simd_masking_ops::{
838838
masked_group_min_i32,
839839
masked_group_min_i32_pair,
840840
masked_group_min_i32_via,
841+
masked_group_power_sums_i32,
842+
masked_group_power_sums_i32_pair,
843+
masked_group_power_sums_i32_via,
841844
masked_group_sum_i32,
842845
masked_group_sum_i32_pair,
843846
masked_group_sum_i32_via,
@@ -863,6 +866,7 @@ pub use crate::simd_masking_ops::{
863866
ternary_match_u64_to_mask_under,
864867
KeyRunCarry,
865868
MortonDir,
869+
PowerSums,
866870
SYM_EMPTY_I64,
867871
};
868872
// The popcount that closes the loop on the masks above: `mask_count` in ABI

0 commit comments

Comments
 (0)