Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .claude/blackboard.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,11 @@
## 2026-09-30 — `group_walk` becomes a (group,row) visitor; `PowerSums` + `masked_group_power_sums_i32{,_via,_pair}`

- **Walker:** the private `group_walk` now takes `groups: usize` and `fold: FnMut(group, row)`; it never sees a sink. Its `i64` slot was an accident of the first folds, not part of what it does. All 15 call sites were migrated mechanically (`|k, i| out[k] = …`, passing `out.len()`); no public signature changed.
- **New:** `#[repr(C)] PowerSums { n: u64, sum: i64, sum_sq: u128 }`, the degree-0/1/2 power sums, exact integers, deliberately free of statistical vocabulary (meaning is jc's job upstream). Three kernels share one fold. Compile-time layout pin: size 32, offsets 0/8/16, align == `align_of::<u128>()` (16 on x86_64/aarch64, 8 on some cross targets; pin `align(16)` explicitly if it ever crosses an ABI), plus the arithmetic Σx row bound `(2^32-1)·2^31 ≤ i64::MAX`.
- **Why this shape (scratch benchmark before the change: 72 cells of resident/via/pair keys × mask density 1/50/100 % × K 1…65 536 × realistic/extreme values, median of 15 reps, two runs):** the existing sum through the new visitor ran at 0.994× the old walker; a slot-generic walker compiles to the same 79-instruction stream as the visitor with a record sink, so the visitor dominates at zero cost. A record (AoS) beats three lanes (SoA) by 1.37× (median) at K = 65 536 and up to 2.5× on pair keys: one cache line per row instead of three, and 2 stack reloads in the hot loop instead of 6. Three separate passes cost 2.17× (median). Unexplained, reproducible in both runs: SoA beats AoS by ~35 % on dense resident/via keys at K = 16. Recorded, not chased; the decision does not depend on it.
- **Evidence:** 9 new tests against an independent per-row i128/u128 oracle (resident/via at 7 lengths × sparse/half/dense masks; pair with both drops; an unreached group stays zero; agreement with the count/sum folds; accumulate; an extreme fixture whose exact Σx² > `u64::MAX`, asserted from the oracle first; 2^20 × `i32::MIN` exact). **Disable run** (Σx² truncated to u64 in the fold): exit 101, the 4 exactness tests red, the 5 others green. Even the ordinary randomized fixture overflows u64, so the width is needed on realistic data too. Full `cargo test --lib` 2516 passed; group doctests 17/17; clippy `-D warnings` and fmt clean.
- **Next (not done):** mask-risc `GroupFold::PowerSumsI32` + `Out::PowerSums`, reusing `Value::GroupReduced` (no new Value variant), with an oracle arm and differential tests; only after that is green does R2IL get a byte. No real-crate codegen witness has been run for the walker yet; the no-regression figure comes from the scratch copy.

## 2026-09-25 (2) — encryption: Argon2 KDF + envelope behind a default-on `kdf` feature

- `crates/encryption`: `argon2` is optional; `kdf = ["dep:argon2"]`, `default = ["kdf"]`,
Expand Down
4 changes: 4 additions & 0 deletions src/simd.rs
Original file line number Diff line number Diff line change
Expand Up @@ -838,6 +838,9 @@ pub use crate::simd_masking_ops::{
masked_group_min_i32,
masked_group_min_i32_pair,
masked_group_min_i32_via,
masked_group_power_sums_i32,
masked_group_power_sums_i32_pair,
masked_group_power_sums_i32_via,
masked_group_sum_i32,
masked_group_sum_i32_pair,
masked_group_sum_i32_via,
Expand All @@ -863,6 +866,7 @@ pub use crate::simd_masking_ops::{
ternary_match_u64_to_mask_under,
KeyRunCarry,
MortonDir,
PowerSums,
SYM_EMPTY_I64,
};
// The popcount that closes the loop on the masks above: `mask_count` in ABI
Expand Down
Loading
Loading