bitwise/statistics: U64x8 popcount/Hamming, early-exit Hamming, batch Welford for the cascade - #323
Conversation
…64x8 popcount_batch_u64 carried a comment promising "POPCNT if available" over a plain count_ones loop with no dispatch. hamming_avx2 was compiled with target_feature(avx2) but contained no AVX2 instruction (its own comment: "No U8x32 polyfill available"). Both now run through the U64x8 polyfill, whose popcnt() every realization lowers natively (VPOPCNTQ with avx512vpopcntdq, vcntq_u8 on NEON, i8x16_popcnt on wasm simd128, per-lane count_ones on the AVX2 and scalar builds). hamming_avx2 is replaced by the safe, ungated hamming_u64x8, which dispatch_hamming now uses as its fallback on every architecture instead of hamming_scalar; hamming_scalar stays as the cfg(test) oracle. Tests: new reference-parity tests for both functions (every tail length across the 8-word / 64-byte chunk boundary, unaligned and unequal-length slices, all-ones worst case); the AVX2-gated tier tests now exercise hamming_u64x8 unconditionally. Note: popcount_batch_u64 is the baseline side of examples/ternlog_fold_probe, so numbers recorded from that probe before this commit are not comparable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (4)
Included review availability: 1 review is currently available. Your included PR review attempts over the past 7 days set your current allowance at 3 reviews per hour. 📝 WalkthroughWalkthroughHamming dispatch now uses ChangesHamming distance checks
Estimated code review effort: 3 (Moderate) | ~20 minutes Suggested reviewers: Merge Risk: ⚪ Minimal · up to No actionable issue is identified in the bounded Hamming and cascade changes; the PR is mergeable after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
A rabbit counts bits by the byte, Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_3a56f8c9-5881-4c26-8ff9-c35e7ab9cf54) |
New `hamming_distance_within(a, b, max) -> Option<u64>`: `Some(d)` iff the distance d <= max, else None. The running total only grows, so once it passes `max` the rest is skipped; the check runs after every 256-byte block, each counted by the same runtime-dispatched kernel hamming_distance_raw uses (VPOPCNTDQ / AVX-512BW / U64x8 polyfill), so no block loses its best lowering. Re-exported from ndarray::simd beside hamming_distance_raw. Wired into the Belichtungsmesser (hpc/cascade.rs) at every EXACT comparison: Cascade::test, the small-vector path, Stroke 2 of query (budget = threshold - prefix distance; a prefix already over it rejects without a scan) and PackedDatabase Stroke 3. The Stroke 1 estimate comparisons (scaled f64) are left untouched: an integer budget there would need a rounding argument, and those loops only read the short prefix anyway. Results are identical by construction (d_prefix + d_rest <= T <=> d_prefix <= T and d_rest <= T - d_prefix). Tests: exactness at max = d, d-1, d+1, 0, u64::MAX across block boundaries and unequal lengths, and a first-block reject. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Stroke-1 survivors with full distances straddling the threshold (8 prefix bits + k tail bits); the exact strokes must return exactly 8 + k <= T. A budget of T instead of T - prefix would admit three extra candidates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
…dits
Appends sections N-S to agnostic-surface-cpu-matrix.md, built from
read-only per-function audits (file:line cited throughout):
- N masking ops per realization (compare/algebra/ternlog/fold, masked
reductions and group-by — 19 scalar by design, with reasons), gaps
and parity coverage; post-#323 rows for popcount/Hamming
- O crypto lane (U32x16/U64x8 ARX per backend, in-tree consumers, gaps)
- P core::simd counterpart of every simd_nightly method
(native / composed / scalar-fallback / missing) and the gaps that
are upstream candidates for portable-simd
- Q staleness corrections to sections A-L (runtime-dispatch shipped,
wasm backend exists, CI matrix, default config flip, ...)
- R AMX cell rule: encoded-only / executed / executed-under-contention
- S consumer audits: Belichtungsmesser (cascade) and Fisher-z
Sections A-L are left as written; Q records what in them is stale.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
…-gni3cw zspace: Fisher-Z entry tax without materializing z (ZGamma codes) + binomial Hamming null; batch Welford recovered from #323
What
1. Two functions that claimed SIMD they didn't have (
8ec54b0)popcount_batch_u64had a comment promising "POPCNT if available", but its body was a plaincount_onesloop with no dispatch at all.hamming_avx2was compiled with#[target_feature(enable = "avx2")]but contained no AVX2 instruction (its own comment: "No U8x32 polyfill available").Both now go through the
U64x8polyfill. Each backend lowers itspopcnt()differently: VPOPCNTQ whenavx512vpopcntdqis compiled in,vcntq_u8on NEON,i8x16_popcnton wasm simd128, per-lanecount_oneson the AVX2 and scalar builds, andcore::simdon nightly.hamming_avx2is replaced byhamming_u64x8, which needs nounsafeand no runtime AVX2 check.dispatch_hammingnow falls back tohamming_u64x8on every architecture.hamming_scalaris now#[cfg(test)]and serves as the test reference.2.
hamming_distance_within: exact Hamming with an early exit (e715acd,88bce53)hamming_distance_within(a, b, max) -> Option<u64>returnsSome(d)if the distanced <= max, andNoneotherwise.maxthe rest of the input is skipped.hamming_distance_rawuses, so every block still gets the fastest kernel the CPU supports.It is wired into the cascade search (
hpc/cascade.rs, the "Belichtungsmesser") wherever the code compares an exact integer distance to the threshold:Cascade::testquery: the remaining bytes may use only the budgetthreshold − prefix distance, and a prefix already over the threshold is rejected without scanning the restPackedDatabase::cascade_query, with the same budget ruleThe Stroke-1 comparisons are unchanged. They compare a scaled
f64estimate rather than an exact integer, so a budget there would first need an argument that rounding can't flip the result. Results are identical by construction:d_prefix + d_rest <= Tholds exactly whend_prefix <= Tandd_rest <= T − d_prefix.3. Batch Welford for the rolling floor (
75cc287,5e031ad)moments_u32(&[u32]) -> MomentsU32 { n, sum, sum_sq }computes exact integer moments, eight lanes at a time throughU64x8.u128totals, before an 8-lane sum could overflowu64.MomentsU32::mergeis plain integer addition. It is exact, associative and commutative, so moments computed per shard in parallel combine to the same result as one sequential pass.mean()/variance()computen·Σx² − (Σx)²exactly inu128whenever it fits, which is always true at Hamming scale. Only the final division rounds.Cascade::observe_batch(&[u32])folds a whole batch into the rolling floor using the parallel-variance merge (Chan, Golub & LeVeque).mu/sigma/observationsas callingobserveonce per element, up to f64 rounding.observeapplies per step.ndarray::simd.Tests
popcount_batch_u64andhamming_u64x8against a simple per-word reference. They cover every leftover length at the chunk boundaries, offset and unequal-length slices, and the all-ones worst case.hamming_distance_withinsetmaxequal to the true distanced, and tod − 1,d + 1,0andu64::MAX, at lengths across the 256-byte block boundaries. They also cover unequal lengths and a case that must be rejected within the first block.queryandPackedDatabase::cascade_querymust return exactly the candidates with8 + k <= T. The existing cascade tests did not catch a wrong budget.moments_u32tests:u128reference, across the chunk boundaries and for full-range values;u32::MAX, where each square nearly fills au64by itself;f64reference.observe_batchtests:observefrom both an empty and an already-populated state;threshold − prefixto the full thresholdobserve_batchtestsTest runs by build
native(AVX-512F/BW, no VPOPCNTDQ)config-v3(AVX2)observe_batch)cargo clippy -p ndarray --lib --tests -- -D warningsandcargo fmt --checkare clean.Measured (single host, AVX-512F/BW without VPOPCNTDQ, release, best of 7–9 runs)
Batch Welford, 1 M
u32distances:moments_u32observe× 1Mobserve_batchCascade query: 20 000 vectors of 2 KiB, 400 planted neighbours. Hit counts are identical before and after.
querybefore → afterPackedDatabase::cascade_querybefore → afterexamples/ternlog_fold_probe, M arm (mask_ternlogfollowed bypopcount_batch_u64), one run per build, median ns. Numbers recorded before this PR aren't directly comparable, becausepopcount_batch_u64is this arm's baseline.4. Coverage matrix (
edbd97f, documentation only)This adds sections N–S to
.claude/knowledge/agnostic-surface-cpu-matrix.md:core::simdcounterpart column, including gaps that could go upstream.Not in this PR
The other Welford implementations in lance-graph and ladybug-rs:
blasgraph/hdr.rs, the integer version with reservoir and empirical auto-switch;perturbation-simRollingFloor;scientific.rsOnlineStats.They could use
moments_u32for batch updates, but they live in other repos.popcount_batch_u64and the Hamming functions are not yet in the cross-backend parity crates.🤖 Generated with Claude Code
https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm