Skip to content

bitwise/statistics: U64x8 popcount/Hamming, early-exit Hamming, batch Welford for the cascade - #323

Merged
AdaWorldAPI merged 4 commits into
masterfrom
claude/llvm-codegen-polyfill-gni3cw
Sep 23, 2026
Merged

AdaWorldAPI merged 4 commits into
masterfrom
claude/llvm-codegen-polyfill-gni3cw

Conversation

@AdaWorldAPI

@AdaWorldAPI AdaWorldAPI commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

What

1. Two functions that claimed SIMD they didn't have (8ec54b0)

  • popcount_batch_u64 had a comment promising "POPCNT if available", but its body was a plain count_ones loop with no dispatch at all.
  • hamming_avx2 was compiled with #[target_feature(enable = "avx2")] but contained no AVX2 instruction (its own comment: "No U8x32 polyfill available").

Both now go through the U64x8 polyfill. Each backend lowers its popcnt() differently: VPOPCNTQ when avx512vpopcntdq is compiled in, vcntq_u8 on NEON, i8x16_popcnt on wasm simd128, per-lane count_ones on the AVX2 and scalar builds, and core::simd on nightly.

  • hamming_avx2 is replaced by hamming_u64x8, which needs no unsafe and no runtime AVX2 check.
  • dispatch_hamming now falls back to hamming_u64x8 on every architecture.
  • hamming_scalar is now #[cfg(test)] and serves as the test reference.

2. hamming_distance_within: exact Hamming with an early exit (e715acd, 88bce53)

hamming_distance_within(a, b, max) -> Option<u64> returns Some(d) if the distance d <= max, and None otherwise.

  • The running total only grows, so once it passes max the rest of the input is skipped.
  • It checks after every 256-byte block. Each block is counted by the same runtime-dispatched kernel hamming_distance_raw uses, so every block still gets the fastest kernel the CPU supports.

It is wired into the cascade search (hpc/cascade.rs, the "Belichtungsmesser") wherever the code compares an exact integer distance to the threshold:

  • Cascade::test
  • the small-vector path
  • Stroke 2 of query: the remaining bytes may use only the budget threshold − prefix distance, and a prefix already over the threshold is rejected without scanning the rest
  • Stroke 3 of PackedDatabase::cascade_query, with the same budget rule

The Stroke-1 comparisons are unchanged. They compare a scaled f64 estimate rather than an exact integer, so a budget there would first need an argument that rounding can't flip the result. Results are identical by construction: d_prefix + d_rest <= T holds exactly when d_prefix <= T and d_rest <= T − d_prefix.

3. Batch Welford for the rolling floor (75cc287, 5e031ad)

  • moments_u32(&[u32]) -> MomentsU32 { n, sum, sum_sq } computes exact integer moments, eight lanes at a time through U64x8.
    • Each square is split into 32-bit halves, so every lane add stays below 2³².
    • Every 2²⁸ chunks the lanes are drained into u128 totals, before an 8-lane sum could overflow u64.
  • MomentsU32::merge is plain integer addition. It is exact, associative and commutative, so moments computed per shard in parallel combine to the same result as one sequential pass.
  • mean() / variance() compute n·Σx² − (Σx)² exactly in u128 whenever it fits, which is always true at Hamming scale. Only the final division rounds.
  • Cascade::observe_batch(&[u32]) folds a whole batch into the rolling floor using the parallel-variance merge (Chan, Golub & LeVeque).
    • It produces the same mu / sigma / observations as calling observe once per element, up to f64 rounding.
    • Shards can be folded in any order.
    • It checks for drift once per batch: an alert fires when μ moves by more than 2σ of the pre-batch state. That is the same test observe applies per step.
  • Both are re-exported from ndarray::simd.

Tests

  • Reference-parity tests check popcount_batch_u64 and hamming_u64x8 against a simple per-word reference. They cover every leftover length at the chunk boundaries, offset and unequal-length slices, and the all-ones worst case.
  • Exactness tests for hamming_distance_within set max equal to the true distance d, and to d − 1, d + 1, 0 and u64::MAX, at lengths across the 256-byte block boundaries. They also cover unequal lengths and a case that must be rejected within the first block.
  • Falsifier tests for the cascade budget. Every candidate differs from the query in 8 prefix bits, so all of them survive Stroke 1, and their full distances straddle the threshold. query and PackedDatabase::cascade_query must return exactly the candidates with 8 + k <= T. The existing cascade tests did not catch a wrong budget.
  • moments_u32 tests:
    • exact against a u128 reference, across the chunk boundaries and for full-range values;
    • worst case of 100 003 values of u32::MAX, where each square nearly fills a u64 by itself;
    • merge is exact at every split point, in both orders;
    • mean and variance match a two-pass f64 reference.
  • observe_batch tests:
    • matches per-element observe from both an empty and an already-populated state;
    • shards folded in any order give the same result;
    • the drift alert fires on a shifted batch and stays silent on a stable one.
  • Disable checks. Each of these deliberate bugs makes the listed tests fail:
deliberate bug tests that fail
leftover-word handling removed in both polyfill functions 5
budget changed from threshold − prefix to the full threshold both budget tests
high half of the square dropped 4 moments tests
merge's cross term dropped 2 observe_batch tests
alert forced to never fire the alert test
alert forced to always fire the alert test

Test runs by build

build result
host native (AVX-512F/BW, no VPOPCNTDQ) all pass
config-v3 (AVX2) 46/46 (bitwise, cascade, moments, observe_batch)
aarch64 NEON under qemu 38/38 (same set)
  • wasm32 simd128 was compile-checked only; its tests were not run.
  • cargo clippy -p ndarray --lib --tests -- -D warnings and cargo fmt --check are clean.

Measured (single host, AVX-512F/BW without VPOPCNTDQ, release, best of 7–9 runs)

Batch Welford, 1 M u32 distances:

build scalar u128 moments moments_u32 observe × 1M observe_batch
native 760 µs 209 µs 5320 µs 178 µs
v3 (AVX2) 485 µs 433 µs 5672 µs 426 µs

Cascade query: 20 000 vectors of 2 KiB, 400 planted neighbours. Hit counts are identical before and after.

threshold query before → after PackedDatabase::cascade_query before → after
400 512 → 460 µs 204 → 193 µs
800 490 → 465 µs 238 → 227 µs
7000 2415 → 2198 µs 750 → 702 µs

examples/ternlog_fold_probe, M arm (mask_ternlog followed by popcount_batch_u64), one run per build, median ns. Numbers recorded before this PR aren't directly comparable, because popcount_batch_u64 is this arm's baseline.

build 8 words 128 16 384 262 144
v3 55 → 43 119 → 87 10 787 → 10 265 407 483 → 412 181
v4 53 → 44 116 → 81 10 154 → 9 430 400 719 → 426 154

4. Coverage matrix (edbd97f, documentation only)

This adds sections N–S to .claude/knowledge/agnostic-surface-cpu-matrix.md:

  • N: masking operations on each backend.
  • O: the crypto lane.
  • P: the core::simd counterpart column, including gaps that could go upstream.
  • Q: corrections to the stale parts of sections A–L.
  • R: the verification rule for AMX cells.
  • S: the Belichtungsmesser and Fisher-z audits.

Not in this PR

  • The other Welford implementations in lance-graph and ladybug-rs:

    • blasgraph/hdr.rs, the integer version with reservoir and empirical auto-switch;
    • perturbation-sim RollingFloor;
    • the duplicated scientific.rs OnlineStats.

    They could use moments_u32 for batch updates, but they live in other repos.

  • popcount_batch_u64 and the Hamming functions are not yet in the cross-backend parity crates.

🤖 Generated with Claude Code

https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm

…64x8

popcount_batch_u64 carried a comment promising "POPCNT if available" over a
plain count_ones loop with no dispatch. hamming_avx2 was compiled with
target_feature(avx2) but contained no AVX2 instruction (its own comment:
"No U8x32 polyfill available").

Both now run through the U64x8 polyfill, whose popcnt() every realization
lowers natively (VPOPCNTQ with avx512vpopcntdq, vcntq_u8 on NEON,
i8x16_popcnt on wasm simd128, per-lane count_ones on the AVX2 and scalar
builds). hamming_avx2 is replaced by the safe, ungated hamming_u64x8, which
dispatch_hamming now uses as its fallback on every architecture instead of
hamming_scalar; hamming_scalar stays as the cfg(test) oracle.

Tests: new reference-parity tests for both functions (every tail length
across the 8-word / 64-byte chunk boundary, unaligned and unequal-length
slices, all-ones worst case); the AVX2-gated tier tests now exercise
hamming_u64x8 unconditionally.

Note: popcount_batch_u64 is the baseline side of examples/ternlog_fold_probe,
so numbers recorded from that probe before this commit are not comparable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: eb2b9e5a-b8cd-4203-a07c-8575453810a3

📥 Commits

Reviewing files that changed from the base of the PR and between 8ec54b0 and edbd97f.

📒 Files selected for processing (4)
  • .claude/knowledge/agnostic-surface-cpu-matrix.md
  • src/bitwise.rs
  • src/hpc/cascade.rs
  • src/simd.rs

Included review availability: 1 review is currently available. Your included PR review attempts over the past 7 days set your current allowance at 3 reviews per hour.


📝 Walkthrough

Walkthrough

Hamming dispatch now uses hamming_u64x8 after the existing AVX-512 checks. Batch popcount uses chunked U64x8 processing. The new hamming_distance_within function stops when the distance exceeds its limit. Cascade query paths use bounded checks against remaining thresholds.

Changes

Hamming distance checks

Layer / File(s) Summary
Chunked Hamming and popcount
src/bitwise.rs
Hamming dispatch uses U64x8 as its fallback. Batch popcount processes chunks with U64x8 and counts the tail scalarly. Tests compare results with references across chunk boundaries, offset slices, unequal lengths, and large inputs.
Bounded distance function and facade
src/bitwise.rs, src/simd.rs
Added and re-exported hamming_distance_within. It processes the common prefix in 256-byte blocks and returns None when the running distance exceeds the limit. Tests cover threshold and block-boundary cases.
Cascade threshold-budget checks
src/hpc/cascade.rs
Cascade test and query paths use bounded checks. Stroke 2 and packed Stroke 3 limit suffix checks to the threshold remaining after the prefix. Added tests for candidate results around the threshold.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Suggested reviewers: claude

Merge Risk: ⚪ Minimal · up to edbd9

No actionable issue is identified in the bounded Hamming and cascade changes; the PR is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 77.42% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 31 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main changes: U64x8-based popcount and Hamming processing, plus exact early-exit Hamming checks for the cascade.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR

A rabbit counts bits by the byte,
Through chunks that run swift and straight.
A threshold says, “Stop when too high,”
Cascade checks as candidates pass by.
Tail bits settle the final score.

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_3a56f8c9-5881-4c26-8ff9-c35e7ab9cf54)

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review September 23, 2026 19:49
New `hamming_distance_within(a, b, max) -> Option<u64>`: `Some(d)` iff the
distance d <= max, else None. The running total only grows, so once it
passes `max` the rest is skipped; the check runs after every 256-byte block,
each counted by the same runtime-dispatched kernel hamming_distance_raw uses
(VPOPCNTDQ / AVX-512BW / U64x8 polyfill), so no block loses its best lowering.
Re-exported from ndarray::simd beside hamming_distance_raw.

Wired into the Belichtungsmesser (hpc/cascade.rs) at every EXACT
comparison: Cascade::test, the small-vector path, Stroke 2 of query (budget
= threshold - prefix distance; a prefix already over it rejects without a
scan) and PackedDatabase Stroke 3. The Stroke 1 estimate comparisons
(scaled f64) are left untouched: an integer budget there would need a
rounding argument, and those loops only read the short prefix anyway.

Results are identical by construction (d_prefix + d_rest <= T  <=>
d_prefix <= T and d_rest <= T - d_prefix). Tests: exactness at max = d,
d-1, d+1, 0, u64::MAX across block boundaries and unequal lengths, and a
first-block reject.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Stroke-1 survivors with full distances straddling the threshold (8 prefix
bits + k tail bits); the exact strokes must return exactly 8 + k <= T. A
budget of T instead of T - prefix would admit three extra candidates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@AdaWorldAPI AdaWorldAPI changed the title bitwise: route popcount_batch_u64 and the AVX2 Hamming tier through U64x8 bitwise: U64x8 popcount/Hamming + exact early-exit Hamming for the cascade Sep 23, 2026
…dits

Appends sections N-S to agnostic-surface-cpu-matrix.md, built from
read-only per-function audits (file:line cited throughout):

- N  masking ops per realization (compare/algebra/ternlog/fold, masked
     reductions and group-by — 19 scalar by design, with reasons), gaps
     and parity coverage; post-#323 rows for popcount/Hamming
- O  crypto lane (U32x16/U64x8 ARX per backend, in-tree consumers, gaps)
- P  core::simd counterpart of every simd_nightly method
     (native / composed / scalar-fallback / missing) and the gaps that
     are upstream candidates for portable-simd
- Q  staleness corrections to sections A-L (runtime-dispatch shipped,
     wasm backend exists, CI matrix, default config flip, ...)
- R  AMX cell rule: encoded-only / executed / executed-under-contention
- S  consumer audits: Belichtungsmesser (cascade) and Fisher-z

Sections A-L are left as written; Q records what in them is stale.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@AdaWorldAPI
AdaWorldAPI merged commit aa31a55 into master Sep 23, 2026
26 checks passed
@AdaWorldAPI AdaWorldAPI changed the title bitwise: U64x8 popcount/Hamming + exact early-exit Hamming for the cascade bitwise/statistics: U64x8 popcount/Hamming, early-exit Hamming, batch Welford for the cascade Sep 23, 2026
AdaWorldAPI added a commit that referenced this pull request Sep 24, 2026
…-gni3cw

zspace: Fisher-Z entry tax without materializing z (ZGamma codes) + binomial Hamming null; batch Welford recovered from #323
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants