Skip to content

simd: keyed-reduction family — group count/min/max over resident or VIA keys - #320

Merged
AdaWorldAPI merged 3 commits into
masterfrom
claude/fold-distillation-pr-wave-s57uj7
Sep 23, 2026
Merged

AdaWorldAPI merged 3 commits into
masterfrom
claude/fold-distillation-pr-wave-s57uj7

Conversation

@AdaWorldAPI

@AdaWorldAPI AdaWorldAPI commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

What

This adds the keyed-reduction family that DuckDB-style GROUP BY needs, as T0 masking primitives surfaced through ndarray::simd:

function SQL seed
masked_group_count_u32 / _via COUNT(*) … GROUP BY k 0
masked_group_min_i32 / _via MIN(v) … GROUP BY k i64::MAX
masked_group_max_i32 / _via MAX(v) … GROUP BY k i64::MIN

The resident variant reads the group of row i as keys[i]. The _via variant reads it as table[index[i]], with the zero-fallback at both hops that masked_group_sum_i32_via already has.

One walk, named instances

Every keyed reduction is the same walk: visit the rows the mask selects, resolve each row's group through a key address, and fold the row into its slot. Only the fold and the address vary. So the walk now exists once:

  • group_walk: the bit loop, the tail clamp and the caller-named panic message;
  • GroupKeyAddr::{Resident, Via}: resolving a row to its group, including both drops.

Each public function is a named closure over the walker. This is the same facade-over-mechanics shape simd.rs uses for its backends. masked_group_sum_i32 and masked_group_sum_i32_via are rebased onto it with unchanged contracts, panic messages and doctests. A new keyed reduction is now one closure, not another copy of the loop.

Why MIN/MAX use an i64 sink seeded outside the i32 range: a slot that still holds its seed is exactly an empty group, which is SQL NULL. That's recoverable without a second counting pass.

Tests

  • All eight family members are checked against an independent longhand scalar reference, over lengths 0, 1, 63, 64, 65, 130 and 1000, for both key addresses.
  • Anti-vacuity checks confirm the fixture really hits:
    • out-of-universe keys;
    • both VIA drops;
    • i32::MIN;
    • a dirty mask tail.
  • Also covered:
    • empty groups keep their seed;
    • calls accumulate rather than overwrite;
    • the caller's name appears in a short-mask panic;
    • a MIN call with mismatched key and value lengths is refused.

Each check was deliberately broken against the committed code, then restored:

disable red
out-of-universe keys admitted reference differential (out-of-bounds panic)
VIA first hop unguarded reference differential
mask tail clamp removed reference differential
MIN folds with max 3 tests

cargo test --lib simd_masking_ops: 122 passed. The new doctests pass, and clippy with -D warnings and fmt --check are clean.

Consumer

Next is lance-graph: mask-risc gets one grouped-reduce terminal, and quack folds GROUP BY COUNT/MIN/MAX to one program instead of K. That runs under the DuckDB differential harness.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG

Summary by CodeRabbit

  • New Features
    • Added masked group count, minimum, and maximum aggregations, with options for direct keys or keys resolved through an index.
    • These aggregations are available through the SIMD interface and accumulate into existing group results. Minimum and maximum preserve the supplied value for groups with no selected rows.
    • Masked group sums now support the same key lookup and row-selection behavior.

…IA keys

Every keyed reduction is the same walk: visit the rows a mask selects,
resolve each row's group through a key address, fold the row into the
group's slot. That walk now exists once (`group_walk`) with one address
type (`GroupKeyAddr::{Resident, Via}`), and each public function is a
named instance — the facade shape `simd.rs` uses over its backends.

New, surfaced through `ndarray::simd`:
- masked_group_count_u32 / _via       COUNT(*) GROUP BY
- masked_group_min_i32 / _via         MIN GROUP BY (caller seeds i64::MAX)
- masked_group_max_i32 / _via         MAX GROUP BY (caller seeds i64::MIN)

MIN/MAX fold into an i64 sink seeded outside the i32 range, so a slot that
still holds its seed is exactly an empty group — SQL NULL — recoverable
without a second counting pass.

masked_group_sum_i32 and masked_group_sum_i32_via are rebased onto the same
walker; their contracts, panic messages and doctests are unchanged.

Tests compare all eight members against an independent longhand scalar
reference over lengths 0..1000, with anti-vacuity asserting the fixture
really hits out-of-universe keys, both VIA drops, i32::MIN and a dirty
mask tail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
@coderabbitai

coderabbitai Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: c39a1c75-17be-4019-a14f-d7ce2900f6bf

📥 Commits

Reviewing files that changed from the base of the PR and between ac39250 and 6bf3d9f.

📒 Files selected for processing (1)
  • src/simd_masking_ops.rs
 ______________________________
< I'm not mad, just debugging. >
 ------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: a0e65a75-feac-4adb-9608-164f72d7e665

📥 Commits

Reviewing files that changed from the base of the PR and between 040b510 and ac39250.

📒 Files selected for processing (1)
  • src/simd_masking_ops.rs

Included review availability: 1 review is currently available. Your included PR review attempts over the past 7 days set your current allowance at 3 reviews per hour.


📝 Walkthrough

Walkthrough

The change adds a shared walker for masked keyed sums and adds count, minimum, and maximum reductions for resident and indirect keys. It re-exports the new reductions and tests boundary cases, invalid keys, accumulation, and input-length errors.

Changes

Masked Group Reductions

Layer / File(s) Summary
Keyed row traversal
src/simd_masking_ops.rs
Both masked sum variants now use a shared walker to scan selected rows, resolve keys, skip invalid groups, and validate mask length.
Group reduction APIs
src/simd_masking_ops.rs, src/simd.rs
Adds count, minimum, and maximum variants for resident and indirect keys. The SIMD facade re-exports them. Tests compare all variants with scalar references and cover boundary cases, invalid keys, accumulation, and length errors.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Feature

Suggested reviewers: claude

Merge Risk: ⚪ Minimal · up to ac392

No actionable merge-blocking issue is identified; proceed after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 79.17% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: adding keyed group COUNT, MIN, and MAX reductions for resident and VIA-resolved keys.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR

A rabbit counts the rows in flight,
Finds smallest, largest, sums just right.
Through resident keys and tables too,
The masked paths now make it through.
It thumps beside the tests tonight.

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Sep 22, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_f469144e-04f5-419d-a15c-91a57adb8fca)

Copy link
Copy Markdown
Owner Author

realization/nightly × x86_64 failed with SIGILL, and the cause is not this PR's code. It is a mismatch between the CI cache and the runner's CPU.

  • This runner has no AVX-512. The job's own header line reads avx512f=false.
  • The cargo +nightly test step reused cached dependencies. It did not recompile num-traits, matrixmultiply and the other deps; they came from the rust-cache hit v0-rust-nightly-nightly-Linux-x64-dd9e01c4-1972360f. Under the default -Ctarget-cpu=native, those artifacts are built for whichever runner filled the cache.
  • Master's passing runs had AVX-512. Both recent master runs of this job (35774211110, 35685830686) report avx512f=true. An AVX-512 cache executed on a runner without AVX-512 gives illegal instructions.
  • The release-profile parity step is fine. It recompiled every dependency on this runner and passed, showing all 13 check groups bit-identical.
  • The new tests passed before the abort. The six group_family_tests all completed.

The workflow header already notes that the runner pool mixes AVX-512 and non-AVX-512 machines per job. This row's cache key (key: nightly) just doesn't include the CPU tier. Proposed patch, kept out of this PR so it doesn't widen it:

      - name: cpu tier for the cache key
        id: tier
        run: echo "t=$(grep -q avx512f /proc/cpuinfo && echo v4 || echo v3)" >> "$GITHUB_OUTPUT"
      - uses: Swatinem/rust-cache@v2
        with:
          key: nightly-${{ steps.tier.outputs.t }}

The same key gap exists on the other unpinned native rows (host-native). I've re-run the failed job once. If it lands on an AVX-512 runner it will pass, which confirms the diagnosis; a failure on an AVX-512 runner would mean this conclusion is wrong.


Generated by Claude Code

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review September 23, 2026 05:00
@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_b95c1598-eb36-44e1-a5a6-26829c195a80)

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 040b5109c1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/simd_masking_ops.rs Outdated
masked_group_count_u32 and masked_group_count_u32_via accumulated with
`+= 1`. On a slot the caller seeded at i64::MAX (allowed by the
accumulate-into-caller-state contract) that panicked in a debug build and
wrapped in release, so behavior depended on the build profile. The sum
family the count claims to share its contract with uses wrapping_add.
Both count variants now use wrapping_add, and the doc states it.

Test counts_wrap_like_sums_at_the_i64_boundary seeds i64::MAX and checks
both count variants (and the sum, for parity) land on i64::MIN. It
panicked at the old `+= 1` before the fix.

Reported by the Codex review on #320.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
Doc comments for the group_family_tests fixture, oracle and the three
tests that lacked one. Documentation only; no behavior change. Brings the
PR's docstring coverage over the reviewer's 80% threshold (was 79.17%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
@AdaWorldAPI
AdaWorldAPI merged commit ef5ffed into master Sep 23, 2026
24 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants