simd: mask_ternlog_popcount / mask_ternlog_any — Boolean membership ends in Count/Any with no mask written - #322
Merged
Conversation
…mbership ends in Count/Any with no mask written Built only from U64x8 methods every realization already carries (ternlog, popcnt, +, |, reduce_sum), so no backend code is added: the missing piece was the slice loop, not an ISA primitive. Word-level contract identical to mask_ternlog + popcount_batch_u64 / mask_any; register padding is never counted (an odd table evaluates padding lanes to all-ones). Probe examples/ternlog_fold_probe.rs: Count 1.3-1.9x, Any 2.6-6.2x over the materializing pair on avx2/avx512. Tests cover all 256 tables x 14 lengths x dense/sparse against the pair they replace; parity crate extended, green on v4, v3, neon-qemu, wasm, wasm-scalar. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
Contributor
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (6)
✨ Finishing Touches📝 Generate docstrings
Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_1a49c2b0-e342-441f-b798-0361c9373764) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Two slice functions in
simd_masking_ops.rs, re-exported fromndarray::simd:mask_ternlog_popcount::<IMM>(a, b, c) -> u64:Σ popcount(ternlog(a, b, c))mask_ternlog_any::<IMM>(a, b, c) -> bool: is any bit ofternlog(a, b, c)set?A 2- or 3-input Boolean membership can now end in Count or Any without writing a mask. Two-input functions are the same call with a table that ignores
c.The question it answers
Is a new T1 primitive needed for "Boolean membership → Count/Any", or can the existing
U64x8operations already do it without touching memory?Measured answer: nothing new is needed below the slice loop.
U64x8::{ternlog, popcnt, +, |, reduce_sum}already exist in every backend (avx512, avx2-polyfill, scalar, neon, wasm). These two functions are loops over those methods, so no backend code,cfgor intrinsics are added.Contract
mask_ternlog+popcount_batch_u64/mask_anyover the materialized mask.IMMsets the dead tail bits of the last word, and those bits are counted. The caller masks the last word, the same rulemask_ternlogalready documents for its output.mask_ternlog_anychecks once per block of 8 chunks, not once per chunk. A per-chunk check lost to the materializing pair on avx2 (0.91×).Evidence
Probe:
examples/ternlog_fold_probe.rs. Truth tableAND2_OR; timings are medians; "M/R" is the speed-up over materializing the mask and then reducing it.config-v3)config-v4, novpopcntdq)A plain scalar fused loop roughly ties the register fold on avx2 and at the memory-bound size on avx512. The win comes from not writing the mask, not from SIMD as such.
Tests
mask_ternlog_folds_match_the_materializing_pair_for_all_256_tablescompares each fold against the exact pair it replaces. It covers all 256 truth tables × 14 lengths (0 … 129) × dense and sparse inputs.mask_ternlog_folds_never_count_register_paddingcovers NOR3 over 9 words (7 padding lanes), hits that sit only in the tail, and hits past the first block.crates/simd-masking-parityslice_ternlog!now checks both folds (codes0x69xcount,0x65xany). It passes on native v4, native v3,neon-qemu,wasmandwasm-scalar.clippy -D warningspasses on v3 and v4.Consumer
lance-graph-mask-riscwill use these for its fused terminal pathMaskOp::{And, Or, Xor, AndNot, Ternlog} → Count/Any, in a follow-up lance-graph PR.🤖 Generated with Claude Code
https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG