Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .claude/blackboard.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,20 @@
## 2026-09-23 (21) — ternlog → Count/Any without a mask: a slice loop, NOT a new ISA primitive

Question asked (lance-graph #1270 prompt): what is the smallest T1 operation that lets an arbitrary 2/3-input Boolean membership END in Count/Any without writing a mask — and does existing `U64x8` composition already do it register-only?

**Answer (MEASURED): the composition exists on every realization; only the slice loop was missing.** `U64x8::{ternlog::<IMM>, popcnt, +, |, reduce_sum}` are present on avx512 / avx2-polyfill / scalar / neon / wasm, so no backend code is added. Added two slice functions to `simd_masking_ops.rs` + the facade, built only from those methods:

- `mask_ternlog_popcount::<IMM>(a, b, c) -> u64` — `Σ popcount(ternlog(a,b,c))`, lane-wise accumulate, one `reduce_sum`.
- `mask_ternlog_any::<IMM>(a, b, c) -> bool` — OR-accumulate, horizontal test once per block of 8 chunks.

Word-level contract (both): every bit of every word counts, exactly as `popcount_batch_u64`/`mask_any` over the materialized `mask_ternlog` result would; an odd `IMM` sets the last word's dead tail bits and they count (caller masks the last word). Register PADDING never counts — the tail runs through the packed op and only the live lanes are read, because `ternlog(0,0,0)` is all-ones for an odd table (pinned by `mask_ternlog_folds_never_count_register_padding`, NOR3 over 9 words).

Evidence: `examples/ternlog_fold_probe.rs` (M = materialize+reduce, R = register fold, S = scalar fused), `AND2_OR`. Count M/R: avx2 1.27–1.50×, avx512 1.90–1.94×. Any M/R (all-zero worst case): avx2 3.2–6.2×, avx512 2.6–4.4×. A per-chunk Any test LOST to M on avx2 at 16K words (0.91×) — hence the block. S ≈ R on avx2 and at the memory-bound avx512 size: the win is not writing the mask, not SIMD per se.

Tests: `mask_ternlog_folds_match_the_materializing_pair_for_all_256_tables` (all 256 tables × 14 lengths × dense/sparse, against the exact pair they replace), padding test, two length-mismatch panics. Parity: `slice_ternlog!` in `crates/simd-masking-parity` now also checks both folds (`0x69x` count, `0x65x` any); green on native v4, native v3, neon-qemu, wasm, wasm-scalar.

Consumer: lance-graph-mask-risc fused terminal `MaskOp::{And,Or,Xor,AndNot,Ternlog} → Count/Any`.

## 2026-09-21 (20) — six index-addressed mask primitives (0xDxx): gather / scatter-or / keyed group-sum (direct + via index) / indexed equality / ORDERED key-run distinct fold

All in `simd_masking_ops.rs` + the `simd::` facade, documented in ADDRESS terms only (`index` / `table` / `keys`; no join, foreign-key, semijoin, table-name or ERP vocabulary — T1 does not know what a consumer means by an index lane). All are deliberately scalar bit-walks: permutations/scatters indexed by data, not a fixed stride, so none of this crate's backends can vector-load them (same shape as `masked_strided_group_sum`, which is NOT a keyed group-by — it sums one record's own byte-groups into a scalar; zero callers of it are affected).
Expand Down
4 changes: 4 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,10 @@ required-features = ["std"]
name = "ternlog_amortization_probe"
required-features = ["std"]

[[example]]
name = "ternlog_fold_probe"
required-features = ["std"]

[[example]]
name = "r2il_column_scan_probe"
required-features = ["std"]
Expand Down
29 changes: 19 additions & 10 deletions crates/simd-masking-parity/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -38,12 +38,12 @@ use ndarray::simd::{
le_i32_to_mask_under, le_u64_to_mask, le_u8_to_mask, lt_i32_to_mask, lt_i32_to_mask_under, lt_u64_to_mask,
lt_u8_to_mask, mask_all, mask_and, mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32,
mask_not, mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_shift_morton,
mask_ternlog, mask_ternlog_assign, mask_xor, mask_xor_assign, masked_group_sum_i32, masked_group_sum_i32_via,
masked_key_run_count_u32, masked_max_i32, masked_min_i32, masked_strided_group_sum, masked_sum_i32,
masked_sum_wrapping_add_i32, ne_i32_to_mask,
ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, ne_u64_to_mask, ne_u8_to_mask,
ternary_match_strided_to_mask, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under,
ternary_match_u64_to_mask, ternary_match_u64_to_mask_under, ternlog, I32x16, KeyRunCarry, MortonDir, U32x16, U64x8,
mask_ternlog, mask_ternlog_any, mask_ternlog_assign, mask_ternlog_popcount, mask_xor, mask_xor_assign,
masked_group_sum_i32, masked_group_sum_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32,
masked_strided_group_sum, masked_sum_i32, masked_sum_wrapping_add_i32, ne_i32_to_mask, ne_i32_to_mask_under,
ne_u32_to_mask, ne_u32_to_mask_under, ne_u64_to_mask, ne_u8_to_mask, ternary_match_strided_to_mask,
ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, ternary_match_u64_to_mask,
ternary_match_u64_to_mask_under, ternlog, I32x16, KeyRunCarry, MortonDir, U32x16, U64x8,
};

/// Number of check groups [`run`] executes (for the log line only).
Expand Down Expand Up @@ -642,6 +642,14 @@ fn check_mask_algebra() -> Result<(), u32> {
if t != dst {
return Err($code | 0x8);
}
// The no-mask folds equal the materializing pair they replace.
let want: u64 = dst.iter().map(|w| u64::from(w.count_ones())).sum();
if mask_ternlog_popcount::<IMM>(&a, &b, &c) != want {
return Err($code | 0x80);
}
if mask_ternlog_any::<IMM>(&a, &b, &c) != mask_any(&dst) {
return Err($code | 0x40);
}
}};
}
slice_ternlog!(ternlog::AND2_OR, 0x610);
Expand Down Expand Up @@ -788,9 +796,7 @@ fn check_masked_reductions() -> Result<(), u32> {
}
let want_wrapping_add = (0..n)
.filter(|&i| (m[i / 64] >> (i % 64)) & 1 == 1)
.fold(0i64, |acc, i| {
acc.wrapping_add(vals[i].wrapping_add(rhs[i]) as i64)
});
.fold(0i64, |acc, i| acc.wrapping_add(vals[i].wrapping_add(rhs[i]) as i64));
if masked_sum_wrapping_add_i32(&vals, &rhs, m) != want_wrapping_add {
return Err(0x850 | k);
}
Expand Down Expand Up @@ -1410,7 +1416,10 @@ fn check_gather_scatter_group() -> Result<(), u32> {
if n >= 2 && keys[0] < keys[n - 1] {
let mut bad = keys[1..].to_vec();
bad.push(keys[0]);
let mut c = KeyRunCarry { key: Some(bad[0]), hit: true };
let mut c = KeyRunCarry {
key: Some(bad[0]),
hit: true,
};
let before = c;
if masked_key_run_count_u32(&bad, &sel_bits, &mut c).is_some() {
return Err(0xD52);
Expand Down
166 changes: 166 additions & 0 deletions examples/ternlog_fold_probe.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,166 @@
//! Can a 2/3-input Boolean membership END in Count/Any without writing a mask?
//!
//! The question this probe answers is narrower than "is a fused primitive
//! faster". It is: **does the existing `U64x8` surface already compose the
//! fold register-only**, or is a new T1 primitive required? Every realization
//! of `U64x8` (avx512 / avx2-polyfill / scalar / neon / wasm) carries
//! `ternlog::<IMM>`, `popcnt`, `+` and `reduce_sum`, so the composition
//! `ternlog → popcnt → accumulate → reduce_sum` type-checks everywhere. What
//! the probe measures is whether that composition is a real win over the
//! materializing path, on each backend this workspace builds for.
//!
//! Three arms compute the IDENTICAL result for `(a & b) | c` (`AND2_OR`, the
//! mask-risc flagship immediate); a correctness gate aborts on disagreement:
//!
//! | arm | shape | derived words written |
//! |---|---|---|
//! | M | `mask_ternlog::<IMM>` into a buffer, then `popcount_batch_u64` | n |
//! | R | chunked `U64x8::ternlog → popcnt`, lane-wise accumulate, one `reduce_sum` | 0 |
//! | S | plain scalar `((a & b) \| c).count_ones()` over zipped slices | 0 |
//!
//! Any is measured the same way: M = materialize + `mask_any`; R =
//! OR-accumulate with a test once per block of 8 chunks; S = scalar `any`. Any
//! is timed on an all-zero result (the worst case: no early exit is possible).
//! A first R arm that tested every chunk LOST to M on avx2 at 16K words (0.91×)
//! — the per-chunk horizontal test cost more than the ternlog it guarded.
//!
//! Measured 2026-09-23 (median ns, avx2 = `config-v3`, avx512 = `config-v4`,
//! a host without `avx512vpopcntdq`, so avx512 `popcnt` is the LUT path):
//!
//! | backend | words | Count M/R | Any M/R |
//! |---|---|---|---|
//! | avx2 | 16 384 | 1.27 | 3.24 |
//! | avx2 | 262 144 | 1.50 | 6.18 |
//! | avx512 | 16 384 | 1.90 | 2.59 |
//! | avx512 | 262 144 | 1.94 | 4.39 |
//!
//! The scalar fused arm S is ~equal to R on avx2 and at the memory-bound size
//! on avx512 — the win here is NOT writing the mask, not SIMD per se.
//!
//! Usage (the backend is compile-time and the default config is
//! `target-cpu=native`, so pin the tier and read the `backend:` line):
//!
//! ```text
//! env -u RUSTFLAGS cargo --config .cargo/config-v3.toml run --release --example ternlog_fold_probe
//! env -u RUSTFLAGS cargo --config .cargo/config-v4.toml run --release --example ternlog_fold_probe
//! ```

use std::time::Instant;

use ndarray::simd::ternlog::AND2_OR;
use ndarray::simd::{mask_any, mask_ternlog, popcount_batch_u64, U64x8};

const L: usize = U64x8::LANES;

fn lcg(s: &mut u64) -> u64 {
*s = s
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
*s
}

/// Arm R: the register-only composition over existing `U64x8` methods.
fn fold_count(a: &[u64], b: &[u64], c: &[u64]) -> u64 {
let (ca, ta) = a.as_chunks::<L>();
let (cb, tb) = b.as_chunks::<L>();
let (cc, tc) = c.as_chunks::<L>();
let mut acc = U64x8::splat(0);
for ((x, y), z) in ca.iter().zip(cb).zip(cc) {
let t = U64x8::from_array(*x).ternlog::<AND2_OR>(U64x8::from_array(*y), U64x8::from_array(*z));
acc += t.popcnt();
}
let mut tail = 0u64;
for ((x, y), z) in ta.iter().zip(tb).zip(tc) {
tail += ((x & y) | z).count_ones() as u64;
}
acc.reduce_sum() + tail
}

fn fold_any(a: &[u64], b: &[u64], c: &[u64]) -> bool {
let (ca, ta) = a.as_chunks::<L>();
let (cb, tb) = b.as_chunks::<L>();
let (cc, tc) = c.as_chunks::<L>();
// OR-accumulate in a register and test once per block of chunks: a
// per-chunk horizontal test costs more than the ternlog it guards.
const BLOCK: usize = 8;
let mut acc = U64x8::splat(0);
for (i, ((x, y), z)) in ca.iter().zip(cb).zip(cc).enumerate() {
acc |= U64x8::from_array(*x).ternlog::<AND2_OR>(U64x8::from_array(*y), U64x8::from_array(*z));
if i % BLOCK == BLOCK - 1 && acc.to_array().iter().any(|&w| w != 0) {
return true;
}
}
if acc.to_array().iter().any(|&w| w != 0) {
return true;
}
ta.iter()
.zip(tb)
.zip(tc)
.any(|((x, y), z)| (x & y) | z != 0)
}

fn median<T>(reps: usize, mut f: impl FnMut() -> T) -> (f64, T) {
let mut ts = Vec::with_capacity(reps);
let mut last = None;
for _ in 0..reps {
let t = Instant::now();
last = Some(std::hint::black_box(f()));
ts.push(t.elapsed().as_nanos() as f64);
}
ts.sort_by(|x, y| x.total_cmp(y));
(ts[reps / 2], last.expect("reps > 0"))
}

fn main() {
let backend = if cfg!(target_feature = "avx512f") {
"avx512"
} else if cfg!(target_feature = "avx2") {
"avx2-polyfill"
} else {
"scalar"
};
println!("backend: {backend}");
println!(
"{:>8} {:>6} {:>12} {:>12} {:>12} {:>8} {:>8}",
"words", "term", "M_ns", "R_ns", "S_ns", "M/R", "S/R"
);
let mut seed = 0x7E4_u64;
for &n in &[8usize, 128, 16_384, 262_144] {
let a: Vec<u64> = (0..n).map(|_| lcg(&mut seed)).collect();
let b: Vec<u64> = (0..n).map(|_| lcg(&mut seed)).collect();
let c: Vec<u64> = (0..n).map(|_| lcg(&mut seed) & lcg(&mut seed)).collect();
let zero = vec![0u64; n];
let mut dst = vec![0u64; n];
let reps = if n > 100_000 { 41 } else { 2001 };

let (m, vm) = median(reps, || {
mask_ternlog::<AND2_OR>(&a, &b, &c, &mut dst);
popcount_batch_u64(&dst)
});
let (r, vr) = median(reps, || fold_count(&a, &b, &c));
let (s, vs) = median(reps, || {
a.iter()
.zip(&b)
.zip(&c)
.map(|((x, y), z)| ((x & y) | z).count_ones() as u64)
.sum::<u64>()
});
assert!(vm == vr && vr == vs, "count disagreement at n={n}: {vm} {vr} {vs}");
println!("{n:>8} {:>6} {m:>12.0} {r:>12.0} {s:>12.0} {:>8.2} {:>8.2}", "Count", m / r, s / r);

// Any over an all-zero result: a = b = c = 0, so no early exit.
let (m, vm) = median(reps, || {
mask_ternlog::<AND2_OR>(&zero, &zero, &zero, &mut dst);
mask_any(&dst)
});
let (r, vr) = median(reps, || fold_any(&zero, &zero, &zero));
let (s, vs) = median(reps, || {
zero.iter()
.zip(&zero)
.zip(&zero)
.any(|((x, y), z)| (x & y) | z != 0)
});
assert!(!vm && !vr && !vs, "any disagreement at n={n}");
println!("{n:>8} {:>6} {m:>12.0} {r:>12.0} {s:>12.0} {:>8.2} {:>8.2}", "Any", m / r, s / r);
}
}
2 changes: 2 additions & 0 deletions src/simd.rs
Original file line number Diff line number Diff line change
Expand Up @@ -823,7 +823,9 @@ pub use crate::simd_masking_ops::{
mask_set_range,
mask_shift_morton,
mask_ternlog,
mask_ternlog_any,
mask_ternlog_assign,
mask_ternlog_popcount,
mask_xor,
mask_xor_assign,
masked_group_count_u32,
Expand Down
Loading
Loading