diff --git a/.claude/board/TECH_DEBT.md b/.claude/board/TECH_DEBT.md index ee7b9afeb..b4aea9521 100644 --- a/.claude/board/TECH_DEBT.md +++ b/.claude/board/TECH_DEBT.md @@ -1,3 +1,43 @@ +## TD-SYM-SUM-MERGE-IS-NOT-ADDITION-1 (2026-09-23) — OPEN, dormant + +**`GroupFold::SumSymI32`'s seed is not an additive identity, so two partial +sinks must never be combined with `+`.** For MIN/MAX the seed IS the lattice +identity (`i64::MAX` for min, `i64::MIN` for max), so partial sinks merge with +plain `min`/`max` and an empty side is absorbed for free. For the `_sym` SUM +the seed `ndarray::simd::SYM_EMPTY_I64` (`i64::MIN`) marks "no row reached this +group"; `a + b` on two partials would add −2⁶³ for every group that one side +never saw, and turn a present group into garbage or a false "empty". + +The correct merge is the fold's own rule, lifted: + +```text +merge(a, b) = if a == SYM_EMPTY_I64 { b } + else if b == SYM_EMPTY_I64 { a } + else { a.wrapping_add(b) } +``` + +**Why it is dormant, not live:** nothing merges partial sinks today. The +mask-risc executor folds every tile sequentially into ONE caller-owned +`Out::I64` (the `Terminal::GroupReduce` arm of `exec.rs`, seeded once before +the first tile); `lance-graph-quack` executes one program per sink. There is no +parallel, per-segment, or per-fragment group reduction in either crate. + +**When it goes live:** the first parallel / multi-segment / multi-fragment +execution of a `GroupReduce { fold: SumSymI32 }` — e.g. a rayon split of the +tile loop, a Lance-fragment fan-out, or combining sinks across versions. That +change must use the merge above (or an equivalent `GroupFold::merge`) and must +carry a test whose partials include a group empty on ONE side only; a fixture +where every group is present on both sides cannot see the defect. + +**Existing guard, and its limit:** `sym_sum_agrees_with_full_range_sum_plus_count` +(`crates/lance-graph-mask-risc/tests/foreign.rs`) compares the `_sym` SUM with +full-range SUM + COUNT over the same data and would go red on a wrong merge — +but only once the merge path runs under it. It does not exist yet, so the test +cannot fire on it today. + +Cross-ref: `.claude/board/entries/2026-09-23-quack-having-sym-sum-presence-mask.md` +(the `_sym` decision and the presence-mask boundary); ndarray #321; lance-graph #1266. + ## TD-JC-CLIPPY-RED-ON-BASE-2 (2026-09-18) — the 1.98 pre-bump lint sweep was WORKSPACE-scoped, and `jc` is workspace-EXCLUDED **`JC Substrate Proof` is RED on `main`** (run `35335429357`, head `568965e9`): diff --git a/.claude/board/entries/2026-09-23-quack-having-sym-sum-presence-mask.md b/.claude/board/entries/2026-09-23-quack-having-sym-sum-presence-mask.md new file mode 100644 index 000000000..ae986e42a --- /dev/null +++ b/.claude/board/entries/2026-09-23-quack-having-sym-sum-presence-mask.md @@ -0,0 +1,87 @@ +# 2026-09-23 — Quack HAVING: the `_sym` SUM, the presence-mask boundary, the twin check + +**Status:** MEASURED (landed on lance-graph #1266) · OPEN (see end) +**Depends on:** ndarray #321 (`masked_group_sum_sym_i32{,_via}`, `SYM_EMPTY_I64`) +**Supersedes:** the W-C "group existence is OPEN" item of +`2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md` (left verbatim there; +that entry is not edited). Its instruction "decide the fused sum+count sink's +shape before writing `HAVING`" is resolved differently below: a reserved code +inside the fold, a mask at the boundary. + +## W-C HAVING landed — the empty-group decision + +**DECISION:** an empty group is marked by a SEED THE FOLD CANNOT REACH, not +by a count carried beside the value. The rule is uniform: **empty ⇔ the slot +still holds `GroupFold::seed`** (`GroupFold::is_empty_slot`). The 2026-09-22 +entry said "decide the fused sum+count sink's shape before writing +`HAVING`"; this supersedes that, because the sentinel closes the NULL gap +with no second sink. + +- Seeds: `Count` → 0 (never empty — a zero count is an answer), `MinI32` → + `i64::MAX`, `MaxI32` → `i64::MIN`, **new `SumSymI32` → `i64::MIN`**. MIN + cannot share SUM's seed (a min fold must start from its identity), so the + rule is "equals its own seed", not "equals `i64::MIN`". +- **BASIS:** the same move as a 4-bit code read as −7..+7 + NaN rather than + −8..+7: give up one representable value to get a NULL code and a + range closed under negation. At 4/8 bit that also removes a median bias; at + i64 the bias is negligible, and what is bought is closure + NULL. +- Cost: exactly one row of range, for this fold only. + `GROUP_SUM_SYM_MAX_ROWS = 2^32 − 1`, because exactly 2^32 rows of + `i32::MIN` sum to `i64::MIN`. `MASKED_SUM_I32_MAX_ROWS` stays `2^32`: the + coalescing sums may legitimately produce `i64::MIN`. +- Kernel: ndarray #321 `masked_group_sum_sym_i32{,_via}`, first row + REPLACES the marker, later rows `wrapping_add`. Two named closures over the + shared `group_walk`; no new bit loop. + +**HAVING is finalization.** `lower_group_having` emits one `GroupReduce` program +per aggregate over the same filter and key; `GroupHavingPlan::finish` is +the O(K) pass that yields a K-bit group mask. A group survives iff it was +REACHED (read off the first sink: count ≠ 0, or slot ≠ seed) and every +comparison holds. There is no per-predicate NULL check. Every sink shares the +filter and the key, so a reached group is non-empty in all of them, and a +guard that cannot fire would be decoration. + +Five DuckDB cases, 34 total: HAVING on the selected aggregate, on a different +aggregate, fk-keyed with a conjunction, and two over a filter that empties +three groups (`HAVING COUNT(*) >= 0` and `HAVING SUM(amount) < 10000` with +the SUM sink carrying reachedness). Disable-verified red: +- reachedness off → both sparse cases; +- non-count emptiness off → the SUM-first sparse case; +- the executor routing `SumI32` through the coalescing kernel → both + mask-risc differentials. + +**Boundary refinement (same day, operator-raised):** the marker is free +inside the fold and a silent wrong answer outside it — any full-range +consumer would sum, sort or negate `i64::MIN`. Two consequences, both +landed: +- **Named, not implied.** ndarray exposes the reservation only under a + `_sym` suffix (`masked_group_sum_sym_i32{,_via}`, `SYM_EMPTY_I64`); every + unsuffixed reduction stays full two's-complement and never treats + `i64::MIN` specially (pinned by a test). mask-risc names the fold + `GroupFold::SumSymI32` with `GROUP_SUM_SYM_MAX_ROWS`. This is the 4-bit + analogue of choosing −7..+7 + NaN by name, never silently narrowing a + −8..+7 consumer. +- **NULL leaves as a mask.** `normalize_group_sink` is the quack boundary: + it returns a K-bit presence mask and zeroes absent slots in place; `finish` + runs it over every sink and returns `{present, keep}`. The raw sink is + internal encoding. Row NULL and group NULL now have the same shape. + Disable-verified: no zeroing → the leak assertion and unit tests fail; + normalizing only the first sink → the multi-sink unit test fails. + +**Two SUM spellings are a deliberate twin, not debt** (operator, same +day). Full-range `GroupSumI32` + `COUNT` and `_sym` `SumSymI32` answer the +same question two independent ways, so either can check the other whenever +something looks off. `sym_sum_agrees_with_full_range_sum_plus_count` keeps +that check permanent: presence ⇔ count ≠ 0, equal values where present, +full-range 0 where absent, on both key addresses. It is also the only check +that can see a real sum colliding with the reserved code. Disable-verified: +routing `SumSymI32` through the full-range kernel turns it red. + +REVISIT WHEN: a single-pass AVG is wanted — the fused sum+count sink is still +the route to that, now for speed only, not for NULL. + +## What is still open (not done here) +- **W-C:** `ORDER BY rid LIMIT n` — a first-n select, a rank/select member. +- **W-D:** multi-key `GROUP BY` via a fused composite address. +- **Debt:** `TD-SYM-SUM-MERGE-IS-NOT-ADDITION-1` — `_sym` partial sinks must + merge by the fold's rule, never `+`; dormant until a parallel reduction. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 5e561939d..7408a85aa 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,10 +25,11 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -147 entries, 2026-08-06 .. 2026-09-22. +148 entries, 2026-08-06 .. 2026-09-23. | date | entry id | finding | file | |---|---|---|---| +| 2026-09-23 | `quack-having-sym-sum-presence-mask` | | [2026-09-23-quack-having-sym-sum-presence-mask.md](2026-09-23-quack-having-sym-sum-presence-mask.md) | | 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) | | 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) | | 2026-09-22 | `E-CATS-FOLD-DOES-NOT-RETAIN-A-POPULATION-BITMAP-1` | CATS aggregate lowers to one tiled grouped terminal; bitmap realization is a requested boundary sink | [2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md](2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md) | diff --git a/crates/lance-graph-mask-risc/src/exec.rs b/crates/lance-graph-mask-risc/src/exec.rs index adc4f57b2..13940b677 100644 --- a/crates/lance-graph-mask-risc/src/exec.rs +++ b/crates/lance-graph-mask-risc/src/exec.rs @@ -30,10 +30,11 @@ use ndarray::simd::{ mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_xor, mask_xor_assign, masked_group_count_u32, masked_group_count_u32_via, masked_group_max_i32, masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_via, masked_group_sum_i32, - masked_group_sum_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32, - masked_sum_i32, ne_i32_to_mask, ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, - popcount_batch_u64, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, - ternary_match_u64_to_mask, ternary_match_u64_to_mask_under, KeyRunCarry, + masked_group_sum_i32_via, masked_group_sum_sym_i32, masked_group_sum_sym_i32_via, + masked_key_run_count_u32, masked_max_i32, masked_min_i32, masked_sum_i32, ne_i32_to_mask, + ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, popcount_batch_u64, + ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, ternary_match_u64_to_mask, + ternary_match_u64_to_mask_under, KeyRunCarry, }; use crate::ir::{ @@ -1042,6 +1043,21 @@ pub fn execute_into( o, ) } + (GroupKey::Lane(k), GroupFold::SumSymI32(v)) => masked_group_sum_sym_i32( + m, + lane_u32(planes, k, t), + lane_i32(planes, v, t), + o, + ), + (GroupKey::Via { fk, key }, GroupFold::SumSymI32(v)) => { + masked_group_sum_sym_i32_via( + m, + lane_u32(planes, fk, t), + foreign_lane_u32(foreign, key), + lane_i32(planes, v, t), + o, + ) + } } } } diff --git a/crates/lance-graph-mask-risc/src/ir.rs b/crates/lance-graph-mask-risc/src/ir.rs index 81c3ed46a..f931e5340 100644 --- a/crates/lance-graph-mask-risc/src/ir.rs +++ b/crates/lance-graph-mask-risc/src/ir.rs @@ -360,10 +360,14 @@ pub enum Terminal { /// identity ([`GroupFold::seed`]): `0` for a count, `i64::MAX` for a /// minimum, `i64::MIN` for a maximum. A MIN/MAX slot still holding its /// seed afterwards is a group no selected row named — the SQL `NULL` of - /// an empty group. No per-group sum is involved, so no row bound applies. + /// an empty group. Only [`GroupFold::SumSymI32`] carries a row bound + /// ([`GROUP_SUM_SYM_MAX_ROWS`]). /// - /// `SUM` keeps its own terminals ([`Terminal::GroupSumI32`] / - /// [`Terminal::GroupSumViaI32`]); this one does not repeat them. + /// The coalescing `SUM` (empty group reads `0`) keeps its own terminals + /// ([`Terminal::GroupSumI32`] / [`Terminal::GroupSumViaI32`]). The + /// NULL-preserving `SUM` lives here as [`GroupFold::SumSymI32`], because the + /// empty-group rule is the same for every seeded fold: a slot still + /// holding [`GroupFold::seed`] is empty ([`GroupFold::is_empty_slot`]). GroupReduce { mask: Operand, key: GroupKey, @@ -391,6 +395,17 @@ pub enum GroupFold { MinI32(u16), /// `MAX(lanes[val])` over an `I32` lane. MaxI32(u16), + /// `SUM(lanes[val])` over an `I32` lane in the SYMMETRIC range — the + /// `_sym` reading of `ndarray::simd`: real sums live in `±(2^63 − 1)` + /// and `ndarray::simd::SYM_EMPTY_I64` (`i64::MIN`) is reserved for "no + /// selected row reached this group". The first row a group sees REPLACES + /// the marker; later rows add. That keeps a group whose values cancel to + /// `0` distinct from an empty one, which the full-range + /// [`Terminal::GroupSumI32`] cannot tell apart. The reservation is + /// named, never implied: every other sum here is full range. Carries the + /// tighter [`GROUP_SUM_SYM_MAX_ROWS`]. The raw sink is internal encoding; + /// a consumer maps the marker away before treating slots as integers. + SumSymI32(u16), } impl GroupFold { @@ -402,10 +417,30 @@ impl GroupFold { GroupFold::Count => 0, GroupFold::MinI32(_) => i64::MAX, GroupFold::MaxI32(_) => i64::MIN, + GroupFold::SumSymI32(_) => ndarray::simd::SYM_EMPTY_I64, + } + } + + /// Whether a slot holding `v` after the fold is a group no selected row + /// reached — the SQL `NULL` of an empty group. The rule is uniform: + /// empty ⇔ the slot still holds [`GroupFold::seed`]. `COUNT` has no + /// empty groups: a zero count is a real answer, not a `NULL`. + pub const fn is_empty_slot(self, v: i64) -> bool { + match self { + GroupFold::Count => false, + _ => v == self.seed(), } } } +/// The widest plane a NULL-preserving [`GroupFold::SumSymI32`] is defined on: +/// `2^32 − 1` rows. One less than [`MASKED_SUM_I32_MAX_ROWS`] because the +/// seed doubles as the empty marker: exactly `2^32` rows of `i32::MIN` sum +/// to `i64::MIN`, a real value that would read back as `NULL`. Below that +/// row count every real sum lies strictly above `i64::MIN`, leaving the +/// symmetric range `±(2^63 − 1)` — closed under negation. +pub const GROUP_SUM_SYM_MAX_ROWS: usize = (1 << 32) - 1; + /// The widest plane [`Terminal::MaskedSumI32`] is defined on: `2^32` rows. /// The binding side is the NEGATIVE one: `2^32 · i32::MIN = −2^63 = i64::MIN` /// exactly, and one more row of `i32::MIN` wraps. The positive side has two diff --git a/crates/lance-graph-mask-risc/src/lib.rs b/crates/lance-graph-mask-risc/src/lib.rs index 54769f07f..1fc256061 100644 --- a/crates/lance-graph-mask-risc/src/lib.rs +++ b/crates/lance-graph-mask-risc/src/lib.rs @@ -124,7 +124,7 @@ pub use exec::{ pub use fuse::{fuse, fuse_program, ternlog_imm, BoolExpr, FuseError, Fused}; pub use ir::{ Foreign, ForeignPlane, GroupFold, GroupKey, LaneRef, MaskOp, Operand, Planes, Pred, Program, - Terminal, MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS, + Terminal, GROUP_SUM_SYM_MAX_ROWS, MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS, }; pub use reference::{ reference_execute, reference_execute_into, reference_scratch, reference_scratch_with_foreign, diff --git a/crates/lance-graph-mask-risc/src/reference.rs b/crates/lance-graph-mask-risc/src/reference.rs index 573b41bef..3738ccc7c 100644 --- a/crates/lance-graph-mask-risc/src/reference.rs +++ b/crates/lance-graph-mask-risc/src/reference.rs @@ -18,7 +18,7 @@ use crate::ir::{ Foreign, GroupFold, GroupKey, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal, - MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS, + GROUP_SUM_SYM_MAX_ROWS, MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS, }; use crate::value::{ExecError, LaneKind, Out, Value}; use crate::words_for; @@ -537,6 +537,12 @@ pub(crate) fn validate( GroupFold::MinI32(v) | GroupFold::MaxI32(v) => { check_lane(planes, v, LaneKind::I32)? } + GroupFold::SumSymI32(v) => { + check_lane(planes, v, LaneKind::I32)?; + if n > GROUP_SUM_SYM_MAX_ROWS { + return Err(ExecError::SumRowBound { n_rows: n }); + } + } } match out { OutShape::I64(len) if len >= 1 => Ok(()), @@ -905,6 +911,7 @@ pub fn reference_execute_into( GroupFold::Count => 0i64, GroupFold::MinI32(_) => i64::MAX, GroupFold::MaxI32(_) => i64::MIN, + GroupFold::SumSymI32(_) => i64::MIN, }; for x in o.iter_mut() { *x = seed; @@ -934,6 +941,15 @@ pub fn reference_execute_into( GroupFold::Count => o[k] + 1, GroupFold::MinI32(v) => o[k].min(i64::from(i32_at(planes, v, r))), GroupFold::MaxI32(v) => o[k].max(i64::from(i32_at(planes, v, r))), + // First row replaces the empty marker; later rows add. + GroupFold::SumSymI32(v) => { + let x = i64::from(i32_at(planes, v, r)); + if o[k] == i64::MIN { + x + } else { + o[k] + x + } + } }; } } diff --git a/crates/lance-graph-mask-risc/tests/foreign.rs b/crates/lance-graph-mask-risc/tests/foreign.rs index 69e227fdb..6c6bcb518 100644 --- a/crates/lance-graph-mask-risc/tests/foreign.rs +++ b/crates/lance-graph-mask-risc/tests/foreign.rs @@ -7,6 +7,7 @@ use lance_graph_mask_risc::reference::{reference_execute_into, reference_scratch use lance_graph_mask_risc::{ scratch_words_for, tile_words_for, words_for, ExecError, Foreign, ForeignPlane, GroupFold, GroupKey, LaneKind, LaneRef, MaskOp, Operand, Out, Planes, Pred, Program, Terminal, Value, + GROUP_SUM_SYM_MAX_ROWS, MASKED_SUM_I32_MAX_ROWS, }; fn lcg(seed: &mut u64) -> u64 { @@ -1037,7 +1038,12 @@ fn group_reduce_matches_the_oracle_for_every_key_and_fold() { multi_tile |= words < words_for(n); for key in [GroupKey::Lane(3), GroupKey::Via { fk: 0, key: 0 }] { - for fold in [GroupFold::Count, GroupFold::MinI32(2), GroupFold::MaxI32(2)] { + for fold in [ + GroupFold::Count, + GroupFold::MinI32(2), + GroupFold::MaxI32(2), + GroupFold::SumSymI32(2), + ] { let p = Program::new( vec![MaskOp::Pred { pred: Pred::NeU32 { lane: 1, v: 2 }, @@ -1086,6 +1092,172 @@ fn group_reduce_matches_the_oracle_for_every_key_and_fold() { assert!(multi_tile, "every case ran in a single tile"); } +/// FAILS IF: the NULL-preserving `SumI32` fold disagrees with the coalescing +/// `GroupSumI32` terminal on a group that selected rows reached, OR fails to +/// tell a group whose values cancel to `0` from a group no row reached. The +/// second half is the whole reason the fold exists: `GroupSumI32` reads both +/// as `0`. +#[test] +fn sym_sum_agrees_with_the_coalesced_sum_and_keeps_empty_groups_null() { + // Rows: group 0 gets +5 and -5 (cancels to 0), group 1 gets 7, group 2 + // is named by a row the mask drops, group 3 is never named. + let keys = [0u32, 0, 1, 2]; + let vals = [5i32, -5, 7, 100]; + let keep = [0b0111u64]; + let lanes = [LaneRef::U32(&keys), LaneRef::I32(&vals)]; + let masks: [&[u64]; 1] = [&keep]; + let planes = Planes { + n_rows: 4, + masks: &masks, + lanes: &lanes, + }; + let foreign = Foreign { + planes: &[], + lanes: &[], + }; + let run = |terminal: Terminal| { + let p = Program::new(vec![], terminal); + let mut out = vec![99i64; 4]; + reference_execute_into(&p, &planes, &foreign, Out::I64(&mut out)).expect("oracle runs"); + let mut got = vec![-99i64; 4]; + let words = tile_words_for(4); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + execute_into(&p, &planes, &foreign, &mut scratch, Out::I64(&mut got)).expect("runs"); + assert_eq!(got, out, "executor vs oracle for {terminal:?}"); + got + }; + let fold = GroupFold::SumSymI32(1); + let seeded = run(Terminal::GroupReduce { + mask: Operand::Plane(0), + key: GroupKey::Lane(0), + fold, + }); + let coalesced = run(Terminal::GroupSumI32 { + mask: Operand::Plane(0), + key: 0, + val: 1, + }); + assert_eq!(seeded, vec![0, 7, i64::MIN, i64::MIN]); + assert_eq!(coalesced, vec![0, 7, 0, 0]); + let empty: Vec = seeded.iter().map(|&v| fold.is_empty_slot(v)).collect(); + assert_eq!( + empty, + vec![false, false, true, true], + "cancelling group 0 must be present" + ); + for (g, (&s, &c)) in seeded.iter().zip(&coalesced).enumerate() { + if !empty[g] { + assert_eq!(s, c, "group {g}: a reached group sums the same either way"); + } + } + // COUNT never reports an empty slot: a zero count is an answer. + assert!(!GroupFold::Count.is_empty_slot(0)); + // The marker costs exactly one row of range. + assert_eq!(GROUP_SUM_SYM_MAX_ROWS + 1, MASKED_SUM_I32_MAX_ROWS); +} + +/// The standing side-by-side: the full-range pair (`GroupSumI32` + +/// `COUNT`) and the symmetric-range `SumSymI32` answer the same question two +/// independent ways, and must agree group for group: +/// +/// - present in `_sym` ⇔ count ≠ 0; +/// - present → `_sym` sum == full-range sum; +/// - absent → full-range sum == 0. +/// +/// The pair also catches the one failure `_sym` cannot see on its own: a real +/// sum landing on its reserved code (only possible at `2^32` rows) would read +/// as absent while the count says present. +/// +/// FAILS IF the two paths ever disagree, on either key address, at any length +/// including multi-tile. +#[test] +fn sym_sum_agrees_with_full_range_sum_plus_count() { + let groups = 4u32; + let mut saw_absent = false; + let mut saw_present = false; + for &n in &[1usize, 63, 64, 130, 1000, 5000] { + let fx = Fixture::new(n, 10, groups, 0x5EED ^ n as u64); + let (lanes, masks) = fx.planes(); + let planes = Planes { + n_rows: n, + masks: &masks, + lanes: &lanes, + }; + let mut s = 0xFACEu64 ^ n as u64; + // Group 3 is never produced by the remap: a real empty group on VIA. + let remap: Vec = (0..fx.foreign_rows) + .map(|_| lcg(&mut s) as u32 % 3) + .collect(); + let foreign_lanes = [LaneRef::U32(&remap)]; + let foreign = Foreign { + planes: &[], + lanes: &foreign_lanes, + }; + let filter = vec![MaskOp::Pred { + pred: Pred::NeU32 { lane: 1, v: 2 }, + under: None, + dst: 0, + }]; + let run = |terminal: Terminal| { + let p = Program::new(filter.clone(), terminal); + let words = tile_words_for(n); + let slots = p.scratch_slots as usize; + let mut buf = vec![0u64; scratch_words_for(words, slots).expect("sized")]; + let mut scratch = Scratch::over(&mut buf, words, slots).expect("carves"); + let mut out = vec![0x5a5a_i64; groups as usize]; + execute_into(&p, &planes, &foreign, &mut scratch, Out::I64(&mut out)).expect("runs"); + out + }; + for key in [GroupKey::Lane(3), GroupKey::Via { fk: 0, key: 0 }] { + let full = match key { + GroupKey::Lane(k) => run(Terminal::GroupSumI32 { + mask: S0, + key: k, + val: 2, + }), + GroupKey::Via { fk, key } => run(Terminal::GroupSumViaI32 { + mask: S0, + fk, + key, + val: 2, + }), + }; + let count = run(Terminal::GroupReduce { + mask: S0, + key, + fold: GroupFold::Count, + }); + let sym_fold = GroupFold::SumSymI32(2); + let sym = run(Terminal::GroupReduce { + mask: S0, + key, + fold: sym_fold, + }); + for g in 0..groups as usize { + let present = !sym_fold.is_empty_slot(sym[g]); + assert_eq!(present, count[g] != 0, "n={n} {key:?} g={g}: presence"); + if present { + assert_eq!(sym[g], full[g], "n={n} {key:?} g={g}: value"); + saw_present = true; + } else { + assert_eq!(full[g], 0, "n={n} {key:?} g={g}: absent reads 0 full-range"); + saw_absent = true; + } + } + } + } + assert!( + saw_present, + "no group was reached: the value half never ran" + ); + assert!( + saw_absent, + "no group was empty: the presence half never ran" + ); +} + /// FAILS IF: `GroupReduce` accepts a wrong-width lane or a missing sink — /// it must refuse before writing, like `GroupSumI32`. #[test] diff --git a/crates/lance-graph-quack/src/lib.rs b/crates/lance-graph-quack/src/lib.rs index 03cd6740a..9db93881f 100644 --- a/crates/lance-graph-quack/src/lib.rs +++ b/crates/lance-graph-quack/src/lib.rs @@ -755,6 +755,19 @@ pub enum LowerError { /// the WHOLE `out` slice, so K groups would leave the last group's blend /// and silently discard K − 1 — refused rather than answered wrongly. GroupedBlend, + /// A `HAVING` predicate names an aggregate index the plan does not + /// carry, or the plan carries no aggregate at all. + HavingAggOutOfRange { + /// The index named. + index: usize, + /// How many aggregates the plan carries. + aggs: usize, + }, + /// A grouped plan over an empty group universe (`groups == 0`). Every + /// sink program must run with an `Out::I64` of exactly `groups` slots, and + /// the executor refuses a zero-length sink — so the plan would be + /// unexecutable. Refused here, where the caller can see why. + EmptyGroupUniverse, } impl core::fmt::Display for LowerError { @@ -769,6 +782,12 @@ impl core::fmt::Display for LowerError { LowerError::GroupedBlend => { write!(f, "a blend writes the whole output; it cannot be grouped") } + LowerError::HavingAggOutOfRange { index, aggs } => { + write!(f, "HAVING names aggregate {index}; the plan carries {aggs}") + } + LowerError::EmptyGroupUniverse => { + write!(f, "a grouped plan needs at least one group (groups == 0)") + } } } } @@ -876,7 +895,9 @@ pub enum Agg { /// a key past it (or an fk naming no foreign row) is dropped. The /// executor seeds the buffer itself: `0` for a count, `i64::MAX` for a /// minimum and `i64::MIN` for a maximum, so a MIN/MAX slot still - /// holding its seed is an empty group — SQL `NULL`. + /// holding its seed is an empty group — SQL `NULL`. That raw buffer is + /// internal encoding: pass it through [`normalize_group_sink`] before any + /// code treats its slots as plain integers. GroupReduce { /// Where each row's group lives. key: GroupAddr, @@ -910,6 +931,24 @@ pub enum GroupAgg { MinI32(Col), /// `MAX(col)` over a signed column. MaxI32(Col), + /// `SUM(col)` over a signed column, NULL-preserving: an empty group + /// reads SQL `NULL`, never `0`, so a group whose values cancel stays + /// distinguishable from a group no row reached. This is the `SUM` a + /// `HAVING` must see; the coalescing `GROUP BY … SUM` stays + /// [`Agg::GroupSumI32`]. + SumI32(Col), +} + +impl GroupAgg { + /// The mask-risc fold this aggregate lowers to. + pub const fn fold(self) -> GroupFold { + match self { + GroupAgg::Count => GroupFold::Count, + GroupAgg::MinI32(c) => GroupFold::MinI32(c.0), + GroupAgg::MaxI32(c) => GroupFold::MaxI32(c.0), + GroupAgg::SumI32(c) => GroupFold::SumSymI32(c.0), + } + } } /// One query: a filter and what to ask of the rows that pass it. @@ -1212,6 +1251,265 @@ pub fn avg_finish(sum: i64, count: u64) -> Option { (count != 0).then(|| sum as f64 / count as f64) } +/// One `HAVING` comparison: aggregate `agg` (an index into +/// [`GroupHaving::aggs`]) against an `i64` constant. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum HavingCmp { + /// `agg > v` + Gt(i64), + /// `agg >= v` + Ge(i64), + /// `agg < v` + Lt(i64), + /// `agg <= v` + Le(i64), + /// `agg = v` + Eq(i64), + /// `agg <> v` + Ne(i64), +} + +impl HavingCmp { + const fn holds(self, x: i64) -> bool { + match self { + HavingCmp::Gt(v) => x > v, + HavingCmp::Ge(v) => x >= v, + HavingCmp::Lt(v) => x < v, + HavingCmp::Le(v) => x <= v, + HavingCmp::Eq(v) => x == v, + HavingCmp::Ne(v) => x != v, + } + } +} + +/// `SELECT key, aggs… FROM t WHERE filter GROUP BY key HAVING p₀ AND p₁ …`. +/// +/// HAVING is not a new primitive and not a population operation. Every +/// aggregate is one K-slot sink program over the SAME filter and key; HAVING +/// is the O(K) finalization over those sinks ([`GroupHavingPlan::finish`]), +/// which yields the demanded result — a K-bit mask of surviving groups. No +/// O(N) mask crosses a program boundary. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct GroupHaving { + /// The row filter (SQL `WHERE`). + pub filter: Filter, + /// Where each row's group lives. + pub key: GroupAddr, + /// The group universe K. + pub groups: u32, + /// The aggregates the query computes, in `SELECT` order. + pub aggs: Vec, + /// The `HAVING` conjunction: `(index into aggs, comparison)`. + pub having: Vec<(usize, HavingCmp)>, +} + +/// The lowered [`GroupHaving`]: one program per aggregate, each executed +/// with an `Out::I64` of exactly `groups` slots, then [`Self::finish`]. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct GroupHavingPlan { + /// One `Terminal::GroupReduce` program per aggregate, in `aggs` order. + pub programs: Vec, + /// The fold each program runs — what [`Self::finish`] reads emptiness by. + pub folds: Vec, + /// The `HAVING` conjunction, carried through. + pub having: Vec<(usize, HavingCmp)>, + /// The group universe K every sink is sized to. + pub groups: u32, +} + +/// Lower a [`GroupHaving`]. +/// +/// # Errors +/// +/// [`LowerError::EmptyGroupUniverse`] when `groups == 0`; +/// [`LowerError::HavingAggOutOfRange`] when `aggs` is empty or a predicate +/// names a missing aggregate; otherwise as [`lower`]. +pub fn lower_group_having(q: &GroupHaving) -> Result { + if q.groups == 0 { + return Err(LowerError::EmptyGroupUniverse); + } + if q.aggs.is_empty() { + return Err(LowerError::HavingAggOutOfRange { index: 0, aggs: 0 }); + } + if let Some(&(index, _)) = q.having.iter().find(|(i, _)| *i >= q.aggs.len()) { + return Err(LowerError::HavingAggOutOfRange { + index, + aggs: q.aggs.len(), + }); + } + let programs = q + .aggs + .iter() + .map(|&agg| { + lower(&Query { + filter: q.filter.clone(), + agg: Agg::GroupReduce { key: q.key, agg }, + }) + }) + .collect::, _>>()?; + Ok(GroupHavingPlan { + programs, + folds: q.aggs.iter().map(|a| a.fold()).collect(), + having: q.having.clone(), + groups: q.groups, + }) +} + +/// Why [`GroupHavingPlan::finish`] refused its sinks. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum HavingFinishError { + /// Not one sink per program. + SinkCount { + /// Programs in the plan. + expected: usize, + /// Sinks supplied. + found: usize, + }, + /// A sink is not exactly `groups` slots long. + SinkLen { + /// Which sink. + index: usize, + /// `groups`. + expected: usize, + /// Its length. + found: usize, + }, + /// The plan itself is inconsistent — only reachable for a plan built by + /// hand rather than by [`lower_group_having`], whose fields are public: + /// no programs, `folds` not one per program, or a `HAVING` index past the + /// aggregates. + MalformedPlan { + /// Which rule the plan breaks. + reason: &'static str, + }, +} + +impl core::fmt::Display for HavingFinishError { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + match self { + HavingFinishError::SinkCount { expected, found } => { + write!(f, "expected {expected} sinks, got {found}") + } + HavingFinishError::SinkLen { + index, + expected, + found, + } => write!(f, "sink {index} has {found} slots, expected {expected}"), + HavingFinishError::MalformedPlan { reason } => write!(f, "malformed plan: {reason}"), + } + } +} + +impl std::error::Error for HavingFinishError {} + +impl GroupHavingPlan { + /// Finish the query in place: map every sink through + /// [`normalize_group_sink`] and return which groups exist and which + /// survive the `HAVING` conjunction. + /// + /// A group survives when (a) at least one selected row reached it — SQL + /// `GROUP BY` never produces a group from no rows — and (b) every + /// `HAVING` comparison holds. Reachedness is one fact read off any sink, + /// because every program shares the filter and the key: it is taken from + /// the first. SQL's "a comparison against `NULL` is false" needs no + /// separate check: an unreached group is already gone, and a reached one + /// is non-empty in every sink. + /// + /// After this returns, every sink holds plain full-range integers: no + /// fold's internal marker survives, and an absent group's slot is `0`. + /// + /// # Errors + /// + /// [`HavingFinishError`] when the plan is malformed or the sinks do not + /// match its shape. + /// Nothing is modified in that case. + pub fn finish(&self, sinks: &mut [&mut [i64]]) -> Result { + // Every shape rule is checked BEFORE the first write, so an error + // really does leave the sinks untouched. + let malformed = |reason| Err(HavingFinishError::MalformedPlan { reason }); + if self.programs.is_empty() { + return malformed("no aggregate programs"); + } + if self.folds.len() != self.programs.len() { + return malformed("folds must be one per program"); + } + if self.having.iter().any(|&(i, _)| i >= self.folds.len()) { + return malformed("a HAVING index names a missing aggregate"); + } + if sinks.len() != self.programs.len() { + return Err(HavingFinishError::SinkCount { + expected: self.programs.len(), + found: sinks.len(), + }); + } + let k = self.groups as usize; + if let Some((index, s)) = sinks.iter().enumerate().find(|(_, s)| s.len() != k) { + return Err(HavingFinishError::SinkLen { + index, + expected: k, + found: s.len(), + }); + } + let present = normalize_group_sink(self.folds[0], sinks[0]); + for (fold, sink) in self.folds.iter().zip(sinks.iter_mut()).skip(1) { + normalize_group_sink(*fold, sink); + } + let mut keep = present.clone(); + for g in 0..k { + let bit = 1u64 << (g % 64); + if keep[g / 64] & bit != 0 + && !self.having.iter().all(|&(i, cmp)| cmp.holds(sinks[i][g])) + { + keep[g / 64] &= !bit; + } + } + Ok(GroupHavingOutput { present, keep }) + } +} + +/// What [`GroupHavingPlan::finish`] returns: two K-bit masks, bit `g` of word +/// `g / 64`, LSB first — the same bit order as every row mask. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct GroupHavingOutput { + /// Groups at least one selected row reached. + pub present: Vec, + /// Groups that are present AND pass every `HAVING` comparison. Always a + /// subset of `present`. + pub keep: Vec, +} + +/// The consumer boundary of a grouped sink. Returns a K-bit presence mask +/// (bit `g` set ⇔ some selected row reached group `g`) and rewrites the sink +/// IN PLACE so every slot is a plain full-range integer: an absent group's +/// slot becomes `0`. +/// +/// Why this exists: a `GroupReduce` sink is internal encoding. MIN/MAX leave +/// an empty group at a seed outside the `i32` range, and the symmetric-range +/// SUM ([`GroupAgg::SumI32`]) leaves it at `ndarray::simd::SYM_EMPTY_I64` — +/// a value that code assuming full range would read as −9.2·10¹⁸ and then +/// sum, sort, or negate (`-i64::MIN` overflows). NULL leaves this crate the +/// way row NULL already does: as a mask beside the values, never as a +/// reserved value inside them. It also works for a future fold with no spare +/// value to reserve. +/// +/// For `COUNT` a slot is present exactly when it is non-zero, and its value +/// is untouched. +pub fn normalize_group_sink(fold: GroupFold, sink: &mut [i64]) -> Vec { + let mut present = vec![0u64; sink.len().div_ceil(64)]; + for (g, slot) in sink.iter_mut().enumerate() { + let reached = match fold { + GroupFold::Count => *slot != 0, + other => !other.is_empty_slot(*slot), + }; + if reached { + present[g / 64] |= 1 << (g % 64); + } else { + *slot = 0; + } + } + present +} + /// Lower one categorical `GROUP BY` with semantic compression before any /// execution scheduling. /// @@ -1721,11 +2019,7 @@ fn terminal_of(agg: Agg, mask: Operand) -> Terminal { key: key.0, }, }, - fold: match agg { - GroupAgg::Count => GroupFold::Count, - GroupAgg::MinI32(c) => GroupFold::MinI32(c.0), - GroupAgg::MaxI32(c) => GroupFold::MaxI32(c.0), - }, + fold: agg.fold(), }, } } @@ -1765,6 +2059,149 @@ mod tests { Planes, Scratch, Value, }; + /// FAILS IF: the consumer boundary lets a fold's internal marker through, + /// or mistakes a real value for an empty group. The `_sym` SUM case is the + /// one a full-range consumer would misread: a group cancelling to `0` is + /// present, a group left at `SYM_EMPTY_I64` is absent and becomes `0`. + /// FAILS IF: `finish` normalizes only the sink it reads reachedness from. + /// Here COUNT is first and the `_sym` SUM second, so the SUM's markers + /// can only be cleared by the loop over the remaining sinks. + #[test] + fn finish_normalizes_every_sink_not_just_the_first() { + let plan = lower_group_having(&GroupHaving { + filter: Filter::cmp(Col(0), Cmp::EqU32(1)), + key: GroupAddr::Local(Col(1)), + groups: 3, + aggs: vec![GroupAgg::Count, GroupAgg::SumI32(Col(2))], + having: vec![(1, HavingCmp::Ge(0))], + }) + .expect("lowers"); + let e = GroupFold::SumSymI32(0).seed(); + let mut count = [2i64, 0, 1]; + let mut sum = [0i64, e, -4]; + let out = plan + .finish(&mut [&mut count[..], &mut sum[..]]) + .expect("shapes match"); + assert_eq!(out.present, vec![0b101]); + assert_eq!( + out.keep, + vec![0b001], + "group 0 sums to 0 (>= 0), group 2 to -4" + ); + assert_eq!(sum, [0, 0, -4], "the second sink's marker is gone"); + } + + /// FAILS IF a zero-group HAVING lowers to programs the executor would + /// refuse, instead of being refused at lowering with a named error. + #[test] + fn lower_group_having_refuses_an_empty_group_universe() { + let q = GroupHaving { + filter: Filter::cmp(Col(0), Cmp::EqU32(1)), + key: GroupAddr::Local(Col(1)), + groups: 0, + aggs: vec![GroupAgg::Count], + having: vec![], + }; + assert_eq!(lower_group_having(&q), Err(LowerError::EmptyGroupUniverse)); + // The same query with one group lowers: the refusal is about K alone. + assert!(lower_group_having(&GroupHaving { groups: 1, ..q }).is_ok()); + } + + /// FAILS IF `finish` panics on a hand-built plan, or modifies the sinks + /// before refusing one. Each case breaks exactly one shape rule; the + /// sink carries a `_sym` marker that a premature normalization would + /// have rewritten to 0. + #[test] + fn finish_refuses_a_malformed_plan_without_touching_the_sinks() { + let good = lower_group_having(&GroupHaving { + filter: Filter::cmp(Col(0), Cmp::EqU32(1)), + key: GroupAddr::Local(Col(1)), + groups: 2, + aggs: vec![GroupAgg::SumI32(Col(2))], + having: vec![(0, HavingCmp::Ge(0))], + }) + .expect("lowers"); + let e = GroupFold::SumSymI32(0).seed(); + let cases = [ + GroupHavingPlan { + programs: vec![], + folds: vec![], + having: vec![], + ..good.clone() + }, + // folds short, no HAVING: only the folds rule can refuse it + // (otherwise `folds[0]` panics). + GroupHavingPlan { + folds: vec![], + having: vec![], + ..good.clone() + }, + // folds long, no HAVING: only the folds rule can refuse it + // (otherwise it silently succeeds with a fold matching no sink). + GroupHavingPlan { + folds: vec![GroupFold::SumSymI32(2), GroupFold::Count], + having: vec![], + ..good.clone() + }, + GroupHavingPlan { + having: vec![(1, HavingCmp::Ge(0))], + ..good.clone() + }, + ]; + for (n, plan) in cases.iter().enumerate() { + let mut sink = [e, 5]; + let sinks: &mut [&mut [i64]] = if plan.programs.is_empty() { + &mut [] + } else { + &mut [&mut sink[..]] + }; + let got = plan.finish(sinks); + assert!( + matches!(got, Err(HavingFinishError::MalformedPlan { .. })), + "case {n}: {got:?}" + ); + assert_eq!(sink, [e, 5], "case {n}: sinks modified before refusal"); + } + // Silence half: the well-formed plan still finishes. + let mut sink = [e, 5]; + assert!(good.finish(&mut [&mut sink[..]]).is_ok()); + assert_eq!(sink, [0, 5]); + } + + #[test] + fn normalize_group_sink_maps_every_marker_to_a_presence_bit() { + let e = GroupFold::SumSymI32(0).seed(); + assert_eq!(e, i64::MIN, "the _sym marker"); + let mut sum = [0, e, 7, e]; + let p = normalize_group_sink(GroupFold::SumSymI32(0), &mut sum); + assert_eq!(p, vec![0b0101]); + assert_eq!(sum, [0, 0, 7, 0]); + + let mut min = [i64::MAX, -3]; + assert_eq!( + normalize_group_sink(GroupFold::MinI32(0), &mut min), + vec![0b10] + ); + assert_eq!(min, [0, -3]); + + // COUNT: present iff non-zero, value untouched. + let mut count = [0, 4]; + assert_eq!( + normalize_group_sink(GroupFold::Count, &mut count), + vec![0b10] + ); + assert_eq!(count, [0, 4]); + + // Word boundary: group 64 lands in the second word. + let mut wide = vec![e; 65]; + wide[64] = -1; + assert_eq!( + normalize_group_sink(GroupFold::SumSymI32(0), &mut wide), + vec![0, 1] + ); + assert!(!wide.contains(&e)); + } + const N: usize = 1000; const VALS: Col = Col(0); diff --git a/crates/lance-graph-quack/tests/duckdb/cases.tsv b/crates/lance-graph-quack/tests/duckdb/cases.tsv index 3c796fc98..c5b0a30b4 100644 --- a/crates/lance-graph-quack/tests/duckdb/cases.tsv +++ b/crates/lance-graph-quack/tests/duckdb/cases.tsv @@ -28,3 +28,8 @@ anti_join_not_country3 SELECT COUNT(*) FROM line l WHERE l.status=1 AND NOT EXIS count_col_discount_posted SELECT COUNT(discount) FROM line WHERE status=1 2260 sum_discount_posted SELECT SUM(discount) FROM line WHERE status=1 35397 avg_discount_posted SELECT AVG(discount) FROM line WHERE status=1 15.662389380530973 +having_sum_cc SELECT cost_center, SUM(amount) FROM line WHERE status=1 GROUP BY cost_center HAVING SUM(amount) > 2600000 ORDER BY cost_center 0:2930321;1:2840790;2:2760912;5:2827013 +having_min_by_count_cc SELECT cost_center, MIN(amount) FROM line WHERE status=1 GROUP BY cost_center HAVING COUNT(*) < 350 ORDER BY cost_center 3:-4997;4:-4904;6:-4941;7:-4937 +having_sparse_count_cc SELECT cost_center, COUNT(*) FROM line WHERE status=2 AND qty>48 GROUP BY cost_center HAVING COUNT(*) >= 0 ORDER BY cost_center 0:3;1:2;3:2;4:1;6:1 +having_sparse_sum_lt_cc SELECT cost_center, COUNT(*) FROM line WHERE status=2 AND qty>48 GROUP BY cost_center HAVING SUM(amount) < 10000 ORDER BY cost_center 1:2;4:1;6:1 +join_having_count_country SELECT p.country, SUM(l.amount) FROM line l JOIN partner p ON p.rid=l.partner_id WHERE l.status=1 GROUP BY p.country HAVING COUNT(*) >= 300 AND SUM(l.amount) > 2200000 ORDER BY p.country 0:4090182;2:5539383;6:2552339;7:2241753 diff --git a/crates/lance-graph-quack/tests/duckdb/oracle.py b/crates/lance-graph-quack/tests/duckdb/oracle.py index 49de4c16b..7db146852 100644 --- a/crates/lance-graph-quack/tests/duckdb/oracle.py +++ b/crates/lance-graph-quack/tests/duckdb/oracle.py @@ -42,6 +42,11 @@ "join_group_min_country", "group_avg_cc", "join_group_avg_country", + "having_sum_cc", + "having_min_by_count_cc", + "having_sparse_count_cc", + "having_sparse_sum_lt_cc", + "join_having_count_country", } BOOLEAN = {"exists_neg"} ROWS = {"rows_proj"} diff --git a/crates/lance-graph-quack/tests/duckdb_differential.rs b/crates/lance-graph-quack/tests/duckdb_differential.rs index 57545a9b4..1ea9124bd 100644 --- a/crates/lance-graph-quack/tests/duckdb_differential.rs +++ b/crates/lance-graph-quack/tests/duckdb_differential.rs @@ -36,8 +36,9 @@ use lance_graph_mask_risc::{ LaneRef, Out, Planes, Scratch, Terminal, Value, }; use lance_graph_quack::{ - avg_finish, lower, lower_avg, lower_group_avg, lower_group_by, Agg, Cmp, Col, Filter, - ForeignLane, GroupAddr, GroupAgg, GroupBy, Query, + avg_finish, lower, lower_avg, lower_group_avg, lower_group_by, lower_group_having, + normalize_group_sink, Agg, Cmp, Col, Filter, ForeignLane, GroupAddr, GroupAgg, GroupBy, + GroupHaving, HavingCmp, Query, }; use fixture::col::{ @@ -407,12 +408,14 @@ fn run_group_reduce( execute_into(&program, planes, foreign, &mut scratch, Out::I64(&mut out)).expect("runs"); let alloc_bytes_exec = alloc_delta(before); assert_eq!(value, Value::GroupReduced, "case {id}"); + // Through the consumer boundary: the raw sink never reaches the encoder. let empty_is_null = !matches!(fold, GroupFold::Count); + let present = normalize_group_sink(fold, &mut out); let encoded = out .iter() .enumerate() .map(|(k, &v)| { - if empty_is_null && v == fold.seed() { + if empty_is_null && (present[k / 64] >> (k % 64)) & 1 == 0 { format!("{k}:NULL") } else { format!("{k}:{v}") @@ -1421,3 +1424,179 @@ fn avg_discount_posted() { let actual = run_avg("avg_discount_posted", &planes, &filter, DISCOUNT); assert_case(&cases, "avg_discount_posted", &actual); } + +// --------------------------------------------------------------------- +// HAVING — O(K) finalization over the aggregate sinks. +// --------------------------------------------------------------------- + +/// `GROUP BY … HAVING` through [`lower_group_having`]: one K-slot sink per +/// aggregate, finished into a K-bit group mask; the surviving groups are +/// encoded `k:v` with `v` read from aggregate `select`. A group no selected +/// row reached is absent, exactly as DuckDB (no key series here) omits it. +fn run_group_having( + id: &str, + planes: &Planes<'_>, + foreign: &Foreign<'_>, + q: &GroupHaving, + select: usize, +) -> String { + let plan = lower_group_having(q).expect("lowers"); + let mut sinks: Vec> = plan + .programs + .iter() + .map(|program| { + let mut scratch = Scratch::for_program(program, planes.n_rows).expect("carves"); + let mut out = vec![0x5a5a_i64; plan.groups as usize]; + let v = execute_into(program, planes, foreign, &mut scratch, Out::I64(&mut out)) + .unwrap_or_else(|e| panic!("case {id}: {e:?}")); + assert_eq!(v, Value::GroupReduced, "case {id}"); + out + }) + .collect(); + let mut refs: Vec<&mut [i64]> = sinks.iter_mut().map(Vec::as_mut_slice).collect(); + let keep = plan.finish(&mut refs).expect("sinks match the plan").keep; + // The boundary guarantee: no fold's internal marker survives `finish`. + for (fold, sink) in plan.folds.iter().zip(&sinks) { + if !matches!(fold, GroupFold::Count) { + assert!( + !sink.contains(&fold.seed()), + "case {id}: {fold:?} leaked its marker past finish" + ); + } + } + (0..plan.groups as usize) + .filter(|&g| (keep[g / 64] >> (g % 64)) & 1 == 1) + .map(|g| format!("{g}:{}", sinks[select][g])) + .collect::>() + .join(";") +} + +fn posted() -> Filter { + Filter::cmp(STATUS, Cmp::EqU32(1)) +} + +/// A filter that leaves cost centers 2, 5 and 7 with no rows at all. +fn sparse() -> Filter { + Filter::and([ + Filter::cmp(STATUS, Cmp::EqU32(2)), + Filter::cmp(QTY, Cmp::GtI32(48)), + ]) +} + +/// `HAVING SUM(amount) > 2600000` — the predicate on the selected aggregate. +#[test] +fn having_sum_cc() { + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let q = GroupHaving { + filter: posted(), + key: GroupAddr::Local(COST_CENTER), + groups: 8, + aggs: vec![GroupAgg::SumI32(AMOUNT)], + having: vec![(0, HavingCmp::Gt(2_600_000))], + }; + let actual = run_group_having("having_sum_cc", &lanes.planes(), &Foreign::NONE, &q, 0); + assert_case(&load_cases(), "having_sum_cc", &actual); +} + +/// `SELECT MIN(amount) … HAVING COUNT(*) < 350` — the predicate on a +/// DIFFERENT aggregate from the one selected. +#[test] +fn having_min_by_count_cc() { + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let q = GroupHaving { + filter: posted(), + key: GroupAddr::Local(COST_CENTER), + groups: 8, + aggs: vec![GroupAgg::MinI32(AMOUNT), GroupAgg::Count], + having: vec![(1, HavingCmp::Lt(350))], + }; + let actual = run_group_having( + "having_min_by_count_cc", + &lanes.planes(), + &Foreign::NONE, + &q, + 0, + ); + assert_case(&load_cases(), "having_min_by_count_cc", &actual); +} + +/// `HAVING COUNT(*) >= 0` over a sparse filter. The predicate holds for +/// EVERY count, so the only thing dropping cost centers 2, 5 and 7 is the +/// reachedness rule: SQL makes no group from no rows. +#[test] +fn having_sparse_count_cc() { + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let q = GroupHaving { + filter: sparse(), + key: GroupAddr::Local(COST_CENTER), + groups: 8, + aggs: vec![GroupAgg::Count], + having: vec![(0, HavingCmp::Ge(0))], + }; + let actual = run_group_having( + "having_sparse_count_cc", + &lanes.planes(), + &Foreign::NONE, + &q, + 0, + ); + assert_case(&load_cases(), "having_sparse_count_cc", &actual); +} + +/// `SELECT COUNT(*) … HAVING SUM(amount) < 10000` over the sparse filter. +/// The SUM sink is FIRST, so it alone carries reachedness: a coalescing SUM +/// would read the empty groups 2, 5, 7 as `0 < 10000` and admit them; the +/// NULL-preserving SUM leaves them at the seed, and they are dropped. +#[test] +fn having_sparse_sum_lt_cc() { + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let q = GroupHaving { + filter: sparse(), + key: GroupAddr::Local(COST_CENTER), + groups: 8, + aggs: vec![GroupAgg::SumI32(AMOUNT), GroupAgg::Count], + having: vec![(0, HavingCmp::Lt(10_000))], + }; + let actual = run_group_having( + "having_sparse_sum_lt_cc", + &lanes.planes(), + &Foreign::NONE, + &q, + 1, + ); + assert_case(&load_cases(), "having_sparse_sum_lt_cc", &actual); +} + +/// `GROUP BY p.country HAVING COUNT(*) >= 300` — the fk-keyed group. +#[test] +fn join_having_count_country() { + let fx = fixture::generate(); + let lanes = fx.line.lanes(); + let country_lane = [LaneRef::U32(&fx.partner.country)]; + let foreign = Foreign { + planes: &[], + lanes: &country_lane, + }; + let q = GroupHaving { + filter: posted(), + key: GroupAddr::Via { + fk: PARTNER_ID, + key: ForeignLane(0), + }, + groups: 8, + aggs: vec![GroupAgg::Count, GroupAgg::SumI32(AMOUNT)], + having: vec![(0, HavingCmp::Ge(300)), (1, HavingCmp::Gt(2_200_000))], + }; + let actual = run_group_having( + "join_having_count_country", + &lanes.planes(), + &foreign, + &q, + 1, + ); + assert_case(&load_cases(), "join_having_count_country", &actual); +}