Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# 2026-09-22 — Quack/DuckDB parity by T0 folding: the keyed-reduction family, and the primitive basis

**Status:** MEASURED (W-B and W-A shipped on this branch) · OPEN (W-C, W-D)
**Depends on:** ndarray #320 (keyed-reduction family), lance-graph #1262 (fold seam)

## What landed (W-B)

`GROUP BY` with `COUNT` / `MIN` / `MAX` now folds before dispatch, as `SUM`
already did (#1262). One program ends in `Terminal::GroupReduce { mask, key:
GroupKey::{Lane, Via}, fold: GroupFold::{Count, MinI32, MaxI32} }` and writes
a K-slot `Out::I64` sink. The previous shape was K+1 programs plus an O(N)
kept mask.

- `GroupFold::seed()` is 0 for COUNT, `i64::MAX` for MIN and `i64::MIN` for
MAX. All three are outside the i32 value range, so a slot still holding its
seed is exactly an empty group, i.e. SQL `NULL`. The harness and the #1262
re-pin both assert this correspondence rather than assuming it.
- DuckDB differential: 21/21 match DuckDB 1.5.5. The 5 new cases are
`group_min_cc`, `group_max_cc`, `group_max_cc_sparse` (3 empty groups →
`NULL`), `join_group_count_country` and `join_group_min_country`. The
existing 16 expected values did not move.
- `Agg::Rows` (a grouped projection) has no keyed-reduction law and stays
forest residue.
- Disable-verified, each turning red:
1. quack declining to fold COUNT
2. resident MIN dispatched to MAX
3. the harness printing the seed instead of `NULL`

## The primitive basis — reuse like the polyfill, expose like lgj-abi

In ndarray, `simd.rs` is a facade of names over backends. Parity work
reuses that same shape: **a small set of FAMILIES, each one private walker
with named public instances.** It does not add one kernel per SQL feature.
The keyed-reduction family is the worked example. There is one private
`group_walk(mask, n, GroupKeyAddr::{Resident, Via}, out, fold)`, and eight
public names are closures over it (sum/count/min/max × resident/via). A ninth
fold (e.g. `ANY`) is one closure and one name. It is not a new kernel.

| family | T0 facade (ndarray) | mask-risc | lgj-abi verb (T2) |
|---|---|---|---|
| predicate: type × cmp, `_under` gate | `{eq,ne,lt,le,gt,ge,ternary_match}_{i32,u32,u64,u8}_to_mask[_under]`, `eq_u32_via_to_mask` | `Pred::*` | `lgj_op_eq_u32`, `lgj_op_gt_i32`, `lgj_op_ternary_match` |
| mask algebra | `mask_{and,or,andnot,xor,not,ternlog}[_assign]`, `mask_set_range` | `MaskOp::*` | `lgj_mask_{and,or,andnot,ternlog}` |
| scalar reduction | `popcount_batch_u64`, `mask_{any,all}`, `masked_{sum,min,max}_i32` | `Terminal::{Count,Any,All,MaskedSum/Min/Max}` | `lgj_mask_count`, `lgj_reduce_*` |
| address movers | `mask_gather_u32`, `mask_scatter_or_u32`, `*_via` | `MaskOp::Gather`, `Terminal::ScatterOrU32` | `lgj_hop` |
| keyed reduction | `masked_group_{sum,count,min,max}_{i32,u32}[_via]` | `GroupSumI32`, `GroupSumViaI32`, `GroupReduce` | `lgj_plan_group_sum_i32` (SUM only) |
| rank / run / range | `masked_key_run_count_u32`, `mask_set_range` | `CountKeyRunsU32`, `Pred::Range` | — |

The lgj-abi column is the ABI analog of the facade column. `plan_lower.rs`
already carries `GroupKey::{Local, Via}`, the same address split as
`GroupKeyAddr`. The follow-up is to widen `lgj_plan_group_sum_i32` into a
`GroupReduce`-shaped plan verb (fold as a parameter). It should not become
three more exports. Java keeps seeing only the boring `sql()` / `sumBy()`
surface; it never names a fold.

## W-A landed — no new primitive, as predicted

Eight more cases, 29/29 against DuckDB 1.5.5, existing values unchanged:
- `SUM(CASE WHEN p THEN x ELSE 0 END)` is the plain masked SUM with `p` in
the mask.
- `AVG` is `lower_avg` / `lower_group_avg`: SUM and COUNT over the same
filter, finished by `avg_finish` outside the hot path. The float digits
match DuckDB exactly (scalar, resident-keyed and via-keyed). It costs two
passes; a fused sum+count sink would make it one.
- `NOT EXISTS` is `Not` around the factored fk predicate (`EqU32Via`), one
program. The earlier prediction ("`mask_andnot` after a hop") was wrong in
shape: no hop and no second mask are needed for the many-to-one direction.
The one-to-many direction (docs with no posted line) still needs the
ordered-key projection, same as `COUNT DISTINCT`.
- NULLs are a validity plane ANDed into the filter; `COUNT(col)`,
`SUM(col)` and `AVG(col)` follow. The fixture gained one derived nullable
column (no RNG draw).

Disable-verified: an off-by-one AVG denominator, a dropped `Not`, a dropped
validity plane, and NULL written as `0` each turned their cases red.

## What is still open (not done here)
- **W-C:**
- `HAVING` over the K-sized sink. Correction to the plan above: filtering
K result slots is finalization over an O(K) sink, like `avg_finish`, so
it needs NO SIMD primitive at realistic K. What it DOES need is **group
existence**, and that is OPEN:
- A MIN/MAX slot at its seed is an empty group, so `HAVING MIN(x) > c`
must skip it. A naive filter on `i64::MAX` would wrongly KEEP it.
- A SUM slot cannot tell an empty group from a group summing to 0.
SQL gives the empty group `NULL` (excluded by any `HAVING`), not 0.
- Both are solved by carrying a count beside the value. That is the
same fused sum+count sink AVG wants to become one pass instead of
two, so one keyed-reduction member closes two gaps. Decide that sink's
shape before writing `HAVING`, not after.
- `ORDER BY rid LIMIT n`. This needs a first-n select, a rank/select member.
- **W-D:** multi-key `GROUP BY` via a fused composite address. This is a
third `GroupKeyAddr` variant; the walker is unchanged.
- **Out of scope for T0 parity:** arbitrary-value sort, m:n hash joins,
window functions, strings.
- **Downstream:** lance-graph CI resolves ndarray via the local path dep, so
this branch goes green only after ndarray #320 merges.
3 changes: 2 additions & 1 deletion .claude/board/entries/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,10 +25,11 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately
opposite directions; the stranding this convention prevents shows up in
exactly one of them, never both.

146 entries, 2026-08-06 .. 2026-09-22.
147 entries, 2026-08-06 .. 2026-09-22.

| date | entry id | finding | file |
|---|---|---|---|
| 2026-09-22 | `quack-duckdb-parity-t0-keyed-reduction` | | [2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md](2026-09-22-quack-duckdb-parity-t0-keyed-reduction.md) |
| 2026-09-22 | `E-W0C-THE-ROW-BRIDGE-IS-A-DIALECT-NOT-AN-INTERPRETER-1` | a merged relational op carried as loco program data reaches the fused executor with no population crossing; the enum explosion is upstream of mask-risc | [2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md](2026-09-22-e-w0c-the-row-bridge-is-a-dialect-not-an-interpreter-1.md) |
| 2026-09-22 | `E-CATS-FOLD-DOES-NOT-RETAIN-A-POPULATION-BITMAP-1` | CATS aggregate lowers to one tiled grouped terminal; bitmap realization is a requested boundary sink | [2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md](2026-09-22-e-cats-fold-does-not-retain-a-population-bitmap-1.md) |
| 2026-08-31 | `E-Q8-THE-SIX-DOES-NO-WORK-A-DEGREE-ABLATION-COLLAPSES-THE-HEX-OVERLAYS-ENTIRE-ADVANTAGE-1` | B passes every pre-registered gate and the pass is unattributable: at degree 1 it scores identically with 5.5× less memory | [2026-08-31-e-q8-the-six-does-no-work-a-degree-ablation-collapses-the-hex-overlays-entire-advantage-1.md](2026-08-31-e-q8-the-six-does-no-work-a-degree-ablation-collapses-the-hex-overlays-entire-advantage-1.md) |
Expand Down
67 changes: 61 additions & 6 deletions crates/lance-graph-mask-risc/src/exec.rs
Original file line number Diff line number Diff line change
Expand Up @@ -28,15 +28,17 @@ use ndarray::simd::{
le_i32_to_mask, le_i32_to_mask_under, lt_i32_to_mask, lt_i32_to_mask_under, mask_all, mask_and,
mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32, mask_not,
mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_xor,
mask_xor_assign, masked_group_sum_i32, masked_group_sum_i32_via, masked_key_run_count_u32,
masked_max_i32, masked_min_i32, masked_sum_i32, ne_i32_to_mask, ne_i32_to_mask_under,
ne_u32_to_mask, ne_u32_to_mask_under, popcount_batch_u64, ternary_match_u32_to_mask,
ternary_match_u32_to_mask_under, ternary_match_u64_to_mask, ternary_match_u64_to_mask_under,
KeyRunCarry,
mask_xor_assign, masked_group_count_u32, masked_group_count_u32_via, masked_group_max_i32,
masked_group_max_i32_via, masked_group_min_i32, masked_group_min_i32_via, masked_group_sum_i32,
masked_group_sum_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32,
masked_sum_i32, ne_i32_to_mask, ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under,
popcount_batch_u64, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under,
ternary_match_u64_to_mask, ternary_match_u64_to_mask_under, KeyRunCarry,
};

use crate::ir::{
Foreign, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal, MAX_SCRATCH_SLOTS,
Foreign, GroupFold, GroupKey, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal,
MAX_SCRATCH_SLOTS,
};
use crate::reference::{out_shape, validate};
use crate::ternlog_dispatch::{ternlog_dispatch, ternlog_dispatch_assign};
Expand Down Expand Up @@ -738,6 +740,7 @@ pub fn execute_into(
o.fill(0)
}
(Terminal::GroupSumI32 { .. } | Terminal::GroupSumViaI32 { .. }, Out::I64(o)) => o.fill(0),
(Terminal::GroupReduce { fold, .. }, Out::I64(o)) => o.fill(fold.seed()),
_ => {}
}
let n_rows = planes.n_rows;
Expand Down Expand Up @@ -991,6 +994,57 @@ pub fn execute_into(
);
}
}
Terminal::GroupReduce { mask, key, fold } => {
// `validate` already refused a missing/too-small `out` and
// every wrong-width lane; one delegation per tile (law L3),
// the sink seeded above with the fold's identity.
if let Out::I64(o) = &mut out {
let m = read(planes, &slots, mask, t);
match (key, fold) {
(GroupKey::Lane(k), GroupFold::Count) => {
masked_group_count_u32(m, lane_u32(planes, k, t), o)
}
(GroupKey::Via { fk, key }, GroupFold::Count) => {
masked_group_count_u32_via(
m,
lane_u32(planes, fk, t),
foreign_lane_u32(foreign, key),
o,
)
}
(GroupKey::Lane(k), GroupFold::MinI32(v)) => masked_group_min_i32(
m,
lane_u32(planes, k, t),
lane_i32(planes, v, t),
o,
),
(GroupKey::Via { fk, key }, GroupFold::MinI32(v)) => {
masked_group_min_i32_via(
m,
lane_u32(planes, fk, t),
foreign_lane_u32(foreign, key),
lane_i32(planes, v, t),
o,
)
}
(GroupKey::Lane(k), GroupFold::MaxI32(v)) => masked_group_max_i32(
m,
lane_u32(planes, k, t),
lane_i32(planes, v, t),
o,
),
(GroupKey::Via { fk, key }, GroupFold::MaxI32(v)) => {
masked_group_max_i32_via(
m,
lane_u32(planes, fk, t),
foreign_lane_u32(foreign, key),
lane_i32(planes, v, t),
o,
)
}
}
}
}
Terminal::Keep { mask } => {
// The demanded mask, one tile at a time. With `Out::None` the
// scratch is single-tile (checked above) and the slot IS the
Expand Down Expand Up @@ -1018,6 +1072,7 @@ pub fn execute_into(
}),
Terminal::CountKeyRunsU32 { .. } => Value::Count(runs + run_carry.finish()),
Terminal::GroupSumI32 { .. } | Terminal::GroupSumViaI32 { .. } => Value::GroupSummed,
Terminal::GroupReduce { .. } => Value::GroupReduced,
Terminal::Keep { mask } => Value::Mask(mask),
})
}
Expand Down
59 changes: 59 additions & 0 deletions crates/lance-graph-mask-risc/src/ir.rs
Original file line number Diff line number Diff line change
Expand Up @@ -346,6 +346,64 @@ pub enum Terminal {
key: u16,
val: u16,
},
/// The rest of the keyed-reduction family — `COUNT(*)`, `MIN(v)`,
/// `MAX(v)` … `GROUP BY` — as ONE terminal parameterised by where each
/// row's group lives ([`GroupKey`]) and what is folded into it
/// ([`GroupFold`]). For every row `i` where `mask` holds, resolves the
/// row's group and folds it into the caller's `Out::I64` buffer, whose
/// length IS the group universe `K`. Delegates tile by tile to the
/// `ndarray::simd::masked_group_{count,min,max}` family; a key past the
/// universe, and (for [`GroupKey::Via`]) an fk naming no foreign row,
/// drop the row rather than erroring.
///
/// The executor seeds the sink before the first tile with the fold's
/// identity ([`GroupFold::seed`]): `0` for a count, `i64::MAX` for a
/// minimum, `i64::MIN` for a maximum. A MIN/MAX slot still holding its
/// seed afterwards is a group no selected row named — the SQL `NULL` of
/// an empty group. No per-group sum is involved, so no row bound applies.
///
/// `SUM` keeps its own terminals ([`Terminal::GroupSumI32`] /
/// [`Terminal::GroupSumViaI32`]); this one does not repeat them.
GroupReduce {
mask: Operand,
key: GroupKey,
fold: GroupFold,
},
}

/// Where a [`Terminal::GroupReduce`] reads each row's group.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum GroupKey {
/// The group of row `i` is `lanes[lane][i]`, a `U32` lane of this table.
Lane(u16),
/// The group of row `i` is `foreign.lanes[key][lanes[fk][i]]` — `fk` a
/// `U32` lane of this table, `key` a `U32` lane of the foreign table.
/// The two hops are fused; no remapped key lane is materialised.
Via { fk: u16, key: u16 },
}

/// What a [`Terminal::GroupReduce`] folds into each group's slot.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum GroupFold {
/// `COUNT(*)`: each selected row adds 1.
Count,
/// `MIN(lanes[val])` over an `I32` lane.
MinI32(u16),
/// `MAX(lanes[val])` over an `I32` lane.
MaxI32(u16),
}

impl GroupFold {
/// The fold's identity — what every slot holds before the first row.
/// For MIN/MAX it lies outside the `i32` range, so it doubles as the
/// empty-group marker.
pub const fn seed(self) -> i64 {
match self {
GroupFold::Count => 0,
GroupFold::MinI32(_) => i64::MAX,
GroupFold::MaxI32(_) => i64::MIN,
}
}
}

/// The widest plane [`Terminal::MaskedSumI32`] is defined on: `2^32` rows.
Expand Down Expand Up @@ -460,6 +518,7 @@ impl Program {
| Terminal::CountKeyRunsU32 { mask, .. }
| Terminal::GroupSumI32 { mask, .. }
| Terminal::GroupSumViaI32 { mask, .. }
| Terminal::GroupReduce { mask, .. }
| Terminal::Keep { mask } => touch(mask),
}
Self {
Expand Down
4 changes: 2 additions & 2 deletions crates/lance-graph-mask-risc/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -123,8 +123,8 @@ pub use exec::{
};
pub use fuse::{fuse, fuse_program, ternlog_imm, BoolExpr, FuseError, Fused};
pub use ir::{
Foreign, ForeignPlane, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal,
MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS,
Foreign, ForeignPlane, GroupFold, GroupKey, LaneRef, MaskOp, Operand, Planes, Pred, Program,
Terminal, MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS,
};
pub use reference::{
reference_execute, reference_execute_into, reference_scratch, reference_scratch_with_foreign,
Expand Down
71 changes: 69 additions & 2 deletions crates/lance-graph-mask-risc/src/reference.rs
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,8 @@
//! path a consumer should reach for.

use crate::ir::{
Foreign, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal, MASKED_SUM_I32_MAX_ROWS,
MAX_SCRATCH_SLOTS,
Foreign, GroupFold, GroupKey, LaneRef, MaskOp, Operand, Planes, Pred, Program, Terminal,
MASKED_SUM_I32_MAX_ROWS, MAX_SCRATCH_SLOTS,
};
use crate::value::{ExecError, LaneKind, Out, Value};
use crate::words_for;
Expand Down Expand Up @@ -522,6 +522,31 @@ pub(crate) fn validate(
}
}
}
Terminal::GroupReduce { mask, key, fold } => {
check_operand(p, planes, mask)?;
written_slots.readable(mask)?;
match key {
GroupKey::Lane(k) => check_lane(planes, k, LaneKind::U32)?,
GroupKey::Via { fk, key } => {
check_lane(planes, fk, LaneKind::U32)?;
check_foreign_lane(foreign, key, LaneKind::U32)?;
}
}
match fold {
GroupFold::Count => {}
GroupFold::MinI32(v) | GroupFold::MaxI32(v) => {
check_lane(planes, v, LaneKind::I32)?
}
}
match out {
OutShape::I64(len) if len >= 1 => Ok(()),
OutShape::None | OutShape::I32(_) | OutShape::I64(_) | OutShape::Mask(_) => {
Err(ExecError::TerminalNeedsOut {
what: "GroupReduce",
})
}
}
}
}
}

Expand Down Expand Up @@ -872,6 +897,48 @@ pub fn reference_execute_into(
}
Value::GroupSummed
}
Terminal::GroupReduce { mask, key, fold } => {
if let Out::I64(o) = out {
// Independent formulation: seed every slot, then walk the
// survivors one row at a time — no ndarray kernel involved.
let seed = match fold {
GroupFold::Count => 0i64,
GroupFold::MinI32(_) => i64::MAX,
GroupFold::MaxI32(_) => i64::MIN,
};
for x in o.iter_mut() {
*x = seed;
}
let remap = match key {
GroupKey::Via { key, .. } => match foreign.lanes.get(usize::from(key)) {
Some(LaneRef::U32(v)) => &v[..],
_ => &[][..],
},
GroupKey::Lane(_) => &[][..],
};
for r in survivors(mask) {
let k = match key {
GroupKey::Lane(lane) => u32_at(planes, lane, r) as usize,
GroupKey::Via { fk, .. } => {
let idx = u32_at(planes, fk, r) as usize;
if idx >= remap.len() {
continue;
}
remap[idx] as usize
}
};
if k >= o.len() {
continue;
}
o[k] = match fold {
GroupFold::Count => o[k] + 1,
GroupFold::MinI32(v) => o[k].min(i64::from(i32_at(planes, v, r))),
GroupFold::MaxI32(v) => o[k].max(i64::from(i32_at(planes, v, r))),
};
}
}
Value::GroupReduced
}
Terminal::Keep { mask } => {
if let Out::Mask(o) = out {
for w in o.iter_mut() {
Expand Down
4 changes: 4 additions & 0 deletions crates/lance-graph-mask-risc/src/value.rs
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,10 @@ pub enum Value {
/// [`crate::Terminal::GroupSumI32`]: the caller's `Out::I64` buffer was
/// written, one slot per group.
GroupSummed,
/// [`crate::Terminal::GroupReduce`]: the caller's `Out::I64` buffer was
/// written, one slot per group; see [`crate::GroupFold::seed`] for what
/// an empty group holds.
GroupReduced,
}

/// The caller's terminal-result destination — one variant per shape a
Expand Down
Loading
Loading