diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 3c59d8f..98d593d 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,35 @@ +## 2026-09-22 — minor 12: `lgj_plan_group_sum_i32`, the grouped sum that never builds a selection (fold-distillation wave 4) + +Upstream first (lance-graph #1256 merged `99cdca38`: the tiled executor, +`Terminal::GroupSumI32`/`GroupSumViaI32`, `Out::I64`; ndarray #318 merged: +the T1 kernels; lance-graph-java #83 merged: the status arms). This is the +Java→Panama→mask-risc proof that wave: a `GROUP BY` crosses once and no +selection exists at any point. + +- **ABI:** ONE new symbol, no new status, no manifest growth — the plan + surface is reused; only the terminal changed (`Keep` → `GroupSumI32`, or + `GroupSumViaI32` when `via_res != 0`). `n_ops == 0` is legal here (the + whole-lane `Range`, the one `Range` this crate emits, pinned to be exactly + `0..n_rows`). `docs/abi.md` §20. +- **Measured through the membrane:** `View.sumByGroup` = **1 crossing** at + 1,024 and 65,536 rows for 16 groups; the two-crossing-per-group path it + replaces measured **32** beside it in the same run (the `bricks` number, + reproduced). `sumByGroupVia` (the fk-keyed form, `SUM(line) GROUP BY + partner.key`) also 1. +- **Gates:** native **191** tests (+8), clippy `-D warnings` + fmt clean; + Java **612** checks (409 + `GroupSumTest` 203) under JDK 28 + `--enable-preview` against `abi 0.12`; six disable arms red-then-green + (§20.6). `OldAbiCompatTest` gains a minor-12 leg, run BOTH ways: 13/13 + against this library, and against a minor-11 library built from `07e044f` + `View.sumByGroup` throws `AbiMismatchException` naming minor 12 — never a + missing-symbol failure. +- **The eighth named materialisation site:** `Engine.groupSumI32`'s + `toArray`, sized by `groups` (the question), pinned in `DoctrineFenceTest` + and listed in root `CLAUDE.md`. `GroupTotals` exposes no array. +- **Still owed:** the `bricks` consumer's `sumBy()` still runs the 32-crossing + path; migrating it to `sumByGroup` is a consumer-wave change, not this + one. `lgj_hop` still holds mask-sized Vecs (pre-existing, unrelated). + ## 2026-09-19 — PR #81 merged (`07aa441`): production IS the Valhalla arm; Panama × Valhalla is one membrane **The frame, because it is easy to file this wrong:** this was not a JDK diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index ad272d7..1da2a74 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -118,3 +118,18 @@ layout wired end to end. Doctrine: `E-LGJ-THE-MIDDLE-TIER-IS-DELETED-NOT-WRAPPED | D-LGJ-W5 | Three consumer examples (trades / bricks / graph) — one plan file each | **trades DONE 2026-08-17** — `consumers/trades/` (own compile unit, core consumed as a third-party would): `Trade` (schema-not-entity: zero public ctors, zero instance fields, reflection-forced construction still throws), `World.open` → the existing lazy `View` under domain names, zero new membrane surface. TradesParityTest 12/12 (chain vs transcribed-generator recomputation at 1K+64K rows; 0 crossings composing / 1 at terminal THROUGH the domain vocabulary; reflection guard). TradesAllocationTest 3/3 — **the poster's number, measured: 240 bytes/query, IDENTICAL at 64K and 1M rows** (row-count independence is the thesis assertion; 64 KiB absolute backstop). Disable-run: VENUE pointed at the wrong lane → the membrane's own LANE_KIND_MISMATCH rejected it (the binding is checked, not trusted); restored green. **bricks DONE 2026-08-17** — `consumers/bricks/` (2 Sonnet workers K1/K2 per `.claude/waves/wave-consumer-bricks.md`): mask-first RBAC where `authorize(Role)` is a real natively-evaluated predicate in the SAME lazy chain as `where(...)` (`Role.EU_ONLY` = `REGION.eq(EU)`, `DENY_ALL` = `REGION.eq(0xFFFF)` — a genuine impossible predicate, not a Java branch), fail-closed (`UnauthorizedQueryException` BEFORE any crossing; no default-allow path exists), aggregate-only egress (every public method returns `BricksQuery`/`long`/`Map` — structurally no row-shaped type). BricksAuthTest **62/62**: parity vs transcribed generator at 1K+64K; RBAC-as-predicate equivalence (EU_ONLY result == GLOBAL+explicit-where); DENY_ALL counts 0 while paying a real crossing; crossing arithmetic — count()=1, sumBy()=**32 crossings (16 groups × 2: plan_eval + lgj_reduce_sum_i32), IDENTICAL at both row counts** (the thesis: crossings ∝ groups, never rows — the measured 32 corrected K1's "1 per group" Javadoc claim, a real finding about sum-terminal cost); reflection guards. Disable-run: `requireAuthorized` short-circuited → **exactly the 3 can-fire fail-closed checks red, 59 green**; restored, 62/62. Core suite unaffected (188/188). **graph DONE 2026-08-18** — `consumers/graph/` (2 Sonnet workers G1/G2 per `.claude/waves/wave-consumer-graph.md`, dispatched only after the substrate was proven complete at all three levels: generator, ABI, public facade — D-LGJ-W6/W7): `Graph`/`Edge` (schema-not-entity `Edge`, immutable-chaining `Graph` — `from`/`hop`/`minus` each return a NEW `Graph`, mirroring `BricksQuery`'s shape rather than `View`'s laziness since every step but `hop` is already zero-cost). The row-set currency is a Java-side `long[]`, not a native `Mask` — checked and confirmed no public mask-from-row-indices constructor exists; ruled as a deliberate, documented simplification (D1 in the wave file) rather than building a FOURTH core-facade capability under time pressure. `GraphHopTest` **43/43**: hop correctness against the pinned regression (19 @ 1 hop, 29 @ 2 hops) via TWO independently-written pure-Java BFS transcriptions that never call into `Graph`; anti-vacuity; zero-serialization (structural + a reflective public-surface type check); `Edge`'s reflection guard. Disable-run: the target-decode offset corrupted by +4 → hop correctness went red exactly as required (a set-equality check caught it even though the coincidental row COUNT still matched — vindicating G2's choice to assert set equality, not just size). Core suite (204/204) and both prior consumers (trades 12+3/12+3, bricks 62/62) unaffected. **A real measured finding caught and fixed before landing, not shipped wrong:** G2's first draft asserted every hop costs an identical number of crossings; measured, hop 1 on a fresh store costs 2 (the `facetMatches` crossing plus a one-time `RowStore.rawLane()` resolution its first payload read triggers) while hop 2 onward costs exactly 1, steady-state — confirmed directly across 4 consecutive hops before touching the shipped test. Both `Graph.hop()`'s javadoc and `GraphHopTest`'s crossing assertions were corrected to state the true, now-precisely-measured relationship instead of the wrong "identical every hop" assumption | | D-LGJ-W6 | Edge-bearing row store ABI addition (`lgj_rowstore_open_with_edges`, minor 2→3, docs/abi.md §12) — the D1b-shaped "must land as its own W-tier PR before the consumer wave" the graph wave itself named | **DONE 2026-08-18** — orchestrator-authored (genuinely new ABI surface, not consumer-scope work): `registry::open_rowstore_with_edges` + `lgj_rowstore_open_with_edges` (mirrors `lgj_rowstore_open` exactly: same resource kind, same lane shape, no new mask op — purely an alternative constructor), `Engine.openRowStoreWithEdges`/`Abi.requireMinor(3)`, `RowStore.openWithEdges`. `cargo test` **93/93** (+3: registry-level open/describe, out-of-range-classid-matches-plain, radius-overflow-rejected), clippy/fmt clean, release build exports the new symbol (`nm -D`). Java: `AllTests` **194/194** (+6, all in `RowStoreParityTest`) — the strongest new result is a cross-language reproduction of the D1a hop mechanism itself: Java facet-matches + raw-lane-0 payload decode (zero new ABI op) reaches the EXACT same measured hop counts already pinned as a Rust regression (10-row seed → 19 at 1 hop → 29 at 2 hops, `n=2000, seed=0xF00D_CAFE, edge_classid=0, gate_mask=0x0, radius=25`) — proving the two sides of the membrane see identical edge structure, not merely identical classids. Two disable-runs, both red-then-green: (1) registry-level, a classid-not-threaded bug (`open_rowstore_with_edges` hardcoded classid `0`) caught by the out-of-range-parity test; (2) Java-level, the hop's classid-match condition forced to always skip → 1-hop/2-hop both went to 0 and the anti-vacuity assertion failed, exactly as expected. Caught mid-dispatch: the ABI-facing symbol did not exist before this pass (only the bare `RowStore::generate_with_edges` Rust function did, from the prior session) — the graph wave's own STOP-condition-RESOLVED note undersold what was still missing; closed here rather than discovered by G1/G2 mid-flight | | D-LGJ-W7 | Core public facade: `RowStore.classidAt`/`payloadLow64At`/`payloadHi32At` — the per-row zero-copy escape hatch a real `Graph.hop()` needs, since neither `maskOfFacetClass` nor `facetMatches` exposes payload bytes and `ApiSurfaceTest` forbids `MemorySegment`/`internal.*` in any consumer-package signature | **DONE 2026-08-18** — found while checking, concretely, how `consumers/graph` (a genuinely external compile unit per every prior consumer wave's own convention) could ever decode a matched facet's target row: it couldn't — D1a's design assumed a zero-copy raw-lane read Java-side, but nothing on the PUBLIC `RowStore` facade exposed one; only `internal.ffm.Engine.describeLane` did, off-limits to a consumer package by construction. Fixed with THREE new primitive-returning `RowStore` methods, zero new ABI surface at all (reuses `lgj_lane_describe`, already ABI minor 1 — a "lifecycle" crossing per abi.md §6, resolved once and cached; every read after is in-process, matching exports.rs's own stated doctrine "if Java wants one row it reads the MemorySegment in-process, with no crossing at all"). `AllTests` **204/204** (+10 over D-LGJ-W6: 6 in `RowStoreParityTest` — the SAME pinned hop numbers (19/29) reproduced a SECOND time, this time through the genuinely public path a real consumer has to use, plus bounds checks; 6 in `RowStoreLifetimeTest` — closed-before-first-read and closed-after-caching, both guarded). **A real self-caught redundancy, not shipped silently**: the first draft carried the closed-store guard in TWO places (`rawLane()` and `rowOffset()`); disabling one was masked by the other and produced a false-negative disable-run (30/30 green under a broken guard) — caught by re-reading WHY the disable didn't fire rather than accepting the green result, traced to Java method-call evaluation order (`rawLane()`, the receiver, evaluates before `rowOffset()`, its argument), de-duplicated to the ONE correct location, disable-run re-run and confirmed genuinely red-then-green. `ApiSurfaceTest` still 3/3 — zero FFM-typed public signature introduced. **The graph-consumer wave is now dispatchable for real**, proven at three independent levels: Rust generator (D-LGJ-W5's own entry below), ABI membrane (D-LGJ-W6), and the public core facade a consumer package can actually compile against (this row) | + +--- + +## fold-distillation wave — the Java → Panama → mask-risc grouped sum (2026-09-22) + +Upstream substrate: lance-graph #1256 (tiled executor, `GroupSumI32` / +`GroupSumViaI32` terminals, `Out::I64`), ndarray #318 (the T1 kernels), both +merged; lance-graph-java #83 (status arms, merged). Consumer witness on the +same lowering: lance-graph #1257 (SAP CATS, merged). + +| D-id | Deliverable | Status | +|---|---|---| +| D-LGJ-FOLD-4 | ABI minor 12 `lgj_plan_group_sum_i32` + `View.sumByGroup` / `sumByGroupVia` + `GroupTotals`: a `GROUP BY … SUM` as ONE program, one crossing, no selection (`docs/abi.md` §20) | **DONE 2026-09-22** — native 191 tests, Java 612 checks (JDK 28), 1 crossing measured beside the 32-crossing path, six disable arms red-then-green | +| D-LGJ-FOLD-5 | Migrate `consumers/bricks` `sumBy()` from the 32-crossing per-group path onto `sumByGroup` (re-pin BricksAuthTest's crossing arithmetic 32 → 1) | Queued | + diff --git a/CLAUDE.md b/CLAUDE.md index 708afa1..484446e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -283,10 +283,12 @@ build): not row count — verified fixed-size on both the Rust and Java sides); `Abi.java`'s `readCarvings` (bounded by `CARVING_SLOTS`, a manifest constant, not n_rows); `Engine.facetSumResolved`'s fixed `long[2]` - result pair; `View.where()`'s `List.copyOf` of the PREDICATE chain; and - `NativePattern.plan()`'s `predicates.stream()…toList()`. None of the - seven is a hidden proportional-to-n_rows population copy — keep this list - exhaustive when an eighth site is added, rather than letting the + result pair; `View.where()`'s `List.copyOf` of the PREDICATE chain; + `NativePattern.plan()`'s `predicates.stream()…toList()`; and + `Engine.groupSumI32`'s `toArray` of the group totals (minor 12 — sized by + `groups`, the key domain the caller asked for, never by rows). None of the + eight is a hidden proportional-to-n_rows population copy — keep this list + exhaustive when a ninth site is added, rather than letting the enumeration silently go stale again. > **⊘ RE-AUDITED 2026-09-16 and the list WAS stale — it claimed five and diff --git a/docs/abi.md b/docs/abi.md index 613f0e4..a58ce1e 100644 --- a/docs/abi.md +++ b/docs/abi.md @@ -62,8 +62,9 @@ cannot disagree with itself. The ABI is a **machine membrane**. It is not the product. The product is the Java semantic API (see `architecture.md`). Therefore: -- It is **small** — currently 29 symbols (minor 11's three additions are argued - in §19, and the same section argues the SEVEN capabilities it deliberately did +- It is **small** — currently 30 symbols (minor 12's one addition — the grouped + sum that never builds a selection — is argued in §20; 29 at minor 11, whose three + additions are argued in §19, and the same section argues the SEVEN capabilities it deliberately did NOT spend a symbol on; 26 at minor 10, whose one addition — the columnar constructor — is argued in §18; 25 at minor 9, whose one addition is argued in §11: a reduction Java was performing on the wrong side of the membrane, @@ -85,7 +86,7 @@ semantic API (see `architecture.md`). Therefore: ``` LGJ_ABI_MAJOR = 0 // incompatible change ⇒ bump; Java refuses to load -LGJ_ABI_MINOR = 11 // additive change ⇒ bump; older Java may still load +LGJ_ABI_MINOR = 12 // additive change ⇒ bump; older Java may still load LGJ_MAGIC = 0x4C_47_4A_5F_41_42_49_00 // "LGJ_ABI\0" big-endian-read ``` @@ -145,6 +146,17 @@ required — a gate that rejected everything would satisfy a rejection-only test ### Minor version history +- **Minor 12** (2026-09-22) — `lgj_plan_group_sum_i32` (§20): the fused plan + run STRAIGHT INTO a grouped-sum terminal. Every reduction before this minor + paid for a selection first — `lgj_plan_eval` into a mask, then a + `lgj_reduce_*` over it — and a `GROUP BY` paid it once per group (the + `bricks` consumer measured 32 crossings for 16). One symbol lowers the plan + and `GROUP BY key SUM(val)` as ONE `mask_risc::Program` whose accumulator is + tile-local scratch and whose only answer-sized state is the caller's `i64` + per group; a second table's key may be read THROUGH this table's key lane + (`via_res`/`via_lane`) with no partner-side mask. **One crossing for any + number of groups and any number of rows.** No new status; no manifest + growth; a minor-11 Java loads and sees none of it. - **Minor 11** (2026-09-14) — the masking-op completion (§19): the ndarray masking facade finished growing, and this minor consumes what it grew. **Three symbols and seven op-codes for fifteen capabilities**, which is the @@ -427,7 +439,7 @@ predicates or rows are involved. The unfused per-predicate ops are retained only so the fused path can be benchmarked *against* something and so parity can be checked predicate-by-predicate. -## 7. The function surface (29 symbols) +## 7. The function surface (30 symbols) All symbols are prefixed `lgj_`. All return `i32` status except the manifest getter. `out_*` parameters are written only on `OK`. @@ -529,6 +541,21 @@ i32 lgj_reduce_i32(u64 res, u32 lane_id, u32 reduce_op, u64 mask, Sums the `I32` lane over set mask bits into a widened `i64` (no overflow for `n_rows ≤ 2^32` on `i32` inputs). +### Grouped reduction (ABI minor ≥ 12) + +``` +i32 lgj_plan_group_sum_i32(u64 res, const LgjOpDesc* ops, u32 n_ops, + u32 group_lane, u32 val_lane, + u64 via_res, u32 via_lane, + i64* out_sums, u64 n_groups) // minor >= 12, §20 +``` + +`GROUP BY group_lane SUM(val_lane)` over the rows the plan selects, in ONE +crossing and ONE program — no selection is evaluated first and no mask handle +is involved. `out_sums[g]` for `g < n_groups`; a key past `n_groups` is +dropped. `n_ops == 0` is legal and means every row. `via_res != 0` reads the +group key THROUGH `group_lane` into `via_res`'s `via_lane` (the fk-keyed form). + ### Parity escape hatch ``` @@ -1667,3 +1694,112 @@ is textually ABSENT from the function body afterwards, not merely that a replacement occurred), and both then went red. The lesson is the one already on record and worth one more instance: assert what the disable REMOVED, never only that an edit landed. + +## 20. The grouped sum that never builds a selection (ABI minor ≥ 12) + +``` +i32 lgj_plan_group_sum_i32(u64 res, const LgjOpDesc* ops, u32 n_ops, + u32 group_lane, u32 val_lane, + u64 via_res, u32 via_lane, + i64* out_sums, u64 n_groups) +``` + +### 20.1 What was wrong before it + +Every reduction this ABI carried until minor 11 consumed a **mask**: `lgj_plan_eval` +landed the plan's answer in a mask handle, and `lgj_reduce_sum_i32` / +`lgj_reduce_i32` / the register sweeps read that handle. Java never HELD the +selection — it lived natively — but it was still a population-sized thing that +existed only to be consumed by the very next fold, which is the intermediate +materialisation the fold algebra exists to remove (lance-graph #1256's ruling: +*any intermediate population is prohibited if an addressable projection can be +consumed directly by the next fold*). A `GROUP BY` multiplied it by the group +count: `where(key.eq(g)).sumOf(value)` per `g`, measured in the `bricks` +consumer at **32 crossings for 16 groups**. + +### 20.2 What it does + +The plan is lowered exactly as `lgj_plan_eval` lowers it — the prefix rewrite +and the survivor skip of `plan_lower` — but the program ends in +`Terminal::GroupSumI32 { mask: acc, key, val }` instead of `Terminal::Keep`. The +executor runs it tile by tile over `SLOTS × tile_words_for(rows)` words of +scratch; for each tile the terminal folds `val[i]` into `out_sums[key[i]]` for +every selected row and moves on. Nothing population-sized exists at any point: +the accumulator is a tile, the sink is `n_groups` integers, and the crossing +returns when the last tile has folded. + +`via_res != 0` is the fk-keyed form. `group_lane` is then a foreign KEY into +`via_res` (another pattern resource) and the group of row `i` is +`via_lane[group_lane[i]]` — `SUM(line.amount) GROUP BY partner.country` in one +program, `Terminal::GroupSumViaI32`, with the indirection fused inside the +terminal: no partner-side mask, no remapped key lane, no second program. A key +that names no row of `via_res` drops the row (zero fallback at both hops, +`mask_risc`'s contract). `via_res` may equal `res`. + +**`n_ops == 0` is legal here** and means every row. `lgj_plan_eval` refuses an +empty plan because its caller can fill the destination mask itself; a grouped +sum has no destination mask, so the whole-lane case is a program — the one +`Pred::Range` this crate emits, always `0..n_rows`, pinned by +`the_group_sum_range_is_the_whole_lane_and_nothing_else` so that +`exec_error_to_status`'s `RangeOutOfBounds` arm stays an internal-bug mapping. + +### 20.3 Contract + +- `out_sums[g] = Σ val_lane[i]` over selected rows `i` with key `g`, for + `g < n_groups`; a selected row whose key is `>= n_groups` names no group and + is dropped, never an error. Every element is written on `OK` (zero for a + group no row names); none on any failure — `execute_into` validates the whole + program before its first write. +- `group_lane` must be a `U32` lane of `res`, `val_lane` an `I32` lane of `res`, + `via_lane` a `U32` lane of `via_res` when given. +- Statuses: `NULL_ARGUMENT` for a null `out_sums`, a null `ops` with + `n_ops > 0`, or `n_groups == 0` (a zero-length sink is no sink); + `INVALID_HANDLE` / `WRONG_RESOURCE_KIND` for `res` or a non-zero `via_res` that + is not a live pattern; every plan defect as `lgj_plan_eval` reports it; + `INVALID_LANE` / `LANE_KIND_MISMATCH` for the three lanes; `LENGTH_OVERFLOW` + past `2^32 - 1` rows (`Pred::Range` and the sum carry are both `u32`-bounded). + **No new status.** +- Bulk (§6): `O(rows)` in one pass over the predicate lanes plus one read of + the key and value lanes for the selected rows. No per-row and no per-group + crossing. + +### 20.4 Why one symbol, and why no scalar twin + +The plan surface (`LgjOpDesc`) is reused unchanged, so the whole cost of the +capability is the terminal, and the terminal's two shapes (local key / key +through a second table) are one parameter (`via_res`) rather than two symbols +— the shape §15 mandated for the reductions. There is no +`lgj_plan_group_sum_i32_scalar`: parity is falsified in the crate against +`mask_risc`'s row-at-a-time reference executor on the identical lowered +program (both key shapes, multi-tile), and through the membrane against the +two-crossing path it replaces — which runs a different kernel behind a +different terminal, so agreement is evidence rather than a tautology. + +### 20.5 The Java spelling + +`View.sumByGroup(key, value, groups)` → `GroupTotals`, and +`View.sumByGroupVia(key, via, viaKey, value, groups)`. `GroupTotals` is +addressed by key (`total(g)`, `groups()`) and exposes no array — it is sized +by the question the caller typed, never by the data, which is the eighth named +materialisation site (`Engine.groupSumI32`'s `toArray`, pinned in +`DoctrineFenceTest`). Measured: **one crossing** at 1,024 and at 65,536 rows +for 16 groups, beside the per-group path measured at **32** in the same run. + +### 20.6 The disable table (red-then-green, or it is not evidence) + +| what was disabled | test that went red | +|---|---| +| the `via` key silently downgraded to the local key | `via_reads_the_second_table_through_the_key_lane` | +| the all-rows range emptied to `0..0` | `an_empty_plan_and_an_all_or_plan_both_group_every_row`, `the_group_sum_range_is_the_whole_lane_and_nothing_else`, `grouped_sums_agree_with_quack_over_the_combine_sweep` | +| the `n_groups == 0` guard removed | `every_refusal_is_a_status_and_leaves_out_sums_untouched` (mask-risc's own refusal surfaced as `ALLOCATION_FAILED` instead) | +| Java: the `toArray` pin reverted to 1 | `DoctrineFenceTest` — *"UNFENCED MATERIALIZATION: Engine.java\|.toArray( 2 (pinned 1)"* | +| Java: `requirePositiveGroups` made a no-op | `GroupSumTest` — `groups = 0` reached the membrane and came back as the ABI's `NULL_ARGUMENT`, not the Java exception | +| Java: `Engine` passes `viaResource = 0` always | `GroupSumTest` — every `via group g` | + +Each restored from the commit (the disable cycle runs only against committed +work) and re-run green. The compatibility direction was run for real, not +assumed: `OldAbiCompatTest` against a minor-11 library built from `07e044f` +reports *"View.sumByGroup (minor 12) reports an ABI mismatch, not a missing +symbol (threw AbiMismatchException)"*, with every minor-11 feature still +working beside it. + diff --git a/java/src/main/java/com/adaworldapi/lancegraph/GroupTotals.java b/java/src/main/java/com/adaworldapi/lancegraph/GroupTotals.java new file mode 100644 index 0000000..b4f50b6 --- /dev/null +++ b/java/src/main/java/com/adaworldapi/lancegraph/GroupTotals.java @@ -0,0 +1,53 @@ +package com.adaworldapi.lancegraph; + +/** + * The answer to a grouped sum: one widened total per group, addressed by the group's key value. + * + *
This is what {@code SELECT key, SUM(value) … GROUP BY key} hands back — the same thing a + * {@code ResultSet} would, without the cursor. {@link #total(int)} is the row for key {@code g}; + * {@link #groups()} is how many keys were asked for. A key no selected row carries has a total of + * {@code 0}, exactly as SQL would report it with an outer join onto the key domain. + * + *
Sized by the question, never by the data: a {@code GroupTotals} over 16 keys is 16 numbers + * whether the view spans a thousand rows or a billion. No row, no selection and no index list is + * behind it — the totals were folded natively in one pass and only the totals crossed. + * + *
Immutable. There is no accessor for the underlying storage; the totals are read one at a
+ * time by key, which is also the only way a caller ever needs them.
+ */
+public final class GroupTotals {
+
+ private final long[] totals;
+
+ GroupTotals(long[] totals) {
+ this.totals = totals;
+ }
+
+ /** How many groups this answer covers — the {@code groups} the caller asked for. */
+ public int groups() {
+ return totals.length;
+ }
+
+ /**
+ * The sum for key {@code group}, widened to 64 bits; {@code 0} when no selected row carried
+ * that key.
+ *
+ * @throws IndexOutOfBoundsException if {@code group} is not in {@code [0, groups())}
+ */
+ public long total(int group) {
+ java.util.Objects.checkIndex(group, totals.length);
+ return totals[group];
+ }
+
+ @Override
+ public String toString() {
+ StringBuilder sb = new StringBuilder("GroupTotals[");
+ for (int g = 0; g < totals.length; g++) {
+ if (g > 0) {
+ sb.append(", ");
+ }
+ sb.append(g).append('=').append(totals[g]);
+ }
+ return sb.append(']').toString();
+ }
+}
diff --git a/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java b/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java
index f1bd991..bbaf50c 100644
--- a/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java
+++ b/java/src/main/java/com/adaworldapi/lancegraph/NativePattern.java
@@ -226,6 +226,41 @@ OptionalLong maxOf(List One crossing, whatever the number of conditions, groups or rows. The
+ * chain and the grouped sum are one fused native program; nothing is selected first and
+ * nothing per group is asked separately. Before this existed the same question cost two
+ * crossings per group — a {@code sumOf} over {@code where(key.eq(g))} for each {@code g} —
+ * which the {@code bricks} consumer measured at 32 for 16 groups; it is now 1 for any number.
+ *
+ * A key no selected row carries totals {@code 0}. A selected row whose key is
+ * {@code >= groups} belongs to no requested group and is dropped, as a {@code WHERE key < n}
+ * would drop it; ask for more groups to see it.
+ *
+ * @param key an unsigned 32-bit column whose values are the group keys
+ * @param value the signed 32-bit column to sum, widened to 64 bits
+ * @param groups how many keys to answer for, {@code > 0}
+ * @throws AbiMismatchException if the loaded library reports ABI minor < 12
+ */
+ public GroupTotals sumByGroup(U32Field key, I32Field value, int groups) {
+ java.util.Objects.requireNonNull(key, "key");
+ java.util.Objects.requireNonNull(value, "value");
+ requirePositiveGroups(groups);
+ return owner.sumByGroup(predicates, key, value, groups);
+ }
+
+ /**
+ * {@code SELECT via.viaKey, SUM(value) FROM this JOIN via ON via.row = this.key … GROUP BY
+ * via.viaKey} — the grouped sum keyed through a second resource, still one
+ * crossing.
+ *
+ * {@code key} holds, for each row here, the row index in {@code via} it refers to; the
+ * group of that row is the {@code viaKey} value found there. The lookup is fused into the
+ * native fold — no selection on either resource, no remapped key column. A {@code key} that
+ * names no row of {@code via} drops the row, as an inner join would.
+ *
+ * @param via the resource whose {@code viaKey} column supplies the group; may be this view's
+ * own resource
+ * @throws AbiMismatchException if the loaded library reports ABI minor < 12
+ */
+ public GroupTotals sumByGroupVia(U32Field key, NativePattern via, U32Field viaKey,
+ I32Field value, int groups) {
+ java.util.Objects.requireNonNull(key, "key");
+ java.util.Objects.requireNonNull(via, "via");
+ java.util.Objects.requireNonNull(viaKey, "viaKey");
+ java.util.Objects.requireNonNull(value, "value");
+ requirePositiveGroups(groups);
+ return owner.sumByGroupVia(predicates, key, via, viaKey, value, groups);
+ }
+
+ private static void requirePositiveGroups(int groups) {
+ if (groups <= 0) {
+ throw new IllegalArgumentException(
+ "groups must be positive (the number of key values to answer for), was "
+ + groups);
+ }
+ }
+
/**
* A projection of one column through this view.
*
diff --git a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java
index 8e8e38e..7b1cc60 100644
--- a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java
+++ b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Downcalls.java
@@ -491,6 +491,21 @@ private static final class Minor11 {
private Minor11() {}
}
+ /**
+ * ABI minor 12 symbol (docs/abi.md §20): the fused plan run straight into a grouped-sum
+ * terminal — one crossing returns one {@code i64} per group and no selection ever exists.
+ * Lazy per the minor-2..11 rule.
+ */
+ private static final class Minor12 {
+ static final MethodHandle PLAN_GROUP_SUM_I32 = mh("lgj_plan_group_sum_i32",
+ FunctionDescriptor.of(ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG,
+ ValueLayout.ADDRESS, ValueLayout.JAVA_INT, ValueLayout.JAVA_INT,
+ ValueLayout.JAVA_INT, ValueLayout.JAVA_LONG, ValueLayout.JAVA_INT,
+ ValueLayout.ADDRESS, ValueLayout.JAVA_LONG));
+
+ private Minor12() {}
+ }
+
/**
* Sum one facet's 12-byte register, under {@code carving}, over the rows a mask selects.
*
@@ -639,6 +654,29 @@ public static long reduceI32(long res, int laneId, int reduceOp, long mask,
return outValue.get(ValueLayout.JAVA_LONG, 0);
}
+ /**
+ * {@code GROUP BY groupLane SUM(valLane)} over the rows the plan selects, in ONE crossing
+ * (docs/abi.md §20, minor 12). {@code outSums} receives {@code nGroups} widened totals; a
+ * selected row whose key is past {@code nGroups} is dropped, never an error. With
+ * {@code viaRes != 0} the key is read THROUGH {@code groupLane} into {@code viaRes}'s
+ * {@code viaLane} (the fk-keyed form). {@code nOps == 0} is legal and means every row.
+ *
+ * Bulk in the §6 sense: one pass over the predicate lanes, one read of the key and value
+ * lanes for the selected rows, no per-row and no per-group crossing.
+ */
+ public static void planGroupSumI32(long res, MemorySegment ops, int nOps, int groupLane,
+ int valLane, long viaRes, int viaLane, MemorySegment outSums, long nGroups) {
+ crossed();
+ int st;
+ try {
+ st = (int) Minor12.PLAN_GROUP_SUM_I32.invokeExact(res, ops, nOps, groupLane, valLane,
+ viaRes, viaLane, outSums, nGroups);
+ } catch (Throwable t) {
+ throw wrap("lgj_plan_group_sum_i32", t);
+ }
+ Status.check("lgj_plan_group_sum_i32", st);
+ }
+
// ── row store (docs/abi.md §11, ABI minor 2) ─────────────────────────────────────────────
//
// Callers above this class are expected to have already checked Abi.requireMinor(2) — these
diff --git a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java
index f7d59a1..4a74f37 100644
--- a/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java
+++ b/java/src/main/java/com/adaworldapi/lancegraph/internal/ffm/Engine.java
@@ -379,6 +379,37 @@ public static ReduceOutcome reduceI32(long resource, int laneId, int reduceOp, l
return new ReduceOutcome(value, present);
}
+ /**
+ * The grouped sum that never builds a selection (docs/abi.md §20, ABI minor 12):
+ * {@code GROUP BY groupLane SUM(valLane)} over the rows {@code plan} selects, evaluated as
+ * ONE native program and returned as one widened total per group. Unlike every reduce
+ * above, no mask handle is involved at any point — the plan is not evaluated INTO anything;
+ * the fold reads it tile by tile and lands only the group totals.
+ *
+ * {@code viaResource != 0} selects the fk-keyed form: the group of a row is
+ * {@code viaLane[groupLane[row]]}, read on {@code viaResource}, with no partner-side
+ * selection and no remapped key lane (the indirection is fused inside the native terminal).
+ *
+ * An empty {@code plan} is legal here and means every row — the ABI has no destination
+ * mask for Java to fill on its own, so the whole-lane case crosses as a plan of zero ops
+ * rather than being answered locally. Requires ABI minor >= 12.
+ *
+ * The returned array is sized by {@code groups} — the shape of the QUESTION the caller
+ * typed, never by the row count — which is what keeps it on the right side of the
+ * materialization rule (the eighth named site; see the repo's zero-copy section).
+ */
+ public static long[] groupSumI32(long resource, List Three claims, each measured rather than asserted:
+ *
+ *
+ *
+ */
+public final class GroupSumTest {
+
+ private GroupSumTest() {}
+
+ private static final int GROUPS = 16; // Pattern.CLASS spans 0..15
+
+ public static void main(String[] args) {
+ System.out.println("GroupSumTest");
+ if (!NativeRuntime.isAvailable()) {
+ System.exit(Checks.reportUnavailable("GroupSumTest"));
+ }
+ Checks c = new Checks("GroupSumTest");
+ run(c);
+ System.exit(c.report());
+ }
+
+ public static void run(Checks c) {
+ if (NativeRuntime.abiMinor() < 12) {
+ c.that("SKIPPED: library minor " + NativeRuntime.abiMinor()
+ + " predates the grouped sum (minor 12); OldAbiCompatTest covers the gate", true);
+ return;
+ }
+
+ c.section("parity: one crossing answers what sixteen pairs of crossings answered");
+ for (long rows : new long[] {1_024L, 65_536L}) {
+ try (NativePattern data = NativePattern.open(rows, 0x51DEL)) {
+ View[] views = {
+ data.view(),
+ data.view().where(Pattern.VALUE.gt(100)),
+ data.view().where(Pattern.VALUE.gt(100)).where(Pattern.VALUE.gt(-50)),
+ data.view().where(Pattern.CLASS.eq(7)).where(Pattern.VALUE.gt(0)),
+ };
+ for (View v : views) {
+ GroupTotals t = v.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ c.eq(rows + " rows, " + v + ": groups()", GROUPS, t.groups());
+ long live = 0;
+ for (int g = 0; g < GROUPS; g++) {
+ long want = v.where(Pattern.CLASS.eq(g)).sumOf(Pattern.VALUE);
+ c.eq(rows + " rows, " + v + ", group " + g, want, t.total(g));
+ if (want != 0) {
+ live++;
+ }
+ }
+ // Anti-vacuity: the totals are not sixteen zeros agreeing with sixteen zeros.
+ c.that(rows + " rows, " + v + ": at least one group is non-zero", live >= 1);
+ }
+ GroupTotals all = views[0].sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ long everything = 0;
+ for (int g = 0; g < GROUPS; g++) {
+ everything += all.total(g);
+ }
+ c.eq(rows + " rows: the empty chain's groups partition the whole sum",
+ views[0].sumOf(Pattern.VALUE), everything);
+ }
+ }
+
+ c.section("a narrower key domain drops, a wider one pads with zero");
+ try (NativePattern data = NativePattern.open(4_096L, 0x51DEL)) {
+ View v = data.view().where(Pattern.VALUE.gt(0));
+ GroupTotals full = v.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ GroupTotals eight = v.sumByGroup(Pattern.CLASS, Pattern.VALUE, 8);
+ GroupTotals twenty = v.sumByGroup(Pattern.CLASS, Pattern.VALUE, 20);
+ for (int g = 0; g < 8; g++) {
+ c.eq("groups=8 keeps group " + g, full.total(g), eight.total(g));
+ }
+ for (int g = 0; g < GROUPS; g++) {
+ c.eq("groups=20 keeps group " + g, full.total(g), twenty.total(g));
+ }
+ c.eq("groups=20: key 16 is unnamed and zero", 0, twenty.total(16));
+ c.eq("groups=20: key 19 is unnamed and zero", 0, twenty.total(19));
+ c.throwsUp("total(groups()) is out of range", IndexOutOfBoundsException.class,
+ () -> eight.total(8));
+ }
+
+ c.section("cost: one crossing for sixteen groups, at any row count");
+ try (NativePattern small = NativePattern.open(1_024L, 0xC0DEL);
+ NativePattern large = NativePattern.open(65_536L, 0xC0DEL)) {
+ View sv = small.view().where(Pattern.VALUE.gt(100));
+ View lv = large.view().where(Pattern.VALUE.gt(100));
+ // Warm: the first terminal op on a resource may resolve lazily-held state.
+ sv.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ lv.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+
+ long a = Diagnostics.crossings();
+ sv.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ long smallCost = Diagnostics.crossings() - a;
+ long b = Diagnostics.crossings();
+ lv.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ long largeCost = Diagnostics.crossings() - b;
+ c.eq("1,024 rows, 16 groups: sumByGroup crosses once", 1, smallCost);
+ c.eq("65,536 rows, 16 groups: sumByGroup crosses once", 1, largeCost);
+
+ // The path it replaces, measured beside it rather than remembered from bricks.
+ lv.where(Pattern.CLASS.eq(0)).sumOf(Pattern.VALUE);
+ long d = Diagnostics.crossings();
+ for (int g = 0; g < GROUPS; g++) {
+ lv.where(Pattern.CLASS.eq(g)).sumOf(Pattern.VALUE);
+ }
+ long perGroupCost = Diagnostics.crossings() - d;
+ c.eq("the per-group path pays two crossings per group", 2L * GROUPS, perGroupCost);
+
+ long e = Diagnostics.crossings();
+ large.view().sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ c.eq("an empty chain still crosses exactly once (the plan of zero ops is the ABI's)",
+ 1, Diagnostics.crossings() - e);
+ }
+
+ c.section("the keyed form regroups through a second resource, in one crossing");
+ try (NativePattern lines = NativePattern.open(4_096L, 0xA11CEL);
+ NativePattern partners = NativePattern.open(64L, 0xB0BL)) {
+ View v = lines.view().where(Pattern.VALUE.gt(100));
+ GroupTotals local = v.sumByGroup(Pattern.CLASS, Pattern.VALUE, GROUPS);
+ GroupTotals via = v.sumByGroupVia(Pattern.CLASS, partners, Pattern.CLASS,
+ Pattern.VALUE, GROUPS);
+
+ // Oracle: partner row k (k < 16, the only rows a line's CLASS can name) carries
+ // CLASS = c(k); the via total for g is the sum of the local totals over every k
+ // with c(k) == g. c(k) is read from `partners` through the one named materialiser.
+ long[] want = new long[GROUPS];
+ for (int g = 0; g < GROUPS; g++) {
+ try (Mask m = partners.view().where(Pattern.CLASS.eq(g)).select()) {
+ for (long k : m.materializeRows()) {
+ if (k < GROUPS) {
+ want[g] += local.total((int) k);
+ }
+ }
+ }
+ }
+ boolean differs = false;
+ for (int g = 0; g < GROUPS; g++) {
+ c.eq("via group " + g, want[g], via.total(g));
+ differs |= via.total(g) != local.total(g);
+ }
+ c.that("the join genuinely regrouped (via totals differ from local totals)", differs);
+
+ long x = Diagnostics.crossings();
+ v.sumByGroupVia(Pattern.CLASS, partners, Pattern.CLASS, Pattern.VALUE, GROUPS);
+ c.eq("sumByGroupVia crosses once", 1, Diagnostics.crossings() - x);
+
+ GroupTotals self = v.sumByGroupVia(Pattern.CLASS, lines, Pattern.CLASS,
+ Pattern.VALUE, GROUPS);
+ c.eq("a self-join answers 16 groups", GROUPS, self.groups());
+ }
+
+ c.section("arguments are checked in Java, before any crossing");
+ try (NativePattern data = NativePattern.open(256L, 1L)) {
+ long before = Diagnostics.crossings();
+ c.throwsUp("groups = 0", IllegalArgumentException.class,
+ () -> data.view().sumByGroup(Pattern.CLASS, Pattern.VALUE, 0));
+ c.throwsUp("groups < 0", IllegalArgumentException.class,
+ () -> data.view().sumByGroup(Pattern.CLASS, Pattern.VALUE, -3));
+ c.throwsUp("null key", NullPointerException.class,
+ () -> data.view().sumByGroup(null, Pattern.VALUE, 4));
+ c.eq("none of those crossed", 0, Diagnostics.crossings() - before);
+
+ NativePattern gone = NativePattern.open(64L, 2L);
+ gone.close();
+ c.throwsUp("a closed via resource is rejected", ClosedResourceException.class,
+ () -> data.view().sumByGroupVia(Pattern.CLASS, gone, Pattern.CLASS,
+ Pattern.VALUE, 4));
+ data.close();
+ c.throwsUp("a closed resource is rejected", ClosedResourceException.class,
+ () -> data.view().sumByGroup(Pattern.CLASS, Pattern.VALUE, 4));
+ } catch (ClosedResourceException expected) {
+ // try-with-resources closes `data` a second time after the test closed it: that
+ // double close is itself an error by contract, and the one this block expects.
+ c.that("double close is an error, not a no-op", true);
+ }
+ }
+}
diff --git a/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java b/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java
index 00a08af..5408e76 100644
--- a/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java
+++ b/java/src/test/java/com/adaworldapi/lancegraph/OldAbiCompatTest.java
@@ -108,6 +108,19 @@ public static void run(Checks c) {
}
});
+ // Minor 12 — the grouped sum that never builds a selection. Like the minor-11 reduce
+ // above it is built entirely on NativePattern, so it is gated unconditionally here rather
+ // than inside the minor-2 block: against a genuinely minor-1 library it must name minor
+ // 12, never die on a missing symbol.
+ gate(c, loaded, 12, "View.sumByGroup", () -> {
+ try (NativePattern p = NativePattern.open(64, 0x1234L)) {
+ GroupTotals t = p.view().sumByGroup(Pattern.CLASS, Pattern.VALUE, 16);
+ if (t.groups() != 16) {
+ throw new IllegalStateException("asked for 16 groups, got " + t.groups());
+ }
+ }
+ });
+
// Minor 4 — mask complement. Needs a minor-2 store to build masks on, so it is only
// meaningful once the library has minor 2 as well.
if (loaded >= 2) {
diff --git a/native/lgj-abi/src/abi.rs b/native/lgj-abi/src/abi.rs
index 12e6de5..2dc629d 100644
--- a/native/lgj-abi/src/abi.rs
+++ b/native/lgj-abi/src/abi.rs
@@ -82,7 +82,18 @@ pub const LGJ_ABI_MAJOR: u32 = 0;
/// [`LgjOpDesc`] op-codes that cost no symbol at all
/// ([`LGJ_OP_NE_U32`] … [`LGJ_OP_TERNARY_MATCH_U32`]). No new status and no
/// manifest growth, so a minor-10 Java loads and sees none of it.
-pub const LGJ_ABI_MINOR: u32 = 11;
+///
+/// **Minor 12** (docs/abi.md §20): [`crate::exports::lgj_plan_group_sum_i32`]
+/// — the fused plan run STRAIGHT INTO a grouped-sum terminal. Until now every
+/// Java reduction paid for a selection first (`lgj_plan_eval` into a mask,
+/// then `lgj_reduce_*` over it); this symbol lowers the plan and the
+/// `GROUP BY key SUM(val)` as ONE `mask_risc::Program`, so the only
+/// population-sized state is a tile-local scratch word and what crosses back
+/// is one `i64` per group. A second table's `u32` lane may be read THROUGH a
+/// key lane of this one (`via_res`/`via_lane`, the fk-keyed group sum) with
+/// no partner-side mask and no second program. No new status and no manifest
+/// growth, so a minor-11 Java loads and sees none of it.
+pub const LGJ_ABI_MINOR: u32 = 12;
/// `"LGJ_ABI\0"` read big-endian.
///
diff --git a/native/lgj-abi/src/exports.rs b/native/lgj-abi/src/exports.rs
index b2fd0a0..15558d5 100644
--- a/native/lgj-abi/src/exports.rs
+++ b/native/lgj-abi/src/exports.rs
@@ -1719,11 +1719,15 @@ fn validate_plan(pattern: &ResourceEntry, ops: &[LgjOpDesc]) -> Result<(), i32>
/// **Every arm here is unreachable through the ABI**, and saying so is worth
/// more than implying otherwise. `validate_plan` runs first and rejects an
/// unknown opcode, a bad combine, an out-of-range lane and a kind mismatch
-/// before the lowering is even built; the lowering names no input plane
-/// (`Planes::masks` is `&[]`), no sum terminal, no blend and no `Pred::Range`.
-/// What is left — the scratch-sizing family, `ScratchReadBeforeWrite`,
-/// `GateAliasesDst`, `RangeOutOfBounds` — would be a bug in THIS file, not in
-/// a caller's plan.
+/// before the lowering is even built; the `Keep` lowering names no input
+/// plane (`Planes::masks` is `&[]`), no sum terminal, no blend and no
+/// `Pred::Range`. The grouped-sum lowering (minor 12) names a sum terminal
+/// and ONE `Range` — always `0..n_rows`, so `RangeOutOfBounds` stays a bug in
+/// this file; its `SumRowBound` needs more than `2^32` rows, which no pattern
+/// this crate can open reaches; and its foreign-lane checks are done here,
+/// before lowering, against the resolved `via` resource. What is left — the
+/// scratch-sizing family, `ScratchReadBeforeWrite`, `GateAliasesDst`,
+/// `RangeOutOfBounds` — would be a bug in THIS file, not in a caller's plan.
///
/// So the map exists to turn such a bug into a status a caller can see
/// instead of a panic, and its arms are deliberately NOT claimed to be
@@ -2137,6 +2141,295 @@ pub unsafe extern "C" fn lgj_reduce_i32(
})
}
+// ───────────────────────────────────────────────────────────────────────────
+// Grouped reduction (ABI minor ≥ 12) — the fused plan run STRAIGHT INTO a
+// grouped-sum terminal (docs/abi.md §20).
+// ───────────────────────────────────────────────────────────────────────────
+
+// Nine arguments because the ABI symbol has nine; a struct would be a
+// second spelling of the same signature with nothing to check it against.
+#[allow(clippy::too_many_arguments)]
+fn plan_group_sum_impl(
+ res: u64,
+ ops: *const LgjOpDesc,
+ n_ops: u32,
+ group_lane: u32,
+ val_lane: u32,
+ via_res: u64,
+ via_lane: u32,
+ out_sums: *mut i64,
+ n_groups: u64,
+) -> i32 {
+ // An empty plan is LEGAL here, unlike `lgj_plan_eval`: there is no
+ // destination mask a caller could fill on its own, so "every row" has to
+ // be a program (the whole-lane `Range`, `plan_lower::lower_group_sum`).
+ // `ops` may therefore be null exactly when `n_ops == 0`.
+ if out_sums.is_null() || (n_ops > 0 && ops.is_null()) {
+ return LGJ_ERR_NULL_ARGUMENT;
+ }
+ // A zero-length `out_sums` is no output buffer: no group can land
+ // anywhere, and `mask_risc` would refuse it as a missing sink. Reported
+ // as the argument defect it is rather than as the executor's refusal.
+ if n_groups == 0 {
+ return LGJ_ERR_NULL_ARGUMENT;
+ }
+ let groups = match usize::try_from(n_groups) {
+ Ok(g) if g <= isize::MAX as usize / std::mem::size_of::