diff --git a/.claude/board/ISSUES.md b/.claude/board/ISSUES.md index 127547d..05e7e51 100644 --- a/.claude/board/ISSUES.md +++ b/.claude/board/ISSUES.md @@ -1,3 +1,43 @@ +## ISS-LGJ-CONSUMERS-HAVE-NO-CI-LINE (2026-09-22) — OPEN + +`main` was RED and nothing said so. `GraphHopTest` reported 1 FAILED / 65 +passed at `bb81d80`, established by running it from a clean worktree, not +inferred. The defect itself was small — a pin asserting a second +`materializeRows()` costs zero crossings, written before `Mask.words()` began +re-validating its cached window per facade call — and is fixed +(D-LGJ-FOLD-5's second commit). **The issue is that it survived a merge.** + +`.github/workflows/` contains exactly one workflow and it gates `lgj-abi` +only: `cargo fmt --check`, `cargo clippy -D warnings`, `cargo test`. Nothing +compiles a single line of Java. The **612-check core suite and all three +consumer suites are LOCAL gates**, run by whoever remembers to run them, which +root `CLAUDE.md` states plainly (*"the merge gate is the Rust suite plus the +409-check Java run"*). A consumer is one step further out still: even a +session that runs `AllTests` religiously never touches `consumers/`. + +This is the **same shape** as the r2il probe step that sat absent from +lance-graph CI until its OGAR dependency reached main — a gate that exists +and is simply never dispatched is indistinguishable from no gate. There the +absence was deliberate and documented; here it is neither. + +**Why it is not just "add a job":** a Java job needs JDK 28 with +`--enable-preview` (JEP 401 is preview-gated, per +`ISS-LGJ-TOOLCHAIN-MUST-BE-JDK28-VALHALLA-PANAMA`) and JDK 28 is an EA build +that apt does not carry — the container obtains it from a GitHub release +download path because the distribution hosts are gateway-blocked. A CI runner +has its own network posture, so the acquisition ladder that works here is not +evidence it works there. Scoping that is this issue's first task, not an +assumption to build on. + +**Minimum that would have caught this one:** a job that builds +`java/src/{main,test}` plus `consumers/*/src` and runs `AllTests` and the +three consumer mains. No new test, no new assertion — only dispatch. + +**Falsifier for any fix:** re-introduce the stale `0` pin on a branch and +confirm the job goes red. A job that builds the consumers but never runs +their mains would pass, and would be the no-gate-with-extra-steps outcome +this entry exists to name. + ## ISS-LGJ-TOOLCHAIN-MUST-BE-JDK28-VALHALLA-PANAMA — UNBLOCKED; JDK 28 installed and the flip proven (2026-09-19) ⊘ The entry below says the migration is blocked because no JDK 28 can be diff --git a/.claude/board/LATEST_STATE.md b/.claude/board/LATEST_STATE.md index 98d593d..1ad05d2 100644 --- a/.claude/board/LATEST_STATE.md +++ b/.claude/board/LATEST_STATE.md @@ -1,3 +1,76 @@ +## 2026-09-22 — D-LGJ-FOLD-5: the consumer stops asking sixteen times, and a red pin on `main` gets root-caused + +The consumer half of minor 12. `BricksQuery.sumBy()` was sixteen queries — +one `sumOf` over `where(REGION.eq(v))` per region, each paying plan evaluation +into a selection mask plus the reduction over it — and is now one fused +`sumByGroup`. No selection mask is built at any point. + +- **Measured through the consumer, not inherited from the core suite:** 1 + crossing at 1,000 and 64,000 rows, and — new — at 1, 16 and 64 groups. + The old path was already invariant in ROWS, which is why its literal was + asserted at both row counts; this one is invariant in GROUPS as well, a + strictly stronger claim, so it gets its own arm instead of riding on the + row-count one. +- **The group-count arm needed a seam, and the seam is the honest kind.** The + public `sumBy()` always asks for exactly `Orders.REGIONS` groups, so nothing + the test could previously reach varied the group count and *"one crossing + whatever the group count"* was a doc comment no input could falsify. A + package-private `sumByGroupCount` serves that arm alone; `getMethods()` does + not see it, so the aggregate-only-egress guard still audits exactly the + public surface the guarantee is about. +- **Disable arm A recovered the old cost formula exactly:** reverting to the + per-group loop measured **2 / 32 / 128 at 1 / 16 / 64 groups**, identical at + both row counts — `2 × groups` — with **parity GREEN throughout**. The two + paths agree on the answer and differ only in cost, which is what makes this a + migration rather than a behaviour change. Note even the degenerate case + discriminates: 1 group costs 2 on the old path, 1 on the new. +- **`Orders.REGIONS` is a mirror, so it is pinned to the fixture and not to a + copied number.** It replaces three literal 16s with one named source for the + native generator's classid cardinality, which the ABI manifest does not + report. A comparison against a second Java literal would drift in lockstep + with whatever edited it and prove nothing; instead every id below REGIONS + must carry rows and id REGIONS itself must carry none. Two-sided for a + reason: arm B (REGIONS = 20) fires the first half only, arm C (REGIONS = 8) + the second only — at 3,960 rows on id 8 — so neither half alone catches + both directions. +- **Return type stays `Map`:** sized by the question, one entry + per group, never by the data. Handing back `GroupTotals` would leak a facade + type into a consumer's public surface for nothing. + +**And a separate finding, which is the part worth remembering: +`GraphHopTest` was RED on `main`.** 1 FAILED / 65 passed at `bb81d80`, +established by running it from a clean worktree rather than inferred from the +diff's shape. Consumers have **no CI line** — the only workflow gates +`lgj-abi` (fmt, clippy, test) — so nothing reported it, the same +no-CI-line shape that hid the r2il probe step upstream. Filed as +`ISS-LGJ-CONSUMERS-HAVE-NO-CI-LINE`. + +Root cause, and it is the design rather than a regression: `Mask.words()` no +longer returns its cached window directly — it re-describes through +`lgj_mask_describe` and compares the returned epoch against its stamp, so a +cached address is never read after the substrate moved under it. Root +`CLAUDE.md` names that as shipped. The pin asserting a second +`materializeRows()` costs **zero** was written against the pure-cache +behaviour that preceded it, so the test moved, not the code; reverting the +number would mean deleting the re-validation. The plan's §3.5 *"describe +cached per Mask; zero crossings steady-state"* is superseded for `Mask`: the +words are still not re-fetched and the population is still never re-scanned, +but the describe is no longer free. Steady-state is **constant per call**. + +The re-pin asserts that rather than flipping a literal: a THIRD call is +measured and required to equal the second. A one-off extra describe would +satisfy a bare `== 1` on the second call; only a per-call cost satisfies +second == third. **Disable arm D proves the two arms are independent** — +restoring the pure cache kills the magnitude arm (expected 1, was 0) and +leaves the constancy arm GREEN at 0, because it tests a different property. + +- **Gates:** core **612/612**, bricks **70/70** (was 62; the six group-count + arms and two cardinality checks are the +8), graph **68/68** (was 65 + 1 + failed; the third-call cost and its content check are the +2), trades + **12/12 + 3/3** — **765 checks**, JDK 28 `--release 28 --enable-preview` + against a freshly built `abi 0.12, ndarray::simd avx512, release`. Four + disable arms, each red-then-green. + ## 2026-09-22 — minor 12: `lgj_plan_group_sum_i32`, the grouped sum that never builds a selection (fold-distillation wave 4) Upstream first (lance-graph #1256 merged `99cdca38`: the tiled executor, @@ -26,9 +99,12 @@ selection exists at any point. - **The eighth named materialisation site:** `Engine.groupSumI32`'s `toArray`, sized by `groups` (the question), pinned in `DoctrineFenceTest` and listed in root `CLAUDE.md`. `GroupTotals` exposes no array. -- **Still owed:** the `bricks` consumer's `sumBy()` still runs the 32-crossing - path; migrating it to `sumByGroup` is a consumer-wave change, not this - one. `lgj_hop` still holds mask-sized Vecs (pre-existing, unrelated). +- **Still owed:** ⊘ the first clause is DISCHARGED (2026-09-22, D-LGJ-FOLD-5, + see the section above) — struck in place, not deleted: it was true when + written. It read *"the `bricks` consumer's `sumBy()` still runs the + 32-crossing path; migrating it to `sumByGroup` is a consumer-wave change, + not this one."* That consumer wave has now run. `lgj_hop` still holds + mask-sized Vecs (pre-existing, unrelated) and remains owed. ## 2026-09-19 — PR #81 merged (`07aa441`): production IS the Valhalla arm; Panama × Valhalla is one membrane diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index 1da2a74..60bdf70 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -131,5 +131,5 @@ same lowering: lance-graph #1257 (SAP CATS, merged). | D-id | Deliverable | Status | |---|---|---| | D-LGJ-FOLD-4 | ABI minor 12 `lgj_plan_group_sum_i32` + `View.sumByGroup` / `sumByGroupVia` + `GroupTotals`: a `GROUP BY … SUM` as ONE program, one crossing, no selection (`docs/abi.md` §20) | **DONE 2026-09-22** — native 191 tests, Java 612 checks (JDK 28), 1 crossing measured beside the 32-crossing path, six disable arms red-then-green | -| D-LGJ-FOLD-5 | Migrate `consumers/bricks` `sumBy()` from the 32-crossing per-group path onto `sumByGroup` (re-pin BricksAuthTest's crossing arithmetic 32 → 1) | Queued | +| D-LGJ-FOLD-5 | Migrate `consumers/bricks` `sumBy()` from the 32-crossing per-group path onto `sumByGroup` (re-pin BricksAuthTest's crossing arithmetic 32 → 1) | **DONE 2026-09-22** — one fused call; **1 crossing at 1,000 and 64,000 rows AND at 1, 16 and 64 groups**. The group-count arm is new and load-bearing: the public `sumBy()` always asks for `Orders.REGIONS` groups, so nothing previously reachable varied the group count and "one crossing whatever the group count" was unfalsifiable; a package-private `sumByGroupCount` exists for that arm alone (invisible to `getMethods()`, so the aggregate-egress guard still audits exactly the public surface). Disable arm A (revert to the per-group loop) measured the old cost formula exactly — **2 / 32 / 128 at 1 / 16 / 64 groups, identical at both row counts, i.e. `2 × groups`** — with parity GREEN throughout, so the two paths agree on the answer and differ only in cost. `Orders.REGIONS` replaces three literal 16s with one named source, a MIRROR of the native generator's classid cardinality (the manifest does not report it) pinned two-sided against the fixture rather than against a copied literal; arms B (20) and C (8) each fire a DIFFERENT half, so neither half alone catches both directions. BricksAuthTest **70/70** (was 62) | diff --git a/consumers/bricks/src/main/java/com/adaworldapi/bricks/BricksQuery.java b/consumers/bricks/src/main/java/com/adaworldapi/bricks/BricksQuery.java index 734aba4..3000365 100644 --- a/consumers/bricks/src/main/java/com/adaworldapi/bricks/BricksQuery.java +++ b/consumers/bricks/src/main/java/com/adaworldapi/bricks/BricksQuery.java @@ -94,35 +94,57 @@ public long sum(com.adaworldapi.lancegraph.I32Field field) { /** * Sum {@code value} grouped by every possible value of {@code group}. * - *

{@code group} is a {@link com.adaworldapi.lancegraph.U32Field}, and this consumer's fixture - * gives such fields exactly 16 distinct values ({@code 0..15}) — see {@link Orders#REGION}. This - * method issues one fused native query per group value (16 total), each narrowing this query's - * already-authorized chain by one more {@code group.eq(v)} condition and summing {@code value} - * over the result. Each of those 16 sums costs two native crossings — plan evaluation - * into the selection mask, then the {@code lgj_reduce_sum_i32} reduction — so the measured - * total is 32 crossings (unlike {@link #count()}, whose plan evaluation returns the count and - * pays one). The crossing count scales with the number of groups, never with the - * number of rows — the same laziness guarantee every other terminal operation in this - * codebase carries, just paid per group instead of once. + *

One crossing, whatever the number of groups or rows. The whole question — + * this query's authorized chain plus the grouped fold — is a single fused native program; no + * selection is built, nothing is asked per group, and only the totals cross. See {@link + * com.adaworldapi.lancegraph.View#sumByGroup}. * - *

Every group value {@code 0..15} appears as a key in the returned map, including groups with - * zero matching rows (mapped to a sum of {@code 0L}): a group's absence from a real dataset is - * itself a legitimate aggregate fact, not something to hide by omitting the key. + *

This used to be sixteen separate queries: one {@code sumOf} over {@code + * where(group.eq(v))} for each {@code v}, each costing two crossings (plan evaluation into a + * selection mask, then the reduction), measured here at 32. That path was + * invariant in the number of rows but proportional to the number of groups; this one is + * invariant in both, which is the stronger claim and is asserted as such — {@code + * BricksAuthTest} pins the cost at 1 across two row counts and three group + * counts: 1, {@link Orders#REGIONS}, and four times {@code REGIONS}. * - *

If measurement ever shows this 32-crossing loop is a bottleneck, a native grouped-aggregate - * kernel (one crossing, sixteen output buckets) is the natural W6-tier follow-up — not built - * here, because nothing has measured a need for it yet. + *

Groups come from {@link Orders#REGIONS}, the fixture's region cardinality. Every id in + * {@code 0..REGIONS-1} appears as a key in the returned map, including ids no selected row + * carries (mapped to {@code 0L}): a group's absence is itself a legitimate aggregate fact, not + * something to hide by omitting the key. + * + *

The returned map is sized by the question — one entry per group — never by the data, so + * the answer for a billion rows is the same sixteen numbers as the answer for a thousand. * * @throws UnauthorizedQueryException if {@link #authorize(Role)} was never called on this chain + * @throws com.adaworldapi.lancegraph.AbiMismatchException if the loaded library reports ABI + * minor < 12, which is where the fused grouped fold arrived */ public Map sumBy( com.adaworldapi.lancegraph.U32Field group, com.adaworldapi.lancegraph.I32Field value) { + return sumByGroupCount(group, value, Orders.REGIONS); + } + + /** + * {@link #sumBy} with the group count as a parameter, so a test can vary it. + * + *

Package-private on purpose. The public {@code sumBy} answers for exactly the fixture's + * regions and a caller has no business asking for a different number; but the claim that this + * costs one crossing whatever the group count is only a claim if something varies the + * group count, and nothing else in this package can. It is not part of the aggregate-egress + * surface the reflection guard audits — that guard reads public methods, which is the surface + * the guarantee is about. + */ + Map sumByGroupCount( + com.adaworldapi.lancegraph.U32Field group, + com.adaworldapi.lancegraph.I32Field value, + int groups) { requireAuthorized("sumBy()"); java.util.Objects.requireNonNull(group, "group"); java.util.Objects.requireNonNull(value, "value"); - Map result = new LinkedHashMap<>(16); - for (int v = 0; v < 16; v++) { - result.put(v, view.where(group.eq(v)).sumOf(value)); + com.adaworldapi.lancegraph.GroupTotals totals = view.sumByGroup(group, value, groups); + Map result = new LinkedHashMap<>(totals.groups()); + for (int v = 0; v < totals.groups(); v++) { + result.put(v, totals.total(v)); } return result; } diff --git a/consumers/bricks/src/main/java/com/adaworldapi/bricks/Orders.java b/consumers/bricks/src/main/java/com/adaworldapi/bricks/Orders.java index 3e79693..20ae898 100644 --- a/consumers/bricks/src/main/java/com/adaworldapi/bricks/Orders.java +++ b/consumers/bricks/src/main/java/com/adaworldapi/bricks/Orders.java @@ -65,9 +65,25 @@ private Orders() { public static final I32Field REVENUE = new I32Field("revenue", LaneId.of(2), Ordinal.of(1)); + /** + * How many distinct {@link #REGION} values the fixture produces: ids {@code 0..15}. + * + *

This is a fixture fact, not a schema choice — the generator derives the lane-1 tag as + * {@code (a >> 33) & (ROWSTORE_CLASS_CARDINALITY - 1)}, so the cardinality is the native + * generator's, mirrored here because the ABI manifest does not report it. A mirror can drift, + * so it is pinned empirically rather than by comparison: {@code BricksAuthTest}'s + * region-cardinality check counts rows for every id in {@code 0..REGIONS} and requires the + * first {@code REGIONS} to be populated and id {@code REGIONS} itself to be empty. Raise this + * constant without the generator changing and that check goes red. + * + *

It is also the {@code groups} argument {@link BricksQuery#sumBy} passes, which is why + * there is one named source for it instead of a literal at each site. + */ + public static final int REGIONS = 16; + /** * The region id used by {@link Role#EU_ONLY} to build its authorization mask. One of the - * fixture's 16 distinct {@link #REGION} values. + * fixture's {@link #REGIONS} distinct {@link #REGION} values. */ public static final int EU = 7; diff --git a/consumers/bricks/src/test/java/com/adaworldapi/bricks/BricksAuthTest.java b/consumers/bricks/src/test/java/com/adaworldapi/bricks/BricksAuthTest.java index c204591..48caf13 100644 --- a/consumers/bricks/src/test/java/com/adaworldapi/bricks/BricksAuthTest.java +++ b/consumers/bricks/src/test/java/com/adaworldapi/bricks/BricksAuthTest.java @@ -246,8 +246,8 @@ private static void checkFailClosed(Checks c) { // ------------------------------------------------------------------ // 3. Compose-not-execute: crossings stay 0 through where()/authorize(), and the exact - // crossing cost of each terminal (count(): 1; sumBy(): 16, group-count-scaled, never - // row-count-scaled). + // crossing cost of each terminal (count(): 1; sumBy(): 1, scaled by neither the row count + // nor the group count — see checkSumByCrossingCost for the 32 it replaced). // ------------------------------------------------------------------ private static void checkComposeNotExecute(Checks c) { @@ -290,15 +290,33 @@ private static void checkSumByCrossingCost(Checks c, int rows) { chain.sumBy(Orders.REGION, Orders.REVENUE); long cost = Diagnostics.crossings() - c0; - // A sum terminal is TWO crossings, not one: plan evaluation into the selection mask - // (lgj_plan_eval) plus the reduction itself (lgj_reduce_sum_i32) — unlike count(), - // whose plan evaluation RETURNS the count and so pays only one. 16 groups x 2 = 32. - // The thesis assertion is the arithmetic's SHAPE: crossings are proportional to the - // number of groups, never to the number of rows — which is why this same literal is - // asserted at BOTH n = 1000 and n = 64000. - c.eq("sumBy() at n = " + rows + " costs exactly 32 crossings" - + " (2 per region group — plan eval + reduce — never one per row)", - 32, cost); + // ONE crossing. The chain and the grouped fold are a single fused native program + // (lgj_plan_group_sum_i32, ABI minor 12): no selection mask is built, so there is no + // second crossing to reduce over, and nothing is asked per group. + // + // This replaces a measured 32 (re-pinned 2026-09-22, D-LGJ-FOLD-5). That number was + // real and was itself a correction: sixteen sumOf calls over where(REGION.eq(v)), each + // paying plan eval PLUS lgj_reduce_sum_i32. It was already invariant in the number of + // rows, which is why the old literal was asserted at both row counts; this path is + // invariant in the number of GROUPS as well, a strictly stronger claim, so that gets + // its own arm below. + c.eq("sumBy() at n = " + rows + " costs exactly one crossing" + + " (one fused grouped fold — no selection built, nothing per group)", + 1, cost); + + // The group-count arm. Without it, "one crossing whatever the group count" would be a + // doc comment no input could falsify: the public sumBy() always asks for exactly + // Orders.REGIONS groups, so the two row counts above vary rows and hold groups fixed. + // A path that still paid per group would read 1 here at 1 group and more at 64. + for (int groups : new int[] {1, Orders.REGIONS, 4 * Orders.REGIONS}) { + chain.sumByGroupCount(Orders.REGION, Orders.REVENUE, groups); // warm-up + long g0 = Diagnostics.crossings(); + chain.sumByGroupCount(Orders.REGION, Orders.REVENUE, groups); + long gCost = Diagnostics.crossings() - g0; + c.eq("sumBy() at n = " + rows + " with " + groups + " groups still costs exactly" + + " one crossing (invariant in groups, not merely in rows)", + 1, gCost); + } } } @@ -311,6 +329,7 @@ private static void checkSumByParity(Checks c) { checkSumByCrossingCost(c, 1_000); checkSumByCrossingCost(c, 64_000); + checkRegionCardinality(c); final int rows = 64_000; Fixture expected = generate(rows, SEED); @@ -360,6 +379,54 @@ private static void checkSumByParity(Checks c) { // 5. Aggregate-only egress: BricksQuery's return-type surface, and Orders as a schema. // ------------------------------------------------------------------ + /** + * {@link Orders#REGIONS} is a MIRROR of the native generator's classid cardinality, which the + * ABI manifest does not report — so it is pinned against the fixture's own behaviour rather + * than against a number copied from the other side. Two-sided on purpose: every id below + * REGIONS must be populated (so the constant is not too LARGE) and id REGIONS itself must be + * empty (so it is not too SMALL). A comparison against a second Java literal would drift in + * lockstep with whatever edited it and prove nothing; this goes red if the generator changes + * under us, in either direction. + * + *

Mirrors the shape of the native side's own cardinality test, which asserts every facet + * lane hits all 16 classids at n = 4096. + */ + private static void checkRegionCardinality(Checks c) { + c.section("Orders.REGIONS is pinned to the fixture, not to a copied literal"); + + final int rows = 64_000; + try (BricksSession session = Bricks.open(rows, SEED)) { + int populated = 0; + for (int region = 0; region < Orders.REGIONS; region++) { + long n = session.query() + .where(Orders.REGION.eq(region)) + .authorize(Role.GLOBAL) + .count(); + if (n > 0) { + populated++; + } else { + c.that("region " + region + " is populated at n = " + rows + + " — an empty id below Orders.REGIONS means the constant" + + " is larger than the generator's cardinality", + false); + } + } + c.eq("every id in 0..Orders.REGIONS-1 carries rows at n = " + rows, + Orders.REGIONS, populated); + + long beyond = session.query() + .where(Orders.REGION.eq(Orders.REGIONS)) + .authorize(Role.GLOBAL) + .count(); + c.eq("id Orders.REGIONS (" + Orders.REGIONS + ") carries no rows — a populated id" + + " here means the constant is smaller than the generator's" + + " cardinality and sumBy() is silently dropping a group", + 0L, beyond); + } + } + + // ------------------------------------------------------------------ + private static void checkAggregateOnlyEgress(Checks c) { c.section("aggregate-only egress — BricksQuery's public return types"); diff --git a/consumers/graph/src/main/java/com/adaworldapi/graph/Graph.java b/consumers/graph/src/main/java/com/adaworldapi/graph/Graph.java index 95df95a..643362d 100644 --- a/consumers/graph/src/main/java/com/adaworldapi/graph/Graph.java +++ b/consumers/graph/src/main/java/com/adaworldapi/graph/Graph.java @@ -215,8 +215,11 @@ public long count() { /** * The current frontier's row indices — the ONE named materialising terminal on this class (root * {@code CLAUDE.md}'s "named exceptions": row ids out). {@code O(n)} in the frontier's - * population; see {@link Mask#materializeRows()} for the exact cost shape (a one-time describe, - * then zero further crossings for repeated calls against the same underlying {@link Mask}). + * population; see {@link Mask#materializeRows()} for the exact cost shape — one lifecycle + * crossing per call, repeated calls included, because the cached window is re-validated + * against the registry every time rather than read on trust. Corrected 2026-09-22: this + * read *"a one-time describe, then zero further crossings for repeated calls"*, which was + * the pure-cache behaviour that preceded the re-validation. */ public long[] materializeRows() { return frontier == null ? new long[0] : frontier.materializeRows(); diff --git a/consumers/graph/src/test/java/com/adaworldapi/graph/GraphHopTest.java b/consumers/graph/src/test/java/com/adaworldapi/graph/GraphHopTest.java index d49794a..0b24244 100644 --- a/consumers/graph/src/test/java/com/adaworldapi/graph/GraphHopTest.java +++ b/consumers/graph/src/test/java/com/adaworldapi/graph/GraphHopTest.java @@ -149,6 +149,26 @@ public static void main(String[] args) { private static final long PREDICTED_COUNT_CROSSINGS = 1; // MEASURED-AND-PINNED private static final long PREDICTED_MINUS_CROSSINGS = 2; // MEASURED-AND-PINNED private static final long PREDICTED_MATERIALIZE_FIRST_CROSSINGS = 1; // MEASURED-AND-PINNED + + /** + * What a materializeRows() call costs AFTER the first, re-pinned 1 (was 0) on 2026-09-22. + * + *

It was 0 while {@code Mask.words()} was a pure cache. It is 1 because that cache now + * re-validates: every top-level facade call re-describes the mask through {@code + * lgj_mask_describe} and compares the returned epoch against the stamp it holds, so a cached + * address can never be read after the substrate has moved under it. That is the shipped + * safety property (root {@code CLAUDE.md}, zero-copy section: {@code Mask} "re-validates the + * cached window against the generation-checked registry at each top-level facade call"), so + * the number is the design's, not a regression — reverting it would mean deleting the + * re-validation. + * + *

The pin this replaces cited the plan's §3.5 "describe cached per Mask; zero crossings + * steady-state". That reading is superseded for {@code Mask}: the describe is still cached in + * the sense that the WORDS are not re-fetched and the population is never re-scanned — + * {@code lgj_mask_describe} fills a descriptor and does no work over the rows — but it is no + * longer free. Steady-state is now "constant per call", which is what the arms below assert. + */ + private static final long PREDICTED_MATERIALIZE_STEADY_CROSSINGS = 1; // MEASURED-AND-PINNED private static final long PREDICTED_IMPORT_CROSSINGS = 2; // MEASURED-AND-PINNED private static long[] seedRows() { @@ -449,9 +469,11 @@ private static void checkCrossingsProportionalToHops(Checks c) { c.eq("minus(Mask) costs exactly " + PREDICTED_MINUS_CROSSINGS + " crossing(s)", PREDICTED_MINUS_CROSSINGS, minusCost); - // materializeRows(): "describe cached per Mask; zero crossings steady-state" (spec - // §3.5) -- so the FIRST call on a given Mask should cost something, the SECOND call - // on the very same Mask should cost nothing further. + // materializeRows(): constant cost per call, not zero after the first -- see + // PREDICTED_MATERIALIZE_STEADY_CROSSINGS for why the spec's "zero crossings + // steady-state" no longer holds and why that is the design rather than a regression. + // A THIRD call is measured too: it is what distinguishes "constant per call" from + // "one extra, once", and only the former is consistent with a per-call re-validation. long beforeFirstMaterialize = Diagnostics.crossings(); long[] materializedOnce = afterMinus.materializeRows(); long firstMaterializeCost = Diagnostics.crossings() - beforeFirstMaterialize; @@ -462,9 +484,21 @@ private static void checkCrossingsProportionalToHops(Checks c) { c.eq("materializeRows() first call costs " + PREDICTED_MATERIALIZE_FIRST_CROSSINGS + " crossing (the one-time describe of this Mask's word lane)", PREDICTED_MATERIALIZE_FIRST_CROSSINGS, firstMaterializeCost); - c.eq("materializeRows() SECOND call on the SAME Mask costs zero further crossings" - + " -- steady-state, per spec §3.5", - 0, secondMaterializeCost); + long beforeThirdMaterialize = Diagnostics.crossings(); + long[] materializedThrice = afterMinus.materializeRows(); + long thirdMaterializeCost = Diagnostics.crossings() - beforeThirdMaterialize; + + c.eq("materializeRows() SECOND call on the SAME Mask costs " + + PREDICTED_MATERIALIZE_STEADY_CROSSINGS + + " further crossing -- the cached window is re-validated against the" + + " registry on every facade call, never read on trust", + PREDICTED_MATERIALIZE_STEADY_CROSSINGS, secondMaterializeCost); + c.eq("materializeRows() THIRD call costs the same as the second -- constant per call," + + " which a one-off extra describe would not produce", + secondMaterializeCost, thirdMaterializeCost); + c.that("the third materializeRows() agrees with the first on content too -- the" + + " re-validation returns the same selection, it does not re-answer it", + sameRowSet(materializedOnce, materializedThrice)); c.that("both materializeRows() calls agree on content (the describe caching did not" + " change the answer)", sameRowSet(materializedOnce, materializedAgain)); diff --git a/java/src/main/java/com/adaworldapi/lancegraph/Mask.java b/java/src/main/java/com/adaworldapi/lancegraph/Mask.java index d3cdfb7..02725c4 100644 --- a/java/src/main/java/com/adaworldapi/lancegraph/Mask.java +++ b/java/src/main/java/com/adaworldapi/lancegraph/Mask.java @@ -40,12 +40,25 @@ public final class Mask implements AutoCloseable { private boolean closed; /** - * The mask's own packed-bit word window, resolved once via {@code lgj_mask_describe} (a - * lifecycle crossing, per abi.md §6, not a bulk one) and cached: mask storage is allocated - * once and never reallocated, resized, or moved while the resource is alive (a hard ABI - * guarantee), the same invariant {@link RowStore#rawLane()} relies on for its own caching. - * Every read or write through {@link #materializeRows()} after the first call is an - * in-process segment access with no further crossing at all. + * The mask's own packed-bit word window, obtained via {@code lgj_mask_describe} (a lifecycle + * crossing, per abi.md §6, never a bulk one) and cached: mask storage is allocated once and + * never reallocated, resized, or moved while the resource is alive (a hard ABI guarantee), + * the same invariant {@link RowStore#rawLane()} relies on for its own caching. + * + *

The cache is re-validated, not trusted, so every facade call pays one lifecycle + * crossing — including calls after the first. {@link #words()} re-describes and + * compares the returned epoch against the stamp it holds, so a cached address can never be + * read after the substrate has moved under it. What the cache still buys is what matters: + * the words are never re-fetched and the population is never re-scanned, because + * {@code lgj_mask_describe} fills a descriptor and does no work over the rows. The cost is + * therefore CONSTANT PER CALL, not zero after the first. + * + *

⊘ This paragraph replaces *"resolved once … and cached: every read or write through + * {@link #materializeRows()} after the first call is an in-process segment access with no + * further crossing at all"*, which described the pure-cache behaviour that preceded the + * re-validation and was false once it landed. Corrected 2026-09-22 alongside the + * {@code GraphHopTest} pin that measured it; found by review, not by a test, because a + * javadoc contract has none. */ private Engine.LaneWindow words;