From 75934e3455bc056678e4821b81f9238ee9096246 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 21:57:24 +0000 Subject: [PATCH 1/7] =?UTF-8?q?plan:=20deepnsm-v2-lexical-address-v1=20?= =?UTF-8?q?=E2=80=94=20harvest=20of=20closed=20#1303?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A word is a 16-bit address into an immutable, versioned COCA codebook; six words per facet; relations driver-supplied via Fisher-z. Carries the measured three-reference correspondence (3 of 4,264 ordinals align), the COCA bake with f/c baked as u8 (no runtime counting), and ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE. D-LXA-1..4 queued. The #1303 plan files and socket sections are not carried. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .claude/board/INTEGRATION_PLANS.md | 16 ++ .claude/board/ISSUES.md | 6 + .claude/board/STATUS_BOARD.md | 9 + ...-reference-sets-are-not-ordinal-aligned.md | 23 +++ .claude/board/entries/README.md | 3 +- .../plans/deepnsm-v2-lexical-address-v1.md | 154 ++++++++++++++++++ 6 files changed, 210 insertions(+), 1 deletion(-) create mode 100644 .claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md create mode 100644 .claude/plans/deepnsm-v2-lexical-address-v1.md diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index a104d2a66..1b5cd83b6 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -1,3 +1,19 @@ +## 2026-09-30 (1) — deepnsm-v2-lexical-address-v1 — a word is a 16-bit address into a versioned, baked COCA codebook → `.claude/plans/deepnsm-v2-lexical-address-v1.md` + +**Status:** PROPOSAL (D-LXA-1..4). No code authorized. Harvested from PR #1303 +(`deepnsm-v2-cam96-pairwise` v1–v5, closed unmerged). The v1–v5 plan files are not +carried; they remain on branch `claude/brave-mayer-65y3cy`. + +- **Six words per facet.** Each `[a,b]` is one 16-bit address into an immutable, + versioned lexical codebook. It is not two poles and not an embedding slice. +- **Relations are driver-supplied** through the Fisher-z LUT. +- **Three COCA references** (4096 / 5k lemma / 20k academic) are measured as neither + nested nor aligned: only 3 of 4,264 shared words keep their ordinal. Reading across + references is refused. +- **The COCA bake** stores frequency, PoS and lemma in the row as a u8 `` prior. + Nothing is counted at runtime. +- Carries `ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE`. + ## 2026-09-29 (2) — deepnsm-v2-lexical-evidence-consumer-v1 rewritten against the DeepNSM → DeepNSM-v2 migration → `.claude/plans/deepnsm-v2-lexical-evidence-consumer-v1.md` **Status:** PROPOSAL (D-LXC-1..10). No code authorized. Supersedes entry (1) diff --git a/.claude/board/ISSUES.md b/.claude/board/ISSUES.md index 7e3df4816..b61ad1cfb 100644 --- a/.claude/board/ISSUES.md +++ b/.claude/board/ISSUES.md @@ -1,3 +1,9 @@ +## ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE — the driver packs p64 plane bits as a CausalMask; the inverse reads bit 2 differently (2026-09-30) + +**Status:** OPEN. **Basis:** VERIFIED-IN-CODE on `main` `0d31c54f`. `cognitive-shader-driver` `driver.rs` (emission stage) packs `CausalMask::from_bits(h.predicates & 0x07)`, where `predicates` is the p64 predicate-plane byte: bit0 CAUSES, bit1 ENABLES, **bit2 SUPPORTS** (`p64-bridge` `SUPPORTS = 2`). `p64-bridge::edge_to_layer_mask` maps mask bit 2 (confounding) to **CONTRADICTS**. Emit and its inverse disagree on bit 2; three mask vocabularies (rung, Pearl, p64) meet at one byte. +- Found during PR #1303 (closed unmerged); recorded with `deepnsm-v2-lexical-address-v1.md` §6, which does not depend on it. +- **What closes it:** one named mapping in `p64-bridge` used by both directions, with a round-trip test that can fail on bit 2. + ## ISS-REPORT-NO-COMPOSITE-KEY-GROUP-FOLD — a 2-D fold costs one pass per partition member (2026-09-23) **Status:** OPEN. **Basis:** VERIFIED-IN-CODE. mask-risc's `GroupKey` is one diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index 80bd4a365..86d2440a4 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -305,6 +305,15 @@ evaluates one; `execute` stays the consumer's call on a scratch it owns. | D-QCK-8 | accumulator gating — an `AND` gates each later child under the accumulator built so far, not only under a resident plane. The same rule `lgj-abi`'s `plan_lower` already had; the two lowerings now implement one law | Shipped. Disable-verified 3×: gate-never-on-accumulator, an `OR` gating its children on its own accumulator, and `hoist_gate_subset` never rotating — each reddens | `the_gate_reaches_every_comparison_and_is_dropped_only_where_it_vanishes` via the new `gate_shape` helper (all-gated / on-accumulator / dropped, three facts not two); `every_lowered_filter_agrees_with_an_independent_per_row_reading` is what catches the DROP going unsound | | D-QCK-9 | A1 `Filter::and_by_skip` — the caller's measured skip score orders an `AND`'s conjuncts. DuckDB's hill-climb NOT ported (see the corrected reason) | Shipped, with the matrix's A1 falsifier RUN (`examples/adaptive_order_probe.rs`, 65 536 rows × 5 conjuncts × all 120 permutations × 4 regimes). Order moves the skip fraction, so A1 is **ADAPT conditionally** — inert in 2 of the 4 regimes, and the pre-registered *representative predicate stream* half was never run (matrix §9). ⊘ **Corrected 2026-09-15:** this row read "the control signal is DEAD WORDS, not selectivity" — the probe never ranks by either signal and never calls `and_by_skip`, so only the between-regime diagnostic is supported; and the "hill-climb explores the wrong landscape" reason was MEASURED FALSE (`[4092, 3069, 2046, 1023, 0]`, deltas all −1023: a monotone ramp). The real reason is that this crate never executes | `skip_ordering_moves_the_work_and_never_the_answer`: the reordering agrees with the oracle (not with the other lowering), the emitted program's first predicate is the highest-scored, and descending/no-sort disables both redden. **The probe itself is run by no gate** — `cargo test` compiles an example and never executes it — so the four-row table is a measured-once observation, not a pinned one ⊘ **2026-09-15, operator ruling (E-256-BY-256-IS-EXACTLY-64K-THE-RAILS-SKIP-UNIT-IS-ITS-HI-BYTE-AND-A-QUARTER-BLOCK-IS-A-REMAINDER-1):** the clustered figures were the `/50` quarter-block cut's; re-measured on the byte boundary (`/48`, the rail's hi byte) → 116 survivors (0.177 %), 99.61 % in both words and 256-row blocks; selective in blocks 0.00 → 61.91 %. The probe now prints both units and the ramp. ⊘ Same day (E-THE-RAIL-IS-A-NEEDLE-NOT-A-MASK-…-1): "the rail's hi byte" as skip unit was my reading — a rail is a 2-byte row ADDRESS, not a mask; the block is the tile's 2-nibble cell; numbers unchanged. ⊘ Same day (E-A-THOUGHT-MASKS-ITSELF-BY-ITS-DISTANCE-FROM-ROOT-…-1): the radius is STEPLESS — the probe sweeps d = 40..=56, 2^(56−d) rows each; words saturate at d = 50, blocks at d = 48 | +## deepnsm-v2-lexical-address (D-ids minted 2026-09-30, `.claude/plans/deepnsm-v2-lexical-address-v1.md`; harvested from closed #1303) + +| D-id | scope | status | gate / falsifier | +|---|---|---|---| +| D-LXA-1 | `LexicalAddress(u16)` newtype + `ReferenceSet {id, version, sha256}`; cross-reference read refused | Queued | G-LEX: every declared entry resolves to its (lemma, PoS); wrong reference refused; bare `u16` is `compile_fail` | +| D-LXA-2 | `lexical_correspondence.tsv` generator (id4096 / id5k / id20k / Exact·Ambiguous·Missing, three digests) | Queued | G-REF: re-derivation reproduces the measured table incl. "3 of 4,264"; one hand-edited ordinal reddens | +| D-LXA-3 | the COCA bake: (lemma, PoS) rows with `f` (reading share) and `c` (log-frequency evidence) baked as u8 offline; unknown flagged, never 0 | Queued — shares `academic_20k.csv` with D-LXC-4 (Blocked); must not duplicate that loader | byte-identical re-derivation; `the` f = 1, `record` f < 1; all-unambiguous fixture stays f = 1 | +| D-LXA-4 | six-slot ClassView reading of a 12-byte facet as six `LexicalAddress`es under one `ReferenceSet` | Queued — home (second facet vs `Identity` tenant) OPEN | X-written / Y-read refused; slot rotation changes the resolved words | + ## mask-risc-executor (PR3 — D-ids minted 2026-09-14, `.claude/plans/mask-risc-executor-v1.md`) | D-id | scope | status | gate / falsifier | diff --git a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md new file mode 100644 index 000000000..6f56a9501 --- /dev/null +++ b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md @@ -0,0 +1,23 @@ +# 2026-09-30 — The three COCA reference sets are neither nested nor ordinal-aligned + +**Status:** MEASURED · OPEN — plan `.claude/plans/deepnsm-v2-lexical-address-v1.md` §2 (D-LXA-2, G-REF). Measured on PR #1303 (closed unmerged); carried here. + +## Measured (python over the committed CSVs and the Tigris artifact) +| reference | rows | distinct keys | key | +|---|---|---|---| +| COCA4096 — `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank (homographs share a rank) | +| COCA5K_LEMMA — `lemmas_5k.csv` | 5,050 | 4,380 lemmas / 5,050 (lemma, PoS) | (lemma, PoS) | +| COCA20K_ACAD — `academic_20k.csv`; Tigris `lance-graph/codebooks/deepnsm-v2-academic-coca-v1/` | 20,845 | 18,559 words / 20,842 (word, Pos) | `word_id` = admission order | + +- 5k ∩ 20k: 4,895 of 5,050 (lemma, PoS); 4,264 words; **116 of the 5k lemmas are absent from the 20k**. +- 4096 ∩ 20k: 3,462 of 3,559 words. +- **Ordinal alignment: 3 of 4,264 shared words carry the same ordinal in the 5k and in the 20k.** +- Ambiguity differs per set: the 20k carve drops 2,286 same-word-different-Pos duplicates to one id (MANIFEST: 18,559 of 20,480 reserved slots; basins 73..79 empty); the 4096 shares ranks across homographs; the 5k keeps (lemma, PoS) distinct. +- The Tigris "academic codebook" TSV is a `PaletteVocab::from_frequency_ranked` carve (`word_id, basin, slot, …`), not a trained centroid codebook. Its `academic_20k.csv` sha256 equals the committed file's. + +## Consequence +A reading contract names (reference id, version, digest); an ordinal is stable only within its reference; correspondence is an explicit (lemma, PoS) map with `Ambiguous{n}` / `Missing` as values; switching references must preserve the declared map and never reuse an ordinal. Gate G-REF re-derives this table and diffs it. + +## Open +- Whether the academic reference gets its own codebook or shares COCA4096's through the correspondence map (plan §5). +- `range`/`disp` are loaded by no `src/` code; `Vocabulary::load` never opens `lemmas_5k.csv`. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 393a6e819..c8c0120f4 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,10 +25,11 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -176 entries, 2026-08-06 .. 2026-09-29. +177 entries, 2026-08-06 .. 2026-09-30. | date | entry id | finding | file | |---|---|---|---| +| 2026-09-30 | `three-reference-sets-are-not-ordinal-aligned` | | [2026-09-30-three-reference-sets-are-not-ordinal-aligned.md](2026-09-30-three-reference-sets-are-not-ordinal-aligned.md) | | 2026-09-29 | `deepnsm-v2-counted-pick-tag-deltas` | | [2026-09-29-deepnsm-v2-counted-pick-tag-deltas.md](2026-09-29-deepnsm-v2-counted-pick-tag-deltas.md) | | 2026-09-26 | `deepnsm-v2-lexical-evidence-survives-routing` | | [2026-09-26-deepnsm-v2-lexical-evidence-survives-routing.md](2026-09-26-deepnsm-v2-lexical-evidence-survives-routing.md) | | 2026-09-25 | `window-scheduling-and-two-level-ternlog` | | [2026-09-25-window-scheduling-and-two-level-ternlog.md](2026-09-25-window-scheduling-and-two-level-ternlog.md) | diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md new file mode 100644 index 000000000..a6eab1e0b --- /dev/null +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -0,0 +1,154 @@ +# deepnsm-v2-lexical-address-v1 — a word is a 16-bit address into a versioned, baked COCA codebook + +> **Status:** PROPOSAL (D-LXA-1..4). Plan only; no code is authorized by this file. +> **Written against:** `main` `0d31c54f` (2026-09-30). +> **Harvested from:** PR #1303 (`deepnsm-v2-cam96-pairwise-v5`), which was **closed without +> merging**. Only four things from it are kept here: +> - its CORRECTION 2 (the lexical-address design); +> - its F23 three-reference measurement; +> - its COCA prior, recast as a bake; +> - one verified code defect. +> +> Everything else stays on branch `claude/brave-mayer-65y3cy` (head `d6c041d7`) and is +> **not** carried: the embedding-subspace design and its gates (already retracted in v5 +> itself), the §11/§11R execution socket, and the "baton" doc-comment edits. +> **Board:** `STATUS_BOARD.md` § deepnsm-v2-lexical-address · entry +> `entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md` · `ISSUES.md` +> `ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE`. + +--- + +## §0 — In plain words + +A facet's 12 payload bytes hold **six words**. Each word is a pair of bytes `[a,b]`, +read together as one **16-bit address**. That address points at one row of an +**immutable, versioned lexical codebook**, and that row says exactly which word it is: +its lemma and part of speech. + +- The two bytes carry no separate meaning. They are not two poles, not + "nearest / second-nearest centroid", and not a slice of an embedding. +- **Relations between words are not in the bytes.** The cognitive shader driver + supplies them, calibrated through the Fisher-z LUT. Fisher-z governs relation values; + it never makes a word's identity fuzzy. +- **Frequency, PoS and lemma are baked into the codebook row**, computed once and + offline. Nothing is counted at runtime. +- Embeddings are at most a comparison instrument. Reconstructing an embedding is **not** + an acceptance criterion. + +Operator, 2026-09-30: + +> *"Six lexical slots, not six embedding subspaces. Each slot's two-byte address resolves +> through an immutable, versioned lexical codebook; the driver supplies the selected +> relational reading."* + +--- + +## §1 — Two readings of `[a,b]`, and which one this is + +| | reading | status | +|---|---|---| +| (1) | `[a,b]` **identifies a word**: an exact, unique 16-bit address | **the representation** | +| (2) | `[a,b]` names two lexical poles whose cell is a relation | **not** a unit of identity anywhere in this plan | + +Earlier Cam96 drafts (v1–v5 on #1303) silently substituted (2) for (1), under the name +"nearest and second-nearest centroid". An artifact check found that nothing shipped, +released, or stored in Tigris ever assigns a word two needles: +- the COCA CAM-PQ gives six per word (one per subspace); +- bgz17 / bgz-tensor `nearest`/`assign` give one. + +The phrase originated in #1303's v1 with no source. + +--- + +## §2 — Three reference sets, never interchangeable (F23, measured) + +An address means something only against **one named reference**. There are three. They +are neither nested nor aligned, measured over the committed CSVs and the Tigris artifact: + +| reference | source | rows | distinct keys | key | +|---|---|---|---|---| +| `COCA4096` | `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank (homographs share a rank) | +| `COCA5K_LEMMA` | `lemmas_5k.csv` | 5,050 | 4,380 lemmas / 5,050 (lemma, PoS) | (lemma, PoS) | +| `COCA20K_ACAD` | `academic_20k.csv`; Tigris `lance-graph/codebooks/deepnsm-v2-academic-coca-v1/` | 20,845 | 18,559 words / 20,842 (word, Pos) | `word_id` = admission order | + +- 5k ∩ 20k = 4,895 of 5,050 (lemma, PoS) and 4,264 words. **116 of the 5k lemmas are + absent from the 20k.** +- 4096 ∩ 20k = 3,462 of 3,559 words. +- **Only 3 of the 4,264 shared words carry the same ordinal in the 5k and the 20k.** +- Each set handles homographs differently: + - the 20k collapses 2,286 same-word-different-Pos duplicates into one id; + - the 4096 shares a rank across homographs; + - the 5k keeps each (lemma, PoS) distinct. +- The Tigris "academic codebook" is a vocabulary carve (`PaletteVocab::from_frequency_ranked`), + not a trained centroid codebook. Its CSV's sha256 equals the committed file's. + +**The rule:** +- A reading contract names `ReferenceSet { id, version, sha256 }`. +- An ordinal is stable **within** one reference, never across references. +- Correspondence between references is an explicit (lemma, PoS) map with + `Ambiguous{n}` and `Missing` as values. It is never assumed to be nested or aligned. +- Reading an address against the wrong reference is **refused**. It is never answered + with a different word. + +--- + +## §3 — The COCA bake (frequency, PoS and lemma built in; nothing counted at runtime) + +Each codebook row is one `(lemma, PoS)` identity. Its fields are computed **once, +offline**, when the artifact is built. Float arithmetic is allowed at build time, which +is the derived side of the no-float rule. The results are stored quantized as `u8`: + +| field | how it is computed | why | +|---|---|---| +| `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | +| **`f`** (u8) | the reading's share among its surface form's readings: `form_count(this) / surface_count(surface)` (from `LexicalEvidence`). `f = 1` for an unambiguous form | an ambiguous form (`record`) gets `f < 1` without anyone counting at query time | +| **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | +| unknown | a missing count stays **unknown** (a flag), never `0` | zero would assert "no evidence"; unknown asserts "not measured" | + +`range` and `disp` are left out. They are collinear with frequency, and `disp` measures +evenness across genres, not evidence. + +At runtime the prior `` is **a read of the resolved row**, never a computation. +Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11). + +--- + +## §4 — Deliverables, in build order + +| D-id | what | gate (can fire / can stay silent) | +|---|---|---| +| **D-LXA-1** | `LexicalAddress` newtype over `u16` plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX:** every declared entry of a reference resolves to exactly its declared (lemma, PoS). A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | +| **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "3 of 4,264". Hand-editing one ordinal must turn the check red | +| **D-LXA-3** | The COCA bake of §3: one artifact per reference, `f`/`c` baked as `u8`, unknown flagged | Re-deriving must reproduce the bake byte for byte. Pinned rows: `the` must have `f = 1` and `record` must have `f < 1`. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | +| **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`. A register in any other shape is refused, never reinterpreted | A facet written under reference X and read under Y is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag) | + +**Collision:** D-LXA-3 reads `academic_20k.csv`, which `D-LXC-4` (the academic loader, +currently Blocked on a ruling about three duplicate (word, PoS) pairs) also owns. D-LXA-3 +waits for that ruling, or takes it as its own first question. It must not duplicate the +loader. + +--- + +## §5 — Open + +- **Where the six slots live:** the row's second facet (bytes 16..32) or a new 16-byte + `Identity` value tenant. Not decided. Either choice is additive and leaves + `NODE_ROW_STRIDE` unchanged. +- **Which relational reading(s) the driver selects** between two resolved words, and the + LUT that carries them. That belongs to the ClassView and is not decided here. +- **Does `COCA20K_ACAD` get its own codebook**, or does it share `COCA4096`'s through the + correspondence map? + +--- + +## §6 — The verified defect carried with this plan + +`ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE` (`ISSUES.md`). The driver packs the p64 +predicate-plane byte straight into a `CausalMask`: `driver.rs`, emission stage, +`CausalMask::from_bits(h.predicates & 0x07)`. In that byte, bit 2 is **SUPPORTS** +(`p64-bridge`: `SUPPORTS = 2`). But `edge_to_layer_mask` maps mask bit 2 to +**CONTRADICTS**. Emitting a mask and inverting it therefore disagree on bit 2. + +This plan does not depend on the defect. It is recorded because it is real and was found +in the same read. The fix belongs to `p64-bridge`: one named mapping used in both +directions, plus a round-trip test that can fail on bit 2. From 575d76fedb856859021f9495ce460a4c702e927a Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 22:11:40 +0000 Subject: [PATCH 2/7] =?UTF-8?q?plan=20lexical-address:=20address=20Codex?= =?UTF-8?q?=20review=20=E2=80=94=20ordinal=200,=20ranks,=20f=20by=20surfac?= =?UTF-8?q?e=20form?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Aligned ordinals are 4 of 4,264 (the/there/care/wage), not 3: the earlier count dropped ordinal 0. G-REF now requires ordinal 0 to be counted. - COCA4096 ranks are unique per (word, PoS); a homograph spans several ranks, it does not share one. - The bake splits into an identity table (lemma, PoS, c) and a surface-form table (form, reading, f): the reading share is a property of the surface form (record -> n 120,048 / v 13,014), not of the lemma row. Re-measured from the committed CSVs. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .claude/board/INTEGRATION_PLANS.md | 2 +- .claude/board/STATUS_BOARD.md | 4 +- ...-reference-sets-are-not-ordinal-aligned.md | 6 +-- .../plans/deepnsm-v2-lexical-address-v1.md | 38 +++++++++++++------ 4 files changed, 32 insertions(+), 18 deletions(-) diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index 1b5cd83b6..f9cd8dc41 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -8,7 +8,7 @@ carried; they remain on branch `claude/brave-mayer-65y3cy`. versioned lexical codebook. It is not two poles and not an embedding slice. - **Relations are driver-supplied** through the Fisher-z LUT. - **Three COCA references** (4096 / 5k lemma / 20k academic) are measured as neither - nested nor aligned: only 3 of 4,264 shared words keep their ordinal. Reading across + nested nor aligned: only 4 of 4,264 shared words keep their ordinal (ordinal 0 counted). Reading across references is refused. - **The COCA bake** stores frequency, PoS and lemma in the row as a u8 `` prior. Nothing is counted at runtime. diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index 86d2440a4..514713315 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -310,8 +310,8 @@ evaluates one; `execute` stays the consumer's call on a scratch it owns. | D-id | scope | status | gate / falsifier | |---|---|---|---| | D-LXA-1 | `LexicalAddress(u16)` newtype + `ReferenceSet {id, version, sha256}`; cross-reference read refused | Queued | G-LEX: every declared entry resolves to its (lemma, PoS); wrong reference refused; bare `u16` is `compile_fail` | -| D-LXA-2 | `lexical_correspondence.tsv` generator (id4096 / id5k / id20k / Exact·Ambiguous·Missing, three digests) | Queued | G-REF: re-derivation reproduces the measured table incl. "3 of 4,264"; one hand-edited ordinal reddens | -| D-LXA-3 | the COCA bake: (lemma, PoS) rows with `f` (reading share) and `c` (log-frequency evidence) baked as u8 offline; unknown flagged, never 0 | Queued — shares `academic_20k.csv` with D-LXC-4 (Blocked); must not duplicate that loader | byte-identical re-derivation; `the` f = 1, `record` f < 1; all-unambiguous fixture stays f = 1 | +| D-LXA-2 | `lexical_correspondence.tsv` generator (id4096 / id5k / id20k / Exact·Ambiguous·Missing, three digests) | Queued | G-REF: re-derivation reproduces the measured table incl. "4 of 4,264" with ordinal 0 counted; one hand-edited ordinal reddens | +| D-LXA-3 | the COCA bake: identity table (lemma, PoS, `c` log-frequency evidence) + surface-form table (form, reading, `f` reading share), u8, offline; unknown flagged, never 0 | Queued — shares `academic_20k.csv` with D-LXC-4 (Blocked); must not duplicate that loader | byte-identical re-derivation; surface `the` f = 1, surface `record` → n/v both 0 < f < 1; all-unambiguous fixture stays f = 1 | | D-LXA-4 | six-slot ClassView reading of a 12-byte facet as six `LexicalAddress`es under one `ReferenceSet` | Queued — home (second facet vs `Identity` tenant) OPEN | X-written / Y-read refused; slot rotation changes the resolved words | ## mask-risc-executor (PR3 — D-ids minted 2026-09-14, `.claude/plans/mask-risc-executor-v1.md`) diff --git a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md index 6f56a9501..ba768d8a1 100644 --- a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md +++ b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md @@ -5,14 +5,14 @@ ## Measured (python over the committed CSVs and the Tigris artifact) | reference | rows | distinct keys | key | |---|---|---|---| -| COCA4096 — `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank (homographs share a rank) | +| COCA4096 — `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank: one row per (word, PoS), each rank unique; a homograph spans several ranks | | COCA5K_LEMMA — `lemmas_5k.csv` | 5,050 | 4,380 lemmas / 5,050 (lemma, PoS) | (lemma, PoS) | | COCA20K_ACAD — `academic_20k.csv`; Tigris `lance-graph/codebooks/deepnsm-v2-academic-coca-v1/` | 20,845 | 18,559 words / 20,842 (word, Pos) | `word_id` = admission order | - 5k ∩ 20k: 4,895 of 5,050 (lemma, PoS); 4,264 words; **116 of the 5k lemmas are absent from the 20k**. - 4096 ∩ 20k: 3,462 of 3,559 words. -- **Ordinal alignment: 3 of 4,264 shared words carry the same ordinal in the 5k and in the 20k.** -- Ambiguity differs per set: the 20k carve drops 2,286 same-word-different-Pos duplicates to one id (MANIFEST: 18,559 of 20,480 reserved slots; basins 73..79 empty); the 4096 shares ranks across homographs; the 5k keeps (lemma, PoS) distinct. +- **Ordinal alignment: 4 of 4,264 shared words carry the same ordinal in the 5k and in the 20k** (`the`, `there`, `care`, `wage`; 0-based first-occurrence order, exact match). ⊘ First recorded as 3: that count dropped ordinal 0, a falsy-zero error (Codex, #1307). +- Ambiguity differs per set: the 20k carve drops 2,286 same-word-different-Pos duplicates to one id (MANIFEST: 18,559 of 20,480 reserved slots; basins 73..79 empty); the 4096 gives each (word, PoS) its own rank; the 5k keeps (lemma, PoS) distinct. - The Tigris "academic codebook" TSV is a `PaletteVocab::from_frequency_ranked` carve (`word_id, basin, slot, …`), not a trained centroid codebook. Its `academic_20k.csv` sha256 equals the committed file's. ## Consequence diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index a6eab1e0b..65f1e89fb 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -67,17 +67,21 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris | reference | source | rows | distinct keys | key | |---|---|---|---|---| -| `COCA4096` | `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank (homographs share a rank) | +| `COCA4096` | `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank: one row per (word, PoS), every rank unique, so a homograph occupies **several** ranks (`to/t` = 6, `to/i` = 12) | | `COCA5K_LEMMA` | `lemmas_5k.csv` | 5,050 | 4,380 lemmas / 5,050 (lemma, PoS) | (lemma, PoS) | | `COCA20K_ACAD` | `academic_20k.csv`; Tigris `lance-graph/codebooks/deepnsm-v2-academic-coca-v1/` | 20,845 | 18,559 words / 20,842 (word, Pos) | `word_id` = admission order | - 5k ∩ 20k = 4,895 of 5,050 (lemma, PoS) and 4,264 words. **116 of the 5k lemmas are absent from the 20k.** - 4096 ∩ 20k = 3,462 of 3,559 words. -- **Only 3 of the 4,264 shared words carry the same ordinal in the 5k and the 20k.** +- **Only 4 of the 4,264 shared words carry the same ordinal in the 5k and the 20k**: + `the`, `there`, `care` and `wage`. Ordinals are 0-based first-occurrence order, and + words are matched exactly. ⊘ #1303 reported "3": that count dropped ordinal 0 (`the`), + a falsy-zero error; `PaletteVocab` treats id 0 as valid. Lower-casing the 20k words + moves the shared count to 4,318 and leaves the aligned count at 4. - Each set handles homographs differently: - the 20k collapses 2,286 same-word-different-Pos duplicates into one id; - - the 4096 shares a rank across homographs; + - the 4096 gives each (word, PoS) its own rank, so a homograph spans several ranks; - the 5k keeps each (lemma, PoS) distinct. - The Tigris "academic codebook" is a vocabulary carve (`PaletteVocab::from_frequency_ranked`), not a trained centroid codebook. Its CSV's sha256 equals the committed file's. @@ -94,21 +98,31 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris ## §3 — The COCA bake (frequency, PoS and lemma built in; nothing counted at runtime) -Each codebook row is one `(lemma, PoS)` identity. Its fields are computed **once, -offline**, when the artifact is built. Float arithmetic is allowed at build time, which -is the derived side of the no-float rule. The results are stored quantized as `u8`: +There are **two baked tables**, because the two parts of the prior are keyed +differently: +- the **identity table**, one row per `(lemma, PoS)`, which the address resolves to; +- the **surface-form table**, one row per `(surface form, reading)`. The reading share + `f` belongs here, because it is a property of a *surface form*, not of a lemma. The + same `record/v` lemma row has a different share for `record`, `recorded` and + `recording`, so baking `f` onto the lemma row would make it depend on which form was + picked at build time. + +Every field is computed **once, offline**, when the artifact is built. Float arithmetic +is allowed at build time, which is the derived side of the no-float rule. The results +are stored quantized as `u8`: | field | how it is computed | why | |---|---|---| -| `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | -| **`f`** (u8) | the reading's share among its surface form's readings: `form_count(this) / surface_count(surface)` (from `LexicalEvidence`). `f = 1` for an unambiguous form | an ambiguous form (`record`) gets `f < 1` without anyone counting at query time | -| **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | +| identity: `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | +| identity: **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | +| surface form: **`f`** (u8) | the reading's share among that surface form's readings, from `word_forms.csv` (`wordFreq`): e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time | | unknown | a missing count stays **unknown** (a flag), never `0` | zero would assert "no evidence"; unknown asserts "not measured" | `range` and `disp` are left out. They are collinear with frequency, and `disp` measures evenness across genres, not evidence. -At runtime the prior `` is **a read of the resolved row**, never a computation. +At runtime the prior `` is **two reads**, never a computation: `f` from the +surface-form row that routed the token, `c` from the identity row it resolved to. Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11). --- @@ -118,8 +132,8 @@ Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11) | D-id | what | gate (can fire / can stay silent) | |---|---|---| | **D-LXA-1** | `LexicalAddress` newtype over `u16` plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX:** every declared entry of a reference resolves to exactly its declared (lemma, PoS). A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | -| **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "3 of 4,264". Hand-editing one ordinal must turn the check red | -| **D-LXA-3** | The COCA bake of §3: one artifact per reference, `f`/`c` baked as `u8`, unknown flagged | Re-deriving must reproduce the bake byte for byte. Pinned rows: `the` must have `f = 1` and `record` must have `f < 1`. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | +| **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "4 of 4,264" **with ordinal 0 counted** (a falsy-zero generator must fail it). Hand-editing one ordinal must turn the check red | +| **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `c`) and a surface-form table (`form`, reading, `f`), all `u8`, unknown flagged | Re-deriving must reproduce the bake byte for byte. Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | | **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`. A register in any other shape is refused, never reinterpreted | A facet written under reference X and read under Y is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag) | **Collision:** D-LXA-3 reads `academic_20k.csv`, which `D-LXC-4` (the academic loader, From f572882f5edff2f8c78eb924116731809c678915 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 22:12:53 +0000 Subject: [PATCH 3/7] plan lexical-address: canonical u8 encoding; per-reference resolution key - Bake encoding: q = round_half_even(x * 254) in f64, 255 = UNKNOWN sentinel; K recorded in the artifact header. - G-LEX checks each reference against its own key: COCA4096 -> (word, PoS), 5k -> (lemma, PoS), the academic carve -> (word, Ambiguous{PoS set}), since the carve merged 2,286 same-word rows. A PoS-exact academic reference is a new version, left open. - Entry: blank line before the table (MD058). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- ...-reference-sets-are-not-ordinal-aligned.md | 1 + .../plans/deepnsm-v2-lexical-address-v1.md | 29 +++++++++++++++++-- 2 files changed, 27 insertions(+), 3 deletions(-) diff --git a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md index ba768d8a1..cd63c9ab6 100644 --- a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md +++ b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md @@ -3,6 +3,7 @@ **Status:** MEASURED · OPEN — plan `.claude/plans/deepnsm-v2-lexical-address-v1.md` §2 (D-LXA-2, G-REF). Measured on PR #1303 (closed unmerged); carried here. ## Measured (python over the committed CSVs and the Tigris artifact) + | reference | rows | distinct keys | key | |---|---|---|---| | COCA4096 — `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank: one row per (word, PoS), each rank unique; a homograph spans several ranks | diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index 65f1e89fb..f95cb24b0 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -86,6 +86,19 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris - The Tigris "academic codebook" is a vocabulary carve (`PaletteVocab::from_frequency_ranked`), not a trained centroid codebook. Its CSV's sha256 equals the committed file's. +**What an address resolves to depends on the reference's own key:** + +| reference | an address resolves to | exactly one (lemma, PoS)? | +|---|---|---| +| `COCA4096` | one rank = one `(word, PoS)` row | yes; ranks are unique per `(word, PoS)` | +| `COCA5K_LEMMA` | one `(lemma, PoS)` row | yes | +| `COCA20K_ACAD` (Tigris carve v1) | one `word_id` = one **word**; the carve merged 2,286 same-word-different-PoS rows | **no.** It resolves to `(word, Ambiguous{PoS set})` | + +G-LEX checks each reference against **its own** declared key. It never demands a +(lemma, PoS) from a reference that does not carry one. A PoS-exact academic reference +is possible: its 20,842 distinct `(word, Pos)` rows fit in 16 bits. It would be minted +as a **new reference version**, not by reinterpreting the carve (§5). + **The rule:** - A reading contract names `ReferenceSet { id, version, sha256 }`. - An ordinal is stable **within** one reference, never across references. @@ -116,7 +129,16 @@ are stored quantized as `u8`: | identity: `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | | identity: **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | | surface form: **`f`** (u8) | the reading's share among that surface form's readings, from `word_forms.csv` (`wordFreq`): e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time | -| unknown | a missing count stays **unknown** (a flag), never `0` | zero would assert "no evidence"; unknown asserts "not measured" | +| unknown | a missing count stays **unknown**, never `0` | zero would assert "no evidence"; unknown asserts "not measured" | + +**The canonical `u8` encoding**, so the byte-for-byte gate is reproducible: +- `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 254)`, computed + in `f64`, which gives `0..=254`. +- **`255` is the UNKNOWN sentinel.** No computed value can produce it, so unknown needs + no separate flag byte and can never be confused with `x = 1` (`254`). +- `ln` is `f64::ln` and `freq_max` is the maximum over the reference being baked. `K` + is written into the artifact header next to the three digests, so it is part of what + is reproduced. `range` and `disp` are left out. They are collinear with frequency, and `disp` measures evenness across genres, not evidence. @@ -131,7 +153,7 @@ Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11) | D-id | what | gate (can fire / can stay silent) | |---|---|---| -| **D-LXA-1** | `LexicalAddress` newtype over `u16` plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX:** every declared entry of a reference resolves to exactly its declared (lemma, PoS). A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | +| **D-LXA-1** | `LexicalAddress` newtype over `u16` plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX:** every declared entry of a reference resolves to exactly its declared reading **under that reference's own key** (§2 table): `(word, PoS)`, `(lemma, PoS)`, or `(word, Ambiguous{PoS set})`. A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | | **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "4 of 4,264" **with ordinal 0 counted** (a falsy-zero generator must fail it). Hand-editing one ordinal must turn the check red | | **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `c`) and a surface-form table (`form`, reading, `f`), all `u8`, unknown flagged | Re-deriving must reproduce the bake byte for byte. Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | | **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`. A register in any other shape is refused, never reinterpreted | A facet written under reference X and read under Y is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag) | @@ -151,7 +173,8 @@ loader. - **Which relational reading(s) the driver selects** between two resolved words, and the LUT that carries them. That belongs to the ClassView and is not decided here. - **Does `COCA20K_ACAD` get its own codebook**, or does it share `COCA4096`'s through the - correspondence map? + correspondence map? And is a PoS-exact academic reference (20,842 `(word, Pos)` + rows) minted as a new version beside the carve? --- From 59401ffee85bd6f6c67dbd66530a5f701961df48 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 22:16:19 +0000 Subject: [PATCH 4/7] =?UTF-8?q?plan=20lexical-address:=20no=20sentinel=20i?= =?UTF-8?q?n=20the=20u8=20=E2=80=94=20unknown=20is=20a=20missing=20row?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 255 = UNKNOWN sentinel was a NaN smuggled into an unsigned byte and was never asked for. f and c now use the full range, round_half_even(x * 255); a reading without a measured count has no row. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .claude/plans/deepnsm-v2-lexical-address-v1.md | 11 ++++++----- 1 file changed, 6 insertions(+), 5 deletions(-) diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index f95cb24b0..4cd54a8a9 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -129,13 +129,14 @@ are stored quantized as `u8`: | identity: `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | | identity: **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | | surface form: **`f`** (u8) | the reading's share among that surface form's readings, from `word_forms.csv` (`wordFreq`): e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time | -| unknown | a missing count stays **unknown**, never `0` | zero would assert "no evidence"; unknown asserts "not measured" | +| unknown | a missing count produces **no row**. Unknown is absence, never a value, and never `0` | zero would assert "no evidence"; a missing row asserts "not measured" | **The canonical `u8` encoding**, so the byte-for-byte gate is reproducible: -- `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 254)`, computed - in `f64`, which gives `0..=254`. -- **`255` is the UNKNOWN sentinel.** No computed value can produce it, so unknown needs - no separate flag byte and can never be confused with `x = 1` (`254`). +- `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 255)`, computed + in `f64`. That uses the full unsigned range `0..=255`, and every byte value is a number. +- **There is no sentinel.** An unsigned byte carries no NaN-like "unknown" value. A + reading with no measured count simply has no row in the table, and a lookup that + finds no row is unknown. - `ln` is `f64::ln` and `freq_max` is the maximum over the reference being baked. `K` is written into the artifact header next to the three digests, so it is part of what is reproduced. From 18837669640c92b87d3a8a39a01f47be96102776 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 22:17:47 +0000 Subject: [PATCH 5/7] plan lexical-address: known/unknown is an alpha bit, not a missing row The baked table is the split tunnel's read-only spine, complete with one row per reference entry; a measured f/c is an overlay write at the same address, and "measured" is that row's bit in the AlphaMask. No sentinel, no holes, no second table. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .claude/plans/deepnsm-v2-lexical-address-v1.md | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index 4cd54a8a9..83e31b247 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -129,14 +129,24 @@ are stored quantized as `u8`: | identity: `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | | identity: **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | | surface form: **`f`** (u8) | the reading's share among that surface form's readings, from `word_forms.csv` (`wordFreq`): e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time | -| unknown | a missing count produces **no row**. Unknown is absence, never a value, and never `0` | zero would assert "no evidence"; a missing row asserts "not measured" | +| known / unknown | **one bit in the alpha mask** over the table (see below). Never a byte value and never `0` | zero would assert "no evidence"; an unset alpha bit asserts "not measured" | **The canonical `u8` encoding**, so the byte-for-byte gate is reproducible: - `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 255)`, computed in `f64`. That uses the full unsigned range `0..=255`, and every byte value is a number. -- **There is no sentinel.** An unsigned byte carries no NaN-like "unknown" value. A - reading with no measured count simply has no row in the table, and a lookup that - finds no row is unknown. +- **There is no sentinel, and no row is left out.** An unsigned byte carries no + NaN-like "unknown" value. +- **"Measured or not" uses the alpha channel's split tunnel** (`contract::alpha`, + `alpha_tunnel`): + - The baked table is the **spine**. It is complete, with one row per reference + entry, it is read-only, and it is read by every reader without a lock. + - Writing a measured `f` / `c` is an **overlay write at the same address**. The + spine never changes. + - Whether a row is measured is its bit in the overlay's **`AlphaMask`**. A reader + tests "known" as a mask operation, alongside every other mask it combines. +- Unknown is therefore an **unset alpha bit**. It is not a byte value, not a missing + row, and not a second table. The table keeps its shape, so addresses stay dense and + no reader ever handles a hole. - `ln` is `f64::ln` and `freq_max` is the maximum over the reference being baked. `K` is written into the artifact header next to the three digests, so it is part of what is reproduced. From 727a66da056c55b80514128d4ab7bc49ccd7149e Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 22:57:54 +0000 Subject: [PATCH 6/7] plan lexical-address: ratified v3 (5+3 council); escalate the alpha fit MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Two keys: WordId (surface form, existing) and LexicalAddress (reading, new), joined by readings(WordId). WordId is not superseded. - ReferenceSet gains key_kind; the classid reading mode pins the full (id, version, sha256), with a version-mismatch refusal gate. - c renamed lemma_evidence: (f, lemma_evidence) is not a truth pair and does not reuse TruthU8 (floor quantization, 128 sentinel). - f is the listed-reading share from form_count; fill byte 0 for unknown rows; one known mask per table; Option accessor. - §3.1: the alpha channel is defined as not a bake, so "known" cannot be an alpha bit as written. Options (a)-(d), recommendation (d). Escalated as ISS-LXA-ALPHA-FIT; D-LXA-3 blocked until ruled. - D-LXA-4 gated on its own contract plan; legacy Cam96/SpoFacet readings stay unchanged. - CodeRabbit: 2,286 = 2,283 extra (word, Pos) keys + 3 duplicate rows; INTEGRATION_PLANS names the two tables separately. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .claude/board/AGENT_LOG.md | 12 + .claude/board/INTEGRATION_PLANS.md | 10 +- .claude/board/ISSUES.md | 16 ++ .claude/board/STATUS_BOARD.md | 6 +- ...-reference-sets-are-not-ordinal-aligned.md | 2 +- .../plans/deepnsm-v2-lexical-address-v1.md | 228 +++++++++++++++--- 6 files changed, 239 insertions(+), 35 deletions(-) diff --git a/.claude/board/AGENT_LOG.md b/.claude/board/AGENT_LOG.md index dbf75b1f6..b450edea3 100644 --- a/.claude/board/AGENT_LOG.md +++ b/.claude/board/AGENT_LOG.md @@ -1,3 +1,15 @@ +## 2026-09-30 — 5+3 council on `deepnsm-v2-lexical-address-v1` (PR #1307) + +- The 5: prior-art, iron-rule, code truth, cascade-impact, creative-explorer. + 0 VIOLATES; every §2 number reproduced from the CSVs. Draft v2: 11 changes. +- The 3: overclaim-auditor, dilution-collapse-sentinel, firewall-warden. + 1 BLOCK (§3.1: option (a) called an alpha bit), 22 FIX. v3: 11 more changes — + two keys (`WordId` surface, `LexicalAddress` reading), `c` renamed + `lemma_evidence` (not a truth pair, no `TruthU8`), per-table known mask with + `Option`, §3.1 rewritten with option (d), D-LXA-4 gated on a contract plan. +- One conflict with a frozen decision (F4 vs the alpha channel's definition) is + escalated as `ISS-LXA-ALPHA-FIT`, not overridden. Plan only; no code. + ## 2026-09-30 — correction to the two D-LXC-1 / D-LXC-11 entries below - Their "70,396 triples", "G6 = 25" and "25 tags moved (G6 exact)" record diff --git a/.claude/board/INTEGRATION_PLANS.md b/.claude/board/INTEGRATION_PLANS.md index f9cd8dc41..5c22aee74 100644 --- a/.claude/board/INTEGRATION_PLANS.md +++ b/.claude/board/INTEGRATION_PLANS.md @@ -1,6 +1,7 @@ ## 2026-09-30 (1) — deepnsm-v2-lexical-address-v1 — a word is a 16-bit address into a versioned, baked COCA codebook → `.claude/plans/deepnsm-v2-lexical-address-v1.md` -**Status:** PROPOSAL (D-LXA-1..4). No code authorized. Harvested from PR #1303 +**Status:** PROPOSAL (D-LXA-1..4), ratified v3 by a 5+3 council. No code authorized. +D-LXA-3 is blocked on the operator escalation `ISS-LXA-ALPHA-FIT`. Harvested from PR #1303 (`deepnsm-v2-cam96-pairwise` v1–v5, closed unmerged). The v1–v5 plan files are not carried; they remain on branch `claude/brave-mayer-65y3cy`. @@ -10,8 +11,11 @@ carried; they remain on branch `claude/brave-mayer-65y3cy`. - **Three COCA references** (4096 / 5k lemma / 20k academic) are measured as neither nested nor aligned: only 4 of 4,264 shared words keep their ordinal (ordinal 0 counted). Reading across references is refused. -- **The COCA bake** stores frequency, PoS and lemma in the row as a u8 `` prior. - Nothing is counted at runtime. +- **The COCA bake** has two tables: an identity table `(lemma, PoS, lemma_evidence)` + and a surface-form table `(form, reading, f)`, all u8, built offline. Nothing is + counted at runtime. `f` and `lemma_evidence` are two statements, not a truth pair. +- **Two keys:** `WordId` (surface form, existing) and `LexicalAddress` (reading, new), + joined by `readings(WordId)`. - Carries `ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE`. ## 2026-09-29 (2) — deepnsm-v2-lexical-evidence-consumer-v1 rewritten against the DeepNSM → DeepNSM-v2 migration → `.claude/plans/deepnsm-v2-lexical-evidence-consumer-v1.md` diff --git a/.claude/board/ISSUES.md b/.claude/board/ISSUES.md index b61ad1cfb..ad6687c1e 100644 --- a/.claude/board/ISSUES.md +++ b/.claude/board/ISSUES.md @@ -1,8 +1,24 @@ +## ISS-LXA-ALPHA-FIT — known/unknown for the COCA bake vs the alpha channel's own definition (2026-09-30) + +**Status:** OPEN — operator escalation. **Basis:** VERIFIED-IN-CODE (5+3 council on +`deepnsm-v2-lexical-address-v1`, §3.1). +- The ruling says known/unknown uses the alpha channel split tunnel. The alpha channel + is defined as not a bake (no digest, discardable whole, `alpha.rs:11-16, 857-861`); + `claim` writes only a stamp (`alpha.rs:683-716`); the overlay sits only over + `&[NodeRow]` (`alpha.rs:531`); its bit means "attended". "Known" is a baked, digested + fact, and F3 (nothing counted at runtime) leaves the write side nothing to carry. +- Options (plan §3.1): (a) baked coverage mask shaped like `AlphaMask`; (b) codebook + entries as `NodeRow`s; (c) a non-`NodeRow` alpha overlay meaning "measured"; + (d) = (a) plus the unchanged overlay as the attention recorder over the codebook. + Council recommendation: (d). +- **What closes it:** the operator's choice. D-LXA-3 is blocked until then. + ## ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE — the driver packs p64 plane bits as a CausalMask; the inverse reads bit 2 differently (2026-09-30) **Status:** OPEN. **Basis:** VERIFIED-IN-CODE on `main` `0d31c54f`. `cognitive-shader-driver` `driver.rs` (emission stage) packs `CausalMask::from_bits(h.predicates & 0x07)`, where `predicates` is the p64 predicate-plane byte: bit0 CAUSES, bit1 ENABLES, **bit2 SUPPORTS** (`p64-bridge` `SUPPORTS = 2`). `p64-bridge::edge_to_layer_mask` maps mask bit 2 (confounding) to **CONTRADICTS**. Emit and its inverse disagree on bit 2; three mask vocabularies (rung, Pearl, p64) meet at one byte. - Found during PR #1303 (closed unmerged); recorded with `deepnsm-v2-lexical-address-v1.md` §6, which does not depend on it. - **What closes it:** one named mapping in `p64-bridge` used by both directions, with a round-trip test that can fail on bit 2. +- **Also (5+3 council, 2026-09-30):** sites `driver.rs:489`, `p64-bridge/src/lib.rs:74-75, 123-124`. Inference type 1 also sets SUPPORTS (`lib.rs:81`), and `driver.rs:706-720` keeps its own local predicate-bit table; the one mapping must cover both. ## ISS-REPORT-NO-COMPOSITE-KEY-GROUP-FOLD — a 2-D fold costs one pass per partition member (2026-09-23) diff --git a/.claude/board/STATUS_BOARD.md b/.claude/board/STATUS_BOARD.md index 514713315..834003b70 100644 --- a/.claude/board/STATUS_BOARD.md +++ b/.claude/board/STATUS_BOARD.md @@ -309,10 +309,10 @@ evaluates one; `execute` stays the consumer's call on a scratch it owns. | D-id | scope | status | gate / falsifier | |---|---|---|---| -| D-LXA-1 | `LexicalAddress(u16)` newtype + `ReferenceSet {id, version, sha256}`; cross-reference read refused | Queued | G-LEX: every declared entry resolves to its (lemma, PoS); wrong reference refused; bare `u16` is `compile_fail` | +| D-LXA-1 | `LexicalAddress(u16)` newtype (reading key, beside the surface-form `WordId`) + `ReferenceSet {id, version, sha256, key_kind}`; cross-reference read refused | Queued | G-LEX (typed values): every declared entry resolves under its reference's own key — `(word, PoS)`, `(lemma, PoS)` or `(word, Ambiguous)`; wrong reference refused; bare `u16` is `compile_fail` | | D-LXA-2 | `lexical_correspondence.tsv` generator (id4096 / id5k / id20k / Exact·Ambiguous·Missing, three digests) | Queued | G-REF: re-derivation reproduces the measured table incl. "4 of 4,264" with ordinal 0 counted; one hand-edited ordinal reddens | -| D-LXA-3 | the COCA bake: identity table (lemma, PoS, `c` log-frequency evidence) + surface-form table (form, reading, `f` reading share), u8, offline; unknown flagged, never 0 | Queued — shares `academic_20k.csv` with D-LXC-4 (Blocked); must not duplicate that loader | byte-identical re-derivation; surface `the` f = 1, surface `record` → n/v both 0 < f < 1; all-unambiguous fixture stays f = 1 | -| D-LXA-4 | six-slot ClassView reading of a 12-byte facet as six `LexicalAddress`es under one `ReferenceSet` | Queued — home (second facet vs `Identity` tenant) OPEN | X-written / Y-read refused; slot rotation changes the resolved words | +| D-LXA-3 | the COCA bake: identity table (lemma, PoS, `lemma_evidence`) + surface-form table (form, reading, `f` listed-reading share), u8, offline; unknown = unset known bit, fill byte 0 | Blocked — on operator escalation `ISS-LXA-ALPHA-FIT`, and shares `academic_20k.csv` with D-LXC-4 (Blocked); must not duplicate that loader | byte-identical re-derivation; surface `the` f = 1, surface `record` → n/v both 0 < f < 1; all-unambiguous fixture stays f = 1 | +| D-LXA-4 | six-slot ClassView reading of a 12-byte facet as six `LexicalAddress`es under one `ReferenceSet`, carried by a new classid / reading mode; existing `Cam96` / `SpoFacet` readings unchanged | Blocked — a contract change needing its own contract plan (not yet written); home (second facet vs `Identity` tenant) OPEN | X-written / Y-read refused; (X, v1)-written / (X, v2)-read refused; slot rotation changes the resolved words | ## mask-risc-executor (PR3 — D-ids minted 2026-09-14, `.claude/plans/mask-risc-executor-v1.md`) diff --git a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md index cd63c9ab6..e5125209d 100644 --- a/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md +++ b/.claude/board/entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md @@ -13,7 +13,7 @@ - 5k ∩ 20k: 4,895 of 5,050 (lemma, PoS); 4,264 words; **116 of the 5k lemmas are absent from the 20k**. - 4096 ∩ 20k: 3,462 of 3,559 words. - **Ordinal alignment: 4 of 4,264 shared words carry the same ordinal in the 5k and in the 20k** (`the`, `there`, `care`, `wage`; 0-based first-occurrence order, exact match). ⊘ First recorded as 3: that count dropped ordinal 0, a falsy-zero error (Codex, #1307). -- Ambiguity differs per set: the 20k carve drops 2,286 same-word-different-Pos duplicates to one id (MANIFEST: 18,559 of 20,480 reserved slots; basins 73..79 empty); the 4096 gives each (word, PoS) its own rank; the 5k keeps (lemma, PoS) distinct. +- Ambiguity differs per set: the 20k carve drops 2,286 rows to one id per word: 2,283 extra `(word, Pos)` keys plus 3 exact duplicate source rows (`wastewater/n`, `disproportionately/r`, `instill/v`) (MANIFEST: 18,559 of 20,480 reserved slots; basins 73..79 empty); the 4096 gives each (word, PoS) its own rank; the 5k keeps (lemma, PoS) distinct. - The Tigris "academic codebook" TSV is a `PaletteVocab::from_frequency_ranked` carve (`word_id, basin, slot, …`), not a trained centroid codebook. Its `academic_20k.csv` sha256 equals the committed file's. ## Consequence diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index 83e31b247..80a79f62a 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -1,6 +1,8 @@ # deepnsm-v2-lexical-address-v1 — a word is a 16-bit address into a versioned, baked COCA codebook > **Status:** PROPOSAL (D-LXA-1..4). Plan only; no code is authorized by this file. +> **Council:** 5+3, ratified v3 (2026-09-30). The change ledger is §7. §3.1 carries one +> **operator escalation** (the alpha-channel fit). D-LXA-3 does not start before it is ruled. > **Written against:** `main` `0d31c54f` (2026-09-30). > **Harvested from:** PR #1303 (`deepnsm-v2-cam96-pairwise-v5`), which was **closed without > merging**. Only four things from it are kept here: @@ -14,7 +16,7 @@ > itself), the §11/§11R execution socket, and the "baton" doc-comment edits. > **Board:** `STATUS_BOARD.md` § deepnsm-v2-lexical-address · entry > `entries/2026-09-30-three-reference-sets-are-not-ordinal-aligned.md` · `ISSUES.md` -> `ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE`. +> `ISS-CE64-EMIT-INVERSE-BIT2-DISAGREE`, `ISS-LXA-ALPHA-FIT` (the §3.1 escalation). --- @@ -58,6 +60,28 @@ released, or stored in Tigris ever assigns a word two needles: The phrase originated in #1303's v1 with no source. +**Prior art: `WordId` is a different key, and it stays.** `deepnsm-v2` already has a +16-bit id: `WordId = u16` in `PaletteVocab` (`vocab.rs:26,47`). It identifies a **surface +form**: `PaletteVocab` is a list of words, readings hang off it (`LexicalReading.form_count` +counts "occurrences of THIS surface form", `lexical.rs:116`), and a lemma "is NOT a +WordId" (`lexical.rs:93`). `LexicalAddress` identifies a **reading**, i.e. a `(lemma, PoS)` +identity. They are two keys: + +| type | identifies | keys | exists | +|---|---|---|---| +| `WordId` | a surface form | the surface-form table (`f`) | yes (`vocab.rs:26`) | +| `LexicalAddress` | a `(lemma, PoS)` reading | the identity table | new (D-LXA-1) | + +They are joined by `readings(WordId) → [LexicalAddress]`. `WordId` is not wrapped and not +superseded. + +`WordId`'s basin/identity split (`vocab.rs:7-21,31-40`) is **not** reading (2). It is a +positional split of one exact id: ids are assigned in frequency order, so the high byte +is a frequency band. The module doc itself says semantic distance "never [comes] from the +id arithmetic itself" (`vocab.rs:19-21`). §0's "the two bytes carry no separate meaning" +is a rule for `LexicalAddress`. It does not govern `WordId`, and nothing here edits +`vocab.rs`. + --- ## §2 — Three reference sets, never interchangeable (F23, measured) @@ -67,7 +91,7 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris | reference | source | rows | distinct keys | key | |---|---|---|---|---| -| `COCA4096` | `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank: one row per (word, PoS), every rank unique, so a homograph occupies **several** ranks (`to/t` = 6, `to/i` = 12) | +| `COCA4096` | `crates/deepnsm/word_frequency/word_rank_lookup.csv`, ranks ≤ 4096 | 4,096 ranks | 3,559 words | rank: one row per (word, PoS), every rank unique, so a homograph occupies **several** ranks (`to` has two rows, at ranks 6 and 12) | | `COCA5K_LEMMA` | `lemmas_5k.csv` | 5,050 | 4,380 lemmas / 5,050 (lemma, PoS) | (lemma, PoS) | | `COCA20K_ACAD` | `academic_20k.csv`; Tigris `lance-graph/codebooks/deepnsm-v2-academic-coca-v1/` | 20,845 | 18,559 words / 20,842 (word, Pos) | `word_id` = admission order | @@ -77,10 +101,13 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris - **Only 4 of the 4,264 shared words carry the same ordinal in the 5k and the 20k**: `the`, `there`, `care` and `wage`. Ordinals are 0-based first-occurrence order, and words are matched exactly. ⊘ #1303 reported "3": that count dropped ordinal 0 (`the`), - a falsy-zero error; `PaletteVocab` treats id 0 as valid. Lower-casing the 20k words - moves the shared count to 4,318 and leaves the aligned count at 4. + a falsy-zero error; `PaletteVocab` treats id 0 as valid. Lower-casing both lists moves + the shared count to 4,318 (lower-casing only the 20k gives 4,316). The aligned count + stays 4 either way. - Each set handles homographs differently: - - the 20k collapses 2,286 same-word-different-Pos duplicates into one id; + - the 20k collapses same-word-different-Pos rows into one id. 2,286 rows collapse + away: 2,283 extra `(word, Pos)` keys (20,842 − 18,559) plus 3 exact duplicate rows + (20,845 − 20,842: `wastewater/n`, `disproportionately/r`, `instill/v`); - the 4096 gives each (word, PoS) its own rank, so a homograph spans several ranks; - the 5k keeps each (lemma, PoS) distinct. - The Tigris "academic codebook" is a vocabulary carve (`PaletteVocab::from_frequency_ranked`), @@ -92,7 +119,7 @@ are neither nested nor aligned, measured over the committed CSVs and the Tigris |---|---|---| | `COCA4096` | one rank = one `(word, PoS)` row | yes; ranks are unique per `(word, PoS)` | | `COCA5K_LEMMA` | one `(lemma, PoS)` row | yes | -| `COCA20K_ACAD` (Tigris carve v1) | one `word_id` = one **word**; the carve merged 2,286 same-word-different-PoS rows | **no.** It resolves to `(word, Ambiguous{PoS set})` | +| `COCA20K_ACAD` (Tigris carve v1) | one `word_id` = one **word**; the carve merged 2,283 extra `(word, Pos)` keys and 3 duplicate rows | **no.** It resolves to `(word, Ambiguous{PoS set})` | G-LEX checks each reference against **its own** declared key. It never demands a (lemma, PoS) from a reference that does not carry one. A PoS-exact academic reference @@ -100,12 +127,24 @@ is possible: its 20,842 distinct `(word, Pos)` rows fit in 16 bits. It would be as a **new reference version**, not by reinterpreting the carve (§5). **The rule:** -- A reading contract names `ReferenceSet { id, version, sha256 }`. +- A reading contract names `ReferenceSet { id, version, sha256, key_kind }`, where + `key_kind ∈ {SurfaceKeyed, IdentityKeyed}`. `COCA20K_ACAD` (a `PaletteVocab` carve) is + surface-keyed. `COCA4096` and `COCA5K_LEMMA` are identity-keyed. The surface-form table + of §3 exists only for a surface-keyed reference, or through the correspondence map. +- The facet bytes carry **no** reference id and no version byte (F5: content-blind). The + reference is carried by the facet's **classid**, through the ClassView's reading mode. + That reading mode pins the **whole** `(id, version, sha256)`. A new codebook version + therefore means a new classid or reading mode, and a facet written under v1 and read + under v2 is refused. - An ordinal is stable **within** one reference, never across references. - Correspondence between references is an explicit (lemma, PoS) map with `Ambiguous{n}` and `Missing` as values. It is never assumed to be nested or aligned. - Reading an address against the wrong reference is **refused**. It is never answered with a different word. +- **The three resolution shapes stay visible in the type.** One `LexicalAddress` type + resolves to `Reading::Exact(word|lemma, PoS)` or `Reading::Ambiguous(word, PoS set)`. + Every caller must match both arms; there is no accessor that returns a single PoS + without handling `Ambiguous`. --- @@ -127,35 +166,117 @@ are stored quantized as `u8`: | field | how it is computed | why | |---|---|---| | identity: `lemma`, `PoS` | from `lemmas_5k.csv` (or the reference's own key) | the declared reading the address resolves to | -| identity: **`c`** (u8) | `c = w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading | -| surface form: **`f`** (u8) | the reading's share among that surface form's readings, from `word_forms.csv` (`wordFreq`): e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time | -| known / unknown | **one bit in the alpha mask** over the table (see below). Never a byte value and never `0` | zero would assert "no evidence"; an unset alpha bit asserts "not measured" | +| identity: **`lemma_evidence`** (u8) | `w / (w+1)`, `w = ln(1+freq) / ln(1+freq_max) · K` (`K` is a labelled, hand-tuned pin) | the amount of evidence behind the reading as a whole. It describes the lemma, not any one surface form's split, so it is **not** the `c` of a truth pair with `f` | +| surface form: **`f`** (u8) | the reading's **listed-reading share**: its share among the readings the source lists for that surface form, from `word_forms.csv` (`wordFreq`) via `LexicalReading.form_count`, not the floored cumulative `coverage` (`lexical.rs:33-39,377-383`). COCA lists are truncated, so `f = 1` means "the only listed reading", not "unambiguous". If any listed reading has no count (`form_count = None`, `lexical.rs:49-50`), the form's share is undefined and the form is marked unknown, never skipped: e.g. surface `record` → `record/n` 120,048 vs `record/v` 13,014, so `f` ≈ 0.90 and ≈ 0.10. `f = 1` for an unambiguous form (`the`) | an ambiguous form gets `f < 1` without anyone counting at query time. The per-form reading list already exists (`LexicalReading.form_count`, `lexical.rs:111-118`; `readings(id)`, `:373`); the bake computes the share from it. The reading-share key follows `deepnsm-v2-coverage-bands-v1` F8 | +| known / unknown | **one bit per row, outside the bytes** (see below and §3.1). Never a byte value and never `0` | zero would assert "no evidence"; an unset bit asserts "not measured". Same principle as `lexical.rs:46-52`, where unknown is `None`, not `0` | + +**`f` is a prior, never a selector.** It never picks a reading. D-LXC-1 / D-LXC-11 found +that choosing a tag by frequency is interference. The driver selects the reading; `f` is +only the prior evidence handed to it. + +**This is not a NARS truth, and it does not reuse `TruthU8`.** `f` answers "given this +surface form, how often is it this reading?". `lemma_evidence` answers "how much evidence +is there for this reading at all?". They are two statements with two evidence masses, on +two tables. `TruthU8 { frequency, confidence }` +(`lance-graph-arm-discovery/src/translator.rs:57-63`) is a truth about **one** statement, +and its encoding differs from the one below in three ways: +- it quantizes by floor division, `(x*255)/n` (`:77,81`); +- it uses `128` as an "unknown ≈ 0.5" sentinel (`:74-75`); +- it is documented as a substrate value, "not a wire DTO", with no LE codec (`:43-49`). + +D-LXA-3 defines its own byte encoding. + +`lexical.rs` counts evidence and states "Not truth … no NARS truth" (`lexical.rs:30-31`). +The bake is a **separate, derived artifact** that crosses that line on purpose, offline. +`lexical.rs` itself is unchanged. **The canonical `u8` encoding**, so the byte-for-byte gate is reproducible: - `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 255)`, computed in `f64`. That uses the full unsigned range `0..=255`, and every byte value is a number. - **There is no sentinel, and no row is left out.** An unsigned byte carries no - NaN-like "unknown" value. -- **"Measured or not" uses the alpha channel's split tunnel** (`contract::alpha`, - `alpha_tunnel`): - - The baked table is the **spine**. It is complete, with one row per reference - entry, it is read-only, and it is read by every reader without a lock. - - Writing a measured `f` / `c` is an **overlay write at the same address**. The - spine never changes. - - Whether a row is measured is its bit in the overlay's **`AlphaMask`**. A reader - tests "known" as a mask operation, alongside every other mask it combines. -- Unknown is therefore an **unset alpha bit**. It is not a byte value, not a missing - row, and not a second table. The table keeps its shape, so addresses stay dense and - no reader ever handles a hole. + NaN-like "unknown" value. A row whose value is unknown still holds a canonical fill + byte, `0`, so the byte-for-byte gate is reproducible. The fill means nothing on its + own: only the known bit says whether the byte is a value. +- **Measured or not is a bit outside the bytes** (operator ruling F4; the mechanism is + subject to §3.1). The baked table is the **spine**: complete, one row per reference + entry, read-only, read by every reader without a lock. Unknown is an **unset bit**. It + is not a byte value and not a missing row. The table keeps its shape, so addresses + stay dense and no reader ever handles a hole. +- There is **one known-mask per table**. The surface-form table and the identity table + are known independently. +- The known bit is forced in the type, like `Exact` / `Ambiguous`. The prior accessor + returns `Option` (the same rule as `lexical.rs:46-52`: unknown is not zero). A + reader cannot use `f` or `lemma_evidence` without having handled the unknown case. +- **How** the bit is carried is escalated (§3.1). The council found that the shipped + alpha API does not fit the first wording of this section. - `ln` is `f64::ln` and `freq_max` is the maximum over the reference being baked. `K` is written into the artifact header next to the three digests, so it is part of what is reproduced. +### §3.1 — Alpha-channel fit: ESCALATED to the operator (`ISS-LXA-ALPHA-FIT`) + +**The ruling, verbatim** (operator, 2026-09-30, on known versus unknown): *"we already have +alpha channel split tunnel trick for that"*. + +**What the code does.** `crates/lance-graph-contract/src/alpha.rs`, `alpha_tunnel.rs`: +- **The alpha channel is defined as not a bake.** It has "no bakes.tsv row, no digest", + and it is discardable whole (`alpha.rs:11-16`, `:857-861`). +- **The split tunnel** (`alpha_tunnel.rs:12-18`) means every lane reads the baked spine + without a lock, and writes go to an overlay at the same addresses. The overlay records + where attention went, one lane per rung. +- **`AlphaOverlay` only sits over `NodeRow` tables.** It is built through + `AlphaAllocation::over`, and only over `&[NodeRow]` (`alpha.rs:531`). +- **`claim` writes only a stamp** (`alpha.rs:683-716`). It has three outcomes: + - a fresh write of an `AlphaStamp` (cycle, seq, rung, visits) into value slot 0; + - on a revisit, `Ok { fresh: false }`, which increments `visits` in the existing row; + - `Unallocated` for an address outside the allocation. + + It never writes a caller's value. +- **The overlay bit means "attended".** It never means "measured". +- **`AlphaMask` is a plain, length-checked bitset.** It has `contains` (`:263`), `and` + (`:345`), `words` (`:478`) and a public constructor `from_words(words, len)` (`:492`). + Its single-bit `set` is private (`:255`). +- A sealed per-cycle persistence rule exists only in a plan + (`.claude/plans/spog-alpha-channel-v1.md`, its frozen decision F1; MedCare-side). It is not in this code. + +**The conflict.** Read literally, "known versus unknown is an alpha-channel bit" +contradicts the alpha channel's own definition. "Known" is a baked, digested, +reproducible fact. An alpha bit is ephemeral, undigested, and means "attended". There is +a second conflict with F3: nothing is counted at runtime, so no runtime write could ever +turn a row from unknown into known. The split tunnel's write side has nothing to carry +for this question. + +**Options:** +- **(a) A baked coverage mask.** The bake emits the table plus one bitset per table + (known = 1), built with `AlphaMask::from_words`. This does **not** implement the split + tunnel the ruling names: it keeps only the `AlphaMask` type and the baked-spine read + side. The mask is part of the bake (digested), so it must not be called an alpha bit. +- **(b) Codebook entries as `NodeRow`s,** with the known flag or values in a value tenant + after the stamp. This uses the real overlay, but it costs 512 B per word, needs a new + tenant, and its bit still means "attended". +- **(c) Extend the alpha API** with a non-`NodeRow` overlay whose bit means "measured". + This is a contract change that redefines alpha, and it needs its own plan. +- **(d) (a) plus the split tunnel for what it is for.** A baked coverage mask carries known + versus unknown. The unchanged alpha overlay is kept as the **attention recorder** over + the codebook spine: which lexical addresses the driver visited, per cycle, per rung. + Its addresses come from a `NodeRow` projection of the codebook, or later from a + non-`NodeRow` allocation. That is option (c) scoped to allocation only, with the + meaning unchanged. + +**Council recommendation: (d).** It keeps both legs of what the ruling points at, each +with the meaning its code gives it. It never stores a value in the overlay. The +refinement of `f` / `lemma_evidence` is never an overlay write (F3; `claim` is stamp-only). + +**The question to the operator:** confirm that "known" is a baked coverage mask shaped +like `AlphaMask` (options a or d), and that the alpha overlay keeps its "attended" +meaning. + `range` and `disp` are left out. They are collinear with frequency, and `disp` measures evenness across genres, not evidence. -At runtime the prior `` is **two reads**, never a computation: `f` from the -surface-form row that routed the token, `c` from the identity row it resolved to. +At runtime the prior is **two reads**, never a computation: `f` from the surface-form row +that routed the token (keyed by `WordId`), and `lemma_evidence` from the identity row it +resolved to (keyed by `LexicalAddress`). Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11). --- @@ -164,16 +285,21 @@ Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11) | D-id | what | gate (can fire / can stay silent) | |---|---|---| -| **D-LXA-1** | `LexicalAddress` newtype over `u16` plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX:** every declared entry of a reference resolves to exactly its declared reading **under that reference's own key** (§2 table): `(word, PoS)`, `(lemma, PoS)`, or `(word, Ambiguous{PoS set})`. A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | +| **D-LXA-1** | `LexicalAddress` newtype over `u16` (a **new** identity key beside the surface-form `WordId`, joined by `readings(WordId) → [LexicalAddress]`, §1) plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX** (over typed `LexicalAddress` values; nothing binds raw facet bytes to a reference before D-LXA-4): every declared entry of a reference resolves to exactly its declared reading **under that reference's own key** (§2 table): `(word, PoS)`, `(lemma, PoS)`, or `(word, Ambiguous{PoS set})`. A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | | **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "4 of 4,264" **with ordinal 0 counted** (a falsy-zero generator must fail it). Hand-editing one ordinal must turn the check red | -| **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `c`) and a surface-form table (`form`, reading, `f`), all `u8`, unknown flagged | Re-deriving must reproduce the bake byte for byte. Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | -| **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`. A register in any other shape is refused, never reinterpreted | A facet written under reference X and read under Y is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag) | +| **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `c`) and a surface-form table (`form`, reading, `f`), all `u8`, with the known bit carried as ruled in §3.1. `c` is baked, never derived from a runtime rung (a rung is not reproducible). **Blocked on §3.1 and on D-LXC-4** | Re-deriving must reproduce the bake byte for byte. Unknown rows hold fill byte `0` with the known bit unset. A form with one `form_count = None` reading must come out unknown (fires); the all-listed fixture must not (stays silent). Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | +| **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`, carried by the classid. A register in any other shape is refused, never reinterpreted. **This is a contract change, gated on its own contract plan (not yet written).** Today no reader can refuse: `SpoFacet::from_register` takes a bare `[u8; 12]` (`awareness_facet.rs:106`), and `Cam96 = [u8; 12]` (`space.rs:163`) appears 68 times in 12 files (grep, counting doc comments). `ReadMode` / `ValueSchema` must first gain a lexical reading. **Existing `Cam96` / `SpoFacet` classids keep their current reading unchanged. The lexical reading exists only under a newly minted classid / reading mode, and nothing re-reads existing rows** (I-LEGACY-API-FEATURE-GATED). D-LXA-1..3 ship without it | A facet written under reference X and read under Y is refused. A facet written under `(X, v1)` and read under `(X, v2)` is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag). These three gates move verbatim into the contract plan; the STATUS_BOARD row carries them until it exists | **Collision:** D-LXA-3 reads `academic_20k.csv`, which `D-LXC-4` (the academic loader, currently Blocked on a ruling about three duplicate (word, PoS) pairs) also owns. D-LXA-3 waits for that ruling, or takes it as its own first question. It must not duplicate the loader. +**Ownership:** the baked tables and coverage masks are read-only artifacts with no +mailbox. Nothing writes them at runtime. If §3.1 rules an alpha overlay over the codebook +(options b, c or d), the overlay's writer is the owning mailbox +(`SoaEnvelope::mailbox_owner`); a consumer never writes as itself. + --- ## §5 — Open @@ -195,8 +321,54 @@ loader. predicate-plane byte straight into a `CausalMask`: `driver.rs`, emission stage, `CausalMask::from_bits(h.predicates & 0x07)`. In that byte, bit 2 is **SUPPORTS** (`p64-bridge`: `SUPPORTS = 2`). But `edge_to_layer_mask` maps mask bit 2 to -**CONTRADICTS**. Emitting a mask and inverting it therefore disagree on bit 2. +**CONTRADICTS**. Emitting a mask and inverting it therefore disagree on bit 2. The +council confirmed each site: `driver.rs:489`, `p64-bridge/src/lib.rs:123-124` +(`SUPPORTS = 2`, `CONTRADICTS = 3`) and `:74-75`. It found one more writer: inference +type 1 also sets SUPPORTS (`lib.rs:81`). `driver.rs:706-720` keeps its own local +predicate-bit table, which the p64-bridge fix must also route through the one mapping. This plan does not depend on the defect. It is recorded because it is real and was found in the same read. The fix belongs to `p64-bridge`: one named mapping used in both directions, plus a round-trip test that can fail on bit 2. + +--- + +## §7 — 5+3 council change ledger (v1 → draft v2) + +The five were prior art, iron rules, code truth, cascade impact and different views. +Code truth reproduced every §2 number from the CSVs (gate G1). + +| # | change | source | +|---|---|---| +| L1 | 2,286 split into 2,283 extra `(word, Pos)` keys + 3 duplicate rows | code truth · CodeRabbit | +| L2 | `WordId` / `PaletteVocab` named as prior art | prior art | +| L3 | The reference id lives in the classid (ClassView reading mode), never in the bytes | iron rules · different views | +| L4 | The `Exact` / `Ambiguous` arms are forced in the type | different views | +| L5 | `f` is a prior, never a selector (D-LXC-1/-11) | prior art | +| L6 | `form_count` / `readings` and the coverage-bands F8 key cited | prior art · cascade | +| L7 | The alpha-overlay wording is withdrawn and **escalated** (§3.1): claim is stamp-only, typed over `NodeRow`, ephemeral, means "attended"; it also conflicts with F3 | code truth · iron rules · different views | +| L8 | Mailbox-owner sentence added | iron rules | +| L9 | D-LXA-4 recast as a contract change on its own plan; 1–3 ship without it | cascade | +| L10 | `c` is never derived from a runtime rung | different views | +| L11 | Bit 2: a second SUPPORTS writer (inference type 1) recorded | code truth | + +### v2 → v3 (the three reviewers: overclaim-auditor, dilution-collapse-sentinel, firewall-warden) + +Verdicts: 1 BLOCK (§3.1, dilution), 22 FIX, the rest PASS. All resolved here. + +| # | change | source | +|---|---|---| +| R1 | Two keys, two types. `WordId` identifies a surface form and keys `f`. `LexicalAddress` identifies a reading. Draft v2's "wrap or supersede `WordId`" and its `vocab.rs` supersession note are **withdrawn**: the byte split is positional, not reading (2) | dilution · overclaim · firewall | +| R2 | `ReferenceSet` gains `key_kind`. The ClassView reading mode pins `(id, version, sha256)`, with a version-mismatch refusal gate | dilution | +| R3 | `c` renamed `lemma_evidence`. `(f, lemma_evidence)` is not a truth pair and does not reuse `TruthU8`; the three encoding differences are named | dilution · overclaim | +| R4 | The bake is a derived artifact crossing `lexical.rs`'s no-truth line | overclaim | +| R5 | `f` is the listed-reading share, from `form_count`. A `None` count makes the form unknown, with a two-sided fixture | dilution | +| R6 | Canonical fill byte `0`, one known-mask per table, and `Option` forcing the unknown case | overclaim · dilution | +| R7 | §3.1 rewritten: the ruling quoted; alpha is defined as not a bake; the three `claim` outcomes; the sealed-batch sentence moved to its plan source; option (a) described as not a split tunnel; option (d) added and recommended; the conflict stated plainly (**the BLOCK**) | dilution · overclaim | +| R8 | D-LXA-4: legacy-reading invariant stated; `Cam96` count corrected to 68 in 12 files; gates carried on the STATUS_BOARD row until the contract plan exists; G-LEX scoped to typed values | firewall · overclaim · dilution | +| R9 | §2 wording: rank positions for `to`; 4,316 vs 4,318; the three duplicate rows named | overclaim | +| R10 | Bit-2 sites cited; `driver.rs:706-720` local table named | overclaim | +| R11 | Board: `ISSUES` row `ISS-LXA-ALPHA-FIT`, STATUS_BOARD rows updated, bit-2 issue extended | firewall | + +**Not adopted:** none. The one frozen-decision conflict (F4 against the alpha channel's +own definition) is escalated (`ISS-LXA-ALPHA-FIT`), not overridden. From 914316ce1b374c01b197cefd70223e0850013931 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 30 Sep 2026 23:03:40 +0000 Subject: [PATCH 7/7] plan lexical-address: lemma_evidence lookup for ambiguous (surface-keyed) addresses - An identity-keyed reference resolves an address to one identity row. - A surface-keyed reference (COCA20K_ACAD) resolves to (word, Ambiguous{PoS set}) and has no identity table of its own: each candidate reading's lemma_evidence comes through the D-LXA-2 correspondence map, Missing -> unknown, never aggregated. - D-LXA-3 and the encoding line use lemma_evidence, not c. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01DCEP2fZdHYdMCtcEVcTpS2 --- .../plans/deepnsm-v2-lexical-address-v1.md | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/.claude/plans/deepnsm-v2-lexical-address-v1.md b/.claude/plans/deepnsm-v2-lexical-address-v1.md index 80a79f62a..e9be9909a 100644 --- a/.claude/plans/deepnsm-v2-lexical-address-v1.md +++ b/.claude/plans/deepnsm-v2-lexical-address-v1.md @@ -191,7 +191,7 @@ The bake is a **separate, derived artifact** that crosses that line on purpose, `lexical.rs` itself is unchanged. **The canonical `u8` encoding**, so the byte-for-byte gate is reproducible: -- `f` and `c` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 255)`, computed +- `f` and `lemma_evidence` are in `[0, 1]`. Each is stored as `q = round_half_even(x × 255)`, computed in `f64`. That uses the full unsigned range `0..=255`, and every byte value is a number. - **There is no sentinel, and no row is left out.** An unsigned byte carries no NaN-like "unknown" value. A row whose value is unknown still holds a canonical fill @@ -275,8 +275,19 @@ meaning. evenness across genres, not evidence. At runtime the prior is **two reads**, never a computation: `f` from the surface-form row -that routed the token (keyed by `WordId`), and `lemma_evidence` from the identity row it -resolved to (keyed by `LexicalAddress`). +that routed the token (keyed by `WordId`), and `lemma_evidence` from the identity row of +each candidate reading. + +**Where the identity rows come from depends on the reference's key kind:** +- **Identity-keyed reference** (`COCA4096`, `COCA5K_LEMMA`). An address resolves to one + `(word|lemma, PoS)` row, and that row holds `lemma_evidence`. It is one read. +- **Surface-keyed reference** (`COCA20K_ACAD`). An address resolves to + `(word, Ambiguous{PoS set})`, so it selects **no single identity row**, and the + reference has no identity table of its own. Each candidate `(lemma, PoS)` in the set is + looked up through the correspondence map (D-LXA-2) in an identity-keyed reference. That + gives one `lemma_evidence` per candidate reading, or unknown for a `Missing` candidate. + The values are **never aggregated** into one number for the word; picking among them is + the driver's job, not the bake's (f is a prior, never a selector). Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11). --- @@ -287,7 +298,7 @@ Frequency rank stays a routing signal, not meaning (ρ ≈ −0.07, archive F11) |---|---|---| | **D-LXA-1** | `LexicalAddress` newtype over `u16` (a **new** identity key beside the surface-form `WordId`, joined by `readings(WordId) → [LexicalAddress]`, §1) plus `ReferenceSet { id ∈ {COCA4096, COCA5K_LEMMA, COCA20K_ACAD}, version, sha256 }`, in `deepnsm-v2`. Resolving through the wrong reference is a refusal. No bare `[u8; 12]` or `u16` enters the path | **G-LEX** (over typed `LexicalAddress` values; nothing binds raw facet bytes to a reference before D-LXA-4): every declared entry of a reference resolves to exactly its declared reading **under that reference's own key** (§2 table): `(word, PoS)`, `(lemma, PoS)`, or `(word, Ambiguous{PoS set})`. A cross-reference read is refused. A `compile_fail` test proves a bare `u16` is not accepted | | **D-LXA-2** | Generator for `lexical_correspondence.tsv`: one row per (lemma, PoS) with `id4096 \| id5k \| id20k \| status ∈ {Exact, Ambiguous{n}, Missing}`, the three digests in its header. Generated, never hand-edited | **G-REF:** re-deriving the file must reproduce §2's numbers exactly, including "4 of 4,264" **with ordinal 0 counted** (a falsy-zero generator must fail it). Hand-editing one ordinal must turn the check red | -| **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `c`) and a surface-form table (`form`, reading, `f`), all `u8`, with the known bit carried as ruled in §3.1. `c` is baked, never derived from a runtime rung (a rung is not reproducible). **Blocked on §3.1 and on D-LXC-4** | Re-deriving must reproduce the bake byte for byte. Unknown rows hold fill byte `0` with the known bit unset. A form with one `form_count = None` reading must come out unknown (fires); the all-listed fixture must not (stays silent). Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | +| **D-LXA-3** | The COCA bake of §3: per reference, an identity table (`lemma`, `PoS`, `lemma_evidence`) for each identity-keyed reference and a surface-form table (`form`, reading, `f`), all `u8`, with the known bit carried as ruled in §3.1. A surface-keyed reference reaches `lemma_evidence` per candidate reading through the D-LXA-2 correspondence map, never as one aggregated value (§3). `lemma_evidence` is baked, never derived from a runtime rung (a rung is not reproducible). **Blocked on §3.1 and on D-LXC-4** | Re-deriving must reproduce the bake byte for byte. Unknown rows hold fill byte `0` with the known bit unset. A form with one `form_count = None` reading must come out unknown (fires); the all-listed fixture must not (stays silent). Pinned rows: surface `the` → `the/a` with `f = 1`; surface `record` → `record/n` with `f < 1` and `record/v` with `f > 0`, the two summing to 1 within quantization. An all-unambiguous fixture must give `f = 1` everywhere (stays silent) | | **D-LXA-4** | The six-slot reading: a ClassView-selected reading of a 12-byte facet as six `LexicalAddress`es under one named `ReferenceSet`, carried by the classid. A register in any other shape is refused, never reinterpreted. **This is a contract change, gated on its own contract plan (not yet written).** Today no reader can refuse: `SpoFacet::from_register` takes a bare `[u8; 12]` (`awareness_facet.rs:106`), and `Cam96 = [u8; 12]` (`space.rs:163`) appears 68 times in 12 files (grep, counting doc comments). `ReadMode` / `ValueSchema` must first gain a lexical reading. **Existing `Cam96` / `SpoFacet` classids keep their current reading unchanged. The lexical reading exists only under a newly minted classid / reading mode, and nothing re-reads existing rows** (I-LEGACY-API-FEATURE-GATED). D-LXA-1..3 ship without it | A facet written under reference X and read under Y is refused. A facet written under `(X, v1)` and read under `(X, v2)` is refused. Rotating the six slots changes the resolved words (this proves the six positions are ordered, not a bag). These three gates move verbatim into the contract plan; the STATUS_BOARD row carries them until it exists | **Collision:** D-LXA-3 reads `academic_20k.csv`, which `D-LXC-4` (the academic loader,