Repository navigation
Add UK target-weight rules, size-stage L2 and a selection experiment harness (#1124) - #1163
juaristi22 wants to merge 21 commits into
Conversation
grain_family_equal, nation_grain_family_equal and their sqrt-count variants, computed from one row carrier built the same way from declared targets and from a stored problem; grain_equal and uniform unchanged bit for bit; the doctrine default and the release-candidate refusal are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the UK size chain (#1124) Search and refit each take a chi-square L2 penalty, off by default and then byte-identical; reuse checks bind a selection's L2 and loss weights; the size checkpoint binds the selection L2 only when it is on; a typed stage weighting lets the experiment harness refit under a named rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reads a finished size build's flat evidence and pool read-only, reproduces the stored refit as the control, runs refit, re-selection and refit-level holdout experiments, and scores D, S0 and candidates against the pre-registered acceptance criteria; engine-free tests on a synthetic finished run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd graph (#1124) prepare_uk_full_solve and the rotated holdout compute the new rules from the solve's own row carrier; the problem binds a weight receipt; posture, CLI, schema and scorer admit the names (national role and release candidates still refuse them); the problem and holdout kernels hash the modules that build the weights. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dout (#1124) Six candidate-only flags, validated and refused by the national role and the release-candidate pre-flight; graph config fields that add node parameters only when on; the doctrine solve and the rotated holdout pass them through. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
README with the ladder, selection rules, acceptance criteria and predictions; the step-1a grid as harness JSON; the run and analysis scripts. Nothing run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#1124) A size build made with --selection-l2-lambda or --refit-l2-lambda stores those penalties in its search options and size receipt. The harness now reads them into the baseline settings: the control re-runs the stored refit under the stored refit L2, and every refit on the stored search names the search's own L2 so the reuse check accepts it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s as results (#1124) The plan's step 0 had no code. `run_uk_size_census` (tool: `census`) reads the stored support and its caps without solving: the refit start's and the stretch cap's mass against D by nation and household type at floors 0, 0.1, 0.5 and 1, per-area ceilings under the relative-collapse rule, the search's cap share and capacity bound k_min, the penalty pressure lambda*P/L of the L2 grid per anchor, each rule's loss shares, and the early size triggers. It checks itself against the stored run (S0 never above its cap, the recomputed start matches the receipt). `run` now records a configuration the solver chain refuses as a failed receipt and moves on, instead of aborting the grid; failed receipts are never scored and are published without their traceback. A refit-only `epochs` override lets a smoke run exercise the setup on a real build cheaply. Scorecard: household type in the pool profile (from `ons_household_type`) and in the representativeness block, nation capacity under the cap per set, a float32-aware tolerance for "on the cap", and publication suppresses small effective sample sizes and median unit counts as well as small counts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nfigurations `run.sh` runs `census` after the control. The README records what step 0's census measures and that no release gate runs in the harness (a refused configuration is a failed receipt, not a stop). `analyze.py` lists refused configurations and summarises the census beside the step-1 table. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`uk_size_step1b_plan` (tool: `plan-step1b`) applies the experiment's selection rules to the scored step-1a refits, so step 1b runs straight after 1a: stage `ae` writes the A×E points (the best A rule x each anchor x the knee of that anchor's floor-0.5 L2 ladder and its two half-decade neighbours, each lambda multiplied by the rule's loss ratio at S0 from the census, at floor 0.5); stage `holdout`, once those are scored, writes the refit-level holdouts of C0, C2, the best E per anchor, the best A and the best A×E. Each pick is recorded with the candidates it was read from, and a pick with no candidate is skipped. Publication includes the plans. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fter 1a The README records the readings `plan-step1b` applies (María's go for steps 0-1, 2026-10-08): best E and the knee on each anchor's floor-0.5 ladder, A×E at floor 0.5 at the knee and its half-decade neighbours, the winner rule's criterion-6 count when none passes it, and the lambda grid kept as pre-registered. The knee rule's "largest" is corrected to "smallest" before any result exists: "largest" makes the knee the ladder's top whenever collapse falls with lambda. `run.sh` now runs steps 0, 1a and 1b in one chain; `analyze.py` reports the step-1b picks and the holdouts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`publish` reduces every `run_dir` in the receipts and the census to the run directory's name, so a published results file names no machine. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Steps 0, 1a and 1b ran from the run tree (v6's code 224d93ee6 plus this branch's #1124 commits) on runs/uk-local-k25-675d2a3fd, 2026-10-07 23:48Z to 2026-10-08 04:30Z: the control reproduced the stored refit bit for bit; 26 step-1a refits, 6 A×E points and 4 refit-level holdouts, none refused. results/results.json holds the disclosure-controlled aggregates; the README's results section is written by analyze.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chosen after step 1 (María, 2026-10-08) and labelled as not pre-registered: grain_family_equal_sqrt_count at floor 0.5 with the uniform anchor at lambda 1e-3, 3e-3 and 1e-2 times its loss ratio at S0, the low-lambda region step 1b's knee rule skipped, plus the holdout of E_unif_f0.5_1e-2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
grain_family_equal_sqrt_count at floor 0.5 with the uniform anchor at lambda 0.0021, 0.0064 and 0.021, and the holdout of E_unif_f0.5_1e-2, re-scored and re-published with steps 0-1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#1124) Step 2 runs a 5-refit grid and holdouts on each new search's selection; until now selection mode kept only its own refit, so every further refit would have searched again (hours each). A selection-mode experiment now saves its search and draw under <out>/<name>/selection in the graph's own encodings, with a manifest of the search's rule, selection L2 and file digests. A refit or refit_holdout experiment names it with `selection_from` and runs on that selection, passing the search's weighting and L2 so the reuse check binds; a refit with the search's own settings reproduces the search's refit bit for bit. The tool checks every `selection_from` up front, and a failed search leaves no saved selection, so its dependents fail with that reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
María's go of 2026-10-08. S-A searches under grain_family_equal_sqrt_count (warm-started at the stored penalty times the rule's loss ratio) and refits five ways on its saved selection; S-AE adds selection-stage L2 at the pool-design anchor and runs only when no S-A configuration passes all six criteria. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
main now carries #1115's warm start, which adds the shared initial_lambda to the search, draw and refit parameters; the default-parameters test names it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The search under grain_family_equal_sqrt_count (2026-10-08 21:12Z to 2026-10-09 06:11Z, 4 probes, selected lambda 1.078e-6) and its five refits on the saved selection, re-scored and re-published with steps 0-1c. S-AE was stopped during its search and left no result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Automated review pass (Claude Code, high effort) — round 1 at
|
María's request of 2026-10-09: the S-A search and its five refits under nation_grain_family_equal_sqrt_count, which gives national rows half the loss, warm-started at S-A's selected penalty scaled to this rule, and the holdout of S-A's best configuration on its saved selection. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The search under nation_grain_family_equal_sqrt_count (2026-10-10 00:26Z to 09:44Z, 4 probes, selected lambda 1.0226e-6), its five refits on the saved selection and the 5-fold holdout of S-A's best configuration (SA_f0.5_u0.021), re-scored and re-published with the earlier steps. None was refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Strategy comparison so far (v6 K=25 baseline, 60,000 households)The table shows one or two configurations per strategy, each scored against the same dense solve D. The full scorecard of all 49 sets is in the experiment README.
Reading the table
What it shows
|
Summary
Part of #1124. This PR adds the code the #1124 experiments run on and records their results so far:
Default builds are unchanged: their size-node parameters and checkpoint identities stay as they were. Release candidates still refuse the new rules and penalties.
Best combination so far
"Best" here means the configuration that passes the most acceptance criteria, ties broken by the fewest areas below 25 % of their dense ESS. That is
SA_f0.5_u0.021:grain_family_equal_sqrt_count(step 2's S-A). It selected L0 λ 1.078e-6, giving 60,106 open-probability mass for the 60,000 households.grain_equalat S0.Against the dense solve D and today's 60k file S0:
The remaining national misses concentrate in HMRC top-income bands. Under family-equal weighting, each row of a large national family carries less of the loss: the HMRC families' 662 rows fall from 18.9 % to 6.7 % of the total. The √count variant then tilts count rows toward large counts. Its 5-fold refit-level holdout (2026-10-10) gives a held-out loss of 0.136 under
grain_equal, with 65.6 % of held local rows within 10 % and 86.5 % within 25 %. On the same folds, today's file S0 scores 0.115, 70.5 % and 89.7 %, and the floor alone (C2) scores 0.202, 53.4 % and 76.9 %.Results so far
grain_family_equal_sqrt_count, 649 of 650 constituencies fall under ESS 50.step2_san.json; 4 probes, 8.4 h, selected λ 1.0226e-6): the same design undernation_grain_family_equal_sqrt_count, which gives national rows half the loss.grain_equalat floor 0.5 with uniform L2 at 1e-2 on this selection passes criteria 1, 4 and 5 (5 national rows past 25 %, households −0.30 %). It misses NI (2.44 %) and the lone-person share (27.23 %, 1.40 points below D), and leaves 107 areas below 25 % of dense ESS.The full tables are in
experiments/uk-sparse-selection-ae-20261007/README.md. The aggregates are in itsresults/results.json, which is disclosure-controlled and holds no unit records.What the PR adds
uk_runtime/target_weights.py:grain_family_equal,nation_grain_family_equaland their_sqrt_countvariants;UK_WIDE_GEOGRAPHY_IDSjoinsgeography_constants.uk_runtime/dataset_size.py,size_checkpoint.pyand the graph/CLI wiring:--{selection,refit}-l2-{lambda,anchor,basis}: chi-square basis,initialoruniformanchor;uk_runtime/size_experiment.py,size_experiment_scorecard.pyandtools/run_uk_size_experiment.py:selection_from);experiments/uk-sparse-selection-ae-20261007/:experiments.json,step1c_exploratory.json,step2_sa.jsonandstep2_sae.json;run.sh,analyze.pyandresults/.Running it on another machine
Data (UKDS-licensed, never in git): copy the v6 baseline run directory's
evidence-index.json, the flat artifacts it lists,numerical.graph.jsonand the pool frame object.graph-store/objects/27/2778ba74e3…, about 7 GB in all. To continue in the same out directory as these results, also copydata/ukds/acceptance/1124-sparse-solve/v6/step01. It holds the receipts, the weights and S-A's saved selection.Environment:
uv sync --all-packages --locked --extra uk.Commands: one solve at a time (
--confirm-exclusive), with out directories outside the repository and the run:Portability check: from this branch, rebased on
main36557ca, the control reproduces v6's stored refit bit for bit (checked 2026-10-09: exact, maximum weight difference 0.0, every receipt field equal). That holds even though the experiments ran from a tree pinned to v6's own code.Costs on a 24 GiB Mac:
Open decisions
step2_sae.json. It is pre-registered to warm-start like S-A, but S-A's selected λ (1.078e-6) would be a closer start.Testing
tools/ci_test_plan.py verifypasses.ruff checkandruff format --checkpass on the changed files.🤖 Generated with Claude Code