Skip to content

Add UK target-weight rules, size-stage L2 and a selection experiment harness (#1124) - #1163

Draft
juaristi22 wants to merge 21 commits into
mainfrom
uk-sparse-selection-ae-1124
Draft

juaristi22 wants to merge 21 commits into
mainfrom
uk-sparse-selection-ae-1124

Conversation

@juaristi22

@juaristi22 juaristi22 commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Part of #1124. This PR adds the code the #1124 experiments run on and records their results so far:

  • named row-metadata target-weight rules for the UK local solve (lever A);
  • per-stage L2 penalties for the dataset-size search and refit (lever E);
  • a selection-only experiment harness that varies both on a finished size build, with the dense solve held fixed;
  • the pre-registered experiment folder with its results on the v6 K=25 baseline (60,000 households).

Default builds are unchanged: their size-node parameters and checkpoint identities stay as they were. Release candidates still refuse the new rules and penalties.

Best combination so far

"Best" here means the configuration that passes the most acceptance criteria, ties broken by the fewest areas below 25 % of their dense ESS. That is SA_f0.5_u0.021:

  • Selection: the size search re-run under grain_family_equal_sqrt_count (step 2's S-A). It selected L0 λ 1.078e-6, giving 60,106 open-probability mass for the 60,000 households.
  • Refit on that selection: the same rule, refit baseline floor 0.5, and chi-square L2 toward uniform weights at λ 0.0214. That is 1e-2 × 2.14, the rule's loss ratio to grain_equal at S0.

Against the dense solve D and today's 60k file S0:

  • Criterion 1 (passes): households −0.34 % of D (S0: −3.47 %).
  • Criterion 2 (passes): Northern Ireland share 2.568 % (D: 2.686 %, floor 2.552 %; S0: 2.01 %).
  • Criterion 3 (passes): lone-person share 27.82 %, 0.80 points below D (S0: 23.04 %).
  • Criterion 4 (passes): no local family loses more than a point of within-10 % fit against S0.
  • Criterion 5 (fails): 35 national rows miss by more than 25 % (D: 3, limit 5; S0: 15).
  • Criterion 6 (fails): 93 constituencies and local authorities fall below 25 % of their dense ESS (S0: 158). 33 constituencies and 13 local authorities sit under ESS 50 (S0: 170 and 20).

The remaining national misses concentrate in HMRC top-income bands. Under family-equal weighting, each row of a large national family carries less of the loss: the HMRC families' 662 rows fall from 18.9 % to 6.7 % of the total. The √count variant then tilts count rows toward large counts. Its 5-fold refit-level holdout (2026-10-10) gives a held-out loss of 0.136 under grain_equal, with 65.6 % of held local rows within 10 % and 86.5 % within 25 %. On the same folds, today's file S0 scores 0.115, 70.5 % and 89.7 %, and the floor alone (C2) scores 0.202, 53.4 % and 76.9 %.

Results so far

  • Step 0:
    • The harness re-ran the stored refit S0 bit for bit.
    • The census found 18 constituencies and 11 local authorities whose kept households cannot reach 25 % of their dense ESS even at equal weights. So no refit on the stored selection can pass criterion 6.
  • Steps 1a and 1b (32 refits and 4 holdouts on the stored selection): no configuration passes criteria 1–5.
    • The refit baseline floor and the new rules restore mass and composition by concentrating weight. At floor 0.5 under grain_family_equal_sqrt_count, 649 of 650 constituencies fall under ESS 50.
    • Uniform L2 spreads the weight and hands back the selection's own composition. At λ 0.1 the lone-person share falls to 21.8 %.
  • Step 1c (exploratory, labelled as such; 3 refits and 1 holdout): the rule with gentle uniform L2 keeps the composition only where ESS still collapses.
  • Step 2, S-A: re-running the search under the rule changes which households are kept.
    • With S0's own refit settings, the new selection lifts the lone-person share from 23.0 % to 26.7 %. It also leaves 47 national rows past 25 %, against S0's 15.
    • The rule refit above passes criteria 1–4. National fit (criterion 5) becomes the binding failure.
  • Step 2, S-A under the nation-grain rule (step2_san.json; 4 probes, 8.4 h, selected λ 1.0226e-6): the same design under nation_grain_family_equal_sqrt_count, which gives national rows half the loss.
    • At floor 0 its selection fits worse than S-A's: households −10.2 % under the rule's own refit and −12.6 % with S0's refit settings, with 95–113 national rows past 25 %.
    • At floor 0.5 without L2 it fits about as well as S-A: it passes criteria 1–4 with 9 national rows past 25 % (S-A: 14), and still collapses ESS (888 areas below 25 % of dense ESS).
    • With uniform L2 it misses fewer national rows than S-A at the same penalty (18 and 29, against 28 and 35), but loses the NI criterion (2.53 % and 2.41 %, against the 2.552 % needed).
    • grain_equal at floor 0.5 with uniform L2 at 1e-2 on this selection passes criteria 1, 4 and 5 (5 national rows past 25 %, households −0.30 %). It misses NI (2.44 %) and the lone-person share (27.23 %, 1.40 points below D), and leaves 107 areas below 25 % of dense ESS.
    • So the nation-grain rule moves the remaining misses from national rows to the NI share; it doesn't remove them. The national misses that remain are mostly HMRC top-income bands and SPI regional £200k+ taxpayer counts.
  • S-AE (S-A plus selection-stage L2) was stopped during its search and has no result.

The full tables are in experiments/uk-sparse-selection-ae-20261007/README.md. The aggregates are in its results/results.json, which is disclosure-controlled and holds no unit records.

What the PR adds

  • uk_runtime/target_weights.py:
    • the row carrier and four rules: grain_family_equal, nation_grain_family_equal and their _sqrt_count variants;
    • the rules are bound through the solve, holdout, graph, CLI, diagnostics schema and scorer;
    • UK_WIDE_GEOGRAPHY_IDS joins geography_constants.
  • uk_runtime/dataset_size.py, size_checkpoint.py and the graph/CLI wiring:
    • --{selection,refit}-l2-{lambda,anchor,basis}: chi-square basis, initial or uniform anchor;
    • the settings enter node parameters and the checkpoint identity only when on.
  • uk_runtime/size_experiment.py, size_experiment_scorecard.py and tools/run_uk_size_experiment.py:
    • the cache, the bit-exact control and the census;
    • refit, selection and holdout modes, and saved selections (selection_from);
    • the step-1b planner, the scorecard with its six pre-registered criteria, and disclosure-controlled publication.
    • No release gate runs in the harness. A configuration the solver chain refuses is recorded as a failed receipt, and the run moves on.
  • experiments/uk-sparse-selection-ae-20261007/:
    • the pre-registration, with dated amendments;
    • experiments.json, step1c_exploratory.json, step2_sa.json and step2_sae.json;
    • run.sh, analyze.py and results/.

Running it on another machine

  • Data (UKDS-licensed, never in git): copy the v6 baseline run directory's evidence-index.json, the flat artifacts it lists, numerical.graph.json and the pool frame object .graph-store/objects/27/2778ba74e3…, about 7 GB in all. To continue in the same out directory as these results, also copy data/ukds/acceptance/1124-sparse-solve/v6/step01. It holds the receipts, the weights and S-A's saved selection.

  • Environment: uv sync --all-packages --locked --extra uk.

  • Commands: one solve at a time (--confirm-exclusive), with out directories outside the repository and the run:

    python tools/run_uk_size_experiment.py build-cache --run-dir RUN --cache-dir CACHE --confirm-exclusive
    python tools/run_uk_size_experiment.py control --run-dir RUN --cache-dir CACHE --out OUT --confirm-exclusive
    python tools/run_uk_size_experiment.py run --run-dir RUN --cache-dir CACHE --out OUT --confirm-exclusive --max-rss-gib 18 --experiments experiments/uk-sparse-selection-ae-20261007/step2_sae.json
    python tools/run_uk_size_experiment.py score --run-dir RUN --cache-dir CACHE --out OUT
    python tools/run_uk_size_experiment.py publish --out OUT --to experiments/uk-sparse-selection-ae-20261007/results
    python experiments/uk-sparse-selection-ae-20261007/analyze.py
  • Portability check: from this branch, rebased on main 36557ca, the control reproduces v6's stored refit bit for bit (checked 2026-10-09: exact, maximum weight difference 0.0, every receipt field equal). That holds even though the experiments ran from a tree pinned to v6's own code.

  • Costs on a 24 GiB Mac:

    • the control takes 8 min;
    • a search takes 4–9 h (S-A took 8.5 h over 4 probes);
    • a refit takes 5 min;
    • peak memory is about 11 GiB.

Open decisions

  • S-AE: re-run step2_sae.json. It is pre-registered to warm-start like S-A, but S-A's selected λ (1.078e-6) would be a closer start.
  • National fit: the nation-grain search (above) cut national misses at the cost of the NI share. The misses that remain sit in the HMRC top-income bands and SPI regional £200k+ counts, which family-equal weighting dilutes in both rules.
  • Holdouts of the S-A configurations take about 28 min each on the saved selection.
  • Dataset size (step 3) comes into play if criterion 6 stays out of reach on new selections.

Testing

  • 274 targeted engine-free tests pass on the rebased branch. They cover the size experiment, L2, the rules and their wiring, the full build CLI, the calibration graph, the graph terminal, local rowwise, the posture, the preflight and diagnostics.
  • tools/ci_test_plan.py verify passes.
  • ruff check and ruff format --check pass on the changed files.

🤖 Generated with Claude Code

juaristi22 and others added 19 commits October 9, 2026 11:30
grain_family_equal, nation_grain_family_equal and their sqrt-count variants,
computed from one row carrier built the same way from declared targets and
from a stored problem; grain_equal and uniform unchanged bit for bit; the
doctrine default and the release-candidate refusal are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the UK size chain (#1124)

Search and refit each take a chi-square L2 penalty, off by default and then
byte-identical; reuse checks bind a selection's L2 and loss weights; the size
checkpoint binds the selection L2 only when it is on; a typed stage weighting
lets the experiment harness refit under a named rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reads a finished size build's flat evidence and pool read-only, reproduces the
stored refit as the control, runs refit, re-selection and refit-level holdout
experiments, and scores D, S0 and candidates against the pre-registered
acceptance criteria; engine-free tests on a synthetic finished run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd graph (#1124)

prepare_uk_full_solve and the rotated holdout compute the new rules from the
solve's own row carrier; the problem binds a weight receipt; posture, CLI,
schema and scorer admit the names (national role and release candidates still
refuse them); the problem and holdout kernels hash the modules that build the
weights.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dout (#1124)

Six candidate-only flags, validated and refused by the national role and the
release-candidate pre-flight; graph config fields that add node parameters
only when on; the doctrine solve and the rotated holdout pass them through.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
README with the ladder, selection rules, acceptance criteria and predictions;
the step-1a grid as harness JSON; the run and analysis scripts. Nothing run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#1124)

A size build made with --selection-l2-lambda or --refit-l2-lambda stores those
penalties in its search options and size receipt. The harness now reads them
into the baseline settings: the control re-runs the stored refit under the
stored refit L2, and every refit on the stored search names the search's own L2
so the reuse check accepts it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s as results (#1124)

The plan's step 0 had no code. `run_uk_size_census` (tool: `census`) reads the
stored support and its caps without solving: the refit start's and the stretch
cap's mass against D by nation and household type at floors 0, 0.1, 0.5 and 1,
per-area ceilings under the relative-collapse rule, the search's cap share and
capacity bound k_min, the penalty pressure lambda*P/L of the L2 grid per
anchor, each rule's loss shares, and the early size triggers. It checks itself
against the stored run (S0 never above its cap, the recomputed start matches
the receipt).

`run` now records a configuration the solver chain refuses as a failed
receipt and moves on, instead of aborting the grid; failed receipts are never
scored and are published without their traceback. A refit-only `epochs`
override lets a smoke run exercise the setup on a real build cheaply.

Scorecard: household type in the pool profile (from `ons_household_type`) and
in the representativeness block, nation capacity under the cap per set, a
float32-aware tolerance for "on the cap", and publication suppresses small
effective sample sizes and median unit counts as well as small counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nfigurations

`run.sh` runs `census` after the control. The README records what step 0's
census measures and that no release gate runs in the harness (a refused
configuration is a failed receipt, not a stop). `analyze.py` lists refused
configurations and summarises the census beside the step-1 table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`uk_size_step1b_plan` (tool: `plan-step1b`) applies the experiment's selection
rules to the scored step-1a refits, so step 1b runs straight after 1a: stage
`ae` writes the A×E points (the best A rule x each anchor x the knee of that
anchor's floor-0.5 L2 ladder and its two half-decade neighbours, each lambda
multiplied by the rule's loss ratio at S0 from the census, at floor 0.5);
stage `holdout`, once those are scored, writes the refit-level holdouts of
C0, C2, the best E per anchor, the best A and the best A×E. Each pick is
recorded with the candidates it was read from, and a pick with no candidate
is skipped. Publication includes the plans.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fter 1a

The README records the readings `plan-step1b` applies (María's go for steps
0-1, 2026-10-08): best E and the knee on each anchor's floor-0.5 ladder, A×E
at floor 0.5 at the knee and its half-decade neighbours, the winner rule's
criterion-6 count when none passes it, and the lambda grid kept as
pre-registered. The knee rule's "largest" is corrected to "smallest" before
any result exists: "largest" makes the knee the ladder's top whenever
collapse falls with lambda. `run.sh` now runs steps 0, 1a and 1b in one
chain; `analyze.py` reports the step-1b picks and the holdouts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`publish` reduces every `run_dir` in the receipts and the census to the run
directory's name, so a published results file names no machine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Steps 0, 1a and 1b ran from the run tree (v6's code 224d93ee6 plus this
branch's #1124 commits) on runs/uk-local-k25-675d2a3fd, 2026-10-07 23:48Z to
2026-10-08 04:30Z: the control reproduced the stored refit bit for bit; 26
step-1a refits, 6 A×E points and 4 refit-level holdouts, none refused.
results/results.json holds the disclosure-controlled aggregates; the README's
results section is written by analyze.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chosen after step 1 (María, 2026-10-08) and labelled as not pre-registered:
grain_family_equal_sqrt_count at floor 0.5 with the uniform anchor at lambda
1e-3, 3e-3 and 1e-2 times its loss ratio at S0, the low-lambda region step
1b's knee rule skipped, plus the holdout of E_unif_f0.5_1e-2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
grain_family_equal_sqrt_count at floor 0.5 with the uniform anchor at
lambda 0.0021, 0.0064 and 0.021, and the holdout of E_unif_f0.5_1e-2,
re-scored and re-published with steps 0-1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#1124)

Step 2 runs a 5-refit grid and holdouts on each new search's selection; until
now selection mode kept only its own refit, so every further refit would have
searched again (hours each). A selection-mode experiment now saves its search
and draw under <out>/<name>/selection in the graph's own encodings, with a
manifest of the search's rule, selection L2 and file digests. A refit or
refit_holdout experiment names it with `selection_from` and runs on that
selection, passing the search's weighting and L2 so the reuse check binds; a
refit with the search's own settings reproduces the search's refit bit for
bit. The tool checks every `selection_from` up front, and a failed search
leaves no saved selection, so its dependents fail with that reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
María's go of 2026-10-08. S-A searches under grain_family_equal_sqrt_count
(warm-started at the stored penalty times the rule's loss ratio) and refits
five ways on its saved selection; S-AE adds selection-stage L2 at the
pool-design anchor and runs only when no S-A configuration passes all six
criteria.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
main now carries #1115's warm start, which adds the shared initial_lambda to
the search, draw and refit parameters; the default-parameters test names it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The search under grain_family_equal_sqrt_count (2026-10-08 21:12Z to
2026-10-09 06:11Z, 4 probes, selected lambda 1.078e-6) and its five refits on
the saved selection, re-scored and re-published with steps 0-1c. S-AE was
stopped during its search and left no result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort) — round 1 at d9ab99ef

Verdict: nothing in the code blocks this. The rules and the L2 penalty are sound and well tested, and release candidates can't reach either one. Before calling SA_f0.5_u0.021 the best configuration, I'd re-run it under a couple more seeds and run its holdout, because three of its passes sit close to their thresholds. The engine-free CI failure comes from main, not from this PR.

What checks out

  • The rules are named functions of row metadata, not per-target vectors.
    • grain_family_equal gives 1 / (G * F_g * n).
    • The nation_ variant gives sub-UK national rows a fourth grain.
    • _sqrt_count only moves count weight inside a (grain, family) cell, so every grain, family and cell budget is unchanged.
    • Every vector sums to one, and the tests check each partition.
  • An undeclared value basis is refused under the _sqrt_count rules, rather than guessed (target_weights.py, uk_target_value_basis).
  • The holdout recomputes the rule on its held rows through the same carrier, and the carrier is identical whether it's built from the declared targets or from the stored problem (test_targets_and_stored_problem_build_the_same_carrier).
  • The L2 penalty is off by default. When it's off, no L2 keyword reaches the solver, and default size-node parameters and checkpoint identities don't change (test_default_size_nodes_keep_their_parameters, test_default_calls_pass_no_l2_keyword_and_keep_the_receipt). The record basis and the design anchor are refused, with the reason given.
  • Release safety holds. --release-candidate refuses any non-default --target-weight-rule (rowwise_cli.py:594). It also refuses --dataset-households, which the L2 flags require, so neither change can reach a release. The dense preflight also fails a manifest with selection_l2, refit_l2 or a rule override.
  • The harness is an experiment harness only.
    • The acceptance code (uk_size_acceptance) matches the six pre-registered criteria exactly.
    • Configurations the solver refuses are recorded as failed receipts.
    • No release gate runs inside it.
  • Disclosure control holds. I scanned the published results.json: there are no area codes or unit records, and every small integer is a target count, fold, category or area count. 179 values are suppressed as <10.
  • Tests: 184 engine-free tests in the touched files pass locally (target weights, L2, L2 graph, wiring, size experiment, posture, preflight, diagnostics, full build CLI), and ruff is clean. The new modules don't exist on main, so the new tests can't pass there.
  • CI: both engine-free failures are test_pre_eligibility_spool_is_not_uploaded_after_upgrade. That's main's date-dependent spool fixture: its events are dated 2 Oct, and they're pruned after 7 days. Every other lane passes.

Should

  1. The headline "best" rests on one seed and no holdout.
    • Every receipt uses seed 42.
    • SA_f0.5_u0.021 clears three criteria by small margins:
      • NI share 2.568 % against a floor of 2.552 %;
      • lone-person share 0.80 points against a limit of 1;
      • households −0.34 % against a limit of 1 %.
    • The C5 learning-rate control moves the NI share by about 0.03 points, which is more than the 0.016-point margin.
    • Suggested fix: re-run that configuration under two more seeds, and run its 5-fold holdout (about 28 min on the saved selection), before the description or the README calls it the best so far. A seed-variance line beside the table would also help readers judge every near-threshold pass.
  2. Results provenance points at commits that aren't on GitHub.
    • The receipts record 98753d5dd (36 receipts), 37c0f7801 (6) and e7dd2f199 (4). These are pre-rebase SHAs that don't exist in the remote.
    • The control receipt has no code block at all (size_experiment.py:752, run_uk_size_control).
    • Suggested fix: add a code block to the control receipt. In the README, map the recorded SHAs to this PR's commits, or point to the 2026-10-09 bit-exact portability check on the rebased branch, so the results stay traceable after merge.

Nits

  1. Ranking rule. The description ranks "best" by most criteria passed, with ties broken by fewest areas below 25 % of dense ESS. That isn't one of the README's pre-registered selection rules (README.md:62-64, :75). It gives the same answer as the A×E winner rule (criteria 1–5, then the criterion-6 count). Say so, so the pick doesn't read as post hoc.
  2. Basis default. uk_size_l2_from_options defaults a missing basis to "record" (dataset_size.py:372), but the UK size stages refuse that basis. Default to chi_square, or raise if the key is missing.
  3. Name-based disclosure control. disclosure_controlled only treats a value as a unit count if its key matches a substring list (size_experiment_scorecard.py:777). A future count published under a new key name would go out unsuppressed. A test that walks the published output and flags small integers under unlisted keys would catch that.

juaristi22 and others added 2 commits October 9, 2026 22:13
María's request of 2026-10-09: the S-A search and its five refits under
nation_grain_family_equal_sqrt_count, which gives national rows half the
loss, warm-started at S-A's selected penalty scaled to this rule, and the
holdout of S-A's best configuration on its saved selection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The search under nation_grain_family_equal_sqrt_count (2026-10-10 00:26Z to
09:44Z, 4 probes, selected lambda 1.0226e-6), its five refits on the saved
selection and the 5-fold holdout of S-A's best configuration
(SA_f0.5_u0.021), re-scored and re-published with the earlier steps. None was
refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Strategy comparison so far (v6 K=25 baseline, 60,000 households)

The table shows one or two configurations per strategy, each scored against the same dense solve D. The full scorecard of all 49 sets is in the experiment README.

Selection Refit (rule · floor · L2 λ) Set Households vs D NI share Lone-person share Constituencies under ESS 50 LAs under ESS 50 Areas below 25 % of dense ESS National rows past 25 % grain_equal loss Held-out loss Criteria passed
Dense solve (reference) all 1.6M pool households D 0.00 % 2.69 % 28.63 % 0 3 — 3 0.0113 — —
Stored v6 selection grain_equal · floor 0 · no L2 (today's 60k file) S0 −3.47 % 2.01 % 23.04 % 170 20 158 15 0.0221 0.115 4
Stored v6 selection grain_equal · floor 0.5 · no L2 C2_f0.5 −0.36 % 2.50 % 27.08 % 591 154 898 4 0.0130 0.202 1, 4, 5
Stored v6 selection grain_equal · floor 0.5 · uniform L2 0.01 E_unif_f0.5_1e-2 −0.76 % 2.38 % 24.86 % 65 9 69 7 0.0163 0.145 1, 4
Stored v6 selection grain_equal · floor 0.5 · uniform L2 0.1 E_unif_f0.5_1e-1 −4.41 % 2.27 % 21.76 % 3 7 32 12 0.0260 — none
Stored v6 selection grain_equal · floor 0.5 · initial L2 0.01 E_init_f0.5_1e-2 −0.99 % 2.16 % 24.86 % 182 23 153 6 0.0164 — 1, 4
Stored v6 selection grain_equal · floor 0 · initial L2 0.1 E_init_f0_1e-1 −26.97 % 2.87 % 25.59 % 315 163 431 393 0.2446 — none
Stored v6 selection gfes · floor 0.5 · no L2 A_gfes_f0.5 −0.06 % 2.63 % 28.97 % 649 268 997 15 0.0193 0.212 1–4
Stored v6 selection ngfes · floor 0.5 · no L2 A_ngfes_f0.5 −0.20 % 2.55 % 28.97 % 649 273 999 11 0.0192 — 1, 3, 4
Stored v6 selection gfes · floor 0.5 · uniform L2 0.021 X_gfes_unif_f0.5_0.021 −0.66 % 2.50 % 25.60 % 79 13 88 15 0.0229 — 1, 4
Stored v6 selection gfes · floor 0.5 · uniform L2 0.21 AE_gfes_unif_f0.5_0.21 −4.21 % 2.52 % 21.75 % 2 6 29 46 0.0496 0.129 none
S-A (search under gfes) grain_equal · floor 0 · no L2 SA_ge_f0 −2.12 % 1.98 % 26.72 % 153 20 143 47 0.0302 — none
S-A (search under gfes) gfes · floor 0.5 · no L2 SA_f0.5 −0.23 % 2.66 % 29.28 % 610 147 859 14 0.0188 — 1–4
S-A (search under gfes) gfes · floor 0.5 · uniform L2 0.021 SA_f0.5_u0.021 −0.34 % 2.57 % 27.82 % 33 13 93 35 0.0296 0.136 1–4
S-A (search under gfes) grain_equal · floor 0.5 · uniform L2 0.01 SA_ge_f0.5_u1e-2 −0.51 % 2.44 % 27.03 % 68 18 122 7 0.0167 — 1, 4
SAN (search under ngfes) grain_equal · floor 0 · no L2 SAN_ge_f0 −12.61 % 1.95 % 25.17 % 192 41 198 95 0.0992 — none
SAN (search under ngfes) ngfes · floor 0.5 · no L2 SAN_f0.5 −0.23 % 2.59 % 29.38 % 609 167 888 9 0.0183 — 1–4
SAN (search under ngfes) ngfes · floor 0.5 · uniform L2 0.022 SAN_f0.5_u0.022 −0.40 % 2.41 % 27.83 % 31 17 83 29 0.0281 — 1, 3
SAN (search under ngfes) grain_equal · floor 0.5 · uniform L2 0.01 SAN_ge_f0.5_u1e-2 −0.30 % 2.44 % 27.23 % 82 20 107 5 0.0169 — 1, 4, 5

Reading the table

  • Criteria (pre-registered):

    1. households within 1 % of D;
    2. each nation's share within 5 % relative of D's (NI needs at least 2.55 %);
    3. lone-person share within 1 point of D's (at least 27.63 %);
    4. no local (grain, family) loses more than 1 point of within-10 % fit against S0;
    5. national rows past 25 % at most D's 3 plus 2;
    6. no constituency or local authority below 25 % of its dense ESS.

    S0 passes criterion 4 by construction.

  • Floor: the refit baseline floor (baseline_pi_floor). Floors 0.1, 0.5 and 1.0 give nearly identical refits, so 0.5 stands in for all three.

  • L2: a chi-square penalty toward uniform weights or toward the refit's starting weights (initial). Under the family-equal rules, λ is scaled by the rule's loss level (×2.14 for gfes, ×2.16 for ngfes).

  • Abbreviations: gfes = grain_family_equal_sqrt_count; ngfes = nation_grain_family_equal_sqrt_count.

  • Held-out loss: the mean grain_equal loss on held-out local rows over 5 rotated refit-level folds. Lower is better.

What it shows

  • The floor restores household mass and fit, and collapses per-area ESS: 591 constituencies fall under ESS 50.
  • Uniform L2 restores ESS and lowers the lone-person and NI shares. The initial anchor at floor 0 and λ 0.1 drops 27 % of households.
  • On the stored selection, the family-equal rules meet the composition criteria only with ESS collapsed (649 constituencies under ESS 50).
  • Re-running the search under gfes (S-A) lets the rule refit keep the composition while spreading weight. SA_f0.5_u0.021 passes criteria 1–4 with 93 areas below 25 % of dense ESS. Its held-out loss (0.136) sits between S0's (0.115) and the floor alone's (0.202).
  • No configuration passes criterion 5 together with 2 and 3. With uniform L2, the search under ngfes (SAN) misses fewer national rows than S-A at the same penalty but loses the NI criterion. The national misses that remain are mostly HMRC top-income bands and SPI regional £200k+ taxpayer counts.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants