Repository navigation
Anchor UK year-file runs on the survey year so the State Pension follows the triple lock - #557
Merged
Merged
Conversation
policyengine-uk takes the first year of a dataset as observed data: its State Pension formulas split state_pension_reported at that year against the year's legislated rates and scale the share to the simulated year's rates (the triple lock). create_datasets cut year files from the projection, and run() passed a year file back as a single-year dataset, so the projected year became the observed year and the State Pension followed the CPI uprating of the reported amount. 2026 State Pension was GBP 127.50bn through policyengine.py against GBP 133.69bn from policyengine-uk on the same certified data (6.2.1). Year files now record their data year and keep its tables. run() projects those tables forward as policyengine-uk does and puts the year file's own tables at the simulated year, matching records by ID so region scoping carries over. Year files without a recorded data year are regenerated by ensure_datasets and refused by load_datasets. Fixes #556 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
create_datasets(years=[]) only materializes and loads the source; reading the data year there is wasted work and broke the runtime tests that mock policyengine-uk's Microsimulation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under row filtering, outputs normalised over the whole dataset (business-rates incidence through shareholding, deciles, relative poverty lines) become region-relative (#567), so the scoped test now checks person and benefit-unit outputs only. policyengine-uk encodes enum columns in place on a multi-year dataset's tables; it copies a single-year dataset when projecting it, so the comment names the path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Test a State Pension reform against a direct run with the same compiled policy, and a year file loaded back from disk; both fail with run() reverted to the single-year construction. - Saved UK outputs record the data year their run anchored on; load() refuses outputs without it, so ensure() reruns pre-fix outputs. - Drop policyengine.py's person and benunit weight columns from the data year's tables before projecting, so the years in between match a direct run; order the projected years. - Write the data-year key last so it marks a complete file; let unreadable files raise; load each year file once. - State the scoping caveat (#567) in the docs and tests: deciles, the relative poverty median and shareholding are calculated over the simulated households. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Refuse a dataset that records a data year without its tables (a simulation output) with a clear error instead of an AttributeError. - Docs: row filtering applies to both sets of tables; weight replacement changes the simulated year's weights only. - Changelog: reforms had no effect under policyengine-uk 2.102.3 and were understated under 2.123. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
marked this pull request as ready for review
October 9, 2026 05:54
Contributor
Author
|
Merged at head Gates
Independent review: Opus via
Why this merges without Max's sign-off: it is a correctness fix that moves published aggregates (2026 State Pension £127.5bn → £133.7bn on 6.2.x). It makes the wrapper match policyengine-uk on the same data, which Max's standing rule covers. Follow-ups split out
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #556
Summary
A UK population run through policyengine.py did not reproduce
policyengine_uk.Microsimulationon the same certified dataset and year. The State Pension followed CPI instead of the triple lock, and everything downstream moved with it. Under the current release's policyengine-uk (2.102.3), State Pension rate reforms didn't reach it at all. With this PR, a run of a UK year file matches the direct policyengine-uk run on every output, record by record, with or without a reform.Mechanism (from the code; issue #556 has the per-record evidence)
create_datasetscut each year file from the projection policyengine-uk makes of the certified file (sim.dataset[year]).run()passed that frame back asUKSingleYearDataset(fiscal_year=year).min(simulation.dataset.years). They splitstate_pension_reportedat the data year against that year's legislated rates, then scale to the period's rates.state_pension_reportedby CPI. So for every recipient, the wrapper's State Pension was the reported amount uprated by CPI rather than by the triple lock.Fix
PolicyEngineUKDatasetgainsdata_yearanddata_year_data.create_datasetsrecords the source's first year and stores its tables in each projected year file (data_year,data_year_person, … HDF5 keys).run()builds the policyengine-uk input a direct run builds. It projects the data year's tables forward with policyengine-uk's ownextend_single_year_dataset, and puts the year file's tables (scoped, if a region is set) at the simulated year. Records are matched by entity ID, so region scoping carries over to the data year; records missing from the data year raise a clear error. The tables are copied before handing them over. policyengine-uk encodes enum columns in place on a multi-year dataset's tables, which this path builds; it copies a single-year dataset when projecting it.person_weight/benunit_weightcolumns are dropped from the data year's tables before projecting, because a direct run calculates them. The projected years are passed in order.ensure_datasetsregenerates year files that lack the data-year record.load_datasetsrefuses them, and now loads each file once. The record is written last, so it marks a complete file, and an unreadable file raises instead of reading as unrecorded. A dataset built in memory withoutdata_yearstill runs as observed data for its own year, which is how policyengine-uk treats a single-year dataset.Simulation.load()refuses UK outputs saved without one, as it refuses US outputs without their marker, andSimulation.ensure()runs them again.Evidence (real runs, 2026)
Record-level check, through the public API:
ensure_datasets→Simulation.run()vspolicyengine_uk.Microsimulationon the same materialized file, over 69 output columns (the default outputs plus the State Pension components, Housing Benefit, CTR and thegov_*totals):Before the fix, the same comparison (a prototype of this construction, against the old one) found 26 of 68 columns differing. They included
in_poverty_*for 313–372 households and the income deciles. A Scotland-scoped run's State Pension components, income tax and Pension Credit matched the direct run's for the records kept with the fix, and differed without it.With a reform: new State Pension £260/wk and basic £200/wk from 2026, against 2026 law of £241.30 / £184.90. The policyengine.py runs use
Simulation(policy=...). The direct runs apply the same compiled policy with policyengine.py's own modifier (reform_impact.py). Changes are against each column's own baseline, in £bn:The reformed totals with this PR equal the direct run's to three decimals: State Pension £144.369bn on 6.2.x and £138.564bn on #555.
How each column was run:
cdad33a, which is the branch merged with main (6.2.2), without this PR, on its own year files. Its baseline (pepy_branch.json) was run on the branch before the main merge, which reports 6.2.1. The merge only brought US changes (Use ACS-local data for US regional simulations #552).Impact on published UK figures
Every UK population result computed through policyengine.py's
Simulationonensure_datasetsyear files carries this error, verified on policyengine 6.2.1 and the #555 branch. Both columns below come from running policyengine.py's own API (pepy_poverty.py:ensure_datasets,Simulation,calculate_uk_poverty_rates/_by_age,Aggregate) before and after the fix. In both cases the fixed result equals the direct policyengine-uk run on every aggregate and on overall poverty, to three decimals.senior)Microsimulation"model versus data" table is unaffected.Invariants
RowFilterStrategy, the tested strategy) it holds, for the records kept, for person and benefit-unit variables that don't depend on dataset-wide normalisation. policyengine-uk calculates income deciles, the relative-poverty median andshareholdingover every household in the simulation (read in its source). A row-filtered run therefore gets region-relative values for those, including household flags mapped to people.WeightReplacementStrategychanges them through the weights. That is pre-existing issue Scoped UK runs give the region UK-wide business rates and region-relative ranks #567, which this PR doesn't change. Tested by example for 2024 / 2025 / 2026 / 2028, by a Hypothesis property over random pensioners (age, sex, reported amount) and years 2025–2030, with a reform, and throughload_datasets. Checked on the real certified data for 2026 under both pe-uk versions (tables above).data_yearand the data year's tables.data_yearafteryear, a projected dataset without its data-year tables, and data-year tables withoutdata_yearare all refused. Outputs record a data year without tables.load()and re-run byensure().Tests:
tests/test_uk_year_file_data_year.py, 20 tests on a seven-person 2024 dataset written in policyengine-uk's own file format, real policyengine-uk, no licensed data, about 50 s. All 20 also pass under policyengine-uk 2.123.4 (#555's pin).run()reverted to the old single-year construction, 7 of the first 15 fail: every differential test for a projected year, the scoped run, the Hypothesis property, regeneration, and record matching. The 8 that still pass don't depend on the anchoring: 2024 is its own data year, plus storage, validation and aliasing. The reform test and the loaded-from-disk test also fail under the mutant.f384ec2, 1214 passed, 9 skipped, 2 failed; origin/main gives 1201 passed and 9 skipped in the same environment. Both failures were intests/test_dataset_runtime.py, which mocks policyengine-uk and callscreate_datasets(years=[]).875924cskips the data-year read when no years are requested, and those tests then pass. On1629ee7b, 237 targeted tests pass: every test file that saves, loads or ensures simulations, plus the UK dataset, region, scoping, run-record and release-manifest tests.Compatibility
ensure_datasetsrewrites them once.load_datasetsrefuses them with a message to regenerate. DirectPolicyEngineUKDataset(filepath=...)opens are not checked, matching the US precedent for legacy-input records.Simulation.load()and re-run bySimulation.ensure().policyengine_uk.data.economic_assumptions.extend_single_year_dataset, which has the same signature in 2.102.3 and 2.123.4.managed_microsimulation, household calculations, and the US path.Related
_policyengine_uk_input(dataset). Whichever PR lands second keeps this body. Test the UK nation filters against policyengine-uk's country formula #568 also calls it on the unscoped dataset to derive filter variables; with this body that dataset includes the data-year tables, which is correct, just costlier.Independent review
Opus (
subfleet run --task review --tier standard) requested changes at9e16ad72, with no blockers. It confirmed the mechanism and the construction by reading both policyengine-uk versions, and confirmed every number but one. All ten findings are addressed in1629ee7b:Round 3 at
efaeb13aapproved the nit fixes. The branch was then brought up to date with main (#565, #559) at0110ff97, with no conflicts; its diff against main is still these six files.Round 2 at
1629ee7bapproved: every finding was fixed in code, with no regressions. Its nits are addressed inefaeb13a: changelog wording, citing the policyengine-uk fix PR (#2163, 2.123.1), how each reform column was run, docs on weight replacement, and a clear error when an output dataset is simulated. Logging expected re-runs below warning level touches the US path too, so it is split out.Follow-ups (split out)
country/*regions filter on acountrycolumn that the certified year file doesn't store, so they raise. Found while testing scoping.ensure_datasetsreuses year files cut from an earlier data release or model version (UK and US).axiom: n/a: wrapper data-loading fix; no policy rule changes
🤖 Generated with Claude Code