Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions changelog.d/556.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
UK population runs now anchor the State Pension on the survey year their data was projected from, as policyengine-uk does, so a run of a UK year file matches policyengine-uk on the certified dataset, with or without a reform. Before, it took the projected year as the survey year: the State Pension was uprated by CPI rather than the triple lock, and State Pension rate reforms had no effect on it under policyengine-uk 2.102.3 and were understated under 2.123. Projected year files now also store the survey year's tables, which doubles their size. `ensure_datasets` regenerates year files from earlier releases and `load_datasets` refuses them; `Simulation.load()` refuses UK outputs saved by earlier releases, and `Simulation.ensure()` runs them again.
39 changes: 39 additions & 0 deletions docs/microsim.md
Original file line number Diff line number Diff line change
Expand Up @@ -231,6 +231,45 @@ not need repository-specific download logic. Authentication or authorization
failures are reported directly and do not cause a retry against another
repository type.

### UK year files and the data year

`ensure_datasets` and `create_datasets` cut one file per requested year from
the projection policyengine-uk makes of the certified dataset. policyengine-uk
takes the first year of a dataset as observed data. Its State Pension formulas
split each person's reported State Pension against that year's legislated
rates, and scale the share to the simulated year's rates, which follow the
triple lock. The projection uprates the reported amount by CPI, so handing a
projected year's tables to policyengine-uk as observed data would make the
State Pension follow CPI instead.

Each projected year file therefore keeps its data year, `dataset.data_year`
(2024 for Enhanced FRS 2024–25), and that year's tables,
`dataset.data_year_data`. `Simulation.run()` projects the data year's tables
forward as policyengine-uk does, and uses the file's own tables for the
simulated year. A run of a year file then gives the same result, record by
record, as `policyengine_uk.Microsimulation` on the certified file, with or
without a reform. Keeping the data year's tables doubles a projected file: the
Enhanced FRS 2026 file is 226 MB, against 113 MB for its own tables. A dataset
built in memory without a data year is observed data for its own year, which
is how policyengine-uk treats a single-year dataset.

Row filtering applies to both sets of tables, matched by entity ID. Weight
replacement changes the simulated year's weights only; the data year and the
years between keep national weights. A row-filtered run is a simulation of the
region's households alone, so
variables that policyengine-uk calculates over every household in the
simulation are calculated over the region: income deciles, the relative
poverty median, and `shareholding`, which spreads corporate taxes across
households (PolicyEngine/policyengine.py#567). Variables of a person or
benefit unit that do not depend on those match the national run.

Year files written by earlier releases have no recorded data year.
`ensure_datasets` writes them again and `load_datasets` refuses them. A year
file opened directly, as `PolicyEngineUKDataset(filepath=...)`, is not
checked. A saved UK simulation output records the data year its run anchored
on; `Simulation.load()` refuses an output saved without one, and
`Simulation.ensure()` runs it again.

## Simulations

A `Simulation` needs a dataset, a tax-benefit model version, and optionally a policy (reform):
Expand Down
15 changes: 15 additions & 0 deletions src/policyengine/tax_benefit_models/common/model_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -443,6 +443,21 @@ def load(self, simulation: Simulation) -> None:
"stored inputs (it has no record of them); run it again"
)

if self.country_code == "uk":
from policyengine.tax_benefit_models.uk.datasets import (
_year_file_records_data_year,
)

# UK outputs saved before runs anchored on the observed data year
# (PolicyEngine/policyengine.py#556) took a projected year as
# observed and uprated the State Pension by CPI, so they are not
# reused. ``Simulation.ensure()`` runs such a simulation again.
if not _year_file_records_data_year(Path(filepath)):
raise ValueError(
"Saved UK simulation predates anchoring on the observed "
"data year (it records none); run it again"
)

simulation.output_dataset = self._dataset_class(
id=simulation.id,
name=simulation.dataset.name,
Expand Down
Loading
Loading