Skip to content

Latest commit

 

History

History
424 lines (337 loc) · 17.2 KB

File metadata and controls

424 lines (337 loc) · 17.2 KB
title Microsimulation

For population-level estimates — budget cost, winners and losers, poverty impact — run a microsimulation over calibrated microdata.

For US decile analysis, calculate_decile_impacts(dataset=..., spm=...) passes the selection to both constructed simulations. When supplying existing baseline and reform simulations, set spm on each simulation directly.

Quick example

import policyengine as pe
from policyengine.core import Simulation
from policyengine.outputs import Aggregate, AggregateType

datasets = pe.us.ensure_datasets(years=[2026])
dataset = next(iter(datasets.values()))

baseline = Simulation(dataset=dataset, tax_benefit_model_version=pe.us.model)
baseline.ensure()

total_snap = Aggregate(
    simulation=baseline,
    variable="snap",
    aggregate_type=AggregateType.SUM,
)
total_snap.run()
total_snap.result

Simulation.ensure() loads a cached result if one exists, or runs and caches on miss. Call Simulation.run() explicitly if you want to bypass the cache.

Canonical US SPM settings and provenance

The coordinated canonical SPM integration resolves measurement settings from the bundle. Its default uses each household's observed county_fips and native SPM membership. To make an explicit national sensitivity calculation:

simulation = pe.Simulation(
    dataset=dataset,
    tax_benefit_model_version=pe.us.model,
    spm={"geography_kind": "national", "scenario": "zero_real"},
)
simulation.run()
settings = simulation.spm_config
receipt = simulation.spm_provenance()
serialized = simulation.model_dump_json(include={"spm", "spm_receipt"})

spm_config exposes all six resolved settings; spm_provenance() returns a detached JSON-compatible measurement receipt after calculation. The stored spm_receipt is a Pydantic model. Before calculation, serializing a partial selection preserves omitted options so they continue to inherit the bundle's defaults after JSON restoration; successful calculation freezes all six settings. Selection participates in simulation identity and cached results. The example serializes the measurement settings and receipt; the existing full simulation model graph contains circular model/variable references and cannot currently be exported with an unrestricted model_dump_json(). Use simulation persistence and run records for saved runs. The same spm= selection is accepted by pe.us.managed_microsimulation; its returned country simulation exposes spm_config and spm_provenance() too. See the household SPM contract for exact keys, the 2022–2035 artifact horizon and explicit geography errors. No provider or forecast-file path can replace the bundle's pinned artifact.

For a small local development dataset, explicitly opt out of managed data selection:

simulation = pe.us.managed_microsimulation(
    dataset=str(local_h5_path),
    allow_unmanaged=True,
    spm={"geography_kind": "county"},
)

This does not certify the local file. The canonical population release must preserve the prior native arrays, memberships and weights while adding the source-backed is_spm_independent_minor_role input. Formula-owned measurement counts, thresholds, geographic factors, SPM resources and poverty outputs must not be present as input columns; the wrapper validates the input DataFrames before passing them to the country model. Generic adult/child counts remain separate. The native source enrichment is not recalibration, and inherited schema-5 calibration diagnostics cannot be relabeled as schema 6.

The wrapper's native pandas HDF loader accepts files that store only calibrated household_weight. It maps missing person and group weight columns through native membership IDs in memory, retaining supplied weight columns and the original row order. Inputs are aligned to country populations by native entity ID, and calculated outputs are aligned back to those IDs before attaching weights. Independently shuffled entity tables therefore retain the correct geography and results. Null or duplicate entity IDs are rejected before weight mapping. Group weights are not sums of person weights. The file is opened read-only; missing or ambiguous links raise an error instead of creating unweighted rows. Core variable/period H5 files use the same mapping.

The packaged 5.3.0 production manifest retains policyengine-us==1.764.6 and does not yet certify this integration. Local-wheel tests use an explicitly uncertified development manifest. Registry publication, a producer-issued data compatibility certification and final bundle promotion remain separate gates; measurement provenance alone does not satisfy them.

Datasets

Microdata is stored as HDF5 on Hugging Face. ensure_datasets downloads, caches, and uprates:

datasets = pe.us.ensure_datasets(
    years=[2024, 2026],
    data_folder="./data",  # local cache directory
)
dataset = datasets["populace_us_2024_2026"]

The default US dataset is Populace US 2024 — a Populace-built dataset calibrated to IRS, CMS, SNAP, Census, and other administrative totals. The current UK certified default is Enhanced FRS 2024–25, supplied by policyengine-uk-data. Populace UK 2023 remains available as a named, non-default bundle dataset.

PolicyEngine.py obtains the repository type, immutable revision, and SHA-256 from the installed release bundle. An existing file in the configured data directory is reused only after hash verification.

List datasets already known to the country:

pe.us.load_datasets()  # or pe.uk.load_datasets()

US local-area dataset

Alongside the certified national default, the bundle registers a US dataset for finer geographic work: populace_us_2024_acs_local. It is a Populace US 2024 build of roughly 1.6 million households on an ACS 2024 multispine, with each household PUMA-assigned to a 119th-Congress congressional district, county, and state, and calibrated to state administrative totals and state and congressional-district population. Its release validation summary records four reviewed limitations, so read that summary before relying on it. It ships in its own immutable release. State and congressional-district region simulations select it through the bundle's region_datasets metadata; direct microsimulations can still load it by name.

Two-line load:

import policyengine as pe

sim = pe.us.managed_microsimulation(dataset="populace_us_2024_acs_local")

Or materialize it as a PolicyEngineUSDataset for Simulation:

datasets = pe.us.ensure_datasets(datasets=["populace_us_2024_acs_local"], years=[2024])
dataset = datasets["populace_us_2024_acs_local_2024"]

Because this file carries PUMA-assigned district, county, and state identifiers calibrated to state and congressional-district population, state and congressional-district breakdowns should filter this dataset rather than the national default. Filter it with the same state_fips / congressional_district_geoid row filters used elsewhere (see Regional analysis):

from policyengine.core import Simulation
from policyengine.core.scoping_strategy import RowFilterStrategy

ca = Simulation(
    dataset=dataset,
    tax_benefit_model_version=pe.us.model,
    scoping_strategy=RowFilterStrategy(variable_name="state_fips", variable_value=6),
)

UK private data and raw h5 access

UK population data uses licensed Family Resources Survey inputs. The default UK release bundle points to the private policyengine/policyengine-uk-data-private Hugging Face repository. Set HUGGING_FACE_TOKEN to a token from a Hugging Face account with access:

export HUGGING_FACE_TOKEN=hf_...

For policyengine.py analyses, use the logical dataset name from the release bundle. ensure_datasets resolves it to the pinned private Hugging Face file, downloads it, caches it locally, and creates year-specific uprated datasets:

import policyengine as pe
from policyengine.core import Simulation

datasets = pe.uk.ensure_datasets(
    datasets=["enhanced_frs_2024_25"],
    years=[2026],
    data_folder="./data",
)
dataset = datasets["enhanced_frs_2024_25_2026"]

simulation = Simulation(
    dataset=dataset,
    tax_benefit_model_version=pe.uk.model,
)
simulation.run()

To materialize the raw certified artifact without creating uprated yearly datasets, use PolicyEngine.py's bundle API:

from policyengine.provenance import materialize_dataset

result = materialize_dataset(
    "uk",
    "enhanced_frs_2024_25",
)

print(result.path)
print(result.bundle_dataset.sha256)

The bundle API uses the repository type recorded in the bundle, so callers do not need repository-specific download logic. Authentication or authorization failures are reported directly and do not cause a retry against another repository type.

UK year files and the data year

ensure_datasets and create_datasets cut one file per requested year from the projection policyengine-uk makes of the certified dataset. policyengine-uk takes the first year of a dataset as observed data. Its State Pension formulas split each person's reported State Pension against that year's legislated rates, and scale the share to the simulated year's rates, which follow the triple lock. The projection uprates the reported amount by CPI, so handing a projected year's tables to policyengine-uk as observed data would make the State Pension follow CPI instead.

Each projected year file therefore keeps its data year, dataset.data_year (2024 for Enhanced FRS 2024–25), and that year's tables, dataset.data_year_data. Simulation.run() projects the data year's tables forward as policyengine-uk does, and uses the file's own tables for the simulated year. A run of a year file then gives the same result, record by record, as policyengine_uk.Microsimulation on the certified file, with or without a reform. Keeping the data year's tables doubles a projected file: the Enhanced FRS 2026 file is 226 MB, against 113 MB for its own tables. A dataset built in memory without a data year is observed data for its own year, which is how policyengine-uk treats a single-year dataset.

Row filtering applies to both sets of tables, matched by entity ID. Weight replacement changes the simulated year's weights only; the data year and the years between keep national weights. A row-filtered run is a simulation of the region's households alone, so variables that policyengine-uk calculates over every household in the simulation are calculated over the region: income deciles, the relative poverty median, and shareholding, which spreads corporate taxes across households (#567). Variables of a person or benefit unit that do not depend on those match the national run.

Year files written by earlier releases have no recorded data year. ensure_datasets writes them again and load_datasets refuses them. A year file opened directly, as PolicyEngineUKDataset(filepath=...), is not checked. A saved UK simulation output records the data year its run anchored on; Simulation.load() refuses an output saved without one, and Simulation.ensure() runs it again.

Simulations

A Simulation needs a dataset, a tax-benefit model version, and optionally a policy (reform):

baseline = Simulation(
    dataset=dataset,
    tax_benefit_model_version=pe.us.model,
)

reformed = Simulation(
    dataset=dataset,
    tax_benefit_model_version=pe.us.model,
    policy={"gov.irs.credits.ctc.amount.base[0].amount": 3_000},
)

policy= accepts the same flat {"param.path": value} dict shape as pe.us.calculate_household(reform=...), or a Policy object with explicit ParameterValue entries. Scale parameters use bracket indexing — see Reforms.

Outputs

Every output has the same lifecycle: instantiate with the simulation(s) and configuration, call .run(), read the typed result fields.

from policyengine.outputs import (
    Aggregate,
    AggregateType,
    ChangeAggregate,
    ChangeAggregateType,
)

snap_cost = Aggregate(
    simulation=baseline,
    variable="snap",
    aggregate_type=AggregateType.SUM,
)
snap_cost.run()

budget = ChangeAggregate(
    baseline_simulation=baseline,
    reform_simulation=reformed,
    variable="household_net_income",
    aggregate_type=ChangeAggregateType.SUM,
)
budget.run()

See Outputs for the full catalog.

Memory and performance

A full Populace US microsimulation uses roughly 4 GB of memory and takes 15-30 seconds on a laptop. For parameter sweeps, reuse the baseline:

baseline = Simulation(dataset=dataset, tax_benefit_model_version=pe.us.model)
for amount in [0, 1_000, 2_000, 3_000]:
    reformed = Simulation(
        dataset=dataset,
        tax_benefit_model_version=pe.us.model,
        policy={"gov.irs.credits.ctc.amount.base[0].amount": amount},
    )
    # each iteration runs only the reform

Smaller custom H5 datasets can be passed explicitly for testing:

datasets = pe.us.ensure_datasets(
    datasets=["/path/to/smoke_test_populace_us_2024.h5"],
    years=[2026],
    allow_unmanaged=True,
)

These run in seconds and are fine for integration tests. Don't use them for production analysis — the weights are not calibration-tuned.

Managed microsimulation

managed_microsimulation constructs a country-package Microsimulation pinned to the policyengine.py release bundle (so the dataset selection is certified, not ad-hoc):

from policyengine.tax_benefit_models.us import managed_microsimulation

sim = managed_microsimulation()
# `sim` is a policyengine_us.Microsimulation — use its API directly

Pass allow_unmanaged=True with a custom dataset= to opt out of the release bundle. Explicit local paths and Hugging Face URIs remain supported in this mode. GCS dataset URIs are not supported.

For managed simulations, sim.policyengine_bundle records the actual source package, repository type, revision, verified SHA-256, and local path.

Renamed inputs in stored data

A country engine sets a stored column as an input only when it defines a variable of that name. When policyengine-us renames an input, data written before the rename keeps the old name, and the engine would skip it. policyengine.tax_benefit_models.us.legacy_inputs.LEGACY_INPUT_RENAMES lists each such rename; today it holds one entry, would_claim_wic → takes_up_wic_if_eligible (the WIC take-up draw, PolicyEngine/microcosm#1026).

Every US load path applies it: Simulation.run(), managed_microsimulation and create_datasets (and so ensure_datasets when it creates year files). A rename applies when the data stores the old name, the engine does not define the old name but does define the new one, and the data does not already store the new name. The new input is then set from the stored values for every month of every dataset year. The stored table must list the simulation's entity IDs in the simulation's order, or loading fails rather than attach values to the wrong people. Only values stored for a whole year are mapped, so a file that stores the old name for part of a year, and not the new name, is refused. The mapping turns itself off once the data stores the new name, or if the engine defines the old name again.

The renames applied are recorded as {old: new} ({} when none applied):

  • Simulation.run(): simulation.output_dataset.metadata["legacy_input_renames"], also shown as simulation.release_bundle["legacy_input_renames"]. It includes renames applied when the input year file was cut, so a run over a create_datasets year file records the rename even though the file already stores the new name. A US save() writes the record into the output file, load() restores it, and a run record's results.json binds it;
  • managed_microsimulation: sim.policyengine_bundle["legacy_input_renames"];
  • create_datasets: each year file, and each returned dataset's metadata["legacy_input_renames"].

{} means that no stored column was mapped. It does not show that the data carried the draw: data that stores neither name runs with the new input's default, which for WIC is full take-up.

Files written before this mapping existed have no record, and are not reused:

  • A saved US output may have been calculated without the mapped inputs. load() refuses it, and ensure() runs the simulation again and saves the new output.
  • A year file that ensure_datasets or create_datasets wrote stores neither name, so the draw is lost from it. ensure_datasets creates such year files again, and load_datasets refuses them. A year file opened directly, as PolicyEngineUSDataset(filepath=...), is not checked.

Pinned model versions

Every policyengine release pins specific country-model and country-data versions so results are reproducible. pe.us.model and pe.uk.model expose the pinned TaxBenefitModelVersion.

If the installed country-package version doesn't match the pinned manifest, managed_microsimulation warns. For strict reproducibility, pin country packages to the versions the policyengine release was built against — see Release bundles.

Next