Repository navigation
Revive the UK leg as an internal audit of policyengine-uk - #212
Merged
Merged
Conversation
- Pin policyengine-uk 2.121.0 (core 3.32.17); a fresh 100-household US sample gives identical references under both cores (2,292 rows). - Re-pin the public transfer artifact to policyengine-uk-data 6b1f80e, which stores PIP component categories (the old pin's pip_*_reported columns were silently ignored after policyengine-uk#1656, zeroing PIP). - Prompt reference-period (uprated 2026-27) inputs instead of the stored 2025 values; income tax references matched the prompted facts in only 24 of 100 households before. - Compute UK references from the prompted facts, one PE-UK simulation per household (the 2026 UC health element is a seeded per-benefit-unit draw); keep the full-microdata calculation as a cross-check. - State derived facts the engine uses: sex and date of birth, education status at ages 16-19, LCWRA held before 6 April 2026. - Fail loudly when the artifact stores a variable the engine does not define or a prompted derived variable is renamed (is_child_or_QYP was removed). - Make the supervisor pass --country and use the country's output set. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
consensus-flags and adversary-prepare take --country; load_payload reads a country's payload (load_us_payload keeps its behaviour). UK cases get UK prompts: fiscal year 2026-27, legislation.gov.uk and GOV.UK sources, pounds, the region, and a 2026-10-10 freeze (policyengine-uk 2.125.1), with the schema's freeze date to match. All 62 committed US cases re-render byte for byte and the US schemas are unchanged. Also pin policyengine-uk 2.125.1, the latest release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
The prompt formats every UK amount with no decimals, but the situation used the unrounded reference-period values (£15,095.16 shown as £15,095), so 150 of the internal run's 700 references sat up to £0.99 from the stated facts. Round each UK numeric input as the formatter does (half to even) when building scenarios; an amount shown as £0 becomes unlisted. For the frozen internal run, a rounded copy of the manifest renders all 100 prompts byte for byte as the models saw them, and its references leave the 59 consensus flags unchanged. Property test: rounding never changes what the prompt displays for an amount that stays in the prompt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ference Address the review of #212: - Relationships: the listed adults go in as claimant and partner and every child as neither (is_claimant_or_partner), and a couple as not married, matching the prompt, which now says so. PE-UK otherwise presumed a 37/18 pair to be parent and child and presumed every couple married, giving Marriage Allowance the prompt never states. On the frozen run this moves ten income tax references by £252 and resolves five consensus cells. - Locality: a private renter gets a Broad Rental Market Area, the area of the region's largest city, stored and stated in the prompt. PE-UK reads the LHA from the BRMA alone and defaults everyone to Maidstone. - Education: someone in non-advanced education is stated (and supplied) to be in full-time education begun before 19; PE-UK's default entry age (1000) failed the qualifying-young-person entry condition. - LCWRA wording: the assessment, with protected status conditional on UC entitlement, so pension-age adults are not told they hold an award. - Boundary: to_pe_uk_situation rejects tax-unit or benefit inputs and person fields the prompt filters out. - Tests: the situation property test uses valid typed values and checks each input against the rendered prompt; new tests for the boundary, the renter BRMA, and (slow, real engine) an unmarried 37/18 couple. - docs/audit.md names each country's freeze; the benchmark card describes the new conventions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Address round 2 of the #212 review. The conventions (whole-number amounts, a private renter's Broad Rental Market Area, a non-advanced-education entry age) applied only when scenarios were extracted from the transfer data, so a scenario loaded from a manifest could still give the prompt and the reference different facts. - canonical_uk_scenario applies them, idempotently; describe_household and to_pe_uk_situation both start from it. - The education entry age is a stated fact: an explicit age is kept and shown, 16 only where the scenario gives none (it overwrote explicit ages). - to_pe_uk_situation also rejects household inputs the prompt filters out. - The BRMA mapping is described as a fixed area within each region (Maidstone is not the South East's largest city). - Tests: a loaded unrounded manifest gives prompt and reference the same rounded facts and BRMA; explicit entry ages 18 and 19; hidden household inputs; idempotence (Hypothesis). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lief method Address round 3 of the #212 review and two conventions the reference audit surfaced. - Explicit zeros are stated facts and are kept (a zero months_since_last_birthday moved the date of birth when dropped); rate fields keep their precision; months render without a currency sign. - canonical_uk_person, shared by describe_person and the reference: someone aged 16 to 19 with no education status is "not in education" (PE-UK imputed enrolment from age), and someone with self-employment income and no gainful self-employment determination has none (PE-UK presumed one and applied the Universal Credit minimum income floor). Both follow the prompt's rule that an unlisted status is false, and both are stated. - Pension contribution labels state the relief method: employee contributions under a net pay arrangement, personal contributions as relief at source. On the frozen run this moves three Universal Credit references; two land on the model consensus and the audit judge's own answer (scenario_044 £8,442.06, scenario_047 £15,873.48). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… shown Round-4 review of the UK revival. - The prompt called personal pension contributions "relief at source", but PolicyEngine UK deducts them from taxable income, so the label described a method the reference does not follow. It now states the convention: a gross amount deducted from taxable income. - A UK rate reached the engine at full precision while the prompt showed four significant figures. The canonical scenario now stores the rate the prompt shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # policybench/cli.py
Round-5 review of the UK revival. PolicyEngine UK deducts both kinds of pension contribution from taxable income but neither from adjusted net income, so they do not lower the High Income Child Benefit Charge or the personal allowance taper (policyengine-uk #2243). Both prompt labels and the benchmark card now say so. The rate round-trip property also covers negative rates and values around the thousands separator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Merging on the gates.
🤖 Generated with Claude Code |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This revives PolicyBench's UK leg as an internal audit tool for policyengine-uk. Nothing is published: the public release stays US-only, and no snapshot, payload or dashboard file changes.
Bit-rot fixed (each one observed, with a test)
9514dfb) storedpip_*_reported, which policyengine-uk stopped reading in #1656, so PIP was 0 for every household. It is re-pinned to uk-data6b1f80e(sha256c663daea…), which stores PIP categories.calculate_uk_transfer_microsimulation_valueskeeps the old path as a cross-check: 691 of 700 cells agree. The 9 that don't are explained: 8 by UC deductions, 1 by the LCWRA draw.is_child_or_QYPwas removed in policyengine-uk#1655 and the loader skipped it without warning. The loader now fails loudly on any stored variable the engine doesn't define, on a renamed derived variable, and on benefit-unit inputs the prompt can't show.--country, so every UK worker failed. It now passes each scenario's country and output set. US workload fingerprints are unchanged (tested).One canonical scenario (review rounds 2-6)
The prompt renderer and the reference builder both start from
canonical_uk_scenario/canonical_uk_person, so a scenario loaded from a manifest gives each the same facts. It applies these conventions, each stated in the prompt:Reference adversary for the UK
The reference adversary from #200 now audits the UK leg:
consensus-flagsandadversary-preparetake--country;load_payloadreads one country's payload.Invariants (tested)
Verification
Run locally on Python 3.13 at the final head:
tests/test_scenarios.py,test_eval_no_tools.py,test_held_references.py,test_reference_adversary.py: 435 passed (5 slow deselected).test_consensus.py,test_cli_helpers.pyandtest_cli_reference_outputs.py, 505 passed.Six rounds of GPT-6.1 Sol review; round 6 approved
869e9bb2. The reviews and prompts are in~/reviews/policybench-uk-revival-2026-10/reviews/.What this found
The internal UK run on 100 transfer-path households cost $11.08 across six models. A blind reference audit of 79 cells (59 consensus, 20 single-model misses) found 21 cells on policyengine-uk defects, 34 on unstated facts (now stated, above), 9 on rounding (fixed here) and 13 where the reference held. Engine fixes from this work:
The full record is
~/reviews/policybench-uk-revival-2026-10/README.md. Whether to publish a UK board is Max's call (d1254).🤖 Generated with Claude Code