Repository navigation
Conversation
…22 reference audit (#174) * Add the Claude Opus 5.5 model card and metadata The API lists claude-opus-5-5 with created_at 2026-09-21. It rejects forced tool use with a 400 and thinks by default (a request with no thinking parameter returns a thinking block), so the card runs the JSON contract whole-scenario on the sync path, as Fable 5.1 does. Priced at $4/$20 per 1M with $0.20 cache reads from the pricing page; litellm has no entry yet, so eval_no_tools registers it locally with the Fable line. Gauntlet: 3/3 and 16/16 parsed, $0.057 per scenario. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Prepare the 40-model board copy, note, and freezer paths for Opus 5.5 The leaderboard banner, benchmark card, and paper now name Claude Opus 5.5 beside Fable 5.1 as a row that rejects forced tool calls and answers as JSON. A note records the debut with facts a test recomputes from the frozen payload; the Astra note keeps its own release once this one supersedes it. The freezer reads the adds202609 publish bundle and the 40-model payload, and the copy sites state 40 models. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Judge the cases Opus 5.5 joined and stage the 40-model annotations The fold disturbed 120 of the 668 judged cases. Claude Opus 5.5 judged them (13 through scripts/run_audit_claude.sh under launchd, 107 as Claude Code Workflow subagents returning the audit schema as structured output), each with the provenance sidecar the CLI runner writes. Case notes are regenerated, not copied, so wrong_model_count is right for every case the new row joined. The freezer records the workflow runner and lets one judge bucket name more than one runner. Twenty-five verdicts flag the reference as suspect; they are recorded as the judge returned them and await adjudication before publication. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Apply the developer adjudications to the 40-model annotations The finish stage regenerates row annotations and case notes from the verdict tree, which does not carry the eleven developer adjudications that remove the ambiguous SSI and Medicare outputs from scoring. This applies scripts/apply_adjudications.py to the regenerated files, as the September publish did, so every substantive row of those outputs carries prompt_ambiguity again and the freezer's verification passes. Two of the eleven outputs (scenario_057 snap, scenario_067 ssi) were re-judged today by Claude Opus 5.5, which returned reference_model_issue_fixed with a reference-suspect flag. Their entries now record that verdict verbatim, as the schema requires, in place of GPT-5.6 Sol's 2026-09-01 llm_error verdict; the exclusion decision and its reasoning are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Add the GPT-6 Sol and GPT-6 Luna model cards and metadata OpenAI announced both on 2026-09-22; the Models API lists gpt-6-sol and gpt-6-luna with created 2026-09-14. The model pages price Sol at $2 / $10 per 1M input/output ($0.20 cached) and Luna at $0.10 / $0.50 ($0.01 cached), with the GPT-5.6 line's 272k-token surcharge, so both join the Responses-API group and its cost path. Both default to medium reasoning. Gauntlet on the forced tool contract, like GPT-6 Astra: Sol 3/3 and 16/16 parsed, $0.016 per scenario; Luna 3/3 and 16/16, $0.002 per scenario. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record engine-defect and later-law exclusions and resolve every reference flag A reference-exclusion entry can now name an engine defect (the engine that produced the reference misapplies the law on stated facts: root cause, defect, law, corrected value, upstream issue) or later-published law (the reference depends on a figure published after the references were frozen), besides an unlisted input. Rows on those outputs take the developer-only classes reference_engine_defect and reference_later_law. An adjudication can carry a reference_verdict (affirmed, engine_defect, unlisted_input, later_law) with the law it rests on; applying it clears the judge's reference_suspect flag on the case and its rows, and verification fails while a flag survives. The freezer refuses a snapshot that still has a flagged case without a verdict, and the manifest counts verdicts and which entries the judge flagged, so the paper can report the audit wave after the flags are cleared. The site's exclusion note states the defect, the law, the corrected value and the upstream issue for engine defects, and the published figure for later-published law. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Report the September 22 audit wave in the paper and settle regenerated references The paper no longer claims the frozen snapshot has no known reference defects. It reports the wave Claude Opus 5.5 judged: the flags it raised, how adjudication settled each (affirmed, unlisted input, regenerated, engine defect), the sandbox-fix sweep that found every output each defect moves, their exclusion for every model, and the three root causes fixed upstream after the freeze. The judge roster is three models, the exclusion paragraph covers both exclusion routes, and the SNAP October-December references state the published-before-the-freeze convention. Counts come from the frozen record through paper_results. An adjudication can record a regenerated reference: the flagged value was replaced under a convention the reference sidecar records, and the output stays scored with a final class. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refreeze the 42-model board on the audited September 22 references Adds GPT-6 Sol and GPT-6 Luna to the board beside Claude Opus 5.5 and refreezes the snapshot on references audited under the publication rule: a scored reference follows from stated facts and law published before the 2026-07-03 freeze. - 23 references regenerated under eight publication conventions (14 SNAP outputs held at FY2026; the rest use amounts published before the freeze, or the last published amount, where the engine had projected a parameter with a price index). A ninth, Wisconsin's published 2026 amounts, moves only outputs already excluded for engine defects. The sidecar records each convention with its fix module's sha256. - 55 outputs excluded from scoring for every model: engine defects and references that depend on inputs or definitions the prompt never states. - 61 developer adjudications cover every reference flag the judges raised. - reference_audit/2026-09-22 commits the root causes, the sandbox fixes, the per-fix sweep of all 1,984 references, the build scripts and the independent verification reports; tests/test_reference_audit.py ties every regeneration and exclusion to them. - The Claude thinking-sensitivity summaries are rescored on the new board (scripts/rescore_sensitivity_summaries.py), and the paper, notes, methodology and serving evidence are updated to match. Board (household-weighted exact match, 1,929 scored outputs): GPT-6 Sol 94.5, Claude Opus 5.5 93.2, GPT-5.6 Sol 92.9, GPT-6 Luna 91.6. The dashboard payload release dashboard-data-20260922 is not uploaded yet, so app CI cannot fetch it until it is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Exclude the engine's SNAP rounding defects instead of regenerating them An independent review of the September 22 refreeze found that the SNAP publication convention's module also corrected the engine's SNAP arithmetic. Under the publication rule a formula defect is excluded, not regenerated, so the module is split: - c_snap_hold_fy2026 now holds the FY2026 schedule only (r13_hold_fy2026_v3): the SNAP uprating index and poverty guideline. - Three engine-defect root causes carry the arithmetic: r26 (the allotment keeps cents where 7 CFR 273.10(e)(2)(ii)(A) requires whole dollars), r27 (California's CalFresh net income rounds to the nearest dollar; the engine floors it) and r28 (the minimum benefit is 8% of the one-person maximum rounded to the nearest dollar, $24; the engine returns $23.84). Each is measured against the convention's reference, the value that would otherwise be published. - Recomputing SNAP with net income rounded to the nearest dollar, and with its cents kept, in every state moves no scored reference. The record is now 66 exclusions (44 engine-defect outputs across 14 root causes, 22 unlisted-input outputs) and 12 regenerated references (3 SNAP), with 72 adjudications. The Opus 5.5 judge re-ran the 14 SNAP cases whose reference changed and flagged two of the rounding defects itself. Every model is scored on 1,918 of 1,984 outputs. Board: GPT-6 Sol 94.8, Claude Opus 5.5 93.6, GPT-5.6 Sol 93.3, GPT-6 Luna 92.1. Review fixes from the same pass: - The CalEITC AGI comparison was fixed upstream in PolicyEngine/policyengine-us#9363 (merged 2026-09-01), not #9542. - The paper and audit note count the engine-defect outputs no judge flagged instead of saying "most"; the Fable 5.1 uplift is computed from the sensitivity summary. - Cost wording now matches eval_no_tools: the token-count reconstruction when one exists, the provider's charge otherwise. - The manifest's response window comes from MODEL_RESPONSE_DATE. - tests/test_reference_audit.py checks each sweep row's class, baseline and arithmetic, that every new exclusion has a qualifying move of its class, and that no qualifying move is left scored. The dashboard payload is dashboard-data-20260922 (sha256 5b738d4b..., 116,403,352 bytes); it is not uploaded yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Regenerate references for defects fixed upstream; exclude heat-and-eat SUA Rule (Max, 2026-09-23): an engine defect fixed in policyengine-us after the 2026-07-03 freeze is regenerated with its fix on the same engine version, not excluded; a defect not fixed upstream stays excluded. Seven root causes qualify: r04 capital gain distributions (#8839), r09 New York's renter cap (#9301), r17 the CalEITC's AGI lookup (#9363), r26 SNAP contribution rounding and r27 SNAP net-income rounding (#9318), and r28 SNAP minimum-allotment rounding and r31 SNAP income-limit rounding (#9162). r27 now rounds net income to the nearest dollar in every state, as #9318 does. A second adversarial review (five lenses, two skeptics per major finding) confirmed 27 findings, among them: - scenario_080's SNAP reference rests on the heat-and-eat standard utility allowance, which P.L. 119-21 sec. 10103 limited to households with an elderly or disabled member; the engine still grants it. New engine defect r30, not fixed upstream: 080 SNAP is excluded. - The SNAP convention no longer holds the poverty guideline: the engine already uses the 2026 HHS guideline, published before the freeze. - Adjudications from an earlier build could outlive the exclusions they recorded; this wave's adjudications are now rebuilt each run, and every flag any judge run of the wave raised stays recorded as raised. - Test hardening: sweep rows check their frozen and baseline values, every convention and upstream fix must have a sidecar revision, and a qualifying move must be excluded or regenerated by its own fix. The sidecar now records one revision per convention and per upstream fix. The Opus 5.5 judge re-ran the 19 cases whose reference changed; it flagged none. Eighteen derivation narratives were rewritten from the engine traces and checked number by number by two skeptics each. Record: 51 exclusions (27 engine-defect outputs across 10 root causes not fixed upstream, 24 unlisted-input), 27 regenerated references (13 SNAP; 16 by upstream fixes), 62 adjudications, 38 judge flags. Every model is scored on 1,933 of 1,984 outputs. Board: GPT-6 Sol 94.2, Claude Opus 5.5 92.9, GPT-5.6 Sol 92.6, GPT-6 Luna 91.3. Thirteen read-only Axiom encoding-prep lanes (reference_audit v9) independently derived the affected households' values from statute text; each supports the corrected values it could derive, some on stated assumptions (Wisconsin's election, New Jersey's coverage facts). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Split the Wisconsin capital gain defect from #8839; fix round-3 review findings Round 3 of the adversarial review (20 confirmed findings) is addressed: - r04 now regenerates with the faithful backport of policyengine-us #8839. The verified r04_v2 module also extended Wisconsin's 30% capital gain subtraction (Wis. Stat. 71.05(6)(b)9) to capital gain distributions, which #8839 does not and upstream main still omits. That part is a new unfixed root cause, r32_wi_capital_gain_distributions, measured on top of the Wisconsin convention plus #8839 (sweep_moves carries the combination's rows as class "baseline"). It moves scenario_091 state tax (895.81 -> 878.52), now excluded, and scenario_042 state tax, already excluded under r06. - An excluded output's corrected value now applies its unfixed fixes on top of every publication convention and upstream fix, as the published references do: scenario_080 SNAP is 3,240.00, a whole-dollar total. - Adjudications take the judge's class from each case's verdict.json, not from case notes that already hold the adjudicated class. The freezer refuses a record whose judge fields differ from verdict.json, with a test. by_judge_verdict is now llm_error 45, reference_engine_defect 11, reference_model_issue_fixed 5, prompt_ambiguity 2. - scenario_080's narrative is regenerated from the frozen engine trace (the rendered trace drops zero nodes, so it never showed the utility allowance); scenario_091 keeps its frozen narrative. The 080 exclusion and the paper state that the prompt's "is disabled" establishes no 7 U.S.C. 2012(j) disabled member, and that the output is excluded on either reading. - Paper: the review table no longer calls upstream-fixed outputs excluded; the publication-rule sentence counts the 23 convention regenerations; the Axiom sentence is narrowed to what v8 tested; the abstract, audit note, benchmark card and README say "recorded as engine-defect exclusions". - README lists the modules that set 2025 values; the SNAP convention docstring names the upstream SNAP fixes; make_issues drafts r32. Record: 52 exclusions (28 engine-defect across 11 unfixed root causes in 20 households, 9 never flagged; 24 unlisted input); 26 regenerated references in 24 households (13 SNAP; 15 by upstream fixes, 23 by conventions); 63 adjudications; every model scored on 1,932 of 1,984 outputs. 080 and 091 re-judged by Claude Opus 5.5 (both llm_error, no flag). Board: GPT-6 Sol 94.2, Claude Opus 5.5 92.9, GPT-5.6 Sol 92.7, GPT-6 Luna 91.4. Checks: pytest 845 passed, 5 skipped; bun test 141 passed; eslint; next build; ruff check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add a disability section to the paper; fix round-4 review findings Disability (Max: "should be in the paper"): a new data subsection, "Disability in the household facts", explains that the prompt states one general `is disabled` fact (the six CPS difficulty items; 33 of the 177 benchmark people) while SSI, SNAP, Medicare and the tax code each apply their own determination, each a separate engine input that no benchmark person carries, so a disabled person's references take the non-disabled path unless another listed fact establishes it. Mechanisms were checked in the 1.755.4 code and the counts come from the frozen scenarios (paper_results, tested). The audit paragraph, limitation bullet and benchmark card point to it. Round 4 (22 confirmed findings): - The 2026-09-05 exclusions keep their classification but their alternative values are recomputed with every convention and upstream fix, through u_* situation patches that reproduce the recorded frozen-engine values exactly (023 SNAP 3,576.00, 057 SNAP 840.00, 100 SNAP 8,844.00); 100's note names r30, which also moves it (8,556.00 to 6,924.00). - A regenerated adjudication keeps the judge's subtype when an upstream fix regenerated it; convention-only ones stay thresholds_rates. - A wave flag that the current verdict no longer raises is recorded with judge_reference_suspect_source; the freezer now checks the flag against verdict.json too. - scenario_091's narrative is regenerated from the frozen trace (the engine leaves the $1,170 of distributions out of gross income); its re-judge now flags the reference for that reason, so the wave has 39 flags (20 engine defect) and 8 engine-defect exclusions were never flagged. - r09 cites #9313 as well as #9301; the r30 note and the paper say the published $3,576 (not the frozen value) would stand under the other reading; Wisconsin no longer appears among the regenerating conventions; the Axiom sentence names the untested New Jersey minimum; r32 moves two outputs; the Massachusetts deposit reading is described accurately; the README and card wording follow. Record unchanged in size: 52 exclusions, 26 regenerated references, 63 adjudications, 1,932 scored outputs; board unchanged. Checks: pytest 846 passed, 5 skipped; bun test 141 passed; eslint; next build; ruff check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix round-5 review findings Round 5 (12 confirmed findings, all minor or nit): - The carried 2026-09-05 SNAP adjudications (023, 057, 100) now state the recomputed alternative that their exclusion records carry, and 100's names r30; build_records replaces rather than re-appends the sentence each run. - scenario_091's narrative now states that the engine leaves the listed $1,170 of capital gain distributions out of gross income; the frozen narrative step checks each required figure and fails if one is missing. The case was re-judged (still flagged for that omission; counts unchanged). - Disability section: SNAP's route is receipt of SSI or Social Security disability benefits (policyengine-us approximates the SSI route with SSI disability status); Medicare's disability route needs 24 months of entitlement; the Medicare exclusions are credited to the unstated months of SSDI receipt (all five heads have SSDI income, three are not flagged disabled), not to the "is disabled" fact. - 080: the value that would stand under the other reading is the one the conventions and upstream fixes give ($3,576), not a published value. - The card and README cite #9313 for r09; the freezer's flag check has tests. Record and board unchanged: 52 exclusions, 26 regenerated references, 63 adjudications, 39 flags, 1,932 scored outputs. Checks: pytest 847 passed, 5 skipped; bun test 141 passed; eslint; next build; ruff check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
#175) * Add a note on the five SNAP BBCE households for the September 22 board The note covers the five scored SNAP households whose reference is the $24 minimum benefit and that qualify only through broad-based categorical eligibility: how the 42 models answer them, what the engine's BBCE rule checks, what the prompt says about take-up and unlisted benefits, and what changed since the September 3 note (the $287.68 reference, scenario_112, the board size, and two corrections to that note's description of BBCE). - scripts/snap_pathways_20260922.py recomputes SNAP pathways for all 100 households on policyengine-us 1.755.4 with the committed SNAP convention and upstream SNAP fixes. It reproduces every scored SNAP reference and, on the unmodified engine, every frozen SNAP value within $1. - scripts/bbce_household_rows.py writes every model's answer for the five households from the frozen payload of dashboard-data-20260922. - tests/test_notes.py recomputes all 39 facts from committed files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address review findings on the five SNAP BBCE households note Note text: - Count the 3 of 210 requested answers that came back with no value and no explanation, and use the 207 explanations as the mention denominator. - Say that all five have gross income above 130% of the poverty guideline and that the three passing the gross test have a member the engine treats as elderly or disabled, which exempts them from it. - Quote the prompt's instruction to treat unlisted numbers as 0 and unlisted statuses as false next to the take-up and no-inference instructions, and tie the scenario_112 exclusion to it. - Turn the $276/$23 amounts, the 40 hours and the 100 households into facts. Evidence and tests: - scripts/snap_pathways_20260922.py also records each household's elderly or disabled flag, its lowest monthly gross-income-to-guideline ratio, and the SNAP gross limit and unearned income sources (financial_assistance among them). Regenerated on policyengine-us 1.755.4; existing columns unchanged. - A slow test reruns that script and compares it with the committed CSV and meta when policyengine-us 1.755.4 is installed; it skips elsewhere. - The rows meta now names run_payload_sha256 (the committed data.json.gz) and release_payload_sha256 (the release asset the manifest pins), and the test checks both. - The note test checks both prompt contract variants, that every quotation in the prompt paragraph is the preface's wording, and which models the prose names for each claim. - The app test that could not fail now renders each note and checks that only the September 3 note has the unrounded-reference footnote. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Exclude the SNAP child support defect (r33) in release 20260922b policyengine-us 1.755.4 reads gov.usda.snap.income.deductions.child_support as the 7 CFR 273.9(c)(17) option to exclude child support paid from SNAP gross income, but its values carry the opposite meaning. For 2026 it excludes the payments in 37 jurisdictions that USDA's SNAP State Options Report lists as deducting them from net income (273.9(d)(5)), Michigan among them (16th edition p. 15, 17th edition p. 21; Michigan BEM 556, line 20). The fix, PolicyEngine/policyengine-us#9586, is open and not merged, so under the standing rule the output it moves is excluded for every model rather than regenerated. - reference_audit: root cause r33_snap_child_support_treatment (engine defect, decided 2026-09-24), the sandbox module embedding #9586's corrected parameter at head 3f15666, and the measurement modules on the SNAP convention and on the published SNAP configuration. Recomputing all 1,984 references under each moves one output, scenario_045 SNAP (287.68 frozen, 288 published, 0 corrected); sweep_moves.csv gains its row. The README records the revision. - Records: 53 exclusions (29 engine defect across 12 root causes) and 64 adjudications. The excluded output keeps its frozen reference and narrative, so its judge case was re-judged (Claude Opus 5.5 through scripts/run_audit_claude.sh, 2026-09-24); the adjudication keeps that verdict. 25 references stay regenerated; every model is scored on 1,931 outputs. - Snapshot: the payload refrozen as dashboard-data-20260922b (sha256 b926e1c1..., 116,198,253 bytes); pointer, freezer pin and manifest updated. The board snapshot label stays 2026-09-22: the models and their answers are unchanged. - Notes keep their own release: the three notes on dashboard-data-20260922 are now on a superseded release, so tests keep their verified facts. - Paper, benchmark card, docs, sensitivity summaries and pinned tests updated; the paper is re-rendered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rewrite the SNAP BBCE note for four households on release 20260922b Release dashboard-data-20260922b excludes the Michigan worker's SNAP output (scenario_045, root cause r33): policyengine-us subtracts the child support the worker pays from SNAP gross income, while Michigan counts it there and deducts it from net income, so the household is over Michigan's 200% BBCE limit and gets $0. The fix, policyengine-us#9586, is open, so the output is excluded for every model rather than regenerated. The note keeps its slug, date, title and board snapshot and moves to release dashboard-data-20260922b. It carries the voice-reviewed text of draft PR #176 (head b501fb3), which already applied that review's findings: the hyphen-tolerant BBCE pattern, "SNAP's benefit formula", "and with it SNAP" in the income-limit sentence, the gross-test exemption ordering, the receiving and asset-test claims in the September 3 correction, and "the sixth household, also in Texas". Edits for four households: - The Texas resident is the one household that fails the ordinary gross income test; GPT-6 Astra's one $0 is that household. - A dated correction: an earlier version, published September 23, counted the Michigan worker; PolicyEngine subtracts the child support, Michigan counts it, the household does not qualify, PolicyBench no longer scores it, and PolicyEngine has not yet merged #9586. - The September 3 correction names the Michigan worker among its six. Facts, recomputed on the 20260922b payload: 168 answers, 134 at $0, 5 exact (GPT-6 Astra 3, Claude Opus 5.5 1, GPT-5.5 1), 19% above $0 against 74% for the four savings households, 21 of 166 explanations mention BBCE and 9 of those end at $0 (6 citing a net income limit), GPT-6 Sol $0 on 3 of 4. Four fact keys lose their "five" prefix (fiveAbove0Share and the three fiveBbce* keys become income*). Data: snap_pathways_20260922.csv is regenerated on policyengine-us 1.755.4 against the 20260922b references (only scenario_045 changes: reference 287.68, not scored); bbce_households_20260922.csv drops the Michigan worker's 42 rows; both BBCE metas record the new release. The data links drop the Michigan worker, point the release link at 20260922b, and add policyengine-us#9586. Tests: tests/test_notes.py pins every sentence, including the correction against the r33 exclusion record, the worker's inputs and the pathway row (under the engine it would be a BBCE household; with the child support counted its gross ratio exceeds Michigan's limit). 18 targeted mutations each fail it. The slow pathway regeneration test passes on 1.755.4. pytest 850 passed, 6 skipped; app bun test 143 pass; eslint, ruff check and format (0.15.2) clean; bun run build succeeds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rewrite the excluded Michigan SNAP derivation, re-judge it, revise the BBCE note Review of PR #177 found two client-facing problems with the excluded scenario_045 SNAP output. The restored frozen-bundle derivation called the annual $287.68 a monthly benefit and an average, credited the October minimum to deduction changes, and left out the child support the engine subtracts from gross income. The re-judged audit notes stated that engine exclusion as Michigan's rule, which is the defect this release excludes. - Derivation: regen_references.py rewrites it from the frozen engine's trace (FROZEN_NARRATIVES, FROZEN_REQUIRED, and HAND_CORRECTED, which replaces a writer draft that credited the October change to the poverty guideline). It attributes the $433.33 monthly exclusion to PolicyEngine's child support parameter and sums twelve monthly minimums, $23.84 for nine months and $24.3744 (8% of the projected $304.68 maximum) for three, to $287.68. tests/test_reference_audit.py ties the published text to those tables. - Audit notes: Claude Opus 5.5 re-judged the case through scripts/run_audit_claude.sh with four engine facts in the unified audit's grounding: the parameter reading, and that the trace cites no Michigan rule. It returned llm_error / taxable_income_or_deductions with no reference flag, and each of the 42 diagnoses compares the answer with the reference. The adjudication keeps the new judge subtype. - App: the prediction dialog adds a line above an excluded output's audit note: the note compares the answer with the frozen reference, which carries the defect, unlisted input or projection the exclusion note describes. - Snapshot: refrozen as dashboard-data-20260922b, payload sha256 f098f11db1a209f8..., 116,205,571 bytes; pointer, freezer pin, manifest, BBCE row metas and paper hashes follow. Scores, exclusions and every count are unchanged; only annotation text moved. - BBCE note, from the voice review: the title becomes "Most models answer $0 for households that qualify for SNAP under a state's looser income rules"; the BBCE paragraph names SNAP's ordinary tests; the prompt paragraph follows the households; PolicyEngine's reading of the Michigan resident's listed employer-sponsored premiums sits beside the net income claim (recomputed with the resident paying them: net income passes and the amount stays $288); the Pennsylvania savings household's excluded SNAP output is named; GPT-5.6 Sol, Claude Fable 5.1 and 10 other models share the savings and income split; the corrections say the SNAP amount, not the household, is unscored and date the revision September 24; the September 3 correction drops the three-month gloss. Numbers under ten render as words through a {key:words} placeholder. - September 3 note: a last paragraph and a link point to this correction. - SNAP pathway meta: each excluded SNAP row's exclusion (scenario_045: r33, #9586 open) and the employer-premium reading, both regenerated by scripts/snap_pathways_20260922.py; the CSV is unchanged. Checks: pytest 851 passed, 6 skipped; the slow pathway regeneration test passes on policyengine-us 1.755.4; app bun test 147 pass; eslint clean; bun run build succeeds; ruff 0.15.2 check and format clean; paper re-rendered with an absolute PYTHONPATH and re-pinned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Revise the BBCE note's prose and point the September 22 notes to 20260922b BBCE note (every sentence still pinned in tests/test_notes.py, every number recomputed there): - Title: "because their states open it to higher incomes" replaces "under a state's looser income rules". - P1 states the $24 minimum where it states the rule; P2 introduces "ordinarily" for the tests the rest of the note calls ordinary; P3 states the elderly-or-disabled exemption before who fails which test, and says the prompt lists the Michigan resident's premiums and PolicyEngine treats the employer as paying them. - P4 says what the prompt's rules tell a model, in one sentence that shows the two rules pulling apart, instead of what models do. - P5 names the heat-and-eat utility allowance, says why the Pennsylvania household's answers count toward the share, and both shares now divide by the answers models give (125 of 166 above $0 is 75%, was 74% of 168; 32 of 166 is 19%), as the explanation counts do. - P8 opens with its point: GPT-6 Sol and Claude Opus 5.5 count less income than PolicyEngine does. - P9 gives the first version's figures (172 of 210 answers at $0) and says "counted Michigan's way". - P10: the zero-hours household gets $0 because its one adult fails SNAP's work requirements (policyengine-us 1.755.4 checks the general and the able-bodied-adult requirements together each month and has no countable-month limit; checked by an engine run of the frozen scenario), not because of the time limit alone; the rounding sentence drops the process clause; the September 3 correction now says PolicyEngine applied BBCE on eligibility. September 3 note: the pointer describes the two households it no longer scores instead of naming scenario tags, and its counts are facts. September 22 notes: each closes with what release dashboard-data-20260922b changes (one more exclusion; 29 engine-defect outputs in 21 households across 12 root causes, 9 never flagged; 25 regenerated, 12 SNAP, 14 by upstream fix; 1,931 scored; GPT-6 Sol 94.3, Claude Opus 5.5 93.0, Claude Opus 5 auto re-run 89.4; every other figure unchanged). A new test recomputes each of those against the frozen snapshot and checks every other figure in the GPT-6 Sol note still holds. The notes intro says a test checked every number against the snapshot of the note's release, which also holds for superseded releases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the State Options Report editions in the r33 record and refreeze 20260922b The r33 alternative reading said "USDA's State Options Reports list Michigan" as a deduction state. The 16th (p. 15) and 17th (p. 21) editions do; the 15th lists Michigan as an exclusion state, which Michigan's BEM 556 contradicts. The record now reads "USDA's 16th and 17th State Options Reports list Michigan among those states (the 15th lists it among the exclusion states)", in root_causes.json, reference_exclusions.json and the payload, and the audit README says the same. Refrozen with the same pipeline (fold, export, freeze_snapshot.py, paper render, --rendered-only). The new release payload differs from the previous one (f098f11d..., 116,205,571 bytes) in that one string and nothing else, checked field by field: dashboard-data-20260922b sha256 b1c4ee340a01328f75158bb4ba40b49030a979d2b03badb2c01d57c16554d3b4 116,205,632 bytes data.json.gz sha256 3d7c4ef5632c... The pointer, manifest, freeze_snapshot.py pin, BBCE row metas and the paper's frozen-export hash prefix follow. No score, count or note fact moves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tighten the September 24 prose and correct the 15th-edition wording BBCE note: the title reads "under their states' higher income limits", so "under" attaches only to "qualify". The Michigan premiums caveat moves to the end of its paragraph and the Pennsylvania caveat after the per-model split, dropping the clause about the share. The take-up rule "can tell" a model the household takes the non-cash benefit up. The first version "reported 172 of 210 answers at $0 across five households". SNAP rounds the minimum to the nearest dollar, as policyengine-us has since July 28. At 0 hours PolicyEngine treats the Texas adult as failing SNAP's work requirements, the engine's result rather than a legal finding. The September 3 note's two misdescriptions get their own paragraph. September 3 note: its closing paragraph names release dashboard-data-20260922b for the figures it states and links it. Reference audit and GPT-6 Sol notes: the closing paragraphs split the child support sentence, say the September 22 release regenerated the Michigan output with an upstream fix, and replace "On it" and "this note". Records: 940934d wrote that BEM 556 contradicts the 15th edition's exclusion listing. A 2025 manual cannot contradict an FY 2023 listing; the README and the test comment now say what policyengine-us#9586 found, that no Michigan policy adopted an exclusion for FY 2023. The r33 fix module's first docstring line names the 17th edition, whose values it applies for 2026; no sweep record pins its hash. The payload, exclusion record and freezer pin are unchanged (sha256 b1c4ee34...). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rerun CI now that release dashboard-data-20260922b is uploaded Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ne-us#9586 has merged: release 20260922c (#178) * Regenerate the SNAP child support reference (r33) now that its fix has merged: release 20260922c policyengine-us#9586, which fixes the inverted SNAP child support parameter, was squash-merged on 2026-09-24 (d9e801df4). Under the audit's rule 2 an engine defect fixed upstream is regenerated with its fix, so scenario_045 SNAP (a one-person Michigan household that pays child support) returns to scoring for every model at the regenerated reference of $0, instead of the exclusion release 20260922b applied while the fix was open. - Audit record: r33 marked upstream_fixed; the r33 module embeds the merged child_support.yaml; the sweep with the merged values moves only scenario_045 SNAP; the exclusion and its 2026-09-24 adjudication are removed; regen_references.py grounds the regenerated $0 in named engine variables; the SNAP net-income sensitivity adds r33 to its three procedures; README records both revisions. - Payload dashboard-data-20260922c (sha256 01e7e72b..., 116,100,222 bytes): 52 exclusions, 26 regenerated references, 1,932 scored outputs per model. Top-12 headline scores are unchanged at one decimal (GPT-6 Sol 94.3, Claude Opus 5.5 93.0); six lower-ranked models move by 0.1. - Notes: the September 22 notes close with the 20260922c figures; the September 3 note points to 20260922c; the BBCE note keeps its 20260922b facts and says the fix merged and 20260922c scores the worker at $0. - Paper, benchmark card, sensitivity evidence and every pinned test recomputed on 20260922c; paper re-rendered and re-pinned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address the release 20260922c review findings - Sensitivity doc: Fable 5.1's auto run on all 1,984 outputs is 87.9 on 20260922c (87.852), not 87.8; note that deltas are unrounded differences. - Paper: the upstream-fix topic list names the SNAP child support option; the r33 paragraph reads as a second, now-fixed SNAP defect, in past tense. - Benchmark card: the regenerated SNAP outputs include the child support fix. - Notes: the Sol note scopes "the other scores ... stay the same" to the note; the BBCE note says PolicyEngine "subtracted" the child support. - New test: the September 3 and BBCE notes' release-c sentences and figures are checked while 20260922c is frozen (their old pins only run on 20260922b). - regen_references.py: narratives() no longer rebinds the outer loop's revision when an output has a REGENERATED_NARRATIVES entry (latent; this run rewrote one narrative). - Audit README: release-c counts include conventions and households; the Louisiana household the merged fix reaches (SNAP $0 either way); the sidecar's revision dates are the audit wave's; c13v3_plus_upstream_snap holds the fixes merged before #9586. - snap_pathways_20260922.py docstring dates its files to release 20260922b. - Paper re-rendered and re-pinned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…the /paper page to the manifest (#179) Follow-ups from the review of #178. No data, payload or score changes. - Paper: the September 2026 wave's SNAP row states each fix's months (FY2026 hold October-December; rounding fixes every month of 2026; one Michigan reference regenerated with the child support fix). The r33 paragraph opens "One more SNAP defect was found after the September 22 release". - /paper page: scripts/freeze_snapshot.py writes app/src/paperSnapshot.json (snapshot date, response window, ?v= keys from the served index.html and PDF hashes) and page.tsx reads its label, description and manuscript URLs from it, replacing hand-kept constants that still named the 2026-09-05 snapshot. A test ties the file to the manifest and rejects hand-kept keys and common date forms anywhere in page.tsx. - Sensitivity doc names both headline tables in its rounding note; a stale output count leaves a docstring; the benchmark card paragraph is re-wrapped; test_notes pins the BBCE correction paragraph in full. - Paper re-rendered and re-pinned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Compute PolicyEngine provenance once per supervised run Each supervised scenario runs in its own worker subprocess, and every worker built its resume sidecar's policyengine_bundles itself. For the US bundle that means `from policyengine.provenance.manifest import ...`, which runs policyengine/__init__.py. That builds the US and UK tax-benefit models, and the UK model's data certification makes a Hugging Face request that returns 401 without HUGGING_FACE_TOKEN and raises. The fallback import (policyengine.core.release_manifest) no longer exists in policyengine 4.16.1, so a process's first attempt finds no manifest and the US bundle is recorded through the unbundled branch. Measured with a mocked LLM on 2026-09-28: about 1 GB peak RSS and 11-14 CPU-seconds per worker before its first request. Now the supervisor starts a fresh interpreter (its worker python and worker environment) that computes the bundles once. It writes them to <run_dir>/policyengine_provenance.json with an import-free fingerprint of their inputs: package versions, direct_url and METADATA hashes, the bundled release manifests, this module's source, and the policyengine import flags. Workers get the path through POLICYBENCH_POLICYENGINE_PROVENANCE and use the file only when the fingerprint matches their own. Otherwise they compute as before, and run_state.json counts them. The writer is a fresh process, as each worker was, so the recorded value equals what a worker computing it itself would record. Sidecars are byte-identical, which the end-to-end test checks against a fresh direct computation. The supervisor checks every sidecar against the bundles it read back and never imports policyengine itself. Per-scenario subprocess isolation is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Guard the provenance handoff per the independent review - Hand workers a provenance file only for single-country runs. Each worker computed its one scenario's country in a fresh process, and which branch records the US bundle can depend on what the same process looked up first, so a multi-country writer could record a different value. - Time out the writer (an hour) and fall back as on any other failure. - Remove an earlier run's provenance file before writing, so a failed write never leaves a stale file behind. - Pin sys.prefix in the fingerprint so everything else installed is covered, and say so in the docstring. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say where the provenance handoff applies, per the delta review - The supervisor docstring, the writer comment and the runbook now say the file is written for single-country runs, and what a run without it records (policyengine_provenance: null; workers and the supervisor compute provenance as before). - The guard comment says "succeeds", which is what depends on lookup order. - The fingerprint docstring says sys.prefix identifies the environment but does not detect in-place upgrades of other packages; a test covers the prefix. - A failed removal of an earlier file no longer aborts the run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…rences on policyengine-us 2.15.17 (release 20260929) (#182) * Add the Claude Sonnet 5.5 model card and metadata The Models API lists claude-sonnet-5-5 with created_at 2026-09-28, the release date the model page gives. It rejects forced tool use with a 400, and a request with no thinking parameter returns a thinking block, so the card runs the JSON contract whole-scenario on the sync path, as Opus 5.5 and Fable 5.1 do, and the row reasons at the API default (adaptive thinking, effort high). Priced at the standard $2/$10 per 1M with $0.20 cache reads and $2.50 cache writes from the pricing page, where no footnote marks the row as introductory. litellm's bundled map lacks the id, so eval_no_tools registers it locally with the Fable line and Opus 5.5. Gauntlet: 3/3 and 16/16 parsed, $0.027 per scenario. The roster test in test_paper_results still pins the display names to the 42-model frozen board; it passes again once the board takes the new row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Register Grok 4.7 and DeepSeek V4.1 Flash Grok 4.7 (xai/grok-4.7, released 2026-09-21 per x.ai/news/grok-4-7) and DeepSeek V4.1 Flash (deepseek-flash on the native API, released 2026-09-10) join the registry beside Claude Sonnet 5.5, with prices from the providers' pages. DeepSeek V4.1 Flash onboarded on the JSON contract (the forced tool call is rejected in thinking mode); Grok 4.7's card is provisional until its onboarding finishes with a longer timeout. The frozen-roster paper test fails until the board is refrozen with the three new rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Set the Grok 4.7 model card from its onboarding Forced tool contract, 3/3 and 16/16 whole-scenario; the timeout is 1800s because the whole-scenario probe took 472s and an earlier attempt timed out at 600s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the stage-2 driver for the September 28 additions finish_adds0928.py folds the Claude Sonnet 5.5, Grok 4.7 and DeepSeek V4.1 Flash runs onto the 20260922c board (42 models, 1,932 scored outputs), keeps every incumbent's modelStats identical, prepares the cases the additions join for the Opus 5.5 judge, gates adjudications, and exports only into a stage directory. It refuses partial runs unless --early --partial is given on copied, synthetic run states, and marks such output PARTIAL. The snapshot and freeze helpers and the design note come from the same preparation; 57 tests cover them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin policyengine.py 6.1.2 and policyengine-us 2.15.17 References now come from the newest policyengine-us, run through policyengine.py (Max, 2026-09-28). policyengine.py 6.1.2 certifies policyengine-us 2.2.1, which still has the SNAP rounding defects fixed since, so policyengine-us is pinned separately; the runtime provenance records model_matches_policyengine_bundle: false, as the July run did (policyengine.py 4.16.1 certified 1.723.0; 1.755.4 was installed). policyengine.py 6.1.2 requires Python >= 3.11; CI runs 3.12. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Let the stage-2 driver release a reviewed reference revision The driver refused any incumbent modelStats drift, which is right for an additions-only release. Regenerating references on the newest policyengine-us changes some scored references, so every incumbent's statistics move. export() now accepts that only when: - the snapshot's references differ from the pinned September 22c files (read from git at 3220a7a and checked against their sha256 pins) by exactly one committed engine_upgrade revision whose `changed` list names every changed output with its value, and nothing else; and - an export of the same bundle on the 22c references reproduces every incumbent's live modelStats exactly, so the references are the only source of drift. It writes incumbent-drift.json with each incumbent's before/after exact rate and score. With the base references committed, behaviour is unchanged. Invariant tested by property (Hypothesis): for any set of changed outputs, the revision is accepted iff it lists exactly those outputs with those values; omissions, extra entries, wrong values, a non-upgrade revision, or a sidecar that does not pin the CSV are refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Regenerate the US references on policyengine-us 2.15.17 Max, 2026-09-28: references always come from the newest policyengine-us. 2.15.17 is the newest release on PyPI (2026-09-29 00:23 UTC). policyengine.py 6.1.2 is recorded for provenance only (its certified bundle is 2.2.1). The nine pre-freeze publication conventions are re-expressed for 2.15.17. The SNAP hold needed a rewrite: reforms now apply after uprating, and FY2027 is carried as published values. The eight upstream fixes the September 22 audit regenerated are all in 2.15.17. An output- scope adapter keeps Maryland county tax (#8888) out of the state income tax output. Result, against release 20260922c: - 4 scored references change: - 008 NJ refundable credits, NJ CTC P.L.2026 c.26 (#8971); - 013 AZ SNAP, AZ expanded categorical eligibility at 200% from March 2026; - 028 PA reduced-price school meals, child support counted, 7 CFR 245.6(a)(5)(ii); - 082 NY refundable credits, Empire State child credit rounding (#9425). - 3 outputs newly excluded as unlisted-input readings: 033, 078 and 117 federal. Whether a listed state and local tax refund is income turns on prior-year facts the prompt does not give; 2.15.17 counts it (#9422, fixing #9122). - 2 move by under $1 (078 and 117 state). - 19 excluded outputs that move were re-reviewed and stay excluded. Scored outputs: 1,932 -> 1,929. Exclusions: 52 -> 55. The builder now also passes stated usual weekly hours to weekly_hours_worked_before_lsr, whose default #9261 changed from 40 to 0; scenario_066's SNAP stays 3,576. reference_audit/2026-09-28 records the modules, the 1,984-output sweep and 16 root-cause clusters, each with an investigation and an independent adversarial review. The one cross-cluster conflict (078 federal) is reconciled explicitly. tests/test_reference_upgrade.py checks every module hash, that the sweep table is the committed reference, and that every change is listed and review-backed. Ruff's target moves to py311 with requires-python. Snapshot, payload, paper and sensitivity artifacts still pin 20260922c; the freeze for this release updates them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rewrite the derivation narratives of the changed references Ten reference narratives change: - the nine outputs the engine upgrade changed (four scored changes, three new exclusions, two moves under $1); - scenario_064 dependent2_chip_eligible, whose frozen narrative said the 18-year-old exceeds CHIP's age limit. CHIP covers children under 19. Dependent 2 fails on income (321.5% of the poverty guideline against Wisconsin's 306%) and on employer-sponsored insurance. That narrative is what the September 28 judge flag on the case repeated. The writer is the usual one (claude-haiku-4-5 via case_reference_explanations._prompt), grounded in each change's reviewed basis and the 2.15.17 trace. Four drafts misstated the trace and are replaced by hand-written text that follows it: - 013 AZ SNAP: $24 a month from March, not a monthly $240 from January. - 008: how the NJ EITC derives from the federal credit. - 082: how the NY CDCC derives from the federal credit. - 028: when child support entered the school-meal income. reference_audit/2026-09-28/scripts/narratives_latest.py records them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Derive the stage-2 exclusion count from the reference revision 52 exclusions on the September 22c references, plus each exclusion the committed engine_upgrade revision lists among its changed outputs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read policyengine.py 6.x's bundle manifest; record the upgrade's adjudications policyengine.py 6.x ships one bundle manifest (policyengine/data/bundle/ manifest.json) in place of per-country release manifests, so the unbundled runtime metadata lost its data package and version. That broke the export's reference provenance check. The runtime now reads data_releases.<country> from the bundle manifest when no per-country manifest exists. The reference sidecar is rebuilt with it: data microcosm-data 0.1.0, build populace-us-2024-spm-20260915, bundle us-6.1.2. The reference CSV is byte-identical (sha256 e8bbba8f...), and the build script now reads the 22c base from git, not from the checkout. Adjudications for the upgrade: - 54 re-judged entries carry the September 28 judge's class verbatim, with the replaced values under judge_previous. Adjudicated classes, exclusions and reasoning are unchanged. - The three new unlisted-input exclusions (033, 078, 117 federal: SALT refund taxability) are adjudicated prompt_ambiguity with reference_verdict unlisted_input. - Two cases whose judge labelled model rows reference_later_law for applying superseded law are resolved as llm_error: OBBBA, enacted 2025-07-04, and NJ P.L.2026 c.26, approved 2026-06-30, both predate the 2026-07-03 freeze. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Freeze release dashboard-data-20260929: 45 models on policyengine-us 2.15.17 Local build only; nothing is uploaded or published. Board and references: - 45 models: Claude Sonnet 5.5, Grok 4.7 and DeepSeek V4.1 Flash join the 42 of 20260922c. - 1,929 scored outputs, 55 exclusions. - Payload sha256 e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6. - Audit: 674 judged cases, 68 adjudications. The export's reference-revision gate passed. Re-exported on the 22c references, every incumbent reproduced its live modelStats exactly, so all drift comes from the reviewed engine upgrade. Every incumbent's exact rate rises 0.36 to 0.54 points; the top eight keep their order. Staged ranks: Claude Sonnet 5.5 4th (91.97%), Grok 4.7 12th (88.23%), DeepSeek V4.1 Flash 15th (87.27%). The release is tagged 20260929. The references were regenerated on 2026-09-29 UTC, and the manifest dates them by regenerated_at_utc when the sidecar has one. freeze_adds0928 first brings the manifest's reference pins up to the committed, reviewed revision (checked by the driver's gate), because freeze_snapshot checks staged references against those pins. Prose, paper render, sensitivity docs, notes and test pins follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rescore the reasoning-sensitivity summaries on release 20260929 scripts/sensitivity_by_variable.py and rescore_sensitivity_summaries.py --release dashboard-data-20260929: 1,929 scored outputs, 55 exclusions, 45 board rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Brief the stage-3 prose, notes, paper and pin updates Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add Max's voice rules to the stage-3 brief Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin release 20260929's counts in the tests, sensitivity doc and app Every new pin is recomputed from the frozen artifacts of dashboard-data-20260929 (45 models, 1,929 scored outputs, 55 exclusions, 68 adjudications, 674 judged cases). - paper_results: regenerated_reference_* keep counting the September 22 convention and upstream-fix regenerations (26); new engine_upgrade_* properties count the September 28 move to policyengine-us 2.15.17 from the sidecar's engine_upgrade revision (4 scored changes, 2 within the tolerance, 3 new exclusions, 19 rechecked exclusions). A 0/1 output counts as moved on any change. - The sensitivity doc's ranks, scores and per-program tables move to the 45-model board; servingSensitivity.ts carries the rescored auto runs. - The app's copy of the serving configuration is refreshed from the frozen file, as prepare-data does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * State the 45-model board and the policyengine-us 2.15.17 references in the card and app copy - Benchmark card: a reference-outputs section states the engine (policyengine-us 2.15.17 through policyengine_us.Simulation, with policyengine.py 6.1.2 recorded for provenance), the pre-freeze publication rule, the Maryland output-scope adapter, and what the upgrade changed. Audit scope, exclusion and adjudication counts move to this release; Claude Sonnet 5.5 and DeepSeek V4.1 Flash join the rows whose provider rejects a forced tool. - Paper guide and artifacts doc: this release's counts and payload size. - Methodology names the engine from the board's own payload and counts ten chunked rows of 45; the leaderboard's serving note names Claude Sonnet 5.5 among the rows that reject forced calls. - data.versions.json describes the live version by its engine. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Update the paper for the 45-model board on policyengine-us 2.15.17, render and re-pin it - The abstract and the snapshot table name the reference engine, policyengine-us 2.15.17; policyengine.py 6.1.2 is listed as recorded for provenance, since the references come from policyengine_us.Simulation. - A reference-credibility paragraph states the move from 1.755.4: the re-expressed publication conventions, the Maryland output-scope adapter, the weekly-hours default, the four scored changes, the three new unlisted-input exclusions (SALT refund taxability, 26 U.S.C. 111(a)), the two moves under $1 and the 19 rechecked exclusions, all read from the sidecar's engine_upgrade revision. The discrepancy table gains a row for it. - The September 22 audit prose keeps its history and names 1.755.4 where it described the engine of the time; the limitations say the engine defects persist in 2.15.17 and that pre-freeze law reaches a reference only once the engine encodes it. Claude Sonnet 5.5 joins Opus 5.5 as a JSON row without a sensitivity run; the judge paragraph covers the September 28 additions (five audit waves). - paper_results names the dataset build the households were sampled from (populace-us-2024-5da5a95-20260611, which the population weights also use), not the reference runtime's later default build, and adds the references' rebuild date. - Rendered with paper/render_paper.py and re-pinned with scripts/freeze_snapshot.py --rendered-only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the September 29 release note and update the BBCE note for release 20260929 - A new note, "Claude Sonnet 5.5 debuts fourth as the references move to the newest PolicyEngine", covers the three additions (ranks, costs, serving treatments, and the deepseek-flash alias every DeepSeek V4.1 Flash answer reports), the move to policyengine-us 2.15.17, the four scored changes, the three new exclusions and the incumbents' drift. tests/test_notes.py recomputes every fact from the frozen snapshot and pins every sentence beside its evidence. The drift baseline rebuilds release 20260922c's scores from this snapshot (the upgrade's changes reverted, its exclusions removed); a second test checks the rebuild against 20260922c's committed payload in git history. - The BBCE note keeps its release-20260922b figures and closes with a dated update: the new models' answers for the four households held back by income and the four held back by savings, and the Arizona household the new references add, at $240 with all 45 models answering $0. scripts/bbce_households_20260929.py writes every model's answers to notes/data/*_20260929.csv, which the test regenerates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Put the actor first in the paper's engine paragraph and the card's engine sentence; re-render and re-pin PolicyBench rebuilt the references with policyengine-us 2.15.17 in place of 1.755.4; policyengine-us #9261 changed the weekly-hours default; the paragraph drops "now" and the convention count. The card names policyengine_us.Simulation as what PolicyBench runs. Rendered with paper/render_paper.py and re-pinned with freeze_snapshot.py --rendered-only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Voice pass on the release-20260929 prose; re-render and re-pin the paper Wording only; every number and fact is unchanged. - Release note: the title names policyengine-us 2.15.17 instead of "the newest PolicyEngine". PolicyBench is now the actor for the sandbox fixes, the freeze and the ported conventions. The #9425 fix and the engine's school-meal change are named. The SALT-refund sentence is split, and the scored-output sentence is active. The list of swaps no longer hangs off "every other model keeps its place". The closing pointer names PolicyBench's September 29 update. - BBCE note update: the Sonnet 5.5 list is punctuated so "citing BBCE" stays with the Connecticut answer. The Arizona sentence is split at "encodes the change". - Paper: the abstract, the discrepancy-table row, the engine-upgrade paragraph and the limitations use active voice. The row label names the version. Numbered sentence starts are gone. The garden path in "read from 40 hours" is fixed. - Card: the four upgrade changes are the list's subject. The judge's "later law" cases are described plainly, and the passive re-review sentences are recast. tests/test_notes.py pins the new sentences. It also checks the sidecar basis for the #9425 and school-meal wording. Rendered with paper/render_paper.py and re-pinned with freeze_snapshot.py --rendered-only. pytest: 955 passed, 6 skipped. bun test: 149 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the stage-3 report and brief the review fixes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Anchor the reference engine to its sweep date and record the 2.17.0 check policyengine-us 2.15.17 stays the recorded engine, stated as a dated fact: the newest release when PolicyBench began sweeping the references on 2026-09-29 (uploaded 00:23 UTC). 2.16.0 (04:10 UTC) predates the 11:57 UTC rebuild, so "newest when the references were rebuilt" was false. - reference_audit/2026-09-28/verification: latest_final_2170.csv and .log, sweep_latest.py --fix fixes/latest_final.py run in the .venv-pe2170 venv (policyengine 6.1.2, policyengine-us 2.17.0, the newest release at publication, uploaded 12:21 UTC). The README gains a Verify step and the dated wording; its title uses the upgrade date, September 29. - tests/test_reference_upgrade.py: the 2.17.0 sweep reproduces every one of the 1,929 scored references exactly, its 19 moved outputs are the rechecked exclusions at their recorded 2.15.17 values, and it agrees with 2.15.17 (sweep_moves.csv final) on all 1,984 outputs. - The card says each scored reference comes from 2.15.17 and that the 55 excluded outputs keep their decided values (52 on 1.755.4, 3 on 2.15.17; the 19 that move on 2.15.17 were re-reviewed and stay excluded). data.versions.json's description says the same. - docs/paper.md: reference_output_refresh records the reference runtime's default dataset (populace-us-2024-spm-20260915), which reference computation does not read; the households came from populace-us-2024-5da5a95-20260611 (new household_dataset block). - "September 28 additions" becomes "September 29 additions"; script and test comments date the upgrade 2026-09-29. The sensitivity note drops a stale "newest JSON-transport row", and the BBCE update says the scored references move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tie the release note's drift, engine, cost and Arizona facts to data - notes/data/release_20260922c_exact.csv (+ .meta.json): the 42 no-tools exact scores release dashboard-data-20260922c published, copied from its asset (dashboard-data.json, sha256 01e7e72b..., the pointer at 3220a7a6) by scripts/release_20260922c_scores.py, which refuses any other file. - tests/test_notes.py computes the drift range, the three neighbor swaps, "GPT-6 Sol still leads" and "every one of the 42 earlier models" from that fixture and the frozen payload, with no git access. The vacuous roster check becomes "the board minus the 22c roster is the three additions". The rebuild test compares the reverted-upgrade rebuild with the fixture instead of a git blob CI never fetched. - Engine paragraph: "the newest policyengine-us release" becomes the dated fact (2.15.17, newest when PolicyBench began sweeping on 2026-09-29, uploaded 00:23 UTC) plus the 2.17.0 check at publication; the upload times are checked against the upgrade record. - DeepSeek V4.1 Flash's cost is stated at DeepSeek's standard list price, the peak rate config.py prices the row at. - Arizona: the fifth household held back by income in the BBCE note's sense (with Connecticut, Texas, Michigan and Wisconsin), not the fifth to qualify only through BBCE; the update covers the four other income-held and the four savings-held households. - The title's ordinal equals facts['sonnetRank'] and its engine the sidecar's; the 2026-07-03 freeze date, the rechecked outputs and the cost basis are asserted against committed records. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the paper's engine, exclusion and cost wording; check the upgrade partition paper/index.qmd: - The abstract and engine paragraph date 2.15.17 as the newest release when PolicyBench began sweeping the references, name 2.17.0 (the publication check) and say each scored reference comes from 2.15.17; the 55 excluded outputs keep their decided values (52 on 1.755.4, 3 on 2.15.17) and the 19 that move were re-reviewed. - "The investigation found, and PolicyBench excluded, ..." replaces "Reviewers excluded"; "five audit waves" becomes "each later audit". - The cost sentence rendered a literal \\$0.002 because Quarto printed the inline string's repr; it now goes through Markdown. The cost table rendered \$0.034 and a DataFrame index column; it now uses md_table. policybench/paper_results.py: - partition_engine_upgrade_changes splits the upgrade's changed outputs into scored changes, within-tolerance changes and new exclusions, and raises unless the three add up to the changed list. The count properties read it. - engine_upgrade_date is the rebuild day (2026-09-29); the revision's own date stays the builder's wave date. - New accessors: excluded outputs by engine version, and the publication check's policyengine-us release (read from the verification sweep). tests/test_paper_results.py pins the partition (4 + 2 + 3 = 9), and a Hypothesis property over synthetic revisions and exclusion sets checks that the partition is exact and disjoint and respects the $1 and 0/1 boundaries. The judge-provenance test explains why the older judges' counts fall. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Compute the app's serving counts and exclusion engines from config and payload - app/src/lib/servingConfig.ts reads the bundled serving configuration: the chunked rows (request shape other than whole scenario) and the Claude rows on the JSON contract (each card that sets it says the API rejects a forced tool call). - Methodology states "<chunked> of the <board> models" from that config and the payload's modelStats, says policyengine-us computes each scored reference, and adds the excluded outputs' engines from the payload's referenceExclusions (55: 52 on 1.755.4, 3 on 2.15.17), with the 19 re-reviewed outputs from app/src/lib/referenceEngine.ts. - ModelLeaderboard builds the list of Claude rows that reject forced calls from the config. - app/tests render both components against data-summary.json and the config instead of hand-typed numbers. - tests/test_disclosures.py: the Methodology source computes its counts; referenceEngine.ts's recheck count and engine equal the sidecar's; the live version description matches the exclusion record; every "newest" in public copy is anchored to a time; the manifest checklist names household_dataset; and the rendered HTML and PDF print dollar amounts without escapes or an index column. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * State how references are computed in the manifest; record the households' build The build_manifest reproducibility note said the references "were generated with policyengine.py X and policyengine-us Y against the certified PolicyEngine US populace dataset (<runtime default build>)". PolicyBench computes each reference with policyengine_us.Simulation from the household's own listed inputs (ground_truth.py) and records policyengine.py only for provenance; no dataset enters a reference. The note now says so, names the households' build from the run's scenarios.csv.meta.json, and says reference_output_refresh's dataset fields describe the reference runtime's default dataset. A new household_dataset block records that build (id, dataset, URI, sha256); the existing keys stay. tests/test_snapshot_artifacts.py checks the block against the scenarios meta, the note's wording, and docs/paper.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Date the adjudication record's judge verdicts from their sidecars The upgrade's refresh gave each judge_previous item a judged_on equal to the entry's adjudicated_on, which paired claude-opus-5-5 with dates before Opus 5.5's release (2026-09-21), left judged_on_utc at the replaced verdict's date, and dated the re-judge 2026-09-28, the local date of a run whose verdicts carry 2026-09-29 UTC timestamps. scripts/date_adds0928_judge_verdicts.py takes every date from the sha256-bound verdict.meta.json sidecars: the stage's current verdict for judge_rejudged_on and judged_on_utc, and the 20260922c audit tree's verdict for each judge_previous item whose judge model and classes match (otherwise it would rename the field adjudicated_on; all 54 match). Run on the staged record: 54 re-judge dates 09-28 -> 09-29, 9 judged_on_utc fixed (scenario_074 had Opus 5.5 on 2026-09-05), 51 previous Opus 5.5 dates -> 2026-09-23 and 2 previous Opus 5 dates -> 2026-09-05. tests/test_adjudications.py checks that every judge date falls on or after its model's release, that a re-judged entry is dated on or after the verdict it replaced, that each current date is one the manifest's judge provenance records, and names the two verdict-less llm_error entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Let a post-freeze re-export gate on the 22c asset The documented chain reruns finish_adds0928 --step export after a fix to the staged annotations. After the freeze the live pointer and the committed snapshot are this release's own, so resolve_base refused ("base pointer changed"). The export only needs the 22c payload to replay the incumbents against. resolve_live_base keeps the pre-freeze path (resolve_base) while the pointer names the 22c base. When it names this release, it reads the 22c run payload from BASE_COMMIT (3220a7a6), requires it to rewrap to the 22c asset's sha256 (01e7e72b...) with 42 models, and refuses any other pointer. The 22c replay gate itself is unchanged, and it passed on the rerun: every incumbent reproduces the 22c modelStats on the 22c references. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say which release 2.17.0 agrees with, and introduce the adapter first - The card, the paper and the upgrade README say 2.17.0 gives the same value "as 2.15.17" for all 1,984 outputs. The card and the paper state it after they introduce the Maryland adapter. - Release note: the check sentence names 2.15.17; the Arizona paragraph splits into three sentences, with the BBCE note's update adding the household as a fifth held back by income. - tests/test_disclosures.py rebuilds the card's engine sentences from the sidecar, the exclusion record and the 2.17.0 sweep, and the cost table test compares the whole header. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refreeze release 20260929 with the dated adjudications and manifest note Reran the documented chain on results/local/adds0928-v3: finish_adds0928 triage and export (the 22c replay gate passed), freeze_adds0928 --dry-run and the freeze, then paper/render_paper.py and freeze_snapshot.py --rendered-only. - annotations/.../us_adjudications.json: the staged record with its judge dates taken from the verdict sidecars (sha256 24176bf6... -> 0ea00ff9...). - manifest.json: the new adjudication pin, the household_dataset block and the rewritten reproducibility note; the rendered paper pins. - The payload is unchanged: dashboard-data.json sha256 e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6, and app/src/data.artifact.json still names it. - Paper render: PDF sha256 e6404ea2...; the cost sentence and table print plain dollar amounts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Unit-test the judge-date script on bound and unbound verdicts A synthetic pair of cases: the current verdicts date the re-judge and judged_on_utc (2026-09-29), a previous verdict whose classes match dates judge_previous (2026-09-23), and one whose class differs leaves the record's date under adjudicated_on. A second pass changes nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rebuild the references with the upgrade dated 2026-09-29 and the engine rule anchored to the sweep The records dated the engine upgrade 2026-09-28, the US Eastern date the wave began, while the sweep began at 01:42 UTC on 2026-09-29 and the three new exclusions are decided_on 2026-09-29. build_references_latest.py now dates the revision, the derivation and the new exclusions' notes 2026-09-29, takes each new-exclusion basis date from its decided_on, and states the rule with a time anchor: "References come from the newest policyengine-us release when PolicyBench begins the reference sweep; at publication PolicyBench checks that the newest release gives the same values." Max's ruling of 2026-09-28 keeps its date, labeled as his. The sidecar and exclusion record are rebuilt by the script from the committed files (its docstring gives the command). The reference CSV is byte-identical (sha256 e8bbba8f...); only record text and regenerated_at_utc (15:04 UTC) change. scripts/install_adds0929_references.py installs a rebuild into the snapshot and both stage copies only when that holds, and keeps the stage receipt's pins. verification/sweep_timing.json records the PyPI upload times (read live), the sweep's install, script and first-output times, and the pin commit; tests check the sweep began between the 2.15.17 and 2.16.0 uploads and the 2.17.0 check ran after its upload. The README states the rule the same way, dates the reviews by time zone, lists the five sweeps re-run on 2.15.17, and names the rebuild. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Bind each recorded judge flag to the verdict its date names The date script bound each judge_previous item to a 22c verdict by judge and classes only, so eight items (005 state, 028, 030, 042 federal, 051, 064 federal, 109, 112) were dated 2026-09-23 from verdicts that do not flag the reference, while still carrying judge_reference_suspect: true from an earlier run of the 2026-09-22 wave. The script now compares the flag: an item whose dated verdict does not raise it keeps it only with judge_reference_suspect_source naming that wave (flagged_sept22_wave.json, now committed), and a flag no run explains stops the script. A top-level source is dropped where the current verdict raises the flag itself (005 state, 112). The record now states its date conventions: adjudicated_on names the audit wave, whose decisions were written up to the day its release was committed; judge dates are UTC days from the verdict sidecars. Entries that were not re-judged get judged_on_utc. Where a later wave replaced the verdict a decision reviewed without keeping it (ten 2026-09-05 decisions), adjudicated_verdict records that verdict as its release published it. The script writes the verdicts it read to verification/judge_verdicts.json. Tests check every date and flag against that evidence, check that each decision's reviewed verdict is dated by its wave's release (46 decisions of the 2026-09-22 wave reviewed verdicts dated 2026-09-23, labeled as intended), check the published verdicts against git, and property-test date_entries (renaming rule, flag rule, idempotence) with Hypothesis. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Scope the card and paper to the sweeps each engine ran; fix the note's adapter and price wording - Card and paper: the defect and reading sweeps ran on 1.755.4; on 2.15.17 PolicyBench re-ran five (r02, NIIT, the SALT refund, mortgage residence and 40-hour readings), which move no scored output but scenario_066's SNAP, whose prompt states 40 hours. The card now states the limit the paper's Limitations gives, and the Limitations bullet names the re-run. - Paper: "this snapshot's scored reference outputs come from the fixed engine"; Claude Opus 5 judged "the 132 cases the September 5 additions joined that no later judge re-judged" (card too); the abstract's engine sentence is split and gives the fallback to the last published amount. - Card: the transport list attributes each JSON row as its model card records it (provider rejection for the Claude, DeepSeek V4.1 Flash, Kimi and Qwen rows; card choice for DeepSeek V4 Pro and GLM-5.2; the Gemini family default), and "before PolicyBench froze the references". Methodology says the same without a row list. A test derives the groups from the serving config and model cards. - Release note: the 2.17.0 check ran "with those conventions and an adapter that keeps Maryland county tax out of state income tax" (tested against the check's fix module), and DeepSeek V4.1 Flash is priced "at DeepSeek's peak list price". The note tests read the PyPI upload times from the timing record. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Derive the leaderboard's re-run counts; property-test the serving and exclusion-engine helpers - ModelLeaderboard states how many Claude rows ran without extended thinking and how many carry an auto re-run from SERVING_SENSITIVITY (servingSensitivityCounts), and the test ties those counts to the sensitivity data files. - fast-check (new dev dependency) properties for excludedOutputsByEngine (counts sum to the exclusions, distinct versions in numeric order, input order irrelevant) and its sentence, for chunkedServingModels and jsonContractClaudeModels, and for the re-run counts. - Differential checks against the Python side: the payload's excluded-by-engine counts equal the ones the rendered paper states (paper_results over reference_exclusions.json), and the serving helpers agree with the paper's serving-configuration table, replacing the test that repeated the implementation's filter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refreeze release 20260929 with the redated records; re-render and re-pin the paper Reran the documented chain on the stage (triage, export with the 22c replay gate, freeze dry-run, freeze), then rendered the paper and re-pinned it (freeze_snapshot.py --rendered-only). The payload changes only through annotation text: the three exclusion notes and 127 row annotations now say "Found in the 2026-09-29 engine upgrade" (130 string fields; no score, value or key moves). - dashboard-data-20260929 payload sha256 e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6 -> d146473d9bd7776638c59a0a20774dbe9026d8bcee0f2201e114146609ddf246 (app/src/data.artifact.json and the manifest) - run data.json.gz 3028225189ae... -> 2a6463c9ff9f... - us_adjudications.json 0ea00ff9480c... -> 0e58570a62aa... - us_case_notes.csv 9a4a3e374b49... -> 864cab96e670... - reference_exclusions.json 66ac8670fcee... -> ae28ade59705... - reference_outputs.csv.meta.json c6f1589da553... -> 5469664726ad... - reference_outputs.csv unchanged (e8bbba8fd3e9...) - the BBCE note data files' run_payload_sha256 follows the payload; their rows are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Scope the reference installer's text mask to the engine upgrade install_adds0929_references.py blanked the field names date, rule, basis, note and derivation at every depth, so it would have accepted a rebuild that redated an older revision or rewrote an older exclusion's note. The guard now blanks only the engine_upgrade revision's date, rule and each change's basis, the sidecar's regenerated_at_utc, the exclusion record's derivation, and the notes of the exclusions the upgrade added (the outputs its changed list names that the record lists). Every older revision and exclusion must be unchanged. The historical 15:04 UTC rebuild still passes the narrower guard. New tests refuse a rebuild that changes an older revision's date, rule or basis, an older exclusion's note, or a new exclusion's decided_on, and check that the committed records pass their own guard with exactly the three new exclusions masked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Require the wave's flag record for a kept top-level judge flag date_entries accepted any existing judge_reference_suspect_source on a top-level flag the current verdict does not raise, without checking that an earlier run of the 2026-09-22 wave flagged the case; only the judge_previous branch checked flagged_sept22_wave.json. The top-level branch now requires the case in the wave's flags too, and drops a stale source wherever the recorded flag matches the verdict's. The Hypothesis strategy now draws the top-level flag, the current verdict's flag and the source for each entry. The property checks that a mismatch stops the script unless the flag is set, sourced and in the wave, that the flag itself never changes, and that a source survives exactly where it explains a mismatch. Removing the wave check makes the property fail. Rerun on the stage, the script changes no date or flag. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Date this wave's adjudications by when they were written judge_verdicts.json and the adjudication record's date_conventions said the 2026-09-29 wave's decisions were written up to the day its release was committed, 2026-09-29, but that release has no commit yet. The wave's entry now gives adjudications_written_on 2026-09-29, with commit and pull_request left null for the lead to fill after the merge. The date conventions state the earlier waves' release commit days (2026-09-05 and 2026-09-23) and that the 2026-09-29 wave's decisions were written on 2026-09-29 UTC, after its reference sweep began, and bound a verdict's lateness by the last day its wave's decisions were written. The date script rewrote the staged record's date_conventions (its only change) and the evidence file. The test reads each wave's day from committed_on or, for an uncommitted release, adjudications_written_on, checks the record states the same days, and checks the sweep began that UTC day. The refreeze copies the record into annotations/. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the conventions, adapter and hours alias in the manifest's reference note The manifest's reproducibility note said each scored reference comes from policyengine_us.Simulation on policyengine-us 2.15.17 and the household's own inputs. The references also depend on the nine publication conventions, the Maryland output-scope adapter and the scenario builder's stated-hours alias. The note now names them "as the reference sidecar's engine_upgrade revision pins them (fix_modules, builder)", and docs/paper.md says the same. build_manifest reads the convention count from the sidecar (the latest_c_*.py modules) and stops the freeze if the revision's other modules or its builder note are not the ones the sentence describes. The manifest test rebuilds the sentence from the sidecar. The refreeze regenerates the manifest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Scope the release's claims to what was checked; record the 2.17.1 check - The abstract names the exclusions the audits found (28 unfixed engine defects, 27 unlisted inputs), in place of every output an unrun fix could move. - On 2.15.17, PolicyBench re-ran four September 22 sweeps and ran one new sweep, for the state and local tax refund reading. None moves a scored output by more than the $1 tolerance, and three move by less. verification/rerun_sweeps.json records what each sweep moves; a test recomputes the claim. - The 2.17.0 check is dated by when PolicyBench read PyPI (14:58 UTC). The paper takes the upload and check times from verification/sweep_timing.json. - A second pre-publication check found policyengine-us 2.17.1 (uploaded 17:24 UTC). It also gives the same value for all 1,984 outputs (verification/latest_final_2171.csv, sweep_timing.json later_checks). - JSON transport wording in Methodology, the paper and the sensitivity note follows the card: the card or its family default selects JSON, mostly because the provider rejects a forced tool call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refreeze with the reference note naming conventions, adapter and alias; re-render The manifest's reproducibility note now names the publication conventions, the Maryland output-scope adapter and the stated-hours alias. The paper is re-rendered and re-pinned. Payload unchanged: sha256 d146473d9bd7776638c59a0a20774dbe9026d8bcee0f2201e114146609ddf246. The abstract test reads the PDF with its page-number lines at page breaks removed. The abstract sentence crosses from page 1 to page 2, and pdftotext put the "2" mid-sentence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fetch full history in the CI test job; name the missing base commit The release gates read the September 22c base files from git at BASE_COMMIT. CI's default shallow checkout lacks that commit, so four tests failed with a bare CalledProcessError. The test job now fetches full history, and base_commit_blob says which commit is missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the pre-merge review of #182 and brief its fixes The in-session review (five lenses, each serious finding checked by a skeptic) confirmed one blocker: scenario_023 head_medicaid_eligible rests on the unlisted SSA-disability input that already excludes the household's SNAP, and the wave's own investigator and reviewer had flagged it without a recorded disposition. The brief applies rule 4 and hardens the gates the review's minors found weak. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the reviewers' full reasons for the 19 rechecked exclusions The sidecar's excluded_outputs_rechecked reasons, and final_actions.json's, were the reviewers' corrected_per_output text cut at 600 characters, which dropped, for example, the unlisted-input basis the scenario_100 SNAP review insisted on. Both now carry the full text from clusters.json. - build_references_latest.py takes each reason from clusters.json (and refuses a final_actions reason that differs), and exposes its record functions: reviewed_reason, audit_exclusions, exclusion_derivation and dump_record. It also carries final_actions.json's audit_exclusions: outputs the audit excludes on review, whose value must not move. - rewrite_reference_records.py applies those functions to the installed records without the 1,984-output recompute, and copies the CSV byte for byte. - install_adds0929_references.py treats the rechecked reasons as record text, accepts only the audit exclusions final_actions.json lists (each exactly as recorded), and keeps every install that changed a staged file in the receipt. It installed the rewrite: the reference CSV and the exclusion record are unchanged; only the sidecar's reasons moved. Tests: each recorded reason equals the reviewer's full text; the builder's record functions give the committed records; the installer accepts a rewritten reason and a listed audit exclusion, and refuses an unlisted or altered one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Exclude scenario_023 head_medicaid_eligible on review (unlisted SSA disability) The excl_snap_ssi_disability investigation flagged this scored output out of cluster and its independent review agreed; nothing recorded a disposition. The head's MAGI is 141.2% of the poverty guideline, above the 138% adult limit, so only a disability pathway leads to Medi-Cal, and California's 250% Working Disabled Program requires the SSA definition of disability (42 CFR 435.540(a)): meets_ssi_disability_criteria, the input that already excludes this household's SNAP. policyengine-us 2.15.17 tests the broad is_disabled flag instead. Rule 4 excludes the output. Computed on 2.15.17 with fixes/latest_final.py (scripts/probe_023_medicaid.py, verification/probe_023_medicaid.json): - reference system: 1 under the stated facts, reading A and reading B (WORKING_DISABLED_BUY_IN, WORKING_DISABLED_BUY_IN, SENIOR_OR_DISABLED); - WDP disability test reading meets_ssi_disability_criteria: reading A 0 (category NONE), reading B 1 (SENIOR_OR_DISABLED). - reference_exclusions.json gains the entry (frozen 1.0, alternative 0.0, 2.15.17, decided 2026-09-29), installed with rewrite_reference_records.py and install_adds0929_references.py into the snapshot and both staged copies; the reference CSV and sidecar are unchanged, and the engine_upgrade revision's changed list keeps exactly the engine changes. - final_actions.json lists it under audit_exclusions; clusters.json records the disposition under reconciliations; sweep_moves.csv marks the row excluded with its flagging cluster; verification/reviews/ pr182_review_023.md gives the review and the computed values. - README: 56 exclusions (28 engine-defect, 28 unlisted-input), 1,928 scored, and an "Also excluded on review" section. The staged adjudication record gains the matching prompt_ambiguity entry (results/local, frozen by the next freeze); triage rebuilt the case note to prompt_ambiguity with the adjudication sentence and moved the 22 zero answers' rows to prompt_ambiguity. Nothing else in the staged annotations changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Gate every reference column, the exact exclusion set and the 22c replay Review minors on the finish driver's gates: - reference_revision() compared only `value`. It now requires the same columns and every other column (impact_weight among them) equal to the 22c base on every row, and compares values NaN-aware, so an unlisted weight change or a value turned NaN is refused. - The exclusion check was a count (52 + added) in resolve_base only. check_exclusions now requires exactly the 22c exclusions (from BASE_COMMIT, each entry unchanged), plus the revision's newly excluded outputs, plus final_actions.json's audit_exclusions, and nothing else; an audit exclusion may not be an engine change. resolve_base runs it first, resolve_live_base runs it after the freeze, the export runs it on the staged record, and the freeze runs it before re-pinning the manifest. - The incumbent replay had no test. New tests drive export() with a revision and an audit exclusion: an incumbent that drifts on the 22c references is refused, the replay writes the 22c bytes, and the no-revision branch refuses any drift. With replay_base_references made a no-op, two tests fail (before: all 64 passed). - The drift report names both causes: the reviewed revision and the audit exclusions, the only ways the references differ from 22c, with the replay showing incumbents reproduce on the 22c references. Mutation checks: a value-only reference_revision fails 2 tests, a count-only exclusion check fails 11 (the swapped case among them). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Close three test gaps: 0/1 flips, the pin commit on main, HTML paper claims - test_every_changed_reference_is_listed_and_reviewed exempted any change of at most 1 as "within $1", so a 0/1 eligibility flip of exactly 1 passed without an approved action. It now uses paper_results.moves_beyond_tolerance (any change for a 0/1 flag), and a change recorded as engine_upgrade_within_1 must stay within the tolerance. The reviewer's mutation (028 PA dropped from approved, relabeled within_1) passed before and fails now. - The pin-commit check against git skips once PR #182 is squash-merged, since the commit leaves main's history; the skip now says so. A new git-free test checks the recorded order: 2.15.17 uploaded, the sweep's first output (01:42:18), the pin commit (01:58:29), 2.16.0 uploaded, the reference build. The pin follows the sweep's start, so it asserts that order. - The abstract-exclusions and engine-times paper tests built their HTML and PDF texts together, so both skipped wherever the PDF extractor is missing (CI). Each is split: the HTML half always runs, and only the PDF half skips. In a CI-like environment the HTML halves now run; the abstract one fails against the committed render, which predates the audit exclusion (28 unlisted-input outputs, render says 27), until the paper is re-rendered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Revert the audit exclusion too when rebuilding the 22c drift baseline test_previous_release_scores_rebuild_from_this_snapshot rebuilds release 20260922c's scores from this snapshot with the engine upgrade reverted, and concludes the note's drift is the reference change alone. The references now also carry the audit exclusion of scenario_023 head_medicaid_eligible, so the rebuild restores it to scoring as well, and the docstring names both causes. The 22c half reproduces all 42 published scores exactly; the current half matches once the payload is re-exported on the new exclusion record. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the review's evidence scripts out of ruff 4b7d0520 committed the pre-merge review's probe scripts (docs/adds0928/review_evidence/probe023.py, rescore023.py) as they ran, and `ruff check .` then failed on them with 28 errors. They are records, like reference_audit/, so ruff excludes the directory rather than rewriting them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refreeze release 20260929 with the scenario_023 Medicaid exclusion Reran the documented chain on the stage (triage, export with the 22c replay gate, freeze). A second freeze reproduced every frozen file byte for byte. The payload is now fdaaa738 (123,359,634 bytes; was d146473d). The frozen annotations carry the 69th adjudication, and the Medicaid output's 24 wrong rows (22 llm_error, 2 parse failures) move from scored misses to excluded-output description. Pins moved to the frozen artifacts: 56 exclusions (28 unlisted-input), 1,928 scored outputs, 7,772 annotated rows, 7,768 exact misses, 652 parse failures (Kimi K2.6 390, Kimi K3 58), 712 contract violations, 2,065 rows on excluded outputs, 802 prompt_ambiguity rows, 69 cases (50 judged llm_error), and 4 exclusions valued on 2.15.17. The app pins move too: leader 95.003 and the median Medicaid accuracy 95.38, now "about 1 in 22 people". judge_verdicts.json records the verdict the new adjudication keeps. The BBCE note's data files re-pin the payload; their rows are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rescore the Claude thinking sensitivity on 1,928 scored outputs Reran scripts/sensitivity_by_variable.py and scripts/rescore_sensitivity_summaries.py --release dashboard-data-20260929 on the refrozen snapshot; the reruns reproduce the committed assets. The rescore script gains the word for 56. Board -> auto, exact within $1: Fable 5 83.625 -> 91.521 (would rank #7, +7.9), Opus 5 84.030 -> 90.013 (#9), Sonnet 5 72.642 -> 84.803 (#19), Fable 5.1 90.828 -> 91.728 (would rank #6, was #7, +0.9). The sensitivity doc, servingSensitivity.ts and its tests carry these. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Claude Sonnet 5.5 debuts fifth: rename the release note, name both drift causes On the refrozen payload GPT-6 Luna (92.0970) passes Claude Sonnet 5.5 (92.0808), so the note becomes 2026-09-29-claude-sonnet-5-5-debuts-fifth (slug, file, notes index, the BBCE note's link and the tests). The title's ordinal still follows facts['sonnetRank']. The note calls the pair about level: Luna leads by under 0.02 points, and the test holds the gap under 0.1. The exclusions paragraph adds the Medicaid output: the head's MAGI (141% of the poverty guideline, from the probe), the 138% expansion limit, the Working Disabled Program's Social Security definition, and the general flag policyengine-us 2.15.17 tests (ca_wdp_disability_eligible reads is_disabled). The drift paragraph names both causes, the new references and the Medicaid exclusion, and gives 20 of 42 incumbents matching that output and 22 missing it. The three neighbour swaps become five risers, derived by _models_moving_up, which asserts that every other pair keeps its order. Facts recomputed from the frozen payload: Sonnet 92.1 #5, Luna 92.1 #4, Grok 4.7 88.3 #11, DeepSeek V4.1 Flash 87.1 #15, GPT-6 Sol 95.0, Opus 5.5 93.7, GPT-5.6 Sol 93.6, 1,928 scored of 1,984, 56 excluded, drift 0.14 to 0.83 points. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Describe the Medicaid exclusion in the paper, card and app copy; re-render The paper, card, docs/paper.md, the app's live version description and the referenceEngine comment now give 56 exclusions (28 unlisted-input, four valued on 2.15.17), 1,928 scored outputs and the audit counts from the refrozen annotations (7,772 annotated rows, 7,768 exact misses, 55 unparsed answers on excluded outputs). The paper's disability section, audit table, engine-upgrade paragraph and exclusion paragraph describe the California output. Its MAGI is above the 138% adult limit, so only a disability pathway reaches Medi-Cal. The Working Disabled Program uses SSI's definition (42 CFR 435.540(a)). policyengine-us 2.15.17's ca_wdp_disability_eligible reads is_disabled and gives 1 under either reading. The exclusion paragraph says every exclusion but this one was derived mechanically. Rendered with paper/render_paper.py and re-pinned with freeze_snapshot.py --rendered-only (pdf 50109331, web index 0d75cce9). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Mark the scenario_023 Medicaid judge diagnosis as description; refreeze The staged adjudication for scenario_023 head_medicaid_eligible now carries the sentence every other excluded output's reasoning uses: "Judge diagnoses are retained as description; the class prompt_ambiguity records that the reference, not the model, is indeterminate here." The case note keeps the judge's buy-in diagnosis as description, and the sentence now says so. Reran the documented chain on the stage: triage, export (22c replay gate held; Sonnet 5.5 rank 5/45, Grok 4.7 11/45, DeepSeek V4.1 Flash 15/45, all on 1,928 scored outputs), freeze dry-run, freeze. Only annotation text moved: the payload is now a5cb9989 (123,362,946 bytes; was fdaaa738, 123,359,634), the run payload 9b807f4a (was 86f35f51). No score, rank or count changed; the analysis files are byte-identical. The BBCE note's data files re-pin the payload through scripts/bbce_households_20260929.py; their rows are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Release note: give the incumbents' rise as the two causes' joint effect The drift sentence now opens "Together, the new references and the Medicaid exclusion raise the exact rate of every one of the 42 earlier models", and the next sentences keep the 20 matched / 22 missed split on the Medicaid output. test_notes pins the new sentence and checks why it must be joint: _exact_under can score the current references with the audit exclusion put back, and the Medicaid exclusion by itself lowers the rate of each of the 20 incumbents that had matched that output. Flipping that assertion fails the test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Card: name the disability test one way, as the paper does The card's exclusion paragraphs called the same test "the SSI disability criterion", "the Social Security definition of disability" and "the Social Security definition". They now say "SSI's definition of disability" throughout, the paper's name for what California's Working Disabled Program requires under 42 CFR 435.540(a) (the SSI program's definition). The unlisted input reads "whether a person meets SSI's definition of disability". No figure changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Paper: name the California Medicaid output and describe both engines' sweeps In the exclusions paragraph: - "Every exclusion but that Medicaid output" now names "the California Medicaid output", here and in the closing sentences. - The colon clause covers both engines. Each exclusion but that one is an output a sweep of every reference moved, on policyengine-us 1.755.4 or, for the state and local tax refund reading, on 2.15.17. On 1.755.4 PolicyBench excluded every output an alternative reading moved and every output a defect not fixed upstream moved by more than a dollar (the old text dropped the dollar threshold, per rule 3 of reference_audit/2026-09-22). On 2.15.17 the new sweep moves three previously scored federal income tax outputs beyond the tolerance, and PolicyBench excluded them. - The closing clause reads plainly: PolicyBench found the output when it re-reviewed the household's excluded SNAP output, and excluded it because the program itself requires SSI's definition of disability. paper_results.rerun_sweep_new_excluded_outputs computes the "three" from verification/rerun_sweeps.json: outputs a sweep first run on 2.15.17 moves beyond the tolerance that no 1.755.4 exclusion held. test_reference_upgrade checks that set equals the engine upgrade's new exclusions, all federal income tax. It also checks that every other output the new sweep moves beyond the tolerance was excluded on 1.755.4. docs/paper.md does not carry this text, so it needs no change. Re-rendered with paper/render_paper.py, which also picks up the refrozen run payload (9b807f4ab9a9). Re-pinned with freeze_snapshot.py --rendered-only (pdf 0b6b40cd, web index 593bd25a). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test the freeze's exclusion gate: an extra or swapped exclusion stops it freeze_adds0928.main calls check_exclusions on the frozen run before it re-pins the manifest, and no test covered that call. A new fixture builds a stage and frozen run that pass every check main makes before the gate. The receipt, recombination, 45-model roster, staged-versus-frozen reference files and incumbent-prediction checks run as written. The schema, treatment, run-state evidence and adjudication validators, which need a full release, are stubbed. freeze_snapshot.main is a sentinel, and the freeze_snapshot attributes main assigns are registered with monkeypatch so they are restored. - Control: with exactly the 22c, revision and audit exclusions, the gate passes and the manifest pins the frozen reference_exclusions.json. - An extra exclusion, or a swapped one with the right count, written to the staged and frozen records alike, passes the dry run. The freeze then raises SystemExit with the gate's "exclusions differ ... unexpected" or "... missing" message, and no file in the workspace changes. Mutation check: with the check_exclusions call removed, both refusal cases fail (the freeze pins the record and reaches freeze_snapshot.main), and the control still passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Release note: name the disability test as the paper and card do Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…venance (#185) * Fingerprint policyengine.py 6.x's bundle manifest in PolicyEngine provenance policyengine_provenance_inputs() hashed only policyengine's per-country release manifests (data/release_manifests/{country}.json). policyengine.py 6.1.2 ships none; it ships one data/bundle/manifest.json, which _load_raw_policyengine_manifest (and policyengine's own get_release_manifest) read each country's release from under data_releases. Under the pinned 6.1.2 the fingerprint recorded {'us': None, 'uk': None}, so an in-place edit of the bundle manifest between the provenance writer and a worker went undetected. The pre-merge delta review of #182 showed it: editing data_releases.us.bundle_id left the fingerprint unchanged. The fingerprint now also records bundle_manifest_sha256 (None when the file is absent). Both manifest paths are module constants shared by the reader and the fingerprint, so they cannot drift apart. The per-country layout (4.16.1) fingerprints as before. The supervisor's module docstring said computing provenance imports policyengine and builds the US and UK tax-benefit systems. That holds only when a country's installed model is the exact version policyengine.py pins (policyengine-us 1.723.0 under 4.16.1: re-measured 11.1 CPU-s, 909 MB max RSS, policyengine_us.system and policyengine_uk.system loaded). With 6.1.2 and policyengine-us 2.15.17 (6.1.2 pins 2.2.1) it loads no PolicyEngine module and reads only package metadata and the bundle manifest. The docstring now describes both cases; computing once per run still bounds the cost in the first. Tests: the reviewer's in-place edit, run for each layout, makes a written provenance file stale and the worker recompute the edited release; each layout's manifest hashes; the installed policyengine.py's manifest is hashed; and a Hypothesis property that, across any mix of the two layouts' files, fingerprints are equal exactly when the files are byte-identical, so equal fingerprints hand every country the same raw release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address review nits: qualify the skip flag, align the runbook, note test coverage From the independent Opus review of #185 (APPROVE, nothing blocking): - supervisor.py: the 4.16.1 import built both country systems only with POLICYENGINE_SKIP_COUNTRY_IMPORTS unset (policyengine/__init__.py gates each country import on it); say so. - docs/runbook.md: workers read the provenance file "instead of computing it", not "instead of importing policyengine", since under policyengine.py 6.1.2 plus policyengine-us 2.15.17 computing it imports nothing either. - tests: note that the manifest paths are spelled out on purpose, and that the property test covers both layouts' files, which is all the raw reader opens. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin the mixed layout and whole-file hashing in the property test From the in-session execution verifier and delta reviewer of #185 (both APPROVE, nothing blocking): - A fingerprint that hashed the bundle manifest only when no per-country file existed survived some Hypothesis seeds (for example --hypothesis-seed=5), because no deterministic test had both layouts present. policyengine.py 6.1.2's own loader reads the bundle manifest even then. An explicit @example now covers it. - The strategy varied only data_releases, so hashing json.dumps(data_releases) instead of the file was caught only by the exact-hash test. Bundles may now carry dataset_overlays, which 6.1.2's loader applies, and an @example pins it. - The relocated_policyengine docstring said _load_policyengine_manifest "returns None" when no model matches its pin; it is not called then. Reword. - The in-place edit is credited to the pre-merge delta review of #182, which was not posted on GitHub. - test_provenance_file_reproduces_direct_computation ids came from str(set), which varies with PYTHONHASHSEED, so a node id collected in one process was not found in the next. Ids are now the sorted country codes. Both mutants are caught on seeds 1 and 5: the mixed-layout one by the property test, the data_releases-only one by it and the exact-hash test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Add the Claude Sonnet 5.5 model card and metadata
The Models API lists claude-sonnet-5-5 with created_at 2026-09-28, the
release date the model page gives. It rejects forced tool use with a 400,
and a request with no thinking parameter returns a thinking block, so the
card runs the JSON contract whole-scenario on the sync path, as Opus 5.5
and Fable 5.1 do, and the row reasons at the API default (adaptive
thinking, effort high). Priced at the standard $2/$10 per 1M with $0.20
cache reads and $2.50 cache writes from the pricing page, where no
footnote marks the row as introductory. litellm's bundled map lacks the
id, so eval_no_tools registers it locally with the Fable line and Opus
5.5. Gauntlet: 3/3 and 16/16 parsed, $0.027 per scenario.
The roster test in test_paper_results still pins the display names to the
42-model frozen board; it passes again once the board takes the new row.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Register Grok 4.7 and DeepSeek V4.1 Flash
Grok 4.7 (xai/grok-4.7, released 2026-09-21 per x.ai/news/grok-4-7) and
DeepSeek V4.1 Flash (deepseek-flash on the native API, released 2026-09-10)
join the registry beside Claude Sonnet 5.5, with prices from the providers'
pages. DeepSeek V4.1 Flash onboarded on the JSON contract (the forced tool
call is rejected in thinking mode); Grok 4.7's card is provisional until its
onboarding finishes with a longer timeout. The frozen-roster paper test fails
until the board is refrozen with the three new rows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Set the Grok 4.7 model card from its onboarding
Forced tool contract, 3/3 and 16/16 whole-scenario; the timeout is 1800s
because the whole-scenario probe took 472s and an earlier attempt timed out
at 600s.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add the stage-2 driver for the September 28 additions
finish_adds0928.py folds the Claude Sonnet 5.5, Grok 4.7 and DeepSeek V4.1
Flash runs onto the 20260922c board (42 models, 1,932 scored outputs), keeps
every incumbent's modelStats identical, prepares the cases the additions join
for the Opus 5.5 judge, gates adjudications, and exports only into a stage
directory. It refuses partial runs unless --early --partial is given on
copied, synthetic run states, and marks such output PARTIAL. The snapshot and
freeze helpers and the design note come from the same preparation; 57 tests
cover them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Pin policyengine.py 6.1.2 and policyengine-us 2.15.17
References now come from the newest policyengine-us, run through
policyengine.py (Max, 2026-09-28). policyengine.py 6.1.2 certifies
policyengine-us 2.2.1, which still has the SNAP rounding defects fixed
since, so policyengine-us is pinned separately; the runtime provenance
records model_matches_policyengine_bundle: false, as the July run did
(policyengine.py 4.16.1 certified 1.723.0; 1.755.4 was installed).
policyengine.py 6.1.2 requires Python >= 3.11; CI runs 3.12.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Let the stage-2 driver release a reviewed reference revision
The driver refused any incumbent modelStats drift, which is right for an
additions-only release. Regenerating references on the newest
policyengine-us changes some scored references, so every incumbent's
statistics move. export() now accepts that only when:
- the snapshot's references differ from the pinned September 22c files
(read from git at 3220a7a and checked against their sha256 pins) by
exactly one committed engine_upgrade revision whose `changed` list
names every changed output with its value, and nothing else; and
- an export of the same bundle on the 22c references reproduces every
incumbent's live modelStats exactly, so the references are the only
source of drift.
It writes incumbent-drift.json with each incumbent's before/after exact
rate and score. With the base references committed, behaviour is
unchanged.
Invariant tested by property (Hypothesis): for any set of changed
outputs, the revision is accepted iff it lists exactly those outputs with
those values; omissions, extra entries, wrong values, a non-upgrade
revision, or a sidecar that does not pin the CSV are refused.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Regenerate the US references on policyengine-us 2.15.17
Max, 2026-09-28: references always come from the newest policyengine-us.
2.15.17 is the newest release on PyPI (2026-09-29 00:23 UTC).
policyengine.py 6.1.2 is recorded for provenance only (its certified
bundle is 2.2.1).
The nine pre-freeze publication conventions are re-expressed for
2.15.17. The SNAP hold needed a rewrite: reforms now apply after
uprating, and FY2027 is carried as published values. The eight upstream
fixes the September 22 audit regenerated are all in 2.15.17. An output-
scope adapter keeps Maryland county tax (#8888) out of the state income
tax output.
Result, against release 20260922c:
- 4 scored references change:
- 008 NJ refundable credits, NJ CTC P.L.2026 c.26 (#8971);
- 013 AZ SNAP, AZ expanded categorical eligibility at 200% from March 2026;
- 028 PA reduced-price school meals, child support counted,
7 CFR 245.6(a)(5)(ii);
- 082 NY refundable credits, Empire State child credit rounding (#9425).
- 3 outputs newly excluded as unlisted-input readings: 033, 078 and 117
federal. Whether a listed state and local tax refund is income turns on
prior-year facts the prompt does not give; 2.15.17 counts it
(#9422, fixing #9122).
- 2 move by under $1 (078 and 117 state).
- 19 excluded outputs that move were re-reviewed and stay excluded.
Scored outputs: 1,932 -> 1,929. Exclusions: 52 -> 55.
The builder now also passes stated usual weekly hours to
weekly_hours_worked_before_lsr, whose default #9261 changed from 40 to
0; scenario_066's SNAP stays 3,576.
reference_audit/2026-09-28 records the modules, the 1,984-output sweep
and 16 root-cause clusters, each with an investigation and an
independent adversarial review. The one cross-cluster conflict (078
federal) is reconciled explicitly.
tests/test_reference_upgrade.py checks every module hash, that the
sweep table is the committed reference, and that every change is listed
and review-backed. Ruff's target moves to py311 with requires-python.
Snapshot, payload, paper and sensitivity artifacts still pin 20260922c;
the freeze for this release updates them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Rewrite the derivation narratives of the changed references
Ten reference narratives change:
- the nine outputs the engine upgrade changed (four scored changes, three
new exclusions, two moves under $1);
- scenario_064 dependent2_chip_eligible, whose frozen narrative said the
18-year-old exceeds CHIP's age limit. CHIP covers children under 19.
Dependent 2 fails on income (321.5% of the poverty guideline against
Wisconsin's 306%) and on employer-sponsored insurance. That narrative
is what the September 28 judge flag on the case repeated.
The writer is the usual one (claude-haiku-4-5 via
case_reference_explanations._prompt), grounded in each change's
reviewed basis and the 2.15.17 trace. Four drafts misstated the trace
and are replaced by hand-written text that follows it:
- 013 AZ SNAP: $24 a month from March, not a monthly $240 from January.
- 008: how the NJ EITC derives from the federal credit.
- 082: how the NY CDCC derives from the federal credit.
- 028: when child support entered the school-meal income.
reference_audit/2026-09-28/scripts/narratives_latest.py records them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Derive the stage-2 exclusion count from the reference revision
52 exclusions on the September 22c references, plus each exclusion the
committed engine_upgrade revision lists among its changed outputs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Read policyengine.py 6.x's bundle manifest; record the upgrade's adjudications
policyengine.py 6.x ships one bundle manifest (policyengine/data/bundle/
manifest.json) in place of per-country release manifests, so the
unbundled runtime metadata lost its data package and version. That broke
the export's reference provenance check. The runtime now reads
data_releases.<country> from the bundle manifest when no per-country
manifest exists. The reference sidecar is rebuilt with it: data
microcosm-data 0.1.0, build populace-us-2024-spm-20260915, bundle
us-6.1.2. The reference CSV is byte-identical (sha256 e8bbba8f...),
and the build script now reads the 22c base from git, not from the
checkout.
Adjudications for the upgrade:
- 54 re-judged entries carry the September 28 judge's class verbatim,
with the replaced values under judge_previous. Adjudicated classes,
exclusions and reasoning are unchanged.
- The three new unlisted-input exclusions (033, 078, 117 federal: SALT
refund taxability) are adjudicated prompt_ambiguity with
reference_verdict unlisted_input.
- Two cases whose judge labelled model rows reference_later_law for
applying superseded law are resolved as llm_error: OBBBA, enacted
2025-07-04, and NJ P.L.2026 c.26, approved 2026-06-30, both predate
the 2026-07-03 freeze.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Freeze release dashboard-data-20260929: 45 models on policyengine-us 2.15.17
Local build only; nothing is uploaded or published.
Board and references:
- 45 models: Claude Sonnet 5.5, Grok 4.7 and DeepSeek V4.1 Flash join the
42 of 20260922c.
- 1,929 scored outputs, 55 exclusions.
- Payload sha256 e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6.
- Audit: 674 judged cases, 68 adjudications.
The export's reference-revision gate passed. Re-exported on the 22c
references, every incumbent reproduced its live modelStats exactly, so
all drift comes from the reviewed engine upgrade. Every incumbent's
exact rate rises 0.36 to 0.54 points; the top eight keep their order.
Staged ranks: Claude Sonnet 5.5 4th (91.97%), Grok 4.7 12th (88.23%),
DeepSeek V4.1 Flash 15th (87.27%).
The release is tagged 20260929. The references were regenerated on
2026-09-29 UTC, and the manifest dates them by regenerated_at_utc when
the sidecar has one. freeze_adds0928 first brings the manifest's
reference pins up to the committed, reviewed revision (checked by the
driver's gate), because freeze_snapshot checks staged references
against those pins.
Prose, paper render, sensitivity docs, notes and test pins follow.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Rescore the reasoning-sensitivity summaries on release 20260929
scripts/sensitivity_by_variable.py and rescore_sensitivity_summaries.py
--release dashboard-data-20260929: 1,929 scored outputs, 55 exclusions,
45 board rows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Brief the stage-3 prose, notes, paper and pin updates
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add Max's voice rules to the stage-3 brief
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Pin release 20260929's counts in the tests, sensitivity doc and app
Every new pin is recomputed from the frozen artifacts of
dashboard-data-20260929 (45 models, 1,929 scored outputs, 55
exclusions, 68 adjudications, 674 judged cases).
- paper_results: regenerated_reference_* keep counting the September 22
convention and upstream-fix regenerations (26); new engine_upgrade_*
properties count the September 28 move to policyengine-us 2.15.17 from
the sidecar's engine_upgrade revision (4 scored changes, 2 within the
tolerance, 3 new exclusions, 19 rechecked exclusions). A 0/1 output
counts as moved on any change.
- The sensitivity doc's ranks, scores and per-program tables move to the
45-model board; servingSensitivity.ts carries the rescored auto runs.
- The app's copy of the serving configuration is refreshed from the
frozen file, as prepare-data does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* State the 45-model board and the policyengine-us 2.15.17 references in the card and app copy
- Benchmark card: a reference-outputs section states the engine
(policyengine-us 2.15.17 through policyengine_us.Simulation, with
policyengine.py 6.1.2 recorded for provenance), the pre-freeze
publication rule, the Maryland output-scope adapter, and what the
upgrade changed. Audit scope, exclusion and adjudication counts move to
this release; Claude Sonnet 5.5 and DeepSeek V4.1 Flash join the rows
whose provider rejects a forced tool.
- Paper guide and artifacts doc: this release's counts and payload size.
- Methodology names the engine from the board's own payload and counts
ten chunked rows of 45; the leaderboard's serving note names Claude
Sonnet 5.5 among the rows that reject forced calls.
- data.versions.json describes the live version by its engine.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Update the paper for the 45-model board on policyengine-us 2.15.17, render and re-pin it
- The abstract and the snapshot table name the reference engine,
policyengine-us 2.15.17; policyengine.py 6.1.2 is listed as recorded
for provenance, since the references come from policyengine_us.Simulation.
- A reference-credibility paragraph states the move from 1.755.4: the
re-expressed publication conventions, the Maryland output-scope adapter,
the weekly-hours default, the four scored changes, the three new
unlisted-input exclusions (SALT refund taxability, 26 U.S.C. 111(a)),
the two moves under $1 and the 19 rechecked exclusions, all read from
the sidecar's engine_upgrade revision. The discrepancy table gains a
row for it.
- The September 22 audit prose keeps its history and names 1.755.4 where
it described the engine of the time; the limitations say the engine
defects persist in 2.15.17 and that pre-freeze law reaches a reference
only once the engine encodes it. Claude Sonnet 5.5 joins Opus 5.5 as a
JSON row without a sensitivity run; the judge paragraph covers the
September 28 additions (five audit waves).
- paper_results names the dataset build the households were sampled from
(populace-us-2024-5da5a95-20260611, which the population weights also
use), not the reference runtime's later default build, and adds the
references' rebuild date.
- Rendered with paper/render_paper.py and re-pinned with
scripts/freeze_snapshot.py --rendered-only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add the September 29 release note and update the BBCE note for release 20260929
- A new note, "Claude Sonnet 5.5 debuts fourth as the references move to
the newest PolicyEngine", covers the three additions (ranks, costs,
serving treatments, and the deepseek-flash alias every DeepSeek V4.1
Flash answer reports), the move to policyengine-us 2.15.17, the four
scored changes, the three new exclusions and the incumbents' drift.
tests/test_notes.py recomputes every fact from the frozen snapshot and
pins every sentence beside its evidence. The drift baseline rebuilds
release 20260922c's scores from this snapshot (the upgrade's changes
reverted, its exclusions removed); a second test checks the rebuild
against 20260922c's committed payload in git history.
- The BBCE note keeps its release-20260922b figures and closes with a
dated update: the new models' answers for the four households held
back by income and the four held back by savings, and the Arizona
household the new references add, at $240 with all 45 models answering
$0. scripts/bbce_households_20260929.py writes every model's answers
to notes/data/*_20260929.csv, which the test regenerates.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Put the actor first in the paper's engine paragraph and the card's engine sentence; re-render and re-pin
PolicyBench rebuilt the references with policyengine-us 2.15.17 in place
of 1.755.4; policyengine-us #9261 changed the weekly-hours default; the
paragraph drops "now" and the convention count. The card names
policyengine_us.Simulation as what PolicyBench runs. Rendered with
paper/render_paper.py and re-pinned with freeze_snapshot.py
--rendered-only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Voice pass on the release-20260929 prose; re-render and re-pin the paper
Wording only; every number and fact is unchanged.
- Release note: the title names policyengine-us 2.15.17 instead of "the
newest PolicyEngine". PolicyBench is now the actor for the sandbox fixes,
the freeze and the ported conventions. The #9425 fix and the engine's
school-meal change are named. The SALT-refund sentence is split, and
the scored-output sentence is active. The list of swaps no longer
hangs off "every other model keeps its place". The closing pointer
names PolicyBench's September 29 update.
- BBCE note update: the Sonnet 5.5 list is punctuated so "citing BBCE"
stays with the Connecticut answer. The Arizona sentence is split at
"encodes the change".
- Paper: the abstract, the discrepancy-table row, the engine-upgrade
paragraph and the limitations use active voice. The row label names
the version. Numbered sentence starts are gone. The garden path in
"read from 40 hours" is fixed.
- Card: the four upgrade changes are the list's subject. The judge's
"later law" cases are described plainly, and the passive re-review
sentences are recast.
tests/test_notes.py pins the new sentences. It also checks the sidecar
basis for the #9425 and school-meal wording. Rendered with
paper/render_paper.py and re-pinned with freeze_snapshot.py
--rendered-only. pytest: 955 passed, 6 skipped. bun test: 149 passed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Record the stage-3 report and brief the review fixes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Anchor the reference engine to its sweep date and record the 2.17.0 check
policyengine-us 2.15.17 stays the recorded engine, stated as a dated fact:
the newest release when PolicyBench began sweeping the references on
2026-09-29 (uploaded 00:23 UTC). 2.16.0 (04:10 UTC) predates the 11:57 UTC
rebuild, so "newest when the references were rebuilt" was false.
- reference_audit/2026-09-28/verification: latest_final_2170.csv and .log,
sweep_latest.py --fix fixes/latest_final.py run in the .venv-pe2170
venv (policyengine 6.1.2, policyengine-us 2.17.0, the newest release at
publication, uploaded 12:21 UTC). The README gains a Verify step and the
dated wording; its title uses the upgrade date, September 29.
- tests/test_reference_upgrade.py: the 2.17.0 sweep reproduces every one of
the 1,929 scored references exactly, its 19 moved outputs are the
rechecked exclusions at their recorded 2.15.17 values, and it agrees
with 2.15.17 (sweep_moves.csv final) on all 1,984 outputs.
- The card says each scored reference comes from 2.15.17 and that the 55
excluded outputs keep their decided values (52 on 1.755.4, 3 on
2.15.17; the 19 that move on 2.15.17 were re-reviewed and stay
excluded). data.versions.json's description says the same.
- docs/paper.md: reference_output_refresh records the reference runtime's
default dataset (populace-us-2024-spm-20260915), which reference
computation does not read; the households came from
populace-us-2024-5da5a95-20260611 (new household_dataset block).
- "September 28 additions" becomes "September 29 additions"; script and
test comments date the upgrade 2026-09-29. The sensitivity note drops
a stale "newest JSON-transport row", and the BBCE update says the
scored references move.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tie the release note's drift, engine, cost and Arizona facts to data
- notes/data/release_20260922c_exact.csv (+ .meta.json): the 42 no-tools
exact scores release dashboard-data-20260922c published, copied from its
asset (dashboard-data.json, sha256 01e7e72b..., the pointer at 3220a7a6)
by scripts/release_20260922c_scores.py, which refuses any other file.
- tests/test_notes.py computes the drift range, the three neighbor swaps,
"GPT-6 Sol still leads" and "every one of the 42 earlier models" from
that fixture and the frozen payload, with no git access. The vacuous
roster check becomes "the board minus the 22c roster is the three
additions". The rebuild test compares the reverted-upgrade rebuild with
the fixture instead of a git blob CI never fetched.
- Engine paragraph: "the newest policyengine-us release" becomes the
dated fact (2.15.17, newest when PolicyBench began sweeping on
2026-09-29, uploaded 00:23 UTC) plus the 2.17.0 check at publication;
the upload times are checked against the upgrade record.
- DeepSeek V4.1 Flash's cost is stated at DeepSeek's standard list price,
the peak rate config.py prices the row at.
- Arizona: the fifth household held back by income in the BBCE note's
sense (with Connecticut, Texas, Michigan and Wisconsin), not the fifth
to qualify only through BBCE; the update covers the four other
income-held and the four savings-held households.
- The title's ordinal equals facts['sonnetRank'] and its engine the
sidecar's; the 2026-07-03 freeze date, the rechecked outputs and the
cost basis are asserted against committed records.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Fix the paper's engine, exclusion and cost wording; check the upgrade partition
paper/index.qmd:
- The abstract and engine paragraph date 2.15.17 as the newest release
when PolicyBench began sweeping the references, name 2.17.0 (the
publication check) and say each scored reference comes from 2.15.17;
the 55 excluded outputs keep their decided values (52 on 1.755.4, 3 on
2.15.17) and the 19 that move were re-reviewed.
- "The investigation found, and PolicyBench excluded, ..." replaces
"Reviewers excluded"; "five audit waves" becomes "each later audit".
- The cost sentence rendered a literal \\$0.002 because Quarto printed
the inline string's repr; it now goes through Markdown. The cost table
rendered \$0.034 and a DataFrame index column; it now uses md_table.
policybench/paper_results.py:
- partition_engine_upgrade_changes splits the upgrade's changed outputs
into scored changes, within-tolerance changes and new exclusions, and
raises unless the three add up to the changed list. The count
properties read it.
- engine_upgrade_date is the rebuild day (2026-09-29); the revision's own
date stays the builder's wave date.
- New accessors: excluded outputs by engine version, and the publication
check's policyengine-us release (read from the verification sweep).
tests/test_paper_results.py pins the partition (4 + 2 + 3 = 9), and a
Hypothesis property over synthetic revisions and exclusion sets checks
that the partition is exact and disjoint and respects the $1 and 0/1
boundaries. The judge-provenance test explains why the older judges'
counts fall.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Compute the app's serving counts and exclusion engines from config and payload
- app/src/lib/servingConfig.ts reads the bundled serving configuration:
the chunked rows (request shape other than whole scenario) and the
Claude rows on the JSON contract (each card that sets it says the API
rejects a forced tool call).
- Methodology states "<chunked> of the <board> models" from that config
and the payload's modelStats, says policyengine-us computes each scored
reference, and adds the excluded outputs' engines from the payload's
referenceExclusions (55: 52 on 1.755.4, 3 on 2.15.17), with the 19
re-reviewed outputs from app/src/lib/referenceEngine.ts.
- ModelLeaderboard builds the list of Claude rows that reject forced calls
from the config.
- app/tests render both components against data-summary.json and the
config instead of hand-typed numbers.
- tests/test_disclosures.py: the Methodology source computes its counts;
referenceEngine.ts's recheck count and engine equal the sidecar's; the
live version description matches the exclusion record; every "newest"
in public copy is anchored to a time; the manifest checklist names
household_dataset; and the rendered HTML and PDF print dollar amounts
without escapes or an index column.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* State how references are computed in the manifest; record the households' build
The build_manifest reproducibility note said the references "were
generated with policyengine.py X and policyengine-us Y against the
certified PolicyEngine US populace dataset (<runtime default build>)".
PolicyBench computes each reference with policyengine_us.Simulation from
the household's own listed inputs (ground_truth.py) and records
policyengine.py only for provenance; no dataset enters a reference.
The note now says so, names the households' build from the run's
scenarios.csv.meta.json, and says reference_output_refresh's dataset
fields describe the reference runtime's default dataset. A new
household_dataset block records that build (id, dataset, URI, sha256);
the existing keys stay. tests/test_snapshot_artifacts.py checks the block
against the scenarios meta, the note's wording, and docs/paper.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Date the adjudication record's judge verdicts from their sidecars
The upgrade's refresh gave each judge_previous item a judged_on equal to
the entry's adjudicated_on, which paired claude-opus-5-5 with dates before
Opus 5.5's release (2026-09-21), left judged_on_utc at the replaced
verdict's date, and dated the re-judge 2026-09-28, the local date of a run
whose verdicts carry 2026-09-29 UTC timestamps.
scripts/date_adds0928_judge_verdicts.py takes every date from the
sha256-bound verdict.meta.json sidecars: the stage's current verdict for
judge_rejudged_on and judged_on_utc, and the 20260922c audit tree's
verdict for each judge_previous item whose judge model and classes match
(otherwise it would rename the field adjudicated_on; all 54 match). Run on
the staged record: 54 re-judge dates 09-28 -> 09-29, 9 judged_on_utc
fixed (scenario_074 had Opus 5.5 on 2026-09-05), 51 previous Opus 5.5
dates -> 2026-09-23 and 2 previous Opus 5 dates -> 2026-09-05.
tests/test_adjudications.py checks that every judge date falls on or
after its model's release, that a re-judged entry is dated on or after
the verdict it replaced, that each current date is one the manifest's
judge provenance records, and names the two verdict-less llm_error
entries.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Let a post-freeze re-export gate on the 22c asset
The documented chain reruns finish_adds0928 --step export after a fix to
the staged annotations. After the freeze the live pointer and the
committed snapshot are this release's own, so resolve_base refused
("base pointer changed"). The export only needs the 22c payload to replay
the incumbents against.
resolve_live_base keeps the pre-freeze path (resolve_base) while the
pointer names the 22c base. When it names this release, it reads the 22c
run payload from BASE_COMMIT (3220a7a6), requires it to rewrap to the 22c
asset's sha256 (01e7e72b...) with 42 models, and refuses any other
pointer. The 22c replay gate itself is unchanged, and it passed on the
rerun: every incumbent reproduces the 22c modelStats on the 22c
references.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Say which release 2.17.0 agrees with, and introduce the adapter first
- The card, the paper and the upgrade README say 2.17.0 gives the same
value "as 2.15.17" for all 1,984 outputs. The card and the paper state
it after they introduce the Maryland adapter.
- Release note: the check sentence names 2.15.17; the Arizona paragraph
splits into three sentences, with the BBCE note's update adding the
household as a fifth held back by income.
- tests/test_disclosures.py rebuilds the card's engine sentences from the
sidecar, the exclusion record and the 2.17.0 sweep, and the cost table
test compares the whole header.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Refreeze release 20260929 with the dated adjudications and manifest note
Reran the documented chain on results/local/adds0928-v3: finish_adds0928
triage and export (the 22c replay gate passed), freeze_adds0928 --dry-run
and the freeze, then paper/render_paper.py and freeze_snapshot.py
--rendered-only.
- annotations/.../us_adjudications.json: the staged record with its judge
dates taken from the verdict sidecars (sha256 24176bf6... -> 0ea00ff9...).
- manifest.json: the new adjudication pin, the household_dataset block and
the rewritten reproducibility note; the rendered paper pins.
- The payload is unchanged: dashboard-data.json sha256
e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6, and
app/src/data.artifact.json still names it.
- Paper render: PDF sha256 e6404ea2...; the cost sentence and table print
plain dollar amounts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Unit-test the judge-date script on bound and unbound verdicts
A synthetic pair of cases: the current verdicts date the re-judge and
judged_on_utc (2026-09-29), a previous verdict whose classes match dates
judge_previous (2026-09-23), and one whose class differs leaves the
record's date under adjudicated_on. A second pass changes nothing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Rebuild the references with the upgrade dated 2026-09-29 and the engine rule anchored to the sweep
The records dated the engine upgrade 2026-09-28, the US Eastern date the
wave began, while the sweep began at 01:42 UTC on 2026-09-29 and the three
new exclusions are decided_on 2026-09-29. build_references_latest.py now
dates the revision, the derivation and the new exclusions' notes
2026-09-29, takes each new-exclusion basis date from its decided_on, and
states the rule with a time anchor: "References come from the newest
policyengine-us release when PolicyBench begins the reference sweep; at
publication PolicyBench checks that the newest release gives the same
values." Max's ruling of 2026-09-28 keeps its date, labeled as his.
The sidecar and exclusion record are rebuilt by the script from the
committed files (its docstring gives the command). The reference CSV is
byte-identical (sha256 e8bbba8f...); only record text and
regenerated_at_utc (15:04 UTC) change. scripts/install_adds0929_references.py
installs a rebuild into the snapshot and both stage copies only when that
holds, and keeps the stage receipt's pins.
verification/sweep_timing.json records the PyPI upload times (read live),
the sweep's install, script and first-output times, and the pin commit;
tests check the sweep began between the 2.15.17 and 2.16.0 uploads and the
2.17.0 check ran after its upload. The README states the rule the same
way, dates the reviews by time zone, lists the five sweeps re-run on
2.15.17, and names the rebuild.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Bind each recorded judge flag to the verdict its date names
The date script bound each judge_previous item to a 22c verdict by judge
and classes only, so eight items (005 state, 028, 030, 042 federal, 051,
064 federal, 109, 112) were dated 2026-09-23 from verdicts that do not
flag the reference, while still carrying judge_reference_suspect: true
from an earlier run of the 2026-09-22 wave. The script now compares the
flag: an item whose dated verdict does not raise it keeps it only with
judge_reference_suspect_source naming that wave (flagged_sept22_wave.json,
now committed), and a flag no run explains stops the script. A top-level
source is dropped where the current verdict raises the flag itself
(005 state, 112).
The record now states its date conventions: adjudicated_on names the audit
wave, whose decisions were written up to the day its release was
committed; judge dates are UTC days from the verdict sidecars. Entries
that were not re-judged get judged_on_utc. Where a later wave replaced the
verdict a decision reviewed without keeping it (ten 2026-09-05 decisions),
adjudicated_verdict records that verdict as its release published it.
The script writes the verdicts it read to verification/judge_verdicts.json.
Tests check every date and flag against that evidence, check that each
decision's reviewed verdict is dated by its wave's release (46 decisions of
the 2026-09-22 wave reviewed verdicts dated 2026-09-23, labeled as
intended), check the published verdicts against git, and property-test
date_entries (renaming rule, flag rule, idempotence) with Hypothesis.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Scope the card and paper to the sweeps each engine ran; fix the note's adapter and price wording
- Card and paper: the defect and reading sweeps ran on 1.755.4; on 2.15.17
PolicyBench re-ran five (r02, NIIT, the SALT refund, mortgage residence
and 40-hour readings), which move no scored output but scenario_066's
SNAP, whose prompt states 40 hours. The card now states the limit the
paper's Limitations gives, and the Limitations bullet names the re-run.
- Paper: "this snapshot's scored reference outputs come from the fixed
engine"; Claude Opus 5 judged "the 132 cases the September 5 additions
joined that no later judge re-judged" (card too); the abstract's engine
sentence is split and gives the fallback to the last published amount.
- Card: the transport list attributes each JSON row as its model card
records it (provider rejection for the Claude, DeepSeek V4.1 Flash, Kimi
and Qwen rows; card choice for DeepSeek V4 Pro and GLM-5.2; the Gemini
family default), and "before PolicyBench froze the references".
Methodology says the same without a row list. A test derives the groups
from the serving config and model cards.
- Release note: the 2.17.0 check ran "with those conventions and an adapter
that keeps Maryland county tax out of state income tax" (tested against
the check's fix module), and DeepSeek V4.1 Flash is priced "at DeepSeek's
peak list price". The note tests read the PyPI upload times from the
timing record.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Derive the leaderboard's re-run counts; property-test the serving and exclusion-engine helpers
- ModelLeaderboard states how many Claude rows ran without extended
thinking and how many carry an auto re-run from SERVING_SENSITIVITY
(servingSensitivityCounts), and the test ties those counts to the
sensitivity data files.
- fast-check (new dev dependency) properties for excludedOutputsByEngine
(counts sum to the exclusions, distinct versions in numeric order,
input order irrelevant) and its sentence, for chunkedServingModels and
jsonContractClaudeModels, and for the re-run counts.
- Differential checks against the Python side: the payload's
excluded-by-engine counts equal the ones the rendered paper states
(paper_results over reference_exclusions.json), and the serving helpers
agree with the paper's serving-configuration table, replacing the test
that repeated the implementation's filter.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Refreeze release 20260929 with the redated records; re-render and re-pin the paper
Reran the documented chain on the stage (triage, export with the 22c
replay gate, freeze dry-run, freeze), then rendered the paper and re-pinned
it (freeze_snapshot.py --rendered-only).
The payload changes only through annotation text: the three exclusion
notes and 127 row annotations now say "Found in the 2026-09-29 engine
upgrade" (130 string fields; no score, value or key moves).
- dashboard-data-20260929 payload sha256
e7d5e056b53c0d6d406bb3afb389aeaf80ebb611932ce72c8b3c3ceaef5ddad6 ->
d146473d9bd7776638c59a0a20774dbe9026d8bcee0f2201e114146609ddf246
(app/src/data.artifact.json and the manifest)
- run data.json.gz 3028225189ae... -> 2a6463c9ff9f...
- us_adjudications.json 0ea00ff9480c... -> 0e58570a62aa...
- us_case_notes.csv 9a4a3e374b49... -> 864cab96e670...
- reference_exclusions.json 66ac8670fcee... -> ae28ade59705...
- reference_outputs.csv.meta.json c6f1589da553... -> 5469664726ad...
- reference_outputs.csv unchanged (e8bbba8fd3e9...)
- the BBCE note data files' run_payload_sha256 follows the payload; their
rows are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Scope the reference installer's text mask to the engine upgrade
install_adds0929_references.py blanked the field names date, rule, basis,
note and derivation at every depth, so it would have accepted a rebuild
that redated an older revision or rewrote an older exclusion's note. The
guard now blanks only the engine_upgrade revision's date, rule and each
change's basis, the sidecar's regenerated_at_utc, the exclusion record's
derivation, and the notes of the exclusions the upgrade added (the
outputs its changed list names that the record lists). Every older
revision and exclusion must be unchanged.
The historical 15:04 UTC rebuild still passes the narrower guard. New
tests refuse a rebuild that changes an older revision's date, rule or
basis, an older exclusion's note, or a new exclusion's decided_on, and
check that the committed records pass their own guard with exactly the
three new exclusions masked.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Require the wave's flag record for a kept top-level judge flag
date_entries accepted any existing judge_reference_suspect_source on a
top-level flag the current verdict does not raise, without checking that
an earlier run of the 2026-09-22 wave flagged the case; only the
judge_previous branch checked flagged_sept22_wave.json. The top-level
branch now requires the case in the wave's flags too, and drops a stale
source wherever the recorded flag matches the verdict's.
The Hypothesis strategy now draws the top-level flag, the current
verdict's flag and the source for each entry. The property checks that a
mismatch stops the script unless the flag is set, sourced and in the
wave, that the flag itself never changes, and that a source survives
exactly where it explains a mismatch. Removing the wave check makes the
property fail. Rerun on the stage, the script changes no date or flag.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Date this wave's adjudications by when they were written
judge_verdicts.json and the adjudication record's date_conventions said
the 2026-09-29 wave's decisions were written up to the day its release
was committed, 2026-09-29, but that release has no commit yet. The wave's
entry now gives adjudications_written_on 2026-09-29, with commit and
pull_request left null for the lead to fill after the merge. The date
conventions state the earlier waves' release commit days (2026-09-05 and
2026-09-23) and that the 2026-09-29 wave's decisions were written on
2026-09-29 UTC, after its reference sweep began, and bound a verdict's
lateness by the last day its wave's decisions were written.
The date script rewrote the staged record's date_conventions (its only
change) and the evidence file. The test reads each wave's day from
committed_on or, for an uncommitted release, adjudications_written_on,
checks the record states the same days, and checks the sweep began that
UTC day. The refreeze copies the record into annotations/.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Name the conventions, adapter and hours alias in the manifest's reference note
The manifest's reproducibility note said each scored reference comes from
policyengine_us.Simulation on policyengine-us 2.15.17 and the household's
own inputs. The references also depend on the nine publication
conventions, the Maryland output-scope adapter and the scenario builder's
stated-hours alias. The note now names them "as the reference sidecar's
engine_upgrade revision pins them (fix_modules, builder)", and
docs/paper.md says the same.
build_manifest reads the convention count from the sidecar (the
latest_c_*.py modules) and stops the freeze if the revision's other
modules or its builder note are not the ones the sentence describes. The
manifest test rebuilds the sentence from the sidecar. The refreeze
regenerates the manifest.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Scope the release's claims to what was checked; record the 2.17.1 check
- The abstract names the exclusions the audits found (28 unfixed engine
defects, 27 unlisted inputs), in place of every output an unrun fix
could move.
- On 2.15.17, PolicyBench re-ran four September 22 sweeps and ran one new
sweep, for the state and local tax refund reading. None moves a scored
output by more than the $1 tolerance, and three move by less.
verification/rerun_sweeps.json records what each sweep moves; a test
recomputes the claim.
- The 2.17.0 check is dated by when PolicyBench read PyPI (14:58 UTC). The
paper takes the upload and check times from verification/sweep_timing.json.
- A second pre-publication check found policyengine-us 2.17.1 (uploaded
17:24 UTC). It also gives the same value for all 1,984 outputs
(verification/latest_final_2171.csv, sweep_timing.json later_checks).
- JSON transport wording in Methodology, the paper and the sensitivity note
follows the card: the card or its family default selects JSON, mostly
because the provider rejects a forced tool call.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Refreeze with the reference note naming conventions, adapter and alias; re-render
The manifest's reproducibility note now names the publication conventions,
the Maryland output-scope adapter and the stated-hours alias. The paper is
re-rendered and re-pinned. Payload unchanged: sha256
d146473d9bd7776638c59a0a20774dbe9026d8bcee0f2201e114146609ddf246.
The abstract test reads the PDF with its page-number lines at page breaks
removed. The abstract sentence crosses from page 1 to page 2, and
pdftotext put the "2" mid-sentence.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Fetch full history in the CI test job; name the missing base commit
The release gates read the September 22c base files from git at
BASE_COMMIT. CI's default shallow checkout lacks that commit, so four
tests failed with a bare CalledProcessError. The test job now fetches
full history, and base_commit_blob says which commit is missing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Register GPT-6.1 Sol with its onboarded model card
OpenAI announced GPT-6.1 Sol at DevDay on 2026-09-29. The model page
prices gpt-6.1-sol at $2 / $10 per 1M input/output with $0.10 cached
input, and gives medium as its default reasoning effort; the Models API
lists the id with created 2026-09-27T23:47Z.
The onboarding gauntlet passed the forced tool contract 3/3 (735
completion tokens) and 16/16 whole-scenario (1,626 tokens) on the
Responses API, so the card matches GPT-6 Sol's treatment. The
frozen-roster paper test fails until the board is refrozen with the new
row.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Brief the GPT-6.1 Sol additions driver
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add the GPT-6.1 Sol additions driver
finish_gpt61sol.py folds the gpt61sol supervised run onto the 45-model board
of dashboard-data-20260929 with the same prepare/judge/triage/export steps
and --early/--partial rehearsal as finish_adds0928.py, minus the reference
revision: the five reference files are pinned and checked in prepare and
again in export, and all 45 incumbent modelStats must serialize byte for
byte as released (Fable 5's batch usage is carried over first, as the
exporter cannot recompute it). A re-export after the freeze reads the base
payload from BASE_COMMIT (PR #182's head until the lead repoints it).
The audit is seeded from the 20260929 stage with the grounding its prompts
were rendered from (pinned by sha256). Only cases GPT-6.1 Sol joins or opens
may change; any incumbent-only prompt change or vanished seed case is
refused, and a sidecar's prompt_sha256 must match the current prompt.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add the GPT-6.1 Sol freeze and rehearsal copy
freeze_gpt61sol.py adapts freeze_adds0928.py to the 46-model board: GPT-6.1
Sol is the only new model, the treatment check is reused unchanged, and
--tag defaults to finish_gpt61sol.RELEASE_TAG. The reference-revision pin
catch-up is gone; instead the staged, committed and manifest-pinned
references must all equal release 20260929's bytes (and the frozen copies
again after the freezer runs). The export receipt must bind the references,
the adjudications and the run state. Adjudications may add decisions or
restate a re-judged class but may not drop one or move an exclusion.
snapshot_gpt61sol.py runs the existing CSV-only rehearsal copier over the
new model without mutating snapshot_adds0928's module state.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Document the GPT-6.1 Sol release procedure
docs/gpt61sol/design.md states what the release is, each gate with the
tests that pin it, where the audit seed and its grounding come from (every
seed prompt re-renders byte-identically from the committed snapshot and
unified_audit/grounding.csv), and the exact command sequence from prepare
through freeze_snapshot.py --rendered-only with the real paths. A test keeps
RELEASE_TAG the only spelling of the new tag in the driver's files.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Brief the repin to release 20260929 as merged and the prepare step
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Repin the GPT-6.1 Sol driver to release 20260929 as merged (d616e67c)
#182 changed after review and merged as d616e67c. Every base value is
recomputed from the files committed there; the payload and predictions also
match the published release assets' digests.
- BASE_COMMIT: f7ced3b3 (#182 head) -> d616e67c, the merge on main.
- BASE_SHA256: d146473d... -> a5cb9989... (123,362,946 bytes; the committed
data.json.gz rewraps to it).
- reference_outputs.csv.meta.json: 54696647... -> 816fef53... (reasons
rewritten); reference_exclusions.json: ae28ade5... -> bf4e6a24...
(scenario_023 head_medicaid_eligible excluded). The CSV and scenarios pins
are unchanged.
- BASE_EXCLUSIONS 55 -> 56, with BASE_OUTPUTS = 1984 and
BASE_SCORED = 1928. resolve_base now checks the scored count, and export
refuses an addition not scored on exactly those 1,928 outputs.
- GROUNDING_SHA256 is unchanged (b1e4a9bc...), re-checked.
After the merge, all 674 seed prompts, all 986 case ids and cases.jsonl
re-render byte-identically from the committed snapshot. That includes the
023 Medicaid case (339d113a..., its verdict's bound prompt_sha256): its
adjudication, row annotations and case note changed, but none enters a
prompt. A slow test keeps that check, and two committed-data tests pin
what triage relies on: the adjudications exclude exactly the 56 scoring
exclusions and keep every seed verdict's class.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Drop the export n gate, which cannot fire; mark the driver brief's base superseded
An independent review showed that modelStats n is the scored-reference grid
crossed with every model (analysis.py _expected_prediction_grid and
compute_metrics). So n is the same for every model, whatever it answered.
The no-drift gate, which runs first, already pins every incumbent's n at
1,928, so an addition-only n check could never refuse anything, and its
design-doc row credited it with a guarantee it did not provide. The
1,928 count stays checked in resolve_base and pinned by the committed-data
test.
brief_driver.md is the historical instruction and its base values are #182's
head. A note now points readers to the merged release's pins in design.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Brief the GPT-6.1 Sol judge/triage step and the driver review
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Restate the adjudications GPT-6.1 Sol's re-judge replaced
prepare re-judges every case GPT-6.1 Sol answers wrong; 54 of the 134
carry a developer adjudication. Triage requires each record to name the
case's current verdict, and the record's date_conventions date it from
the verdict's sidecar. scripts/restate_gpt61sol_adjudications.py restates
those entries the way release 20260929 restated its own re-judges:
- the replaced seed verdict is appended to judge_previous, dated from its
sha256-bound sidecar, and the stage's Opus 5.5 verdict goes on top,
dated by judge_rejudged_on (and judged_on_utc where present);
- the 2026-09-22 wave-flag rule the committed record follows is kept at
every level, and an existing flag source keeps its wording;
- no decision field changes value or place; judge_previous becomes its
earlier items plus the seed verdict; a new flag no reference verdict
answers stops the script.
It refuses a seed that is not the stage's (carried-over verdicts must
match byte for byte), a stage verdict older than the seed, and an entry
whose dates disagree. It writes only after the restated record loads,
passes triage's verbatim-verdict check and keeps the scoring exclusions,
and it is idempotent.
Invariants, checked by example and Hypothesis tests: decisions unchanged
in value and order; top = latest stage verdict under the wave-flag rule;
judge_previous = earlier items + seed verdict; dates follow the sidecars;
idempotent; restating after each re-judge equals restating once after
the last (up to the wording of a source an intermediate flag cleared).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Restate an entry that still names its seed verdict, even when an earlier item reads alike
restate_gpt61sol_adjudications.py decided "already restated" from
judge_previous[-1] == the seed item alone. An entry never restated whose
last judge_previous item happened to read exactly like the seed verdict
(same judge, classes, flag, flag source and UTC day, e.g. an earlier
same-day same-class re-judge) took that branch: the earlier item was
dropped and the seed never appended, breaking the stated invariant that
judge_previous becomes its earlier items plus the seed verdict.
- An entry whose top level names the seed verdict is now restated for the
first time; the judge_previous[-1] test is only the fallback.
- The script never writes an entry whose top level names the seed: a
re-judge that repeats the seed verdict on the seed's own UTC day stops
it, since content alone could not tell that entry from one never
restated.
- The property test now draws an earlier item that echoes the seed
verdict (the pre-fix code fails it and Hypothesis shrinks to echo=True)
and expects the new refusal exactly where the script refuses (checked
exhaustively by the reviewer: 20,736 cases, 0 mismatches); two example
tests pin both behaviours.
No effect on the GPT-6.1 Sol stage: restating the committed record with
the fixed code reproduces the staged record byte for byte (sha256
8e068b74...f46a), a rerun restates nothing, and triage's outputs are
unchanged. Independent review: APPROVE.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Brief the fixes for the driver review and the judge report
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Check a restated adjudication record in memory before writing anything
restate_gpt61sol_adjudications.py wrote the restated record to a
.restating file beside the staged record and only then loaded and checked
it (review finding 7). It now serializes the record, parses that exact
text in memory with policybench.adjudications.parse_adjudications (the
validation load_adjudications runs, split out so a record can be checked
before it exists on disk), checks triage's verbatim-verdict rule and the
scoring exclusions, and only then writes the record atomically and its
evidence file.
The new test spies on the verbatim check and fails on the old code: a
.restating file was already in the stage when the check ran.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Refuse a GPT-6.1 Sol export receipt that names another base
The freeze's verify_receipt checked the payload hash, the model count,
partial and the release tag, but not the base_tag and base_sha256 that
export writes (review finding 5). A receipt from an export against any
other base passed. It now requires base_tag == BASE_TAG and base_sha256
== BASE_SHA256; two new cases of the receipt test fail on the old code.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Let a staged adjudication record change only where GPT-6.1 Sol re-opened a case
Review finding 1: the freeze compared only the staged record's case keys
and exclusions with the committed record, and triage compared nothing, so
a staged record could rewrite any decision (class, reasoning, reference
verdict) of a case GPT-6.1 Sol never joined, and apply_adjudications would
publish it on every incumbent row. Finding 4: the freeze's baseline was the
working-tree record, which the freeze itself overwrites.
- verify_adjudication_changes (driver) is the one gate, used by triage and
the freeze. Against release 20260929's record read from git at
BASE_COMMIT (base_adjudications), every committed entry keeps every
field byte for byte, key order and entry order included. A case in
prompt-changes.json's changed or added lists may rewrite only its judge
fields (JUDGE_FIELDS from the restate script, the one definition) and
its reasoning exactly as a listed wording amendment says. A new entry
may decide only a re-opened case; none may be dropped.
- Wording amendments (item 9 of the brief): <stage>/wording-amendments.json
lists case id, field, old text, new text and reason. Only three fields
qualify: an entry's reasoning, the case note and one model's row
annotation (which names the model). The case must be re-opened, the old
text must occur exactly once, and nothing that carries a class, an
exclusion or a score can be named. Triage applies the reasoning
amendments to the staged record (checked in memory first, idempotent)
and the case-note and row amendments to the annotations, then re-checks
that every case note still carries its exact adjudication sentence.
- Export binds prompt-changes.json, the amendments and stage.json in
release-ready.json; the freeze refuses a receipt that does not bind them
(the amendments whenever the stage has them), checks the record with the
same gate against git, checks each case-note and row amendment is in the
staged CSVs, and commits the amendment list beside the record.
test_triage_may_restate_a_rejudged_class, which pinned the permissive
behaviour, is replaced by test_a_rejudged_case_may_be_restated and
test_a_rewrite_of_an_incumbent_only_case_is_refused (judge class, decision
class, reasoning, key order, entry order). The baseline test simulates a
freeze that stopped after overwriting the working-tree record; it failed
on the old code. On the real stage, the old triage and freeze both
accepted a rewritten decision of us__scenario_007__head_medicare_eligible
(results/local/gpt61sol-v1-ops/a4/verify_findings.py).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Bind every judged verdict to its own bytes and carried-over ones to the seed
Review finding 2: validate_verdicts checked verdict_sha256 only for
verdicts naming GPT-6.1 Sol, and a prompt binding only where a sidecar
happened to record one, so an edited carried-over verdict still passed
(on the real stage, an edited us__scenario_000__head_medicaid_eligible
verdict validated). Finding 6: the tests pinning the seed skipped off this
machine.
- Every judged verdict's sidecar must now carry the verdict's own sha256.
- A case whose prompt is the seed's carries its seed verdict over and must
keep the seed's verdict bytes; a sidecar prompt_sha256, where the seed
recorded one, must match. Any failure there is refused outright (a
re-judge cannot restore it) and nothing is set aside.
- Every other verdict is new and must record the sha256 of the prompt it
judged; that covers every verdict naming GPT-6.1 Sol (finding 3's gate).
- The seed is docs/gpt61sol/seed_digest.csv: case id, prompt sha256 and
verdict sha256 for each of the 20260929 audit's 674 judged cases, its
bytes pinned by SEED_DIGEST_SHA256. prepare refuses any other seed and
binds it in stage.json; judge and triage read the binding and re-check
it against the pin. --step bind-seed binds it for a stage prepared
before prepare did (this one), after checking the stage case by case:
kept cases keep the seed's prompt and verdict bytes, changed prompts
differ, and no seed case vanished.
- Invalid or hedged verdicts are moved to <stage>/rejected-verdicts/
instead of deleted, so stricter validation never destroys evidence.
Not done as the review worded it: 453 of the seed's 674 sidecars (the
Codex, Opus 5 and Workflow runners') never recorded prompt_sha256, so a
sidecar prompt hash cannot be required of every case without rewriting
the seed's provenance. The seed binding carries the same guarantee for
them: the committed digest records the prompt beside each verdict.
Runs anywhere: the committed digest hashes to its pin, lists 674 cases
once each, records 023 Medicaid's 339d113a prompt, and every committed
adjudication decides a case it lists. Local only: the digest equals the
real seed. All the new tests fail or error on the previous commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Isolate each Claude judge and bill only the lane; bind new verdicts to prompts
Brief item 8 and review finding 3. Of the 134 GPT-6.1 Sol re-judges, 17
used file tools (9 read developer decision records), because the runner
left Read, Grep and Glob on and ran each judge from inside the worktree;
22 kept verdicts billed the desktop login, because a keychain-token lane's
token never reached the child calls. No sidecar recorded the prompt.
scripts/run_audit_claude.sh now:
- runs each judge from a fresh empty mktemp directory, refused if it sits
inside any git repository, and removes it afterwards;
- removes every built-in tool (--tools "", checked on CLI 2.1.284: the
structured answer still arrives and the judge reports no tools), denies
the file, search, web and shell tools by name as well, and keeps
--strict-mcp-config, --disable-slash-commands and safe mode;
- copies each session transcript beside its verdict and rejects a verdict
whose transcript shows any tool call but the structured answer, or that
has no single transcript;
- refuses to start unless CLAUDE_CONFIG_DIR names a directory that is not
the desktop login's ($HOME/.claude) and `claude auth status` reports a
login there, and unsets ANTHROPIC_API_KEY and ANTHROPIC_AUTH_TOKEN. A
token lane points it at an empty directory and passes its token; with no
token such a directory logs nothing in (checked), so nothing falls back;
- records in the sidecar the sha256 of the prompt it piped, the login the
CLI reports (method, email and org where reported; a token login reports
no email), AUDIT_ACCOUNT as the declared account, and the isolation;
- takes AUDIT_ONLY to judge named cases only.
tests/test_run_audit_claude.py drives the runner with a fake CLI that
records every call's argv, working directory and environment: 14 of its
15 tests fail on the old runner.
scripts/stamp_gpt61sol_prompt_bindings.py stamps prompt_sha256 on stage
verdicts judged before the runner recorded it, but only where the judge's
own session transcript shows prompt.md's exact text as its prompt; one
missing, ambiguous or different transcript stops it with nothing written.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Find already-valid audit verdicts in one pass, not one interpreter per case
run_audit_claude.sh checked each case's verdict with its own Python
process, before judging and again for the final tally: 674 starts each
time on this audit, several minutes, which pushed a five-case re-judge
past ten minutes and made a validation-only `finish_gpt61sol.py --step
judge` take as long. valid_cases now validates every verdict in one
process with validate_verdict.verdict_errors; the loop skips those cases
and the tally counts them. A validation-only judge step on the GPT-6.1
Sol stage now takes 16 seconds. The runner tests pass unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Document the stricter GPT-6.1 Sol gates, the isolated judges and local-only checks
design.md now states what each new gate enforces and which test pins it:
every verdict bound to its bytes and carried-over ones to the committed
seed digest; judges isolated from files and billed only to the lane; the
adjudication record checked against release 20260929 in git, changing
only re-opened cases' judge fields and listed reasoning amendments; the
wording-amendment list and its limits; and the receipt's base and new
bindings. It lists the four checks that can run only on this machine and
what stands in for them elsewhere, and the bind-seed, judge-credential,
restate and amendment steps in the commands.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Close the review's bypasses of the verdict and adjudication gates
An independent review of the fixes pushed two bypasses through triage,
export and the freeze dry run on a scratch copy of the stage:
- a kept case whose prompt.md, verdict and sidecar were all edited passed
as a new verdict, since "carried over" was decided by the prompt on disk;
- a case moved from kept to changed in prompt-changes.json could then
rewrite its decision's judge fields and take a reasoning amendment,
because nothing re-checked the file after prepare.
Now:
- load_seed, which judge, triage and rejudged_cases (so the freeze and the
stamp script) all go through, re-derives kept, changed and added from the
stage's prompts and the bound seed with prepare's own check_prompt_changes
and refuses unless prompt-changes.json says the same. A kept prompt that
drifts, or a case moved between the lists, stops every step.
- The staged record's bytes must be exactly its parsed content in the
committed form, so duplicate keys cannot hide text the freeze would
commit, and its note, schema and date conventions must stay 20260929's.
- A re-opened entry whose judge fields changed must name the case's current
bound Opus 5.5 verdict as its judge, be dated by that sidecar's UTC day
and keep 20260929's judge_previous with exactly one item appended, as the
restate script writes it (triage and the freeze).
- The freeze reads 20260929's predictions and serving configuration from
git, like the adjudication record, not from the working tree it rewrites.
- set_aside keeps the judge's envelope, log and transcript with the verdict.
- A malformed amendments file and a re-opened case without a sidecar in the
stamp script are refusals with their own messages.
- A Hypothesis test drives single mutations (value, deletion, added key,
added judge key, swapped decision keys) of every committed entry, re-opened
or not: a judge-field change passes only on a re-opened case.
Each bypass has a test that fails on the previous commit. The real stage
passes triage under the new gates.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Give each Claude judge an allowlisted environment and accept only a lane login
An independent review of the hardened runner found that a missing lane
token could still reach the desktop login: the CLI names its keychain entry
from CLAUDE_SECURESTORAGE_CONFIG_DIR when set, and an empty value selects
the desktop's entry (confirmed with `claude auth status`, no model call).
Any login method also passed the check, including Bedrock, Vertex, an API
key helper and a settings API key.
- Every claude call (the login check, --version and each judge) now runs
under env -i with an allowlist: PATH, HOME, user, locale, temp dir, the
lane's config dir, its token and the effort level, plus safe mode. No API
key, base URL, provider switch or keychain override reaches it.
- The login check runs the same way from an empty directory and requires a
first-party login, by the lane's token when one is set, or a claude.ai
home login that is not the desktop's account; a token login reports no
email, so AUDIT_ACCOUNT must name it.
- The desktop directory is compared by file identity, under $HOME and the
account's own home, so a symlink, a case variant or a HOME override is
refused.
- A transcript is read strictly: any part whose type ends in tool_use but
the structured answer, a context attachment outside the listed kinds, an
environment inside a git repository, an unreadable line, or other than one
transcript rejects the verdict.
- AUDIT_PARALLEL must be a positive integer; AUDIT_ONLY is never
glob-expanded and names that match no case are reported; the effort level
is recorded in the sidecar. The header no longer claims safe mode drops
all settings.
The fake CLI now reads its settings from a file beside it (the runner
strips the environment) and records every call's environment. New tests
cover each refusal and rejection above, a schema-invalid verdict that must
be judged again, and the stale evidence it replaces.
Co-Aut…
* Recompute SNAP pathways and BBCE household rows on release 20260930
scripts/snap_pathways_20260930.py runs the references' configuration
(policyengine-us 2.15.17 with latest_final) on the 100 frozen households
and records each month's SNAP and TANF non-cash tests, allotment parts and
pathway, the parameters and engine formulas the BBCE note cites, and the
Michigan premium reading. It reproduces every scored SNAP reference.
scripts/bbce_households_20260930.py derives the five households held back
by income (Arizona's from March) and the four held back by savings from
that recomputation and writes every model's answer on the 46-model board.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Record the TANF non-cash tests' formulas in the 20260930 pathway meta
The BBCE note says PolicyEngine applies BBCE to any household eligible for the
TANF-funded non-cash benefit without checking receipt. The meta now records,
from policyengine-us 2.15.17, the formulas of the three tests that eligibility
combines and the gross limit and elderly-or-disabled status they read, so the
note test can check that none reads receipt. The CSV is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Republish the BBCE note as of October 5 on release 20260930
The note on SNAP households that qualify only through broad-based
categorical eligibility moves from 2026-09-23-five-snap-households-bbce to
2026-10-05-five-snap-households-bbce, dated October 5, on release
dashboard-data-20260930 (46 models). It counts five households held back by
income (Arizona's from March, after Arizona raised its BBCE limit to 200% of
poverty) and four held back by savings, drops the dated revision and update
paragraphs, and states its corrections of the September 3 note in one closing
paragraph.
test_bbce_households_note_facts recomputes every fact from the frozen
snapshot and the 2.15.17 pathway recomputation and pins every sentence beside
its evidence, including the engine formulas behind the BBCE mechanism claim.
The September 3 note, the reference audit note and the September 29 release
note link the new slug; the September 3 note's closing paragraph and the
September 29 note's two sentences about the BBCE note describe the October 5
note. The 20260929 rows keep a regeneration test.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Redirect the September 23 BBCE URL to the October 5 note
next.config.ts redirects /notes/2026-09-23-five-snap-households-bbce
permanently (308) to /notes/2026-10-05-five-snap-households-bbce, so links to
the note's first URL keep working. A new app test checks the redirect, that
its source is no note's slug, that its destination is a note, and that no note
links a redirected URL. The notes test pins the October 5 note's text and the
September 3 note's pointer to it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Check that the BBCE note's repository links resolve to committed files
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* State the BBCE note's minimum under law and under the references
Apply review findings to the October 5 BBCE note:
- Paragraph 0 states the minimum as $24 through September 2026 and $25
from October, and says the references follow law published before the
July 3 freeze and keep $24 through December (c_snap_hold_fy2026). The
test checks the hold rule's dates, the 2.15.17 hold module's nodes and
window, the upgrade sweep's raw 2.15.17 values ($3 above each
reference), and, on the pinned engine, the minimum's parameters.
- "$0 means the household does not qualify" replaces "treats the household
as ineligible", which two Gemini explanations contradict.
- The explanation counts split into disjoint groups (6 + 5 + 2 = 13), with
netLimitOnly replacing incomeBbceZeroWithNetLimit and
overLimitWithNetLimit.
- Three models, not two, give Arizona's pre-March 185% limit; the test
checks the set is complete.
- The correction paragraph names PolicyEngine's child support fix and
Michigan's rule, says four of the September 3 note's six are here, and
renames previous* facts to sept3*.
- Serial commas throughout; "match the reference".
- The note test checks links unconditionally and recomputes facts only
while the note's release is frozen.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tighten the September 3 and September 29 notes' pointers to the BBCE note
- The September 3 note says PolicyBench scores its six households as it
does "from release dashboard-data-20260922c on", so the sentence after
the October 5 pointer does not read as that note's release. Its test
checks every later reference revision and only the six's exclusions,
so a further revision does not break it.
- The September 29 release note's pointer drops the dangling possessive.
- The 20260930 pathway test's docstring says the repository pins
policyengine-us 2.15.17.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tighten the BBCE note's corrections, freeze date and explanation counts
- Name July 3 as the law cutoff, not the date of the references, which
the engine upgrade rebuilt on September 29.
- Tie the September 3 receipt correction to how PolicyEngine computed
that note's amounts, with evidence from the 1.755.4 recomputation, and
name the four states the asset-test sentence corrects.
- Say only the September 23 version counted the Michigan worker, checked
against every later version in git history.
- Say PolicyEngine rounds the minimum too, read from 2.15.17's formula.
- Count the 16 BBCE explanations above $0: six match, seven use the $23
minimum in force through September 2025 and answer $276.
- Drop repeated household-size and Arizona-month clauses, the net-income
aside, and the azMonths fact; reword the $0 sentence so it cannot read
as saying the households do not qualify.
- Fix the September 3 note's closing referent ("PolicyBench no longer
scores a Texas household's SNAP amount").
- Pin the September 29 release note's link to the BBCE note, and quote
7 CFR 273.10(e)(2)(ii)(C) with its initial-month qualifier.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Describe the 20260929 BBCE files' current readers
The September 29 update these files served is gone with the old note;
the September 29 release note's facts count their households and the
October 5 note reads the 20260930 files.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Finish the October 5 BBCE note: split the opening, correct the farm-rent claim
- Split the opening paragraph so the page description and social card stop after the
household count (558 characters, down from 1,017). The references and freeze sentences
move into their own paragraph.
- Name all four models whose $0 explanations say the Arizona resident's income exceeds
the state's limit (GPT-6 Astra, GPT-6.1 Sol, Gemini 3.5 Flash, Gemini 3.8 Flash), and
separate them from the three Gemini models that give the pre-March 185% limit.
- Correct the September 3 note's claim that SNAP excludes the second Texas household's
farm rent. SNAP counts rent as unearned income under 7 CFR 273.9(b)(2), or as
self-employment income when the landlord manages the property 20 or more hours a week
(policyengine-us #9671).
- Put the Michigan child support correction in the present tense, in this note and in
the September 3 note's closing paragraph.
- Clarify "that note's amounts"; reword the $0 sentence; count Arizona "among five held
back by income" in the September 29 note.
- Rename previousMinimum and previousMinimumAnnual to fy2025Minimum and
fy2025MinimumAnnual.
Every changed sentence is re-pinned beside its evidence in tests/test_notes.py and
app/tests/notes.test.ts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Apply the final review: farm-rent law, the zero-hours household, two referents
- Farm rent: cite 7 CFR 273.9(b)(2)(ii) (unearned) and 273.9(b)(1)(ii) (earned when a household member manages the property 20 or more hours a week), net of business costs. Say that PolicyEngine counts farm rent only from policyengine-us #9671, which postdates the references' engine.
- Restore the old note's point that at 0 hours, the prompt's rule for unlisted numbers, PolicyEngine gives the second Texas household $0, the September 3 note's models' answer. Pinned to the exclusion record's alternative value and the three models' predictions.
- Michigan: restore the cause (PolicyEngine subtracted the child support from gross income until policyengine-us #9586).
- Arizona: name Gemini 3.5 Flash and Gemini 3.8 Flash instead of "the two Gemini models".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TDD: the 16 new cases fail against the checkpoint; its 76 cases pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cover d963 SALT payments, d974 Medicare premiums, d972 employer choices, premium payer labels, dependent tax-return scope and loaded labels. Render all 100 fixtures and pin the 2.1.0 source and v2-definition identity. Validation: 276 relevant tests pass; repository lint and format checks pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ds (#165) Reuse the mergeable sweep without activating its CLI or changing v1. Replay all 1,984 recorded baselines and 70 perturbations over 100 fixtures. The 16 report tests pass; explicit unknowns override named conventions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep proposed definitions in v2 instead of the byte-hashed published spec. State sales-tax conventions and before-response hours, clarify prior spouses, validate payment/choice/date inputs, and cite the audited IRMAA lagged reader. Require household and affected-person fact markers in the report; pin all legacy residual and compound counts without gating on legacy gaps. Regressions were observed failing before fixes; 296 tests, lint, and format checks passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…165) Resolve the approved review's minor follow-ups: sales-table year/version scope, dependent local unknowns, explicit before-response hours unknowns, IRMAA assumption wording, and required release-sweep registry additions. Prune ineffective registry entries and document the one-tax-unit boundary. TDD: eight cases fail first; all 306 final cases, lint and format now pass. Refresh the source identity and golden after rendering all 100 fixtures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Extends the unactivated PolicyBench v2 prompt contract for issue #165 with the October 5 decisions d963, d972, and d974. These statements make the audited reference assumptions explicit; publishing or activating v2 remains Max's decision. This PR stays a draft.
The contract now states:
state_withheld_income_tax, or the named withholding-estimate convention; IRS optional sales-table and local-sales proxy conventions are explicit too.The proposed output wording is isolated in the v2 module.
benchmark_specs.jsonis byte-for-byte unchanged, preserving published v1 wording, its raw spec hash, payload metadata, and resume behavior. No published prompt, reference, exclusion, board, evaluator path, or activation changes.Contract version is 2.1.0. The identity hashes source and canonical v2 definitions; the golden for scenario 074 is updated and all 100 frozen fixtures render without mutation. Existing program-specific disability, integer SSDI duration, take-up, provenance, and premium requirements remain.
The required-facts test reuses byte-exact comparison/report logic from mergeable #196 without integrating its CLI into v1. It replays actual recorded evidence over 100 fixtures / 1,984 outputs / 70 moving rows; it reports legacy gaps rather than requiring zero. Coverage requires the exact stated marker in the affected household and person, respecting unknowns and unsupported values.
Remaining individually moving inputs are
county,meets_ssi_disability_criteria,months_receiving_social_security_disability,first_home_mortgage_origination_year, andsecond_home_mortgage_origination_year: 12 distinct outputs, all already excluded; zero scored residuals. All four original compound readings require a fresh v2 sweep even though their constituent inputs are now covered. Before activation, adapt the reference builder and rerun the complete registered-input sweep, including #196's later Part B not-enrolled reading, employer-choice support, and registry additions forstate_sales_tax,mt_withheld_income_tax, andmedicare_irmaa_magi_two_years_prior. Recorded replay is not a fresh simulation or an exhaustive input-discovery proof.Validation: 306 tests passed (113 v2 contract, 27 report, 159 existing evaluator including 20 prompt tests, and 7 spec tests) in 61.58 seconds. Repository Ruff lint and format checks pass (preserved diagnostic scripts and task-local Git metadata excluded). Tests were written and observed failing before implementation. Golden identity/text and the published v1 spec fingerprint are pinned. All four hosted checks (app, lint, paper, test) passed on final commit
41061fca; the full hosted Python suite passed 1,692 tests.Independent review: approve, with no remaining defects, in the final independent
subfleet run --task review --tier standardreview20261006-030320-pb-v2-oct5-review-r3(Claude Opus 5.5). It confirmed the minor follow-ups from the preceding approval were resolved. The initial review's v1 spec-hash issue and coverage/validation findings were addressed; the disputed IRMAA reader was verified against installed policyengine-us 2.15.17 and cited with immutable source links.axiom: n/a — benchmark prompt contract and audit reporting only; no country-model policy encoding changes.
Related to #165. The remaining data/provenance and activation gates are unresolved; this draft does not claim completion of those gates. Do not merge or activate as part of this task.