Skip to content

Correct scenario_031's Medicaid annotations to the engine's California disregard - #197

Closed
MaxGhenis wants to merge 2 commits into
mainfrom
part-b-annotation-031-20261005
Closed

MaxGhenis wants to merge 2 commits into
mainfrom
part-b-annotation-031-20261005

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Summary

The case note, reference explanation and all 43 row annotations for scenario_031 head_medicaid_eligible (California, single head age 67, reference: eligible) explained the reference by subtracting a Medicare Part B premium from countable income. policyengine-us 2.15.17, the reference engine, subtracts no premium. This PR rewrites that text to the engine's actual mechanism.

It is wording only:

  • no reference, exclusion, failure class or score changes;
  • the frozen run directory is untouched.

The finding came from the Part B audit (#193). Details are in reference_audit/2026-10-05-medicaid-031-annotations/README.md.

What the engine does (read in 2.15.17 source this session)

  • Category. SENIOR_OR_DISABLED. is_ssi_aged is age ≥ 65, so age alone meets the category condition, with no SSI receipt or disability needed.
  • Adult expansion group. is_adult_for_medicaid_nfc requires age 19 to 64 and no Medicare eligibility, so this head is outside it. It is the only formula in variables/gov/hhs/medicaid that bars on Medicare.
  • Countable income.
    • medicaid_optional_senior_or_disabled_countable_income counts SSI unearned income: $9,445.47 Social Security (in full), $13,608 pension and $800 IRA, for $23,853.47. The $1,165 of alimony paid is not subtracted.
    • The state disregard senior_or_disabled.income.disregard.individual['CA'] is $230 a month. It is passed to _apply_ssi_exclusions as general_exclusion, so it replaces SSI's $20 exclusion: $23,853.47 − $2,760 = $21,093.47.
  • Limit. 1.38 × the 2026 guideline for one person ($15,960) = $22,024.80. Income at or below the limit passes.
  • Assets. $4,200 of countable resources, under CA's $130,000 limit.
  • Part B. The engine models a $2,434.80 Part B premium, but countable income does not read it. With the premium at 0, or Medicare enrollment off, countable income is still $21,093.47 and the result is the same.

The old text compared "about $21,400" with "about $21,600", which is 138% of the 2025 guideline. Gross income is $1,828.67 over the real limit, and gross less only SSI's $20 exclusion is still over it. So the models that skipped California's disregard still erred.

Root cause. The judge prompt behind the published text stated the wrong mechanism twice:

  • the case's grounding line (untracked grounding.csv, pinned by GROUNDING_SHA256): "a non-MAGI pathway whose income counting deducts health insurance premiums (including Medicare Part B)";
  • the old reference explanation.

No other case's grounding carries that phrase. This PR fixes the explanation. The grounding line should be regenerated before the next audit stage that re-judges this case.

Changes

  • rewrites.json: 45 whole-field rewrites (1 case note, 1 reference explanation, 43 rows).
  • scripts/apply_rewrites.py:
    • applies the rewrites and re-pins the three files' sha256 in paper/snapshot/20260501/manifest.json;
    • refuses a missing or ambiguous row, or a field that holds neither text;
    • round-trips the CSVs byte for byte;
    • with --check, writes nothing and confirms every rewrite is in force.
  • scripts/engine_values.py → verification/engine_values.json: recomputes every figure on 2.15.17 + latest_final and asserts each relation the text states, under the modeled premium, no_part_b and not_enrolled.
  • scripts/payload_diff.py → verification/payload_diff.json: rebuilds the payload from the frozen run with export_country, once with main's annotations and once with these.
  • verification/row_review.json: per-row review record (quoted model evidence, problems found, edits after review).
  • tests/test_annotation_rewrites.py: keeps the ledger in force, so a release rebuilding these files from an older base can't silently restore the old text.

Invariants (stated and checked)

  • Wording only. Against origin/main, exactly 45 fields differ, all on scenario_031 head_medicaid_eligible. Keys, order, quoting, failure_source, failure_subtype and reference_suspect are identical (checked by a field-by-field diff; test_rewritten_rows_keep_their_classes).
  • No score change. The two payload rebuilds differ in 132 paths, all in that one cell's text fields: 43 annotation, 43 caseAnnotation and 46 referenceExplanation. No modelStats, programStats, heatmap or other cell value moves.
  • Base reproduction. The base rebuild reproduces the frozen data.json.gz, except claude-fable-5's four usage fields, which the drivers carry (CARRIED_USAGE).
  • Mechanism. engine_values.py asserts each of these:
    • countable income = gross − 12 × $230;
    • the limit is 1.38 × $15,960;
    • countable income ≤ the limit < gross − $240;
    • the premium-off and not-enrolled readings give identical countable income and eligibility;
    • the published reference is 1.
  • Subtypes kept. The decisive-error tally (29 income counting / 4 threshold / 3 asset / 7 category) equals the recorded subtype counts (29 / 4 / 3 / 7).

Not changed here, and why

Coordination with release tooling

Tests

  • pytest -m "not slow": see the comment below for the run on this head; the baseline on origin/main was 1563 passed, 8 skipped.
  • ruff check . and ruff format --check . are clean.
  • apply_rewrites.py --check: 45 rewrites in force.

Review

An independent Opus 5.5 Subfleet review is running against 19d28a54. Its report will be added as verification/independent_review.md.

🤖 Generated with Claude Code

…a disregard

The case note, reference explanation and 43 row annotations for
scenario_031 head_medicaid_eligible explained the reference (eligible) by
subtracting a Medicare Part B premium from countable income. policyengine-us
2.15.17 subtracts no premium. It counts $23,853.47 of SSI unearned income and
applies California's $230 monthly disregard in place of SSI's $20 exclusion.
That gives $21,093.47 against $22,024.80 (138% of the 2026 guideline).

Wording-only: every row stays llm_error with its subtype. The ledger
(reference_audit/2026-10-05-medicaid-031-annotations/rewrites.json) is
applied by scripts/apply_rewrites.py, which re-pins the manifest hashes, and
tests/test_annotation_rewrites.py keeps it in force. engine_values.py
asserts every figure on the reference system. payload_diff.py shows a
re-export changes only this cell's text fields. The frozen run dir is
untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
policybench-site Ready Ready Preview Oct 5, 2026 6:40pm UTC

Request Review

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Local run on 19d28a54 (Python 3.12, uv sync --locked --extra dev, the CI command pytest -m "not slow"): 1568 passed, 8 skipped, 14 deselected. The baseline on origin/main 8b4c0ca1 was 1563 passed, 8 skipped; the difference is the 5 new tests in tests/test_annotation_rewrites.py. ruff check . and ruff format --check . are clean. apply_rewrites.py --check: 45 rewrites in force. engine_values.py (2.15.17 venv) and payload_diff.py ran with every assertion passing.

MaxGhenis added a commit that referenced this pull request Oct 5, 2026
…_114 row check

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions

Narrows 'holds only if the disregard is skipped' to the engine's limit, names
gemini-3.5-flash's source for $23,853, drops 'about' before exact figures,
rewords claude-fable-5, the case note's category clause and the explanation's
limit phrase. Records immigration_status in engine_values.json and tightens
payload_diff's reproduction check. Adds the review as
verification/independent_review.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 6, 2026
The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 9, 2026
…last answer (release dashboard-data-20261006) (#202)

* Date the model response window from the last answer, not the release date

The release drivers set the manifest's model_response_date to
"2026-06-12 to <release date>". Release dashboard-data-20260930 recorded
an end of 2026-09-30, but the last answer on that board (GPT-6.1 Sol)
completed at 2026-09-29 22:36:32 UTC and no model answered on the 30th.

freeze_snapshot.model_response_window now derives the end as the UTC date
of the latest request_completed_at in the predictions the freeze gzips
into the snapshot. main() derives it before its first write and passes it
to build_manifest. The start stays configured (MODEL_RESPONSE_START),
because the first waves' rows carry no timestamps and the run label gives
their date. A driver that still sets MODEL_RESPONSE_DATE must name the
derived window, or the freeze refuses. Both drivers stop setting it, and
freeze_gpt61sol derives the window in its preflight so --dry-run reports
it.

The live manifest is not changed here: correcting it is a published-claim
change queued for Max (cos decision d831). A strict xfail test records the
live mismatch and fails once a freeze corrects the manifest, so that
freeze must remove the mark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Audit the state income tax withheld in federal SALT; propose three exclusions

policyengine-us 2.15.17 fills the state income tax part of the federal SALT
deduction with state_withheld_income_tax, an AGI-formula estimate of
withholding, and no prompt states state income tax withheld or paid. This adds
the audit (reference_audit/2026-10-05): a sweep of all 1,984 outputs under the
reference, liability, net and zero readings (baseline reproduces all 1,928
scored references), the side readings the records quote, three proposed
reference_depends_on_unlisted_input records (022 CA, 081 MA, 114 VA federal
income tax), and the leaderboard impact scored through the analyze CLI on
scratch copies.

No reference, exclusion record or published number changes here. The
exclusions change published scores and wait for Max's ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Audit Louisiana's 2026 standard deduction behind two scored references

scenario_051 and scenario_077 state income tax rest on policyengine-us's
explicit 2026 Louisiana standard deduction of $12,835 (PR #8411), which
no Louisiana publication states. The audit sources what Louisiana
published, sweeps all 1,984 outputs under six candidate deductions on
policyengine-us 2.15.17 with latest_final, and scores each option
through the analyze CLI. It changes no reference.

Before the 2026-07-03 freeze the Department of Revenue published only a
provisional $12,875 (withholding tables and estimated-tax worksheet),
saying the return amount would differ; RIB 26-019 (2026-09-28) set it
at $12,838, which moves the references by $0.09 and no exact match.
The audit recommends keeping both references and drafts the
documentation; the ruling is Max's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep the sweep, scoring and check logs the audit README cites

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Audit optional employer pass-through in the payroll tax references; draft both records

Four scored payroll_tax references (scenario_032 MN, 043 CO, 081 MA, 082 NY) count
an employee share of a state paid-leave premium that the law lets the employer
deduct but does not require. policyengine-us 2.15.17 counts the largest share the
employer may deduct. The output asks for "mandatory employee state payroll taxes".

reference_audit/2026-10-05-payroll classifies every state program in the payroll
references from primary law (two adversarial reviews per program, none refuted),
sweeps all 1,984 outputs with an output-scope adapter that drops the optional
shares (four outputs move, nothing else), and drafts both candidate records:
proposed_exclusions.json and proposed_regenerations.json. It scores each through
the analyze CLI; the unchanged copy reproduces the published payload.

Changes no reference or published number. Draft README; review pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Audit the Medicare Part B premium in medical expenses; propose excluding scenario_114's Virginia income tax

policyengine-us 2.15.17 treats every Medicare-eligible person as enrolled
(takes_up_medicare_if_eligible defaults to true) and adds the modeled Part B
premium, $2,434.80 for 2026, to medical expenses. No prompt states Medicare
enrollment or a premium. A sweep of all 1,984 outputs on the reference system
with the premium at 0, and with enrollment off, moves exactly two: scenario_114
federal (10,729.61 -> 11,265.27, already proposed in #191) and Virginia
(3,514.15 -> 3,654.15). No SNAP or eligibility output moves by any amount.

Adds reference_audit/2026-10-05-medicare-part-b: the sweep, a per-household
explanation, the drafted records (one for the Virginia output, and a fallback
for the federal output if #191's record is not adopted) and the leaderboard
impact under each combination with #191. Changes no reference or published
number; the exclusion waits for Max's ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Rest the payroll recommendation on the reference rules; record the independent review

An eight-agent review of 8714770 (differential impact and sweep, mechanism,
advocates for keep and regenerate with a judge, release checklist, claim audit)
agreed on every number: 2,438 impact values, all 1,984 swept outputs bit for bit,
and the 36 hand-computed payroll parts. The judge upheld exclusion but struck the
draft's board-effect and model-behaviour arguments, which no rule uses.

- README: the recommendation now rests on 2026-09-22/09-28 rule 4 and the r24,
  r25 and r14 precedents; regeneration is the strict reading of "mandatory".
  Adds both options' board effects side by side, the counter-evidence the
  reviewers found, the Acts 2026 c. 101 line mapping, the review summary and
  the release checklist. Fixes the claim audit's five errors and thirteen
  imprecisions.
- proposed_exclusions.json: one shared unlisted input; the basis cites the
  rules and precedents (no engine input exists for the employer's choice).
- payroll_mandatory_scope.py: docstring states its scope (DE, ME, VT and WA left
  for upstream). Sweep, proposals and impact rerun on the new hash; values
  unchanged.
- verification/independent_reviews.json: the eight reports.

Changes no reference or published number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name the payroll decision (cos d972) in the audit README

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Answer the take-up reading in the Part B records; add the legacy impact summary and independent reviews

- State the prompt's "assume program take-up" sentence in both records and the
  README, and why it does not settle Medicare enrollment (the benchmark's
  DEFAULT_TAKEUP_INPUTS has no Medicare entry; the request asks eligibility only).
- Name both ACA coverage tests in the mechanism section.
- leaderboard_impact.py recomputes the freeze's legacy impact_summary_by_model.csv
  in every case (the unchanged copy reproduces the frozen file). The stale-weight
  check now reports that file, which reorders 10 of 46 models with 2.15.17's
  weights, so the finding is narrowed to "no change in the analyze payload".
- Read #191's records from a sha256-checked copy at its head 8af912a, so this
  directory does not depend on #191's branch.
- Add the four workflow reviewers' reports (verification/independent_reviews.json).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Correct scenario_031's Medicaid annotations to the engine's California disregard

The case note, reference explanation and 43 row annotations for
scenario_031 head_medicaid_eligible explained the reference (eligible) by
subtracting a Medicare Part B premium from countable income. policyengine-us
2.15.17 subtracts no premium. It counts $23,853.47 of SSI unearned income and
applies California's $230 monthly disregard in place of SSI's $20 exclusion.
That gives $21,093.47 against $22,024.80 (138% of the 2026 guideline).

Wording-only: every row stays llm_error with its subtype. The ledger
(reference_audit/2026-10-05-medicaid-031-annotations/rewrites.json) is
applied by scripts/apply_rewrites.py, which re-pins the manifest hashes, and
tests/test_annotation_rewrites.py keeps it in force. engine_values.py
asserts every figure on the reference system. payload_diff.py shows a
re-export changes only this cell's text fields. The frozen run dir is
untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Point the Part B audit's annotation findings to #197 and the scenario_114 row check

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Apply the independent review's wording fixes to scenario_031's annotations

Narrows 'holds only if the disregard is skipped' to the engine's limit, names
gemini-3.5-flash's source for $23,853, drops 'about' before exact figures,
rewords claude-fable-5, the case note's category clause and the explanation's
limit phrase. Records immigration_status in engine_values.json and tightens
payload_diff's reproduction check. Adds the review as
verification/independent_review.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep CSV timestamps exact and test the real manifest builder's window

Answers the hard-tier review of #188 (b6b385c):
- model_response_window parses with float_precision=round_trip; pandas' default
  parser rounded the last double before 2027-02-01 UTC up to midnight and dated
  a January 31 answer February 1. New end-to-end CSV regression.
- A test now calls the real build_manifest with a window no release has used
  and requires it in model_response_date and both reproducibility notes; a
  builder hard-coded to September 30 fails it (mutation f).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the release 20261006 driver, its spec and its tests

scripts/release_20261006.py stages, exports and freezes release
dashboard-data-20261006 from release 20260930 (read from git at 8b4c0ca):
the eight ruled exclusion records and their adjudications, the record edits
the three audits wrote, the annotation rewrites, and the derived response
window. docs/release_20261006/spec.json holds its inputs.

Gates: the base stage is 20260930's; the annotations differ from 20260930's
only on the excluded and reworded outputs; two exports agree byte for byte;
the payload differs from 20260930's only where the release allows; rebuilt
with 20260930's exclusion record, the staged bundle reproduces 20260930's
statistics exactly; the freeze leaves references, predictions and serving
configuration as they were and changes the manifest only where the release
allows.

tests/test_release_20261006.py rebuilds each committed record from git and
the spec, rescores the payload's rows independently, and checks the gates
with Hypothesis properties.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Date the 2026-10-05 wave's decisions in UTC, document the driver, widen the in-force test

- spec: the new wave's adjudications were written on 2026-10-06 UTC;
- DATE_CONVENTIONS_NEW names Max's rulings of 2026-10-05 US Eastern time;
- docs/release_20261006/design.md documents the driver's inputs, gates and order;
- the in-force test strips any adjudication sentence before re-appending it;
- rescore_sensitivity_summaries knows the word for 64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* WIP: partial content edits written against the trial freeze

Test pins, paper and docs prose, and a draft release note from the stopped
content pass. Some are reviewed and some are not; the content pass after the
fresh freeze finishes them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the release's annotation rewrite ledger and its evidence

306 whole-field rewrites in four groups, each drafted from its audit's
evidence, checked by an adversarial verifier and revised:
- salt (81): scenario_022 and scenario_081 federal income tax; the SALT
  state income tax is the engine's withholding estimate, not the household's;
- s114 (67): scenario_114 federal and Virginia income tax; the withholding
  estimate and the imputed Medicare Part B premium;
- payroll (151): the four paid-leave payroll outputs; optional employer
  pass-through shares, and scenario_032's and 043's note errors;
- louisiana (7): scenario_051 and 077 case notes and rows; sourced
  $12,835 chain, no change to scored references.

Each group's evidence file quotes the model explanations and engine
values behind every changed number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the d994 hold in the release spec

* Freeze dashboard-data-20261006 and rescore sensitivity evidence

* Pin the release note to the draft v2 contract and document the d994 hold

* Render the paper, pin it, record d994's ruling and the validation logs

Renders the paper outside the sandbox with the installed Quarto 1.9.36 and
re-pins it with freeze_snapshot.py --rendered-only (pdf sha256 9403564b...).
The Louisiana audit README now records d994 (ruled 2026-10-06, after this
release was built): the convention is kept as written, and both Louisiana
references are excluded in the release after this one.

Validation logs: render exit 0; app lint, 170 bun tests and the build exit
0; full pytest 1,651 passed, 9 skipped, 1 failed. The failure is
test_local_claude_models_resolve_without_remote_cost_map[claude-sonnet-5-5],
which fails on main too once litellm loads its remote cost map. PR #201
fixes it, and this branch takes the fix from main before it merges.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the reviewed head

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name the published rates behind Massachusetts's maximum; restore the pre-freeze tree

The independent review of head 8b831f2 found that the release note and 32
annotation rewrites called $805.01 (0.46%) the most a Massachusetts employer
may deduct. The payroll audit records a competing reading: a literal
application of Acts 2026 c. 101 s. 45 would let a large employer deduct up to
0.772% ($1,351.02) for 2026, while the Department's rates page kept 0.46%.
The texts now say "the most Massachusetts's published 2026 rates let the
employer deduct": the ledger's scenario_081 payroll texts, its adjudication
reasoning, a fifth record edit for the exclusion record, and the note. The
exclusion and every score are unchanged.

Also from the review: verify_payload now requires the staged payload's
households and outputs to be the base's before it compares cells, with a
test that a dropped cell or household is refused; the design note records
d994's ruling instead of calling it pending.

The freeze's outputs (5435bab) are restored to their pre-freeze state so
prepare, export and freeze can run again on the corrected ledger.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Freeze dashboard-data-20261006 again on the corrected Massachusetts wording

Payload f03441f64316997de843e82c27272442c794148a121c25e1343fbddbe38574a0,
127,224,818 bytes. Against the reviewed head 8b831f2 the freeze changes 31
row annotations, one case note, one adjudication's reasoning and one
exclusion record's alternative reading, all for scenario_081's payroll
output, and the hashes that follow. The analysis files and the sensitivity
evidence are byte-identical, so no score moved. The paper is rendered and
pinned again (pdf sha256 f3ccebc8...).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Point the effects record at the corrected payload; drop the session handoff drafts

The summary CSV the effects record binds is byte-identical after the
re-freeze, so every expected effect still holds; only the payload hash it
names changes. The handoff drafts were one session's working notes, not
repo documentation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the validation logs of the corrected freeze

Full pytest on the merged head: 1,662 passed, 9 skipped, exit 0. App lint,
170 bun tests and the build: exit 0. Paper render: exit 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the reviewed head

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Qualify Massachusetts's caps in the two remaining exclusion-record sentences

The round-2 review found two record sentences that still stated
Massachusetts's limits without the published-rates qualification:
scenario_081's federal-tax note ('the largest share the employer may
deduct') and its payroll record, which attributed the 100% family and 40%
medical caps directly to s. 6(c). Acts 2026 c. 101 ss. 25-26 and 45 can be
read to swap those caps for 2026. The note now names the published 2026
rates, and a sixth record edit qualifies the caps the same way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Restore the pre-freeze tree for the round-3 freeze

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Freeze dashboard-data-20261006 on the round-3 record edits

Payload aa34e5c9ea926848dc7af460a98f17021bc42dae54a8015c9988a85e79a5d462,
127,225,036 bytes. Against round 2 only scenario_081's two exclusion-record
sentences change; the analysis files and sensitivity evidence remain
byte-identical to the first freeze, so no score moved. The release asset is
replaced and verified against the pointer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the round-3 validation logs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the reviewed head

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Shipped in release dashboard-data-20261006 (PR #202, squash-merged as 9ce4ade8, now the latest data release and live on policybench.org). This PR's commits were folded into that release as a merge commit, so its content is on main; closing it rather than merging it a second time. Its scenario_031 Medicaid annotation rewrites ship with the release's re-export.

@MaxGhenis MaxGhenis closed this Oct 9, 2026
MaxGhenis added a commit that referenced this pull request Oct 9, 2026
…no score changes) (#200)

* Add the reference-adversary audit pass: consensus trigger, blind adversary prompt and runners, definition conformance, publication sources (WIP: adversary not yet run)

Work from the reference-adversary session (task_8940406e), which stopped at the 2026-10-05 account cutoff before committing. 361 tests pass (test_consensus, test_definition_conformance, test_publication_sources, test_reference_adversary, test_reference_adversary_runner, test_audit). The consensus flags (61 of 1,928 cells), the definition-conformance report and the publication-source report are generated; the adversary judges have not run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Point scenario_031's skip at #197 and record the rebased test run

The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run reference-adversary cases in a rolling pool, not fixed batches

Both runners started AUDIT_PARALLEL cases and waited for all of them before starting more, so one slow case held every slot: the first live batch sat 15 minutes on one SNAP case while three slots idled. The runners now start the next case as soon as any running case finishes, never running more than AUDIT_PARALLEL at once (bash 3.2 has no wait -n, so they count running jobs with jobs -pr; a finished job not yet noticed only delays a start). The Claude runner still checks the stop flag before every start.

test_claude_runner_refills_a_slot_without_exceeding_the_parallel_cap holds scenario_001's stage 1 until scenario_003's first call starts: a batch runner never releases it, the pool does, and the call timeline never shows more than two calls at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run the reference adversary over the 52 consensus-flagged cells another audit does not cover

Opus 5.5 at xhigh judged each case in two blind stages on Subfleet lane claude-9 (oauth token from the lane's keychain entry, an empty lane config dir, never the desktop login or an API key): 104 judge calls, all accepted, none contaminated, invalid or refused. Verdicts: 45 reference_holds, 6 reference_wrong (scenario_018 AZ, 025 OH, 026 child1 and child2 NC, 043 CO, 082 NY), 1 definition_mismatch (123 PA). adversary-collect reports 0 missing and 0 inconsistent, and queues 7 cases for developer adjudication. No verdict changes a score.

runs/ holds every case's prompts, outputs, provenance sidecars and session transcripts, the runner logs, the collected tables and run_record.json. scripts/verdict_table.py renders verification/verdict_table.md and verdict_counts.json; scripts/engine_probe.py rebuilds a frozen-run household on the reference system for engine evidence (verification/probes/).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the independent verification of the seven non-holding adversary verdicts (WIP: lane stopped before the report and PR)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Write the reference-adversary report, record Max's rulings and the final test run

README.md reports the pass: design, flag parameters, the 61 flagged / 9 skipped / 52 run
cells, the verdicts, the independent verification of the seven that did not hold (four
confirmed engine defects, one ambiguous definition scope over four cells, NC 026 refuted),
leaderboard impact, the definition-conformance and publication-source findings, judge cost
(104 calls, $91.22 reported by the CLI, 5.25 judge-hours) and the rulings.

proposed_changes.json status now records d1022 (exclude the eight cells in the next
release after dashboard-data-20261006; regenerate the four defect cells on fixed
policyengine-us) and d994 (exclude Louisiana 051 and 077; keep the published-amounts
convention). The root-cause records are unchanged. verification/pytest_final.txt: 362
passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Cite d1029, the queued household-scope decision, in the report

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep the diagnosis judge's template unchanged in this PR

Dropping the stale "bugs were fixed before this run" sentence from the judge
prompt changes all 674 seed prompts. The fold drivers carry a seed verdict
only when its prompt re-renders byte-identically, so the next model addition
would refuse or need a full re-judge. Restore the template and its test to
origin/main's, and note in the report that the sentence's removal waits for a
versioned judge template, in a separate PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Address the code review of the reference adversary

Brings over the review fixes a stopped lane left on
reference-adversary-review-wip (4c1334e), with its failing test fixed and
lint cleared:

- Blinding: the Claude transcript audit now reads each WebSearch result and
  rejects an output whose search results list a blocked URL or name
  PolicyEngine or PolicyBench (WebSearch has no deny rule, so results reach
  the judge). The source rule asks the judge to pass blocked_domains on
  every search.
- adversary-collect refuses a repeated judge label, queues every case with
  an inconsistent verdict whatever its class, and exits non-zero when a case
  has no usable verdict unless --allow-missing is given.
- collect_adversary requires a verdict sidecar bound to the current stage 1.
- Consensus: a member whose own answer is within the tolerance counts as
  exact and leaves the wrong cluster (consensus_flags.json and the prototype
  file reproduce byte-identically under the new rule).
- check_login refuses a token lane that declares the desktop login's account.
- Codex runner: an allowlisted environment, now with LC_CTYPE as the Claude
  runner has, and a refusal of a Codex home holding an AGENTS.md.
- leaderboard_impact.py refuses a --scratch inside the repository and an
  unknown --only variant; build_proposals.py records each record's ruling
  (d1022), separates the pass's own expectations from the rulings, and
  states the Colorado basis as the verifier did. proposed_changes.json is
  its output, reproduced byte-identically.

The runner test test_codex_runner_passes_only_allowlisted_variables failed
on LC_CTYPE. The runner was right: with no LANG in the allowlisted
environment, the fake codex (a Python script) sets LC_CTYPE=C.UTF-8 itself
under PEP 538 locale coercion. The test now allows LC_CTYPE, which the
runner also forwards for parity with the Claude runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Re-judge the 10 cases whose search results reached a blocked source

The review found WebSearch results listing PolicyEngine and GitHub pages in
run 2's transcripts. Under the audit that now reads search results, 13 of
the 104 accepted outputs, in 10 cases, would have been rejected: 28 exposed
searches, mostly github.com/PolicyEngine and TheAxiomFoundation issues, and
one policyengine.org page.

A lane re-judged those 10 cases on 2026-10-06 (runs/claude-rejudge) with the
source rule that asks for blocked_domains on every search; it left the
evidence uncommitted. This records it:

- scripts/search_exposure.py runs the current transcript audit over both
  runs and writes verification/search_exposure.{json,md}, and rewrites
  runs/rejudge_flags.json (consensus_flags.json restricted to the 10 cells)
  byte-identically to the file the re-judge was prepared from.
- runs/collected-rejudge is adversary-collect's output for the re-judge:
  10 verdicts, 0 missing, 0 inconsistent, Colorado 043 queued.
- runs/run_record.json records the re-judge's lane, login, settings, times
  and cost ($24.82 reported, 4.70 judge-hours).
- The runner logs the README and verdict_table.py cite were git-ignored
  (*.log) and never committed; they are now, with the re-judge's.

Outcome: 0 of the 20 re-judge transcripts flagged; verdict unchanged in all
10. One stage 1 moved: scenario_013 SNAP now finds $288, and stage 2 still
holds the $240 reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Report the search exposure and re-judge, and fix the report review's findings

README:
- A section on search results from blocked sources: what the audit missed,
  the 13 flagged transcripts in 10 cases by source, the re-judge and its
  outcome (every verdict unchanged), and scenario_013 SNAP, whose blind
  stage 1 could not date Arizona's 200% limit and found $288.
- Counts with the re-judged outputs in place (verdicts unchanged; 39 high
  and 13 medium; stage 1 41/6/3/2), the re-judge's cost, files and
  reproduce steps.
- Report review: 35 amount and 17 eligibility derivations, not 42 and 10;
  score ranks shift more than exact ranks and within-1% ranks less (no 5/10%
  ranks exist); d1022's bullets keep to the ruling, and the pass's own
  expectations move out of them; the AZ filing-status leaves end at their
  2025 values, not all at $15,750.
- Open items: the four engine fixes' policyengine-us PRs (Ohio's premiums
  fix has none yet), scenario_013's effective date, and the Codex runner's
  blindness to search results.

docs/audit.md and the Claude runner header now say that WebSearch has no
deny rule, that the audit reads its results, and that the Codex runner
cannot; and describe the consensus rule's exact-member exclusion and the
collect command's queue, sidecar binding, labels and --allow-missing.

search_exposure.py writes runs/rejudge_flags.json before the re-judge
exists, so the reproduce steps run in order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run: 376 passed at 10b3d16

The six test files (the five adversary files plus tests/test_audit.py),
with ruff check and ruff format clean at the locked ruff 0.15.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Fix the round-2 review's findings on the reference adversary

The independent review of a29cb3c found four gaps:

- The Codex runner refused a Codex home holding AGENTS.md but not
  AGENTS.override.md, which Codex loads first. It now refuses either, and
  the refusal test covers both.
- check_login compared AUDIT_ACCOUNT to the desktop login's email verbatim,
  so a lane declaring "claude:<desktop email>" passed. The comparison now
  drops a provider prefix; a prefixed regression covers it.
- collect_adversary and adversary-collect accepted a directory with no
  cases.jsonl and collected it as a judge with no cases. Both now refuse it.
- adversary-collect split LABEL=DIR at the last "=", so a directory
  containing "=" broke. It now splits at the first "=" when the text before
  it is a label, and keeps a bare DIR whole otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at b46c4a5 and document the round-2 fixes

pytest_final.txt: the six test files (the five adversary files plus
tests/test_audit.py), run one file at a time on a loaded machine: 379
passed, ruff check and ruff format clean at the locked ruff 0.15.2.

docs/audit.md: adversary-collect splits LABEL=DIR at the first "=" and
refuses a directory without cases.jsonl; the Codex runner refuses a Codex
home holding AGENTS.override.md or AGENTS.md. A missing blank line had
folded the paragraph after the collect rules into the last list item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Fix the round-3 review's findings on the reference adversary

The independent review of a6323f8 found one major gap and five smaller
ones:

- The Codex runner checked only the AGENTS files, but Codex adds other text
  to a session without a tool event: config.toml's developer_instructions
  and model_instructions_file, memories, and the skills under
  $HOME/.agents/skills. Each judge call now skips the lane's config.toml
  (--ignore-user-config; auth still comes from CODEX_HOME), runs with
  memories off (--disable memories), and gets a fresh, empty HOME, removed
  afterwards (the login check too). CODEX_HOME is always passed, defaulting
  to ~/.codex. The runner also refuses a Codex home holding any skill but
  the bundled .system ones. The header, docs/audit.md and the README name
  what it still cannot control: what Codex bundles, an administrator's
  /etc/codex, and apps or plugins on the ChatGPT account (whose calls are
  MCP tool events, which the audit rejects). No committed run used this
  runner.
- check_login refused an empty AUDIT_ACCOUNT but accepted "claude:" or
  blank space, which declare no account; it now tests the parsed account.
  Its docstring says it is stricter than run_audit_claude.sh, not the same.
- adversary-collect compared judge labels case-sensitively, so claude and
  Claude shared files on a case-insensitive file system, and accepted the
  label "merged", whose files are the merged table's. Both are refused, as
  is an empty DIR (claude=). The docs and help say to pass a relative bare
  DIR containing "=" as ./adv=2, and a test covers it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at bc5925e: 386 passed

The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The round-3 fixes add four tests to
test_reference_adversary.py and three to the runner tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read the pass's payload from its commit, and fix the round-4 review's notes

CI on 8f968e8 failed six tests once main carried release
dashboard-data-20261006 (#202): they read the working tree's payload,
which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran
on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus
and frozen-run tests now read the payload and the reference explanations
as 8b4c0ca (#187) committed them, and the README says the Reproduce steps
need those inputs.

The round-4 review approved 8f968e8 with two notes:

- The Codex runner's skills check parsed `ls` output, so an entry named a
  bare newline, or ".system\n.system", passed it. It now walks the
  directory with globs and refuses every entry but a real .system
  directory (a symlink named .system too), naming it with %q; tests cover
  both names and the symlink.
- The AUDIT_MODEL default said "the lane's"; with the lane's config.toml
  skipped it is Codex's own.

The review left hooks as an unconfirmed channel. Each Codex call now also
turns the hooks, plugins and apps features off, as codex-cli 0.159.0's
`codex features list` names them, beside memories; the docs name what
remains outside the runner's control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at 4a62b1e: 391 passed

The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The skills-check tests now cover five
names and a symlinked .system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — c484ea40 Deployed Oct 5, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant