Skip to content

Stop scoring eight tax outputs and date the response window from the last answer (release dashboard-data-20261006) - #202

Merged
MaxGhenis merged 39 commits into
mainfrom
release-exclusions-batch
Oct 9, 2026
Merged

MaxGhenis merged 39 commits into
mainfrom
release-exclusions-batch

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Release dashboard-data-20261006. It applies Max's rulings of 2026-10-05 to release dashboard-data-20260930. No model, answer, reference value or serving row changes; the board snapshot stays "Snapshot 2026-09-30".

Ruling Change Evidence
d963 Exclude federal income tax before refundable credits for scenario_022 (CA), 081 (MA) and 114 (VA). Their SALT deduction rests on policyengine-us's formula estimate of state income tax withheld, which no prompt states. reference_audit/2026-10-05/ (#191)
d974 Exclude scenario_114's Virginia income tax. Its medical deduction counts a modeled Medicare Part B premium, which no prompt states. reference_audit/2026-10-05-medicare-part-b/ (#193)
d972 Exclude payroll_tax for scenario_032 (MN), 043 (CO), 081 (MA) and 082 (NY). Each counts the employee share of a state paid-leave premium that the employer may deduct but need not. reference_audit/2026-10-05-payroll/ (#194)
d831 Date the model response window from the last answer: 2026-06-12 to 2026-09-29. scripts/freeze_snapshot.py (#188)
Louisiana (#192) Keep scenario_051 and 077 at the $12,835 deduction for this release; ship the wording-only case-note rewrites. reference_audit/2026-10-05-louisiana/

This PR folds #188, #191, #192 (without its held sentence), #193, #194 and #197, plus 306 wording-only annotation rewrites (docs/release_20261006/annotation_rewrites.json, with evidence). d994 was ruled on 2026-10-06, after this release was built: keep the convention, and exclude Louisiana 051/077 in the next release. The Louisiana README records it.

Effects (reproduced in docs/release_20261006/validation/effects.json)

Driver, gates, invariants

scripts/release_20261006.py (prepare / export / freeze), docs/release_20261006/spec.json and design doc docs/release_20261006/design.md. The invariants are in tests/test_release_20261006.py:

  • the frozen exclusion record is build_exclusions(20260930's, spec) byte for byte;
  • adjudications are 20260930's plus eight, each keeping its bound judge verdict;
  • the annotation files rebuild byte for byte from 20260930's, plus Correct scenario_031's Medicaid annotations to the engine's California disregard #197's ledger, the rewrites and the adjudications;
  • references and predictions are unchanged;
  • every model is scored on 1,920 outputs;
  • an independent re-aggregation of the payload reproduces every model's exact rate;
  • Hypothesis properties: rewrites move only their fields and are idempotent; a rewrite of text that isn't there is refused; manifest diffs are symmetric and empty only for equal manifests.

Asset

Uploaded as non-latest before CI: https://github.com/PolicyEngine/policybench/releases/tag/dashboard-data-20261006

  • dashboard-data.json: 127,218,158 B, sha256 1780d2ec37f90b654265f8c7a191ba57f5ec5c4d28624bf9d2e2de8a48d2f871. Downloaded back and verified; it equals app/src/data.artifact.json.
  • predictions.csv.gz: sha256 ca2c4c48…, byte-identical to the frozen snapshot copy.

Validation (docs/release_20261006/validation/)

  • Paper render: exit 0 (HTML and PDF), re-pinned with --rendered-only.
  • App: lint, 170/170 bun tests and the build, exit 0. Ruff check and format pass.
  • Full pytest: 1,651 passed, 9 skipped, 1 failed. The failure, test_local_claude_models_resolve_without_remote_cost_map[claude-sonnet-5-5], fails on main too once litellm loads its remote cost map. Add the Claude Haiku 5.5 model card and metadata #201 fixes it, and this branch merges main after Add the Claude Haiku 5.5 model card and metadata #201 lands.
  • Independent review: subfleet 20261008-175317-pb-20261006-final-review (Opus 5.5, standard tier). The result will be posted here.

After merge: gh release edit dashboard-data-20261006 --latest, a prod check on policybench.org, and closing folded PRs #188, #191, #192, #193, #194 and #197 with pointers.

🤖 Generated with Claude Code

MaxGhenis and others added 28 commits October 2, 2026 08:13
…date

The release drivers set the manifest's model_response_date to
"2026-06-12 to <release date>". Release dashboard-data-20260930 recorded
an end of 2026-09-30, but the last answer on that board (GPT-6.1 Sol)
completed at 2026-09-29 22:36:32 UTC and no model answered on the 30th.

freeze_snapshot.model_response_window now derives the end as the UTC date
of the latest request_completed_at in the predictions the freeze gzips
into the snapshot. main() derives it before its first write and passes it
to build_manifest. The start stays configured (MODEL_RESPONSE_START),
because the first waves' rows carry no timestamps and the run label gives
their date. A driver that still sets MODEL_RESPONSE_DATE must name the
derived window, or the freeze refuses. Both drivers stop setting it, and
freeze_gpt61sol derives the window in its preflight so --dry-run reports
it.

The live manifest is not changed here: correcting it is a published-claim
change queued for Max (cos decision d831). A strict xfail test records the
live mismatch and fails once a freeze corrects the manifest, so that
freeze must remove the mark.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…clusions

policyengine-us 2.15.17 fills the state income tax part of the federal SALT
deduction with state_withheld_income_tax, an AGI-formula estimate of
withholding, and no prompt states state income tax withheld or paid. This adds
the audit (reference_audit/2026-10-05): a sweep of all 1,984 outputs under the
reference, liability, net and zero readings (baseline reproduces all 1,928
scored references), the side readings the records quote, three proposed
reference_depends_on_unlisted_input records (022 CA, 081 MA, 114 VA federal
income tax), and the leaderboard impact scored through the analyze CLI on
scratch copies.

No reference, exclusion record or published number changes here. The
exclusions change published scores and wait for Max's ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scenario_051 and scenario_077 state income tax rest on policyengine-us's
explicit 2026 Louisiana standard deduction of $12,835 (PR #8411), which
no Louisiana publication states. The audit sources what Louisiana
published, sweeps all 1,984 outputs under six candidate deductions on
policyengine-us 2.15.17 with latest_final, and scores each option
through the analyze CLI. It changes no reference.

Before the 2026-07-03 freeze the Department of Revenue published only a
provisional $12,875 (withholding tables and estimated-tax worksheet),
saying the return amount would differ; RIB 26-019 (2026-09-28) set it
at $12,838, which moves the references by $0.09 and no exact match.
The audit recommends keeping both references and drafts the
documentation; the ruling is Max's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…raft both records

Four scored payroll_tax references (scenario_032 MN, 043 CO, 081 MA, 082 NY) count
an employee share of a state paid-leave premium that the law lets the employer
deduct but does not require. policyengine-us 2.15.17 counts the largest share the
employer may deduct. The output asks for "mandatory employee state payroll taxes".

reference_audit/2026-10-05-payroll classifies every state program in the payroll
references from primary law (two adversarial reviews per program, none refuted),
sweeps all 1,984 outputs with an output-scope adapter that drops the optional
shares (four outputs move, nothing else), and drafts both candidate records:
proposed_exclusions.json and proposed_regenerations.json. It scores each through
the analyze CLI; the unchanged copy reproduces the published payload.

Changes no reference or published number. Draft README; review pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing scenario_114's Virginia income tax

policyengine-us 2.15.17 treats every Medicare-eligible person as enrolled
(takes_up_medicare_if_eligible defaults to true) and adds the modeled Part B
premium, $2,434.80 for 2026, to medical expenses. No prompt states Medicare
enrollment or a premium. A sweep of all 1,984 outputs on the reference system
with the premium at 0, and with enrollment off, moves exactly two: scenario_114
federal (10,729.61 -> 11,265.27, already proposed in #191) and Virginia
(3,514.15 -> 3,654.15). No SNAP or eligibility output moves by any amount.

Adds reference_audit/2026-10-05-medicare-part-b: the sweep, a per-household
explanation, the drafted records (one for the Virginia output, and a fallback
for the federal output if #191's record is not adopted) and the leaderboard
impact under each combination with #191. Changes no reference or published
number; the exclusion waits for Max's ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dependent review

An eight-agent review of 8714770 (differential impact and sweep, mechanism,
advocates for keep and regenerate with a judge, release checklist, claim audit)
agreed on every number: 2,438 impact values, all 1,984 swept outputs bit for bit,
and the 36 hand-computed payroll parts. The judge upheld exclusion but struck the
draft's board-effect and model-behaviour arguments, which no rule uses.

- README: the recommendation now rests on 2026-09-22/09-28 rule 4 and the r24,
  r25 and r14 precedents; regeneration is the strict reading of "mandatory".
  Adds both options' board effects side by side, the counter-evidence the
  reviewers found, the Acts 2026 c. 101 line mapping, the review summary and
  the release checklist. Fixes the claim audit's five errors and thirteen
  imprecisions.
- proposed_exclusions.json: one shared unlisted input; the basis cites the
  rules and precedents (no engine input exists for the employer's choice).
- payroll_mandatory_scope.py: docstring states its scope (DE, ME, VT and WA left
  for upstream). Sweep, proposals and impact rerun on the new hash; values
  unchanged.
- verification/independent_reviews.json: the eight reports.

Changes no reference or published number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ct summary and independent reviews

- State the prompt's "assume program take-up" sentence in both records and the
  README, and why it does not settle Medicare enrollment (the benchmark's
  DEFAULT_TAKEUP_INPUTS has no Medicare entry; the request asks eligibility only).
- Name both ACA coverage tests in the mechanism section.
- leaderboard_impact.py recomputes the freeze's legacy impact_summary_by_model.csv
  in every case (the unchanged copy reproduces the frozen file). The stale-weight
  check now reports that file, which reorders 10 of 46 models with 2.15.17's
  weights, so the finding is narrowed to "no change in the analyze payload".
- Read #191's records from a sha256-checked copy at its head 8af912a, so this
  directory does not depend on #191's branch.
- Add the four workflow reviewers' reports (verification/independent_reviews.json).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a disregard

The case note, reference explanation and 43 row annotations for
scenario_031 head_medicaid_eligible explained the reference (eligible) by
subtracting a Medicare Part B premium from countable income. policyengine-us
2.15.17 subtracts no premium. It counts $23,853.47 of SSI unearned income and
applies California's $230 monthly disregard in place of SSI's $20 exclusion.
That gives $21,093.47 against $22,024.80 (138% of the 2026 guideline).

Wording-only: every row stays llm_error with its subtype. The ledger
(reference_audit/2026-10-05-medicaid-031-annotations/rewrites.json) is
applied by scripts/apply_rewrites.py, which re-pins the manifest hashes, and
tests/test_annotation_rewrites.py keeps it in force. engine_values.py
asserts every figure on the reference system. payload_diff.py shows a
re-export changes only this cell's text fields. The frozen run dir is
untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…_114 row check

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions

Narrows 'holds only if the disregard is skipped' to the engine's limit, names
gemini-3.5-flash's source for $23,853, drops 'about' before exact figures,
rewords claude-fable-5, the case note's category clause and the explanation's
limit phrase. Records immigration_status in engine_values.json and tightens
payload_diff's reproduction check. Adds the review as
verification/independent_review.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Answers the hard-tier review of #188 (b6b385c):
- model_response_window parses with float_precision=round_trip; pandas' default
  parser rounded the last double before 2027-02-01 UTC up to midnight and dated
  a January 31 answer February 1. New end-to-end CSV regression.
- A test now calls the real build_manifest with a window no release has used
  and requires it in model_response_date and both reproducibility notes; a
  builder hard-coded to September 30 fails it (mutation f).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scripts/release_20261006.py stages, exports and freezes release
dashboard-data-20261006 from release 20260930 (read from git at 8b4c0ca):
the eight ruled exclusion records and their adjudications, the record edits
the three audits wrote, the annotation rewrites, and the derived response
window. docs/release_20261006/spec.json holds its inputs.

Gates: the base stage is 20260930's; the annotations differ from 20260930's
only on the excluded and reworded outputs; two exports agree byte for byte;
the payload differs from 20260930's only where the release allows; rebuilt
with 20260930's exclusion record, the staged bundle reproduces 20260930's
statistics exactly; the freeze leaves references, predictions and serving
configuration as they were and changes the manifest only where the release
allows.

tests/test_release_20261006.py rebuilds each committed record from git and
the spec, rescores the payload's rows independently, and checks the gates
with Hypothesis properties.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…en the in-force test

- spec: the new wave's adjudications were written on 2026-10-06 UTC;
- DATE_CONVENTIONS_NEW names Max's rulings of 2026-10-05 US Eastern time;
- docs/release_20261006/design.md documents the driver's inputs, gates and order;
- the in-force test strips any adjudication sentence before re-appending it;
- rescore_sensitivity_summaries knows the word for 64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test pins, paper and docs prose, and a draft release note from the stopped
content pass. Some are reviewed and some are not; the content pass after the
fresh freeze finishes them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
306 whole-field rewrites in four groups, each drafted from its audit's
evidence, checked by an adversarial verifier and revised:
- salt (81): scenario_022 and scenario_081 federal income tax; the SALT
  state income tax is the engine's withholding estimate, not the household's;
- s114 (67): scenario_114 federal and Virginia income tax; the withholding
  estimate and the imputed Medicare Part B premium;
- payroll (151): the four paid-leave payroll outputs; optional employer
  pass-through shares, and scenario_032's and 043's note errors;
- louisiana (7): scenario_051 and 077 case notes and rows; sourced
  $12,835 chain, no change to scored references.

Each group's evidence file quotes the model explanations and engine
values behind every changed number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Renders the paper outside the sandbox with the installed Quarto 1.9.36 and
re-pins it with freeze_snapshot.py --rendered-only (pdf sha256 9403564b...).
The Louisiana audit README now records d994 (ruled 2026-10-06, after this
release was built): the convention is kept as written, and both Louisiana
references are excluded in the release after this one.

Validation logs: render exit 0; app lint, 170 bun tests and the build exit
0; full pytest 1,651 passed, 9 skipped, 1 failed. The failure is
test_local_claude_models_resolve_without_remote_cost_map[claude-sonnet-5-5],
which fails on main too once litellm loads its remote cost map. PR #201
fixes it, and this branch takes the fix from main before it merges.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
policybench-site Ready Ready Preview Oct 9, 2026 12:32am UTC

Request Review

…pre-freeze tree

The independent review of head 8b831f2 found that the release note and 32
annotation rewrites called $805.01 (0.46%) the most a Massachusetts employer
may deduct. The payroll audit records a competing reading: a literal
application of Acts 2026 c. 101 s. 45 would let a large employer deduct up to
0.772% ($1,351.02) for 2026, while the Department's rates page kept 0.46%.
The texts now say "the most Massachusetts's published 2026 rates let the
employer deduct": the ledger's scenario_081 payroll texts, its adjudication
reasoning, a fifth record edit for the exclusion record, and the note. The
exclusion and every score are unchanged.

Also from the review: verify_payload now requires the staged payload's
households and outputs to be the base's before it compares cells, with a
test that a dropped cell or household is refused; the design note records
d994's ruling instead of calling it pending.

The freeze's outputs (5435bab) are restored to their pre-freeze state so
prepare, export and freeze can run again on the corrected ledger.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 5 commits October 8, 2026 19:24
…ording

Payload f03441f64316997de843e82c27272442c794148a121c25e1343fbddbe38574a0,
127,224,818 bytes. Against the reviewed head 8b831f2 the freeze changes 31
row annotations, one case note, one adjudication's reasoning and one
exclusion record's alternative reading, all for scenario_081's payroll
output, and the hashes that follow. The analysis files and the sensitivity
evidence are byte-identical, so no score moved. The paper is rendered and
pinned again (pdf sha256 f3ccebc8...).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…andoff drafts

The summary CSV the effects record binds is byte-identical after the
re-freeze, so every expected effect still holds; only the payload hash it
names changes. The handoff drafts were one session's working notes, not
repo documentation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Full pytest on the merged head: 1,662 passed, 9 skipped, exit 0. App lint,
170 bun tests and the build: exit 0. Paper render: exit 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Independent review, round 1: REQUEST_CHANGES at 8b831f2c (GPT-6.1 Sol, Subfleet 20261008-175903-pb-20261006-final-review-sol; the Opus lanes had no capacity). It was a read-only review of the pushed head and the saved validation evidence.

Blocking, now fixed (4c60c97e). The release note and 32 annotation rewrites called $805.01 (0.46%) the most a Massachusetts employer may deduct. The payroll audit records a competing reading: a literal application of Acts 2026 c. 101 s. 45 would allow 0.772% ($1,351.02) for 2026, while the Department's rates page kept 0.46%. For scenario_081's payroll output, the texts now say "the most Massachusetts's published 2026 rates let the employer deduct". That covers the ledger, the adjudication reasoning, a fifth record edit for the exclusion record, the note and its pins. The exclusion and every score are unchanged.

Non-blocking, also fixed.

  • verify_payload now requires the staged payload's households and outputs to equal the base's before comparing cells. A new test refuses a dropped cell or household.
  • The design note records d994's ruling instead of calling it pending.

Otherwise confirmed. The reviewer checked that the eight records and eight adjudications match the rulings exactly, the scope re-export, the manifest limits, all 46 rows of the expected effects and the 306 rewrites against their evidence.

Re-freeze. New payload f03441f64316997de843e82c27272442c794148a121c25e1343fbddbe38574a0, 127,224,818 bytes. It replaces the PR body's 1780d2ec…. The non-latest release asset was replaced and downloaded back to verify. Against the reviewed head, the freeze changes only scenario_081's payroll texts and the hashes that follow. The analysis and sensitivity files are byte-identical, so no score moved.

Validation on the pushed head 2abd54c3 (main merged in, with #201):

  • full pytest: 1,662 passed, 9 skipped, exit 0;
  • app lint, 170 bun tests and the build: exit 0;
  • paper render: exit 0.

Round 2, a delta re-review by the same reviewer, is running: Subfleet 20261008-193852-pb-20261006-review-r2.

MaxGhenis and others added 5 commits October 8, 2026 20:16
…ntences

The round-2 review found two record sentences that still stated
Massachusetts's limits without the published-rates qualification:
scenario_081's federal-tax note ('the largest share the employer may
deduct') and its payroll record, which attributed the 100% family and 40%
medical caps directly to s. 6(c). Acts 2026 c. 101 ss. 25-26 and 45 can be
read to swap those caps for 2026. The note now names the published 2026
rates, and a sixth record edit qualifies the caps the same way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Payload aa34e5c9ea926848dc7af460a98f17021bc42dae54a8015c9988a85e79a5d462,
127,225,036 bytes. Against round 2 only scenario_081's two exclusion-record
sentences change; the analysis files and sensitivity evidence remain
byte-identical to the first freeze, so no score moved. The release asset is
replaced and verified against the pointer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Independent review, round 3: APPROVE at 9edf8bac (GPT-6.1 Sol, Subfleet 20261009-064206-pb-20261006-review-r3b). The PR head 5e9ee75f differs from it only by docs/release_20261006/validation/HEAD.txt, which names the reviewed parent.

  • Round 2, 20261008-193852-pb-20261006-review-r2: REQUEST_CHANGES on two scenario_081 sentences in the exclusion record: the federal note's "largest share the employer may deduct", and the payroll record's caps. Both are fixed by spec record edits that cite the published 2026 rates and flag the competing reading of Acts 2026 c. 101. Round-3 freeze aa34e5c9….
  • Round 3: "Both round-2 blockers are resolved in the frozen record." The Massachusetts wording is consistent across the record, adjudication, annotations, reference explanation, release note and case note. The analysis, impact summaries and sensitivity files are unchanged since the first freeze. Pins and receipts agree on aa34e5c9…, 127,225,036 bytes. One earlier round-3 attempt could not read files on its lane; this one used GitHub and the evidence written into the brief.
  • Validation at the reviewed head:
    • pytest: 1,662 passed, 9 skipped;
    • app: 170 bun tests, the lint and the build;
    • paper: rendered to HTML and PDF.
  • Asset: the non-latest release asset is the round-3 payload, downloaded back and verified against the pointer.

Merging on the gates: gh pr checks exit 0, MERGEABLE, not a draft, no changes requested, independent review approved. Then gh release edit --latest and a prod check.

@MaxGhenis
MaxGhenis merged commit 9ce4ade into main Oct 9, 2026
6 checks passed
@MaxGhenis
MaxGhenis deleted the release-exclusions-batch branch October 9, 2026 11:10
MaxGhenis added a commit that referenced this pull request Oct 9, 2026
… notes

CI on 8f968e8 failed six tests once main carried release
dashboard-data-20261006 (#202): they read the working tree's payload,
which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran
on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus
and frozen-run tests now read the payload and the reference explanations
as 8b4c0ca (#187) committed them, and the README says the Reproduce steps
need those inputs.

The round-4 review approved 8f968e8 with two notes:

- The Codex runner's skills check parsed `ls` output, so an entry named a
  bare newline, or ".system\n.system", passed it. It now walks the
  directory with globs and refuses every entry but a real .system
  directory (a symlink named .system too), naming it with %q; tests cover
  both names and the symlink.
- The AUDIT_MODEL default said "the lane's"; with the lane's config.toml
  skipped it is Codex's own.

The review left hooks as an unconfirmed channel. Each Codex call now also
turns the hooks, plugins and apps features off, as codex-cli 0.159.0's
`codex features list` names them, beside memories; the docs name what
remains outside the runner's control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 9, 2026
…no score changes) (#200)

* Add the reference-adversary audit pass: consensus trigger, blind adversary prompt and runners, definition conformance, publication sources (WIP: adversary not yet run)

Work from the reference-adversary session (task_8940406e), which stopped at the 2026-10-05 account cutoff before committing. 361 tests pass (test_consensus, test_definition_conformance, test_publication_sources, test_reference_adversary, test_reference_adversary_runner, test_audit). The consensus flags (61 of 1,928 cells), the definition-conformance report and the publication-source report are generated; the adversary judges have not run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Point scenario_031's skip at #197 and record the rebased test run

The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run reference-adversary cases in a rolling pool, not fixed batches

Both runners started AUDIT_PARALLEL cases and waited for all of them before starting more, so one slow case held every slot: the first live batch sat 15 minutes on one SNAP case while three slots idled. The runners now start the next case as soon as any running case finishes, never running more than AUDIT_PARALLEL at once (bash 3.2 has no wait -n, so they count running jobs with jobs -pr; a finished job not yet noticed only delays a start). The Claude runner still checks the stop flag before every start.

test_claude_runner_refills_a_slot_without_exceeding_the_parallel_cap holds scenario_001's stage 1 until scenario_003's first call starts: a batch runner never releases it, the pool does, and the call timeline never shows more than two calls at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run the reference adversary over the 52 consensus-flagged cells another audit does not cover

Opus 5.5 at xhigh judged each case in two blind stages on Subfleet lane claude-9 (oauth token from the lane's keychain entry, an empty lane config dir, never the desktop login or an API key): 104 judge calls, all accepted, none contaminated, invalid or refused. Verdicts: 45 reference_holds, 6 reference_wrong (scenario_018 AZ, 025 OH, 026 child1 and child2 NC, 043 CO, 082 NY), 1 definition_mismatch (123 PA). adversary-collect reports 0 missing and 0 inconsistent, and queues 7 cases for developer adjudication. No verdict changes a score.

runs/ holds every case's prompts, outputs, provenance sidecars and session transcripts, the runner logs, the collected tables and run_record.json. scripts/verdict_table.py renders verification/verdict_table.md and verdict_counts.json; scripts/engine_probe.py rebuilds a frozen-run household on the reference system for engine evidence (verification/probes/).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the independent verification of the seven non-holding adversary verdicts (WIP: lane stopped before the report and PR)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Write the reference-adversary report, record Max's rulings and the final test run

README.md reports the pass: design, flag parameters, the 61 flagged / 9 skipped / 52 run
cells, the verdicts, the independent verification of the seven that did not hold (four
confirmed engine defects, one ambiguous definition scope over four cells, NC 026 refuted),
leaderboard impact, the definition-conformance and publication-source findings, judge cost
(104 calls, $91.22 reported by the CLI, 5.25 judge-hours) and the rulings.

proposed_changes.json status now records d1022 (exclude the eight cells in the next
release after dashboard-data-20261006; regenerate the four defect cells on fixed
policyengine-us) and d994 (exclude Louisiana 051 and 077; keep the published-amounts
convention). The root-cause records are unchanged. verification/pytest_final.txt: 362
passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Cite d1029, the queued household-scope decision, in the report

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep the diagnosis judge's template unchanged in this PR

Dropping the stale "bugs were fixed before this run" sentence from the judge
prompt changes all 674 seed prompts. The fold drivers carry a seed verdict
only when its prompt re-renders byte-identically, so the next model addition
would refuse or need a full re-judge. Restore the template and its test to
origin/main's, and note in the report that the sentence's removal waits for a
versioned judge template, in a separate PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Address the code review of the reference adversary

Brings over the review fixes a stopped lane left on
reference-adversary-review-wip (4c1334e), with its failing test fixed and
lint cleared:

- Blinding: the Claude transcript audit now reads each WebSearch result and
  rejects an output whose search results list a blocked URL or name
  PolicyEngine or PolicyBench (WebSearch has no deny rule, so results reach
  the judge). The source rule asks the judge to pass blocked_domains on
  every search.
- adversary-collect refuses a repeated judge label, queues every case with
  an inconsistent verdict whatever its class, and exits non-zero when a case
  has no usable verdict unless --allow-missing is given.
- collect_adversary requires a verdict sidecar bound to the current stage 1.
- Consensus: a member whose own answer is within the tolerance counts as
  exact and leaves the wrong cluster (consensus_flags.json and the prototype
  file reproduce byte-identically under the new rule).
- check_login refuses a token lane that declares the desktop login's account.
- Codex runner: an allowlisted environment, now with LC_CTYPE as the Claude
  runner has, and a refusal of a Codex home holding an AGENTS.md.
- leaderboard_impact.py refuses a --scratch inside the repository and an
  unknown --only variant; build_proposals.py records each record's ruling
  (d1022), separates the pass's own expectations from the rulings, and
  states the Colorado basis as the verifier did. proposed_changes.json is
  its output, reproduced byte-identically.

The runner test test_codex_runner_passes_only_allowlisted_variables failed
on LC_CTYPE. The runner was right: with no LANG in the allowlisted
environment, the fake codex (a Python script) sets LC_CTYPE=C.UTF-8 itself
under PEP 538 locale coercion. The test now allows LC_CTYPE, which the
runner also forwards for parity with the Claude runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Re-judge the 10 cases whose search results reached a blocked source

The review found WebSearch results listing PolicyEngine and GitHub pages in
run 2's transcripts. Under the audit that now reads search results, 13 of
the 104 accepted outputs, in 10 cases, would have been rejected: 28 exposed
searches, mostly github.com/PolicyEngine and TheAxiomFoundation issues, and
one policyengine.org page.

A lane re-judged those 10 cases on 2026-10-06 (runs/claude-rejudge) with the
source rule that asks for blocked_domains on every search; it left the
evidence uncommitted. This records it:

- scripts/search_exposure.py runs the current transcript audit over both
  runs and writes verification/search_exposure.{json,md}, and rewrites
  runs/rejudge_flags.json (consensus_flags.json restricted to the 10 cells)
  byte-identically to the file the re-judge was prepared from.
- runs/collected-rejudge is adversary-collect's output for the re-judge:
  10 verdicts, 0 missing, 0 inconsistent, Colorado 043 queued.
- runs/run_record.json records the re-judge's lane, login, settings, times
  and cost ($24.82 reported, 4.70 judge-hours).
- The runner logs the README and verdict_table.py cite were git-ignored
  (*.log) and never committed; they are now, with the re-judge's.

Outcome: 0 of the 20 re-judge transcripts flagged; verdict unchanged in all
10. One stage 1 moved: scenario_013 SNAP now finds $288, and stage 2 still
holds the $240 reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Report the search exposure and re-judge, and fix the report review's findings

README:
- A section on search results from blocked sources: what the audit missed,
  the 13 flagged transcripts in 10 cases by source, the re-judge and its
  outcome (every verdict unchanged), and scenario_013 SNAP, whose blind
  stage 1 could not date Arizona's 200% limit and found $288.
- Counts with the re-judged outputs in place (verdicts unchanged; 39 high
  and 13 medium; stage 1 41/6/3/2), the re-judge's cost, files and
  reproduce steps.
- Report review: 35 amount and 17 eligibility derivations, not 42 and 10;
  score ranks shift more than exact ranks and within-1% ranks less (no 5/10%
  ranks exist); d1022's bullets keep to the ruling, and the pass's own
  expectations move out of them; the AZ filing-status leaves end at their
  2025 values, not all at $15,750.
- Open items: the four engine fixes' policyengine-us PRs (Ohio's premiums
  fix has none yet), scenario_013's effective date, and the Codex runner's
  blindness to search results.

docs/audit.md and the Claude runner header now say that WebSearch has no
deny rule, that the audit reads its results, and that the Codex runner
cannot; and describe the consensus rule's exact-member exclusion and the
collect command's queue, sidecar binding, labels and --allow-missing.

search_exposure.py writes runs/rejudge_flags.json before the re-judge
exists, so the reproduce steps run in order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run: 376 passed at 10b3d16

The six test files (the five adversary files plus tests/test_audit.py),
with ruff check and ruff format clean at the locked ruff 0.15.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Fix the round-2 review's findings on the reference adversary

The independent review of a29cb3c found four gaps:

- The Codex runner refused a Codex home holding AGENTS.md but not
  AGENTS.override.md, which Codex loads first. It now refuses either, and
  the refusal test covers both.
- check_login compared AUDIT_ACCOUNT to the desktop login's email verbatim,
  so a lane declaring "claude:<desktop email>" passed. The comparison now
  drops a provider prefix; a prefixed regression covers it.
- collect_adversary and adversary-collect accepted a directory with no
  cases.jsonl and collected it as a judge with no cases. Both now refuse it.
- adversary-collect split LABEL=DIR at the last "=", so a directory
  containing "=" broke. It now splits at the first "=" when the text before
  it is a label, and keeps a bare DIR whole otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at b46c4a5 and document the round-2 fixes

pytest_final.txt: the six test files (the five adversary files plus
tests/test_audit.py), run one file at a time on a loaded machine: 379
passed, ruff check and ruff format clean at the locked ruff 0.15.2.

docs/audit.md: adversary-collect splits LABEL=DIR at the first "=" and
refuses a directory without cases.jsonl; the Codex runner refuses a Codex
home holding AGENTS.override.md or AGENTS.md. A missing blank line had
folded the paragraph after the collect rules into the last list item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Fix the round-3 review's findings on the reference adversary

The independent review of a6323f8 found one major gap and five smaller
ones:

- The Codex runner checked only the AGENTS files, but Codex adds other text
  to a session without a tool event: config.toml's developer_instructions
  and model_instructions_file, memories, and the skills under
  $HOME/.agents/skills. Each judge call now skips the lane's config.toml
  (--ignore-user-config; auth still comes from CODEX_HOME), runs with
  memories off (--disable memories), and gets a fresh, empty HOME, removed
  afterwards (the login check too). CODEX_HOME is always passed, defaulting
  to ~/.codex. The runner also refuses a Codex home holding any skill but
  the bundled .system ones. The header, docs/audit.md and the README name
  what it still cannot control: what Codex bundles, an administrator's
  /etc/codex, and apps or plugins on the ChatGPT account (whose calls are
  MCP tool events, which the audit rejects). No committed run used this
  runner.
- check_login refused an empty AUDIT_ACCOUNT but accepted "claude:" or
  blank space, which declare no account; it now tests the parsed account.
  Its docstring says it is stricter than run_audit_claude.sh, not the same.
- adversary-collect compared judge labels case-sensitively, so claude and
  Claude shared files on a case-insensitive file system, and accepted the
  label "merged", whose files are the merged table's. Both are refused, as
  is an empty DIR (claude=). The docs and help say to pass a relative bare
  DIR containing "=" as ./adv=2, and a test covers it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at bc5925e: 386 passed

The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The round-3 fixes add four tests to
test_reference_adversary.py and three to the runner tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read the pass's payload from its commit, and fix the round-4 review's notes

CI on 8f968e8 failed six tests once main carried release
dashboard-data-20261006 (#202): they read the working tree's payload,
which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran
on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus
and frozen-run tests now read the payload and the reference explanations
as 8b4c0ca (#187) committed them, and the README says the Reproduce steps
need those inputs.

The round-4 review approved 8f968e8 with two notes:

- The Codex runner's skills check parsed `ls` output, so an entry named a
  bare newline, or ".system\n.system", passed it. It now walks the
  directory with globs and refuses every entry but a real .system
  directory (a symlink named .system too), naming it with %q; tests cover
  both names and the symlink.
- The AUDIT_MODEL default said "the lane's"; with the lane's config.toml
  skipped it is Codex's own.

The review left hooks as an unconfirmed channel. Each Codex call now also
turns the hooks, plugins and apps features off, as codex-cli 0.159.0's
`codex features list` names them, beside memories; the docs name what
remains outside the runner's control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the final test run at 4a62b1e: 391 passed

The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The skills-check tests now cover five
names and a symlinked .system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 5e9ee75f Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant