Repository navigation
Conversation
…a disregard The case note, reference explanation and 43 row annotations for scenario_031 head_medicaid_eligible explained the reference (eligible) by subtracting a Medicare Part B premium from countable income. policyengine-us 2.15.17 subtracts no premium. It counts $23,853.47 of SSI unearned income and applies California's $230 monthly disregard in place of SSI's $20 exclusion. That gives $21,093.47 against $22,024.80 (138% of the 2026 guideline). Wording-only: every row stays llm_error with its subtype. The ledger (reference_audit/2026-10-05-medicaid-031-annotations/rewrites.json) is applied by scripts/apply_rewrites.py, which re-pins the manifest hashes, and tests/test_annotation_rewrites.py keeps it in force. engine_values.py asserts every figure on the reference system. payload_diff.py shows a re-export changes only this cell's text fields. The frozen run dir is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Local run on |
…_114 row check Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions Narrows 'holds only if the disregard is skipped' to the engine's limit, names gemini-3.5-flash's source for $23,853, drops 'about' before exact figures, rewords claude-fable-5, the case note's category clause and the explanation's limit phrase. Records immigration_status in engine_values.json and tightens payload_diff's reproduction check. Adds the review as verification/independent_review.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…last answer (release dashboard-data-20261006) (#202) * Date the model response window from the last answer, not the release date The release drivers set the manifest's model_response_date to "2026-06-12 to <release date>". Release dashboard-data-20260930 recorded an end of 2026-09-30, but the last answer on that board (GPT-6.1 Sol) completed at 2026-09-29 22:36:32 UTC and no model answered on the 30th. freeze_snapshot.model_response_window now derives the end as the UTC date of the latest request_completed_at in the predictions the freeze gzips into the snapshot. main() derives it before its first write and passes it to build_manifest. The start stays configured (MODEL_RESPONSE_START), because the first waves' rows carry no timestamps and the run label gives their date. A driver that still sets MODEL_RESPONSE_DATE must name the derived window, or the freeze refuses. Both drivers stop setting it, and freeze_gpt61sol derives the window in its preflight so --dry-run reports it. The live manifest is not changed here: correcting it is a published-claim change queued for Max (cos decision d831). A strict xfail test records the live mismatch and fails once a freeze corrects the manifest, so that freeze must remove the mark. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Audit the state income tax withheld in federal SALT; propose three exclusions policyengine-us 2.15.17 fills the state income tax part of the federal SALT deduction with state_withheld_income_tax, an AGI-formula estimate of withholding, and no prompt states state income tax withheld or paid. This adds the audit (reference_audit/2026-10-05): a sweep of all 1,984 outputs under the reference, liability, net and zero readings (baseline reproduces all 1,928 scored references), the side readings the records quote, three proposed reference_depends_on_unlisted_input records (022 CA, 081 MA, 114 VA federal income tax), and the leaderboard impact scored through the analyze CLI on scratch copies. No reference, exclusion record or published number changes here. The exclusions change published scores and wait for Max's ruling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Audit Louisiana's 2026 standard deduction behind two scored references scenario_051 and scenario_077 state income tax rest on policyengine-us's explicit 2026 Louisiana standard deduction of $12,835 (PR #8411), which no Louisiana publication states. The audit sources what Louisiana published, sweeps all 1,984 outputs under six candidate deductions on policyengine-us 2.15.17 with latest_final, and scores each option through the analyze CLI. It changes no reference. Before the 2026-07-03 freeze the Department of Revenue published only a provisional $12,875 (withholding tables and estimated-tax worksheet), saying the return amount would differ; RIB 26-019 (2026-09-28) set it at $12,838, which moves the references by $0.09 and no exact match. The audit recommends keeping both references and drafts the documentation; the ruling is Max's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the sweep, scoring and check logs the audit README cites Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Audit optional employer pass-through in the payroll tax references; draft both records Four scored payroll_tax references (scenario_032 MN, 043 CO, 081 MA, 082 NY) count an employee share of a state paid-leave premium that the law lets the employer deduct but does not require. policyengine-us 2.15.17 counts the largest share the employer may deduct. The output asks for "mandatory employee state payroll taxes". reference_audit/2026-10-05-payroll classifies every state program in the payroll references from primary law (two adversarial reviews per program, none refuted), sweeps all 1,984 outputs with an output-scope adapter that drops the optional shares (four outputs move, nothing else), and drafts both candidate records: proposed_exclusions.json and proposed_regenerations.json. It scores each through the analyze CLI; the unchanged copy reproduces the published payload. Changes no reference or published number. Draft README; review pending. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Audit the Medicare Part B premium in medical expenses; propose excluding scenario_114's Virginia income tax policyengine-us 2.15.17 treats every Medicare-eligible person as enrolled (takes_up_medicare_if_eligible defaults to true) and adds the modeled Part B premium, $2,434.80 for 2026, to medical expenses. No prompt states Medicare enrollment or a premium. A sweep of all 1,984 outputs on the reference system with the premium at 0, and with enrollment off, moves exactly two: scenario_114 federal (10,729.61 -> 11,265.27, already proposed in #191) and Virginia (3,514.15 -> 3,654.15). No SNAP or eligibility output moves by any amount. Adds reference_audit/2026-10-05-medicare-part-b: the sweep, a per-household explanation, the drafted records (one for the Virginia output, and a fallback for the federal output if #191's record is not adopted) and the leaderboard impact under each combination with #191. Changes no reference or published number; the exclusion waits for Max's ruling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rest the payroll recommendation on the reference rules; record the independent review An eight-agent review of 8714770 (differential impact and sweep, mechanism, advocates for keep and regenerate with a judge, release checklist, claim audit) agreed on every number: 2,438 impact values, all 1,984 swept outputs bit for bit, and the 36 hand-computed payroll parts. The judge upheld exclusion but struck the draft's board-effect and model-behaviour arguments, which no rule uses. - README: the recommendation now rests on 2026-09-22/09-28 rule 4 and the r24, r25 and r14 precedents; regeneration is the strict reading of "mandatory". Adds both options' board effects side by side, the counter-evidence the reviewers found, the Acts 2026 c. 101 line mapping, the review summary and the release checklist. Fixes the claim audit's five errors and thirteen imprecisions. - proposed_exclusions.json: one shared unlisted input; the basis cites the rules and precedents (no engine input exists for the employer's choice). - payroll_mandatory_scope.py: docstring states its scope (DE, ME, VT and WA left for upstream). Sweep, proposals and impact rerun on the new hash; values unchanged. - verification/independent_reviews.json: the eight reports. Changes no reference or published number. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the payroll decision (cos d972) in the audit README Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Answer the take-up reading in the Part B records; add the legacy impact summary and independent reviews - State the prompt's "assume program take-up" sentence in both records and the README, and why it does not settle Medicare enrollment (the benchmark's DEFAULT_TAKEUP_INPUTS has no Medicare entry; the request asks eligibility only). - Name both ACA coverage tests in the mechanism section. - leaderboard_impact.py recomputes the freeze's legacy impact_summary_by_model.csv in every case (the unchanged copy reproduces the frozen file). The stale-weight check now reports that file, which reorders 10 of 46 models with 2.15.17's weights, so the finding is narrowed to "no change in the analyze payload". - Read #191's records from a sha256-checked copy at its head 8af912a, so this directory does not depend on #191's branch. - Add the four workflow reviewers' reports (verification/independent_reviews.json). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Correct scenario_031's Medicaid annotations to the engine's California disregard The case note, reference explanation and 43 row annotations for scenario_031 head_medicaid_eligible explained the reference (eligible) by subtracting a Medicare Part B premium from countable income. policyengine-us 2.15.17 subtracts no premium. It counts $23,853.47 of SSI unearned income and applies California's $230 monthly disregard in place of SSI's $20 exclusion. That gives $21,093.47 against $22,024.80 (138% of the 2026 guideline). Wording-only: every row stays llm_error with its subtype. The ledger (reference_audit/2026-10-05-medicaid-031-annotations/rewrites.json) is applied by scripts/apply_rewrites.py, which re-pins the manifest hashes, and tests/test_annotation_rewrites.py keeps it in force. engine_values.py asserts every figure on the reference system. payload_diff.py shows a re-export changes only this cell's text fields. The frozen run dir is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Point the Part B audit's annotation findings to #197 and the scenario_114 row check Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Apply the independent review's wording fixes to scenario_031's annotations Narrows 'holds only if the disregard is skipped' to the engine's limit, names gemini-3.5-flash's source for $23,853, drops 'about' before exact figures, rewords claude-fable-5, the case note's category clause and the explanation's limit phrase. Records immigration_status in engine_values.json and tightens payload_diff's reproduction check. Adds the review as verification/independent_review.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep CSV timestamps exact and test the real manifest builder's window Answers the hard-tier review of #188 (b6b385c): - model_response_window parses with float_precision=round_trip; pandas' default parser rounded the last double before 2027-02-01 UTC up to midnight and dated a January 31 answer February 1. New end-to-end CSV regression. - A test now calls the real build_manifest with a window no release has used and requires it in model_response_date and both reproducibility notes; a builder hard-coded to September 30 fails it (mutation f). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the release 20261006 driver, its spec and its tests scripts/release_20261006.py stages, exports and freezes release dashboard-data-20261006 from release 20260930 (read from git at 8b4c0ca): the eight ruled exclusion records and their adjudications, the record edits the three audits wrote, the annotation rewrites, and the derived response window. docs/release_20261006/spec.json holds its inputs. Gates: the base stage is 20260930's; the annotations differ from 20260930's only on the excluded and reworded outputs; two exports agree byte for byte; the payload differs from 20260930's only where the release allows; rebuilt with 20260930's exclusion record, the staged bundle reproduces 20260930's statistics exactly; the freeze leaves references, predictions and serving configuration as they were and changes the manifest only where the release allows. tests/test_release_20261006.py rebuilds each committed record from git and the spec, rescores the payload's rows independently, and checks the gates with Hypothesis properties. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Date the 2026-10-05 wave's decisions in UTC, document the driver, widen the in-force test - spec: the new wave's adjudications were written on 2026-10-06 UTC; - DATE_CONVENTIONS_NEW names Max's rulings of 2026-10-05 US Eastern time; - docs/release_20261006/design.md documents the driver's inputs, gates and order; - the in-force test strips any adjudication sentence before re-appending it; - rescore_sensitivity_summaries knows the word for 64. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * WIP: partial content edits written against the trial freeze Test pins, paper and docs prose, and a draft release note from the stopped content pass. Some are reviewed and some are not; the content pass after the fresh freeze finishes them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the release's annotation rewrite ledger and its evidence 306 whole-field rewrites in four groups, each drafted from its audit's evidence, checked by an adversarial verifier and revised: - salt (81): scenario_022 and scenario_081 federal income tax; the SALT state income tax is the engine's withholding estimate, not the household's; - s114 (67): scenario_114 federal and Virginia income tax; the withholding estimate and the imputed Medicare Part B premium; - payroll (151): the four paid-leave payroll outputs; optional employer pass-through shares, and scenario_032's and 043's note errors; - louisiana (7): scenario_051 and 077 case notes and rows; sourced $12,835 chain, no change to scored references. Each group's evidence file quotes the model explanations and engine values behind every changed number. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the d994 hold in the release spec * Freeze dashboard-data-20261006 and rescore sensitivity evidence * Pin the release note to the draft v2 contract and document the d994 hold * Render the paper, pin it, record d994's ruling and the validation logs Renders the paper outside the sandbox with the installed Quarto 1.9.36 and re-pins it with freeze_snapshot.py --rendered-only (pdf sha256 9403564b...). The Louisiana audit README now records d994 (ruled 2026-10-06, after this release was built): the convention is kept as written, and both Louisiana references are excluded in the release after this one. Validation logs: render exit 0; app lint, 170 bun tests and the build exit 0; full pytest 1,651 passed, 9 skipped, 1 failed. The failure is test_local_claude_models_resolve_without_remote_cost_map[claude-sonnet-5-5], which fails on main too once litellm loads its remote cost map. PR #201 fixes it, and this branch takes the fix from main before it merges. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the reviewed head Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the published rates behind Massachusetts's maximum; restore the pre-freeze tree The independent review of head 8b831f2 found that the release note and 32 annotation rewrites called $805.01 (0.46%) the most a Massachusetts employer may deduct. The payroll audit records a competing reading: a literal application of Acts 2026 c. 101 s. 45 would let a large employer deduct up to 0.772% ($1,351.02) for 2026, while the Department's rates page kept 0.46%. The texts now say "the most Massachusetts's published 2026 rates let the employer deduct": the ledger's scenario_081 payroll texts, its adjudication reasoning, a fifth record edit for the exclusion record, and the note. The exclusion and every score are unchanged. Also from the review: verify_payload now requires the staged payload's households and outputs to be the base's before it compares cells, with a test that a dropped cell or household is refused; the design note records d994's ruling instead of calling it pending. The freeze's outputs (5435bab) are restored to their pre-freeze state so prepare, export and freeze can run again on the corrected ledger. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Freeze dashboard-data-20261006 again on the corrected Massachusetts wording Payload f03441f64316997de843e82c27272442c794148a121c25e1343fbddbe38574a0, 127,224,818 bytes. Against the reviewed head 8b831f2 the freeze changes 31 row annotations, one case note, one adjudication's reasoning and one exclusion record's alternative reading, all for scenario_081's payroll output, and the hashes that follow. The analysis files and the sensitivity evidence are byte-identical, so no score moved. The paper is rendered and pinned again (pdf sha256 f3ccebc8...). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Point the effects record at the corrected payload; drop the session handoff drafts The summary CSV the effects record binds is byte-identical after the re-freeze, so every expected effect still holds; only the payload hash it names changes. The handoff drafts were one session's working notes, not repo documentation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the validation logs of the corrected freeze Full pytest on the merged head: 1,662 passed, 9 skipped, exit 0. App lint, 170 bun tests and the build: exit 0. Paper render: exit 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the reviewed head Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Qualify Massachusetts's caps in the two remaining exclusion-record sentences The round-2 review found two record sentences that still stated Massachusetts's limits without the published-rates qualification: scenario_081's federal-tax note ('the largest share the employer may deduct') and its payroll record, which attributed the 100% family and 40% medical caps directly to s. 6(c). Acts 2026 c. 101 ss. 25-26 and 45 can be read to swap those caps for 2026. The note now names the published 2026 rates, and a sixth record edit qualifies the caps the same way. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Restore the pre-freeze tree for the round-3 freeze Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Freeze dashboard-data-20261006 on the round-3 record edits Payload aa34e5c9ea926848dc7af460a98f17021bc42dae54a8015c9988a85e79a5d462, 127,225,036 bytes. Against round 2 only scenario_081's two exclusion-record sentences change; the analysis files and sensitivity evidence remain byte-identical to the first freeze, so no score moved. The release asset is replaced and verified against the pointer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the round-3 validation logs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the reviewed head Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
Shipped in release |
…no score changes) (#200) * Add the reference-adversary audit pass: consensus trigger, blind adversary prompt and runners, definition conformance, publication sources (WIP: adversary not yet run) Work from the reference-adversary session (task_8940406e), which stopped at the 2026-10-05 account cutoff before committing. 361 tests pass (test_consensus, test_definition_conformance, test_publication_sources, test_reference_adversary, test_reference_adversary_runner, test_audit). The consensus flags (61 of 1,928 cells), the definition-conformance report and the publication-source report are generated; the adversary judges have not run yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Point scenario_031's skip at #197 and record the rebased test run The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Run reference-adversary cases in a rolling pool, not fixed batches Both runners started AUDIT_PARALLEL cases and waited for all of them before starting more, so one slow case held every slot: the first live batch sat 15 minutes on one SNAP case while three slots idled. The runners now start the next case as soon as any running case finishes, never running more than AUDIT_PARALLEL at once (bash 3.2 has no wait -n, so they count running jobs with jobs -pr; a finished job not yet noticed only delays a start). The Claude runner still checks the stop flag before every start. test_claude_runner_refills_a_slot_without_exceeding_the_parallel_cap holds scenario_001's stage 1 until scenario_003's first call starts: a batch runner never releases it, the pool does, and the call timeline never shows more than two calls at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Run the reference adversary over the 52 consensus-flagged cells another audit does not cover Opus 5.5 at xhigh judged each case in two blind stages on Subfleet lane claude-9 (oauth token from the lane's keychain entry, an empty lane config dir, never the desktop login or an API key): 104 judge calls, all accepted, none contaminated, invalid or refused. Verdicts: 45 reference_holds, 6 reference_wrong (scenario_018 AZ, 025 OH, 026 child1 and child2 NC, 043 CO, 082 NY), 1 definition_mismatch (123 PA). adversary-collect reports 0 missing and 0 inconsistent, and queues 7 cases for developer adjudication. No verdict changes a score. runs/ holds every case's prompts, outputs, provenance sidecars and session transcripts, the runner logs, the collected tables and run_record.json. scripts/verdict_table.py renders verification/verdict_table.md and verdict_counts.json; scripts/engine_probe.py rebuilds a frozen-run household on the reference system for engine evidence (verification/probes/). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the independent verification of the seven non-holding adversary verdicts (WIP: lane stopped before the report and PR) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Write the reference-adversary report, record Max's rulings and the final test run README.md reports the pass: design, flag parameters, the 61 flagged / 9 skipped / 52 run cells, the verdicts, the independent verification of the seven that did not hold (four confirmed engine defects, one ambiguous definition scope over four cells, NC 026 refuted), leaderboard impact, the definition-conformance and publication-source findings, judge cost (104 calls, $91.22 reported by the CLI, 5.25 judge-hours) and the rulings. proposed_changes.json status now records d1022 (exclude the eight cells in the next release after dashboard-data-20261006; regenerate the four defect cells on fixed policyengine-us) and d994 (exclude Louisiana 051 and 077; keep the published-amounts convention). The root-cause records are unchanged. verification/pytest_final.txt: 362 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Cite d1029, the queued household-scope decision, in the report Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the diagnosis judge's template unchanged in this PR Dropping the stale "bugs were fixed before this run" sentence from the judge prompt changes all 674 seed prompts. The fold drivers carry a seed verdict only when its prompt re-renders byte-identically, so the next model addition would refuse or need a full re-judge. Restore the template and its test to origin/main's, and note in the report that the sentence's removal waits for a versioned judge template, in a separate PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address the code review of the reference adversary Brings over the review fixes a stopped lane left on reference-adversary-review-wip (4c1334e), with its failing test fixed and lint cleared: - Blinding: the Claude transcript audit now reads each WebSearch result and rejects an output whose search results list a blocked URL or name PolicyEngine or PolicyBench (WebSearch has no deny rule, so results reach the judge). The source rule asks the judge to pass blocked_domains on every search. - adversary-collect refuses a repeated judge label, queues every case with an inconsistent verdict whatever its class, and exits non-zero when a case has no usable verdict unless --allow-missing is given. - collect_adversary requires a verdict sidecar bound to the current stage 1. - Consensus: a member whose own answer is within the tolerance counts as exact and leaves the wrong cluster (consensus_flags.json and the prototype file reproduce byte-identically under the new rule). - check_login refuses a token lane that declares the desktop login's account. - Codex runner: an allowlisted environment, now with LC_CTYPE as the Claude runner has, and a refusal of a Codex home holding an AGENTS.md. - leaderboard_impact.py refuses a --scratch inside the repository and an unknown --only variant; build_proposals.py records each record's ruling (d1022), separates the pass's own expectations from the rulings, and states the Colorado basis as the verifier did. proposed_changes.json is its output, reproduced byte-identically. The runner test test_codex_runner_passes_only_allowlisted_variables failed on LC_CTYPE. The runner was right: with no LANG in the allowlisted environment, the fake codex (a Python script) sets LC_CTYPE=C.UTF-8 itself under PEP 538 locale coercion. The test now allows LC_CTYPE, which the runner also forwards for parity with the Claude runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Re-judge the 10 cases whose search results reached a blocked source The review found WebSearch results listing PolicyEngine and GitHub pages in run 2's transcripts. Under the audit that now reads search results, 13 of the 104 accepted outputs, in 10 cases, would have been rejected: 28 exposed searches, mostly github.com/PolicyEngine and TheAxiomFoundation issues, and one policyengine.org page. A lane re-judged those 10 cases on 2026-10-06 (runs/claude-rejudge) with the source rule that asks for blocked_domains on every search; it left the evidence uncommitted. This records it: - scripts/search_exposure.py runs the current transcript audit over both runs and writes verification/search_exposure.{json,md}, and rewrites runs/rejudge_flags.json (consensus_flags.json restricted to the 10 cells) byte-identically to the file the re-judge was prepared from. - runs/collected-rejudge is adversary-collect's output for the re-judge: 10 verdicts, 0 missing, 0 inconsistent, Colorado 043 queued. - runs/run_record.json records the re-judge's lane, login, settings, times and cost ($24.82 reported, 4.70 judge-hours). - The runner logs the README and verdict_table.py cite were git-ignored (*.log) and never committed; they are now, with the re-judge's. Outcome: 0 of the 20 re-judge transcripts flagged; verdict unchanged in all 10. One stage 1 moved: scenario_013 SNAP now finds $288, and stage 2 still holds the $240 reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Report the search exposure and re-judge, and fix the report review's findings README: - A section on search results from blocked sources: what the audit missed, the 13 flagged transcripts in 10 cases by source, the re-judge and its outcome (every verdict unchanged), and scenario_013 SNAP, whose blind stage 1 could not date Arizona's 200% limit and found $288. - Counts with the re-judged outputs in place (verdicts unchanged; 39 high and 13 medium; stage 1 41/6/3/2), the re-judge's cost, files and reproduce steps. - Report review: 35 amount and 17 eligibility derivations, not 42 and 10; score ranks shift more than exact ranks and within-1% ranks less (no 5/10% ranks exist); d1022's bullets keep to the ruling, and the pass's own expectations move out of them; the AZ filing-status leaves end at their 2025 values, not all at $15,750. - Open items: the four engine fixes' policyengine-us PRs (Ohio's premiums fix has none yet), scenario_013's effective date, and the Codex runner's blindness to search results. docs/audit.md and the Claude runner header now say that WebSearch has no deny rule, that the audit reads its results, and that the Codex runner cannot; and describe the consensus rule's exact-member exclusion and the collect command's queue, sidecar binding, labels and --allow-missing. search_exposure.py writes runs/rejudge_flags.json before the re-judge exists, so the reproduce steps run in order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the final test run: 376 passed at 10b3d16 The six test files (the five adversary files plus tests/test_audit.py), with ruff check and ruff format clean at the locked ruff 0.15.2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the round-2 review's findings on the reference adversary The independent review of a29cb3c found four gaps: - The Codex runner refused a Codex home holding AGENTS.md but not AGENTS.override.md, which Codex loads first. It now refuses either, and the refusal test covers both. - check_login compared AUDIT_ACCOUNT to the desktop login's email verbatim, so a lane declaring "claude:<desktop email>" passed. The comparison now drops a provider prefix; a prefixed regression covers it. - collect_adversary and adversary-collect accepted a directory with no cases.jsonl and collected it as a judge with no cases. Both now refuse it. - adversary-collect split LABEL=DIR at the last "=", so a directory containing "=" broke. It now splits at the first "=" when the text before it is a label, and keeps a bare DIR whole otherwise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the final test run at b46c4a5 and document the round-2 fixes pytest_final.txt: the six test files (the five adversary files plus tests/test_audit.py), run one file at a time on a loaded machine: 379 passed, ruff check and ruff format clean at the locked ruff 0.15.2. docs/audit.md: adversary-collect splits LABEL=DIR at the first "=" and refuses a directory without cases.jsonl; the Codex runner refuses a Codex home holding AGENTS.override.md or AGENTS.md. A missing blank line had folded the paragraph after the collect rules into the last list item. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the round-3 review's findings on the reference adversary The independent review of a6323f8 found one major gap and five smaller ones: - The Codex runner checked only the AGENTS files, but Codex adds other text to a session without a tool event: config.toml's developer_instructions and model_instructions_file, memories, and the skills under $HOME/.agents/skills. Each judge call now skips the lane's config.toml (--ignore-user-config; auth still comes from CODEX_HOME), runs with memories off (--disable memories), and gets a fresh, empty HOME, removed afterwards (the login check too). CODEX_HOME is always passed, defaulting to ~/.codex. The runner also refuses a Codex home holding any skill but the bundled .system ones. The header, docs/audit.md and the README name what it still cannot control: what Codex bundles, an administrator's /etc/codex, and apps or plugins on the ChatGPT account (whose calls are MCP tool events, which the audit rejects). No committed run used this runner. - check_login refused an empty AUDIT_ACCOUNT but accepted "claude:" or blank space, which declare no account; it now tests the parsed account. Its docstring says it is stricter than run_audit_claude.sh, not the same. - adversary-collect compared judge labels case-sensitively, so claude and Claude shared files on a case-insensitive file system, and accepted the label "merged", whose files are the merged table's. Both are refused, as is an empty DIR (claude=). The docs and help say to pass a relative bare DIR containing "=" as ./adv=2, and a test covers it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the final test run at bc5925e: 386 passed The six test files (the five adversary files plus tests/test_audit.py), run one at a time on a loaded machine, with ruff check and ruff format clean at the locked ruff 0.15.2. The round-3 fixes add four tests to test_reference_adversary.py and three to the runner tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read the pass's payload from its commit, and fix the round-4 review's notes CI on 8f968e8 failed six tests once main carried release dashboard-data-20261006 (#202): they read the working tree's payload, which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus and frozen-run tests now read the payload and the reference explanations as 8b4c0ca (#187) committed them, and the README says the Reproduce steps need those inputs. The round-4 review approved 8f968e8 with two notes: - The Codex runner's skills check parsed `ls` output, so an entry named a bare newline, or ".system\n.system", passed it. It now walks the directory with globs and refuses every entry but a real .system directory (a symlink named .system too), naming it with %q; tests cover both names and the symlink. - The AUDIT_MODEL default said "the lane's"; with the lane's config.toml skipped it is Codex's own. The review left hooks as an unconfirmed channel. Each Codex call now also turns the hooks, plugins and apps features off, as codex-cli 0.159.0's `codex features list` names them, beside memories; the docs name what remains outside the runner's control. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the final test run at 4a62b1e: 391 passed The six test files (the five adversary files plus tests/test_audit.py), run one at a time on a loaded machine, with ruff check and ruff format clean at the locked ruff 0.15.2. The skills-check tests now cover five names and a symlinked .system. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Summary
The case note, reference explanation and all 43 row annotations for scenario_031
head_medicaid_eligible(California, single head age 67, reference: eligible) explained the reference by subtracting a Medicare Part B premium from countable income. policyengine-us 2.15.17, the reference engine, subtracts no premium. This PR rewrites that text to the engine's actual mechanism.It is wording only:
The finding came from the Part B audit (#193). Details are in
reference_audit/2026-10-05-medicaid-031-annotations/README.md.What the engine does (read in 2.15.17 source this session)
SENIOR_OR_DISABLED.is_ssi_agedis age ≥ 65, so age alone meets the category condition, with no SSI receipt or disability needed.is_adult_for_medicaid_nfcrequires age 19 to 64 and no Medicare eligibility, so this head is outside it. It is the only formula invariables/gov/hhs/medicaidthat bars on Medicare.medicaid_optional_senior_or_disabled_countable_incomecounts SSI unearned income: $9,445.47 Social Security (in full), $13,608 pension and $800 IRA, for $23,853.47. The $1,165 of alimony paid is not subtracted.senior_or_disabled.income.disregard.individual['CA']is $230 a month. It is passed to_apply_ssi_exclusionsasgeneral_exclusion, so it replaces SSI's $20 exclusion: $23,853.47 − $2,760 = $21,093.47.The old text compared "about $21,400" with "about $21,600", which is 138% of the 2025 guideline. Gross income is $1,828.67 over the real limit, and gross less only SSI's $20 exclusion is still over it. So the models that skipped California's disregard still erred.
Root cause. The judge prompt behind the published text stated the wrong mechanism twice:
grounding.csv, pinned byGROUNDING_SHA256): "a non-MAGI pathway whose income counting deducts health insurance premiums (including Medicare Part B)";No other case's grounding carries that phrase. This PR fixes the explanation. The grounding line should be regenerated before the next audit stage that re-judges this case.
Changes
rewrites.json: 45 whole-field rewrites (1 case note, 1 reference explanation, 43 rows).scripts/apply_rewrites.py:paper/snapshot/20260501/manifest.json;--check, writes nothing and confirms every rewrite is in force.scripts/engine_values.py→verification/engine_values.json: recomputes every figure on 2.15.17 +latest_finaland asserts each relation the text states, under the modeled premium,no_part_bandnot_enrolled.scripts/payload_diff.py→verification/payload_diff.json: rebuilds the payload from the frozen run withexport_country, once with main's annotations and once with these.verification/row_review.json: per-row review record (quoted model evidence, problems found, edits after review).tests/test_annotation_rewrites.py: keeps the ledger in force, so a release rebuilding these files from an older base can't silently restore the old text.Invariants (stated and checked)
origin/main, exactly 45 fields differ, all on scenario_031head_medicaid_eligible. Keys, order, quoting,failure_source,failure_subtypeandreference_suspectare identical (checked by a field-by-field diff;test_rewritten_rows_keep_their_classes).annotation, 43caseAnnotationand 46referenceExplanation. NomodelStats,programStats,heatmapor other cell value moves.data.json.gz, except claude-fable-5's four usage fields, which the drivers carry (CARRIED_USAGE).engine_values.pyasserts each of these:Not changed here, and why
payload_diff.jsonshows exactly what a re-export will change.Coordination with release tooling
test_the_committed_reference_explanations_are_20260929s, which pins the committedus_case_reference_explanations.csvat HEAD to release 20260929's bytes. Whichever of Harden the GPT-6.1 Sol driver: pin kept sidecars, reference explanations, re-opened flags and the export receipt #190 and this PR merges second will see that test fail until it is scoped. Suggested scoping: check the blob at the release commit, or allow rewrites listed in this ledger. Commented on Harden the GPT-6.1 Sol driver: pin kept sidecars, reference explanations, re-opened flags and the export receipt #190.judge-isolation-20260930b, no PR yet) builds from release 20260930's committed annotation bytes. Its freeze refuses a working tree that differs from them, so it fails loudly rather than reverting this. It will need to treat this ledger as part of its base. The owning session has been told.Tests
pytest -m "not slow": see the comment below for the run on this head; the baseline on origin/main was 1563 passed, 8 skipped.ruff check .andruff format --check .are clean.apply_rewrites.py --check: 45 rewrites in force.Review
An independent Opus 5.5 Subfleet review is running against
19d28a54. Its report will be added asverification/independent_review.md.🤖 Generated with Claude Code