Skip to content

Add the reference adversary and record its first run (evidence only; no score changes) - #200

Merged
MaxGhenis merged 19 commits into
mainfrom
reference-adversary
Oct 9, 2026
Merged

MaxGhenis merged 19 commits into
mainfrom
reference-adversary

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Adds the reference adversary, a pass whose only job is to attack PolicyBench's US references, and records its first run on the frozen 20260612 run (46 models, 1,928 scored cells). Evidence and tooling only: no reference, exclusion record or score changes. The full report is reference_audit/2026-10-05-reference-adversary/README.md.

What it found

  • Flagged and run. The consensus trigger flagged 61 cells: a wrong cluster of at least 15 models, or 3 of the top 5. Another audit owns 9 of them (Audit Louisiana's 2026 standard deduction behind two scored references #192 Louisiana, Audit optional employer pass-through in the payroll tax references; propose excluding four outputs #194 payroll, Correct scenario_031's Medicaid annotations to the engine's California disregard #197 scenario_031), so 52 were run.
  • Verdicts. The blind, law-first, two-stage adversary returned 45 reference_holds, 6 reference_wrong and 1 definition_mismatch.
  • Search exposure and re-judge. The code review found that WebSearch results listing PolicyEngine and GitHub pages reached the judge. WebFetch had a deny rule, but WebSearch did not, and the run's audit read only the queries.
    • The audit now reads every search result. Run over the original transcripts, it flags 13 of 104, in 10 cases.
    • Those 10 were re-judged, with the prompt asking for blocked_domains on every search. None of the 20 new transcripts is flagged, and every verdict is unchanged.
    • One stage 1 moved. scenario_013 SNAP (Arizona) can no longer date the 185% → 200% eligibility change behind the $240 reference. Blind, it finds $288; stage 2 still holds $240. This is listed as an open item (verification/search_exposure.md).
  • Independent verification of the 7 (verification/independent/):
    • CONFIRMED engine defects, all still on policyengine-us main 2.29.11:

      Cell Defect Reference Law
      AZ 018 Standard deduction not indexed under A.R.S. 43-1041(H) $1,146.05 $1,137.30
      OH 025 R.C. 5747.01(A)(10)(b) premiums dropped $1,921.57 $1,916.60
      CO 043 Sales tax refund paid without the FY 2025-26 surplus condition $19 $0
      NY 082 Pre-2026 child care credit instead of Tax Law 606(c-2) $667 $1,187.61
    • AMBIGUOUS: PA 123. The definitions don't say whether a dependent's own required return counts. The same question reaches 4 cells: PA 123 and MO 093, state and federal.

    • REFUTED: NC 026 child1/child2. SPA NC-23-0009's 42 CFR 435.218 election covers insured children, so both references stand.

  • Leaderboard impact (verification/leaderboard_impact.json, scored with policybench analyze). Excluding the 8 cells raises every model's exact rate by 0.40–0.65 pp and swaps claude-sonnet-5.5 and gpt-6-luna at ranks 4 and 5.
  • Engine-side checks.
  • Judge cost (reported by the CLI as API-price estimates; the calls billed subscription lanes):
    • the run: 104 calls, $91.22, 5.25 judge-hours;
    • the re-judge: 20 calls, $24.82, 4.70 judge-hours.

Rulings (recorded in proposed_changes.json status)

  • d1022 (Max, 2026-10-06): exclude the 8 cells in the next release after dashboard-data-20261006, then regenerate the four defect cells once fixed policyengine-us versions land. The fixes are open policyengine-us PRs #9928 (AZ), #9946 (CO) and #9948 (NY). For Ohio, #9925 fixes a double count, and the premiums fix has no PR yet.
  • d994 (Max, 2026-10-06): exclude Louisiana 051 and 077, and keep the published-amounts convention.

This PR installs neither ruling; the next release does.

Review fixes since the first review

The independent code review requested changes (blinding: search results). The report review found count and wording errors. Both are addressed:

  • Blinding. claude_transcript_audit rejects a search whose result lists a blocked URL or names PolicyEngine or PolicyBench. The source rule asks for blocked_domains on every search. Docs and both runner headers now say the Codex runner cannot see search results.
  • Collect. adversary-collect:
    • refuses a repeated judge label;
    • queues every case with an inconsistent verdict, whatever its class;
    • requires a verdict sidecar bound to the current stage 1;
    • exits non-zero on missing verdicts unless --allow-missing is given.
  • Consensus. A model whose own answer is within the tolerance counts as exact and leaves the wrong cluster. consensus_flags.json and the prototype file reproduce byte-identically under the new rule.
  • Credentials.
    • Codex runner: an allowlisted environment, with the Claude runner's locale set, and it refuses a Codex home holding an AGENTS.md.
    • check_login refuses a token lane that declares the desktop account.
  • Scripts.
    • leaderboard_impact.py refuses a --scratch inside the repo and an unknown --only variant.
    • build_proposals.py records each record's ruling (decision: d1022) and keeps the pass's own expectations out of the rulings. It also states the Colorado basis as the verifier did; proposed_changes.json reproduces byte-identically.
  • Report.
    • The derivations are 35 amount and 17 eligibility outputs, not 42 and 10.
    • The rank-shift text is corrected: no 5/10% ranks exist.
    • The d1022 bullets now keep to the ruling.
    • The AZ filing-status values are fixed.
  • Evidence. runs/collected/ reproduces byte-identically under the new collect. The runner logs the README cites were git-ignored and are now committed.
  • Diagnosis judge template. It is unchanged here (738d6a6). Dropping its stale "bugs were fixed before this run" sentence waits for a versioned judge template in a separate PR, so seed verdicts keep re-rendering byte-identically.

Invariants

The tests check each of these, with Hypothesis property tests where marked:

  • Consensus trigger (tests/test_consensus.py):
    • its flags match a brute-force recount, which applies the exact-member rule (property);
    • no member of a wrong cluster is within the tolerance of an amount reference (property);
    • raising any threshold never adds a flag (property);
    • prediction order does not matter (property);
    • the report is deterministic and leaves the payload untouched (property).
  • Adversary (tests/test_reference_adversary.py):
    • a stage-1 prompt never contains the derivation (property);
    • apply_adversary_flags changes nothing but the flags it sets (property);
    • the adjudication queue holds every non-holding verdict and every inconsistent one;
    • a verdict counts only with a sidecar bound to the current stage 1;
    • a changed stage-1 prompt invalidates everything downstream;
    • the audit rejects a search result that reaches a blocked source, and passes a clean one.
  • Definition conformance: the component tree is deterministic and independent of graph order, and every node's signed path multiplies correctly (properties). On the real reference system, all 1,928 scored references reproduce within 1e-3, and 8,128 node-sum checks have 0 failures.
  • Publication sources: citation dates are never after the text's stated years, and classification is total and deterministic (properties).
  • Runner (tests/test_reference_adversary_runner.py, with fake CLIs):
    • blinding;
    • transcript rejection;
    • credential refusal;
    • the Codex runner's allowlisted environment and its AGENTS.md refusal;
    • a rolling pool that never exceeds its parallel cap.

Tests

reference_audit/2026-10-05-reference-adversary/verification/pytest_final.txt holds the final run:

  • Files: the five adversary test files plus tests/test_audit.py.
  • Result: 376 passed, at 10b3d16. The later commit a29cb3c changes only that file and the README.
  • ruff check . and ruff format --check . pass with the locked ruff 0.15.2.

The previously failing test_codex_runner_passes_only_allowlisted_variables was a test bug. With no LANG, the fake codex, a Python script, sets LC_CTYPE itself under PEP 538, so the runner was right.

axiom: n/a: PolicyBench evidence and tooling; the engine defects are fixed in policyengine-us under d1022.

🤖 Generated with Claude Code

MaxGhenis and others added 6 commits October 5, 2026 21:38
…rsary prompt and runners, definition conformance, publication sources (WIP: adversary not yet run)

Work from the reference-adversary session (task_8940406e), which stopped at the 2026-10-05 account cutoff before committing. 361 tests pass (test_consensus, test_definition_conformance, test_publication_sources, test_reference_adversary, test_reference_adversary_runner, test_audit). The consensus flags (61 of 1,928 cells), the definition-conformance report and the publication-source report are generated; the adversary judges have not run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both runners started AUDIT_PARALLEL cases and waited for all of them before starting more, so one slow case held every slot: the first live batch sat 15 minutes on one SNAP case while three slots idled. The runners now start the next case as soon as any running case finishes, never running more than AUDIT_PARALLEL at once (bash 3.2 has no wait -n, so they count running jobs with jobs -pr; a finished job not yet noticed only delays a start). The Claude runner still checks the stop flag before every start.

test_claude_runner_refills_a_slot_without_exceeding_the_parallel_cap holds scenario_001's stage 1 until scenario_003's first call starts: a batch runner never releases it, the pool does, and the call timeline never shows more than two calls at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er audit does not cover

Opus 5.5 at xhigh judged each case in two blind stages on Subfleet lane claude-9 (oauth token from the lane's keychain entry, an empty lane config dir, never the desktop login or an API key): 104 judge calls, all accepted, none contaminated, invalid or refused. Verdicts: 45 reference_holds, 6 reference_wrong (scenario_018 AZ, 025 OH, 026 child1 and child2 NC, 043 CO, 082 NY), 1 definition_mismatch (123 PA). adversary-collect reports 0 missing and 0 inconsistent, and queues 7 cases for developer adjudication. No verdict changes a score.

runs/ holds every case's prompts, outputs, provenance sidecars and session transcripts, the runner logs, the collected tables and run_record.json. scripts/verdict_table.py renders verification/verdict_table.md and verdict_counts.json; scripts/engine_probe.py rebuilds a frozen-run household on the reference system for engine evidence (verification/probes/).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y verdicts (WIP: lane stopped before the report and PR)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nal test run

README.md reports the pass: design, flag parameters, the 61 flagged / 9 skipped / 52 run
cells, the verdicts, the independent verification of the seven that did not hold (four
confirmed engine defects, one ambiguous definition scope over four cells, NC 026 refuted),
leaderboard impact, the definition-conformance and publication-source findings, judge cost
(104 calls, $91.22 reported by the CLI, 5.25 judge-hours) and the rulings.

proposed_changes.json status now records d1022 (exclude the eight cells in the next
release after dashboard-data-20261006; regenerate the four defect cells on fixed
policyengine-us) and d994 (exclude Louisiana 051 and 077; keep the published-amounts
convention). The root-cause records are unchanged. verification/pytest_final.txt: 362
passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
policybench-site Ready Ready Preview Oct 9, 2026 3:29pm UTC

Request Review

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dropping the stale "bugs were fixed before this run" sentence from the judge
prompt changes all 674 seed prompts. The fold drivers carry a seed verdict
only when its prompt re-renders byte-identically, so the next model addition
would refuse or need a full re-judge. Restore the template and its test to
origin/main's, and note in the report that the sentence's removal waits for a
versioned judge template, in a separate PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Removed the judge-template edit from this PR (commit 738d6a6). Dropping the stale sentence changes all 674 seed prompts and breaks byte-identical verdict carry-over at the next model addition (flagged by the release session building 20261006 and Haiku 5.5). The sentence's removal moves to a separate PR that versions the judge template. Review fixes left uncommitted by a stopped lane are parked on reference-adversary-review-wip: one runner test fails and lint is pending. A queued lane finishes them before this PR merges.

MaxGhenis and others added 4 commits October 8, 2026 18:50
Brings over the review fixes a stopped lane left on
reference-adversary-review-wip (4c1334e), with its failing test fixed and
lint cleared:

- Blinding: the Claude transcript audit now reads each WebSearch result and
  rejects an output whose search results list a blocked URL or name
  PolicyEngine or PolicyBench (WebSearch has no deny rule, so results reach
  the judge). The source rule asks the judge to pass blocked_domains on
  every search.
- adversary-collect refuses a repeated judge label, queues every case with
  an inconsistent verdict whatever its class, and exits non-zero when a case
  has no usable verdict unless --allow-missing is given.
- collect_adversary requires a verdict sidecar bound to the current stage 1.
- Consensus: a member whose own answer is within the tolerance counts as
  exact and leaves the wrong cluster (consensus_flags.json and the prototype
  file reproduce byte-identically under the new rule).
- check_login refuses a token lane that declares the desktop login's account.
- Codex runner: an allowlisted environment, now with LC_CTYPE as the Claude
  runner has, and a refusal of a Codex home holding an AGENTS.md.
- leaderboard_impact.py refuses a --scratch inside the repository and an
  unknown --only variant; build_proposals.py records each record's ruling
  (d1022), separates the pass's own expectations from the rulings, and
  states the Colorado basis as the verifier did. proposed_changes.json is
  its output, reproduced byte-identically.

The runner test test_codex_runner_passes_only_allowlisted_variables failed
on LC_CTYPE. The runner was right: with no LANG in the allowlisted
environment, the fake codex (a Python script) sets LC_CTYPE=C.UTF-8 itself
under PEP 538 locale coercion. The test now allows LC_CTYPE, which the
runner also forwards for parity with the Claude runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The review found WebSearch results listing PolicyEngine and GitHub pages in
run 2's transcripts. Under the audit that now reads search results, 13 of
the 104 accepted outputs, in 10 cases, would have been rejected: 28 exposed
searches, mostly github.com/PolicyEngine and TheAxiomFoundation issues, and
one policyengine.org page.

A lane re-judged those 10 cases on 2026-10-06 (runs/claude-rejudge) with the
source rule that asks for blocked_domains on every search; it left the
evidence uncommitted. This records it:

- scripts/search_exposure.py runs the current transcript audit over both
  runs and writes verification/search_exposure.{json,md}, and rewrites
  runs/rejudge_flags.json (consensus_flags.json restricted to the 10 cells)
  byte-identically to the file the re-judge was prepared from.
- runs/collected-rejudge is adversary-collect's output for the re-judge:
  10 verdicts, 0 missing, 0 inconsistent, Colorado 043 queued.
- runs/run_record.json records the re-judge's lane, login, settings, times
  and cost ($24.82 reported, 4.70 judge-hours).
- The runner logs the README and verdict_table.py cite were git-ignored
  (*.log) and never committed; they are now, with the re-judge's.

Outcome: 0 of the 20 re-judge transcripts flagged; verdict unchanged in all
10. One stage 1 moved: scenario_013 SNAP now finds $288, and stage 2 still
holds the $240 reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…findings

README:
- A section on search results from blocked sources: what the audit missed,
  the 13 flagged transcripts in 10 cases by source, the re-judge and its
  outcome (every verdict unchanged), and scenario_013 SNAP, whose blind
  stage 1 could not date Arizona's 200% limit and found $288.
- Counts with the re-judged outputs in place (verdicts unchanged; 39 high
  and 13 medium; stage 1 41/6/3/2), the re-judge's cost, files and
  reproduce steps.
- Report review: 35 amount and 17 eligibility derivations, not 42 and 10;
  score ranks shift more than exact ranks and within-1% ranks less (no 5/10%
  ranks exist); d1022's bullets keep to the ruling, and the pass's own
  expectations move out of them; the AZ filing-status leaves end at their
  2025 values, not all at $15,750.
- Open items: the four engine fixes' policyengine-us PRs (Ohio's premiums
  fix has none yet), scenario_013's effective date, and the Codex runner's
  blindness to search results.

docs/audit.md and the Claude runner header now say that WebSearch has no
deny rule, that the audit reads its results, and that the Codex runner
cannot; and describe the consensus rule's exact-member exclusion and the
collect command's queue, sidecar binding, labels and --allow-missing.

search_exposure.py writes runs/rejudge_flags.json before the re-judge
exists, so the reproduce steps run in order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py),
with ruff check and ruff format clean at the locked ruff 0.15.2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 2 commits October 8, 2026 20:17
The independent review of a29cb3c found four gaps:

- The Codex runner refused a Codex home holding AGENTS.md but not
  AGENTS.override.md, which Codex loads first. It now refuses either, and
  the refusal test covers both.
- check_login compared AUDIT_ACCOUNT to the desktop login's email verbatim,
  so a lane declaring "claude:<desktop email>" passed. The comparison now
  drops a provider prefix; a prefixed regression covers it.
- collect_adversary and adversary-collect accepted a directory with no
  cases.jsonl and collected it as a judge with no cases. Both now refuse it.
- adversary-collect split LABEL=DIR at the last "=", so a directory
  containing "=" broke. It now splits at the first "=" when the text before
  it is a label, and keeps a bare DIR whole otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in #201 (Claude Haiku 5.5's model card), whose corrected Sonnet 5.5
cache-read price fixes the test_eval_no_tools failure on this branch's CI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pytest_final.txt: the six test files (the five adversary files plus
tests/test_audit.py), run one file at a time on a loaded machine: 379
passed, ruff check and ruff format clean at the locked ruff 0.15.2.

docs/audit.md: adversary-collect splits LABEL=DIR at the first "=" and
refuses a directory without cases.jsonl; the Codex runner refuses a Codex
home holding AGENTS.override.md or AGENTS.md. A missing blank line had
folded the paragraph after the collect rules into the last list item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent review of a6323f8 found one major gap and five smaller
ones:

- The Codex runner checked only the AGENTS files, but Codex adds other text
  to a session without a tool event: config.toml's developer_instructions
  and model_instructions_file, memories, and the skills under
  $HOME/.agents/skills. Each judge call now skips the lane's config.toml
  (--ignore-user-config; auth still comes from CODEX_HOME), runs with
  memories off (--disable memories), and gets a fresh, empty HOME, removed
  afterwards (the login check too). CODEX_HOME is always passed, defaulting
  to ~/.codex. The runner also refuses a Codex home holding any skill but
  the bundled .system ones. The header, docs/audit.md and the README name
  what it still cannot control: what Codex bundles, an administrator's
  /etc/codex, and apps or plugins on the ChatGPT account (whose calls are
  MCP tool events, which the audit rejects). No committed run used this
  runner.
- check_login refused an empty AUDIT_ACCOUNT but accepted "claude:" or
  blank space, which declare no account; it now tests the parsed account.
  Its docstring says it is stricter than run_audit_claude.sh, not the same.
- adversary-collect compared judge labels case-sensitively, so claude and
  Claude shared files on a case-insensitive file system, and accepted the
  label "merged", whose files are the merged table's. Both are refused, as
  is an empty DIR (claude=). The docs and help say to pass a relative bare
  DIR containing "=" as ./adv=2, and a test covers it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The round-3 fixes add four tests to
test_reference_adversary.py and three to the runner tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… notes

CI on 8f968e8 failed six tests once main carried release
dashboard-data-20261006 (#202): they read the working tree's payload,
which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran
on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus
and frozen-run tests now read the payload and the reference explanations
as 8b4c0ca (#187) committed them, and the README says the Reproduce steps
need those inputs.

The round-4 review approved 8f968e8 with two notes:

- The Codex runner's skills check parsed `ls` output, so an entry named a
  bare newline, or ".system\n.system", passed it. It now walks the
  directory with globs and refuses every entry but a real .system
  directory (a symlink named .system too), naming it with %q; tests cover
  both names and the symlink.
- The AUDIT_MODEL default said "the lane's"; with the lane's config.toml
  skipped it is Codex's own.

The review left hooks as an unconfirmed channel. Each Codex call now also
turns the hooks, plugins and apps features off, as codex-cli 0.159.0's
`codex features list` names them, beside memories; the docs name what
remains outside the runner's control.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py),
run one at a time on a loaded machine, with ruff check and ruff format
clean at the locked ruff 0.15.2. The skills-check tests now cover five
names and a symlinked .system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis merged commit 4db91b5 into main Oct 9, 2026
6 checks passed
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merged at the reviewed head 395d81f (squash 4db91b5).

  • Review. The round-5 independent review (GPT-6.1 Sol) approved this head: ~/reviews/policybench-reference-adversary-2026-10/pr200-r5/review_r5_out.md. Round 4 approved 8f968e8; its minor note and its nit were fixed in 4a62b1e.
  • CI. gh pr checks exited 0 on this head: test, app, lint, paper, and the Vercel checks.
  • Score impact. None. The PR adds tooling and evidence only, and changes no published reference, exclusion or score.
  • Follow-up. The review's one remaining minor is open: reference_audit/2026-10-05-reference-adversary/scripts/leaderboard_impact.py regenerates its impact evidence from the working tree's run bundle, not from the pinned 8b4c0ca inputs. The committed results are accurate. It is tracked as a follow-up task.
  • Rulings. d1022 governs the exclusions this evidence supports. The Haiku 5.5 release session will take this PR's proposed_changes.json when it rebases.

MaxGhenis added a commit that referenced this pull request Oct 10, 2026
…6 (release dashboard-data-20261010, 47 models) (#208)

* WIP: start the Haiku 5.5 release driver from finish_gpt61sol.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin Claude Haiku 5.5's finished run inputs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Stage the ten exclusions Max ruled on 2026-10-06 in the Haiku 5.5 driver

docs/haiku55/spec.json lists d1022's eight records (pinned by sha256 in
reference_audit/2026-10-05-reference-adversary/proposed_changes.json) and
d994's two Louisiana records (reference_audit/2026-10-05-louisiana/
proposed_exclusions.json, reference_law_published_after_freeze: the 2026
return amount Louisiana published, $12,838, came after the freeze), with
their adjudication classes and reasoning.

- install-exclusions writes release 20261006's 64 records plus the ten into
  the stage and rebinds the record in stage.json.
- adjudicate-exclusions decides the ten from each case's bound Opus 5.5
  verdict, restating scenario_051's 2026-09-22 regeneration in place.
- Triage accepts exactly those entries and the 2026-10-06 wave in the date
  conventions.
- Export requires 1,910 scored outputs per model and, with 20261006's record
  put back, byte-identical incumbent modelStats.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name Claude Haiku 5.5 and release 20261006 in the driver's messages

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Describe a later-law exclusion's frozen value as the engine's own figure

The excluded-output note said the frozen value of a
reference_law_published_after_freeze record "used the engine's
projection". Louisiana's $12,835 (d994), the first such record, is a
computation from published CPI-U, not a projection, so the note now says
"the engine's own figure", which fits both. The Louisiana records state the
ruling as the earlier records do ("decision d994"). Adds Claude Haiku 5.5's
display name to the paper and app rosters for the release that folds it in.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test the local Claude registration in a fresh interpreter without the remote cost map

The card review of #201 found that the parametrized registration test reads
litellm's already-loaded map, which the remote fetch can fill, so it passes
without the local entry. This test disables the remote map in a subprocess,
checks the bundled backup lacks the newest ids, and requires each locally
registered model at its override prices.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the Haiku 5.5 release's freeze driver, tests and design note; fix three driver bugs

- scripts/freeze_haiku55.py adapts freeze_gpt61sol.py. It rebuilds the
  payload from the receipt-bound stage, checks the installed exclusions, and
  writes the snapshot (74-record exclusion record), annotations, pointer and
  version label. It appends this release's wording amendments to release
  20261006's committed list rather than replacing it.
- tests/test_finish_haiku55.py ports the GPT-6.1 Sol driver tests and covers
  the new steps (reworded cases, spec records, exclusion build, ruled
  adjudications, date conventions, install, the scope gate), with Hypothesis
  properties. tests/test_freeze_haiku55.py does the same for the freeze, on
  fixtures only.
- docs/haiku55/design.md: procedure, gates and invariants.

Driver fixes the new tests caught:
- a restated scenario_051 was rebuilt with release 20261006's judge fields,
  so triage, export and the freeze would refuse it;
- reworded_since_seed compared rows with their pandas index, so a reordered
  file would read as reworded;
- a re-export after the freeze was refused at the committed exclusion
  record, which the freeze replaces with the release's.

The spec's record_edits now name each engine defect's upstream fix (#9928,
#9946, #9948, merged 2026-10-09; Ohio #10020, open, #9925 related). The
later-law docstring no longer calls the frozen value a projection.
592 passed, 2 skipped, 1 strict xfail (the spec's adjudications_written_on,
set when the adjudication step runs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Write the judge provenance record from the stage's sidecars

--step provenance builds docs/haiku55/judge_provenance.json from each
re-opened case's sidecar: its hashes, its sidecar fields with addresses
withheld, and its group by declared account (the pb-judge setup-token
login, or subfleet lane claude-18's token for the cases judged after that
login's weekly limit). It then runs verify_judge_provenance on the result,
so export's gate holds by construction. Tests cover the grouping, the
refusals and, locally, the live stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin the Haiku 5.5 driver to release 20261006 as shipped

BASE_COMMIT is #202's merge on main (9ce4ade), BASE_SHA256 the shipped
payload aa34e5c9..., and the exclusion pin the shipped record (92741dfd...,
whose two scenario_081 sentences the release's review reworded). Reference
explanations, values, scenarios and predictions are unchanged, so every
staged prompt and verdict still binds. The design note's base table follows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the 2026-10-06 wave's decision date and the judge provenance

The spec names 2026-10-09 as the day the ten ruled outputs' decisions were
written (adjudicate-exclusions). docs/haiku55/judge_provenance.json lists the
231 new Opus 5.5 verdicts: 210 on the pb-judge setup-token login and 21 on
subfleet lane claude-18's token after that login's weekly limit. Each was
judged isolated with a token login, and every transcript passes the
runner's checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the reference builder for a second engine upgrade (2.15.17 to a newer policyengine-us)

Generalized from reference_audit/2026-09-28/scripts/build_references_latest.py
(#182). Reads release 20261006 from git, recomputes all 1,984 outputs with the
pinned pre-freeze conventions, and gates every move against an actions file:
approved scored moves, regenerated engine-defect exclusions, new exclusions,
and rechecked excluded outputs that keep their decided values (rule 5).
Narratives regenerate the explanation of each changed output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Let the paper's reference facts follow any number of engine upgrades

EngineUpgrade and its accessors key the paper's figures to the last upgrade,
keep the September (2.15.17) facts pinned to that revision, and rebuild the
exclusion record and references as they stood on any date. A mock second
upgrade (tests/second_engine_upgrade.py, labelled mock data) exercises them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add --step install-references to the Haiku 5.5 driver

Gates a reference build against release 20261006 before writing anything,
renders the audit it gives in memory, installs it in place, sets aside only
the verdicts whose prompts change, and holds every later step (exclusions,
adjudications, triage, export) to the installed build. The freeze refuses an
upgraded stage until it learns to carry one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Aim a drifted regeneration at its fix modules' pre-fix value, not a stale record

A 20261006 engine-defect record decided on 1.755.4 keeps that engine's
corrected value, which other engine changes have since moved (scenario_005
federal: record 107,198.34; the fix module on 2.35.3 gives 108,525.58). A
regeneration now names its audited target: the record's alternative_value
(kind record), or its fix modules' corrected value on an older engine that
still has the defect, from a committed evidence file (kind fix_modules). The
builder applies the modules again on the new engine and refuses if they still
move the output; the driver re-derives the target from the evidence file and
the modules committed at BASE_COMMIT. Evidence mode writes that file.

Also: copy r19_irs_sales_tax_2025.json beside the conventions it serves, and
write the judge provenance record with indent=1 as docs/gpt61sol's is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the fix modules' pre-fix evidence on policyengine-us 2.37.2

24 engine-defect outputs, each on 2.37.2 with the pinned conventions alone and
with its audited fix modules after them. Every module still moves its output
beyond $1 there, so each defect is present on 2.37.2; the corrected values are
the targets a later engine's regeneration must land on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add a pre-merge check of a policyengine-us checkout against the fix evidence

Runs the builder's regeneration gate early: each evidence item must land
within $1 of its corrected value, and its fix modules must leave it
unchanged, on the policyengine-us that PYTHONPATH puts first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Decide the outputs an engine upgrade newly excludes

docs/haiku55/spec.json's upgrade_adjudications holds one item per output the
installed build newly excludes (the Indiana county cells), in the ruled
adjudications' shape. adjudicate-exclusions appends each one's entry, built
from its item, its build record and its case's bound verdict and dated by the
record; the date conventions then name the upgrade's wave, with the day the
spec says its decisions were written. Triage's gate holds the staged record
to exactly those entries, so every excluded output stays decided.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Point the paper's September upgrade facts at that upgrade

Renders byte for byte as before. The upgrade facts move to
r.september_upgrade and r.excluded_outputs_by_engine_version_phrase, so a
second engine upgrade can add its own without changing them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name the upstream fixes behind WI 064 and VA 039 state, by bisecting their merges

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Rule the engine upgrade into the release spec

Max, 2026-10-09: wait for the fixed engine and regenerate the cells it fixes.
The spec now names d1022's four defect cells as regenerated by the upgrade,
decides the two Indiana county outputs the upgrade newly excludes
(upgrade_adjudications), and dates those decisions. The Idaho approval's basis
cites policyengine-us #9810 and Idaho Code 63-3082.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read a regeneration's removed record under the builder's key

The builder and policybench.paper_results name it "record"; the driver read
"removed_record", which only its mock build wrote, so a real build failed
install-references. A local test now runs the 2.37.2 rehearsal build through
load_build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Reuse an earlier build's narrative when its writer inputs are unchanged

A rebuild on a newer engine rewrites every changed output's narrative, and
judges' prompts render it, so each rewrite costs a verdict. --reuse-from keeps
the earlier narrative when the value, cause, grounding (its engine's name
aside), PolicyEngine variable and trace are byte-identical and the narrative
does not name the earlier engine; the reused rows are listed beside the
output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name the build's own engine in the drafted Indiana county record

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Decide triage flags from the spec; scope-export the upgrade without the old annotations

On the upgraded references the judge flagged WI 064 state, regenerated by
#9801: it applied a $500 Wisconsin capital loss limit. The 2025 Wisconsin
Schedule WD limits a net capital loss to $3,000 ($1,500 married filing
separately) or less, as policyengine-us does, so the reference stands. The
spec's triage_adjudications now carries such decisions (affirmed references
only, on re-opened cases, dated by a wave the release names, read only with
an upgrade installed); adjudicate-exclusions appends each from its case's
bound verdict after dropping the regenerated records' decisions, and the
gate holds it exact. A dropped case's new entry is new, not a restatement.

The upgrade's scope export puts release 20261006's references back but keeps
the stage's annotations, which follow the new references, so it now checks
the dashboard schema without requiring an annotation on every wrong answer;
it is read only for modelStats. The judge provenance record describes the
working stage that carries the upgrade.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Freeze a stage that carries an engine upgrade of the references

freeze_haiku55.py now freezes an upgraded stage as strictly as a plain one:
it re-gates the installed build and the recorded drops, requires the export
receipt to bind every retained build file, holds the staged, committed and
manifest reference pins to the build (scenarios stay release 20261006's),
takes the build's explanations, exclusion record and counts (69 records,
1,915 scored outputs on the 2.37.2 rehearsal), names exactly the manifest
fields the last upgrade may change, counts exclusions by their recorded
engines in the version text, and checks the adjudication record with the
upgrade's new-exclusion and triage decisions. freeze_snapshot.py follows the
sidecar's last engine_upgrade revision, not the first. A stage without an
upgrade freezes byte for byte as before.

Built on a Subfleet build lane (GPT-6.1 Sol); its final regression run was
stopped under a disk alert and re-run here. The scope export keeps calling
release_20261006.export_payload, now with require_failure_annotations=False.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Answer the machinery review: pin fix-module siblings, bind approvals and tolerances, correct eight narratives

From the GPT-6.1 Sol review of the upgrade machinery (findings 3-8; 1 and 2
were already fixed in the scope export and triage commits):

- A fix module that loads a sibling (r01, r02 and r03's v2 modules) now pins it:
  the builder and the evidence record each sibling's bytes as committed at the
  base, and the driver re-derives them. The 2.37.2 evidence is re-run with
  them; its values are unchanged.
- Non-finite engine, module, evidence and action values are refused.
- The installer requires an approval for every scored move beyond $1, and
  holds each regeneration to its own tolerance (at most $1).
- Narrative reuse checks the earlier narratives file is that build's (each
  changed row once, for the US, no error, at the build's value), and keeps a
  narrative naming the engine only when the engine is the same.
- A narrative states its amount only as a dollar figure equal to the cent; a
  bare "2026" no longer passes for $0.
- Eight narratives that misstated a figure or the mechanism behind the change
  are hand-corrected from the engine's own values
  (hand_corrected_narratives.json, with the reasons in its .md).

Also re-pins the spec to #200's merged proposed_changes.json (each record now
names decision d1022; CO 043's reading is reworded), names d994 on the two
Louisiana records, and drops an unverified Idaho statute citation from the
drafted approval basis.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record WI 064's answered flag in the design note, not as a decision

The case's re-judged verdict, after its narrative was corrected, raises no
flag, so the release records no triage decision; the design note keeps the
earlier verdict's $500 capital-loss hypothesis and why the reference stands.
The judge provenance record follows the re-judged stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Stage Claude Haiku 5.5's tool_choice: auto sensitivity run

The board row forces the answer tool, which leaves Claude's extended thinking
off (the onboarding probe's forced-tool response carries no thinking block);
the 2026-10-08 re-run with tool_choice: auto declares the tool and leaves it
to the model. Its predictions (results/local/haiku55-stage/auto-run, model
renamed claude-haiku-5.5-thinking) are committed as a deterministic gzip, and
the by-variable and rescore scripts take it. Its scores, would-rank and
by-variable asset are written by the rescore after the 47-model board is
frozen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Draft the Haiku 5.5 release note and paper changes on the 2.37.2 rehearsal

Both are drafts: every bracketed value is the rehearsal's and becomes a fact a
test recomputes from the frozen release; the note and paper proper replace
them once the release on the fixed policyengine-us is frozen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add the paper's accessors for the 2026-10-06 rulings and the second upgrade's counts

ruled_records gathers the records the 2026-10-06 rulings decided, from the
exclusion record and from a later upgrade's regenerated_exclusions, and
refuses one that names no ruling; the counts split them by ruling and by
whether the upgrade regenerated them. Word forms of the last upgrade's counts
use digits above ten. Tested on the real 2.37.2 rehearsal build (local) and
the mock second upgrade, whose 2026-10-06 record now names its ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name each regeneration's upstream fix from a committed map

upstreams.json maps the 20261006 engine-defect root causes to the 2026-10-09
fix pull requests, with the bisected merges (#9801 for WI 064, #9633 for VA
039 state) as per-output entries; actions_from_cells.py --upstreams applies
it, so the build on the final engine needs no hand edits of the draft. On the
2.37.2 rehearsal it names the same fixes the reviewed actions did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* PRACTICE FREEZE on the policyengine-us 2.37.2 rehearsal (replaced at T)

The 47-model snapshot frozen from results/local/rehearsal-2372/stage, so the
paper, the note, the sensitivity rescore and the older releases' test pins can
be brought up to a frozen release before the fixed policyengine-us ships. The
freeze on that release replaces every file here; the squash merge keeps only
it. Not for publication.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Address the delta review's findings 1-5

1. The driver binds each regenerated output to its one reviewed action:
   the revision must restate the action's tolerance, upstream, basis and
   target (kind, fix modules, evidence), the action's alternative_value
   must be the record's, and the action's tolerance is the one enforced,
   for landing and for module inertness.
2. The freeze verifies every fix module the final engine_upgrade pins:
   its path is the module's, its sha256 the bytes committed there, no
   module twice, and a later upgrade keeps the earlier upgrade's pins and
   conventions.
3. policybench/fix_module_closure.py derives a fix module's local
   dependency closure from its syntax tree, shared by the builder and the
   driver. Path(__file__).with_name("<literal>") is the one supported
   loader; a module reaching local files any other way (the c13v3_*
   wrappers' _HERE / f"{name}.py") is refused rather than pinned short.
   On every committed module it agrees with literal discovery or refuses.
4. Narrative reuse requires finite source values; "nan" no longer passes.
5. NY 082's hand narrative: the child credit phases out at federal AGI
   ($117,652.65); the care credit's percentage is about 29.354%.

Also ruff format/check across the release branch, and drop a duplicated
paper_results accessor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Rescore the sensitivity runs on the 47-model board and add Haiku 5.5's

- sensitivity_by_variable.py and rescore_sensitivity_summaries.py on the
  frozen board; Claude Haiku 5.5's tool_choice auto run scores 90.386
  (would rank #11) against its 80.950 board row (#37).
- claude-thinking-2026-08.md: every score, rank and per-program table
  re-read from the frozen board; a Claude Haiku 5.5 section. The scores on
  all 1,984 outputs move with the references, and a new test recomputes
  them.
- The app's sensitivity chips take the rescored values and a Haiku 5.5
  entry; the app tests' release pins move to this board, and the engine
  sentences' tests handle any number of engines.
- A test holds ENGINE_UPGRADE_RECHECK to the frozen sidecar.
- The freeze's pin check reads committed bytes from the script's own
  repository, so a scratch ROOT in tests still sees them; the freeze tests
  read the version list as release 20261006 left it.

Values are the 2.37.2 rehearsal's; they are re-derived at the final freeze.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Carry the release chain's records and tests to release 20261009

- pyproject/uv.lock pin policyengine-us to the reference engine (2.37.2 in
  the rehearsal) and policyengine-core to the builder's 3.32.29.
- scripts/date_haiku55_judge_verdicts.py: writes the verdicts this
  release's restated decisions name (docs/haiku55/judge_verdicts_20261009.json)
  and brings verification/judge_verdicts.json up to it, as release
  20261006's driver did: the 2026-10-05 wave's release is now #202's commit,
  and this release's 2026-10-06 and 2026-10-09 waves join uncommitted.
- tests/test_adjudications.py: earlier releases' checks read their records
  and evidence at their commits; the counts move to this record; a new test
  checks this release's restatements, drops and additions against 20261006's
  record.
- The September 22 audit's tests know a later engine upgrade can regenerate
  a record; the 2026-09-29 installer and driver tests read release 20260929's
  commit; a re-judged case supersedes the 2026-10-05 wording rewrites its new
  verdict writes, without the corrected error returning.
- reference_audit/2026-10-09-engine-upgrade/scripts/sweep_timing.py: the
  2026-10-09 move's timing record and publication check, refusing an engine
  that was not the newest release when the sweep began.

Counts are the 2.37.2 rehearsal's; they are re-derived at the final freeze.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Move the paper, notes and cost tests to release 20261009

Release 20261006's differential tests run on its committed files
(_release_20261006_full); notes on it recompute from its commit; the frozen
release's board snapshot is its manifest's; the pins move to the 47-model
board (rehearsal values, re-derived at the final freeze).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hold the snapshot tests to the real builder's setup (delta review finding 6)

The freeze mocks record the setup build_references_upgrade.py writes on a
later upgrade (inherited modules with paths, latest_final.py and the
sales-tax table, its builder note), and a test checks it against the real
2.37.2 rehearsal build. The manifest test accepts that setup, with the
stated-hours alias from the earlier upgrade and the inherited conventions.
The snapshot audit counts move to the rehearsal's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Write the release's paper text, card, guide and note

- paper/index.qmd: the October 2026 upgrade row and paragraph name the
  restored outputs' fixes, the changed reference and the new exclusions from
  the sidecar (new paper_results accessors, which refuse an unnamed root
  cause); a paragraph on the 2026-10-06 rulings (the reference adversary's
  eight records and Louisiana's two); the abstract and the audit classes gain
  the third exclusion class, law published after the freeze; the sentence on
  mechanically derived exclusions excepts the 2026-10-06 records; the judges
  sentence adds the October 9 addition. Rendered.
- docs/benchmark_card.md and docs/paper.md: the engine, counts and the two
  October paragraphs.
- app/src/notes: the Claude Haiku 5.5 release note, every figure a fact that
  tests/test_notes.py recomputes from the frozen release.
- tests: the card's engine test reads each move's timing record; the paper's
  September timing test names the September engine.

Rehearsal values (policyengine-us 2.37.2); the timing sentences and final
numbers are written at the final freeze.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* PRACTICE re-pin of the rendered paper on the 2.37.2 rehearsal (replaced at T)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* State the October sweep's timing in the paper from its timing record

engine_upgrade_timing_sentence reads the 2026-10-09 audit's
verification/sweep_timing.json: the sweep's day, the engine's PyPI upload
time and what the publication check found. It is empty until the record is
written, and refuses a check that moved outputs or a record of another
engine. MOCK-record tests cover each case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count scored and excluded outputs apart in the publication check

A newer policyengine-us release that fixes a defect behind an excluded output
moves that output's computed value without touching a scored reference. The
check records both counts; the paper's sentence states the scored outputs are
unchanged and how many excluded ones move, and still refuses a release that
moves a scored output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Derive the upgrade's day and release-named files in tests

The final build's sweep falls on a later UTC day than the rehearsal's, so the
tests read the engine upgrade's day from the sidecar, find the release note
and the verdict evidence by name pattern, and the evidence file takes its
name from the driver's release tag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count engine-defect records whose defect is now fixed apart

A record the upgrade rechecked whose new-engine value is its corrected value
has a fixed defect and stays excluded for the unstated input its note names.
The abstract and the limitations no longer call those defects unfixed: they
count with the unstated inputs, and the audit paragraph and the card say how
many there are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Expect the action-binding refusal where it now fires first

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Say the earlier models' rates fall, and how restored targets are known

The note's last paragraph claimed the new references raise every earlier
model's rate. On the release with the IRA fix they fall: the restored
outputs release 20261006 had excluded are ones few models get right. The
paragraph now states the fall, its cause and the counterfactual, with facts
recomputed from the frozen predictions, and its test fails until the frozen
release bears the claim out.

The paper now says how many restored outputs are held to the corrected
value their record carries and how many to the value the audit's fix
modules give on a later engine that still has the defect.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Drop the superseded note and paper drafts

The release note and the paper carry the text now; the drafts held
rehearsal numbers. Two comments named the engine move by a day the
release may not bear.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Bring the design note up to the engine upgrade

The note predated the ruling to wait for the fixed engine. It now
describes the upgrade: which engine, the build and its actions, the two
kinds of regeneration target, the install, and what later steps and the
freeze hold an installed build to, with the tests that pin each gate.
The open items are resolved and removed, and statements that tests and
scripts were planned or uncommitted are gone.

Three comments it listed as stale are corrected in the driver, the
restate script and the driver's tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Say which restored outputs return to scoring and which stay scored

The note and the card counted every regenerated record as an output
PolicyBench had stopped scoring. Four of them, the October 6 engine
defects, were still scored in release 20261006; the new engine fixes
them, so they stay scored instead of being excluded. The note now splits
the restored outputs into the ones release 20261006 excluded and those
four, and its test checks the split against release 20261006's record.

The note test's per-output match count compared the payload's exact mark
with 1.0, but the payload writes a hit as 100.0, so it counted no match.
It now compares with 100.0 and checks the scale.

The fixes behind the restored outputs are also available as a phrase
(engine_upgrade_restored_fix_phrase), which the paper's sentence uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hold the note's facts at the pre-T rehearsal's expected values

Every fact is rewritten from the frozen release at T; until then they are
the expected values from the rehearsal on the code T will carry, so the
draft reads consistently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Tie the publication check to the release's build (pre-T review 1, 6)

sweep_timing.py check now takes the reference build directory instead of
two free paths. It refuses a build whose sidecar does not name the sweep's
engine, pin its CSV and name the committed builder; whose reference CSV
and exclusion record are not the committed snapshot's; or whose
computed.csv does not give every scored reference. The record pins each
input by sha256, and check exits non-zero, after writing it, when the
newest release moves a scored output.

Both steps refuse a sweep whose output does not follow its engine (wheel
upload, then install, then output), so an older file cannot backdate the
sweep. The record lists newer releases whose wheels are all yanked.

The timing sentence dates the engine's upload when it was not on the
sweep's day, and the card test checks its timing record last, so a missing
record no longer hides the September and exclusion-group checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Match the builder's target rule; separate unfixed defects (pre-T review 4, 5, 11)

engine_upgrade_restored_targets now judges a fix-module target the way the
builder's regeneration_target does: the modules must move the output
beyond the exact-match tolerance on their engine, so a flag moved from 0
to 1 counts, and a regenerated flag must equal its target. The refusal and
the 'because' clause the builder has no counterpart for are gone.

The paper's September audit paragraph no longer says every engine-defect
record stays excluded until a fixed engine replaces it: it states how many
the reference engine still has the defect behind, and that an output also
resting on an unstated input stays excluded. The October table row names
how many restored outputs were decided on 2026-10-06.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run the note's claim checks on recomputed facts (pre-T review 2)

The facts-equality check failed on the practice snapshot and hid every
check after it. It is now its own test, with the direction of the last
paragraph. The claim checks run on the facts recomputed from the frozen
release, and new ones cover the sentences nothing checked: the dependent's
,000 of wages and own return, Louisiana's price-index computation and
September 28 bulletin, the unstated Indiana county, Idaho's permanent
building fund tax, the serving claims (from the model card) and the
unchanged conventions. 'The cause is' now rests on the incumbents'
answers and the household weights being release 20261006's. The note's
gain is the difference of the rates it prints.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test the card's hand-written counts against the frozen records (pre-T review 3)

The card's October paragraph (what the move restores, from where, the
fixes behind them, the scored change and the new exclusions) and its
exclusion breakdown (excluded outputs and households, engine-defect
records, the unfixed ones and their root causes, the fixed-but-kept
sentence, unstated inputs and later law) are now recomputed by a test.
The root causes the card names as not fixed must be ones the reference
engine still has.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hold the judge evidence to the decisions it names (pre-T review 7)

scripts/date_haiku55_judge_verdicts.py now refuses, before writing, a
decision whose judge, classes, day or reference flag depart from the
verdict it names (a flag set from an earlier run needs its
judge_reference_suspect_source), and a set of new waves other than the
2026-10-06 rulings' and the engine upgrade's day. The adjudication test
checks the flag of each restated decision and reads the 'current'
evidence of every new-wave decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Refuse more ways a fix module can reach local files (pre-T review 8)

The closure now also refuses directory listings (glob, rglob, iterdir,
listdir, scandir, walk), the module's own location reached another way
(__spec__, __loader__, an attribute .__file__, inspect, globals(),
vars()) and imports of this repository's own package. Of the committed
fix modules this refuses only r14_unlisted_weekly_hours_v2_v2.py, which
imports policybench.scenarios and is no engine-upgrade target; no module
the upgrade's evidence applies is refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin regeneration targets at the freeze; test the mock against every real build (pre-T review 9, 10)

The freeze now checks that each regenerated output held to the audit's
fix modules publishes the committed bytes of every module and sibling its
target pins, and its committed evidence file. The mock-builder test
compares the mock setup with the committed snapshot's engine upgrade from
2.15.17, so it runs everywhere once a freeze commits one, and with every
local build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name later law among the sensitivity doc's exclusions; test its gain claim (pre-T review 11)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Run the publication check's sweep itself; require builder and sidecar (pre-T re-review 1-3)

A file's birth time cannot show which engine computed it, so check no
longer takes a check sweep's output: it runs the sweep, a first pass of
this audit's builder in the venv holding PyPI's newest release, which the
builder refuses unless the venv holds that release, and records what ran.
The build's sidecar must name the committed builder and be the frozen
sidecar byte for byte, and every computed and reference value must be
finite before any comparison.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Close the re-review's nits (pre-T re-review 4-7)

The fix-module closure also refuses a path built from anything but
__file__ (Path(stem + '.py'), the working directory, the home directory)
and opening a file whose name is not a literal; of the committed modules
this refuses only r18_hold_all_projections.py, which no evidence applies,
and the docstring now says what a static read cannot see. The freeze's
regeneration-target check no longer lets an absent file and an absent pin
compare equal. The note test checks the unforced probe its comparison
rests on, and the card-count test refuses a fixed-but-kept sentence when
there are no such records.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Check the imported engine is the wheel; keep the serving probes (pre-T review round 3)

A version label can sit beside older code: an editable install whose
checkout moved, or a source tree ahead of the wheel on the import path.
Both timing steps now run a check under the venv's interpreter, in the
sweep's environment: the version, that policyengine_us is imported from
the installed package, that it is not editable, that every package file
matches the wheel's RECORD sha256, and that no unrecorded module sits
beside them. On the 2.37.2 venv it verifies 18,280 files in 3 s. The venv
path is resolved once, so a relative --check-venv means the same venv for
discovery, the check and the sweep, and the scratch directory is created
if absent.

The two serving probes behind the note's thinking claims are committed
as a minimal fixture (docs/haiku55/serving_probes.json: request settings,
response block types, usage), and the note test reads it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Revert "PRACTICE re-pin of the rendered paper on the 2.37.2 rehearsal (replaced at T)"

This reverts commit 6a758b8.

* Revert "PRACTICE FREEZE on the policyengine-us 2.37.2 rehearsal (replaced at T)"

This reverts commit 41ba29f.

* Roll the release's day to 2026-10-10

The release is built on the policyengine-us release that carries #10032,
which reaches PyPI on 2026-10-10 UTC: the tag, snapshot date, upgrade wave
day, note, judge-evidence file and the card and test names follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Check the engine in the process that computes (pre-T review round 4)

The engine check and the builder now run in one process under the venv's
interpreter (ENGINE_RUNNER via run_on_engine), launched with -P and a
fresh, empty bytecode-cache prefix: no script or working directory joins
the import path, and no __pycache__ beside the sources is read, so every
module compiles from the verified source. The check also refuses
unrecorded importable files (sourceless bytecode, native extensions) and
accepts sha384 and sha512 RECORD hashes. sweep_timing.py run runs the
reference build the same way.

Regressions run the real runner on MOCK installs: stale unchecked-hash
bytecode that a plain interpreter executes, a package beside the builder
script, an unrecorded native extension, each permitted hash algorithm, and
a missing scratch parent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Refuse symlinks in the engine; take only fresh receipts and outputs (pre-T review round 5)

The engine check refuses any symlink in the installed package: a linked
directory hid its files from the scan and could shadow a recorded module,
so installs must be copies (uv's default link modes). sweep_timing.py run
refuses an existing receipt or an out-dir already holding computed.csv,
and refuses after the run unless the builder wrote a new one; path
arguments are resolved in the parent, which a new test showed mattered:
the child runs in the repository root, so a relative --out-dir named
another directory there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read the builder's arguments as the builder does (pre-T review round 6)

sweep_timing.py run now parses the builder's arguments with the builder's
own flags (equals forms, abbreviations, last occurrence), refuses unknown
ones, resolves every path and passes them on in one canonical form. The
file it holds to freshness is the one the builder writes: --computed-csv
when given, else <out-dir>/computed.csv. A test keeps the replica's flags
equal to the builder's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Correct the reference narratives of all 24 outputs the 2.38.6 move changes

The narrative writer's drafts were checked against each output's trace by an
independent reviewer; these texts state the final engine's value and figures,
and the notes give what each draft got wrong.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the judges of the cases the 2.38.6 references re-opened

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep CA 099's state income tax excluded: California disallows educator expenses

The release's judge flagged the regenerated reference: California does not
conform to the federal educator expense deduction (FTB Instructions for
Schedule CA (540), Section C, line 11), and policyengine-us 2.38.6 still
allows it in California AGI, so the reference is about $26.90 low. The
output keeps its record (the build's actions list it under
excluded_outputs_rechecked), and its corrected narrative comes out: the
release installs 23 corrected narratives, not 24.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Record the judge of CA 099's re-opened case

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin policyengine-us 2.38.6, the references' engine

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Freeze release dashboard-data-20261010: Claude Haiku 5.5 on policyengine-us 2.38.6 references

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Date the release's judge verdicts; rescore the sensitivity runs on the 2.38.6 board

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count fixed engine defects by the audit's evidence; separate a further defect

An engine-defect record whose output later engine changes moved is judged
fixed against the corrected value the audit's committed fix-module
evidence gives, not only the record's own value: CA 005 and CT 120 federal
income tax land on theirs, so the IRA phase-out they rest on is no longer
counted unfixed. A record whose recorded defect is fixed but whose
reference rests on a further defect the engine has (CA 099's California
income tax) is its own kind: counted with the defects upstream has not
fixed in the abstract and limitations, and said so in the paper. A landed
record that names neither reason is refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* State the 2.38.6 release in the card, guide, app and note

The card's engine, October move and exclusion breakdown, the paper guide's
counts, the app's engine constant and the release note's facts follow the
frozen release; the note's match count is the lower median (half of the
newly scored outputs are answered within $1 by at most one model).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Render the paper on the 2.38.6 references and pin it

Re-pins the rendered PDF and web files in the manifest and paperSnapshot,
syncs the app's serving config with the frozen copy, and commits the sweep
timing and publication-check record the card and paper cite.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin the tests to the 2.38.6 board and drop a stale judge wave

The judge-verdict dating script built on the working tree's evidence, so a
rehearsal pass's 2026-10-09 wave (release dashboard-data-20261009) survived
into the T evidence beside the real 2026-10-10 wave. The script now keeps
release 20261006's waves and rewrites this release's on every pass, with a
regression test, and the evidence is regenerated from the T stage (only
that wave's six lines change; the restatements file is byte-identical).

Test pins move from the rehearsal to T: 70 adjudications (14 regenerated
records' decisions dropped), 58 excluded outputs in 42 households, 1,926
scored per model, 659 parse failures (644 less four on the new exclusions
plus 19 on the 14 restored outputs), 718 contract violations, 8,001/7,997/
2,138 audit counts, GPT-6 Sol's 95.6% headline, and Fable 5.1's program
rates (federal 67/84 after seven restored outputs and two removed). The
restored-target mock tests now count the sidecar's own two fix-module
targets beside the mocked ones, and the judge-day check for the 2026-10-06
wave is back to the 2026-10-08 and 2026-10-09 runs it reviewed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Answer the final review: CI pins, the 14-of-18 split, and the actions file

Final review (20261010-032143-pb-haiku-final-review) and the PR's first CI
run:

- Blocker: test_reference_exclusions and test_snapshot_artifacts still held
  rehearsal counts. They now pin T's: 58 exclusions (42 unlisted input, 14
  engine defect, 2 later law; 38/18/2 by engine), 8,001 annotated rows,
  7,997 misses, 10,139 below full bounded score, 7,342 llm_error and 659
  parse failures. CI also failed test_freeze_haiku55's rehearsal snapshot
  dates (now the spec's 2026-10-10) and its other-tag test, which the date
  roll had left naming the driver's own tag (it now derives the day after).
- sweep_timing.py read st_birthtime directly, which Linux's stat lacks, so
  every timing test crashed on CI. birth_time() now refuses on such a
  platform rather than date the record by a modification time, and the
  tests use a labeled MOCK modification-time source there only (they create
  each dated file once). Checked under a simulated Linux stat locally.
- Should-fix: the paper said the move returns all 18 fixed outputs to
  scoring. A new accessor, engine_upgrade_restored_split_sentence (a pure
  restored_split_sentence with a Hypothesis test), says 14 that earlier
  releases excluded return and the four ruled on 2026-10-06 while still
  scored stay scored; the table cell says the same. The 2026-10-06
  paragraphs in the paper and card now say PolicyBench ruled to exclude ten.
  Paper re-rendered and re-pinned.
- Nit: the build's final_actions.json is now committed byte for byte (the
  sidecar's actions_sha256 ef23b50e...), as the 2026-09-29 upgrade's is. Its
  note predates CA 099's move to the rechecks and says 19 and 16; the new
  README's erratum says so, and a test holds the arrays at 18 and 17 against
  the sidecar.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 395d81fc Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant