Repository navigation
Add the reference adversary and record its first run (evidence only; no score changes) - #200
Merged
Merged
Conversation
…rsary prompt and runners, definition conformance, publication sources (WIP: adversary not yet run) Work from the reference-adversary session (task_8940406e), which stopped at the 2026-10-05 account cutoff before committing. 361 tests pass (test_consensus, test_definition_conformance, test_publication_sources, test_reference_adversary, test_reference_adversary_runner, test_audit). The consensus flags (61 of 1,928 cells), the definition-conformance report and the publication-source report are generated; the adversary judges have not run yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The scenario_031 head_medicaid_eligible skip now cites #197, the PR that rewrote its annotations, instead of the pre-PR worktree. The six reference-adversary test files plus tests/test_audit.py pass on the rebase onto origin/main 2d39998 (361 passed); the output is in verification/pytest_rebased.txt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both runners started AUDIT_PARALLEL cases and waited for all of them before starting more, so one slow case held every slot: the first live batch sat 15 minutes on one SNAP case while three slots idled. The runners now start the next case as soon as any running case finishes, never running more than AUDIT_PARALLEL at once (bash 3.2 has no wait -n, so they count running jobs with jobs -pr; a finished job not yet noticed only delays a start). The Claude runner still checks the stop flag before every start. test_claude_runner_refills_a_slot_without_exceeding_the_parallel_cap holds scenario_001's stage 1 until scenario_003's first call starts: a batch runner never releases it, the pool does, and the call timeline never shows more than two calls at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er audit does not cover Opus 5.5 at xhigh judged each case in two blind stages on Subfleet lane claude-9 (oauth token from the lane's keychain entry, an empty lane config dir, never the desktop login or an API key): 104 judge calls, all accepted, none contaminated, invalid or refused. Verdicts: 45 reference_holds, 6 reference_wrong (scenario_018 AZ, 025 OH, 026 child1 and child2 NC, 043 CO, 082 NY), 1 definition_mismatch (123 PA). adversary-collect reports 0 missing and 0 inconsistent, and queues 7 cases for developer adjudication. No verdict changes a score. runs/ holds every case's prompts, outputs, provenance sidecars and session transcripts, the runner logs, the collected tables and run_record.json. scripts/verdict_table.py renders verification/verdict_table.md and verdict_counts.json; scripts/engine_probe.py rebuilds a frozen-run household on the reference system for engine evidence (verification/probes/). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y verdicts (WIP: lane stopped before the report and PR) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nal test run README.md reports the pass: design, flag parameters, the 61 flagged / 9 skipped / 52 run cells, the verdicts, the independent verification of the seven that did not hold (four confirmed engine defects, one ambiguous definition scope over four cells, NC 026 refuted), leaderboard impact, the definition-conformance and publication-source findings, judge cost (104 calls, $91.22 reported by the CLI, 5.25 judge-hours) and the rulings. proposed_changes.json status now records d1022 (exclude the eight cells in the next release after dashboard-data-20261006; regenerate the four defect cells on fixed policyengine-us) and d994 (exclude Louisiana 051 and 077; keep the published-amounts convention). The root-cause records are unchanged. verification/pytest_final.txt: 362 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Merged
Dropping the stale "bugs were fixed before this run" sentence from the judge prompt changes all 674 seed prompts. The fold drivers carry a seed verdict only when its prompt re-renders byte-identically, so the next model addition would refuse or need a full re-judge. Restore the template and its test to origin/main's, and note in the report that the sentence's removal waits for a versioned judge template, in a separate PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Removed the judge-template edit from this PR (commit 738d6a6). Dropping the stale sentence changes all 674 seed prompts and breaks byte-identical verdict carry-over at the next model addition (flagged by the release session building 20261006 and Haiku 5.5). The sentence's removal moves to a separate PR that versions the judge template. Review fixes left uncommitted by a stopped lane are parked on |
Brings over the review fixes a stopped lane left on reference-adversary-review-wip (4c1334e), with its failing test fixed and lint cleared: - Blinding: the Claude transcript audit now reads each WebSearch result and rejects an output whose search results list a blocked URL or name PolicyEngine or PolicyBench (WebSearch has no deny rule, so results reach the judge). The source rule asks the judge to pass blocked_domains on every search. - adversary-collect refuses a repeated judge label, queues every case with an inconsistent verdict whatever its class, and exits non-zero when a case has no usable verdict unless --allow-missing is given. - collect_adversary requires a verdict sidecar bound to the current stage 1. - Consensus: a member whose own answer is within the tolerance counts as exact and leaves the wrong cluster (consensus_flags.json and the prototype file reproduce byte-identically under the new rule). - check_login refuses a token lane that declares the desktop login's account. - Codex runner: an allowlisted environment, now with LC_CTYPE as the Claude runner has, and a refusal of a Codex home holding an AGENTS.md. - leaderboard_impact.py refuses a --scratch inside the repository and an unknown --only variant; build_proposals.py records each record's ruling (d1022), separates the pass's own expectations from the rulings, and states the Colorado basis as the verifier did. proposed_changes.json is its output, reproduced byte-identically. The runner test test_codex_runner_passes_only_allowlisted_variables failed on LC_CTYPE. The runner was right: with no LANG in the allowlisted environment, the fake codex (a Python script) sets LC_CTYPE=C.UTF-8 itself under PEP 538 locale coercion. The test now allows LC_CTYPE, which the runner also forwards for parity with the Claude runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The review found WebSearch results listing PolicyEngine and GitHub pages in
run 2's transcripts. Under the audit that now reads search results, 13 of
the 104 accepted outputs, in 10 cases, would have been rejected: 28 exposed
searches, mostly github.com/PolicyEngine and TheAxiomFoundation issues, and
one policyengine.org page.
A lane re-judged those 10 cases on 2026-10-06 (runs/claude-rejudge) with the
source rule that asks for blocked_domains on every search; it left the
evidence uncommitted. This records it:
- scripts/search_exposure.py runs the current transcript audit over both
runs and writes verification/search_exposure.{json,md}, and rewrites
runs/rejudge_flags.json (consensus_flags.json restricted to the 10 cells)
byte-identically to the file the re-judge was prepared from.
- runs/collected-rejudge is adversary-collect's output for the re-judge:
10 verdicts, 0 missing, 0 inconsistent, Colorado 043 queued.
- runs/run_record.json records the re-judge's lane, login, settings, times
and cost ($24.82 reported, 4.70 judge-hours).
- The runner logs the README and verdict_table.py cite were git-ignored
(*.log) and never committed; they are now, with the re-judge's.
Outcome: 0 of the 20 re-judge transcripts flagged; verdict unchanged in all
10. One stage 1 moved: scenario_013 SNAP now finds $288, and stage 2 still
holds the $240 reference.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…findings README: - A section on search results from blocked sources: what the audit missed, the 13 flagged transcripts in 10 cases by source, the re-judge and its outcome (every verdict unchanged), and scenario_013 SNAP, whose blind stage 1 could not date Arizona's 200% limit and found $288. - Counts with the re-judged outputs in place (verdicts unchanged; 39 high and 13 medium; stage 1 41/6/3/2), the re-judge's cost, files and reproduce steps. - Report review: 35 amount and 17 eligibility derivations, not 42 and 10; score ranks shift more than exact ranks and within-1% ranks less (no 5/10% ranks exist); d1022's bullets keep to the ruling, and the pass's own expectations move out of them; the AZ filing-status leaves end at their 2025 values, not all at $15,750. - Open items: the four engine fixes' policyengine-us PRs (Ohio's premiums fix has none yet), scenario_013's effective date, and the Codex runner's blindness to search results. docs/audit.md and the Claude runner header now say that WebSearch has no deny rule, that the audit reads its results, and that the Codex runner cannot; and describe the consensus rule's exact-member exclusion and the collect command's queue, sidecar binding, labels and --allow-missing. search_exposure.py writes runs/rejudge_flags.json before the re-judge exists, so the reproduce steps run in order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py), with ruff check and ruff format clean at the locked ruff 0.15.2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
The independent review of a29cb3c found four gaps: - The Codex runner refused a Codex home holding AGENTS.md but not AGENTS.override.md, which Codex loads first. It now refuses either, and the refusal test covers both. - check_login compared AUDIT_ACCOUNT to the desktop login's email verbatim, so a lane declaring "claude:<desktop email>" passed. The comparison now drops a provider prefix; a prefixed regression covers it. - collect_adversary and adversary-collect accepted a directory with no cases.jsonl and collected it as a judge with no cases. Both now refuse it. - adversary-collect split LABEL=DIR at the last "=", so a directory containing "=" broke. It now splits at the first "=" when the text before it is a label, and keeps a bare DIR whole otherwise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in #201 (Claude Haiku 5.5's model card), whose corrected Sonnet 5.5 cache-read price fixes the test_eval_no_tools failure on this branch's CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pytest_final.txt: the six test files (the five adversary files plus tests/test_audit.py), run one file at a time on a loaded machine: 379 passed, ruff check and ruff format clean at the locked ruff 0.15.2. docs/audit.md: adversary-collect splits LABEL=DIR at the first "=" and refuses a directory without cases.jsonl; the Codex runner refuses a Codex home holding AGENTS.override.md or AGENTS.md. A missing blank line had folded the paragraph after the collect rules into the last list item. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent review of a6323f8 found one major gap and five smaller ones: - The Codex runner checked only the AGENTS files, but Codex adds other text to a session without a tool event: config.toml's developer_instructions and model_instructions_file, memories, and the skills under $HOME/.agents/skills. Each judge call now skips the lane's config.toml (--ignore-user-config; auth still comes from CODEX_HOME), runs with memories off (--disable memories), and gets a fresh, empty HOME, removed afterwards (the login check too). CODEX_HOME is always passed, defaulting to ~/.codex. The runner also refuses a Codex home holding any skill but the bundled .system ones. The header, docs/audit.md and the README name what it still cannot control: what Codex bundles, an administrator's /etc/codex, and apps or plugins on the ChatGPT account (whose calls are MCP tool events, which the audit rejects). No committed run used this runner. - check_login refused an empty AUDIT_ACCOUNT but accepted "claude:" or blank space, which declare no account; it now tests the parsed account. Its docstring says it is stricter than run_audit_claude.sh, not the same. - adversary-collect compared judge labels case-sensitively, so claude and Claude shared files on a case-insensitive file system, and accepted the label "merged", whose files are the merged table's. Both are refused, as is an empty DIR (claude=). The docs and help say to pass a relative bare DIR containing "=" as ./adv=2, and a test covers it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py), run one at a time on a loaded machine, with ruff check and ruff format clean at the locked ruff 0.15.2. The round-3 fixes add four tests to test_reference_adversary.py and three to the runner tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… notes CI on 8f968e8 failed six tests once main carried release dashboard-data-20261006 (#202): they read the working tree's payload, which #202 rewrote (1920 scored cells, sha256 b1da3eae), while the pass ran on release dashboard-data-20260930's (1928 cells, 1e029aaa). The consensus and frozen-run tests now read the payload and the reference explanations as 8b4c0ca (#187) committed them, and the README says the Reproduce steps need those inputs. The round-4 review approved 8f968e8 with two notes: - The Codex runner's skills check parsed `ls` output, so an entry named a bare newline, or ".system\n.system", passed it. It now walks the directory with globs and refuses every entry but a real .system directory (a symlink named .system too), naming it with %q; tests cover both names and the symlink. - The AUDIT_MODEL default said "the lane's"; with the lane's config.toml skipped it is Codex's own. The review left hooks as an unconfirmed channel. Each Codex call now also turns the hooks, plugins and apps features off, as codex-cli 0.159.0's `codex features list` names them, beside memories; the docs name what remains outside the runner's control. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The six test files (the five adversary files plus tests/test_audit.py), run one at a time on a loaded machine, with ruff check and ruff format clean at the locked ruff 0.15.2. The skills-check tests now cover five names and a symlinked .system. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Merged at the reviewed head 395d81f (squash 4db91b5).
|
This was referenced Oct 9, 2026
MaxGhenis
added a commit
that referenced
this pull request
Oct 10, 2026
…6 (release dashboard-data-20261010, 47 models) (#208) * WIP: start the Haiku 5.5 release driver from finish_gpt61sol.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin Claude Haiku 5.5's finished run inputs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Stage the ten exclusions Max ruled on 2026-10-06 in the Haiku 5.5 driver docs/haiku55/spec.json lists d1022's eight records (pinned by sha256 in reference_audit/2026-10-05-reference-adversary/proposed_changes.json) and d994's two Louisiana records (reference_audit/2026-10-05-louisiana/ proposed_exclusions.json, reference_law_published_after_freeze: the 2026 return amount Louisiana published, $12,838, came after the freeze), with their adjudication classes and reasoning. - install-exclusions writes release 20261006's 64 records plus the ten into the stage and rebinds the record in stage.json. - adjudicate-exclusions decides the ten from each case's bound Opus 5.5 verdict, restating scenario_051's 2026-09-22 regeneration in place. - Triage accepts exactly those entries and the 2026-10-06 wave in the date conventions. - Export requires 1,910 scored outputs per model and, with 20261006's record put back, byte-identical incumbent modelStats. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name Claude Haiku 5.5 and release 20261006 in the driver's messages Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Describe a later-law exclusion's frozen value as the engine's own figure The excluded-output note said the frozen value of a reference_law_published_after_freeze record "used the engine's projection". Louisiana's $12,835 (d994), the first such record, is a computation from published CPI-U, not a projection, so the note now says "the engine's own figure", which fits both. The Louisiana records state the ruling as the earlier records do ("decision d994"). Adds Claude Haiku 5.5's display name to the paper and app rosters for the release that folds it in. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test the local Claude registration in a fresh interpreter without the remote cost map The card review of #201 found that the parametrized registration test reads litellm's already-loaded map, which the remote fetch can fill, so it passes without the local entry. This test disables the remote map in a subprocess, checks the bundled backup lacks the newest ids, and requires each locally registered model at its override prices. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the Haiku 5.5 release's freeze driver, tests and design note; fix three driver bugs - scripts/freeze_haiku55.py adapts freeze_gpt61sol.py. It rebuilds the payload from the receipt-bound stage, checks the installed exclusions, and writes the snapshot (74-record exclusion record), annotations, pointer and version label. It appends this release's wording amendments to release 20261006's committed list rather than replacing it. - tests/test_finish_haiku55.py ports the GPT-6.1 Sol driver tests and covers the new steps (reworded cases, spec records, exclusion build, ruled adjudications, date conventions, install, the scope gate), with Hypothesis properties. tests/test_freeze_haiku55.py does the same for the freeze, on fixtures only. - docs/haiku55/design.md: procedure, gates and invariants. Driver fixes the new tests caught: - a restated scenario_051 was rebuilt with release 20261006's judge fields, so triage, export and the freeze would refuse it; - reworded_since_seed compared rows with their pandas index, so a reordered file would read as reworded; - a re-export after the freeze was refused at the committed exclusion record, which the freeze replaces with the release's. The spec's record_edits now name each engine defect's upstream fix (#9928, #9946, #9948, merged 2026-10-09; Ohio #10020, open, #9925 related). The later-law docstring no longer calls the frozen value a projection. 592 passed, 2 skipped, 1 strict xfail (the spec's adjudications_written_on, set when the adjudication step runs). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Write the judge provenance record from the stage's sidecars --step provenance builds docs/haiku55/judge_provenance.json from each re-opened case's sidecar: its hashes, its sidecar fields with addresses withheld, and its group by declared account (the pb-judge setup-token login, or subfleet lane claude-18's token for the cases judged after that login's weekly limit). It then runs verify_judge_provenance on the result, so export's gate holds by construction. Tests cover the grouping, the refusals and, locally, the live stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin the Haiku 5.5 driver to release 20261006 as shipped BASE_COMMIT is #202's merge on main (9ce4ade), BASE_SHA256 the shipped payload aa34e5c9..., and the exclusion pin the shipped record (92741dfd..., whose two scenario_081 sentences the release's review reworded). Reference explanations, values, scenarios and predictions are unchanged, so every staged prompt and verdict still binds. The design note's base table follows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the 2026-10-06 wave's decision date and the judge provenance The spec names 2026-10-09 as the day the ten ruled outputs' decisions were written (adjudicate-exclusions). docs/haiku55/judge_provenance.json lists the 231 new Opus 5.5 verdicts: 210 on the pb-judge setup-token login and 21 on subfleet lane claude-18's token after that login's weekly limit. Each was judged isolated with a token login, and every transcript passes the runner's checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the reference builder for a second engine upgrade (2.15.17 to a newer policyengine-us) Generalized from reference_audit/2026-09-28/scripts/build_references_latest.py (#182). Reads release 20261006 from git, recomputes all 1,984 outputs with the pinned pre-freeze conventions, and gates every move against an actions file: approved scored moves, regenerated engine-defect exclusions, new exclusions, and rechecked excluded outputs that keep their decided values (rule 5). Narratives regenerate the explanation of each changed output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Let the paper's reference facts follow any number of engine upgrades EngineUpgrade and its accessors key the paper's figures to the last upgrade, keep the September (2.15.17) facts pinned to that revision, and rebuild the exclusion record and references as they stood on any date. A mock second upgrade (tests/second_engine_upgrade.py, labelled mock data) exercises them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add --step install-references to the Haiku 5.5 driver Gates a reference build against release 20261006 before writing anything, renders the audit it gives in memory, installs it in place, sets aside only the verdicts whose prompts change, and holds every later step (exclusions, adjudications, triage, export) to the installed build. The freeze refuses an upgraded stage until it learns to carry one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Aim a drifted regeneration at its fix modules' pre-fix value, not a stale record A 20261006 engine-defect record decided on 1.755.4 keeps that engine's corrected value, which other engine changes have since moved (scenario_005 federal: record 107,198.34; the fix module on 2.35.3 gives 108,525.58). A regeneration now names its audited target: the record's alternative_value (kind record), or its fix modules' corrected value on an older engine that still has the defect, from a committed evidence file (kind fix_modules). The builder applies the modules again on the new engine and refuses if they still move the output; the driver re-derives the target from the evidence file and the modules committed at BASE_COMMIT. Evidence mode writes that file. Also: copy r19_irs_sales_tax_2025.json beside the conventions it serves, and write the judge provenance record with indent=1 as docs/gpt61sol's is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the fix modules' pre-fix evidence on policyengine-us 2.37.2 24 engine-defect outputs, each on 2.37.2 with the pinned conventions alone and with its audited fix modules after them. Every module still moves its output beyond $1 there, so each defect is present on 2.37.2; the corrected values are the targets a later engine's regeneration must land on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add a pre-merge check of a policyengine-us checkout against the fix evidence Runs the builder's regeneration gate early: each evidence item must land within $1 of its corrected value, and its fix modules must leave it unchanged, on the policyengine-us that PYTHONPATH puts first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Decide the outputs an engine upgrade newly excludes docs/haiku55/spec.json's upgrade_adjudications holds one item per output the installed build newly excludes (the Indiana county cells), in the ruled adjudications' shape. adjudicate-exclusions appends each one's entry, built from its item, its build record and its case's bound verdict and dated by the record; the date conventions then name the upgrade's wave, with the day the spec says its decisions were written. Triage's gate holds the staged record to exactly those entries, so every excluded output stays decided. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Point the paper's September upgrade facts at that upgrade Renders byte for byte as before. The upgrade facts move to r.september_upgrade and r.excluded_outputs_by_engine_version_phrase, so a second engine upgrade can add its own without changing them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the upstream fixes behind WI 064 and VA 039 state, by bisecting their merges Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rule the engine upgrade into the release spec Max, 2026-10-09: wait for the fixed engine and regenerate the cells it fixes. The spec now names d1022's four defect cells as regenerated by the upgrade, decides the two Indiana county outputs the upgrade newly excludes (upgrade_adjudications), and dates those decisions. The Idaho approval's basis cites policyengine-us #9810 and Idaho Code 63-3082. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read a regeneration's removed record under the builder's key The builder and policybench.paper_results name it "record"; the driver read "removed_record", which only its mock build wrote, so a real build failed install-references. A local test now runs the 2.37.2 rehearsal build through load_build. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Reuse an earlier build's narrative when its writer inputs are unchanged A rebuild on a newer engine rewrites every changed output's narrative, and judges' prompts render it, so each rewrite costs a verdict. --reuse-from keeps the earlier narrative when the value, cause, grounding (its engine's name aside), PolicyEngine variable and trace are byte-identical and the narrative does not name the earlier engine; the reused rows are listed beside the output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the build's own engine in the drafted Indiana county record Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Decide triage flags from the spec; scope-export the upgrade without the old annotations On the upgraded references the judge flagged WI 064 state, regenerated by #9801: it applied a $500 Wisconsin capital loss limit. The 2025 Wisconsin Schedule WD limits a net capital loss to $3,000 ($1,500 married filing separately) or less, as policyengine-us does, so the reference stands. The spec's triage_adjudications now carries such decisions (affirmed references only, on re-opened cases, dated by a wave the release names, read only with an upgrade installed); adjudicate-exclusions appends each from its case's bound verdict after dropping the regenerated records' decisions, and the gate holds it exact. A dropped case's new entry is new, not a restatement. The upgrade's scope export puts release 20261006's references back but keeps the stage's annotations, which follow the new references, so it now checks the dashboard schema without requiring an annotation on every wrong answer; it is read only for modelStats. The judge provenance record describes the working stage that carries the upgrade. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Freeze a stage that carries an engine upgrade of the references freeze_haiku55.py now freezes an upgraded stage as strictly as a plain one: it re-gates the installed build and the recorded drops, requires the export receipt to bind every retained build file, holds the staged, committed and manifest reference pins to the build (scenarios stay release 20261006's), takes the build's explanations, exclusion record and counts (69 records, 1,915 scored outputs on the 2.37.2 rehearsal), names exactly the manifest fields the last upgrade may change, counts exclusions by their recorded engines in the version text, and checks the adjudication record with the upgrade's new-exclusion and triage decisions. freeze_snapshot.py follows the sidecar's last engine_upgrade revision, not the first. A stage without an upgrade freezes byte for byte as before. Built on a Subfleet build lane (GPT-6.1 Sol); its final regression run was stopped under a disk alert and re-run here. The scope export keeps calling release_20261006.export_payload, now with require_failure_annotations=False. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Answer the machinery review: pin fix-module siblings, bind approvals and tolerances, correct eight narratives From the GPT-6.1 Sol review of the upgrade machinery (findings 3-8; 1 and 2 were already fixed in the scope export and triage commits): - A fix module that loads a sibling (r01, r02 and r03's v2 modules) now pins it: the builder and the evidence record each sibling's bytes as committed at the base, and the driver re-derives them. The 2.37.2 evidence is re-run with them; its values are unchanged. - Non-finite engine, module, evidence and action values are refused. - The installer requires an approval for every scored move beyond $1, and holds each regeneration to its own tolerance (at most $1). - Narrative reuse checks the earlier narratives file is that build's (each changed row once, for the US, no error, at the build's value), and keeps a narrative naming the engine only when the engine is the same. - A narrative states its amount only as a dollar figure equal to the cent; a bare "2026" no longer passes for $0. - Eight narratives that misstated a figure or the mechanism behind the change are hand-corrected from the engine's own values (hand_corrected_narratives.json, with the reasons in its .md). Also re-pins the spec to #200's merged proposed_changes.json (each record now names decision d1022; CO 043's reading is reworded), names d994 on the two Louisiana records, and drops an unverified Idaho statute citation from the drafted approval basis. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record WI 064's answered flag in the design note, not as a decision The case's re-judged verdict, after its narrative was corrected, raises no flag, so the release records no triage decision; the design note keeps the earlier verdict's $500 capital-loss hypothesis and why the reference stands. The judge provenance record follows the re-judged stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Stage Claude Haiku 5.5's tool_choice: auto sensitivity run The board row forces the answer tool, which leaves Claude's extended thinking off (the onboarding probe's forced-tool response carries no thinking block); the 2026-10-08 re-run with tool_choice: auto declares the tool and leaves it to the model. Its predictions (results/local/haiku55-stage/auto-run, model renamed claude-haiku-5.5-thinking) are committed as a deterministic gzip, and the by-variable and rescore scripts take it. Its scores, would-rank and by-variable asset are written by the rescore after the 47-model board is frozen. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Draft the Haiku 5.5 release note and paper changes on the 2.37.2 rehearsal Both are drafts: every bracketed value is the rehearsal's and becomes a fact a test recomputes from the frozen release; the note and paper proper replace them once the release on the fixed policyengine-us is frozen. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the paper's accessors for the 2026-10-06 rulings and the second upgrade's counts ruled_records gathers the records the 2026-10-06 rulings decided, from the exclusion record and from a later upgrade's regenerated_exclusions, and refuses one that names no ruling; the counts split them by ruling and by whether the upgrade regenerated them. Word forms of the last upgrade's counts use digits above ten. Tested on the real 2.37.2 rehearsal build (local) and the mock second upgrade, whose 2026-10-06 record now names its ruling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name each regeneration's upstream fix from a committed map upstreams.json maps the 20261006 engine-defect root causes to the 2026-10-09 fix pull requests, with the bisected merges (#9801 for WI 064, #9633 for VA 039 state) as per-output entries; actions_from_cells.py --upstreams applies it, so the build on the final engine needs no hand edits of the draft. On the 2.37.2 rehearsal it names the same fixes the reviewed actions did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * PRACTICE FREEZE on the policyengine-us 2.37.2 rehearsal (replaced at T) The 47-model snapshot frozen from results/local/rehearsal-2372/stage, so the paper, the note, the sensitivity rescore and the older releases' test pins can be brought up to a frozen release before the fixed policyengine-us ships. The freeze on that release replaces every file here; the squash merge keeps only it. Not for publication. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address the delta review's findings 1-5 1. The driver binds each regenerated output to its one reviewed action: the revision must restate the action's tolerance, upstream, basis and target (kind, fix modules, evidence), the action's alternative_value must be the record's, and the action's tolerance is the one enforced, for landing and for module inertness. 2. The freeze verifies every fix module the final engine_upgrade pins: its path is the module's, its sha256 the bytes committed there, no module twice, and a later upgrade keeps the earlier upgrade's pins and conventions. 3. policybench/fix_module_closure.py derives a fix module's local dependency closure from its syntax tree, shared by the builder and the driver. Path(__file__).with_name("<literal>") is the one supported loader; a module reaching local files any other way (the c13v3_* wrappers' _HERE / f"{name}.py") is refused rather than pinned short. On every committed module it agrees with literal discovery or refuses. 4. Narrative reuse requires finite source values; "nan" no longer passes. 5. NY 082's hand narrative: the child credit phases out at federal AGI ($117,652.65); the care credit's percentage is about 29.354%. Also ruff format/check across the release branch, and drop a duplicated paper_results accessor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rescore the sensitivity runs on the 47-model board and add Haiku 5.5's - sensitivity_by_variable.py and rescore_sensitivity_summaries.py on the frozen board; Claude Haiku 5.5's tool_choice auto run scores 90.386 (would rank #11) against its 80.950 board row (#37). - claude-thinking-2026-08.md: every score, rank and per-program table re-read from the frozen board; a Claude Haiku 5.5 section. The scores on all 1,984 outputs move with the references, and a new test recomputes them. - The app's sensitivity chips take the rescored values and a Haiku 5.5 entry; the app tests' release pins move to this board, and the engine sentences' tests handle any number of engines. - A test holds ENGINE_UPGRADE_RECHECK to the frozen sidecar. - The freeze's pin check reads committed bytes from the script's own repository, so a scratch ROOT in tests still sees them; the freeze tests read the version list as release 20261006 left it. Values are the 2.37.2 rehearsal's; they are re-derived at the final freeze. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Carry the release chain's records and tests to release 20261009 - pyproject/uv.lock pin policyengine-us to the reference engine (2.37.2 in the rehearsal) and policyengine-core to the builder's 3.32.29. - scripts/date_haiku55_judge_verdicts.py: writes the verdicts this release's restated decisions name (docs/haiku55/judge_verdicts_20261009.json) and brings verification/judge_verdicts.json up to it, as release 20261006's driver did: the 2026-10-05 wave's release is now #202's commit, and this release's 2026-10-06 and 2026-10-09 waves join uncommitted. - tests/test_adjudications.py: earlier releases' checks read their records and evidence at their commits; the counts move to this record; a new test checks this release's restatements, drops and additions against 20261006's record. - The September 22 audit's tests know a later engine upgrade can regenerate a record; the 2026-09-29 installer and driver tests read release 20260929's commit; a re-judged case supersedes the 2026-10-05 wording rewrites its new verdict writes, without the corrected error returning. - reference_audit/2026-10-09-engine-upgrade/scripts/sweep_timing.py: the 2026-10-09 move's timing record and publication check, refusing an engine that was not the newest release when the sweep began. Counts are the 2.37.2 rehearsal's; they are re-derived at the final freeze. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Move the paper, notes and cost tests to release 20261009 Release 20261006's differential tests run on its committed files (_release_20261006_full); notes on it recompute from its commit; the frozen release's board snapshot is its manifest's; the pins move to the 47-model board (rehearsal values, re-derived at the final freeze). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hold the snapshot tests to the real builder's setup (delta review finding 6) The freeze mocks record the setup build_references_upgrade.py writes on a later upgrade (inherited modules with paths, latest_final.py and the sales-tax table, its builder note), and a test checks it against the real 2.37.2 rehearsal build. The manifest test accepts that setup, with the stated-hours alias from the earlier upgrade and the inherited conventions. The snapshot audit counts move to the rehearsal's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Write the release's paper text, card, guide and note - paper/index.qmd: the October 2026 upgrade row and paragraph name the restored outputs' fixes, the changed reference and the new exclusions from the sidecar (new paper_results accessors, which refuse an unnamed root cause); a paragraph on the 2026-10-06 rulings (the reference adversary's eight records and Louisiana's two); the abstract and the audit classes gain the third exclusion class, law published after the freeze; the sentence on mechanically derived exclusions excepts the 2026-10-06 records; the judges sentence adds the October 9 addition. Rendered. - docs/benchmark_card.md and docs/paper.md: the engine, counts and the two October paragraphs. - app/src/notes: the Claude Haiku 5.5 release note, every figure a fact that tests/test_notes.py recomputes from the frozen release. - tests: the card's engine test reads each move's timing record; the paper's September timing test names the September engine. Rehearsal values (policyengine-us 2.37.2); the timing sentences and final numbers are written at the final freeze. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * PRACTICE re-pin of the rendered paper on the 2.37.2 rehearsal (replaced at T) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * State the October sweep's timing in the paper from its timing record engine_upgrade_timing_sentence reads the 2026-10-09 audit's verification/sweep_timing.json: the sweep's day, the engine's PyPI upload time and what the publication check found. It is empty until the record is written, and refuses a check that moved outputs or a record of another engine. MOCK-record tests cover each case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Count scored and excluded outputs apart in the publication check A newer policyengine-us release that fixes a defect behind an excluded output moves that output's computed value without touching a scored reference. The check records both counts; the paper's sentence states the scored outputs are unchanged and how many excluded ones move, and still refuses a release that moves a scored output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Derive the upgrade's day and release-named files in tests The final build's sweep falls on a later UTC day than the rehearsal's, so the tests read the engine upgrade's day from the sidecar, find the release note and the verdict evidence by name pattern, and the evidence file takes its name from the driver's release tag. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Count engine-defect records whose defect is now fixed apart A record the upgrade rechecked whose new-engine value is its corrected value has a fixed defect and stays excluded for the unstated input its note names. The abstract and the limitations no longer call those defects unfixed: they count with the unstated inputs, and the audit paragraph and the card say how many there are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Expect the action-binding refusal where it now fires first Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say the earlier models' rates fall, and how restored targets are known The note's last paragraph claimed the new references raise every earlier model's rate. On the release with the IRA fix they fall: the restored outputs release 20261006 had excluded are ones few models get right. The paragraph now states the fall, its cause and the counterfactual, with facts recomputed from the frozen predictions, and its test fails until the frozen release bears the claim out. The paper now says how many restored outputs are held to the corrected value their record carries and how many to the value the audit's fix modules give on a later engine that still has the defect. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Drop the superseded note and paper drafts The release note and the paper carry the text now; the drafts held rehearsal numbers. Two comments named the engine move by a day the release may not bear. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Bring the design note up to the engine upgrade The note predated the ruling to wait for the fixed engine. It now describes the upgrade: which engine, the build and its actions, the two kinds of regeneration target, the install, and what later steps and the freeze hold an installed build to, with the tests that pin each gate. The open items are resolved and removed, and statements that tests and scripts were planned or uncommitted are gone. Three comments it listed as stale are corrected in the driver, the restate script and the driver's tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Say which restored outputs return to scoring and which stay scored The note and the card counted every regenerated record as an output PolicyBench had stopped scoring. Four of them, the October 6 engine defects, were still scored in release 20261006; the new engine fixes them, so they stay scored instead of being excluded. The note now splits the restored outputs into the ones release 20261006 excluded and those four, and its test checks the split against release 20261006's record. The note test's per-output match count compared the payload's exact mark with 1.0, but the payload writes a hit as 100.0, so it counted no match. It now compares with 100.0 and checks the scale. The fixes behind the restored outputs are also available as a phrase (engine_upgrade_restored_fix_phrase), which the paper's sentence uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hold the note's facts at the pre-T rehearsal's expected values Every fact is rewritten from the frozen release at T; until then they are the expected values from the rehearsal on the code T will carry, so the draft reads consistently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tie the publication check to the release's build (pre-T review 1, 6) sweep_timing.py check now takes the reference build directory instead of two free paths. It refuses a build whose sidecar does not name the sweep's engine, pin its CSV and name the committed builder; whose reference CSV and exclusion record are not the committed snapshot's; or whose computed.csv does not give every scored reference. The record pins each input by sha256, and check exits non-zero, after writing it, when the newest release moves a scored output. Both steps refuse a sweep whose output does not follow its engine (wheel upload, then install, then output), so an older file cannot backdate the sweep. The record lists newer releases whose wheels are all yanked. The timing sentence dates the engine's upload when it was not on the sweep's day, and the card test checks its timing record last, so a missing record no longer hides the September and exclusion-group checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Match the builder's target rule; separate unfixed defects (pre-T review 4, 5, 11) engine_upgrade_restored_targets now judges a fix-module target the way the builder's regeneration_target does: the modules must move the output beyond the exact-match tolerance on their engine, so a flag moved from 0 to 1 counts, and a regenerated flag must equal its target. The refusal and the 'because' clause the builder has no counterpart for are gone. The paper's September audit paragraph no longer says every engine-defect record stays excluded until a fixed engine replaces it: it states how many the reference engine still has the defect behind, and that an output also resting on an unstated input stays excluded. The October table row names how many restored outputs were decided on 2026-10-06. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Run the note's claim checks on recomputed facts (pre-T review 2) The facts-equality check failed on the practice snapshot and hid every check after it. It is now its own test, with the direction of the last paragraph. The claim checks run on the facts recomputed from the frozen release, and new ones cover the sentences nothing checked: the dependent's ,000 of wages and own return, Louisiana's price-index computation and September 28 bulletin, the unstated Indiana county, Idaho's permanent building fund tax, the serving claims (from the model card) and the unchanged conventions. 'The cause is' now rests on the incumbents' answers and the household weights being release 20261006's. The note's gain is the difference of the rates it prints. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test the card's hand-written counts against the frozen records (pre-T review 3) The card's October paragraph (what the move restores, from where, the fixes behind them, the scored change and the new exclusions) and its exclusion breakdown (excluded outputs and households, engine-defect records, the unfixed ones and their root causes, the fixed-but-kept sentence, unstated inputs and later law) are now recomputed by a test. The root causes the card names as not fixed must be ones the reference engine still has. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hold the judge evidence to the decisions it names (pre-T review 7) scripts/date_haiku55_judge_verdicts.py now refuses, before writing, a decision whose judge, classes, day or reference flag depart from the verdict it names (a flag set from an earlier run needs its judge_reference_suspect_source), and a set of new waves other than the 2026-10-06 rulings' and the engine upgrade's day. The adjudication test checks the flag of each restated decision and reads the 'current' evidence of every new-wave decision. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refuse more ways a fix module can reach local files (pre-T review 8) The closure now also refuses directory listings (glob, rglob, iterdir, listdir, scandir, walk), the module's own location reached another way (__spec__, __loader__, an attribute .__file__, inspect, globals(), vars()) and imports of this repository's own package. Of the committed fix modules this refuses only r14_unlisted_weekly_hours_v2_v2.py, which imports policybench.scenarios and is no engine-upgrade target; no module the upgrade's evidence applies is refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin regeneration targets at the freeze; test the mock against every real build (pre-T review 9, 10) The freeze now checks that each regenerated output held to the audit's fix modules publishes the committed bytes of every module and sibling its target pins, and its committed evidence file. The mock-builder test compares the mock setup with the committed snapshot's engine upgrade from 2.15.17, so it runs everywhere once a freeze commits one, and with every local build. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name later law among the sensitivity doc's exclusions; test its gain claim (pre-T review 11) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Run the publication check's sweep itself; require builder and sidecar (pre-T re-review 1-3) A file's birth time cannot show which engine computed it, so check no longer takes a check sweep's output: it runs the sweep, a first pass of this audit's builder in the venv holding PyPI's newest release, which the builder refuses unless the venv holds that release, and records what ran. The build's sidecar must name the committed builder and be the frozen sidecar byte for byte, and every computed and reference value must be finite before any comparison. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Close the re-review's nits (pre-T re-review 4-7) The fix-module closure also refuses a path built from anything but __file__ (Path(stem + '.py'), the working directory, the home directory) and opening a file whose name is not a literal; of the committed modules this refuses only r18_hold_all_projections.py, which no evidence applies, and the docstring now says what a static read cannot see. The freeze's regeneration-target check no longer lets an absent file and an absent pin compare equal. The note test checks the unforced probe its comparison rests on, and the card-count test refuses a fixed-but-kept sentence when there are no such records. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Check the imported engine is the wheel; keep the serving probes (pre-T review round 3) A version label can sit beside older code: an editable install whose checkout moved, or a source tree ahead of the wheel on the import path. Both timing steps now run a check under the venv's interpreter, in the sweep's environment: the version, that policyengine_us is imported from the installed package, that it is not editable, that every package file matches the wheel's RECORD sha256, and that no unrecorded module sits beside them. On the 2.37.2 venv it verifies 18,280 files in 3 s. The venv path is resolved once, so a relative --check-venv means the same venv for discovery, the check and the sweep, and the scratch directory is created if absent. The two serving probes behind the note's thinking claims are committed as a minimal fixture (docs/haiku55/serving_probes.json: request settings, response block types, usage), and the note test reads it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Revert "PRACTICE re-pin of the rendered paper on the 2.37.2 rehearsal (replaced at T)" This reverts commit 6a758b8. * Revert "PRACTICE FREEZE on the policyengine-us 2.37.2 rehearsal (replaced at T)" This reverts commit 41ba29f. * Roll the release's day to 2026-10-10 The release is built on the policyengine-us release that carries #10032, which reaches PyPI on 2026-10-10 UTC: the tag, snapshot date, upgrade wave day, note, judge-evidence file and the card and test names follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Check the engine in the process that computes (pre-T review round 4) The engine check and the builder now run in one process under the venv's interpreter (ENGINE_RUNNER via run_on_engine), launched with -P and a fresh, empty bytecode-cache prefix: no script or working directory joins the import path, and no __pycache__ beside the sources is read, so every module compiles from the verified source. The check also refuses unrecorded importable files (sourceless bytecode, native extensions) and accepts sha384 and sha512 RECORD hashes. sweep_timing.py run runs the reference build the same way. Regressions run the real runner on MOCK installs: stale unchecked-hash bytecode that a plain interpreter executes, a package beside the builder script, an unrecorded native extension, each permitted hash algorithm, and a missing scratch parent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Refuse symlinks in the engine; take only fresh receipts and outputs (pre-T review round 5) The engine check refuses any symlink in the installed package: a linked directory hid its files from the scan and could shadow a recorded module, so installs must be copies (uv's default link modes). sweep_timing.py run refuses an existing receipt or an out-dir already holding computed.csv, and refuses after the run unless the builder wrote a new one; path arguments are resolved in the parent, which a new test showed mattered: the child runs in the repository root, so a relative --out-dir named another directory there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read the builder's arguments as the builder does (pre-T review round 6) sweep_timing.py run now parses the builder's arguments with the builder's own flags (equals forms, abbreviations, last occurrence), refuses unknown ones, resolves every path and passes them on in one canonical form. The file it holds to freshness is the one the builder writes: --computed-csv when given, else <out-dir>/computed.csv. A test keeps the replica's flags equal to the builder's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Correct the reference narratives of all 24 outputs the 2.38.6 move changes The narrative writer's drafts were checked against each output's trace by an independent reviewer; these texts state the final engine's value and figures, and the notes give what each draft got wrong. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the judges of the cases the 2.38.6 references re-opened Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep CA 099's state income tax excluded: California disallows educator expenses The release's judge flagged the regenerated reference: California does not conform to the federal educator expense deduction (FTB Instructions for Schedule CA (540), Section C, line 11), and policyengine-us 2.38.6 still allows it in California AGI, so the reference is about $26.90 low. The output keeps its record (the build's actions list it under excluded_outputs_rechecked), and its corrected narrative comes out: the release installs 23 corrected narratives, not 24. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Record the judge of CA 099's re-opened case Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin policyengine-us 2.38.6, the references' engine Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Freeze release dashboard-data-20261010: Claude Haiku 5.5 on policyengine-us 2.38.6 references Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Date the release's judge verdicts; rescore the sensitivity runs on the 2.38.6 board Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Count fixed engine defects by the audit's evidence; separate a further defect An engine-defect record whose output later engine changes moved is judged fixed against the corrected value the audit's committed fix-module evidence gives, not only the record's own value: CA 005 and CT 120 federal income tax land on theirs, so the IRA phase-out they rest on is no longer counted unfixed. A record whose recorded defect is fixed but whose reference rests on a further defect the engine has (CA 099's California income tax) is its own kind: counted with the defects upstream has not fixed in the abstract and limitations, and said so in the paper. A landed record that names neither reason is refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * State the 2.38.6 release in the card, guide, app and note The card's engine, October move and exclusion breakdown, the paper guide's counts, the app's engine constant and the release note's facts follow the frozen release; the note's match count is the lower median (half of the newly scored outputs are answered within $1 by at most one model). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Render the paper on the 2.38.6 references and pin it Re-pins the rendered PDF and web files in the manifest and paperSnapshot, syncs the app's serving config with the frozen copy, and commits the sweep timing and publication-check record the card and paper cite. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin the tests to the 2.38.6 board and drop a stale judge wave The judge-verdict dating script built on the working tree's evidence, so a rehearsal pass's 2026-10-09 wave (release dashboard-data-20261009) survived into the T evidence beside the real 2026-10-10 wave. The script now keeps release 20261006's waves and rewrites this release's on every pass, with a regression test, and the evidence is regenerated from the T stage (only that wave's six lines change; the restatements file is byte-identical). Test pins move from the rehearsal to T: 70 adjudications (14 regenerated records' decisions dropped), 58 excluded outputs in 42 households, 1,926 scored per model, 659 parse failures (644 less four on the new exclusions plus 19 on the 14 restored outputs), 718 contract violations, 8,001/7,997/ 2,138 audit counts, GPT-6 Sol's 95.6% headline, and Fable 5.1's program rates (federal 67/84 after seven restored outputs and two removed). The restored-target mock tests now count the sidecar's own two fix-module targets beside the mocked ones, and the judge-day check for the 2026-10-06 wave is back to the 2026-10-08 and 2026-10-09 runs it reviewed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Answer the final review: CI pins, the 14-of-18 split, and the actions file Final review (20261010-032143-pb-haiku-final-review) and the PR's first CI run: - Blocker: test_reference_exclusions and test_snapshot_artifacts still held rehearsal counts. They now pin T's: 58 exclusions (42 unlisted input, 14 engine defect, 2 later law; 38/18/2 by engine), 8,001 annotated rows, 7,997 misses, 10,139 below full bounded score, 7,342 llm_error and 659 parse failures. CI also failed test_freeze_haiku55's rehearsal snapshot dates (now the spec's 2026-10-10) and its other-tag test, which the date roll had left naming the driver's own tag (it now derives the day after). - sweep_timing.py read st_birthtime directly, which Linux's stat lacks, so every timing test crashed on CI. birth_time() now refuses on such a platform rather than date the record by a modification time, and the tests use a labeled MOCK modification-time source there only (they create each dated file once). Checked under a simulated Linux stat locally. - Should-fix: the paper said the move returns all 18 fixed outputs to scoring. A new accessor, engine_upgrade_restored_split_sentence (a pure restored_split_sentence with a Hypothesis test), says 14 that earlier releases excluded return and the four ruled on 2026-10-06 while still scored stay scored; the table cell says the same. The 2026-10-06 paragraphs in the paper and card now say PolicyBench ruled to exclude ten. Paper re-rendered and re-pinned. - Nit: the build's final_actions.json is now committed byte for byte (the sidecar's actions_sha256 ef23b50e...), as the 2026-09-29 upgrade's is. Its note predates CA 099's move to the rechecks and says 19 and 16; the new README's erratum says so, and a test holds the arrays at 18 and 17 against the sidecar. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the reference adversary, a pass whose only job is to attack PolicyBench's US references, and records its first run on the frozen 20260612 run (46 models, 1,928 scored cells). Evidence and tooling only: no reference, exclusion record or score changes. The full report is
reference_audit/2026-10-05-reference-adversary/README.md.What it found
reference_holds, 6reference_wrongand 1definition_mismatch.blocked_domainson every search. None of the 20 new transcripts is flagged, and every verdict is unchanged.verification/search_exposure.md).verification/independent/):CONFIRMED engine defects, all still on policyengine-us main 2.29.11:
AMBIGUOUS: PA 123. The definitions don't say whether a dependent's own required return counts. The same question reaches 4 cells: PA 123 and MO 093, state and federal.
REFUTED: NC 026 child1/child2. SPA NC-23-0009's 42 CFR 435.218 election covers insured children, so both references stand.
verification/leaderboard_impact.json, scored withpolicybench analyze). Excluding the 8 cells raises every model's exact rate by 0.40–0.65 pp and swaps claude-sonnet-5.5 and gpt-6-luna at ranks 4 and 5.Rulings (recorded in
proposed_changes.jsonstatus)This PR installs neither ruling; the next release does.
Review fixes since the first review
The independent code review requested changes (blinding: search results). The report review found count and wording errors. Both are addressed:
claude_transcript_auditrejects a search whose result lists a blocked URL or names PolicyEngine or PolicyBench. The source rule asks forblocked_domainson every search. Docs and both runner headers now say the Codex runner cannot see search results.adversary-collect:--allow-missingis given.consensus_flags.jsonand the prototype file reproduce byte-identically under the new rule.AGENTS.md.check_loginrefuses a token lane that declares the desktop account.leaderboard_impact.pyrefuses a--scratchinside the repo and an unknown--onlyvariant.build_proposals.pyrecords each record's ruling (decision: d1022) and keeps the pass's own expectations out of the rulings. It also states the Colorado basis as the verifier did;proposed_changes.jsonreproduces byte-identically.runs/collected/reproduces byte-identically under the new collect. The runner logs the README cites were git-ignored and are now committed.Invariants
The tests check each of these, with Hypothesis property tests where marked:
tests/test_consensus.py):tests/test_reference_adversary.py):apply_adversary_flagschanges nothing but the flags it sets (property);tests/test_reference_adversary_runner.py, with fake CLIs):AGENTS.mdrefusal;Tests
reference_audit/2026-10-05-reference-adversary/verification/pytest_final.txtholds the final run:tests/test_audit.py.ruff check .andruff format --check .pass with the locked ruff 0.15.2.The previously failing
test_codex_runner_passes_only_allowlisted_variableswas a test bug. With noLANG, the fake codex, a Python script, setsLC_CTYPEitself under PEP 538, so the runner was right.axiom: n/a: PolicyBench evidence and tooling; the engine defects are fixed in policyengine-us under d1022.
🤖 Generated with Claude Code