Repository navigation
Conversation
leaderboard_impact.py copied the working tree's run files and compared against the working tree's payload, so regenerating after release dashboard-data-20261006 (#202) scored that release's payload and exclusions and silently rewrote 19 of the 23 committed verification/leaderboard_impact* files (1,928 -> 1,920 scored cells). The script now stages the run bundle from 8b4c0ca (release dashboard-data-20260930) and proposed_changes.json from 4db91b5 (#200) with git show, checks each file against its pinned sha256 and stops before scoring on any mismatch. It reads only the staged copies. --out-dir writes the evidence elsewhere. tests/test_reference_adversary_impact.py checks the pins, the refusals (including #202's payload), that a checkout whose working-tree run and proposals are junk still scores the pinned bytes, and (slow) that a full regeneration reproduces all 23 committed files byte for byte. The README lists every pinned input with its sha256 and the scripts that read it, and the Reproduce steps restore the whole run and annotations directory from 8b4c0ca for the scripts that still read the working tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Hypothesis test flips, truncates or extends any pinned run file and requires git_input to refuse it without writing the target, and to accept the unchanged bytes. A differential check requires the payload pin to equal definition_conformance.py's and publication_sources.py's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
definition_conformance.py, publication_sources.py, engine_probe.py and build_proposals.py read the working tree's run, which release dashboard-data-20261006 (#202) rewrote (payload and exclusion record), and the two engine-side scripts built their reference system from the working tree's convention modules, one of which #204 rewrote. Each now stages every input through scripts/pass_inputs.py, which writes `git show <commit>:<path>` to scratch and stops unless the bytes match the pinned sha256, before the script computes or writes anything. leaderboard_impact.py (#205) uses the same module, so the pins live in one place. build_proposals.py now writes to a required --out and refuses the committed proposed_changes.json, which is pinned at 3a6e5920. No evidence file, reference, exclusion or score changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…inked --out _identical_but now requires the regenerated file to carry the same key, indentation and punctuation on each volatile line, so only that one JSON value may differ; a new test feeds it a changed value elsewhere, a renamed key, re-indentation, a changed list and a dropped line. The build_proposals refusal test adds a relative path and a symlink to the committed proposed_changes.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Contributor
Author
|
Merge-order note, from a local merge reported by the session stacking on both PRs:
|
MaxGhenis
added a commit
that referenced
this pull request
Oct 9, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
marked this pull request as draft
October 10, 2026 00:20
…stage before the engine Review of b4f2aac (GPT-6.1 Sol, hard tier) asked for these changes. - policybench loads benchmark_specs.json itself, as policybench.scenarios is imported, so the engine-side scripts computed with the checkout's definitions while recording the pinned copy's hash. Each process of definition_conformance.py, publication_sources.py and engine_probe.py, workers included, now stages the pinned definitions and points policybench at them (pass_inputs.use_staged_specs) before importing anything else from the package. The test fence no longer exempts the package's own read for these scripts. - publication_sources.py loses --from-facts, which built both reports from a saved file with no pin check. - The conventions and definitions are staged before the engine is imported, hooked or handed to a worker, not after. - Workers are started with the spawn method and an initializer that stages, so a forked worker cannot inherit the main process's system. - build_proposals.py creates --out exclusively, so a hard link to the committed proposed_changes.json or another checkout's copy is refused like any existing file. - leaderboard_impact.py keeps the caller's PYTHONPATH for its scoring child, and its slow regeneration now runs fenced in both processes. - Each script checks that the pass_inputs it imported is the one beside it. No evidence file, reference, exclusion or score changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 10, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This is stacked on #205, which pinned
leaderboard_impact.py. It targetsmainso CI runs, which means #205's two commits (59c32d60,95aaa8fb) appear in this diff until #205 merges. This PR's own change is95aaa8fb..HEAD. Merge #205 first; this branch is then rebased onto main.It gives the reference adversary's other scripts the same staged, sha256-checked inputs:
definition_conformance.pypublication_sources.pyengine_probe.pybuild_proposals.pyThe pins and the check now live in one module,
scripts/pass_inputs.py, andleaderboard_impact.pyuses it too.No committed evidence file, reference, exclusion or score changes, and
proposed_changes.jsonstays at3a6e5920…. The diff touches only the audit's scripts and README, plus tests.The bug, reproduced
These runs used the scripts as of #205's head (
95aaa8fb), from a clean checkout of main. On main, release dashboard-data-20261006 (#202) has rewritten the run's payload and exclusion record, and #204 has rewritten one convention module. Record:bug_on_main.txtin the review folder named under Review.engine_probe.pyexits 0 and silently writes different evidence. Run with the arguments the committedco_sales_tax_refund_scenario_043.jsonrecords, its probe has nopayroll_taxreproduction row. Stop scoring eight tax outputs and date the response window from the last answer (release dashboard-data-20261006) #202 later excluded that cell, and the script reads the working tree's exclusion record. By the same mechanism scenario_082's probe would lose itspayroll_taxrow too, since Stop scoring eight tax outputs and date the response window from the last answer (release dashboard-data-20261006) #202 excluded that cell as well (deduced from the code, not run).definition_conformance.py, run after following Pin the reference adversary's leaderboard impact to the pass's inputs #205's README Reproduce step (restore the run and annotations from8b4c0ca1), also exits 0. It recordslatest_alt_snap_mortgage_residence.pyas5e701f8f…(Stop passing the mortgage interest inputs policyengine-us #9605 deletes #204's rewrite) instead of the pass's00cd1388…. The engine-side scripts copied the reference system's convention modules from the working tree.build_proposals.pywrites the pinnedproposed_changes.jsonin place. Today the bytes come out identical, but only because none of its eight cells is among Stop scoring eight tax outputs and date the response window from the last answer (release dashboard-data-20261006) #202's new exclusions.definition_conformance.pyandpublication_sources.pystop at their payload pin on an unrestored tree. They read every other input from the working tree unchecked.policybenchloadsbenchmark_specs.jsonitself aspolicybench.scenariosis imported, so the engine helpers looked up each output's engine variable and aggregation in the checkout's file. Its bytes equal the pass's today, so nothing has drifted yet.The fix
scripts/pass_inputs.py(new) holds every pin, as commit, path and sha256:8b4c0ca1;policybench/benchmark_specs.jsonat8b4c0ca1;latest_finalis built from, at8b4c0ca1;049f4f09;proposed_changes.jsonandverification/definition_conformance.jsonat4db91b5f.git_inputwritesgit show <commit>:<path>to scratch and stops (SystemExit) unless the bytes match the pin.stage_run,stage_fixesandstage_specsstage sets of files.use_staged_specspointspolicybenchat the staged output definitions for the current process.Each script stages everything it reads before it computes or writes anything, and reads only the staged copy. The three engine-side scripts also stage before they import the engine or start a worker:
definition_conformance.py:policybenchat the staged definitions before importing anything else from it;available: False;publication_sources.pydoes the same for the payload, references, sidecar, exclusion record and scenarios. It loses--from-facts, which built both reports from a saved file with no pin check.engine_probe.pystages the scenarios, references and exclusion record, then the conventions and definitions throughdefinition_conformance's builder.build_proposals.py:--out, which it creates exclusively, so it never writes a file that already exists;proposed_changes.jsonhowever the path is spelled.leaderboard_impact.py(Pin the reference adversary's leaderboard impact to the pass's inputs #205) imports the shared pins.pass_inputs()becamestage_inputs(), and its scoring child now keeps the caller'sPYTHONPATHafter the checkout. It still scores with the checkout'spolicybench, output definitions included; its published-reproduction check covers that.pass_inputsit imported is the file beside it.README.
Tests.
tests/test_reference_adversary_inputs.py(new) holds the new tests below.tests/working_tree_fence.py(new) is an audit hook that fails any open of the working tree's copy of a pinned input,benchmark_specs.jsonincluded. Armed through a generatedsitecustomize, it reaches subprocesses and spawned workers.tests/test_reference_adversary_impact.pyis adapted to the shared module, and its slow regeneration now runs fenced. Pin the reference adversary's leaderboard impact to the pass's inputs #205's byte-refusal property test moved to the new file, where it covers every pin.Verification
Byte-for-byte regeneration. Each script, at head
f6774fd1, ran against main's working tree (#202's run, #204's module) and wrote to scratch, with the fence armed in every process (pytest_slow_r2.txt, 11 passed in 27 minutes on a heavily loaded machine):verification/definition_conformance.mdverification/definition_conformance.jsonrun.seconds(wall time) andrun.script_sha256(the script's own hash; the committed file records33ce4c3e…, the script as #200 merged it)verification/publication_sources.mdverification/publication_sources.jsonmeta.secondsverification/probes/*.json(8)verification/leaderboard_impact*(23)proposed_changes.json(via--outto scratch, fast suite)3a6e5920…A baseline of the unmodified scripts, with the run and modules restored from
8b4c0ca1, differs from the committed files in the samesecondslines only (manual_regeneration.txt). So the environment (uv sync --locked --extra dev --python 3.12, policyengine-us 2.15.17) reproduces the evidence, and the scripts' changes alter none of it.Tests.
f6774fd1(50 passed). The full non-slow suite last ran locally before the review fixes (2,057 passed at54a6a2ba,pytest_fast.txt); CI runs it at head.Mutation check. Partly run; the rest is pending.
f6774fd1(mutants_r2.txt): 12 mutants that each reintroduce a working-tree read, drop the sha256 check, stage after engine work, drop theproposed_changes.jsonrefusal, or put a pin in a script. Every one fails a fast test.mutants_r2.py): 17 mutants that reintroduce the review's findings one at a time. The machine is saturated and local test runs are on hold.mainoralways_zeroreading the working tree'sscenarios.csv, whose bytes equal the pin). Before the review fixes, the first of these passed the fast suite and failed the fenced slow regeneration withPermissionError(mutant_slow_only.txt).Invariants
build_proposals.pyandengine_probe.pywhole and each script's staging and reading in-process;benchmark_specs.json, not the checkout's. A copy of the package with a changed mapping, put first on the path, does not change the lookup. A process that has already loaded other definitions is refused._identical_butrequires the same key, indentation and punctuation on those lines, and the regeneratedscript_sha256to be the current script's hash.build_proposals.pycreates its output exclusively and refusesproposed_changes.jsonby path, so neither a hard link nor another checkout's copy is written.pass_inputs.pyholds a sha256 or names the run's path. The README tables list each pin with exactly the scripts that stage it, checked by recording which pins each script stages.Review
subfleet run --task review --tier hard, atb4f2aac8): REQUEST_CHANGES, with two blockers and five other findings (review_sol.md).f6774fd1;response_r1.mdmaps each finding to its fix and tests.Everything cited above is in
~/reviews/policybench-reference-adversary-2026-10/other-scripts-pinned/.Follow-ups
The Louisiana, Medicare Part B and payroll audits' impact scripts have the same working-tree pattern; #206 pins them. Their engine-side sweeps build
latest_finalfrom the working tree's convention modules and are a separate follow-up, which can reusepass_inputs.pyand the fence.axiom: n/a: audit tooling only; no policy encoded or changed
🤖 Generated with Claude Code