Skip to content

Pin the reference adversary's other scripts to the pass's inputs - #207

Draft
MaxGhenis wants to merge 7 commits into
mainfrom
reference-adversary-pinned-inputs
Draft

MaxGhenis wants to merge 7 commits into
mainfrom
reference-adversary-pinned-inputs

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This is stacked on #205, which pinned leaderboard_impact.py. It targets main so CI runs, which means #205's two commits (59c32d60, 95aaa8fb) appear in this diff until #205 merges. This PR's own change is 95aaa8fb..HEAD. Merge #205 first; this branch is then rebased onto main.

It gives the reference adversary's other scripts the same staged, sha256-checked inputs:

  • definition_conformance.py
  • publication_sources.py
  • engine_probe.py
  • build_proposals.py

The pins and the check now live in one module, scripts/pass_inputs.py, and leaderboard_impact.py uses it too.

No committed evidence file, reference, exclusion or score changes, and proposed_changes.json stays at 3a6e5920…. The diff touches only the audit's scripts and README, plus tests.

The bug, reproduced

These runs used the scripts as of #205's head (95aaa8fb), from a clean checkout of main. On main, release dashboard-data-20261006 (#202) has rewritten the run's payload and exclusion record, and #204 has rewritten one convention module. Record: bug_on_main.txt in the review folder named under Review.

The fix

scripts/pass_inputs.py (new) holds every pin, as commit, path and sha256:

git_input writes git show <commit>:<path> to scratch and stops (SystemExit) unless the bytes match the pin. stage_run, stage_fixes and stage_specs stage sets of files. use_staged_specs points policybench at the staged output definitions for the current process.

Each script stages everything it reads before it computes or writes anything, and reads only the staged copy. The three engine-side scripts also stage before they import the engine or start a worker:

  • definition_conformance.py:
    • stages the payload, references, sidecar, scenarios and the law table, then the conventions and output definitions;
    • points policybench at the staged definitions before importing anything else from it;
    • starts its workers with the spawn method and an initializer, so each worker stages its own conventions and definitions and none inherits the main process's system;
    • refuses a law-table mismatch instead of degrading to available: False;
    • still names each input by its repository path in the run record.
  • publication_sources.py does the same for the payload, references, sidecar, exclusion record and scenarios. It loses --from-facts, which built both reports from a saved file with no pin check.
  • engine_probe.py stages the scenarios, references and exclusion record, then the conventions and definitions through definition_conformance's builder.
  • build_proposals.py:
    • stages the references, the exclusion record and the conformance scan;
    • writes only to a required --out, which it creates exclusively, so it never writes a file that already exists;
    • refuses the committed proposed_changes.json however the path is spelled.
  • leaderboard_impact.py (Pin the reference adversary's leaderboard impact to the pass's inputs #205) imports the shared pins. pass_inputs() became stage_inputs(), and its scoring child now keeps the caller's PYTHONPATH after the checkout. It still scores with the checkout's policybench, output definitions included; its published-reproduction check covers that.
  • Each script checks that the pass_inputs it imported is the file beside it.

README.

  • The Inputs tables list every pin with its commit, sha256 and the scripts that stage it, including the 20 convention modules.
  • "Who checks the pins" now says every script does, and how the output definitions are handled.
  • Reproduce keeps the restore step only for the CLI commands, which record the payload path they are given. It lists the scripts' commands into scratch, and says what each regeneration test compares.

Tests.

  • tests/test_reference_adversary_inputs.py (new) holds the new tests below.
  • tests/working_tree_fence.py (new) is an audit hook that fails any open of the working tree's copy of a pinned input, benchmark_specs.json included. Armed through a generated sitecustomize, it reaches subprocesses and spawned workers.
  • tests/test_reference_adversary_impact.py is adapted to the shared module, and its slow regeneration now runs fenced. Pin the reference adversary's leaderboard impact to the pass's inputs #205's byte-refusal property test moved to the new file, where it covers every pin.

Verification

Byte-for-byte regeneration. Each script, at head f6774fd1, ran against main's working tree (#202's run, #204's module) and wrote to scratch, with the fence armed in every process (pytest_slow_r2.txt, 11 passed in 27 minutes on a heavily loaded machine):

Output Result
verification/definition_conformance.md identical
verification/definition_conformance.json identical except the values of run.seconds (wall time) and run.script_sha256 (the script's own hash; the committed file records 33ce4c3e…, the script as #200 merged it)
verification/publication_sources.md identical
verification/publication_sources.json identical except the value of meta.seconds
verification/probes/*.json (8) all identical, each from the arguments it records
verification/leaderboard_impact* (23) all identical
proposed_changes.json (via --out to scratch, fast suite) identical, sha256 3a6e5920…

A baseline of the unmodified scripts, with the run and modules restored from 8b4c0ca1, differs from the committed files in the same seconds lines only (manual_regeneration.txt). So the environment (uv sync --locked --extra dev --python 3.12, policyengine-us 2.15.17) reproduces the evidence, and the scripts' changes alter none of it.

Tests.

  • Fast suite: the two reference-adversary test files pass locally at f6774fd1 (50 passed). The full non-slow suite last ran locally before the review fixes (2,057 passed at 54a6a2ba, pytest_fast.txt); CI runs it at head.
  • The slow regenerations are deselected in CI.

Mutation check. Partly run; the rest is pending.

  • Run at f6774fd1 (mutants_r2.txt): 12 mutants that each reintroduce a working-tree read, drop the sha256 check, stage after engine work, drop the proposed_changes.json refusal, or put a pin in a script. Every one fails a fast test.
  • Written but not yet run (mutants_r2.py): 17 mutants that reintroduce the review's findings one at a time. The machine is saturated and local test runs are on hold.
  • Also pending: two mutants that only the fenced slow tests can catch (a main or always_zero reading the working tree's scenarios.csv, whose bytes equal the pin). Before the review fixes, the first of these passed the fast suite and failed the fenced slow regeneration with PermissionError (mutant_slow_only.txt).

Invariants

  • Pinned inputs. Every input byte a script reads equals the pinned bytes, whatever the working tree holds. Covered by:
    • the fence tests, which run build_proposals.py and engine_probe.py whole and each script's staging and reading in-process;
    • the slow regenerations, fenced in every process;
    • the differential test that the committed evidence records the same hashes as the pins.
  • Pinned output definitions. The engine helpers look up each output in the pinned benchmark_specs.json, not the checkout's. A copy of the package with a changed mapping, put first on the path, does not change the lookup. A process that has already loaded other definitions is refused.
  • Refusal before use. Any byte that is not the pinned byte is refused, and nothing is written. A script pointed at Stop scoring eight tax outputs and date the response window from the last answer (release dashboard-data-20261006) #202's run, or given a wrong convention or definitions pin, stops before any engine import, system build, worker or output write. Covered by a Hypothesis property over all 31 pins and per-script refusal tests whose engine entry points raise if reached.
  • Workers stage their own inputs. Both pools use the spawn context and an initializer that stages, whatever the platform default.
  • Reproduction. Each script regenerates its committed outputs byte for byte, except the value on the volatile lines named above. _identical_but requires the same key, indentation and punctuation on those lines, and the regenerated script_sha256 to be the current script's hash.
  • Existing files are never rewritten. build_proposals.py creates its output exclusively and refuses proposed_changes.json by path, so neither a hard link nor another checkout's copy is written.
  • One source of pins. No script other than pass_inputs.py holds a sha256 or names the run's path. The README tables list each pin with exactly the scripts that stage it, checked by recording which pins each script stages.

Review

  • Round 1 (GPT-6.1 Sol, subfleet run --task review --tier hard, at b4f2aac8): REQUEST_CHANGES, with two blockers and five other findings (review_sol.md).
  • All seven are addressed in f6774fd1; response_r1.md maps each finding to its fix and tests.
  • Round 2 has not been requested yet. This PR stays a draft until it has passed.

Everything cited above is in ~/reviews/policybench-reference-adversary-2026-10/other-scripts-pinned/.

Follow-ups

The Louisiana, Medicare Part B and payroll audits' impact scripts have the same working-tree pattern; #206 pins them. Their engine-side sweeps build latest_final from the working tree's convention modules and are a separate follow-up, which can reuse pass_inputs.py and the fence.

axiom: n/a: audit tooling only; no policy encoded or changed

🤖 Generated with Claude Code

MaxGhenis and others added 5 commits October 9, 2026 14:06
leaderboard_impact.py copied the working tree's run files and compared
against the working tree's payload, so regenerating after release
dashboard-data-20261006 (#202) scored that release's payload and
exclusions and silently rewrote 19 of the 23 committed
verification/leaderboard_impact* files (1,928 -> 1,920 scored cells).

The script now stages the run bundle from 8b4c0ca (release
dashboard-data-20260930) and proposed_changes.json from 4db91b5 (#200)
with git show, checks each file against its pinned sha256 and stops
before scoring on any mismatch. It reads only the staged copies.
--out-dir writes the evidence elsewhere.

tests/test_reference_adversary_impact.py checks the pins, the refusals
(including #202's payload), that a checkout whose working-tree run and
proposals are junk still scores the pinned bytes, and (slow) that a
full regeneration reproduces all 23 committed files byte for byte.

The README lists every pinned input with its sha256 and the scripts
that read it, and the Reproduce steps restore the whole run and
annotations directory from 8b4c0ca for the scripts that still read
the working tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A Hypothesis test flips, truncates or extends any pinned run file and
requires git_input to refuse it without writing the target, and to
accept the unchanged bytes. A differential check requires the payload
pin to equal definition_conformance.py's and publication_sources.py's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
definition_conformance.py, publication_sources.py, engine_probe.py and
build_proposals.py read the working tree's run, which release
dashboard-data-20261006 (#202) rewrote (payload and exclusion record),
and the two engine-side scripts built their reference system from the
working tree's convention modules, one of which #204 rewrote. Each now
stages every input through scripts/pass_inputs.py, which writes
`git show <commit>:<path>` to scratch and stops unless the bytes match
the pinned sha256, before the script computes or writes anything.
leaderboard_impact.py (#205) uses the same module, so the pins live in
one place.

build_proposals.py now writes to a required --out and refuses the
committed proposed_changes.json, which is pinned at 3a6e5920.

No evidence file, reference, exclusion or score changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…inked --out

_identical_but now requires the regenerated file to carry the same key,
indentation and punctuation on each volatile line, so only that one JSON
value may differ; a new test feeds it a changed value elsewhere, a renamed
key, re-indentation, a changed list and a dropped line. The
build_proposals refusal test adds a relative path and a symlink to the
committed proposed_changes.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
policybench-site Ready Ready Preview Oct 10, 2026 12:56am UTC

Request Review

@MaxGhenis
MaxGhenis changed the base branch from impact-evidence-pinned-inputs to main October 9, 2026 19:10
@MaxGhenis MaxGhenis closed this Oct 9, 2026
@MaxGhenis MaxGhenis reopened this Oct 9, 2026
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merge-order note, from a local merge reported by the session stacking on both PRs:

MaxGhenis added a commit that referenced this pull request Oct 9, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis marked this pull request as draft October 10, 2026 00:20
MaxGhenis and others added 2 commits October 9, 2026 20:20
…stage before the engine

Review of b4f2aac (GPT-6.1 Sol, hard tier) asked for these changes.

- policybench loads benchmark_specs.json itself, as policybench.scenarios
  is imported, so the engine-side scripts computed with the checkout's
  definitions while recording the pinned copy's hash. Each process of
  definition_conformance.py, publication_sources.py and engine_probe.py,
  workers included, now stages the pinned definitions and points
  policybench at them (pass_inputs.use_staged_specs) before importing
  anything else from the package. The test fence no longer exempts the
  package's own read for these scripts.
- publication_sources.py loses --from-facts, which built both reports
  from a saved file with no pin check.
- The conventions and definitions are staged before the engine is
  imported, hooked or handed to a worker, not after.
- Workers are started with the spawn method and an initializer that
  stages, so a forked worker cannot inherit the main process's system.
- build_proposals.py creates --out exclusively, so a hard link to the
  committed proposed_changes.json or another checkout's copy is refused
  like any existing file.
- leaderboard_impact.py keeps the caller's PYTHONPATH for its scoring
  child, and its slow regeneration now runs fenced in both processes.
- Each script checks that the pass_inputs it imported is the one beside
  it.

No evidence file, reference, exclusion or score changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 609d5a7f Deployed Oct 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant