Skip to content

feat(parsec): add the h4 arm driver and three trial-integrity detectors - #695

Open
bcarmeli wants to merge 1 commit into
mainfrom
parsec/h4-t1-e3-driver
Open

bcarmeli wants to merge 1 commit into
mainfrom
parsec/h4-t1-e3-driver

Conversation

@bcarmeli

@bcarmeli bcarmeli commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Adds scripts/parsec/h4_t1_e3/ — the driver for the h4 held-out arms, plus three detectors for failure modes that currently produce results indistinguishable from real ones.

The detectors are the part worth reviewing

This suite has three distinct ways a bad trial looks like a valid score. The harness catches only the first.

failure how it looks detector
1 harness error reward null, errored_trials increments built in
2 agent cancelled scored 0.0, errored_trials still 0 audit_trials.py
3 simulator errored or invented data plausible partial credit audit_absence_traps.py

(2) A cancelled agent still writes a reward.json; the verifier's completion gate multiplies it by zero. The only reliable signal is the agent's own final result event, exactly as tests/verify.py reads it.

Do not detect this by grepping transcripts for error strings. A structured error_during_execution contains no matching text, so the search returns nothing and its silence reads as proof the trial was sound. I made that mistake and it produced a confidently wrong "regression" claim.

Correcting one such trial moved a task from 0.800 to 1.000 — which then matched the reference exactly.

(3) The trial completes, the agent asks exactly the right question, tool calls score 1.0, and the reward is believable. Observed: the simulator fabricated a complete Azure subscription — tenant id, resource group, status — for an entity absent from its own seed, in a task whose entire trap is that the subscription does not exist. The agent reported what it was handed and lost the assertion.

These failures target the tasks that test whether the agent is honest about missing data, so they make an honest agent look worse than a credulous one. Counts rose 2 → 8 → 11 across three arms on byte-identical seed data.

compare_arms.py applies detector (2) to every arm uniformly — each arm carries a dead trial of its own, so filtering one and not the others would compare different things — and prints per-trial vectors with the distribution named. A single mean per task is actively misleading: of three shapes present (stable / stable-plus-outlier / bimodal), the mean is the measurement for only the first.

@osher — detectors (2) and (3) are arguably not parsec-specific; any benchmark with a completion gate can score a dead trial as 0.0. They're under scripts/parsec/ because that's where they work today and because compare_arms/build_queue import them. If there's a better home for generic trial-integrity checks, happy to move them.

Bundle location

$PARSEC_H4_BUNDLE, no fallback default — same rule parsec_paths.py states for PARSEC_V4N: there is no single "the" bundle to guess at, and guessing wrong seeds a run from one arm's prompts while reporting it as another's. Bundles live on parsec-history (artifacts/h4/<arm>/, beside v1/v2/v4) so main carries no parsec artifacts; a run needs that branch checked out too. Verified: unset → clear RuntimeError, bogus path → named, valid path → proceeds.

Adapted, not copied

  • path depths bumped for scripts/parsec/ (parents[2] → [3]), adapters source re-pointed after chore: move parsec v4 design docs to parsec-history #581's move
  • 37 stale v4_t4_e1 references corrected, two functional rather than cosmetic: --dry-run printed project paths that would never be created, and the generated project header named the wrong arm
  • two dead v4 task-list constants removed (44 unused lines)
  • collect_cost_time.py dropped entirely — a verbatim v4_t4_e1 copy, never adapted, that would have read the wrong glob and written to the wrong path

Independent of #694 (which touches v4_t2_e1's adapter); both can land in either order.

The driver (scaffold_projects, run_one_task, preflight_check, collect_results,
compare_e1_e2) mirrors scripts/parsec/v4_t2_e1's set. run_one_task gains
--n-trials so a task that lost trials to infrastructure can be topped up to
five rather than re-run: an infra-killed trial yields no reward at all, so the
survivors stay valid measurements.

The prompt bundle is located by $PARSEC_H4_BUNDLE, with no fallback default,
for the reason parsec_paths.py gives about PARSEC_V4N — there is no single
"the" bundle to guess at, and guessing wrong would seed a run from one arm's
prompts while reporting it as another's. The bundles live on parsec-history
(artifacts/h4/<arm>/, beside v1/v2/v4) so main carries no parsec artifacts,
which means a run needs that branch checked out too.

The three detectors exist because this suite has three distinct ways for a bad
trial to look like a real score, and the harness catches only the first:

  1. harness error — reward null, errored_trials increments. Already handled.

  2. agent did not finish — audit_trials.py. A cancelled agent still writes a
     reward.json; the verifier's completion gate multiplies it by zero, so the
     trial reads as a legitimate 0.0 while errored_trials stays 0. The only
     reliable signal is the agent's own final result event, exactly as
     tests/verify.py reads it. Do NOT substitute a text search of the
     transcript: a structured error_during_execution contains no matching
     string, so the search returns nothing and its silence reads as proof the
     trial was sound. Correcting one such trial moved a task from 0.800 to
     1.000, which then matched the reference exactly.

  3. simulator answered with an error or an invention — audit_absence_traps.py.
     The trial completes, the agent asks the right question and scores 1.0 on
     tool calls, and the reward is a believable partial score. Seen as a
     fabricated Azure subscription in a task whose whole trap is that the
     subscription does not exist. These failures target the tasks that test
     whether the agent is honest about missing data, so they make an honest
     agent look worse than a credulous one. Counts rose 2 -> 8 -> 11 across
     three arms on byte-identical seed data.

compare_arms.py applies detector 2 to every arm uniformly — each arm carries a
dead trial of its own, so filtering one and not the others would compare
different things — and prints per-trial vectors with the distribution named. A
single mean per task is misleading here: of three shapes present (stable,
stable-plus-outlier, bimodal) the mean is the measurement for only the first.

build_queue.py emits each task's shortfall of CLEAN trials, so a top-up pass
asks only for what is missing.

Adapted for this location and arm, not copied verbatim: path depths bumped for
scripts/parsec/ (parents[2] -> [3]), the adapters source re-pointed after #581's
move, and 37 stale v4_t4_e1 references corrected — two of which were functional
rather than cosmetic (--dry-run printed project paths that would never be
created, and the generated project header named the wrong arm). Two dead v4
task-list constants (44 unused lines) removed, and collect_cost_time.py dropped
entirely: it was a verbatim v4_t4_e1 copy, never adapted, that would have read
the wrong glob and written to the wrong path.

Signed-off-by: boazc <boazc@il.ibm.com>
@skillberry-bot skillberry-bot added enhancement New feature or request honesty Honesty/consistency of reported results (brand-critical) paper-experiment Publication-grade held-out experiments for the NAACL/ACL paper observability Live run visibility, logging, tracing tech-debt Dead code, duplication, refactors labels Oct 8, 2026
@skillberry-bot

Copy link
Copy Markdown
Contributor

🏷️ Automatic Labeling

I've analyzed this pull request and added the following labels:

  • enhancement - honesty - paper-experiment - observability - enhancement - honesty - paper-experiment - observability - tech-debt

These labels were selected based on the PR title, description, and changed files. If you believe any labels are incorrect or missing, feel free to adjust them manually.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request honesty Honesty/consistency of reported results (brand-critical) observability Live run visibility, logging, tracing paper-experiment Publication-grade held-out experiments for the NAACL/ACL paper tech-debt Dead code, duplication, refactors

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants