Repository navigation
Conversation
The driver (scaffold_projects, run_one_task, preflight_check, collect_results,
compare_e1_e2) mirrors scripts/parsec/v4_t2_e1's set. run_one_task gains
--n-trials so a task that lost trials to infrastructure can be topped up to
five rather than re-run: an infra-killed trial yields no reward at all, so the
survivors stay valid measurements.
The prompt bundle is located by $PARSEC_H4_BUNDLE, with no fallback default,
for the reason parsec_paths.py gives about PARSEC_V4N — there is no single
"the" bundle to guess at, and guessing wrong would seed a run from one arm's
prompts while reporting it as another's. The bundles live on parsec-history
(artifacts/h4/<arm>/, beside v1/v2/v4) so main carries no parsec artifacts,
which means a run needs that branch checked out too.
The three detectors exist because this suite has three distinct ways for a bad
trial to look like a real score, and the harness catches only the first:
1. harness error — reward null, errored_trials increments. Already handled.
2. agent did not finish — audit_trials.py. A cancelled agent still writes a
reward.json; the verifier's completion gate multiplies it by zero, so the
trial reads as a legitimate 0.0 while errored_trials stays 0. The only
reliable signal is the agent's own final result event, exactly as
tests/verify.py reads it. Do NOT substitute a text search of the
transcript: a structured error_during_execution contains no matching
string, so the search returns nothing and its silence reads as proof the
trial was sound. Correcting one such trial moved a task from 0.800 to
1.000, which then matched the reference exactly.
3. simulator answered with an error or an invention — audit_absence_traps.py.
The trial completes, the agent asks the right question and scores 1.0 on
tool calls, and the reward is a believable partial score. Seen as a
fabricated Azure subscription in a task whose whole trap is that the
subscription does not exist. These failures target the tasks that test
whether the agent is honest about missing data, so they make an honest
agent look worse than a credulous one. Counts rose 2 -> 8 -> 11 across
three arms on byte-identical seed data.
compare_arms.py applies detector 2 to every arm uniformly — each arm carries a
dead trial of its own, so filtering one and not the others would compare
different things — and prints per-trial vectors with the distribution named. A
single mean per task is misleading here: of three shapes present (stable,
stable-plus-outlier, bimodal) the mean is the measurement for only the first.
build_queue.py emits each task's shortfall of CLEAN trials, so a top-up pass
asks only for what is missing.
Adapted for this location and arm, not copied verbatim: path depths bumped for
scripts/parsec/ (parents[2] -> [3]), the adapters source re-pointed after #581's
move, and 37 stale v4_t4_e1 references corrected — two of which were functional
rather than cosmetic (--dry-run printed project paths that would never be
created, and the generated project header named the wrong arm). Two dead v4
task-list constants (44 unused lines) removed, and collect_cost_time.py dropped
entirely: it was a verbatim v4_t4_e1 copy, never adapted, that would have read
the wrong glob and written to the wrong path.
Signed-off-by: boazc <boazc@il.ibm.com>
Contributor
|
🏷️ Automatic Labeling I've analyzed this pull request and added the following labels:
These labels were selected based on the PR title, description, and changed files. If you believe any labels are incorrect or missing, feel free to adjust them manually. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
scripts/parsec/h4_t1_e3/— the driver for the h4 held-out arms, plus three detectors for failure modes that currently produce results indistinguishable from real ones.The detectors are the part worth reviewing
This suite has three distinct ways a bad trial looks like a valid score. The harness catches only the first.
null,errored_trialsincrementserrored_trialsstill 0audit_trials.pyaudit_absence_traps.py(2) A cancelled agent still writes a
reward.json; the verifier's completion gate multiplies it by zero. The only reliable signal is the agent's own final result event, exactly astests/verify.pyreads it.Correcting one such trial moved a task from 0.800 to 1.000 — which then matched the reference exactly.
(3) The trial completes, the agent asks exactly the right question, tool calls score 1.0, and the reward is believable. Observed: the simulator fabricated a complete Azure subscription — tenant id, resource group, status — for an entity absent from its own seed, in a task whose entire trap is that the subscription does not exist. The agent reported what it was handed and lost the assertion.
These failures target the tasks that test whether the agent is honest about missing data, so they make an honest agent look worse than a credulous one. Counts rose 2 → 8 → 11 across three arms on byte-identical seed data.
compare_arms.pyapplies detector (2) to every arm uniformly — each arm carries a dead trial of its own, so filtering one and not the others would compare different things — and prints per-trial vectors with the distribution named. A single mean per task is actively misleading: of three shapes present (stable / stable-plus-outlier / bimodal), the mean is the measurement for only the first.@osher — detectors (2) and (3) are arguably not parsec-specific; any benchmark with a completion gate can score a dead trial as 0.0. They're under
scripts/parsec/because that's where they work today and becausecompare_arms/build_queueimport them. If there's a better home for generic trial-integrity checks, happy to move them.Bundle location
$PARSEC_H4_BUNDLE, no fallback default — same ruleparsec_paths.pystates forPARSEC_V4N: there is no single "the" bundle to guess at, and guessing wrong seeds a run from one arm's prompts while reporting it as another's. Bundles live onparsec-history(artifacts/h4/<arm>/, besidev1/v2/v4) so main carries no parsec artifacts; a run needs that branch checked out too. Verified: unset → clearRuntimeError, bogus path → named, valid path → proceeds.Adapted, not copied
scripts/parsec/(parents[2]→[3]), adapters source re-pointed after chore: move parsec v4 design docs to parsec-history #581's movev4_t4_e1references corrected, two functional rather than cosmetic:--dry-runprinted project paths that would never be created, and the generated project header named the wrong armcollect_cost_time.pydropped entirely — a verbatim v4_t4_e1 copy, never adapted, that would have read the wrong glob and written to the wrong pathIndependent of #694 (which touches
v4_t2_e1's adapter); both can land in either order.