Skip to content

agent-optimize: rewrite SKILL.md loop to use graph/Pareto/targeted-eval/model-routing - #671

Merged
OsherElhadad merged 3 commits into
mainfrom
agent-optimize-rewrite-665-ws3
Oct 6, 2026
Merged

OsherElhadad merged 3 commits into
mainfrom
agent-optimize-rewrite-665-ws3

Conversation

@OsherElhadad

Copy link
Copy Markdown
Collaborator

Summary

Issue #665's final workstream (ws3). A forensic run (run_20261003_184253) showed that the
agent-optimize SKILL.md's loop was being followed in prose but not in behavior, because the
engine pieces that landed on main from PRs #666-#669 (model_routing, gate_mode: pareto,
CandidateGraph/plan_round.py/EvaluationPlan) were never wired into the instructions an agent
actually reads:

  • 8 iterations, 1 candidate per round, zero parallelism, with no stated reason N was always 1.
  • 7 of 8 candidates paid the full 90-rollout val gate regardless of how narrow their hypothesis was.
  • 2 accepted candidates targeting disjoint clusters never got merged, despite merge_compliance_warning
    firing twice.

This PR rewrites the loop so following it mechanically produces the new behavior:

  1. Branch planning: step 2 now calls plan_round.py on the diagnose output to get a branch plan
    before any candidate dir is created, and creates exactly the number of slots it proposes (varies
    with root-cause structure — not a fixed bucket-A/B default).
  2. Staged, targeted evaluation: step 5 builds and persists an EvaluationPlan
    (build_evaluation_plan()/persist_evaluation_plan()) per candidate and screens at its own stage
    (affected_tasks/regression_sentinels), escalating only on evidence of a wider footprint — instead
    of defaulting straight to full val.
  3. Required merge before sealing: merging disjoint-cluster accepted candidates is now an explicit,
    non-optional step framed as a visible protocol violation to skip, not advisory prose.
  4. Model routing: added guidance to call resolve_model(role, spec) at each judgment point (plan,
    root-cause, propose, implement, evaluation-analysis, merge, synthesis) and report the resolved model.
  5. Pareto-gate guidance: when capevolve.yaml declares objectives, use gate_mode: pareto via
    gate_check.py/commit.py (never round.py, whose screen/merge machinery stays single-metric) and
    read the result as a frontier (better/worse/tied per objective), not a disguised boolean.
  6. references/algorithm.md updated to match — new "Branch planning", "Evaluation plans", "Model
    routing" and "Pareto acceptance" sections, and the old "default to N≥3" framing corrected to point at
    plan_round.py instead of asserting a fixed constant.

Backward compatible by construction: a single-objective run with no objectives declared, or no
model_routing block, behaves exactly as before; plan_round.py legitimately proposes N=1 for a
trivial single-cluster round (that is not the anti-pattern — never running it, or overriding its
estimate unstated, is).

Code changes beyond the two docs

  • skills/algorithms/agent-optimize/scripts/gate_check.py: --mode pareto was declared in core
    (cap_evolve.gate, PR gate: multi-objective Pareto acceptance mode #667) but never reachable from any script — added it to GATE_MODES plus
    --objectives/--metrics-candidate/--metrics-current/--metrics-stderr-candidate/
    --metrics-stderr-current, wired through to gate.decide(), with ParetoObjectiveError surfaced as
    a clean JSON refusal (exit 2) rather than a traceback. round.py still forwards --mode verbatim
    from gate_check.GATE_MODES, so a round that mistakenly tries --mode pareto fails loudly
    (GateCheckFailed) rather than silently — the SKILL.md explicitly documents pareto mode as reachable
    only through gate_check.py+commit.py by hand, never through round.py.
  • templates/project/capevolve.yaml: documented gate_mode: pareto and the optional objectives:
    block (comment-only; the spec loader already reads arbitrary keys generically, so no schema code
    changes were needed).
  • core/tests/test_skill_code_claims.py: extended with tests pinning (a) plan_round.py's JSON output
    shape against what SKILL.md's step 2 reads, (b) model_routing.ROLES against the roles table in
    SKILL.md, (c) gate_check.py's pareto-mode CLI flags existing and actually reaching
    gate.decide()'s kwargs (not just accepted and ignored), (d) evaluation_plan.py's stage constants
    against SKILL.md's stage guidance, (e) merge_search.check_merge_compliance still being the function
    measure.py calls at finalize time, per SKILL.md's "Stop & seal" claim.

Known limitation (left out of scope, by design)

Full automatic Pareto-gate wiring through round.py's batched screen/merge/null-control cascade was
not attempted — that machinery is built around one scalar delta end to end, and bolting a
multi-objective path onto it is a larger change than this workstream's scope (SKILL.md + the two scripts
it names). A pareto-gated candidate is always gated by hand, one tag at a time, through gate_check.py
directly — documented explicitly in both SKILL.md and algorithm.md rather than left as a silent gap.

Test plan

  • python -m pytest core/tests -q — 1884 passed, 1 skipped, 2 failed (both pre-existing and
    environment-dependent — test_tau2_airline_eval_cost.py's cost_source assertions fail
    identically on unmodified main in this environment, confirmed by stashing this PR's changes and
    re-running; unrelated to this change).
  • python skills/algorithms/agent-optimize/scripts/check.py → "ok": true.
  • python skills/_registry/lint_skills.py skills → no agent-optimize errors (SKILL.md body kept
    under the repo's 500-line/5000-token budget).

…al/model-routing

Issue #665's forensic run showed the SKILL.md prose never used the new
machinery (plan_round.py, EvaluationPlan, gate_mode: pareto, model_routing,
merge_search.py) even though it already existed: 8 rounds, 1 candidate each,
full-val paid regardless of hypothesis scope, two mergeable accepts never
merged.

- Step 2 now calls plan_round.py for a branch plan (varying N, not a fixed
  bucket-A/B default) before any candidate dir is created.
- Step 5 builds and persists an EvaluationPlan per candidate and screens at
  its stage (affected_tasks/regression_sentinels), escalating only on
  evidence, instead of defaulting to full val.
- Merging disjoint-cluster accepted candidates before sealing is now an
  explicit REQUIRED step, framed as a visible protocol violation to skip.
- Added model-routing guidance (resolve_model per decision role) and
  Pareto-gate guidance (gate_mode: pareto via gate_check.py/commit.py,
  reading a frontier result instead of a single accept/reject).
- references/algorithm.md updated to match: new "Branch planning",
  "Evaluation plans", "Model routing" and "Pareto acceptance" sections,
  and the old "default to N>=3" framing corrected to point at plan_round.py.
- gate_check.py: wired --mode pareto end to end (new GATE_MODES entry,
  --objectives/--metrics-*/--metrics-stderr-* flags reaching gate.decide).
- templates/project/capevolve.yaml: documented gate_mode: pareto + objectives.
- core/tests/test_skill_code_claims.py: new tests pinning plan_round.py's
  JSON shape, the model_routing roles table, gate_check.py's pareto wiring,
  the evaluation_plan stage constants, and merge_search's finalize-time hook
  against SKILL.md's claims.

SKILL.md stayed under the repo's 500-line/5000-token body budget by moving
rationale into references/algorithm.md and trimming prose elsewhere;
mechanics are unchanged.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
@skillberry-bot skillberry-bot added algorithm Optimization algorithms: GEPA / SkillOpt / hill-climb enhancement New feature or request documentation Improvements or additions to documentation tech-debt Dead code, duplication, refactors priority-p1 High impact labels Oct 6, 2026
@skillberry-bot

Copy link
Copy Markdown
Contributor

🏷️ Automatic Labeling

I've analyzed this pull request and added the following labels:

  • algorithm - enhancement - documentation - tech-debt - algorithm - enhancement - documentation - tech-debt - priority-p1

These labels were selected based on the PR title, description, and changed files. If you believe any labels are incorrect or missing, feel free to adjust them manually.

round.py's --mode choices were gate_check.GATE_MODES verbatim, so adding "pareto"
to that shared list (this PR) made `round.py --mode pareto` CLI-acceptable. But
round.py's _gate() never forwards --objectives/--metrics-* to gate_check.py, so
`--mode pareto` always crashed at runtime (gate_check.py raises ParetoObjectiveError
and exits 2, round.py raises GateCheckFailed) -- after a wasted gate_check.py
subprocess call.

round.py now derives its own choices as gate_check.GATE_MODES minus "pareto", so
this is rejected at argument-parsing time instead, with a comment explaining why
pareto is excluded on purpose. Also fixes test_round_mode_choices_are_exactly_gate_check_mode_choices
(now test_round_mode_choices_are_gate_check_mode_choices_minus_pareto), which
asserted the two lists were identical -- that invariant is now "identical except
for pareto" by design, and the comment in gate_check.py describing round.py's
--mode as "restricted to the single-metric modes" is now accurate instead of
aspirational.

Adds a regression test confirming `round.py --mode pareto` is rejected at
argument-parsing time, before any gate_check subprocess call.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Osher-Elhadad <Osher.Elhadad@ibm.com>
@OsherElhadad

Copy link
Copy Markdown
Collaborator Author

Fixed a confirmed review finding: round.py's own --mode choices were gate_check.GATE_MODES verbatim, so this PR's addition of "pareto" to that shared list made round.py --mode pareto CLI-acceptable. But round.py's _gate() never forwards --objectives/--metrics-* to gate_check.py, so --mode pareto always crashed at runtime (gate_check.py raises ParetoObjectiveError and exits 2, round.py raises GateCheckFailed) — after a wasted gate_check.py subprocess call.

Fix (skills/algorithms/agent-optimize/scripts/round.py):

  • --mode choices now derive as [m for m in gate_check.GATE_MODES if m != "pareto"], with a comment explaining the exclusion (both why --objectives/--metrics-* aren't forwarded, and that round.py's batched cascade assumes one scalar delta per candidate). round.py --mode pareto now fails fast at argument-parsing time with argparse's own clear "invalid choice" error, instead of after a wasted subprocess call.
  • Updated test_round_mode_choices_are_exactly_gate_check_mode_choices (now test_round_mode_choices_are_gate_check_mode_choices_minus_pareto in core/tests/test_a_failed_gate_is_not_a_verdict.py) to assert the new intended invariant: identical to gate_check.GATE_MODES except for the deliberately-excluded "pareto".
  • Added core/tests/test_round_rejects_pareto_mode.py, a regression test confirming round.py --mode pareto is rejected at argument-parsing time (exit code 2, before any run-dir/project/gate_check work).
  • The existing comment in gate_check.py claiming round.py's --mode "stays restricted to the single-metric modes" was aspirational rather than accurate before this fix; it's now true and needed no further edit.

Test results: python -m pytest core/tests -q → 1885 passed, 1 skipped, 2 pre-existing unrelated failures in test_tau2_airline_eval_cost.py (not touched by this change).

argparse's invalid-choice error quotes each choice individually on some
Python versions (3.12+) and not others (3.11), so the single literal
substring 'paired, significant, strict, threshold' matched locally but
not in CI's Python 3.11. Check each mode name independently instead.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
@OsherElhadad
OsherElhadad merged commit 1d40d16 into main Oct 6, 2026
15 checks passed
@OsherElhadad
OsherElhadad deleted the agent-optimize-rewrite-665-ws3 branch October 6, 2026 09:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

algorithm Optimization algorithms: GEPA / SkillOpt / hill-climb documentation Improvements or additions to documentation enhancement New feature or request priority-p1 High impact tech-debt Dead code, duplication, refactors

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants