Skip to content

optimizer: native Pareto archive gate + epsilon-constraint mode in round.py - #690

Merged
OsherElhadad merged 2 commits into
mainfrom
pareto-archive-gate-684
Oct 8, 2026
Merged

OsherElhadad merged 2 commits into
mainfrom
pareto-archive-gate-684

Conversation

@OsherElhadad

Copy link
Copy Markdown
Collaborator

Summary

Closes the #684 item 1 bug: round.py excluded "pareto" from its own --mode choices because _gate() never forwarded --objectives/--metrics-* to gate_check.py, so a spec declaring gate_mode: pareto was silently gated single-objective (reward-only). This fixes it properly end-to-end rather than excluding it, and adds epsilon_constraint as a second, simpler multi-objective mode (item 2).

  • core/cap_evolve/pareto_archive.py (new): ParetoArchive — a persistent bounded non-dominated frontier (Cheng & Li 1997, cited in the MOO survey grounding issue agent-optimize v3: native Pareto gating, archive-based search, real merging, cost-reduction levers #684). Insertion requires non-domination against the archive AND a significant win on >= 1 objective, both using the SAME per-objective significance floor gate.py's gate: multi-objective Pareto acceptance mode #667 fix established (reuses gate._objective_state) and the SAME dominance arithmetic (selection.dominates) rather than reimplementing either. Over capacity (default 15), evicts the point with the smallest NSGA-II crowding distance. Persisted as versioned JSON at $RUN_DIR/pareto_archive.json, so it survives across rounds and process restarts.
  • core/cap_evolve/gate.py: new gate_mode: epsilon_constraint (Haimes, Lasdon & Wismer 1971's bounded-objective-function method) — maximize reward subject to constraints: [{name, max}], with the same noise-floor discipline: a constraint satisfied only nominally (its own measurement noise could put it over the ceiling) is rejected, not accepted.
  • core/cap_evolve/harness.py: candidate_cost_objective() — derives a candidate's cost mean/SE from its own already-persisted per-task cost_usd data (no new field).
  • skills/algorithms/agent-optimize/scripts/round.py: removed the ROUND_MODES exclusion of "pareto"; _gate() now forwards --objectives/--metrics-*/--constraints; for --mode pareto, round.py builds/updates the persistent archive and a candidate's accept/reject is now literally "did the archive give it a slot", overriding the one-shot pairwise verdict gate_check.py alone would have produced.
  • skills/algorithms/agent-optimize/scripts/gate_check.py: epsilon_constraint added to GATE_MODES, new --constraints flag.
  • Existing single-metric modes (paired/significant/strict/threshold) are completely untouched — no code path they use was touched.

Design decisions

  • Archive capacity default: 15 (DEFAULT_CAPACITY in pareto_archive.py) — a flat, unmeasured default (ponytail: comment), configurable later via capevolve.yaml if a real run shows it binds.
  • Eviction rule: classic NSGA-II crowding distance — for each objective, sort archive points by that objective's value, give the two boundary (extreme) points infinite distance, and sum each interior point's normalized gap to its neighbors across every objective. The point with the smallest total distance is evicted. Boundary/extreme points on every objective are therefore never evicted first.
  • Where objective/metric values come from: only "cost" is derivable today — harness.py's existing per-task cost_usd (from issue agent-optimize v2: dashboard multi-objective visibility, optimizer-diagnosis flow view, bigger/more-parallel candidates #676's instrumentation) is the only secondary metric persisted per task alongside reward; neither "latency" nor "tokens" is persisted per-task, so declaring either without an explicit --metrics-* value raises the same ParetoObjectiveError refusal gate.py's own cost→latency→tokens fallback chain already uses (refuse, don't silently degrade).
  • round.py's cascade (screen/merge/null-control) is unchanged for pareto/epsilon_constraint — those steps only ever triaged on scalar reward anyway. Only the full-val gate step is pareto-native; two purely-diagnostic re-gate loops (control_relative, verdict_by_reference) are skipped for the two new modes since they don't yet thread objective/constraint metrics through (documented with a ponytail:-style comment — add if a multi-objective run's drift turns out to matter).

Test plan

  • core/tests/test_pareto_archive.py: archive insertion/eviction correctness on synthetic objective vectors (dominance, non-dominated tradeoffs, maintenance pruning, crowding-distance eviction keeps the frontier spread), the significance floor refusing a noise-level cost "win" (the same adversarial case as the earlier gate.py bug — tie on reward + float-noise cost delta must NOT get archived), and persistence surviving a simulated process restart (write, reload, confirm state matches).
  • core/tests/test_gate_epsilon_constraint.py: accept/reject correctness, the noise-level-constraint-satisfaction adversarial case, missing value/stderr refusals, paired-vs-significant fallback.
  • core/tests/test_round_pareto_archive.py: end-to-end — round.py --mode pareto with a declared objectives config produces real pareto_archive/dominance fields in its output (not paired-mode scalar fields), asserted directly since that's the exact symptom that proved the real-run bug; the archive persists across SEPARATE round.py subprocess invocations (genuine process restarts); epsilon_constraint requires a non-empty constraints: config and runs end-to-end.
  • Updated test_a_failed_gate_is_not_a_verdict.py's mode-list invariant (round.py's --mode choices now equal gate_check.py's, no gap) and removed test_round_rejects_pareto_mode.py (asserted the now-removed exclusion).

python -m pytest core/tests -q: 1933 passed, 2 failed, 5 skipped — the 2 failures are the pre-existing, unrelated test_tau2_airline_eval_cost.py failures named in the issue (an environment/model-name mismatch in that test's fixture, nothing to do with this change).

Confirmed directly (not just inferred): round.py --mode pareto end-to-end now produces real pareto_archive/dominance output — see test_round_pareto_archive.py::test_round_pareto_produces_archive_fields_not_paired_scalar_fallback.

Closes #684 items 1-2.

…und.py

round.py excluded "pareto" from its own --mode choices because _gate() never
forwarded --objectives/--metrics-* to gate_check.py, so a spec declaring
gate_mode: pareto was silently gated single-objective. Fix it end-to-end
instead of excluding it (#684 items 1-2):

- cap_evolve/pareto_archive.py: a persistent bounded ParetoArchive (Cheng &
  Li 1997) — non-domination ranking (reuses selection.dominates), the same
  per-objective significance floor gate.py's #667 fix established (reuses
  gate._objective_state), and NSGA-II crowding-distance eviction once over
  capacity (default 15). Persisted as versioned JSON in the run dir.
- cap_evolve/gate.py: new gate_mode: epsilon_constraint (Haimes et al. 1971
  bounded-objective-function method) — maximize reward subject to
  constraints: [{name, max}], with the same noise-floor discipline on the
  ceiling check.
- cap_evolve/harness.py: candidate_cost_objective(), deriving a candidate's
  cost mean/SE from its own already-persisted per-task cost_usd data.
- round.py: removed the ROUND_MODES exclusion; _gate() forwards
  --objectives/--metrics-*/--constraints; a pareto-gated candidate's
  accept/reject is now decided by the archive's own insertion result, not a
  one-shot pairwise read.
- gate_check.py: epsilon_constraint added to GATE_MODES, --constraints flag.

Existing single-metric gate modes (paired/significant/strict/threshold) are
untouched.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
@skillberry-bot skillberry-bot added algorithm Optimization algorithms: GEPA / SkillOpt / hill-climb bug Something isn't working enhancement New feature or request documentation Improvements or additions to documentation labels Oct 8, 2026
@skillberry-bot

Copy link
Copy Markdown
Contributor

🏷️ Automatic Labeling

I've analyzed this pull request and added the following labels:

  • algorithm - bug - enhancement - algorithm - bug - enhancement - documentation

These labels were selected based on the PR title, description, and changed files. If you believe any labels are incorrect or missing, feel free to adjust them manually.

p1 = _round(run_dir, project, "cand_1", "--mode", "pareto")
assert p1.returncode == 0, p1.stdout + p1.stderr
out1 = json.loads(p1.stdout)
size_after_round_1 = out1["pareto_archive"]["size_total"] if "size_total" in out1["pareto_archive"] \
…ence

try_insert only validated the NEW candidate's values against the declared
objectives; an existing archive point (e.g. round.py's baseline seed when
harness.candidate_cost_objective reports no cost) could already be missing
one, so the next candidate with full values crashed _states_vs with a raw
KeyError instead of the designed ParetoObjectiveError refusal.

Also reuse rundir.py's _file_lock + _atomic_write (the same pattern used for
state.json's identical load->mutate->save cycle) in ParetoArchive.save/
load_or_create instead of a plain write_text/read_text, so two processes
touching the same run dir's archive can't tear or lose a write.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
@OsherElhadad

Copy link
Copy Markdown
Collaborator Author

Pushed a fix for both confirmed review findings in `core/cap_evolve/pareto_archive.py` (commit 2e91f49):

1. Crash bug. `_states_vs` now validates that an existing archive point also has every declared objective before diffing against it, raising the designed `ParetoObjectiveError` (with a clear "existing archive point is missing declared objective ... (have: ...)" message, same shape as gate.py's own missing-metric refusal) instead of letting a bare dict lookup raise `KeyError`. This covers the case directly, independent of round.py's seeding path, since the archive point itself can be incomplete for other reasons too (e.g. a version-mismatch reload, or any future caller that appends a point directly).

Verified the exact reviewer repro no longer crashes:
```python
archive.points.append(ArchivePoint(tag='baseline', values={'reward': 0.5}, stderr={'reward': 0.0}))
archive.try_insert('cand_1', {'reward': 0.6, 'cost': 1.0}, {'reward': 0.0, 'cost': 0.1})

-> now raises ParetoObjectiveError, not KeyError

```
Added as a regression test: `test_insert_against_existing_point_missing_declared_objective_raises_not_keyerror` in `core/tests/test_pareto_archive.py`.

2. Persistence locking. `ParetoArchive.save()`/`load_or_create()` now reuse `rundir.py`'s existing `_file_lock` + `_atomic_write` helpers (the same pattern `state.json`'s read-modify-write methods already use) instead of plain `write_text`/`read_text`, so a concurrent writer can't tear `pareto_archive.json` or race another process's save.

Tests: `python -m pytest core/tests -q` → 1938 passed, 1 skipped, 2 failed — the 2 failures are the pre-existing unrelated `test_tau2_airline_eval_cost.py::test_priced_run_is_reported_as_before` / `test_rollout_says_whether_its_cost_was_measured` (cost_source `"tau2"` vs `"partial_models"`), untouched by this change.

@OsherElhadad
OsherElhadad merged commit a46ff17 into main Oct 8, 2026
15 checks passed
OsherElhadad added a commit that referenced this pull request Oct 8, 2026
… add utopia-nadir normalization, fix optimizer-cost telemetry gap (#691)

Issue #684 items 9 and 10 (dashboard verification + optimizer-cost telemetry fix).

Telemetry (item 10): optimizer_seconds/optimizer_usd were always 0 in agent-mode runs
because meter.py's automatic metering only works under host.py's headless driver (it
reads a claude-code session log that only exists there) -- a SKILL.md compliance gap,
not a framework instrumentation gap. commit.py now logs an optimizer_cost_warning when
real wall-clock time clearly passed since the run's previous decision but the new
commit still carries zero optimizer cost, surfaced in commit.py's own warnings list
and on the dashboard's node detail. SKILL.md's guidance was strengthened and trimmed to
stay within the repo's 5000-token skill-body budget; the long rationale moved to
algorithm.md's new "Optimizer cost telemetry" section.

Dashboard verification (item 9): built a synthetic multi-candidate reward+cost fixture
with genuine dominance relationships (two non-dominated frontier candidates, one
strictly dominated candidate) and confirmed dashboard.reduce_run's existing node shape
(val/cost_usd/per_task/per_task_metrics) correctly feeds the frontend's already-tested
Pareto frontier logic and per-objective Tasks-tab deltas. Also wired through round.py's
real pareto-gate output schema (confirmed against the now-merged-adjacent PR #690):
objective_values/objective_stderrs/pareto_archive per candidate, read off the same
gate-table lookup dashboard.py already used for verdict_stable -- no round.py or
commit.py changes needed.

Utopia/nadir normalization: added an optional normalized-axes toggle to the Pareto
scatter (lib/pareto.ts's normalizePoints, per the MOO survey's eq. 7), approximating
the utopia/nadir points from best/worst OBSERVED values across the run rather than
solving for them exactly. Raw units stay the default view.

Per-task ownership view: added a Tasks-tab panel showing which candidate currently
best-scores each task (GEPA's per-instance ownership idea), computed client-side from
each GraphNode's existing per_task data. Verified against PR #687's actual
compute_ownership() schema (owners/best_score/ownership_count) -- the client-side
computation matches it exactly, so no new backend field or dependency on that PR
landing first was needed.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
Co-authored-by: Osher Elhadad <Osher.Elhadad@ibm.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
OsherElhadad pushed a commit that referenced this pull request Oct 8, 2026
…t budget

PR #690 added ~70 tokens to two paragraphs documenting the native pareto/
epsilon_constraint gate modes, pushing SKILL.md's body to ~5007 tokens and
failing test_skill_authoring_lint.py::test_the_repo_has_no_authoring_violations
on main. Tightened the same two paragraphs without losing any information —
now ~4998 tokens, lint clean.

Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

algorithm Optimization algorithms: GEPA / SkillOpt / hill-climb bug Something isn't working documentation Improvements or additions to documentation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

agent-optimize v3: native Pareto gating, archive-based search, real merging, cost-reduction levers

2 participants