Repository navigation
optimizer: native Pareto archive gate + epsilon-constraint mode in round.py - #690
Conversation
…und.py round.py excluded "pareto" from its own --mode choices because _gate() never forwarded --objectives/--metrics-* to gate_check.py, so a spec declaring gate_mode: pareto was silently gated single-objective. Fix it end-to-end instead of excluding it (#684 items 1-2): - cap_evolve/pareto_archive.py: a persistent bounded ParetoArchive (Cheng & Li 1997) — non-domination ranking (reuses selection.dominates), the same per-objective significance floor gate.py's #667 fix established (reuses gate._objective_state), and NSGA-II crowding-distance eviction once over capacity (default 15). Persisted as versioned JSON in the run dir. - cap_evolve/gate.py: new gate_mode: epsilon_constraint (Haimes et al. 1971 bounded-objective-function method) — maximize reward subject to constraints: [{name, max}], with the same noise-floor discipline on the ceiling check. - cap_evolve/harness.py: candidate_cost_objective(), deriving a candidate's cost mean/SE from its own already-persisted per-task cost_usd data. - round.py: removed the ROUND_MODES exclusion; _gate() forwards --objectives/--metrics-*/--constraints; a pareto-gated candidate's accept/reject is now decided by the archive's own insertion result, not a one-shot pairwise read. - gate_check.py: epsilon_constraint added to GATE_MODES, --constraints flag. Existing single-metric gate modes (paired/significant/strict/threshold) are untouched. Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
|
🏷️ Automatic Labeling I've analyzed this pull request and added the following labels:
These labels were selected based on the PR title, description, and changed files. If you believe any labels are incorrect or missing, feel free to adjust them manually. |
| p1 = _round(run_dir, project, "cand_1", "--mode", "pareto") | ||
| assert p1.returncode == 0, p1.stdout + p1.stderr | ||
| out1 = json.loads(p1.stdout) | ||
| size_after_round_1 = out1["pareto_archive"]["size_total"] if "size_total" in out1["pareto_archive"] \ |
…ence try_insert only validated the NEW candidate's values against the declared objectives; an existing archive point (e.g. round.py's baseline seed when harness.candidate_cost_objective reports no cost) could already be missing one, so the next candidate with full values crashed _states_vs with a raw KeyError instead of the designed ParetoObjectiveError refusal. Also reuse rundir.py's _file_lock + _atomic_write (the same pattern used for state.json's identical load->mutate->save cycle) in ParetoArchive.save/ load_or_create instead of a plain write_text/read_text, so two processes touching the same run dir's archive can't tear or lose a write. Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
|
Pushed a fix for both confirmed review findings in `core/cap_evolve/pareto_archive.py` (commit 2e91f49): 1. Crash bug. `_states_vs` now validates that an existing archive point also has every declared objective before diffing against it, raising the designed `ParetoObjectiveError` (with a clear "existing archive point is missing declared objective ... (have: ...)" message, same shape as gate.py's own missing-metric refusal) instead of letting a bare dict lookup raise `KeyError`. This covers the case directly, independent of round.py's seeding path, since the archive point itself can be incomplete for other reasons too (e.g. a version-mismatch reload, or any future caller that appends a point directly). Verified the exact reviewer repro no longer crashes: -> now raises ParetoObjectiveError, not KeyError``` 2. Persistence locking. `ParetoArchive.save()`/`load_or_create()` now reuse `rundir.py`'s existing `_file_lock` + `_atomic_write` helpers (the same pattern `state.json`'s read-modify-write methods already use) instead of plain `write_text`/`read_text`, so a concurrent writer can't tear `pareto_archive.json` or race another process's save. Tests: `python -m pytest core/tests -q` → 1938 passed, 1 skipped, 2 failed — the 2 failures are the pre-existing unrelated `test_tau2_airline_eval_cost.py::test_priced_run_is_reported_as_before` / `test_rollout_says_whether_its_cost_was_measured` (cost_source `"tau2"` vs `"partial_models"`), untouched by this change. |
… add utopia-nadir normalization, fix optimizer-cost telemetry gap (#691) Issue #684 items 9 and 10 (dashboard verification + optimizer-cost telemetry fix). Telemetry (item 10): optimizer_seconds/optimizer_usd were always 0 in agent-mode runs because meter.py's automatic metering only works under host.py's headless driver (it reads a claude-code session log that only exists there) -- a SKILL.md compliance gap, not a framework instrumentation gap. commit.py now logs an optimizer_cost_warning when real wall-clock time clearly passed since the run's previous decision but the new commit still carries zero optimizer cost, surfaced in commit.py's own warnings list and on the dashboard's node detail. SKILL.md's guidance was strengthened and trimmed to stay within the repo's 5000-token skill-body budget; the long rationale moved to algorithm.md's new "Optimizer cost telemetry" section. Dashboard verification (item 9): built a synthetic multi-candidate reward+cost fixture with genuine dominance relationships (two non-dominated frontier candidates, one strictly dominated candidate) and confirmed dashboard.reduce_run's existing node shape (val/cost_usd/per_task/per_task_metrics) correctly feeds the frontend's already-tested Pareto frontier logic and per-objective Tasks-tab deltas. Also wired through round.py's real pareto-gate output schema (confirmed against the now-merged-adjacent PR #690): objective_values/objective_stderrs/pareto_archive per candidate, read off the same gate-table lookup dashboard.py already used for verdict_stable -- no round.py or commit.py changes needed. Utopia/nadir normalization: added an optional normalized-axes toggle to the Pareto scatter (lib/pareto.ts's normalizePoints, per the MOO survey's eq. 7), approximating the utopia/nadir points from best/worst OBSERVED values across the run rather than solving for them exactly. Raw units stay the default view. Per-task ownership view: added a Tasks-tab panel showing which candidate currently best-scores each task (GEPA's per-instance ownership idea), computed client-side from each GraphNode's existing per_task data. Verified against PR #687's actual compute_ownership() schema (owners/best_score/ownership_count) -- the client-side computation matches it exactly, so no new backend field or dependency on that PR landing first was needed. Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com> Co-authored-by: Osher Elhadad <Osher.Elhadad@ibm.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
…t budget PR #690 added ~70 tokens to two paragraphs documenting the native pareto/ epsilon_constraint gate modes, pushing SKILL.md's body to ~5007 tokens and failing test_skill_authoring_lint.py::test_the_repo_has_no_authoring_violations on main. Tightened the same two paragraphs without losing any information — now ~4998 tokens, lint clean. Signed-off-by: Osher Elhadad <Osher.Elhadad@ibm.com>
Summary
Closes the #684 item 1 bug:
round.pyexcluded"pareto"from its own--modechoices because_gate()never forwarded--objectives/--metrics-*togate_check.py, so a spec declaringgate_mode: paretowas silently gated single-objective (reward-only). This fixes it properly end-to-end rather than excluding it, and addsepsilon_constraintas a second, simpler multi-objective mode (item 2).core/cap_evolve/pareto_archive.py(new):ParetoArchive— a persistent bounded non-dominated frontier (Cheng & Li 1997, cited in the MOO survey grounding issue agent-optimize v3: native Pareto gating, archive-based search, real merging, cost-reduction levers #684). Insertion requires non-domination against the archive AND a significant win on >= 1 objective, both using the SAME per-objective significance floorgate.py's gate: multi-objective Pareto acceptance mode #667 fix established (reusesgate._objective_state) and the SAME dominance arithmetic (selection.dominates) rather than reimplementing either. Over capacity (default 15), evicts the point with the smallest NSGA-II crowding distance. Persisted as versioned JSON at$RUN_DIR/pareto_archive.json, so it survives across rounds and process restarts.core/cap_evolve/gate.py: newgate_mode: epsilon_constraint(Haimes, Lasdon & Wismer 1971's bounded-objective-function method) — maximize reward subject toconstraints: [{name, max}], with the same noise-floor discipline: a constraint satisfied only nominally (its own measurement noise could put it over the ceiling) is rejected, not accepted.core/cap_evolve/harness.py:candidate_cost_objective()— derives a candidate's cost mean/SE from its own already-persisted per-taskcost_usddata (no new field).skills/algorithms/agent-optimize/scripts/round.py: removed theROUND_MODESexclusion of"pareto";_gate()now forwards--objectives/--metrics-*/--constraints; for--mode pareto, round.py builds/updates the persistent archive and a candidate's accept/reject is now literally "did the archive give it a slot", overriding the one-shot pairwise verdictgate_check.pyalone would have produced.skills/algorithms/agent-optimize/scripts/gate_check.py:epsilon_constraintadded toGATE_MODES, new--constraintsflag.paired/significant/strict/threshold) are completely untouched — no code path they use was touched.Design decisions
DEFAULT_CAPACITYinpareto_archive.py) — a flat, unmeasured default (ponytail:comment), configurable later via capevolve.yaml if a real run shows it binds.harness.py's existing per-taskcost_usd(from issue agent-optimize v2: dashboard multi-objective visibility, optimizer-diagnosis flow view, bigger/more-parallel candidates #676's instrumentation) is the only secondary metric persisted per task alongside reward; neither "latency" nor "tokens" is persisted per-task, so declaring either without an explicit--metrics-*value raises the sameParetoObjectiveErrorrefusalgate.py's own cost→latency→tokens fallback chain already uses (refuse, don't silently degrade).control_relative,verdict_by_reference) are skipped for the two new modes since they don't yet thread objective/constraint metrics through (documented with aponytail:-style comment — add if a multi-objective run's drift turns out to matter).Test plan
core/tests/test_pareto_archive.py: archive insertion/eviction correctness on synthetic objective vectors (dominance, non-dominated tradeoffs, maintenance pruning, crowding-distance eviction keeps the frontier spread), the significance floor refusing a noise-level cost "win" (the same adversarial case as the earlier gate.py bug — tie on reward + float-noise cost delta must NOT get archived), and persistence surviving a simulated process restart (write, reload, confirm state matches).core/tests/test_gate_epsilon_constraint.py: accept/reject correctness, the noise-level-constraint-satisfaction adversarial case, missing value/stderr refusals, paired-vs-significant fallback.core/tests/test_round_pareto_archive.py: end-to-end —round.py --mode paretowith a declared objectives config produces realpareto_archive/dominance fields in its output (not paired-mode scalar fields), asserted directly since that's the exact symptom that proved the real-run bug; the archive persists across SEPARATEround.pysubprocess invocations (genuine process restarts);epsilon_constraintrequires a non-emptyconstraints:config and runs end-to-end.test_a_failed_gate_is_not_a_verdict.py's mode-list invariant (round.py's--modechoices now equalgate_check.py's, no gap) and removedtest_round_rejects_pareto_mode.py(asserted the now-removed exclusion).python -m pytest core/tests -q: 1933 passed, 2 failed, 5 skipped — the 2 failures are the pre-existing, unrelatedtest_tau2_airline_eval_cost.pyfailures named in the issue (an environment/model-name mismatch in that test's fixture, nothing to do with this change).Confirmed directly (not just inferred):
round.py --mode paretoend-to-end now produces realpareto_archive/dominance output — seetest_round_pareto_archive.py::test_round_pareto_produces_archive_fields_not_paired_scalar_fallback.Closes #684 items 1-2.