Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
59 commits
Select commit Hold shift + click to select a range
b4ad00a
calibrate(ddd-weather-discount): rubric v2 from 13-trial human review
Jul 7, 2026
3f220e5
feat(evaluator): per-task dimensions + deterministic precheck hook
Jul 7, 2026
92c17a5
feat(evaluator): hand every judge the agent's full diff
Jul 7, 2026
ace02cd
feat(evaluator): diff-first evaluation procedure in the base judge pr…
Jul 7, 2026
06b3597
feat(cli): judge-matrix run plan for calibration round 2
Jul 7, 2026
760f197
fix(evaluator): don't hard-fail auth preflight when env credentials a…
Jul 7, 2026
cb7e275
docs(calibration): pending v2.1 rubric proposal + live round-2 state
Jul 7, 2026
68d6758
chore(calibration): preserve the manual-gating eval helpers from round 2
Jul 7, 2026
a025b9d
docs(calibration): fable subset complete (10/10) - formal check 5 PAS…
Jul 7, 2026
67739b7
calibrate(ddd-weather-discount): rubric v2.1 - M1/M4/M5 recalibrated …
Jul 7, 2026
6da3bf7
chore(calibration): eval helpers compute the dimensions fingerprint d…
Jul 7, 2026
7373ec3
docs(calibration): v2.1 applied - record before/after commits and the…
Jul 7, 2026
201de89
feat(calibration): claude-supple variant + live v2.1 verification record
Jul 8, 2026
4fb65ad
calibrate(ddd-weather-discount): rubric v2.2 - live-verification find…
Jul 8, 2026
dc35afc
docs(calibration): handover snapshot - rubric lineage, live results, …
Jul 8, 2026
849414a
calibrate(ddd-weather-discount): rubric v2.3 - M5 direction-neutral
Jul 8, 2026
65a379e
calibrate(ddd-weather-discount): v2.3 amendment - configurable intera…
Jul 8, 2026
0d63ab2
calibrate(ddd-weather-discount): name M5's top tier MAX
Jul 8, 2026
ecd55d3
calibrate(ddd-weather-discount): sharpen MAX with a two-part evidence…
Jul 8, 2026
2f52df8
variant(claude-supple): instruction v2 - study-the-model phrasing + u…
Jul 8, 2026
65f3983
variant(claude-supple): instruction v3 - explain supple design, add a…
Jul 8, 2026
de6bdb1
docs(calibration): handover 2026-07-08 evening - v2.2/v2.3 verified l…
Jul 8, 2026
4279c64
docs(calibration): ayg7ckA v2.3 = 0.75; v2.4 candidate - M5 idiom-vs-…
Jul 9, 2026
22ccc1f
docs(calibration): #13 v2.3 control clean - 0.59 identical to v2.2, b…
Jul 9, 2026
6a6292a
pricing: add claude-fable-5 ($10/$50, cache read $1) per official pag…
Jul 9, 2026
c77f9a2
pricing: claude-opus-4-8 output 15->25 per official page (owner-appro…
Jul 9, 2026
b8e5ced
feat(runner): variant-level instruction override + claude-supple-hint…
Jul 9, 2026
51ad484
revert: variant-level instruction override (breaks the benchmark inva…
Jul 9, 2026
b83a601
variant(claude-deeper-hint): minimal think-deeper hint in CLAUDE.md, …
Jul 9, 2026
006df89
docs(calibration): deeper-hint probe 0.78-0.80 + Opus judge pilot res…
Jul 9, 2026
f381949
docs(calibration): vanilla#2 Ss2F6dR - 0.73 Fable / 0.67 Opus, M4 NON…
Jul 9, 2026
094a8c2
docs(calibration): deeper-hint#2 b5jkaWo - Fable 0.74, Opus 0.73; arm…
Jul 9, 2026
77db677
docs(calibration): deeper-hint arm complete - 0.79/0.74/0.805 (Fable)…
Jul 9, 2026
4626498
docs(calibration): matrix complete - vanilla#3 env-artifact reward an…
Jul 9, 2026
8427540
docs(calibration): matrix fully symmetric - 24 v2.3 evals, arm gap +0…
Jul 9, 2026
3fcbeb0
docs(calibration): night chain complete - arms n=4, record 0.87, Opus…
Jul 10, 2026
f25422f
fix(envs): disable MSBuild node reuse in all dotnet task environments
Jul 10, 2026
48b5657
docs(calibration): ntcoding arm opened (0.81, R1 NONE reformer profil…
Jul 10, 2026
edc904f
docs(calibration): ntcoding#2 qpCRH3A 0.78 - opposite shape to #1 (re…
Jul 10, 2026
8e3e280
docs(calibration): ntcoding arm complete n=3 (~0.81) - best-of-both-s…
Jul 10, 2026
16f3871
docs(calibration): ntcoding#4 0.84 - all Fable arms n=4; final landsc…
Jul 10, 2026
e64c646
docs(calibration): Opus batch complete - full model x configuration g…
Jul 10, 2026
92fcc1c
docs(calibration): Opus cross-cells n=2 - hint effect on Opus revised…
Jul 10, 2026
868bd69
docs(calibration): GRID COMPLETE - n=4 every cell; final theses recorded
Jul 11, 2026
7433800
viz(calibration): raw-measurements grid + judge test-retest plots (PL…
Jul 13, 2026
7b674a7
tools(calibration): persist the battle-tested trial-probe and eval-to…
Jul 13, 2026
18a5282
docs(calibration): register the persisted tooling in the handover map
Jul 13, 2026
41a0440
docs: register calibration tooling in root CLAUDE.md and the runner s…
Jul 13, 2026
26159a5
viz: cost x quality plane for the 24-trial grid (PL+EN); pricing re-v…
Jul 13, 2026
0e00f7f
viz: verdict heatmap 18 checks x 6 arms (PL+EN), parsed from judge re…
Jul 13, 2026
611dbc1
viz: heatmap polish — config labels on top, per-group hues (M violet …
Jul 14, 2026
22a4e3b
feat: cache-aware cost is THE cost (ADR-014)
Jul 15, 2026
9dd150e
docs: handover note - ADR-014 cost repricing, old dollar figures belo…
Jul 15, 2026
2800a8d
docs: purge the old full-rate cost story everywhere (ADR-014 sweep)
Jul 15, 2026
139be77
viz: drop run count from cost-quality chart title
Jul 16, 2026
78bfd81
calibration: move session ops artifacts to private nasde-calibration …
Jul 17, 2026
541087e
fix: pass precheck workspace path POSIX-style for JSON-safe output on…
Jul 20, 2026
62337d6
fix: locate Git Bash on Windows for precheck.sh execution
Jul 20, 2026
4258b44
test: skip precheck happy-path test on Windows (PATH bash may be the …
Jul 20, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .claude/skills/nasde-benchmark-calibration/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,19 @@ The rubric to edit lives at `evals/<source>/tasks/<task>/assessment_criteria.md`
dimension's scale/description, not one task's thresholds). The `<source>` and `<task>` come from the
trial's `result.json` (`source`, `task_name`).

Calibration can also restructure a single task's dimensions: an `assessment_dimensions.json`
placed next to the task's `assessment_criteria.md` overrides the benchmark-wide file for that task
only (different fingerprint — old and new evaluations are never mixed in one summary group). For
change-related checks, the evaluator already hands every judge the agent's full diff
(`<trial>/agent_changes.diff` + inline diffstat) — rubrics should direct the judge to answer
"what did the agent change" questions from that diff. When a task additionally needs hard mechanical
enforcement (e.g. a disqualification score cap), ship a `precheck.sh` — the evaluator runs it,
injects its JSON into the judge prompt as ground facts, and enforces its optional
`normalized_score_cap`. See `examples/ddd-architectural-challenges/tasks/ddd-weather-discount/` for
the reference calibrated task and `CALIBRATION_ROUND2_2026-07-07.md` for the loop's acceptance
criteria (repeatability, judge-model agreement, human-ranking correlation, dimension disjointness,
regression assertions).

Show the user a concrete **diff of the rubric** — the specific threshold/description change that would
have moved the judge toward the human's score — and **wait for approval before writing**. Never edit
the rubric silently.
Expand Down
32 changes: 31 additions & 1 deletion ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,36 @@ flowchart TB

---

### Agent diff — universal judge input

For every trial, the evaluator materializes the agent's full diff (start state → final
workspace: `git diff HEAD` + untracked files, via `workspace_diff.capture_patch` — the
same capture used for `changes.patch` in exports) into `<trial>/agent_changes.diff`,
and injects an "Agent diff" prompt section: the diffstat inline plus the file path for
Read/Grep (the `ClaudeSubprocessBackend` grants `--add-dir` on the trial dir when the
file is present). Rationale: the judge sees only the final state and — like a human
reviewer without a diff — cannot see removals or out-of-feature edits; the diff is the
universal, task-agnostic reference point for every "what did the agent change" check.
Skipped gracefully when the workspace has no git repo or nothing changed.

### Per-task rubric inputs

Three optional files next to a task's `assessment_criteria.md` refine its evaluation:

- `assessment_dimensions.json` — overrides the benchmark-wide dimensions file **for that task only**
(`resolve_dimensions_path`). A different dimensions file yields a different fingerprint, so
evaluations under old and new dimensions are never mixed in one summary group.
- `ground_truth_decisions.json` — reference decisions injected verbatim into the judge prompt.
- `precheck.sh` — an optional deterministic policy layer on top of the agent diff: the evaluator
runs it on the host before judging (`bash precheck.sh <workspace-path>`). Its stdout must be one
JSON object; it is injected into the judge prompt as "Deterministic pre-check signals" (facts the
judge must stay consistent with), recorded in `assessment_eval_*.json` under `precheck`, and its
optional `normalized_score_cap` (0..1) is enforced on the trial's normalized score (cap
application is recorded, so a capped score is always explainable). Use it when a task wants hard,
mechanical enforcement (e.g. a disqualification cap) rather than judge interpretation of the diff.
Any precheck failure degrades to "no precheck" with a warning. All three are also bundled into
`.calibration/` by `nasde calibrate publish`.

## Evaluator configuration

The evaluator agent is configurable via `[evaluation]` in `nasde.toml`. All options are optional — defaults provide a working evaluator out of the box.
Expand Down Expand Up @@ -213,7 +243,7 @@ When `mcp_config` is set, its path is passed through to the backend CLI (`--mcp-

### Token & cost economics ([ADR-011](docs/adr/011-token-cost-metrics.md))

Independently of the LLM judge, each trial's **token usage and cost** are read from the agent's `agent/trajectory.json` `final_metrics` (which Harbor writes for both Claude and Codex). A single extractor (`token_metrics.py`) computes `input = total_prompt_tokens` (full, cache included), `output = total_completion_tokens + reasoning_output_tokens`, and a USD cost at the **full catalog rate with no cache discount** ("as if every run were the first" — deterministic, order-independent). It derives `token_efficiency` (score per 1M tokens) and `cost_efficiency` (score per USD), using the dominant evaluator cluster's `normalized_score_mean`. The same extractor feeds both the run path (`evaluator.py` → `assessment_summary.json`) and the export path (`results_exporter.py` → `metrics.json`), so they cannot diverge. Prices come from a bundled, versioned `pricing.toml`; an unpriced model leaves cost null (token metrics still computed). `nasde run` prints a per-`(agent, model)` cost table after assessment completes.
Independently of the LLM judge, each trial's **token usage and cost** are read from the agent's `agent/trajectory.json` `final_metrics` (which Harbor writes for both Claude and Codex). A single extractor (`token_metrics.py`) computes `input = total_prompt_tokens` (full, cache included), `output = total_completion_tokens + reasoning_output_tokens`, the cache read/write volumes, and a **cache-aware USD cost** (ADR-014): fresh input at the full rate, cache writes at the cache-write rate, cache reads at the cached rate — what the API would bill for the run. The raw volumes stay in `token_usage`, so the cache-free ceiling is derivable offline; the scalar efficiency ratios were removed (ADR-011) — models are compared as a quality-vs-cost Pareto front. The same extractor feeds both the run path (`evaluator.py` → `assessment_summary.json`) and the export path (`results_exporter.py` → `metrics.json`), so they cannot diverge. Prices come from a bundled, versioned `pricing.toml`; an unpriced model leaves cost null (token metrics still computed). `nasde run` prints a per-`(agent, model)` cost table after assessment completes.

---

Expand Down
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,18 @@ See [docs/RELEASING.md](docs/RELEASING.md) for the release procedure.
## [Unreleased]

### Changed
- **Cost is now cache-aware ([ADR-014](docs/adr/014-cache-aware-cost.md)) — supersedes ADR-011's
"as if every run were the first" formula.** `cost_usd` bills fresh input at the
full rate, cache writes at the new per-model `cache_write_per_1m` (Anthropic
1-hour mode: 2× input), cache reads at `cached_input_per_1m` (0.1×), and output
at the output rate — matching what the API would bill (verified against Harbor's
per-step accounting to the cent on 20 of 24 grid trials). Rationale: measured
cache read ratios are a stable 93–98% of input across a full 24-trial grid, and
the old full-rate figure sat ~4.4× above a real bill. `token_usage` gains
`cache_write_tokens`; the cache-free ceiling is no longer stored (derivable as
`input × input_rate + output × output_rate`). A model entry missing cache rates
bills those volumes at the full input rate — conservative, never a silent
discount. Historical exports need a one-shot economics backfill to reprice.
- **Harbor bumped from 0.13 to 0.19** (`harbor[daytona,modal,e2b,runloop,gke]>=0.19,<0.20`).
The Python-API surface nasde drives (`JobConfig.model_validate`, `Job.create`,
`job.run()`, `AgentConfig` `import_path`/`kwargs`/`skills`/`mcp_servers`/`env`)
Expand Down
Loading