Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,12 +15,15 @@ agents/
flutter-reviewer.md # Read-only Flutter code reviewer subagent
docs/ # Gitignored, local only
plan/ # Planning and design documents
evals/ # `claude plugin eval` suite — all 15 skills, 101 cases
evals/ # `claude plugin eval` suite — all 15 skills, 102 cases
README.md # Case format, grader reference, how to add a case
_fixture/
fixture.sh # The only copy of the neutral Flutter skeleton; every case symlinks here
mocks/ # MCP stand-ins, one directory per server name in .mcp.json
<server>/
_tools.json # The tools/list result the model sees; may be narrower than the server
<tool>.md # Frontmatter optional; the body is the tool result
<skill>/ # One directory per skill, named for the skill it covers
NOTES.md # What each case discriminates and why it exists
<case-name>/ # One directory per case, named for the case
prompt.md # Frontmatter: run limits and tools. Body: the prompt
case.yaml # schema_version, name, and the scaffold hook
Expand Down
93 changes: 72 additions & 21 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Does Claude route to the skill, and does the output follow it? Cases run through
[`claude plugin eval`](https://code.claude.com/docs/en/plugin-evals) against real models.

```bash
claude plugin eval . --scaffold # all 101 cases, both arms
claude plugin eval . --scaffold # all 102 cases, both arms
claude plugin eval . --scaffold --tag bloc # one skill
claude plugin eval . --scaffold --ablation none # with-plugin arm only, half the cost
```
Expand Down Expand Up @@ -92,7 +92,7 @@ threshold is `1.0`, so **pass `--threshold 0.8` or every imperfect case exits 1*
| ------------- | ------------------------------------- | ------------------------------------------------------------------------ |
| `tool_used` | `tool`, `input_match`, `min`, `max` | Routing. The defaults are wrong for this suite, see below |
| `regex` | `pattern`, `flags`, `match`, `target` | `match: not_contains` for absence. Case-insensitivity goes in `flags: i` |
| `tool_order` | `before`, `after` | Unused here |
| `tool_order` | `before`, `after` | `before` runs first, `after` second. Gate ordering in `green-gate` only |
| `file_exists` | `path`, `exists` | Unused here: cases are graded on the reply, not on files |
| `llm` | `criteria`, `focus` | Frontmatter is just `type: llm`; the file body is the rubric |
| `baseline` | `baseline_file`, `criteria` | Unused here |
Expand All @@ -105,7 +105,7 @@ made to a [mocked MCP tool](#mocking-the-mcp-servers).

### Routing graders

A hundred of the 101 cases carry one:
All but one of the 102 cases carry one:

```markdown
---
Expand Down Expand Up @@ -142,12 +142,30 @@ Negative controls use the same grader with `min: 0` and `max: 0`.
- **Check both arms answered.** A no-plugin arm that asks a clarifying question instead of
doing the work makes every grader look discriminating. That is a prompt that is not
self-contained, not a result.
- **The plugin's SessionStart hook can fire inside a run.** `warn-missing-mcp.sh` injects
"Very Good CLI is not installed" whenever `check_vgv_cli` returns `not_installed`, and a
tool-driven case then answers with that blocker instead of the question. The
`unverifiable` status keeps it quiet. Two prompts still say the session cannot reach a
- **The plugin's SessionStart hook fires inside every run.** `check_vgv_cli` returns
`unverifiable` there — `very_good` resolves on PATH but `very_good --version` answers
nothing, because it is a shim that execs `dart` under a throwaway `$HOME`. That does not
keep the hook quiet; it emits a notice, and an earlier wording of it asserted the MCP
server "will not start", which is false under mocks. Cases then opened with that blocker
instead of the answer and their rubrics failed. The notice now states only what is known
and says to carry on if the tools answer. Re-read it before blaming a case. Two prompts still say the session cannot reach a
real toolchain, which is a different problem and true regardless.

- **A FAIL clause can punish a better answer than the PASS clause asked for.** Three
green-gate rubrics failed replies that did more than required: one listed every failure
and then asked for fuller analyzer text, one offered the required test and also noted the
method might be dead code, one stated the format rule for a run the prompt forbade. Each
FAIL clause was narrow enough that the extra thoroughness tripped it. Write the FAIL
clause for the wrong answer, not for any departure from the shortest right one.
- **`PASS%` in the results table is runs scoring exactly 1.00**, not runs clearing
`--threshold`. A case at 0.90 shows `33%` while passing the threshold on every run. Read
`SCORE` against the threshold; read `PASS%` only when hunting flaky graders.
- **Judges are not stable.** Twelve judged runs of unchanged green-gate code left no `llm`
grader at 12/12; the range was 4/12 to 11/12. A case with several `llm` graders will fail
something most runs whatever the skill did. When a judge fails a reply that plainly
satisfies its rubric, convert the check to a `regex` on the skill's vocabulary rather than
rewording the rubric.

Beyond that: write prompts as a user would send them, name no skill in a prompt, grade
mechanically where you can, include the cases where the skill must say no, and keep a
negative control's rubric to the absence of the skill's vocabulary.
Expand Down Expand Up @@ -201,6 +219,13 @@ Match the skill's own vocabulary. Keep `llm` graders for short output; for anyth
Every run starts in an empty workspace. `_fixture/fixture.sh` recreates the neutral Flutter
skeleton, hooked up through `context.scaffold_script`, and runs **only with `--scaffold`**.

It writes `pubspec.yaml`, `lib/counter.dart` and `test/counter_test.dart`. The two source
files exist so a case that drives the MCP tools has something real on disk to act on, and
they are deliberately the most boring code that satisfies that: a plain class with no
Flutter import, and one `test()` with one `expect()`. Nothing there is a VGV convention,
because everything there is visible to the no-plugin arm of every other skill's cases. Do
not grow them into a widget, a bloc, a `pumpApp` or a mocked dependency.

`context.scaffold_script` will not take a path that leaves the case directory, but it does
resolve a symlink inside it, so each case's `fixture.sh` is a symlink to the one script. A
new case needs its own:
Expand All @@ -221,8 +246,8 @@ and runs fail at scaffold time. CI runs on Linux and is unaffected.
every run reports on a `mocked:` line. Real servers start only with `--allow-real-servers`
or `--mocks off`, neither of which is used here.

This applies to eval runs only. A real session still reaches the real `very_good mcp`
server.
This applies to eval runs only. A real session still reaches the real `very_good mcp` and
`dart mcp-server` servers.

```text
evals/mocks/very-good-cli/
Expand All @@ -231,10 +256,14 @@ evals/mocks/very-good-cli/
├── packages_check_licenses.md
├── packages_get.md
└── test.md
evals/mocks/dart/
├── _tools.json # the analyze_files and dart_format entries only
├── analyze_files.md
└── dart_format.md
```

Frontmatter is optional and the body is the tool result. Three of the four here have no
frontmatter, which is the same as `type: fixed`.
Frontmatter is optional and the body is the tool result. Five of the six bodies here have
no frontmatter, which is the same as `type: fixed`.

| Key | Default | Purpose |
| ------------ | ------- | --------------------------------------------------------------- |
Expand All @@ -261,8 +290,14 @@ Four things that are not obvious:
that is `create`, so only `create.md` carries an `expect`. Grade argument *choices* with
`tool_used` or a `regex` against `mock_calls`.
- **`_tools.json` drives the schema the model sees.** It is the `tools/list` result, the
object with the `tools` array. Regenerate it after a Very Good CLI release by speaking
MCP to `very_good mcp` over stdio; a stale one teaches a schema the CLI no longer has.
object with the `tools` array. Regenerate it by speaking MCP over stdio to the server
itself — `very_good mcp` for one, `dart mcp-server --enable dart_format` for the other —
and capturing the `tools/list` result; a stale one teaches a schema the CLI no longer
has. Answer the server's `roots/list` request back to the client or `dart` blocks.
- **A `_tools.json` may be narrower than the server.** `dart` exposes 15 tools; the mock
ships the 2 `green-gate` names. That removes the listed-but-unmocked case entirely, and
keeps `roots` out of reach — a separate tool whose real handshake a mock cannot perform.
Adding one later is additive: capture it from the same probe, drop in a `<tool>.md`.

The mocks are reachable only when `check_vgv_cli` returns `unverifiable`. In a run
`very_good` resolves on PATH but `very_good --version` answers nothing, because it is a
Expand All @@ -278,6 +313,21 @@ answer to the no-plugin arm, exactly as a non-neutral fixture does.
`packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive
licenses, so a case has something real to flag.

**The mock set is green everywhere else**, so a case can drive every gate and reach an
exit. `analyze_files` returns no errors, `dart_format` reports `0 changed`, and `test`
passes at 100%. **Keep those numbers consistent with the fixture on disk.** `dart format`
on the seeded package really does print `Formatted 2 files (0 changed)`, and the test body
reports the one test and three executable lines the fixture actually holds. An earlier
version claimed 38 tests and 378 lines; the model read the workspace, caught the mock
lying, and refused to call the package green — so a mock that contradicts the fixture does
not merely go unread, it fails the case. A failure-path case supplies its own `mocks/`
override rather than reddening the shared set.

One body is deliberately lossy. Real `dart_format` output opens with
`dart format in <absolute root>:`, which a fixed body cannot know, so the mock drops that
line and keeps the `Formatted N files (M changed)` summary the gate is actually read
from.

**Converting a case to drive a tool invalidates every rubric that read the narration.** The
model stops describing the call and just makes it, so a blind judge sees no evidence and
fails a rubric that was passing. Four rubrics went stale this way in one pass, two of them
Expand All @@ -289,8 +339,8 @@ only those judging something still in the reply.
prompt it replaced: it routes, calls, reads the answer, then writes the reply.
`ui-package-scaffolds-with-app-ui-package-template` measured 11 to 15 turns where the
suite's usual `max_turns: 12` and `timeout_seconds: 600` had been ample, and hit both
limits. The four tool-driving cases carry `max_turns: 20` and `timeout_seconds: 900`. A cap
breach scores the case 0 with no failing grader, so it reads as a content failure.
limits. A tool-driving case therefore carries `max_turns: 20` and `timeout_seconds: 900`.
A cap breach scores the case 0 with no failing grader, so it reads as a content failure.

**A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails
for free and takes any grader that needs the tool's output with it. Unlike
Expand All @@ -305,7 +355,7 @@ plugin and 0.00 without, on three runs each.

```bash
E="claude plugin eval . --scaffold"
$E # all 101, both arms
$E # all 102, both arms
$E --tag bloc --tag testing # two skills
$E --case bloc-writes-sealed-events-and-states # one case
$E --ablation none # with-plugin arm only, half the cost
Expand All @@ -331,7 +381,7 @@ Read the two arm scores, not the total. The without-arm is supposed to score bad
baseline, so its Δ reads low. Routing is decided before the switch, so the pin never
explains a routing miss.

Costs: **$13** for 100 cases in one arm, **$25** for both, roughly **$0.12 per run**. At
Costs: **$13** for 102 cases in one arm, **$25** for both, roughly **$0.12 per run**. At
`--runs 3` a two-arm sweep is six runs per case, so budget around **$75**. `-j` up to 8
shortens wall clock.

Expand All @@ -353,7 +403,7 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on-
- CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`.
- `--ablation none` and `--ablation with-without` weight the baseline differently. Compare
runs from one mode at a time.
- The job has a one-hour ceiling. 100 cases in one arm measured roughly 35 minutes at
- The job has a one-hour ceiling. 102 cases in one arm measured roughly 35 minutes at
`-j 4`. A two-arm run at `--runs 3` is 600 runs and does not fit.

---
Expand All @@ -365,9 +415,10 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on-
`aggregate-result.json`, only a `tracePath` into a sandbox deleted unless `--keep-temp`
is passed. `report.html` does show the judged text.
- **Judge calibration.** Most graders are `llm` with no human-labelled gold set.
- **Tool execution, mostly.** One case drives a mocked tool,
`license-compliance-runs-check-with-full-license-info`. The other tool-driven skills are
still graded on the calls they narrate.
- **Tool execution, partly.** Six cases assert a call to a mocked tool, across
`create-project`, `green-gate`, `license-compliance` and `ui-package` — find them by
grepping the graders for `mcp__plugin`. Every other tool-driven case is still graded on
the calls it narrates.
- **Stable routing.** Whether a skill activates is nondeterministic, which is why routing
is a `tool_used` grader rather than inferred from content.
- **Prose in a `SKILL.md`.** Deliberate: an earlier version asserted a hundred `contains`
Expand Down
27 changes: 25 additions & 2 deletions evals/_fixture/fixture.sh
Original file line number Diff line number Diff line change
Expand Up @@ -3,10 +3,33 @@
#
# THE ONLY COPY. Every case's fixture.sh is a symlink to this file. Edit it here.
#
# KEEP THIS NEUTRAL. The pubspec must not name anything a skill teaches.
# KEEP THIS NEUTRAL. Neither the pubspec nor the seeded source may name anything a
# skill teaches: no widget, no bloc, no mocktail, no golden, no pumpApp. The source
# exists so a case that drives the MCP tools has something real on disk, nothing more.
set -euo pipefail
mkdir -p lib test
touch lib/.gitkeep test/.gitkeep
cat > lib/counter.dart <<'DART'
/// Keeps a running total.
class Counter {
int _value = 0;

int get value => _value;

void add(int amount) => _value += amount;
}
DART
cat > test/counter_test.dart <<'DART'
import 'package:eval_fixture_app/counter.dart';
import 'package:flutter_test/flutter_test.dart';

void main() {
test('add increases the value', () {
final counter = Counter();
counter.add(3);
expect(counter.value, 3);
});
}
DART
cat > pubspec.yaml <<'YAML'
name: eval_fixture_app
description: Scratch Flutter app used as working-directory context for behavior evals.
Expand Down
Loading
Loading