Skip to content

feat(evals): mock the dart MCP server and make green-gate a reliable regression check - #168

Merged
ryzizub merged 15 commits into
mainfrom
session/stoic-pelican-3u8q
Oct 5, 2026
Merged

ryzizub merged 15 commits into
mainfrom
session/stoic-pelican-3u8q

Conversation

@ryzizub

@ryzizub ryzizub commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Description

Mocks the dart MCP server for evals and makes the green-gate suite a reliable regression check.

What changed

  • evals/mocks/dart/ — analyze_files and dart_format, bodies captured from the real server.
  • New case green-gate-runs-the-four-gates-on-a-green-package grades real tool calls (tool_used, tool_order, mock_calls) instead of narration.
  • Shared fixture seeds a plain Counter class and one test so tool-driving cases have something on disk. Mock numbers match it.
  • warn-missing-mcp.sh no longer claims the MCP server "will not start" when it only failed to read a version. Under eval that claim was false and every case opened with it.
  • Green-gate's five least stable llm graders are now one-line regexes on the skill's vocabulary. Judges were at 4/12 to 11/12 on unchanged code.
  • The fifteen NOTES.md files are removed. Nothing validated them and eleven were already wrong. General guidance lives once in evals/README.md.

Green-gate, --threshold 0.8, 3 runs

                                  main    now
budgets-per-package               0.79   0.95
escalates-when-loop-stops         0.92   1.00
excludes-generated-files          1.00   1.00
plans-the-four-gates-in-order     0.87   1.00
refuses-to-carry-green-forward    0.86   0.95
refuses-to-weaken-coverage-gate   0.83   1.00
runs-the-four-gates (new)           -    0.98
stays-out-of-plain-function-work  1.00   1.00

Reviewer notes

  • Read the new case's with-arm score, not Δ; most of its graders need a mocked tool that is absent without the plugin.
  • create-project-does-not-over-ask moves from max_turns: 12 to 20. It drives a mocked tool and was the only such case under the tool-driving caps. Unrelated to the dart mock; revert if you prefer.
  • The hook change affects real sessions too: the notice now says to carry on if the MCP tools answer.
  • Merging triggers the full 15-skill post-merge run (~$13), since evals/_fixture/ and evals/mocks/ changed.

Type of Change

  • New feature (feat)
  • Bug fix (fix)
  • Code refactor (refactor)
  • Documentation (docs)
  • CI change (ci)
  • Chore (chore)

🤖 Generated with Claude Code

@ryzizub
ryzizub requested a review from a team as a code owner September 30, 2026 07:53
@unicoderbot

unicoderbot Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Heads up — I'm auto-updating this PR by merging the latest main into this branch. No action needed; I'll comment again only if this update hits a conflict or an unexpected error.

@ryzizub
ryzizub marked this pull request as draft September 30, 2026 07:59
@ryzizub
ryzizub marked this pull request as ready for review September 30, 2026 13:43
@ryzizub ryzizub changed the title feat(evals): add a dart MCP mock and a green-gate case that drives the gates feat(evals): mock the dart MCP server and make green-gate a reliable regression check Sep 30, 2026
@unicoderbot

unicoderbot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Heads up — I'm auto-updating this PR by merging the latest main into this branch. No action needed; I'll comment again only if this update hits a conflict or an unexpected error.

@ryzizub
ryzizub requested a review from marcossevilla October 1, 2026 12:40
@unicoderbot

unicoderbot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Heads up — I'm auto-updating this PR by merging the latest main into this branch. No action needed; I'll comment again only if this update hits a conflict or an unexpected error.

@unicoderbot

unicoderbot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Heads up — I'm auto-updating this PR by merging the latest main into this branch. No action needed; I'll comment again only if this update hits a conflict or an unexpected error.

ryzizub and others added 4 commits October 5, 2026 15:20
Mocks analyze_files and dart_format with the real tools/list schema captured
from `dart mcp-server --enable dart_format`, both returning a green result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The shared fixture now writes a plain Counter class and a single-expect test, so
a case that drives the MCP tools has something real on disk. `dart format` on it
reports "Formatted 2 files (0 changed)", matching the dart_format mock exactly.
The very-good-cli test mock moves from 94.2% to 100% so the mock set no longer
contradicts itself with a failing coverage gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers the skill's one-pass no-op path: an already-green package, all four gates
run, green confirmed from observed numbers, nothing edited. Grades tool choice,
arguments and order through tool_used, tool_order and mock_calls rather than
through narration. Edit and Write are granted so the no-edit graders can fail;
Bash is granted per the standing green-gate invariant. The seven existing cases
are untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds the dart block and the narrower-than-the-server _tools.json rule to
evals/README.md, documents the fixture's seeded source and the green mock set,
and rewrites the green-gate NOTES paragraph that still claimed no MCP server was
available. Corrects the case count to 101 in five places, the tool_order grader
row that read "Unused here", and the "one case drives a mocked tool" line.

CI gains a Plugin Validation step asserting every tool in a _tools.json has a
body file beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@unicoderbot

unicoderbot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

I tried to auto-update this PR with main but ran into conflicts. The branch is unchanged — please merge it manually.

ryzizub and others added 11 commits October 5, 2026 15:22
Seeding the fixture with source falsified "the fixture has no source in lib/"
in seven sibling NOTES files the change never touched; each now states what the
fixture actually holds, and the advice to paste code into prompts stands.

Also: the sixth stale case count in the README Running block, a mock-body count
that treated _tools.json as a body, an overclaim that every mock number matches
the fixture (only dart_format does), and green-gate NOTES counting its negative
control as a narration case.

Grader fixes: min_coverage now asserts the value, not just the key, so a model
passing 80 no longer scores; exclude_coverage gains the grader SKILL.md's fourth
required argument was missing; the blind rubric names 100% rather than "the
target", which a judge that never sees the prompt cannot resolve.

The CI step now catches a malformed _tools.json, which process substitution was
swallowing at status 0, and checks both directions so an orphan body file is an
error too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lable

license-compliance, create-project and ui-package each opened by telling case
authors the MCP tools are unreachable and never to assert a tool call. #159
mocked both servers and added 13 graders across these three skills that assert
exactly that, so the guidance contradicted the cases beside it.

Each now records which of its cases drive the mocked tool, which assert it did
not fire, and that a mocked tool is absent in the no-plugin arm so its Δ is
misleading. What genuinely stays unmeasurable is kept: license-compliance never
resolves a real dependency graph, and ui-package has no package on disk to edit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The case could pass with its own point missing. Eleven graders totalling weight
15 meant a miss on calls-analyze-with-fixes scored exactly 0.8000, equal to the
threshold. That grader and skill-fired now carry weight 5.

Grader corrections: min_coverage also matched "1000" and 100.0; exclude_coverage
passed on an empty value; format-after-analyze now covers the precedence rule
analyze-before-test left untested; calls-the-format-gate carries max: 1, the only
thing enforcing the one-pass no-op path.

Bash is no longer granted to the case. The skill reserves it for parsing
coverage/lcov.info, the mocked test tool writes no such file, and granting it let
`cat >` write files behind the edits-nothing and writes-nothing graders. The
"green-gate must declare Bash" invariant governs the skill's frontmatter, not a
case's allowed_tools.

The CI step rejected `_server.md`, a mock form evals/README.md documents, and
skipped any mock directory with no _tools.json — the exact case it exists to
catch. It now uses jq, matching evals.yaml and the hook scripts.

Docs: "Two cases drive a mocked tool" was six; "Every tool-driving case carries
max_turns: 20" was false for create-project-does-not-over-ask, whose caps are
raised rather than the claim narrowed; evals/testing/NOTES.md was an eleventh
file still saying the harness configures no MCP servers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The check was wrong twice in one review pass — it rejected `_server.md`, a mock
form evals/README.md documents, and skipped exactly the manifest-less directory
it existed to catch. A broken mock surfaces on the next eval run anyway, so the
guard was not paying for its own subtlety.

Reverts config/cspell.json too; shopt, nullglob and esac were added only for that
step and cspell does not read shell scripts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Grader counts, weight arithmetic and a turn measurement copied from the README
all go stale the moment a grader is added — one of them already did inside this
branch, "three test-call-* graders" against four files. Case names replace counts
where the point was which cases do what, and the weighting warning keeps its
lesson without the arithmetic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Running the case found two bugs that static review had dismissed.

The test mock claimed 38 tests and 378 lines against a package holding one test
and three lines. The model read the workspace, caught the mock lying, and refused
to call the package green — correctly. It now reports what the fixture actually
has, and the README records that mock numbers must track the fixture, with this
as the worked example.

The llm rubric then failed replies that reported all four gates correctly but
added that 100% of three lines is shallow coverage. Its FAIL clause invited the
judge to read that caveat as a number contradicting a gate. The rubric now says
outright that such commentary is not a gate failure.

The case scores 1.00 across three runs, up from 0.95.

tool_order semantics are confirmed rather than assumed: `before` runs first,
`after` second, evidenced by the trace indices the grader reports. The NOTES
to-do asking a future reader to verify this is replaced by the result, and the
README grader table records it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…down

check_vgv_cli returns `unverifiable` whenever dart is off the PATH the hook
inherits, which is always true under `claude plugin eval`. The warning then
claimed the Very Good CLI MCP server "will not start" — false there, since the
server is mocked and answers every call. Cases opened by reporting that blocker
instead of answering the question, and their rubrics failed on output that never
addressed the prompt.

Hooks and the MCP client do not share a PATH, so an unreadable version was never
evidence the server is down. The notice now states only what is known and says to
carry on if the tools answer.

Measured on the green-gate tag, 3 runs, against main:

  plans-the-four-gates-in-order    0.87 -> 1.00
  budgets-per-package              0.79 -> 0.88
  refuses-to-carry-green-forward   0.86 -> 0.90
  refuses-to-weaken-coverage-gate  0.83 -> 0.89

Failing graders across the tag drop from 10 to 4. evals/README.md claimed the
`unverifiable` status kept the hook quiet; it never did, and now says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these failed a response that satisfied it. per-failure-detail listed all
four failures as `code @ file:line` and then asked for the analyzer's full message
text, which the judge read as an omission. format-judged-by-changed-count was
phrased as if judging an executed run on a prompt that forbids running.
offers-the-missing-test demanded the test be the only alternative offered, so
noting that deleting dead code would also close the gap failed it, and it graded a
verbal claim of following testing standards rather than the test shown.
one-shared-coverage-target contradicted itself: it forbade a per-package override
while permitting the package to be handled separately, and a separate invocation
with its own target is both.

calls-the-format-gate moves from max 1 to max 2. SKILL.md says a round that
rewrites files is confirmed green on the next round, so a second format call is
correct, and a run that made one scored a false red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Twelve judged runs of unchanged code left no llm grader in green-gate stable:
the best passed 11/12, the worst 4/12. A case with several of them failed
something most runs whatever the skill did, and three rounds of rubric rewording
did not move the aggregate.

The five least stable graders are now a regex on the skill's own vocabulary,
each checking one thing. Against every recorded reply they would have scored
10 to 12 of 12 where the judges managed 4 or 5. One compound rubric is cut to
a single condition. discovers-roots-by-pubspec-walk is deleted: the prompt hands
over the package layout, so it asked for discovery of something already given.

Green-gate now clears --threshold 0.8 on all eight cases, five at 1.00, lowest
0.95. On main it was six of seven, lowest 0.79.

evals/README.md records three traps this surfaced: FAIL clauses that punish
thorough answers, PASS% meaning runs at exactly 1.00 rather than runs clearing
the threshold, and judge instability with the numbers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fifteen files, 1,550 lines, nothing validating any of them. Eleven carried claims
this branch falsified or that #159 had already falsified, and correcting them cost
four rounds across the branch. The parts that changed a decision today were two
lines of skill-level history; the per-case entries never did, and prompt.md's
description already says what each case asks for. What is general — noise floor,
judge instability, the trap list — lives once in evals/README.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#165 added a case and bumped 100 to 101; this branch did the same for its own.
Together that is 102. Both sides' edits merged cleanly to 101, which was wrong.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@ryzizub
ryzizub force-pushed the session/stoic-pelican-3u8q branch from c9a22d9 to 3597224 Compare October 5, 2026 13:24
@ryzizub
ryzizub merged commit 9651736 into main Oct 5, 2026
5 checks passed
@ryzizub
ryzizub deleted the session/stoic-pelican-3u8q branch October 5, 2026 13:50
@vgvbot vgvbot mentioned this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants