Repository navigation
feat(evals): mock the dart MCP server and make green-gate a reliable regression check - #168
Conversation
|
Heads up — I'm auto-updating this PR by merging the latest |
|
Heads up — I'm auto-updating this PR by merging the latest |
|
Heads up — I'm auto-updating this PR by merging the latest |
|
Heads up — I'm auto-updating this PR by merging the latest |
Mocks analyze_files and dart_format with the real tools/list schema captured from `dart mcp-server --enable dart_format`, both returning a green result. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The shared fixture now writes a plain Counter class and a single-expect test, so a case that drives the MCP tools has something real on disk. `dart format` on it reports "Formatted 2 files (0 changed)", matching the dart_format mock exactly. The very-good-cli test mock moves from 94.2% to 100% so the mock set no longer contradicts itself with a failing coverage gate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers the skill's one-pass no-op path: an already-green package, all four gates run, green confirmed from observed numbers, nothing edited. Grades tool choice, arguments and order through tool_used, tool_order and mock_calls rather than through narration. Edit and Write are granted so the no-edit graders can fail; Bash is granted per the standing green-gate invariant. The seven existing cases are untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds the dart block and the narrower-than-the-server _tools.json rule to evals/README.md, documents the fixture's seeded source and the green mock set, and rewrites the green-gate NOTES paragraph that still claimed no MCP server was available. Corrects the case count to 101 in five places, the tool_order grader row that read "Unused here", and the "one case drives a mocked tool" line. CI gains a Plugin Validation step asserting every tool in a _tools.json has a body file beside it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
I tried to auto-update this PR with |
Seeding the fixture with source falsified "the fixture has no source in lib/" in seven sibling NOTES files the change never touched; each now states what the fixture actually holds, and the advice to paste code into prompts stands. Also: the sixth stale case count in the README Running block, a mock-body count that treated _tools.json as a body, an overclaim that every mock number matches the fixture (only dart_format does), and green-gate NOTES counting its negative control as a narration case. Grader fixes: min_coverage now asserts the value, not just the key, so a model passing 80 no longer scores; exclude_coverage gains the grader SKILL.md's fourth required argument was missing; the blind rubric names 100% rather than "the target", which a judge that never sees the prompt cannot resolve. The CI step now catches a malformed _tools.json, which process substitution was swallowing at status 0, and checks both directions so an orphan body file is an error too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lable license-compliance, create-project and ui-package each opened by telling case authors the MCP tools are unreachable and never to assert a tool call. #159 mocked both servers and added 13 graders across these three skills that assert exactly that, so the guidance contradicted the cases beside it. Each now records which of its cases drive the mocked tool, which assert it did not fire, and that a mocked tool is absent in the no-plugin arm so its Δ is misleading. What genuinely stays unmeasurable is kept: license-compliance never resolves a real dependency graph, and ui-package has no package on disk to edit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The case could pass with its own point missing. Eleven graders totalling weight 15 meant a miss on calls-analyze-with-fixes scored exactly 0.8000, equal to the threshold. That grader and skill-fired now carry weight 5. Grader corrections: min_coverage also matched "1000" and 100.0; exclude_coverage passed on an empty value; format-after-analyze now covers the precedence rule analyze-before-test left untested; calls-the-format-gate carries max: 1, the only thing enforcing the one-pass no-op path. Bash is no longer granted to the case. The skill reserves it for parsing coverage/lcov.info, the mocked test tool writes no such file, and granting it let `cat >` write files behind the edits-nothing and writes-nothing graders. The "green-gate must declare Bash" invariant governs the skill's frontmatter, not a case's allowed_tools. The CI step rejected `_server.md`, a mock form evals/README.md documents, and skipped any mock directory with no _tools.json — the exact case it exists to catch. It now uses jq, matching evals.yaml and the hook scripts. Docs: "Two cases drive a mocked tool" was six; "Every tool-driving case carries max_turns: 20" was false for create-project-does-not-over-ask, whose caps are raised rather than the claim narrowed; evals/testing/NOTES.md was an eleventh file still saying the harness configures no MCP servers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The check was wrong twice in one review pass — it rejected `_server.md`, a mock form evals/README.md documents, and skipped exactly the manifest-less directory it existed to catch. A broken mock surfaces on the next eval run anyway, so the guard was not paying for its own subtlety. Reverts config/cspell.json too; shopt, nullglob and esac were added only for that step and cspell does not read shell scripts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Grader counts, weight arithmetic and a turn measurement copied from the README all go stale the moment a grader is added — one of them already did inside this branch, "three test-call-* graders" against four files. Case names replace counts where the point was which cases do what, and the weighting warning keeps its lesson without the arithmetic. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Running the case found two bugs that static review had dismissed. The test mock claimed 38 tests and 378 lines against a package holding one test and three lines. The model read the workspace, caught the mock lying, and refused to call the package green — correctly. It now reports what the fixture actually has, and the README records that mock numbers must track the fixture, with this as the worked example. The llm rubric then failed replies that reported all four gates correctly but added that 100% of three lines is shallow coverage. Its FAIL clause invited the judge to read that caveat as a number contradicting a gate. The rubric now says outright that such commentary is not a gate failure. The case scores 1.00 across three runs, up from 0.95. tool_order semantics are confirmed rather than assumed: `before` runs first, `after` second, evidenced by the trace indices the grader reports. The NOTES to-do asking a future reader to verify this is replaced by the result, and the README grader table records it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…down check_vgv_cli returns `unverifiable` whenever dart is off the PATH the hook inherits, which is always true under `claude plugin eval`. The warning then claimed the Very Good CLI MCP server "will not start" — false there, since the server is mocked and answers every call. Cases opened by reporting that blocker instead of answering the question, and their rubrics failed on output that never addressed the prompt. Hooks and the MCP client do not share a PATH, so an unreadable version was never evidence the server is down. The notice now states only what is known and says to carry on if the tools answer. Measured on the green-gate tag, 3 runs, against main: plans-the-four-gates-in-order 0.87 -> 1.00 budgets-per-package 0.79 -> 0.88 refuses-to-carry-green-forward 0.86 -> 0.90 refuses-to-weaken-coverage-gate 0.83 -> 0.89 Failing graders across the tag drop from 10 to 4. evals/README.md claimed the `unverifiable` status kept the hook quiet; it never did, and now says so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these failed a response that satisfied it. per-failure-detail listed all four failures as `code @ file:line` and then asked for the analyzer's full message text, which the judge read as an omission. format-judged-by-changed-count was phrased as if judging an executed run on a prompt that forbids running. offers-the-missing-test demanded the test be the only alternative offered, so noting that deleting dead code would also close the gap failed it, and it graded a verbal claim of following testing standards rather than the test shown. one-shared-coverage-target contradicted itself: it forbade a per-package override while permitting the package to be handled separately, and a separate invocation with its own target is both. calls-the-format-gate moves from max 1 to max 2. SKILL.md says a round that rewrites files is confirmed green on the next round, so a second format call is correct, and a run that made one scored a false red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Twelve judged runs of unchanged code left no llm grader in green-gate stable: the best passed 11/12, the worst 4/12. A case with several of them failed something most runs whatever the skill did, and three rounds of rubric rewording did not move the aggregate. The five least stable graders are now a regex on the skill's own vocabulary, each checking one thing. Against every recorded reply they would have scored 10 to 12 of 12 where the judges managed 4 or 5. One compound rubric is cut to a single condition. discovers-roots-by-pubspec-walk is deleted: the prompt hands over the package layout, so it asked for discovery of something already given. Green-gate now clears --threshold 0.8 on all eight cases, five at 1.00, lowest 0.95. On main it was six of seven, lowest 0.79. evals/README.md records three traps this surfaced: FAIL clauses that punish thorough answers, PASS% meaning runs at exactly 1.00 rather than runs clearing the threshold, and judge instability with the numbers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fifteen files, 1,550 lines, nothing validating any of them. Eleven carried claims this branch falsified or that #159 had already falsified, and correcting them cost four rounds across the branch. The parts that changed a decision today were two lines of skill-level history; the per-case entries never did, and prompt.md's description already says what each case asks for. What is general — noise floor, judge instability, the trap list — lives once in evals/README.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#165 added a case and bumped 100 to 101; this branch did the same for its own. Together that is 102. Both sides' edits merged cleanly to 101, which was wrong. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
c9a22d9 to
3597224
Compare
Description
Mocks the
dartMCP server for evals and makes the green-gate suite a reliable regression check.What changed
evals/mocks/dart/—analyze_filesanddart_format, bodies captured from the real server.green-gate-runs-the-four-gates-on-a-green-packagegrades real tool calls (tool_used,tool_order,mock_calls) instead of narration.Counterclass and one test so tool-driving cases have something on disk. Mock numbers match it.warn-missing-mcp.shno longer claims the MCP server "will not start" when it only failed to read a version. Under eval that claim was false and every case opened with it.llmgraders are now one-line regexes on the skill's vocabulary. Judges were at 4/12 to 11/12 on unchanged code.NOTES.mdfiles are removed. Nothing validated them and eleven were already wrong. General guidance lives once inevals/README.md.Green-gate,
--threshold 0.8, 3 runsReviewer notes
create-project-does-not-over-askmoves frommax_turns: 12to20. It drives a mocked tool and was the only such case under the tool-driving caps. Unrelated to the dart mock; revert if you prefer.evals/_fixture/andevals/mocks/changed.Type of Change
feat)fix)refactor)docs)ci)chore)🤖 Generated with Claude Code