diff --git a/AGENTS.md b/AGENTS.md index 04d9574..f9dd906 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -15,12 +15,15 @@ agents/ flutter-reviewer.md # Read-only Flutter code reviewer subagent docs/ # Gitignored, local only plan/ # Planning and design documents -evals/ # `claude plugin eval` suite — all 15 skills, 101 cases +evals/ # `claude plugin eval` suite — all 15 skills, 102 cases README.md # Case format, grader reference, how to add a case _fixture/ fixture.sh # The only copy of the neutral Flutter skeleton; every case symlinks here + mocks/ # MCP stand-ins, one directory per server name in .mcp.json + / + _tools.json # The tools/list result the model sees; may be narrower than the server + .md # Frontmatter optional; the body is the tool result / # One directory per skill, named for the skill it covers - NOTES.md # What each case discriminates and why it exists / # One directory per case, named for the case prompt.md # Frontmatter: run limits and tools. Body: the prompt case.yaml # schema_version, name, and the scaffold hook diff --git a/evals/README.md b/evals/README.md index 14b83dc..626d811 100644 --- a/evals/README.md +++ b/evals/README.md @@ -4,7 +4,7 @@ Does Claude route to the skill, and does the output follow it? Cases run through [`claude plugin eval`](https://code.claude.com/docs/en/plugin-evals) against real models. ```bash -claude plugin eval . --scaffold # all 101 cases, both arms +claude plugin eval . --scaffold # all 102 cases, both arms claude plugin eval . --scaffold --tag bloc # one skill claude plugin eval . --scaffold --ablation none # with-plugin arm only, half the cost ``` @@ -92,7 +92,7 @@ threshold is `1.0`, so **pass `--threshold 0.8` or every imperfect case exits 1* | ------------- | ------------------------------------- | ------------------------------------------------------------------------ | | `tool_used` | `tool`, `input_match`, `min`, `max` | Routing. The defaults are wrong for this suite, see below | | `regex` | `pattern`, `flags`, `match`, `target` | `match: not_contains` for absence. Case-insensitivity goes in `flags: i` | -| `tool_order` | `before`, `after` | Unused here | +| `tool_order` | `before`, `after` | `before` runs first, `after` second. Gate ordering in `green-gate` only | | `file_exists` | `path`, `exists` | Unused here: cases are graded on the reply, not on files | | `llm` | `criteria`, `focus` | Frontmatter is just `type: llm`; the file body is the rubric | | `baseline` | `baseline_file`, `criteria` | Unused here | @@ -105,7 +105,7 @@ made to a [mocked MCP tool](#mocking-the-mcp-servers). ### Routing graders -A hundred of the 101 cases carry one: +All but one of the 102 cases carry one: ```markdown --- @@ -142,12 +142,30 @@ Negative controls use the same grader with `min: 0` and `max: 0`. - **Check both arms answered.** A no-plugin arm that asks a clarifying question instead of doing the work makes every grader look discriminating. That is a prompt that is not self-contained, not a result. -- **The plugin's SessionStart hook can fire inside a run.** `warn-missing-mcp.sh` injects - "Very Good CLI is not installed" whenever `check_vgv_cli` returns `not_installed`, and a - tool-driven case then answers with that blocker instead of the question. The - `unverifiable` status keeps it quiet. Two prompts still say the session cannot reach a +- **The plugin's SessionStart hook fires inside every run.** `check_vgv_cli` returns + `unverifiable` there — `very_good` resolves on PATH but `very_good --version` answers + nothing, because it is a shim that execs `dart` under a throwaway `$HOME`. That does not + keep the hook quiet; it emits a notice, and an earlier wording of it asserted the MCP + server "will not start", which is false under mocks. Cases then opened with that blocker + instead of the answer and their rubrics failed. The notice now states only what is known + and says to carry on if the tools answer. Re-read it before blaming a case. Two prompts still say the session cannot reach a real toolchain, which is a different problem and true regardless. +- **A FAIL clause can punish a better answer than the PASS clause asked for.** Three + green-gate rubrics failed replies that did more than required: one listed every failure + and then asked for fuller analyzer text, one offered the required test and also noted the + method might be dead code, one stated the format rule for a run the prompt forbade. Each + FAIL clause was narrow enough that the extra thoroughness tripped it. Write the FAIL + clause for the wrong answer, not for any departure from the shortest right one. +- **`PASS%` in the results table is runs scoring exactly 1.00**, not runs clearing + `--threshold`. A case at 0.90 shows `33%` while passing the threshold on every run. Read + `SCORE` against the threshold; read `PASS%` only when hunting flaky graders. +- **Judges are not stable.** Twelve judged runs of unchanged green-gate code left no `llm` + grader at 12/12; the range was 4/12 to 11/12. A case with several `llm` graders will fail + something most runs whatever the skill did. When a judge fails a reply that plainly + satisfies its rubric, convert the check to a `regex` on the skill's vocabulary rather than + rewording the rubric. + Beyond that: write prompts as a user would send them, name no skill in a prompt, grade mechanically where you can, include the cases where the skill must say no, and keep a negative control's rubric to the absence of the skill's vocabulary. @@ -201,6 +219,13 @@ Match the skill's own vocabulary. Keep `llm` graders for short output; for anyth Every run starts in an empty workspace. `_fixture/fixture.sh` recreates the neutral Flutter skeleton, hooked up through `context.scaffold_script`, and runs **only with `--scaffold`**. +It writes `pubspec.yaml`, `lib/counter.dart` and `test/counter_test.dart`. The two source +files exist so a case that drives the MCP tools has something real on disk to act on, and +they are deliberately the most boring code that satisfies that: a plain class with no +Flutter import, and one `test()` with one `expect()`. Nothing there is a VGV convention, +because everything there is visible to the no-plugin arm of every other skill's cases. Do +not grow them into a widget, a bloc, a `pumpApp` or a mocked dependency. + `context.scaffold_script` will not take a path that leaves the case directory, but it does resolve a symlink inside it, so each case's `fixture.sh` is a symlink to the one script. A new case needs its own: @@ -221,8 +246,8 @@ and runs fail at scaffold time. CI runs on Linux and is unaffected. every run reports on a `mocked:` line. Real servers start only with `--allow-real-servers` or `--mocks off`, neither of which is used here. -This applies to eval runs only. A real session still reaches the real `very_good mcp` -server. +This applies to eval runs only. A real session still reaches the real `very_good mcp` and +`dart mcp-server` servers. ```text evals/mocks/very-good-cli/ @@ -231,10 +256,14 @@ evals/mocks/very-good-cli/ ├── packages_check_licenses.md ├── packages_get.md └── test.md +evals/mocks/dart/ +├── _tools.json # the analyze_files and dart_format entries only +├── analyze_files.md +└── dart_format.md ``` -Frontmatter is optional and the body is the tool result. Three of the four here have no -frontmatter, which is the same as `type: fixed`. +Frontmatter is optional and the body is the tool result. Five of the six bodies here have +no frontmatter, which is the same as `type: fixed`. | Key | Default | Purpose | | ------------ | ------- | --------------------------------------------------------------- | @@ -261,8 +290,14 @@ Four things that are not obvious: that is `create`, so only `create.md` carries an `expect`. Grade argument *choices* with `tool_used` or a `regex` against `mock_calls`. - **`_tools.json` drives the schema the model sees.** It is the `tools/list` result, the - object with the `tools` array. Regenerate it after a Very Good CLI release by speaking - MCP to `very_good mcp` over stdio; a stale one teaches a schema the CLI no longer has. + object with the `tools` array. Regenerate it by speaking MCP over stdio to the server + itself — `very_good mcp` for one, `dart mcp-server --enable dart_format` for the other — + and capturing the `tools/list` result; a stale one teaches a schema the CLI no longer + has. Answer the server's `roots/list` request back to the client or `dart` blocks. +- **A `_tools.json` may be narrower than the server.** `dart` exposes 15 tools; the mock + ships the 2 `green-gate` names. That removes the listed-but-unmocked case entirely, and + keeps `roots` out of reach — a separate tool whose real handshake a mock cannot perform. + Adding one later is additive: capture it from the same probe, drop in a `.md`. The mocks are reachable only when `check_vgv_cli` returns `unverifiable`. In a run `very_good` resolves on PATH but `very_good --version` answers nothing, because it is a @@ -278,6 +313,21 @@ answer to the no-plugin arm, exactly as a non-neutral fixture does. `packages_check_licenses` returns one `GPL-3.0` and one `unknown` among twelve permissive licenses, so a case has something real to flag. +**The mock set is green everywhere else**, so a case can drive every gate and reach an +exit. `analyze_files` returns no errors, `dart_format` reports `0 changed`, and `test` +passes at 100%. **Keep those numbers consistent with the fixture on disk.** `dart format` +on the seeded package really does print `Formatted 2 files (0 changed)`, and the test body +reports the one test and three executable lines the fixture actually holds. An earlier +version claimed 38 tests and 378 lines; the model read the workspace, caught the mock +lying, and refused to call the package green — so a mock that contradicts the fixture does +not merely go unread, it fails the case. A failure-path case supplies its own `mocks/` +override rather than reddening the shared set. + +One body is deliberately lossy. Real `dart_format` output opens with +`dart format in :`, which a fixed body cannot know, so the mock drops that +line and keeps the `Formatted N files (M changed)` summary the gate is actually read +from. + **Converting a case to drive a tool invalidates every rubric that read the narration.** The model stops describing the call and just makes it, so a blind judge sees no evidence and fails a rubric that was passing. Four rubrics went stale this way in one pass, two of them @@ -289,8 +339,8 @@ only those judging something still in the reply. prompt it replaced: it routes, calls, reads the answer, then writes the reply. `ui-package-scaffolds-with-app-ui-package-template` measured 11 to 15 turns where the suite's usual `max_turns: 12` and `timeout_seconds: 600` had been ample, and hit both -limits. The four tool-driving cases carry `max_turns: 20` and `timeout_seconds: 900`. A cap -breach scores the case 0 with no failing grader, so it reads as a content failure. +limits. A tool-driving case therefore carries `max_turns: 20` and `timeout_seconds: 900`. +A cap breach scores the case 0 with no failing grader, so it reads as a content failure. **A mocked tool is not there in the no-plugin arm**, so a `tool_used` grader on one fails for free and takes any grader that needs the tool's output with it. Unlike @@ -305,7 +355,7 @@ plugin and 0.00 without, on three runs each. ```bash E="claude plugin eval . --scaffold" -$E # all 101, both arms +$E # all 102, both arms $E --tag bloc --tag testing # two skills $E --case bloc-writes-sealed-events-and-states # one case $E --ablation none # with-plugin arm only, half the cost @@ -331,7 +381,7 @@ Read the two arm scores, not the total. The without-arm is supposed to score bad baseline, so its Δ reads low. Routing is decided before the switch, so the pin never explains a routing miss. -Costs: **$13** for 100 cases in one arm, **$25** for both, roughly **$0.12 per run**. At +Costs: **$13** for 102 cases in one arm, **$25** for both, roughly **$0.12 per run**. At `--runs 3` a two-arm sweep is six runs per case, so budget around **$75**. `-j` up to 8 shortens wall clock. @@ -353,7 +403,7 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on- - CI needs `ANTHROPIC_API_KEY`, having no Claude Code session, and `--trust-plugin`. - `--ablation none` and `--ablation with-without` weight the baseline differently. Compare runs from one mode at a time. -- The job has a one-hour ceiling. 100 cases in one arm measured roughly 35 minutes at +- The job has a one-hour ceiling. 102 cases in one arm measured roughly 35 minutes at `-j 4`. A two-arm run at `--runs 3` is 600 runs and does not fit. --- @@ -365,9 +415,10 @@ scoped by `--tag` to the changed skills, with-plugin arm only, and `continue-on- `aggregate-result.json`, only a `tracePath` into a sandbox deleted unless `--keep-temp` is passed. `report.html` does show the judged text. - **Judge calibration.** Most graders are `llm` with no human-labelled gold set. -- **Tool execution, mostly.** One case drives a mocked tool, - `license-compliance-runs-check-with-full-license-info`. The other tool-driven skills are - still graded on the calls they narrate. +- **Tool execution, partly.** Six cases assert a call to a mocked tool, across + `create-project`, `green-gate`, `license-compliance` and `ui-package` — find them by + grepping the graders for `mcp__plugin`. Every other tool-driven case is still graded on + the calls it narrates. - **Stable routing.** Whether a skill activates is nondeterministic, which is why routing is a `tool_used` grader rather than inferred from content. - **Prose in a `SKILL.md`.** Deliberate: an earlier version asserted a hundred `contains` diff --git a/evals/_fixture/fixture.sh b/evals/_fixture/fixture.sh index 45778ef..956eb0c 100755 --- a/evals/_fixture/fixture.sh +++ b/evals/_fixture/fixture.sh @@ -3,10 +3,33 @@ # # THE ONLY COPY. Every case's fixture.sh is a symlink to this file. Edit it here. # -# KEEP THIS NEUTRAL. The pubspec must not name anything a skill teaches. +# KEEP THIS NEUTRAL. Neither the pubspec nor the seeded source may name anything a +# skill teaches: no widget, no bloc, no mocktail, no golden, no pumpApp. The source +# exists so a case that drives the MCP tools has something real on disk, nothing more. set -euo pipefail mkdir -p lib test -touch lib/.gitkeep test/.gitkeep +cat > lib/counter.dart <<'DART' +/// Keeps a running total. +class Counter { + int _value = 0; + + int get value => _value; + + void add(int amount) => _value += amount; +} +DART +cat > test/counter_test.dart <<'DART' +import 'package:eval_fixture_app/counter.dart'; +import 'package:flutter_test/flutter_test.dart'; + +void main() { + test('add increases the value', () { + final counter = Counter(); + counter.add(3); + expect(counter.value, 3); + }); +} +DART cat > pubspec.yaml <<'YAML' name: eval_fixture_app description: Scratch Flutter app used as working-directory context for behavior evals. diff --git a/evals/accessibility/NOTES.md b/evals/accessibility/NOTES.md deleted file mode 100644 index 7aa4613..0000000 --- a/evals/accessibility/NOTES.md +++ /dev/null @@ -1,117 +0,0 @@ -# accessibility eval notes - -## Grading - -Mixed grading. `accessibility-wraps-cupertino-and-announces-on-ios` and -`accessibility-adds-non-drag-alternative-to-dismissible` are graded on the artifact, -because they ask for a widget class. The first four cases are graded on what the response -narrates, the question it asks, the report it writes, the request it declines. - -Prompts name no skill, so the routing grader catches a routing failure directly. - -Prompts are self-contained. The fixture has no source in `lib/`, so a prompt about "my -SettingsView" earns a refusal and every grader fails for an unrelated reason. The widget -under audit is pasted in. - -Phases 1 and 2 of the workflow call `AskUserQuestion`, which neither arm exposes, since -`allowed_tools` is `Read, Glob, Grep, Skill`. So the level-and-platform gate is graded on -the question the response *narrates* in text, not on a tool call, the same treatment the -MCP-backed skills get. - -Most cases state the WCAG level and the platforms up front. Without them the skill -correctly stops at Phase 1 and asks, which would leave nothing to grade in the -remediation cases. Exactly one case, -`accessibility-asks-level-and-platform-before-auditing`, omits them on purpose, to grade -the gate. - -WCAG criterion IDs are graded mechanically, `2.5.8`, `2.1.1` and so on, because Core -Standards pins a specific ID to each rule, and the without-arm tends to answer in generic -"add a label" prose without citing WCAG 2.2 at all. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### accessibility-asks-level-and-platform-before-auditing - -**Discriminates.** The rules under test are SKILL.md's "Never assume AA" and "Begin every -audit by asking which of the six platforms are targeted". An unaided model assumes AA, -audits the widget on the spot, and talks about iOS and Android only. `AA + selected AAA` -is verbatim from the Phase 1 option list and is not phrasing a bare model volunteers. - -**Notes.** `names-linux-platform` grades the six-platform list from Phase 2. Linux is the -tell, because an unaided answer talks about iOS and Android, not about Orca on Linux -desktop. - -### accessibility-audits-with-criterion-ids-and-severities - -**Discriminates.** The report shape comes from `references/audit-templates.md`: per finding -a WCAG ID plus name plus level, a CRITICAL/MAJOR/MINOR severity from the severity guide, -and a before/after fix. The planted 16x16 close button is the 2.5.8 Target Size (Minimum) -violation and its GestureDetector is 2.1.1. An unaided model returns generic "add a label" -prose with no WCAG 2.2 IDs, and grades severity as High/Medium/Low or axe-style -Serious/Moderate rather than the skill's three labels. - -**Notes.** `uses-critical-severity-label` grades the upper-case severity labels, which are -the skill's own vocabulary, from the severity guide in `references/audit-templates.md`. - -### accessibility-declines-gesture-detector-tap-target - -**Discriminates.** The rule is SKILL.md's "Never use bare GestureDetector for tap targets -... GestureDetector is pointer-only and unreachable via keyboard or switch access". The -prompt asks for it by name and supplies a plausible excuse, so complying is the failure -mode. An unaided model obliges, ships the GestureDetector version as its recommendation -with at most a passing caveat, and never cites 2.1.1 Keyboard as the criterion being -broken. - -**Notes.** `names-accessible-tap-widget` is the documented replacement list from the -Gesture Detector core standard. - -### accessibility-declines-exclude-semantics-on-actionable-content - -**Discriminates.** Two rules at once: "Never use ExcludeSemantics on non-decorative -content", and the MergeSemantics correctness rule from `references/examples.md`, which -merges the static label and value pair only and keeps the interactive child independently -focusable. An unaided model either complies with the ExcludeSemantics request or merges -the whole Row, folding the button's role away. - -**Notes.** `fix-uses-merge-semantics` includes the opening paren, so the fix has to arrive -as code. The without-arm named MergeSemantics in prose, as a possible cause rather than -the fix. - -### accessibility-wraps-cupertino-and-announces-on-ios - -**Discriminates.** Two skill facts carry this case: the Cupertino Semantics core standard -("Cupertino widgets ship with weaker semantic defaults ... Always wrap them in -Semantics(label:, value:, button:)"), and the iOS reference's gotcha that -`liveRegion: true` does not auto-announce on iOS, Flutter issue #45968, so -SemanticsService.announce is required as well. An unaided model adds liveRegion and treats -it as sufficient, and either leaves the CupertinoSwitch bare or swaps it for a Material -Switch instead of wrapping it. - -### accessibility-adds-non-drag-alternative-to-dismissible - -**Discriminates.** Only the criterion ID separates the arms here. The rule is WCAG 2.2 -2.5.7, "Every dragging-based function must offer a non-drag alternative on the same -screen. `Dismissible` needs an explicit delete button", and the pattern in -`references/examples.md` keeps the Dismissible and adds a tooltipped IconButton in the -ListTile's trailing slot. A bare model reaches that shape on its own but does not name the -criterion it is remediating. The skill does, everywhere. - -**Notes.** `keeps-the-dismissible` exists because the reference pattern keeps the -Dismissible and adds the button beside it, so the swipe must survive the fix. -`labels-the-alternative-control` exists because an icon-only control needs a Tooltip or -Semantics label of its own, from the Icon Buttons core standard and WCAG 4.1.2. -`cites-2-5-7-dragging-movements` is the one assertion the without-arm misses, since Core -Standards pins 2.5.7 to drag alternatives and `references/examples.md` titles the pattern -with it. - -### accessibility-stays-out-of-plain-dart-work - -**Discriminates.** What must not appear is the skill firing at all, and any Semantics, -semanticLabel, WCAG, screen-reader or accessibility vocabulary in an mm:ss formatter. - -**Notes.** `answers-the-question` grades task success mechanically rather than by the -judge. The judge never sees the prompt, so "did it answer the question" is unanswerable -from the output alone. diff --git a/evals/animations/NOTES.md b/evals/animations/NOTES.md deleted file mode 100644 index 83ce4ec..0000000 --- a/evals/animations/NOTES.md +++ /dev/null @@ -1,115 +0,0 @@ -# animations eval notes - -## Grading - -Graded on the artifact. Five cases ask for Dart and grade the code that comes back. -`animations-reviews-planted-violations` is the exception, because it asks for a list of -findings and is graded on the review prose. The negative control asks for Dart too, but -grades the absence of motion vocabulary. - -Prompts name no skill, so the routing grader catches a routing failure directly. - -The load-bearing thing this skill teaches is Material 3 motion tokens. SKILL.md says never -hardcode `Duration(milliseconds: ...)` or use `Curves.*` for new code, and the bare model -hardcodes both every time. So most cases grade tokens twice, once positively with -`Durations.` and `Easing.`, and once as the absence of the forbidden form. - -Trap: every hardcoded-value negative is anchored to the argument position, `duration: -Duration(` or `curve: Curves.`, never to the bare type name. An unanchored -`Duration\(milliseconds:` also fires on a response that names the bad form in a comment -while writing the good one, which is a false failure. - -Not measurable here: the Core Standard "Clarify visual intent when the request is -ambiguous". A single-shot harness gives the model nobody to ask, so a deliberately vague -prompt grades the harness rather than the skill. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### animations-uses-m3-motion-tokens-for-implicit-animation - -**Discriminates.** Without the skill the model reaches for a StatefulWidget with an -AnimationController and hardcodes 300ms and Curves.easeOut. - -**Notes.** `no-hardcoded-duration-or-curve` is the Anti-Patterns section's "Hardcoded -magic values" example, inverted. Four cases carry a grader of this name and the four -patterns deliberately differ: every one is `not_contains`, so widening them to a shared -union would make each stricter and fail correct answers. `centralizes-motion-constants` -alone forbids `= Duration(milliseconds`, because there the point is that the constant -belongs in one place; a case that may legitimately declare a constant must not inherit -that arm. Reviewed and left diverging on purpose. `no-controller-constructed` anchors on `vsync:`, the -mechanical proof a controller was constructed, with `\s*` because the argument is -sometimes written without the space. - -### animations-declines-controller-for-simple-fade - -**Discriminates.** Without the skill the model complies. It writes the StatefulWidget, the -controller and the AnimatedBuilder exactly as asked. - -**Notes.** `uses-animated-opacity` comes from SKILL.md, Anti-Patterns, "Using explicit -when implicit suffices", where the good form for this exact request is AnimatedOpacity. -`no-controller-constructed` is anchored on `vsync:`, the one argument a constructed -controller cannot omit. Matching AnimationController by name would fail a decline that -explains itself. - -### animations-staggers-on-a-single-controller - -**Discriminates.** Without the skill the model spins up a controller per animation and -hardcodes each duration, so the two-controller negative fires. - -**Notes.** `staggers-with-interval` comes from Performance, Do Not: "Do not create -multiple AnimationController instances for animations that share timing, use Interval on -a single controller." `only-one-controller` is that same rule in mechanical form, two -`AnimationController(` constructions anywhere in the response, with `[\s\S]*` because the -two are lines apart. `uses-single-ticker-mixin` is Core Standards, one controller means -SingleTickerProviderStateMixin. `uses-duration-tokens` tracks -`references/staggered-animations.md`, which drives the controller with `Durations.long2` -and curves each Interval with `Easing.`. - -### animations-custom-page-transition-via-go-route-data - -**Discriminates.** Without the skill the model wraps the page body in a FadeTransition -inside build, or passes a hand-rolled Duration and Curves value. - -**Notes.** `overrides-build-page` grades the mechanism the skill prescribes, an override -of buildPage on the GoRouteData rather than a wrapper widget inside the page's build -method. `uses-emphasized-easing` follows `references/page-transitions.md`, which curves -every page transition with `Easing.emphasizedDecelerate`. - -### animations-reviews-planted-violations - -**Discriminates.** Four violations are planted: a hardcoded duration and curve, the wrong -ticker mixin, an ExpensiveChart rebuilt every frame, and an animated width. Without the -skill the review stops at the missing dispose and does not mention M3 tokens, the mixin -choice, or the layout cost. - -**Note.** Graded on prose, not Dart, because the prompt asks for findings. - -**Notes.** `uses-single-ticker-mixin` is Core Standards, SingleTickerProviderStateMixin -for one controller and TickerProviderStateMixin only for several. - -### animations-centralizes-motion-constants - -**Discriminates.** Without the skill the model either inlines timings per feature or -writes a constants class full of raw `Duration(milliseconds: ...)`, which the third arm of -the negative catches. - -**Notes.** `declares-app-motion` comes from Core Standards: "durations, curves, and -offsets go in named constants or a centralized AppMotion class, not inline." -`constants-are-duration-tokens` exists because the constants must themselves be M3 tokens, -not a rename of magic numbers. - -### animations-stays-out-of-non-motion-work - -**Discriminates.** What must not appear is the skill firing at all, any animation widget -such as AnimationController, AnimatedBuilder, TweenAnimationBuilder or -CustomTransitionPage, and the token vocabulary `Durations.`, `Easing.` and `AppMotion`. -Nothing else catches the skill firing where it should not. - -**Notes.** `returns-a-color` and `parses-the-hex-string` grade task success mechanically -rather than by the judge. The judge never sees the prompt, so "did it parse the hex -string" is unanswerable from the output. `parses-the-hex-string` carries `(try)?` because -a correct answer that returns null on bad input uses `int.tryParse`, which a bare -`int\.parse` would fail. diff --git a/evals/bloc/NOTES.md b/evals/bloc/NOTES.md deleted file mode 100644 index 9810a33..0000000 --- a/evals/bloc/NOTES.md +++ /dev/null @@ -1,83 +0,0 @@ -# bloc eval notes - -## Grading - -Cases 1, 2, 3 and 5 are graded on the artifact: they ask for Dart and the graders read -the code that comes back. Case 4 is graded on the decision it narrates, two rubrics and -nothing mechanical, because the right answer is a refusal plus an alternative rather -than a snippet. - -Prompts name no skill, so the routing grader catches a routing failure directly instead -of leaving it to show up as unexplained content failures downstream. - -Prompts are self-contained. The fixture has no source in `lib/`, so a prompt about "my -AuthService" earns a refusal and every grader then fails for an unrelated reason. Paste -the class into the prompt rather than adding source to the fixture, which would leak -answers to the without-plugin arm. - -Trap specific to this skill: the fixture pubspec must not list bloc, flutter_bloc, -equatable or mocktail. An earlier version did, the baseline read the dependency list and -inferred the conventions these cases exist to measure, and bloc's measured lift -collapsed from +10 points to +1. - -Measured baseline, first full run under the previous harness: 5/5 with the plugin -against 1/5 without it, negative control excluded because a model without the skill -passes the negative routing grader for free. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### bloc-writes-sealed-events-and-states - -**Discriminates.** Sealed classes, Equatable, props and past-tense events are baseline -knowledge, so an unaided model writes them. `LoginInProgress` is the one it misses, -reaching for `LoginLoading` instead, so that is the grader carrying this case. - -**Note.** The state naming table pins the four names for exactly this bloc: -`LoginInitial`, `LoginInProgress`, `LoginSuccess`, `LoginFailure`. -`pinned-in-progress-name` is the grader that reads the one the baseline misses. - -**Note.** `past-tense-event-names` judges the concrete event subclasses only. An earlier -wording said "every event class name", which the sealed base class `LoginEvent` can never -satisfy, so a correct answer failed on its own base class. - -### bloc-tests-with-bloc-test-and-mocktail - -**Discriminates.** Without the skill the model writes raw `test()` calls that subscribe -to bloc.stream by hand, and may reach for mockito over mocktail. - -**Note.** `imports-mocktail` matches `mocktail/mocktail.dart` rather than the full -`package:mocktail/...` import. The prefix was dropped to work around a previous-harness -limitation that no longer exists, and the pattern is kept unchanged so the numbers stay -comparable. - -**Note.** `declares-mock-class` allows a leading underscore because -`skills/bloc/references/testing.md` still shows a public `MockTodoRepository`, so either -form is a faithful reading of the skill. - -### bloc-separates-page-from-view - -**Discriminates.** Without the skill the model writes one widget that both creates the -bloc and builds the form, which `page-view-split` rejects by name. - -### bloc-refuses-bloc-to-bloc-dependency - -**Discriminates.** Without the skill the model obliges with a `BlocProvider.value` or a -constructor taking AuthBloc, and never says the pattern is banned. - -### bloc-upgrades-cubit-to-bloc-for-transforms - -**Discriminates.** Without the skill the model keeps the Cubit and debounces with a -Timer inside it, or pushes the debounce into the widget. - -### bloc-stays-out-of-plain-dart-work - -**Discriminates.** Must NOT appear: the skill in the tool calls, any flutter_bloc widget -or `extends Bloc<`/`extends Cubit<`, blocTest, or state-management framing in the prose. -Nothing else catches the skill over-firing. - -**Note.** Task success is graded mechanically by `answers-the-question`, not by the -judge. The judge never sees the prompt, so "did it answer the question" is unanswerable -from the output alone. diff --git a/evals/create-project/NOTES.md b/evals/create-project/NOTES.md deleted file mode 100644 index 5c15364..0000000 --- a/evals/create-project/NOTES.md +++ /dev/null @@ -1,94 +0,0 @@ -# create-project eval notes - -## Grading - -No MCP server is available to these runs, so the cases grade the decisions the skill -drives — template inference, name normalization, what it declines — not the create call. -Only assert template names that `SKILL.md` teaches, never ones the MCP server supplies at -run time. - -Consequence worth knowing before writing a case here: with no `create` tool the skill -cannot execute, and every response says so and hands the user a command to run instead. -Do not assert that the response *uses* the MCP tools — that is unsatisfiable in this -environment, and asserting it measured the harness rather than the skill. Wiring the real -server in is not the fix either: `create` would scaffold a project into the fixture on -every run. Mocking the server in the native harness is a separate, later piece of work, -and none of these cases assume it. - -`create-project-infers-dart-package-for-api-client` carries no routing grader, and that is -deliberate: a one-word template question does not activate the skill, measured at 0/2 with -the answer correct both times, so grading routing there would only ever report the harness. -Every other case grades routing. - -Measured baseline, first full run under the previous harness: 5 of 6 with the skill against 0 of 6 without -it, negative control excluded. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### create-project-infers-dart-package-for-api-client - -**Discriminates.** The inference is that a package with no Flutter dependency takes -dart_package, which is also the layered-architecture rule for data and repository layers. -A bare model reaches for flutter_package because the monorepo is a Flutter one, or answers -with prose instead of a template name. - -### create-project-asks-for-organization-when-required - -**Discriminates.** The Key Domain Knowledge under test is that app, plugin and game -templates require an organization and silently take a placeholder when it is skipped, so -the skill asks rather than proceeding. A bare model invents com.example, or scaffolds with -the placeholder the template falls back to, and never raises the question. - -### create-project-scopes-dependency-install-to-the-new-project - -**Discriminates.** A bare model runs the install at the monorepo root, names `flutter -create`, and invents an organization without comment. - -**Grader notes.** `names-packages-get` is a regex because both surface forms count: the -MCP tool is `packages_get`, the shell equivalent is `very_good packages get`. An earlier -substring check on `packages_get` failed a correct answer that used the shell form, which -tested syntax rather than knowledge. `directory-points-at-new-project` is the load-bearing -detail in either form, `--directory apps/storefront` or `directory: 'apps/storefront'`; -this is the part a bare model does not know. `no-invented-organization` accepts either -asking for the org or flagging it as required, so it does not demand a specific phrasing. - -### create-project-asks-when-the-template-is-ambiguous - -**Discriminates.** A bare model picks a template and starts scaffolding, or asks "which -template do you want, dart_package or flutter_package?" — the exact phrasing the skill's -anti-pattern table rules out. - -### create-project-does-not-over-ask - -**Discriminates.** A bare model runs a questionnaire for description, output directory and -application id before it will do anything. - -### create-project-plans-dependency-install - -**Discriminates.** The plan has to scaffold through Very Good CLI with an explicit -template and not stop at creation, because dependencies get installed too. A bare model -plans `flutter create my_store --org com.example` and calls it done, leaving the install -out entirely. - -**Grader notes.** Both judged graders scored 0 on a response that listed -`very_good create flutter_app my_store --org com.example` followed by `very_good packages -get`, which is the answer the case is looking for. Two rubric holes caused it. The -scaffolding rubric failed on the literal string `flutter create`, which the response only -mentioned as a labeled fallback, so its FAIL clause was wider than the complement of its -PASS clause; it now fails only on the step the response settles on. Both rubrics also said -"the plan", and the response opened by saying it could not execute anything in this -environment, which read to the judge as no plan at all. Both now state that a response -narrating steps it cannot run is judged on the steps it lists. - -### create-project-stays-out-of-existing-project-work - -**Discriminates.** create-project must not be invoked, and none of its vocabulary may -appear: no `very_good create`, no template name, no `--org`, and no suggestion to scaffold -a new project or package around the class. - -**Grader notes.** Task success is graded mechanically by `adds-retry-count-field`, not by -the judge. The judge never sees the prompt, so "did it edit the class" is unanswerable -from the output alone. diff --git a/evals/create-project/create-project-does-not-over-ask/prompt.md b/evals/create-project/create-project-does-not-over-ask/prompt.md index 44bb024..bbc7af7 100644 --- a/evals/create-project/create-project-does-not-over-ask/prompt.md +++ b/evals/create-project/create-project-does-not-over-ask/prompt.md @@ -1,6 +1,6 @@ --- -max_turns: 12 -timeout_seconds: 600 +max_turns: 20 +timeout_seconds: 900 allowed_tools: [Read, Glob, Grep, Skill] tags: [create-project] description: The other half of the asking rule. With name and organization in hand, nothing optional gets interrogated. diff --git a/evals/dart-flutter-sdk-upgrade/NOTES.md b/evals/dart-flutter-sdk-upgrade/NOTES.md deleted file mode 100644 index bd03715..0000000 --- a/evals/dart-flutter-sdk-upgrade/NOTES.md +++ /dev/null @@ -1,109 +0,0 @@ -# dart-flutter-sdk-upgrade eval notes - -## Grading - -Graded on narration, never on an artifact. The skill's real job is editing -`.github/workflows/*.yml` and `pubspec.yaml`, running `pub get` / `analyze`, and checking -the diff before a PR. None of that can happen here: the fixture is a bare app with no -`.github/`, no monorepo and no lockfile, and the cases grant no Edit, Bash, or MCP tools. -So every case grades the decisions the response narrates — which key it writes in which -file, which version format goes where, what it verifies, what it refuses to put in the PR. - -The load-bearing distinction throughout, and the thing the no-plugin arm gets wrong, is that -the same upgrade is written three different ways: CI Flutter is `MAJOR.MINOR.x` (literal -wildcard, no caret), CI Dart is an exact patch under a `dart_sdk` key, and pubspec is -`^MAJOR.MINOR.PATCH`. - -Trap for anything added here: do not assert a Dart-for-Flutter version mapping as fact. -The skill resolves it from , and no arm can -fetch that page, so "Flutter 3.41.0 ships Dart 3.11.0" is unknowable rather than wrong. -Where a case needs both numbers to grade a file edit, the prompt supplies both; the one -case that withholds the Dart version grades the lookup itself, not the answer. - -These responses are YAML, not Dart, so no syntax grader applies to any case in this -skill. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### dart-flutter-sdk-upgrade-pins-ci-flutter-with-patch-wildcard - -**Discriminates.** A bare model writes the exact patch or a caret into CI, and treats the -Flutter number as the Dart number. The pasted workflow says `flutter_channel: stable` on -purpose — a workflow that already read `flutter_version: "3.35.x"` would hand both arms -the format the skill exists to teach, and the case would measure copying. - -- `ci-flutter-patch-wildcard` — `[^\n^]{0,3}` absorbs the optional quoting, and excluding - `^` from the class makes `flutter_version: "^3.41.x"` fail. There is deliberately no - mirroring negative grader on `3.41.0`: a correct answer writes the exact patch in order - to reject it. -- `flutter-version-is-not-dart-version` — a one-sided guard against conflation, not a - demand that Dart be discussed. The prompt is CI-only and `flutter_version` takes the - Flutter number, so a correct minimal answer has nowhere to put a Dart version and never - mentions one. Two successive rubrics failed that answer anyway: the first required it to - name the pubspec `sdk:` constraint, and the replacement still failed a response that - "never draws the distinction at all". The rubric now passes on silence and fails only on - the conflation itself — the Flutter number asserted as the Dart number, or written into - a slot the response labels `sdk:` or `dart_sdk:`. That is still what the bare arm does - when it volunteers a pubspec edit, so the discriminator survives the narrowing. Note - also that the case tolerates one miss: routing at weight 3 plus three content graders is - weight 6, so a single content miss scores 5/6 = 0.83 and passes. - -### dart-flutter-sdk-upgrade-uses-dart-sdk-key-for-pure-dart-package - -**Discriminates.** A bare model reaches for `flutter_version` and `flutter_package.yml`, -or adds a `flutter:` constraint to a package that has no Flutter dependency. The workflow -body is described rather than pasted for that reason: those two names are what the skill -knows and a bare model guesses at. - -- `ci-dart-sdk-exact-patch` — `^` is excluded from the class rather than bounded by - length: with a plain `[^\n]{0,3}` the wrong answer `dart_sdk: "^3.11.0"` matches, since the - space, the `"` and the `^` are three characters. -- `pubspec-caret-constraint` — the pubspec form of the same version. `\b` is what keeps - this off `dart_sdk:` — without it a caret-in-CI answer satisfies both patterns and - scores full marks. - -### dart-flutter-sdk-upgrade-resolves-bundled-dart-version-before-editing - -**Discriminates.** A bare model states a bundled Dart version from memory, or reuses the -Flutter number, and starts editing. This is the only case that withholds the Dart -version; no arm can fetch the archive, so the graded behavior is asking rather than -answering. - -### dart-flutter-sdk-upgrade-declines-dependency-bumps-and-lint-fixes - -**Discriminates.** A bare model is helpful and ships all three changes in one pubspec, or -caveats the dependency bump while still writing it. Both target versions are supplied so -the in-scope half stays gradeable mechanically. - -### dart-flutter-sdk-upgrade-reports-pub-get-conflict-instead-of-resolving - -**Discriminates.** A bare model unblocks the user by raising `very_good_analysis` and -moving on, which is the mixed diff the skill exists to prevent. - -### dart-flutter-sdk-upgrade-updates-every-pubspec-and-checks-diff-scope - -**Discriminates.** The plan has four parts: every package's pubspec edited individually, -the shared CI workflow once, `pub get` and `analyze` per package, then the diff scope -check before opening the PR. A bare model edits a root pubspec, verifies once at the repo -root, and never checks that the changed-file list holds only workflows and pubspecs. - -- `diff-name-only-check` — the documented scope check. Both halves matter: `git diff` - with a name-only listing is what proves the PR touched nothing else. - -### dart-flutter-sdk-upgrade-stays-out-of-dependency-edits - -Negative control. - -**Discriminates.** Must NOT appear: the skill firing, `flutter_version` / `dart_sdk` / -`very_good_workflows` / the release archive, or a bumped `sdk:` constraint. Deliberately -adjacent — a pubspec edit with version numbers in it — so a skill that fires on any -pubspec work is caught here. - -- `answers-the-question` — task success is graded mechanically, not by the judge: the - judge never sees the prompt, so "did it add the dependency" is unanswerable from the - output alone. -- `sdk-constraint-untouched` — the SDK constraints must come back untouched. `\b` so a - `dart_sdk:` line — itself a leak — cannot satisfy this. diff --git a/evals/green-gate/NOTES.md b/evals/green-gate/NOTES.md deleted file mode 100644 index f28cb96..0000000 --- a/evals/green-gate/NOTES.md +++ /dev/null @@ -1,108 +0,0 @@ -# green-gate eval notes - -## Grading - -Graded on narration, not on the artifact. No MCP server is available to these runs and the -fixture has no source in `lib/` or `test/`, so the loop this skill exists to run cannot -execute here. The cases grade the decisions the skill narrates: which tool it says it -would call for each gate, the arguments it would pass, the order it would run them in, -what it refuses to weaken, and when it stops and escalates. Never assert that a response -*called* a tool — unsatisfiable here, and it measures the harness. Mocking the MCP servers -in the native harness is a separate, later piece of work, and none of these cases assume -it. - -Every prompt therefore ends by asking for a plan or a verdict rather than for a run, and -every scenario is stated in the prompt rather than left on disk. A prompt that says "fix -my package" earns "there is no code here" and every grader then fails for an unrelated -reason. Prompts describing a repo also say outright that it is not on disk; without that -the model spends its answer on "lib/ only has a .gitkeep". - -Routing is the dominant failure mode here. On a full run 4 of 6 positive cases missed -routing, and every one of those was a prompt asking *about* the gates rather than for a -run. The cause was legible: on `green-gate-plans-the-four-gates-in-order` the model named -the skill in its own answer ("this is what the green-gate skill automates end to end. Want -me to invoke it?") and then improvised a `flutter analyze` / `flutter test --coverage` -shell plan, because the skill read as a runner and the prompt said not to run anything. -The skill's `description` now claims gate *configuration* questions too, and Core -Standards carries a plan-only directive. - -Two things are deliberately absent: - -- No Dart-parse grader. No case demands Dart code, so no case has fenced dart blocks to - parse. -- No negative check on `flutter test`. The skill teaches *why* that command is - hook-blocked, so a correct answer often names it in order to reject it. The tool-routing - rule is graded by the MCP tool names instead. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### green-gate-plans-the-four-gates-in-order - -**Discriminates.** Unrouted, the bare model improvises a `flutter analyze` / `flutter test ---coverage` shell plan and puts format before analyze. - -**Grader notes.** `names-analyze-files` exists because the analyze gate goes through the -Dart MCP server, not `dart analyze` in Bash. `names-check-ignore` and `names-min-coverage` -are the coverage triple with `applyFixes`; `check_ignore` is the one a bare model never -produces, and omitting it makes the `// coverage:ignore` remedy a silent no-op. - -### green-gate-refuses-to-weaken-the-coverage-gate - -**Discriminates.** Without the skill the model is agreeable — it drops the threshold to -90, adds the ignore comment, and declares the package clean. - -### green-gate-refuses-to-carry-green-forward - -**Discriminates.** The rules are "Never cache green" and "exit only on observed numbers": -all four gates re-run in one final round, and the format gate is judged by changed count. -Without the skill the model takes the user's word for the earlier analyze and format -results, checks coverage only, and declares green. - -**Grader notes.** `format-judged-by-changed-count` covers the format gate's own trap: the -format tool reports success whether or not it rewrote anything, so the changed count is -the only signal it is green. - -### green-gate-excludes-generated-files-instead-of-ignoring-them - -**Discriminates.** Without the skill the model endorses the teammate's ignore comments and -suggests settling at a 95–97% target. - -**Grader notes.** `exclude_coverage` is the skill's parameter name, not general Dart -knowledge. - -**Grader notes.** Two cases here have moved between 2/3 and 3/3 across runs with no skill -change at all. That band is this suite's noise floor, so a single red rep is not evidence: -resist "fixing" green-gate itself off one. - -### green-gate-escalates-when-the-loop-stops-making-progress - -**Discriminates.** The no-progress trigger has three parts: an unchanged failure -fingerprint stops the loop, a standing "keep retrying" instruction does not override it, -and the escalation report carries per-failure detail plus a decision ask. Without the -skill the model accepts the blanket permission and keeps grinding, or stops with a bare -"4 errors remain" and no decision ask. - -**Grader notes.** `names-the-no-progress-trigger` matches the skill's own term for the -comparison that makes "no progress" decidable. - -### green-gate-budgets-per-package-across-a-monorepo - -**Discriminates.** The monorepo rules are a per-package iteration budget, -continue-on-failure with a per-package report, one shared `min_coverage` with no -per-package override, and pubspec.yaml-walk discovery shared by analyze and test. Without -the skill the model gives a generic "run the tests in each package" plan, invents a -per-package coverage override for the 62% package, and aborts on the first red one. - -### green-gate-stays-out-of-plain-function-work - -**Discriminates.** Nothing else catches the skill firing where it should not. Must NOT -appear: the skill's parameter vocabulary (`min_coverage`, `exclude_coverage`, -`check_ignore`, `analyze_files`, `lcov`), the phrase "quality gate", a verify-fix-rerun -loop, coverage targets, or the four-gate sequence. - -**Grader notes.** Task success is graded mechanically by `answers-the-question`, not by -the judge. The judge never sees the prompt, so "is the merge logic right" is unanswerable -from the output alone. diff --git a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/continues-past-a-red-package.md b/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/continues-past-a-red-package.md index 9c9a546..fd01a22 100644 --- a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/continues-past-a-red-package.md +++ b/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/continues-past-a-red-package.md @@ -2,15 +2,6 @@ type: llm --- -The response plans quality-gate work across several packages in one repository. +PASS if a package whose gates stay red does not stop work on the other packages. -PASS if both hold: - -- a package whose gates stay red does not stop work on the rest — the plan keeps going - and every package is attempted; and -- each package's own outcome is surfaced at the end, as a per-package rundown. It counts - whether that rundown is presented as results already in hand or as the summary the run - will produce when it finishes. - -FAIL if the first package that stays red stops the whole run, or if outcomes are only -reported in aggregate with no per-package breakdown at the end. +FAIL if the first red package halts the whole run. diff --git a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/discovers-roots-by-pubspec-walk.md b/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/discovers-roots-by-pubspec-walk.md deleted file mode 100644 index f872941..0000000 --- a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/discovers-roots-by-pubspec-walk.md +++ /dev/null @@ -1,16 +0,0 @@ ---- -type: llm ---- - -The response plans quality-gate work across several packages in one repository. - -PASS if both hold: - -- package roots are found by walking the repository for `pubspec.yaml` files; and -- that one discovered set of roots is what both the analyze step and the test step work - over. A sentence explicitly asserting the two sets are identical is not required — - feeding the same discovered roots into both gates is enough. - -FAIL if roots come from something other than a `pubspec.yaml` walk (a hardcoded list, a -directory-name convention, or no stated method at all), or if the analyze step and the -test step are scoped to visibly different sets of packages. diff --git a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/one-shared-coverage-target.md b/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/one-shared-coverage-target.md index 44d8ae4..247d7be 100644 --- a/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/one-shared-coverage-target.md +++ b/evals/green-gate/green-gate-budgets-per-package-across-a-monorepo/graders/one-shared-coverage-target.md @@ -1,7 +1,5 @@ --- -type: llm +type: regex +pattern: 'no per-package|per-package (override|min_coverage|target|threshold)' +flags: i --- - -PASS if the response states that one coverage target applies to every package, with no per-package override available, and calls that out as the reason the 62% package needs an explicit decision, either a lowered shared target or that package handled separately. - -FAIL if it silently applies two different targets, if it offers a per-package override, or if it states the shared target without calling out that the 62% package needs an explicit decision between a lowered shared target and handling that package separately. diff --git a/evals/green-gate/green-gate-escalates-when-the-loop-stops-making-progress/graders/per-failure-detail.md b/evals/green-gate/green-gate-escalates-when-the-loop-stops-making-progress/graders/per-failure-detail.md index e877244..402a2ef 100644 --- a/evals/green-gate/green-gate-escalates-when-the-loop-stops-making-progress/graders/per-failure-detail.md +++ b/evals/green-gate/green-gate-escalates-when-the-loop-stops-making-progress/graders/per-failure-detail.md @@ -1,7 +1,4 @@ --- -type: llm +type: regex +pattern: 'undefined_method\s+@\s+lib/src/sync_worker\.dart:104' --- - -PASS if all four remaining failures appear individually, each with its diagnostic code and its file and line number. Any layout counts: a list of `code @ file:line` entries, or a table with the code in one column and the location in another. - -FAIL if the response gives only a count such as "4 errors remain", or if any of the four is missing its diagnostic code, its file, or its line number. diff --git a/evals/green-gate/green-gate-refuses-to-carry-green-forward/graders/format-judged-by-changed-count.md b/evals/green-gate/green-gate-refuses-to-carry-green-forward/graders/format-judged-by-changed-count.md index 37b7e19..6f43e17 100644 --- a/evals/green-gate/green-gate-refuses-to-carry-green-forward/graders/format-judged-by-changed-count.md +++ b/evals/green-gate/green-gate-refuses-to-carry-green-forward/graders/format-judged-by-changed-count.md @@ -1,7 +1,5 @@ --- -type: llm +type: regex +pattern: '\b0 changed\b|zero changes?\b|changed count|files? changed|changed files?' +flags: i --- - -PASS if the format gate is judged by how many files the formatter changed, with zero changed being green. - -FAIL if the format step is treated as green because the call succeeded, without reading the changed count. diff --git a/evals/green-gate/green-gate-refuses-to-weaken-the-coverage-gate/graders/offers-the-missing-test.md b/evals/green-gate/green-gate-refuses-to-weaken-the-coverage-gate/graders/offers-the-missing-test.md index 789d425..6719752 100644 --- a/evals/green-gate/green-gate-refuses-to-weaken-the-coverage-gate/graders/offers-the-missing-test.md +++ b/evals/green-gate/green-gate-refuses-to-weaken-the-coverage-gate/graders/offers-the-missing-test.md @@ -1,7 +1,4 @@ --- -type: llm +type: regex +pattern: '```dart[\s\S]*?formatFree' --- - -PASS if the response offers the documented alternative: write the missing test for the uncovered method, following the project's testing standards, and re-run the gate. - -FAIL if it refuses with no alternative, if the alternative it offers is anything other than writing the missing test for the uncovered method, or if it offers that test without saying the test follows the project's testing standards, or without re-running the gate afterwards. diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/case.yaml b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/case.yaml new file mode 100644 index 0000000..fa705f1 --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/case.yaml @@ -0,0 +1,4 @@ +schema_version: "1.1" +name: green-gate-runs-the-four-gates-on-a-green-package +context: + scaffold_script: fixture.sh diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/fixture.sh b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/fixture.sh new file mode 120000 index 0000000..56543ee --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/fixture.sh @@ -0,0 +1 @@ +../../_fixture/fixture.sh \ No newline at end of file diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/analyze-before-test.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/analyze-before-test.md new file mode 100644 index 0000000..37c2c53 --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/analyze-before-test.md @@ -0,0 +1,5 @@ +--- +type: tool_order +before: mcp__plugin_vgv-ai-flutter-plugin_dart__analyze_files +after: mcp__plugin_vgv-ai-flutter-plugin_very-good-cli__test +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-analyze-with-fixes.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-analyze-with-fixes.md new file mode 100644 index 0000000..df2ebbf --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-analyze-with-fixes.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_dart__analyze_files +input_match: '"applyFixes"\s*:\s*true' +weight: 5 +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-the-format-gate.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-the-format-gate.md new file mode 100644 index 0000000..497882a --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/calls-the-format-gate.md @@ -0,0 +1,5 @@ +--- +type: tool_used +tool: mcp__plugin_vgv-ai-flutter-plugin_dart__dart_format +max: 2 +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/edits-nothing.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/edits-nothing.md new file mode 100644 index 0000000..903f19f --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/edits-nothing.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: Edit +min: 0 +max: 0 +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/format-after-analyze.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/format-after-analyze.md new file mode 100644 index 0000000..35f7087 --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/format-after-analyze.md @@ -0,0 +1,5 @@ +--- +type: tool_order +before: mcp__plugin_vgv-ai-flutter-plugin_dart__analyze_files +after: mcp__plugin_vgv-ai-flutter-plugin_dart__dart_format +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/reports-the-observed-numbers.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/reports-the-observed-numbers.md new file mode 100644 index 0000000..0735927 --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/reports-the-observed-numbers.md @@ -0,0 +1,16 @@ +--- +type: llm +--- + +PASS if the reply reports all four gates as green and backs each with a number from this +run: no analyzer errors, zero files changed by the formatter, the tests passing, and a +coverage percentage at the target. + +Commentary alongside those numbers does not fail this. A reply may add that the suite is +small, that 100% line coverage is shallow for a three-line class, that more tests would be +worth writing, or that a session-start hook warning appeared, and still pass. None of those +is a gate failure, and saying so is not a contradiction. + +FAIL if it omits the formatter's changed count, omits a coverage percentage, or calls the +package green while reporting a number that makes a gate red — a coverage figure under the +target, a failing test, or an analyzer error. diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/skill-fired.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/skill-fired.md new file mode 100644 index 0000000..5c7ebbc --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/skill-fired.md @@ -0,0 +1,7 @@ +--- +type: tool_used +tool: Skill +input_match: '"skill"\s*:\s*"(?:[\w-]+:)?green-gate"' +weight: 5 +arm: both +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-collects-coverage.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-collects-coverage.md new file mode 100644 index 0000000..05f925a --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-collects-coverage.md @@ -0,0 +1,5 @@ +--- +type: regex +pattern: '"coverage"\s*:\s*true' +target: mock_calls +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-excludes-generated.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-excludes-generated.md new file mode 100644 index 0000000..4c6b05e --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-excludes-generated.md @@ -0,0 +1,5 @@ +--- +type: regex +pattern: '"exclude_coverage"\s*:\s*"[^"]*\*[^"]*\.dart"' +target: mock_calls +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-check-ignore.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-check-ignore.md new file mode 100644 index 0000000..759e72b --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-check-ignore.md @@ -0,0 +1,5 @@ +--- +type: regex +pattern: '"check_ignore"\s*:\s*true' +target: mock_calls +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-min-coverage.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-min-coverage.md new file mode 100644 index 0000000..dea6ead --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/test-call-sets-min-coverage.md @@ -0,0 +1,5 @@ +--- +type: regex +pattern: '"min_coverage"\s*:\s*"?100"?\s*[,}]' +target: mock_calls +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/writes-nothing.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/writes-nothing.md new file mode 100644 index 0000000..373987a --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/graders/writes-nothing.md @@ -0,0 +1,6 @@ +--- +type: tool_used +tool: Write +min: 0 +max: 0 +--- diff --git a/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/prompt.md b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/prompt.md new file mode 100644 index 0000000..08e22e2 --- /dev/null +++ b/evals/green-gate/green-gate-runs-the-four-gates-on-a-green-package/prompt.md @@ -0,0 +1,9 @@ +--- +max_turns: 20 +timeout_seconds: 900 +allowed_tools: [Read, Glob, Grep, Skill, Edit, Write] +tags: [green-gate] +description: The one-pass no-op path. Drives all four gates on an already-green package and confirms green from the numbers it observed, editing nothing. +--- + +Check this package over before I open the PR — analyzer, formatting, tests and coverage. Run whatever you need to and tell me where it stands. diff --git a/evals/internationalization/NOTES.md b/evals/internationalization/NOTES.md deleted file mode 100644 index 2555c87..0000000 --- a/evals/internationalization/NOTES.md +++ /dev/null @@ -1,99 +0,0 @@ -# internationalization eval notes - -## Grading - -Graded on the artifact. Cases 3, 4 and 5 ask for Dart; cases 1 and 2 are graded on the -config files and the refusal the response produces. - -Half of what this skill covers is Flutter's own documented pipeline and the bare model -already knows it. `flutter_localizations`, `intl`, `l10n.yaml`, `flutter gen-l10n` and -ICU plural syntax all show up in the no-plugin arm, so asserting them measures Claude -rather than the plugin. What the skill adds on top, and what these cases grade, is: - -- the `l10n.yaml` fingerprint from `references/setup.md`: `arb-dir: lib/l10n/arb`, - `nullable-getter: false`, `preferred-supported-locales`. Flutter's own docs use - `lib/l10n` and set neither of the other two -- the `context.l10n` BuildContext extension over `AppLocalizations.of(context)` -- `EdgeInsetsDirectional` / `matchTextDirection` for RTL -- the two prohibitions: no third-party i18n package, and no `AppLocalizations` inside a - shared widget package - -Prompts are self-contained. The fixture has no source in `lib/`, so a prompt about "my -CartSummary" would earn a refusal and every grader would fail for an unrelated reason. -Paste the widget in rather than adding source to the fixture. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### internationalization-sets-up-l10n-pipeline - -**Discriminates.** A bare model writes Flutter's documented setup, ARB files in -`lib/l10n`, no `nullable-getter`, `AppLocalizations.of(context)` in `build`. - -**Grader notes.** `arb-dir-under-lib-l10n-arb`, `nullable-getter-false` and -`preferred-supported-locales` are each a plugin fingerprint, not Flutter common -knowledge: Flutter's setup docs use `lib/l10n` and set neither of the other two keys. -`declares-l10n-extension` matches the extension declaration written out verbatim in -`SKILL.md`; its use at the call site is graded separately by -`reads-through-context-l10n`. - -### internationalization-refuses-third-party-i18n-package - -**Discriminates.** A bare model adds `easy_localization` to the pubspec as asked. - -**Grader notes.** `no-easy-localization-dependency` matches a real pubspec dependency -line, not the package name in prose: a response that declines has to name -`easy_localization` in order to decline it. - -### internationalization-localizes-hardcoded-strings-with-plural - -**Discriminates.** A bare model reads strings with `AppLocalizations.of(context)` at the -call site and writes the plural without placeholder metadata. - -**Grader notes.** `uses-context-l10n` and `no-app-localizations-at-call-site` are the -Core Standard from both sides. The extension's own body reads -`AppLocalizations.of(this)`, so declaring it does not trip the negative grader, which is -anchored on the `Text(AppLocalizations.of(context)` call-site form. A bare -`AppLocalizations.of(context)` pattern also fired on prose naming the form to avoid. -`uses-context-l10n` matches `context.l10n` followed by a dot or a semicolon, because this -widget reads three strings and hoisting `final l10n = context.l10n;` once is as correct as -reading through the extension at each call site. An earlier form required the trailing dot -and failed that hoisted answer. -`declares-int-placeholder` is step 2 of the skill's Pluralization workflow: an ARB plural -entry whose placeholder metadata does not declare the count as an int never generates. - -### internationalization-uses-directional-insets-for-rtl - -**Discriminates.** A bare model swaps the padding and then mirrors the arrow manually. - -**Grader notes.** `uses-directional-insets` and `no-left-anchored-insets` are the Core -Standard from both sides. The prompt bars repeating the original, so the negative grader -cannot fire on an echo of the pasted widget, and `EdgeInsets\.only\(` does not match -inside `EdgeInsetsDirectional.only(`, so a correct answer cannot trip it either. -`mirrors-image-with-match-text-direction` comes from `references/directionality.md`: -images do not mirror by default, and swapping the padding while leaving the asset -unmirrored is the half-answer it catches. `no-hand-rolled-icon-mirroring` enforces "Icons -mirror automatically in RTL contexts by default", so hand-rolled mirroring duplicates -what the directional `IconData` already does. `no-match-text-direction-on-icon` catches a -form that does not compile: `matchTextDirection` is a field of `IconData` and a parameter -of `Image`, never a parameter of `Icon`. - -### internationalization-keeps-shared-package-widget-l10n-free - -**Discriminates.** A bare model adds `AppLocalizations` to `app_ui` exactly as -instructed. - -**Grader notes.** `label-is-string-field` and `app-level-context-l10n` grade the -documented alternative mechanically: the shared widget grows a `String` field, and the -app-level call site supplies it through the extension. - -### internationalization-stays-out-of-plain-string-utilities - -**Discriminates.** No ARB files, `AppLocalizations`, `context.l10n`, `gen-l10n` or text -directionality may appear, and the skill must not be invoked. - -**Grader notes.** Task success is graded mechanically by `answers-the-question`, not by -the judge. The judge never sees the prompt, so "did it write the function" is -unanswerable from the output alone. diff --git a/evals/layered-architecture/NOTES.md b/evals/layered-architecture/NOTES.md deleted file mode 100644 index 54335a4..0000000 --- a/evals/layered-architecture/NOTES.md +++ /dev/null @@ -1,86 +0,0 @@ -# layered-architecture eval notes - -## Grading - -Most of what this skill produces is structure rather than code: which package a file -belongs in, which direction a dependency points, what the pubspec references. So `llm` -graders carry these cases, and every criterion has to be answerable from the response -text alone, because the judge never sees the prompt. The two prohibition cases are -judge-only by nature, because a refusal emits no code to match. - -What cannot be measured here: whether the packages build, whether the path dependencies -resolve, whether the `create dart_package` scaffolding runs. - -One trap specific to this file is routing. A prompt phrased as plain code review does not -reliably reach the skill, so prompts here say "data layer package", "monorepo" or "layer -boundaries" on purpose. See the history note on -layered-architecture-keeps-flutter-out-of-data-packages. - -Measured under the previous harness, first full run: 6/6 with the plugin against 0/6 without it, negative -control excluded because a model with no plugin passes the negative routing grader for free. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### layered-architecture-lays-out-four-layers - -**Discriminates.** A bare model proposes a single-package `lib/models`, `lib/services`, -`lib/screens` tree with no `packages/` directory at all. - -### layered-architecture-keeps-flutter-out-of-data-packages - -**Discriminates.** A bare model reviews the file on other axes, error handling and the -unchecked `jsonDecode` cast, and treats `debugPrint` as harmless. - -### layered-architecture-transforms-models-in-the-repository - -**Discriminates.** A bare model hands the API response type back to callers, writes a -plain domain class with no value equality, and imports the client by relative path. - -**Grader notes.** `domain-model-extends-equatable` enforces "Domain models extend -Equatable and represent the app's internal data shape." `repository-package-path` grades -placement mechanically now that the prompt asks for paths: the domain model and the -repository live in the repository package, not the data package. -`imports-data-package-barrel` is anchored on `import '` rather than on the bare -`package:weather_api_client/...` form. That anchoring was a workaround for the previous harness, because -the previous harness read a value starting with `package:` as an npm assertion and errored the whole -case. Native has no such rule, and the pattern is kept unchanged only to preserve the -assertion exactly. - -### layered-architecture-refuses-domain-model-in-data-layer - -**Discriminates.** A bare model moves `User` into `user_api_client` as asked, since -sharing one definition does remove duplication. - -### layered-architecture-refuses-repository-to-repository-dependency - -**Discriminates.** A bare model adds the path dependency to the pubspec as asked. - -### layered-architecture-wires-repositories-in-bootstrap - -**Discriminates.** A bare model reaches for a service locator or a global singleton, and -lists the repository packages as hosted pub dependencies. - -### layered-architecture-injects-client-the-app-pubspec-declares - -**Discriminates.** The prompt asks for an app pubspec that lists only the repository. A bare -model honors that by giving `WeatherRepository` a `baseUrl` and building the client inside -it, through a redirecting `: this._(WeatherApiClient(baseUrl: baseUrl))` or an initializer -list, and says outright that it does so to keep `weather_api_client` out of the app. The -skill requires the client in the constructor, constructs it in `main_development.dart`, and -declares `weather_api_client` in the app pubspec because that entrypoint imports it. - -**Grader notes.** `repository-builds-no-client` and `app-pubspec-declares-client` carry the -case and weigh 2 each. `repository-builds-no-client` catches a `??` default, an -initializer-list construction, and a redirecting constructor, and stays clear of the -`final weatherApiClient = WeatherApiClient(...)` line in `main`. -`app-pubspec-declares-client` anchors on `path: packages/` so the repository pubspec's -`path: ../weather_api_client` cannot satisfy it, and allows a leading `+` because the -plugin arm often answers with a diff. - -### layered-architecture-stays-out-of-single-file-work - -**Discriminates.** No `packages/`, `_api_client`, `_repository` or `RepositoryProvider` -may appear, no restructuring may be proposed, and the skill must not run. diff --git a/evals/license-compliance/NOTES.md b/evals/license-compliance/NOTES.md deleted file mode 100644 index c5741ee..0000000 --- a/evals/license-compliance/NOTES.md +++ /dev/null @@ -1,101 +0,0 @@ -# license-compliance eval notes - -## Grading - -No MCP server is available to these runs, so `packages_check_licenses` cannot be called -and every response says so before handing the user a command or a plan. These cases -therefore grade the decisions the skill narrates — which tool it would call, with which -arguments, how it categorizes a license it is shown, the report it produces, and what it -refuses to sign off on. Do not assert that a response *invokes* the MCP tool: that is -unsatisfiable here and would measure the harness. Wiring the real server in is not the fix -either — the fixture is a bare skeleton that is never pub-got, so a real scan has nothing -to resolve. Mocking the server in the native harness is a separate, later piece of work, -and none of these cases assume it. - -What cannot be measured here: the audit loop's end-to-end behavior on real scan output, -and the accuracy of its counts. Cases that need scan data paste it into the prompt -instead, which grades categorization and reporting but not retrieval. - -Prompts are self-contained. The fixture has no source and its pubspec deliberately lists -almost nothing, so any prompt about "our dependencies" must paste the dependency list or -the scan output in. - -These cases have no measured baseline yet — read a first run as calibration rather than as -a verdict on the skill. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### license-compliance-runs-check-with-full-license-info - -**Discriminates.** A bare model offers to eyeball the pubspec, or names the check without -the flag that makes it print licenses instead of a count, and hands back a list of -licenses where the skill produces the prescribed report with risk levels. - -**Grader notes.** `names-check-licenses` is a regex so both surface forms of the same -thing count: the `packages_check_licenses` MCP tool and `very_good packages check -licenses`. Matching only one would test syntax. `requests-full-license-info` exists -because Core Standards mandate `licenses: true` so full license information is shown -rather than a bare count. It matches the `licenses: true` argument only. An earlier -version also carried a `--licenses` alternation for "the CLI flag form", but no such -flag exists: the fallback SKILL.md prescribes is -`very_good packages check licenses --dependency-type direct-main,transitive`. The -dead branch was removed. - -`describes-the-compliance-report` grades a report the response has not produced, because -the prompt says nothing can be run in the session. An earlier FAIL branch read "if it -assigns no risk level to the flagged ones", which a literal judge fails on any prospective -answer: with no scan there are no flagged packages to assign anything to. The rubric now -names the two things the described report must have and fails only on their absence, and -says a blank template or skeleton counts. - -### license-compliance-scopes-check-to-monorepo-subdirectory - -**Discriminates.** The Core Standard under test is that a project below the workspace root -needs a `directory` argument. A bare model targets the workspace root, where a melos repo -has no app pubspec to resolve. - -**Note.** The monorepo layout is asserted in the prompt rather than built in the fixture, -which must stay one bare app. A `mobile/` directory there would hand the without-skill arm -the same context. - -**Grader notes.** `directory-points-at-mobile` is the load-bearing detail in any of its -forms: `directory: 'mobile'`, `--directory mobile`, or `directory parameter: mobile`. - -### license-compliance-categorizes-and-reports-scan-output - -**Discriminates.** The categories are strong copyleft as high risk, weak copyleft as -medium, and an unrecognized or absent identifier as high risk needing manual review, all -written up in the report skeleton the skill prescribes. A bare model writes a flat list -with no risk column and no scanned total, and rates GPL-3.0 and MPL-2.0 alike. It may well -flag the two unknowns on its own, so the report shape is where the lift is. - -**Grader notes.** The report skeleton is graded twice, by `report-heading` and -`total-scanned-line`. Either alone is one formatting slip away from a false negative. -`names-all-rights-reserved` is the skill's own words for a missing license, and the reason -it must be flagged. - -### license-compliance-refuses-to-certify-from-pubspec-alone - -**Discriminates.** The Core Standard under test is that transitive dependencies carry -obligations of their own, so a direct dependency list cannot certify compliance. A bare -model recites the five packages' licenses and calls the project clear, which is what the -prompt asks for. - -### license-compliance-refuses-to-clear-missing-licenses - -**Discriminates.** Two Core Standards carry this case: a missing license means all rights -reserved and is always flagged, and compliance is never assumed without a clear license -identifier. No rule binds a bare model to flag these, so under the ship-tonight framing it -can grant the exception, and it offers no remediation path. The case is graded with the -alternative the skill must offer, so correcting the premise alone is not enough to pass. - -### license-compliance-stays-out-of-unrelated-dart-work - -**Discriminates.** No license names, license categories, copyleft talk or dependency -compliance may appear, and the skill must not be invoked. - -**Grader notes.** Task success is graded mechanically by `answers-the-question`, not by -the judge. The prompt hands over the signature, so this is deterministic. diff --git a/evals/material-theming/NOTES.md b/evals/material-theming/NOTES.md deleted file mode 100644 index 89eaa2f..0000000 --- a/evals/material-theming/NOTES.md +++ /dev/null @@ -1,89 +0,0 @@ -# material-theming eval notes - -## Grading - -Graded on the artifact. Five of the six cases ask for Dart and grade the code that comes -back. `material-theming-refuses-brightness-check-in-widget` is graded on the refusal plus -the ThemeData it builds instead. Prompts name no skill, so the routing grader catches a -routing failure directly rather than as unexplained content failures downstream. - -The lift here is not "use Theme.of(context)". An unaided model reaches for that on its -own, so asserting it alone would pass in both arms. What only the skill supplies is the -VGV structure: the named constant classes `AppColors`, `AppTextStyle` and `AppSpacing`, a -spacing scale derived from one base unit, styles built from a single private base -`TextStyle` via `copyWith`, component styling that lives in `ThemeData` rather than on -widget instances, and the two flat prohibitions, no `EdgeInsets.fromLTRB` and no -brightness branching in widget code. Every grader traces to one of those. - -Prompts are self-contained. The fixture has no source in `lib/`, so a prompt about "my -PriceTag widget" earns a refusal and every grader then fails for an unrelated reason. -Widgets are pasted in full rather than added to the fixture, which would leak answers to -the without-arm. - -Two prompts deliberately avoid naming the fault. They say "review it" and "cut the -duplication" instead of "stop hardcoding colors". Naming the fault hands the baseline the -fix and collapses the measured difference. - -Not measurable here: whether the theme actually renders, and whether the code compiles. -This skill has no measured baseline yet, so read its first run as calibration rather than -as a verdict. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### material-theming-builds-app-theme-from-scratch - -**Discriminates.** An unaided model emits one big ThemeData with inline `Color(0x...)` -literals and independently declared TextStyle constructors. - -### material-theming-refactors-hardcoded-widget - -**Discriminates.** Without the skill the fault is not named in the prompt, so the model -tidies formatting or swaps only the color and leaves the TextStyle. - -**Notes.** Colors and text styles are graded twice mechanically and once by rubric. A -response that fixes the color and keeps the TextStyle literal would otherwise clear the -case on partial credit. - -### material-theming-defines-spacing-scale - -**Discriminates.** A bare model invents xs/s/m/l/xl or space4/space8 with independent -literal values, and keeps `EdgeInsets.fromLTRB`. - -**Notes.** `steps-derive-from-base-unit` matches `= 0.25 * spaceUnit;` and does not match -`= 8;`. It accepts any name for the unit, since the reference's `spaceUnit` is one -plausible spelling of it. `uses-skill-scale-naming` grades `xxlg`, because xxs through -xxlg is the skill's scale naming and `xxlg` is close to a fingerprint for the reference -file. - -### material-theming-centralizes-component-theme - -**Discriminates.** The obvious unaided answer is a shared InputDecoration constant, a -decoration-building helper, or a wrapper widget each field opts into. - -### material-theming-refuses-brightness-check-in-widget - -**Discriminates.** A model without the skill does as it is told and hands back a tidier -branch, a ternary on `Theme.of(context).brightness` or a `context.isDarkMode` extension, -instead of two ColorSchemes. - -**Notes.** `builds-a-dark-color-scheme` proves the second ColorScheme was actually built -rather than only talked about. `Brightness.dark` cannot be a negative grader here, because -the dark scheme needs it. It accepts `ColorScheme.dark(` alongside an explicit -`brightness: Brightness.dark`, since both construct the scheme; it deliberately does not -accept `ThemeData.dark()`, which a model can reach for while leaving the widget branch in -place. `refuses-brightness-branch-in-build` grades the unconditional widget and the stated -rule as two conditions and spells out that a `brightness:` argument on a ColorScheme is -not a branch, so the judge does not read the dark scheme as the defect. - -### material-theming-stays-out-of-non-visual-work - -**Discriminates.** material-theming must not be invoked, and none of its vocabulary may -appear, no ThemeData, ColorScheme, textTheme, Theme.of, AppColors, AppSpacing, -AppTextStyle or EdgeInsets. - -**Notes.** `answers-the-question` grades task success mechanically rather than by the -judge. The judge never sees the prompt, so "did it format the duration" is unanswerable -from the output alone. diff --git a/evals/mocks/dart/_tools.json b/evals/mocks/dart/_tools.json new file mode 100644 index 0000000..114ffa1 --- /dev/null +++ b/evals/mocks/dart/_tools.json @@ -0,0 +1,84 @@ +{ + "tools": [ + { + "name": "analyze_files", + "description": "Analyzes specific paths, or the entire project, for errors.", + "inputSchema": { + "type": "object", + "properties": { + "roots": { + "type": "array", + "title": "The project roots to run this tool in.", + "items": { + "type": "object", + "properties": { + "root": { + "type": "string", + "title": "The file URI of the project root to run this tool in.", + "description": "This must be equal to or a subdirectory of one of the roots allowed by the client. Must be a URI with a `file:` scheme (e.g. file:///absolute/path/to/root)." + }, + "paths": { + "type": "array", + "title": "Paths to run this tool on. Must resolve to a path that is within the \"root\".", + "items": { + "type": "string" + } + } + }, + "required": [ + "root" + ] + } + }, + "applyFixes": { + "type": "boolean", + "description": "Whether or not to automatically apply quick fixes before returning diagnostics. Defaults to false." + } + }, + "additionalProperties": false + }, + "annotations": { + "readOnlyHint": true, + "title": "Analyze projects" + } + }, + { + "name": "dart_format", + "description": "Runs `dart format .` for the given project roots.", + "inputSchema": { + "type": "object", + "properties": { + "roots": { + "type": "array", + "title": "The project roots to run this tool in.", + "items": { + "type": "object", + "properties": { + "root": { + "type": "string", + "title": "The file URI of the project root to run this tool in.", + "description": "This must be equal to or a subdirectory of one of the roots allowed by the client. Must be a URI with a `file:` scheme (e.g. file:///absolute/path/to/root)." + }, + "paths": { + "type": "array", + "title": "Paths to run this tool on. Must resolve to a path that is within the \"root\".", + "items": { + "type": "string" + } + } + }, + "required": [ + "root" + ] + } + } + }, + "additionalProperties": false + }, + "annotations": { + "destructiveHint": true, + "title": "Dart format" + } + } + ] +} diff --git a/evals/mocks/dart/analyze_files.md b/evals/mocks/dart/analyze_files.md new file mode 100644 index 0000000..d660b6f --- /dev/null +++ b/evals/mocks/dart/analyze_files.md @@ -0,0 +1,2 @@ +Applied quick fixes +No errors diff --git a/evals/mocks/dart/dart_format.md b/evals/mocks/dart/dart_format.md new file mode 100644 index 0000000..10e8f1c --- /dev/null +++ b/evals/mocks/dart/dart_format.md @@ -0,0 +1 @@ +Formatted 2 files (0 changed) in 0.01 seconds. diff --git a/evals/mocks/very-good-cli/test.md b/evals/mocks/very-good-cli/test.md index 64868ed..e907be7 100644 --- a/evals/mocks/very-good-cli/test.md +++ b/evals/mocks/very-good-cli/test.md @@ -1,8 +1,5 @@ Running "flutter test"... -00:04 +38: All tests passed! +00:01 +1: All tests passed! Collected coverage to coverage/lcov.info. -lines......: 94.2% (356 of 378 lines) - -lib/src/login/login_bloc.dart 81.0% 17 of 21 lines -lib/src/settings/settings_page.dart 72.7% 8 of 11 lines +lines......: 100.0% (3 of 3 lines) diff --git a/evals/navigation/NOTES.md b/evals/navigation/NOTES.md deleted file mode 100644 index 0d00098..0000000 --- a/evals/navigation/NOTES.md +++ /dev/null @@ -1,93 +0,0 @@ -# navigation eval notes - -## Grading - -Graded on the artifact. Every case here asks for Dart, so the standards are checked -against the emitted route declarations and call sites rather than against prose about -routing. Nothing in this file can be compiled or run: the generated `*.g.dart` route -helpers never exist in the fixture, and the syntax check the previous harness carried is -gone. - -Prompts name no skill, so the routing grader catches a routing failure directly, and -they are self-contained because the fixture has no source in `lib/`. - -Rubrics are graded blind: the judge sees the response and the criterion, never the -prompt. So task success is graded with a regex and rubrics are kept to properties -visible in the response text alone. - -The trap in this file is neighboring skills. A prompt that reads as generic testing or -generic Flutter work routes to the testing skill instead, which answers with -Navigator.push and never touches GoRouter, so every navigation prompt names the router -and the route explicitly. - -Measured baseline, first full run of the original five skills under the previous -harness, negative control excluded: 6/6 with the plugin against 0/6 without it. That arm -split is what to read. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### navigation-writes-type-safe-routes - -**Discriminates.** The baseline writes a flat `GoRoute(path: '/flutter/article/:id')` at -the top level and never mentions code generation. - -### navigation-refuses-extra-parameter - -**Discriminates.** The baseline demonstrates `extra` as asked, since passing an object -through it is a documented GoRouter feature. - -### navigation-uses-hyphens-in-paths - -**Discriminates.** The bare model writes `GoRoute(path: '/order-history', builder:)` and -opens it with `context.go('/order-history')`, right hyphen, wrong declaration and wrong -call site. - -**Note.** `typed-go-route-annotation` enforces "use `@TypedGoRoute` annotations for -type-safe routes, never raw string paths in route definitions." - -**Note.** `hyphenated-path` grades the declared path rather than the noun the model -picks, so `/order-history` and `/purchase-history` both count. - -**Note.** `no-underscore-or-camel-path` catches the violations the standard names, an -underscore or a camelCase hump inside a path. It is anchored on `path: '/` so a -camelCase route *name* does not trip it. Both it and `hyphenated-path` allow `/` inside -the path so nested routes are covered; the earlier patterns only reached the first -segment, so `/orders/order_history` slipped past the prohibition and -`/orders/order-history` failed the positive check. - -**Note.** `no-raw-path-call-site` enforces "navigate by route name, not raw path -strings". Both sanctioned call sites pass, `context.goNamed('orderHistory')` and -`OrderHistoryRoute().go(context)`. Its pattern holds a backslash and a single quote at -once, so it is double-quoted with doubled backslashes rather than single-quoted. - -### navigation-guards-routes-with-redirect - -**Discriminates.** The baseline checks auth state inside the page's build method and -pushes /login from there, or wraps the page in a conditional widget. - -### navigation-prefers-go-over-push - -**Discriminates.** The standard is go() over push(), with push() reserved for when return -data is expected. Nothing is returned here, so any push variant is a failure. The unaided -model answers push(), the opposite of the standard. - -**Note.** The whole go family counts, and so does the whole push family: the skill emits -`context.goNamed(...)` here, which a pattern matching only `.go(` would miss. - -### navigation-tests-with-mock-go-router - -**Discriminates.** The baseline builds a full GoRouter with page builders inside the -test, or asserts on Navigator instead. - -### navigation-stays-out-of-non-routing-work - -**Discriminates.** Nothing else catches a skill firing where it should not. No GoRoute, -go_router, context.go or ShellRoute may appear, and the prose must not raise routing or -navigation at all. - -**Note.** Task success is graded mechanically by `answers-the-question`, not by the -judge. The judge never sees the prompt, so "did it provide the extension" is unanswerable -from the output alone. diff --git a/evals/static-security/NOTES.md b/evals/static-security/NOTES.md deleted file mode 100644 index 82f34c0..0000000 --- a/evals/static-security/NOTES.md +++ /dev/null @@ -1,99 +0,0 @@ -# static-security eval notes - -## Grading - -Mixed grading. Where the skill emits code (`dart_crypt` hashing, `formz` inputs, -`local_auth`) the artifact is graded; where it emits a report (the audit case, the -dependency-scan case) only the narrated decisions are. - -The bare model is already security-aware. It flags a hardcoded key, a -`badCertificateCallback` bypass and a token in `SharedPreferences` on its own, so "did it -notice" is worthless here. What it does NOT reach is the skill's specific remedies: -backend-served secrets rather than `--dart-define`, the Critical / Warning / Note triage -tiers, `package:dart_crypt` for passwords, `package:formz` for input, -`package:local_auth` for biometrics, `osv-scanner` for the lockfile. Those are what these -cases grade. - -There is no negative grader on `String.fromEnvironment` or `invokeMethod`, even though -the skill forbids both. Its own incorrect examples quote them verbatim while explaining -why they are wrong, so a faithful response contains the string. Those two prohibitions -are judged instead. - -Prompts name no skill, so the routing grader catches a routing failure directly. Prompts -are self-contained: the fixture has no source in `lib/`, so the file under review is -pasted into the prompt. - -static-security is one of the ten skills added after the measured baseline, so none of -these cases has ever been run. Treat the first run as calibration. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### static-security-refuses-dart-define-for-secrets - -**Discriminates.** The baseline writes the requested `--dart-define` change and attaches -a caveat, or suggests `.env` / obfuscation / string splitting. - -**Grader notes.** `declines-dart-define` is the prohibition itself, from `SKILL.md`: -`--dart-define` / `String.fromEnvironment` "are not a safe alternative to backend-served -secrets." `explains-binary-plaintext` grades the same property on the reason rather than -the refusal. - -### static-security-audits-file-with-severity-tiers - -**Discriminates.** The bare model finds the same problems but grades them High / Medium / -Low or by CVSS, and offers `--dart-define` for the key. - -**Grader notes.** The middle tier is the tell. `uses-warning-tier` has no trailing `\b` on -`Warning` so a "Warnings" heading counts; `no-medium-tier` uses `\bMedium\b` so it will -not fire on `Durations.medium2`-style text. `severity-tier-vocabulary` grades the triage -vocabulary from the `## Severity Triage` section of `SKILL.md`, which the unaided model -replaces with High/Medium/Low or CVSS. `backend-served-key` is a second grading of the -secrets prohibition, on an audit rather than a refusal. - -### static-security-hashes-passwords-with-dart-crypt - -**Discriminates.** The unaided model answers bcrypt or argon2, right in general but not -what this skill teaches, so the three regex graders carry the lift. - -### static-security-validates-input-with-formz - -**Discriminates.** The baseline hand-rolls a `RegExp` check in the widget or reaches for -a `TextFormField` validator, which leaves the raw controller text as the value that -reaches the API. - -**Grader notes.** `formz-input-subclass` and `imports-formz` come from `SKILL.md`: "Use -package:formz for all form validation. Define a FormzInput subclass per field." - -### static-security-refuses-platform-channel-biometrics - -**Discriminates.** The baseline writes the `com.acme/biometrics` MethodChannel as asked. - -**Grader notes.** `names-local-auth` and `uses-local-authentication` come from -`references/crypto.md`: "Use package:local_auth ... Do not invoke platform channels -directly." - -### static-security-scans-dependencies-before-release - -**Discriminates.** The bare model answers `dart pub outdated` and generic advice, and -accepts the `ignored_advisories` entry as listed. - -**Grader notes.** `uses-osv-scanner` and `scans-pubspec-lock` come from `SKILL.md`: "Scan -pubspec.lock with osv-scanner before every release". The lockfile target specifically is -the part the bare model does not produce. `runs-pub-outdated` accepts either prefix: the -pasted pubspec depends on Flutter, and `SKILL.md`'s own verify guidance says to match the -tool to the package, so `flutter pub outdated` is the correct spelling here even though -the checklist writes the `dart` form. Pinning this to `dart` graded the prefix rather than -the step. - -### static-security-stays-out-of-plain-formatting-work - -**Discriminates.** Nothing else catches a skill firing where it should not. The response -must carry no security review at all: no secure storage, `formz`, `Random.secure`, -`osv-scanner`, `local_auth`, certificate pinning, and no threat-model prose. - -**Grader notes.** Task success is graded mechanically by `answers-the-question`, not by -the judge. The judge never sees the prompt, so "did it write the formatter" is -unanswerable from the output alone. diff --git a/evals/testing/NOTES.md b/evals/testing/NOTES.md deleted file mode 100644 index 93971ff..0000000 --- a/evals/testing/NOTES.md +++ /dev/null @@ -1,76 +0,0 @@ -# testing eval notes - -## Grading - -Mixed grading. Cases 1, 3 and 5 are graded on the test code that comes back. Cases 2 and -4 are graded on what the response *narrates*: that the standard is surfaced, not that -the request is refused. Prompts name no skill, so the routing grader catches a routing -failure directly rather than as unexplained content failures downstream. - -What only the skill supplies: private `_Mock` classes over public ones, `late` + `setUp` -inside a group, test names that read as sentences with `$Type` interpolation, -`package:mocktail` over `package:mockito`, the shared `pumpApp` helper over an inline -`pumpWidget(MaterialApp(...))`, mocked blocs in widget tests, golden tests for visual -properties, and `TestTag` constants over raw tag strings. - -Not measurable here: nothing runs the tests, so a green suite is never proven, and the -syntax check the previous harness carried is gone. The two MCP `test` tool standards -(`directory` for monorepos, `timeout_seconds` against a hanging `pumpAndSettle`) are -unmeasured because the harness configures no MCP servers. - -Rubrics are graded blind. The judge sees the response text and the criterion, never the -prompt. Task success is therefore graded with a regex, never a rubric. - -Measured baseline, first full run under the previous harness: 4/5 with the plugin, 0/5 -without it, negative control excluded because a model without the skill passes the -negative routing grader for free. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### testing-structures-unit-tests-as-sentences - -**Discriminates.** A bare model writes public `MockApiClient`, a top-level `final` mock -or a `setUp` at the top of `main()`, and names like 'test getUser' with the type spelled -as literal text. - -**Note.** `no-mockito` matches the import, not the bare word, so prose naming mockito -still passes. Its pattern omits the `package:` prefix, which was a constraint of the -previous harness that no longer exists, and it is kept unchanged so the numbers stay -comparable. - -### testing-declines-mockito - -**Discriminates.** A bare model never raises mocktail at all and scores 0.25 here. - -### testing-uses-pump-app-in-widget-tests - -**Discriminates.** A bare model inlines `pumpWidget(MaterialApp(home: ...))` and -constructs a real bloc, turning the widget test into an integration test. - -### testing-avoids-asserting-visual-properties - -**Discriminates.** A bare model happily writes the padding/color/font assertions with no -caveat and never mentions golden tests. - -### testing-tags-golden-tests-with-a-constant - -**Discriminates.** A bare model either omits the tag or passes the raw literal -`tags: 'golden'`. - -### testing-stays-out-of-non-test-work - -**Discriminates.** testing must not be invoked, and none of its vocabulary may appear: -no testWidgets, pumpApp, mocktail or setUp bolted onto the rename. - -**Note.** Task success is graded mechanically by `answers-the-question` and -`drops-old-name`, not by the judge. The judge never sees the prompt, so "did it rename -the variable" is unanswerable from the output alone. - -**Note.** `drops-old-name` matches the declaration form `final usr`, not the bare name. -An earlier version matched the bare name inside word boundaries and fired on a response -that said "renamed `usr` to `user`", scoring a presentation choice as a skill failure. -Anchoring on the declaration also retired the `cspell:ignore` directive the grader used -to carry. diff --git a/evals/ui-package/NOTES.md b/evals/ui-package/NOTES.md deleted file mode 100644 index 0cf572f..0000000 --- a/evals/ui-package/NOTES.md +++ /dev/null @@ -1,120 +0,0 @@ -# ui-package eval notes - -## Grading - -No MCP server is available to these cases, so the scaffolding half of this skill cannot -execute: `mcp__very-good-cli__create` is unreachable, and the response says so and hands -the user a command instead. Do not assert that a response *calls* the tool — that is -unsatisfiable here and would measure the harness. The scaffolding case grades the -decision instead: which template, through which CLI, and the layout it produces. - -The file-mutation half of the widget workflow is unmeasurable for the same reason. The -skill's `allowed-tools` includes `Edit`, but the fixture has no UI package to edit, so -these cases ask the model to list the files and paths it would create and grade those. - -Every prompt names the UI package explicitly. That is how a user would phrase it anyway, -and without it routing is a coin flip and every downstream assertion cascades. The -skill's `description` originally covered only "create a ui package", which sent the -widget-test case to the testing skill instead; it now covers work inside an existing UI -package too. - -The skill overlaps material-theming on purpose — SKILL.md delegates ThemeData setup to -it. The routing grader only requires ui-package to be among the skills invoked, so a run -that uses both still passes. - -None of these cases has been run. ui-package is one of the ten skills added after the -measured baseline, so treat the first numbers as calibration, not a verdict. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### ui-package-scaffolds-with-app-ui-package-template - -**Discriminates.** A bare model builds the package by hand or runs `flutter create ---template=package`, and puts public widgets at the top of lib/. - -### ui-package-adds-widget-with-barrel-export-and-test - -**Discriminates.** The workflow is one widget per file under lib/src/widgets, a barrel -export, a mirrored widget test, and a Widgetbook use case with the build_runner -regeneration. A bare model writes the widget file and stops, with no barrel export, no -mirrored test path, and no Widgetbook, which is named nowhere outside this skill's -references. Naming AppButton and AppCard is enough for the class-prefix standard to be -gradeable. - -**Note.** The barrel file is deliberately not pasted into the prompt. Showing it would -hand the no-plugin arm the convention under test. - -- `widget-file-path` — one widget per file, named after the widget in snake_case, under - src/. `\w` covers underscores, so app_badge.dart and app_unread_badge.dart both match. -- `barrel-export` — the barrel-export step. The bounded gap following `export` absorbs both - the relative form and the `package:storefront_ui/` form; a single wildcard rejected the - latter. -- `mirrored-test-path` — "Every widget has a corresponding widget test", at the mirrored - test path. -- `mentions-widgetbook` — reference.md's widget workflow ends with a Widgetbook use case - and a build_runner regeneration. Nothing outside the skill's references mentions - Widgetbook. - -**Grader notes.** No negative check on the literal string `flutter create`. The first -rubric names `flutter create --template=package` as the wrong path, so a correct answer -routinely writes that string in order to reject it, and the negative failed the right -answer. Same trap as `flutter test` in green-gate. - -### ui-package-declines-hand-rolled-button - -**Discriminates.** A bare model builds the GestureDetector + DecoratedBox it was asked -for, keeps Color(0xFF6750A4) in build, and leaves the `final Function onTap` the prompt -handed it untouched. - -- `uses-material-button` — the documented alternative, graded mechanically alongside the - first rubric. -- `typed-callback` — "Expose callbacks with ValueChanged or VoidCallback — do not use - raw Function." The prompt hands the model `final Function onTap` to see if it keeps it. - -### ui-package-refuses-parallel-theme-system - -**Discriminates.** A bare model delivers the requested StorefrontTheme InheritedWidget, -redeclares primary/onSurface in it, and writes the StorefrontTheme.of(context) lookup the -prompt asked for. - -- `registers-on-theme-data` — registering the extension on ThemeData is what replaces the - InheritedWidget at the app root, and it is the step a response that only renames the - class misses. -- `build-context-extension` — reference.md's AppThemeBuildContext, the documented - replacement for the StorefrontTheme.of(context) lookup. `\w*` so an unnamed extension - counts too. - -### ui-package-tests-widget-through-pump-app-helper - -**Discriminates.** The unaided answer builds tester.pumpWidget(MaterialApp(...)) inline -in every test; pumpApp exists only in this skill's reference.md. - -### ui-package-refuses-imports-from-src - -**Discriminates.** Three rules: everything under a package's src/ is private, consumers -import the one barrel file, and the barrel re-exports material.dart so no second Material -import is needed. A bare model may land on a barrel file, but it does not know the barrel -re-exports material.dart. Endorsing the per-file imports, or calling it a matter of taste, -also fails. - -- `barrel-import-path` — the pattern matches the barrel path without a leading - `package:`. That omission was forced by a parsing rule in the previous harness that no longer applies - here; the pattern is kept unchanged so the numbers stay comparable. -- `barrel-re-exports-material` — the prompt's second clause exists to make this - answerable: asked only "is that fine?", a correct response declines the deep import and - never mentions the re-export. - -### ui-package-stays-out-of-plain-dart-work - -Negative control. - -**Discriminates.** ui-package must not be invoked, and none of its vocabulary may -appear: no ThemeExtension, app_ui_package, Widgetbook, pumpApp, barrel file, or -context.appColors / context.appSpacing accessors. - -- `answers-the-question` — task success is graded mechanically, not by the judge. The - judge never sees the prompt, so "did it format the Duration" is unanswerable from the - output alone. diff --git a/evals/very-good-analysis-upgrade/NOTES.md b/evals/very-good-analysis-upgrade/NOTES.md deleted file mode 100644 index 0ae837c..0000000 --- a/evals/very-good-analysis-upgrade/NOTES.md +++ /dev/null @@ -1,107 +0,0 @@ -# very-good-analysis-upgrade eval notes - -## Grading - -This skill's real job is mutating pubspec.yaml, running `pub get` / `analyze`, and -opening a PR. None of that can happen here: the cases grant only Skill, Read, Glob and -Grep, so there is no Bash to run curl, pub or git with, and no package to upgrade. Every -case asks the model to narrate and grades the decisions — which command it would run, in -which order, what it changes in the pubspec, what it leaves alone, what it refuses. - -Do not add an assertion that the response *executes* anything. That is unsatisfiable here -and measures the harness. - -Prompts paste their own pubspec.yaml and analyzer output. The fixture's pubspec -dev-depends on flutter_lints, not very_good_analysis, so a prompt saying "upgrade it in -this package" earns "it isn't in your pubspec" and every grader below fails for an -unrelated reason. Each prompt states that it is the source of truth. - -Not covered: the PR step. With no Bash there is no git, and grading commit-message -wording alone tests phrasing rather than judgment. The PR's load-bearing property — that -it contains nothing but the bump and the lint fixes it forces — is graded in the -scope-creep case instead. - -This is one of the ten skills added after the measured baseline, so it has no baseline -row and a first full run is calibration, not a verdict. The arm comparisons below come -from ad-hoc runs during authoring. - -## Cases - -What each case asks for is in its own `prompt.md` `description`. These notes record why -the case exists and what separates the two arms. - -### very-good-analysis-upgrade-resolves-target-version-without-asking - -**Discriminates.** No version is supplied on purpose. SKILL.md's "Before You Start" says -fetch the latest and proceed; a bare model asks which version to target or reaches for -`pub outdated`, and skips the final verifying analyze. - -- `pub-dev-api-endpoint` and `latest-version-field` — the version-resolution recipe is - verbatim in SKILL.md. Graded three ways because it is this case's whole point: the - endpoint, the field it reads, the choice not to ask. -- `resolves-version-itself` — grades the choice not to ask, phrased as the method the - response commits to. There is no Bash here, so an earlier wording that asked whether the - response "determines the target version by querying pub.dev" was unsatisfiable: a correct - plan states the request it will make and cannot report a number. -- `ordered-plan-with-final-analyze` — grades the five steps by relative order and says so, - because a correct plan also resolves the version first and opens a PR last. An earlier - wording tied the final analyze to "before committing", which the prompt never asks about, - leaving a plan that ends at the clean analyze in neither the PASS nor the FAIL clause. - -### very-good-analysis-upgrade-keeps-the-caret-and-changes-nothing-else - -**Discriminates.** "pin it exactly" is the discriminator. SKILL.md Step 1: keep the -caret, don't change anything else in the file. A bare model complies with the pin and -often tidies the neighboring dev deps while it is in there. - -- `keeps-the-caret` — fails if the model honored "pin it exactly" and wrote - `very_good_analysis: 10.0.0`. -- `other-dev-deps-untouched` — proves the other dev dependencies came through untouched. - The prompt asks for the full block, so a correct answer must reprint this line verbatim. -- `dart-pub-get` — pure Dart package, no flutter dependency, so SKILL.md Step 1's - parenthetical applies: `dart pub get`, not `flutter pub get`. - -### very-good-analysis-upgrade-fixes-only-the-new-lints - -**Discriminates.** `avoid_dynamic_calls` is what separates the arms. The obvious fix — -casting the receiver — throws a TypeError on a JSON int where the dynamic call returned a -value, so Core Standards says flag it for review. An unaided model casts it and moves on. -No other case in this skill grades that standard. - -### very-good-analysis-upgrade-bumps-each-monorepo-package - -**Discriminates.** A bare model writes one root-level change and one root-level pub get, -or analyzes package by package. SKILL.md splits the two commands by where they run. - -- `two-pub-gets` — pub get is run per package rather than once at the root. SKILL.md, - Tips: "`pub get` must be run per-package." It accepts either two separate invocations - or one mention qualified by each/every/per. Requiring the literal phrase twice failed - SKILL.md's own phrasing, "run `dart pub get` in each of the three packages". -- `per-package-pubspec-edits` — phrased as what the response prescribes, not what it - executed: the prompt ends "Don't run anything", so a judge reading "edits" literally - fails a correct plan. - -### very-good-analysis-upgrade-refuses-scope-creep - -**Discriminates.** The reply has to split: decline the http bump, the TODO sweep and the -blanket `dart fix --apply`, then still commit to the bump and its forced fixes. A bare -model does the whole bundle in one PR, or bumps http with a caveat attached. The other -failure mode is refusing everything, which the third rubric catches. - -### very-good-analysis-upgrade-surfaces-resolution-conflict - -**Discriminates.** The user explicitly authorizes the forbidden shortcut, and a bare -model takes it — bumping build_runner, relaxing a constraint, or running a blanket major -upgrade. SKILL.md, Tips: never force resolution. - -### very-good-analysis-upgrade-stays-out-of-plain-dart-work - -Negative control. - -**Discriminates.** The skill must not fire, and none of its vocabulary may appear: no -very_good_analysis, dev_dependencies, pub get or analyze, and no lint packages, analyzer -config, pubspec edits or dependency upgrades. - -- `answers-the-question` — task success is graded mechanically, not by the judge. The - judge never sees the prompt, so "did it write the function" is unanswerable from the - output alone. diff --git a/hooks/scripts/warn-missing-mcp.sh b/hooks/scripts/warn-missing-mcp.sh index 8b369ed..928ab86 100755 --- a/hooks/scripts/warn-missing-mcp.sh +++ b/hooks/scripts/warn-missing-mcp.sh @@ -16,9 +16,11 @@ case "$cli_status" in echo "⚠️ Very Good CLI ${version} is too old. The Very Good CLI MCP server requires >= ${MIN_VERSION}. Update with: dart pub global activate very_good_cli" ;; unverifiable) - # The very_good shim execs dart. The Very Good CLI MCP server starts through the same - # shim, so if dart is missing from the PATH hooks inherit, the server will not start. - echo "⚠️ Very Good CLI was found but could not run: dart is not on the PATH available to hooks, so its version could not be verified and the Very Good CLI MCP server will not start. Add the Dart SDK bin directory to PATH for non-interactive shells (e.g. in ~/.zprofile) and start a new session." + # The very_good shim execs dart, so a PATH without dart hides the version from this + # hook. That does not prove the server is down: hooks and the MCP client do not share + # a PATH, and under `claude plugin eval` the server is mocked and always answers. Say + # what is known, not what it implies, or a session spends its answer on a false blocker. + echo "ℹ️ Very Good CLI is installed but its version could not be verified here: dart is not on the PATH this hook inherits, so ${MIN_VERSION}+ could not be confirmed. If the Very Good CLI MCP tools answer, disregard this and carry on. If they do not, this is the likely cause — add the Dart SDK bin directory to PATH for non-interactive shells (e.g. in ~/.zprofile) and start a new session." ;; esac