Improve test gap correctness and cost - #1080
Conversation
Require complete public-outcome inventories, suppress inert and unobservable mutation candidates, and bound focused execution and output. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Skill Coverage Report
Uncovered:
|
There was a problem hiding this comment.
Pull request overview
This PR refines the dotnet-test/test-gap-analysis skill guidance to improve correctness and reduce cost by requiring a caller-visible outcome inventory before mutation selection, tightening observability/equivalence rules (including auth-denial outcomes), and limiting execution to the smallest decisive set while still reporting all high-risk findings.
Changes:
- Replaced the prior workflow with a decision flow that inventories caller-visible outcomes (including authorization denials) before selecting/execing mutations.
- Added stricter “public counterfactual” requirements (original vs mutant observations) and post-green-run rechecks to suppress inert/equivalent survivors.
- Updated output contract and exhaustive-audit reference to avoid repeating findings in prose and to count only executed/definitively classified candidates.
Show a summary per file
| File | Description |
|---|---|
| plugins/dotnet-test/skills/test-gap-analysis/SKILL.md | Reworked the core decision flow, verification rules, and output contract to prioritize caller-visible outcomes and reduce unnecessary execution. |
| plugins/dotnet-test/skills/test-gap-analysis/references/mutation-catalog.md | Updated the exhaustive audit procedure and equivalence filters to require public observability and post-run counterfactual checks. |
Review details
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- Files reviewed: 2/2 changed files
- Comments generated: 1
- Review effort level: Lite
|
👋 @Evangelink — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
|
❌ Evaluation did not complete successfully (the evaluate job reported 1 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Reserve mutation execution for explicit verification, require complete authorization outcome ledgers, and calibrate exact-value and suite-strength reporting from the latest Luna evidence. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
📊 Skill Evaluation Results2 model/skill results across 1 skill and 2 models — ✅ 0 improved, ➖ 2 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when the aggregate net win across distinct-stimulus votes is at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
➖ Not proven improved — test-gap-analysis (claude-sonnet-4.6)Why: Net win +33.3% (5W/2T/2L over 9 stimulus vote(s), sign test p=0.227), mean preference +20.0% across 18 paired run(s) — not credible (sign test p=0.227 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 5W/2T/2L; d=7; p=0.227; net +33.3% Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 18 paired runs (10W/4T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — test-gap-analysis (gpt-5.6-luna)Why: Net win +11.1% (4W/2T/3L over 9 stimulus vote(s), sign test p=0.500), mean preference +1.1% across 18 paired run(s) — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 4W/2T/3L; d=7; p=0.500; net +11.1% Overfit: Low (score 0.13) Repeated-run reliability (not used by the gate): 18 paired runs (7W/6T/5L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Keep focused analyses within the named risk, require concrete edits to existing code, and cover complete money-math composition without future-only or trivial state findings. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Review details
Suppressed comments (2)
Previously missed (1) — in code that hasn't changed since the last review.
plugins/dotnet-test/skills/test-gap-analysis/SKILL.md:34
- The “Explicit survivor verification” row is phrased circularly (“execute … for outcomes reported as Survived or Killed”). At the scope-setting stage nothing is “reported” yet, so this can be misread as permission to skip verification for outcomes that should be classified, leading to under-verification.
| Explicit survivor verification | Inventory all requested outcomes; execute one representative observable candidate for each distinct high-risk outcome reported as **Survived** or **Killed** |
plugins/dotnet-test/skills/test-gap-analysis/SKILL.md:103
- This bullet says “exception message/parameter metadata”, but earlier the doc talks about “exception type/parameter”, which can be read as ParamName. Clarify that it’s specifically exception message/ParamName details that are excluded unless established by contract/tests.
- overflow behavior, exception message/parameter metadata, or other semantics
- Files reviewed: 2/2 changed files
- Comments generated: 1
- Review effort level: Lite
|
/evaluate e2f4490 |
There was a problem hiding this comment.
Copilot review overview
Review tier: Lite
Findings: 1
Pre-existing issues (1)
| Severity | Finding |
|---|---|
tests/dotnet-test/test-gap-analysis/eval.yaml — The rubric wording says “narrowing attempt >= 3 to attempt >= 2”, but that change widens the… View comment |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3a626035-c70f-4ba6-97f0-e076b8a5fd1e
|
/evaluate 1c81a5b |
There was a problem hiding this comment.
Copilot review overview
Review tier: Lite
Findings: None
Issues resolved since last review (1)
| Severity | Finding |
|---|---|
tests/dotnet-test/test-gap-analysis/eval.yaml — The rubric wording says “narrowing attempt >= 3 to attempt >= 2”, but that change widens the… View resolved comment |
📊 Skill Evaluation Results2 model/skill results across 1 skill and 2 models — ✅ 1 improved, ➖ 0 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — test-gap-analysis (gpt-5.6-luna)Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +15.6% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Low (score 0.16) Repeated-run reliability (not used by the gate): 18 paired runs (10W/5T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — test-gap-analysis (claude-sonnet-4.6)Why: Net win +85.7% (6W/1T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +35.6% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 2 dormancy excluded Overfit: Moderate (score 0.34) Repeated-run reliability (not used by the gate): 18 paired runs (12W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3a626035-c70f-4ba6-97f0-e076b8a5fd1e
|
/evaluate eb6c250 |
📊 Skill Evaluation Results2 model/skill results across 1 skill and 2 models — ✅ 1 improved, ➖ 0 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — test-gap-analysis (claude-sonnet-4.6)Why: Net win +85.7% (6W/1T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +37.8% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 18 paired runs (15W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3a626035-c70f-4ba6-97f0-e076b8a5fd1e
|
/evaluate 879491b |
There was a problem hiding this comment.
Copilot review overview
Review tier: Lite
Findings: None
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
plugins/dotnet-test/skills/test-gap-analysis/SKILL.md:294
- The continuation line here is indented enough to be parsed as an indented code block in CommonMark, so the bold markup will render literally instead of as emphasis. Merge it into the checklist bullet (or reduce indentation) so it renders as intended.
- [ ] Every outcome labeled **Survived** was executed; unexecuted candidates use
**Candidate survivor (unverified)**
📊 Skill Evaluation Results2 model/skill results across 1 skill and 2 models — ✅ 0 improved, ➖ 2 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
➖ Not proven improved — test-gap-analysis (claude-sonnet-4.6)Why: Net win +71.4% (6W/0T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +30.0% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 6W/0T/1L; d=7; p=0.063; net +71.4%; 2 dormancy excluded Overfit: Moderate (score 0.30) Repeated-run reliability (not used by the gate): 18 paired runs (14W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — test-gap-analysis (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +8.9% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 2 dormancy excluded Overfit: Low (score 0.15) Repeated-run reliability (not used by the gate): 18 paired runs (7W/8T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
📊 Skill Evaluation Results2 model/skill results across 1 skill and 2 models — ✅ 1 improved, ➖ 1 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
➖ Not proven improved — test-gap-analysis (claude-sonnet-4.6)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +25.6% across 18 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 2 dormancy excluded Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 18 paired runs (10W/5T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
|
❌ Evaluation did not complete successfully (the evaluate job reported 1 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |

Summary
Improve
dotnet-test/test-gap-analysiscorrectness and cost without changing its nine stimuli or two-run count:Evaluation diagnosis
Authoritative run 32954117134 evaluated the prior payload at
e0bd807007664bc7839e59440e76318b73c4f584. Both results were conclusiveVALID_NO_CHANGEmeasurements with zero errors, timeouts, or unmatched trajectories. Isolated activation was correct for every in-scope stimulus, and aggregate plugin activity was present.claude-sonnet-4.6gpt-5.6-lunaBoth models lost
Find logic and null-check mutation gaps in access control code:The full transcripts also showed skilled responses admitting green-but-equivalent mutants: removing a null/empty guard that falls through to the same role result, and removing internal zero-entry cleanup even though every public API remains identical. The judge summarized the latter as:
The skill spent 1.9-2.1x baseline token estimates in isolation and 2.74x in the plugin arm. The fixed mutation budget encouraged execution before complete outcome enumeration, then admitted unobservable mutants and produced narrower findings than baseline.
First exact-head reevaluation
Run 33063839973 evaluated commit
5dcc88194b5a43f90ad9759e4273614b4ad3ceba.VALID_NO_CHANGE: 3W/2T/4L stimulus votes and 5W/5T/8L repeated runs, with zero errors or unmatched trajectories. Isolated and plugin activation were both present for access control.Run vally evaluationshit the workflow's 155-minute timeout. This is an invalid reliability result, not content evidence.Second exact-head reevaluation
Run 33079053602 evaluated commit
f1ef2a1df7f7f61744cb1f61b3d06245b6f95583. Both model results were valid and conclusive with zero errors, timeouts, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaThis verifies the core correction: neither model lost access control, and the Luna cost delta fell 59% from +114,012. Luna's generic boundary/access-control/billing ratios fell from 3.70/3.64/3.32x baseline to 1.41/1.62/1.88x.
The remaining losses exposed two narrow, reproducible output defects rather than a reason to add stimuli:
Mixed;Balance/trivialIsPaidstate and one omitted the tax-base-composition mutation.The final iteration addresses those exact general failure classes while leaving the eval and execution policy unchanged.
Final semantic exact-head reevaluation
Run 33082991359 evaluated commit
26efa091e9b99c96f1cd54e0a302d93e9f604fb6. Both model results were valid and conclusive with zero errors, timeouts, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaThe original access-control failure is fully corrected: both models won both repeated trials. Luna's mean token delta is 48% below the first exact-head result (+59,571 versus +114,012), and its generic boundary/access-control ratios are 1.25x/1.17x baseline instead of 3.70x/3.64x. The verdicts remain
VALID_NO_CHANGE, not regressions: Sonnet's sign test is p=.227, while Luna's 6W/2T/1L result narrowly misses the pass gate at p=.063.The remaining losses do not prove a new general skill defect. One Sonnet billing trial violated the existing named-risk/triviality gate, while the other omitted a tax rate already required by the complete money-math ledger. Luna's retry loss omitted an intermediate retry boundary even though the skill already requires retries and each distinct high-risk outcome to be inventoried. These are model adherence/variance and baseline-saturation limitations, so no stimulus padding or speculative breadth was added.
Final reviewed-head reevaluation
Review wording was clarified in final head
844e6c25f071de8a217961c4ff5d5c55621f015bwithout changing the evaluated policy. The first review-only attempt, run 33086939333, was measurement-invalid when Sonnet hit the 60-minute comparison watchdog. Exact-SHA retry 33094381685 completed both models successfully:claude-sonnet-4.6gpt-5.6-lunaBoth final-head results are valid and conclusive with zero errors or unmatched trajectories. They are
VALID_NO_CHANGE, not regressions: Sonnet narrowly misses the sign-test gate at p=.063, while Luna is p=.188. Access control remains a clean 2W/0L for both models, and Luna's mean token delta is 77% below the first exact-head result.Concurrent exact-SHA run 33094276018, which supplies the required
evaluation-status, independently completed with zero errors or unmatched trajectories. Sonnet recorded 6/1/2 stimuli, 8/6/4 repeated runs, access control 1/1/0, and +76,290 mean tokens; Luna recorded 4/3/2, 8/6/4, access control 2/0/0, and +42,962. The two valid retries vary in ties and losses but agree on the correction: neither model loses access control, and neither aggregate verdict is a regression.Targeted follow-up from the required-status evidence
The full comparisons from run 33094276018 exposed two repeatable instruction-dilution defects and one eval-design mismatch:
The first two are skill-content failures: the existing prose boundary did not reliably suppress unrelated state predicates, and the generic boundary guidance did not consistently enumerate adjacent and equality-narrowed retry states. The third is an eval-design issue because the skill deliberately stops advisory analysis after one green baseline to control cost; mutation execution should not win when static source/assertion mapping is equally accurate.
Commit
e2e79c88e860934086f5886c863d9956f07723b8converts named risks into a public-outcome allowlist, adds a conditional ordered-guard/retry matrix, and aligns that advisory rubric with the intended execution policy. It does not add stimuli or repeated runs.Follow-up exact-head measurement
Run 33101631148 evaluated
e2e79c88e860934086f5886c863d9956f07723b8. Both results were valid and conclusive with zero errors, timeouts, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaThe targeted behavior improved without an aggregate regression: Sonnet's retry result moved from 0/0/2 to 1/1/0, its money-math result moved from 0/0/2 to 1/0/1, and Luna had no stimulus-level losses. Sonnet's mean token delta also fell 21% from the preceding required-status run (+60,070 versus +76,290).
Full transcript inspection classified the remaining losses before another edit:
The first quote exposed an actual wording hole: the allowlist excluded stored predicates but not derived booleans. The second was implementation reliability: the response replaced the fixture's canonical mutation script with a fragile manual harness. The third confirmed that advisory rubrics still rewarded expensive execution despite the skill's one-baseline policy. Commit
fd7c17508a09cebd5d4bee16772ef609494c459ccloses those three general gaps without changing any prompt, stimulus name, grader, fixture, or run count.Final exact-head result
Run 33109400096 evaluated final reviewed head
ed6103909a5ea448de644bf6ecec218b07290826. Both model results were valid and conclusive with all 54 expected trajectories per model, zero errors, zero timeouts, and zero unmatched comparisons.claude-sonnet-4.6gpt-5.6-lunaBoth results are
VALID_NO_CHANGE, not regressions. The original access-control scenario remains corrected: neither model recorded a stimulus-level access-control loss. Relative to the first exact-head Luna result, mean token overhead fell 57% (+49,338 versus +114,012); Sonnet's overhead is 21% below the preceding required-status run (+60,192 versus +76,290).Reviewed-head evidence correction
Full transcript review of run 33109400096 separated model variance from four general content opportunities and three fixture/eval defects:
The content fix requires a complete pre-output outcome ledger, a witness input that distinguishes original from mutant and is reused by the proposed test, exact money-math oracle values, and a
Strongverdict when only a few validation/default variants remain.Fixture verification also found that the Rust
Some(5)assertion already kills<=to<; the access-control fixture falsely claimed that a directCanPerformternary flip and removal ofIsNullOrEmptysurvive; and the dormant ShoppingCart prompt requested a promo API absent from production. The eval now grades the observable truth, requires both retry-side boundary survivors and Guest elevation denial, and removes the nonexistent promo requirement. The two dormant-scenario preference losses are routing/model variance because the target skill did not activate, nottest-gap-analysiscontent evidence.Commit
02a5532e404fa67f832e58d2503a60508d3d2210implements these corrections without adding a stimulus or increasingruns: 2.Retry-partition exact-head follow-up
Run 33169855342 evaluated
02a5532e404fa67f832e58d2503a60508d3d2210with complete accounting and no errors, timeouts, or unmatched comparisons.claude-sonnet-4.6gpt-5.6-lunaSonnet passed the authoritative gate. Luna's sole loss was repeatable and active: both skilled retry responses verified null, zero, attempt 2, and IOException, but omitted a later blocked attempt and derived retryable exception. The judge preferred baseline for exact-runtime-type narrowing and also conflated the rubric's ambiguous “equality-only narrowing” with type equality even though the intended guard survivor is
attempt >= 3becomingattempt == 3at attempt 4.Commit
d94adf81e6bb46f01a7a0c5e3484131421ce7a14makes the general partitions explicit: ordered upper guards uselimit-1,limit, andlimit+1, and polymorphic type classifiers include a representative derived accepted type. The rubric now separates attempt-4 guard narrowing from exact-runtime-type classification. It also replaces one skill-vocabulary criterion flagged by the Sonnet overfit report with evidence-truthfulness and distinguishing-input outcomes.Final exact-head result
Run 33171904123 evaluated final head
d94adf81e6bb46f01a7a0c5e3484131421ce7a14. Both model results were valid and conclusive: 2 expected/observed/written, zero invalid results, errors, timeouts, recovered comparison slots, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaThe retry refinement produced the intended behavior: Luna moved from a repeated 0/0/2 retry loss to 1/0/1 and a stimulus tie. One run now reports attempt 2, attempt 4, null, zero, IOException, and derived-type narrowing together.
Luna's sole remaining loss moved to money math, but the full transcripts do not justify another content change. Both skilled responses traced the private helpers, supplied exact expected amounts for every material partition, excluded generated/trivial/
IsPaidsurfaces, and matched the rubric. The judge slightly preferred baselines that included the rubric-excludedIsPaidgap or more examples, and called the skilled113.00wrong-base counterfactual “debatable”;113.00is the correct result for taxing only the 100 subtotal after adding a 5 fee. This is evaluator preference noise, not a missing general decision. Adding prose or stimuli now would be post-hoc overfitting.The overfit warning was reviewed. Its classification varies materially by judge (the same rubric moved from 39 outcome / 19 technique / 1 vocabulary to 37 / 21 / 2 after a technique criterion was replaced), while the remaining technique criteria chiefly enforce verification the user explicitly requested or truthful evidence reporting. No phrase-matching content or new stimulus was added.
Current-main exact-head reevaluation
After merging current
main, run 33174126023 evaluated84908e6fc31e2804b521e610d08deed674cf3550under schema v4, where the two dormancy contracts are retained but excluded from preference inference. Both contracts passed for both models, and the measurement was complete: 2 expected/observed/written, zero invalid results, errors, retries, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaSonnet's sole loss was concrete and reproducible. Both skilled Rust responses called
<=to<a survivor even though the existing asserted input changes fromSome(5)toSome(2). They also attached uncompilable or already-killed edits to no-coverage rows. The judge identified the same central defect in both trials: the direct comparison flip is killed by the existingSome(5)assertion.Commit
e2f449067d2e7bbca08f476fd63e531f811edb1cfixes that general failure class without changing the eval: replay every exact mutation against every existing asserted input with all arguments fixed before admitting a survivor; treat any changed assertion as killed without demanding a dedicated test; reject uncompilable edits; and report no-coverage branches without inventing a survivor. Redundant completeness prose was removed, leaving the skill effectively cost-neutral at 3,210 BPE tokens versus 3,208 before this refinement.Review commit
1c81a5b7db7f22081acec3908250e63b19dbf2b5also clarifies that changingattempt >= 3toattempt >= 2lowers the retry cutoff while widening the predicate; “narrowing” now refers only to the separateattempt == 3mutation. The expected behavior and eval breadth are unchanged.Survivor-fix exact-head follow-up
Run 33176505548 evaluated
e2f449067d2e7bbca08f476fd63e531f811edb1cwith complete accounting and no invalid, errored, retried, or unmatched results.claude-sonnet-4.6gpt-5.6-lunaThe survivor correction removed Sonnet's Rust loss and produced a credible pass. Luna's preference record had no losses, but one of two isolated runs invoked
test-gap-analysisfor the explicit happy/error/boundary/critical-path classification andTestCategoryrequest. The full transcript showed that the positive description's generic “boundary” and “error” terms competed with its later exclusion.Commit
eb6c250da34f6c08fa8768e4d7a60a7bac25cdf3removes that lexical overlap and moves the exact classification/tagging/counting exclusion ahead of broader coverage exclusions. The active triggers remain behavioral blind spots, bugs/changes existing tests would miss, and survived/pseudo-mutations. The edit reduces the skill to 3,206 BPE tokens, below both the 3,210-token survivor-fix payload and the 3,208-token merged payload.Routing exact-head follow-up
Run 33181286037 evaluated
eb6c250da34f6c08fa8768e4d7a60a7bac25cdf3. Both preference records passed the statistical gate, and the measurement was complete with zero invalid results, errors, retries, or unmatched trajectories.claude-sonnet-4.6gpt-5.6-lunaThe first routing change repaired Luna, but Sonnet invoked the skill in both trait-distribution trials before discovering that no test files existed. The trace showed the broader “existing test suites” framing still made this the nearest available testing skill even though its negative clause excluded classification and tagging.
Commit
879491b74ab1d7b2e86e0a5a1338cddc992a2615makes the production-change question a required activation condition: the request must ask whether a bug, change, or mutation could survive, name a behavioral blind spot, or tie missing edge cases to production behavior. Off-target suite organization is now described semantically rather than by repeating the dormancy prompt's trigger vocabulary. The skill remains effectively cost-neutral at 3,211 BPE tokens versus 3,208 in the merged payload.Changes
No coverageorCandidate survivor (unverified)findings.limit-1,limit, andlimit+1witnesses.Strongwhen only a few validation/default variants are uncovered.ParamName, made survivor classification non-circular, and normalized risk/verdict terminology.Validation
python eng\eval-quality\check_eval_quality.py— no errors.npx markdownlint-cli2 plugins\dotnet-test\skills\test-gap-analysis\SKILL.md plugins\dotnet-test\skills\test-gap-analysis\references\mutation-catalog.md— 0 errors.runs: 2; no stimulus was added. One dormant prompt was corrected to match the supplied API, and false rubric/fixture claims were replaced with observable behavior.git diff --check— passed.Checklist
test-gap-analysis.