Opus 5: pricing entry + charts with the Opus 5 vanilla arm - #77
Merged
Merged
Conversation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
Arm n=4 (i9ivEds, m82SdTE, Yi8S8wA, Wipg7XN): $7.03 / 0.839. Judge panel frozen at 2x Fable + 2x Opus 4.8 via explicit JUDGES set so the opus-5-judge pilot evals never enter the means. Third coder color #d97706 validated (CVD worst-pair dE 10.7). as_of bumped to 2026-07-24 (all three Claude entries re-verified against the pricing page). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
Judge panel frozen via JUDGES set (opus-5-judge pilot evals share the
fingerprint and must not enter cells). New parser pattern for unquantified
deduction lists ('Loses points on M3 (...) and M7 (...)') -> PARTIAL; flags
back to the 3 known eval_4 restraint mismatches. Notable: T2 (single fetch
asserted) 47% on Opus 5 vs 0% across the whole grid.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
Hint arm n=4 (UBsUYEb, ycb5f8S, 24pWHEZ, bC6TqMF): $7.47 / 0.866. Heatmap notable: T1 100% and M5 100% on Opus 5 hint; T2 22% (vs 47% vanilla). 128/128 evals parsed, flags unchanged (3 known). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
Skill arm n=4 (mvexh8R, X8Kb8T6, f7DZgwd, jRUpT56): $7.17 / 0.785 — below
Opus 5 vanilla (0.839); hint (0.866) stays the best arm. Two parser fixes:
trailing colon after point assignments ('T1=9:'), and P_V_POINTS now
requires a closing paren so prose counters ('FULL (0 annotations removed)')
are not read as points — the latter also corrects two historical cells
(e.g. R4 Opus 4.8 vanilla 75%->100%). 144/144 evals parsed; flags = the
3 known eval_4 restraint mismatches plus a 4th of the same shape on
mvexh8R (v2.4 material, recorded in the day log).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
…c composite Retitle to 'solution quality' and state the composition in the footnote (model fit 50 + boundaries & restraint 25 + test quality 25, rubric v2.3). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 Generated with Claude Code
https://claude.ai/code/session_0112rjaT9dWVPXhqm2FCQQPh