Make eval configuration modular with set overrides and explicit relaxed matching - #6
Merged
Merged
Conversation
jonabur
marked this pull request as ready for review
October 2, 2026 08:25
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Eval configuration is editable one benchmark at a time, and named sets select whole eval groups with concise overrides and exclusions. Python and the standalone dashboard resolve the same rules, with explicit strict/relaxed matching and warnings when differing few-shot settings contribute.
This is an abbreviated illustration; the shipped set lists every required eval. Omitted settings inherit from the catalogue. This PR remains draft for review.
Configuration authoring
configs/evals/*.yamlfile contains the eval's interpretation, normalization, warnings, component weights, and task-language mappings.configs/catalogue.yamlis a small directory manifest; adding an eval file needs no second registration list.language_defaultsremoves repeated evidence/notes and allows local exceptions. Scope is inferred from explicit language fields, with explicit overrides supported. Clipping defaults to true and is omitted from ordinary YAML.metric_filter; CSV exports still usefilter. Blank and literalnoneare distinct exact matches. Existing imported catalogues using the oldfilterkey need that field renamed.metric,metric_filter, andshotsfor an entire eval group. The UI identifies overridden fields. Catalogue exports retain the global defaults; set exports retain their overrides/exclusions.exclude_languagescombine, using explicit language mappings. Either translation endpoint excludes the pair. Whole component language groups may be excluded; partial component selection remains an error.Matching and intentional behavior changes
Strict matching is the default. Relaxed matching prefers the expected shot count, then a uniquely closest alternative; ties are excluded with warnings. It never chooses by score, substitutes metrics/filters, or pairs different harnesses/backends. Component groups still require one consistent actual protocol. A prominent inconsistency notice and structured expected/actual-shot diagnostics appear only for included relaxed mismatches.
Strict shot mismatches are grouped into one warning per eval, model, and expected/actual shot count, with an affected-task count and expandable list. Duplicate missing-setting/no-selected-score/missing-suite warnings for the same cause are suppressed; coverage remains incomplete, and genuinely missing tasks or other scoring failures still warn. Shared tests and browser checks cover the grouping, counts, task highlights, and unchanged exclusion policy.
Working catalogue expectations are ARC Challenge 10 shots, PIQA 10, and MGSM 5. Their 0-shot translated/global variants need relaxed matching unless the set overrides the expectation. These defaults are explicit review points.
The following sample comparison uses Original weights and standard aggregation; it compares against the configuration at
c1659fa:Relaxed freeform coverage and scores are unchanged. Relaxed flagship coverage removes exactly six Georgian tasks (including both FLORES200 directions) and English X-CSQA. All retained measurements have unchanged normalized scores. Strict flagship is visibly incomplete; its different score reflects the remaining shared subset, not improved model performance. Counts are derived from configuration, not hardcoded.
Validation
Documentation impact
Updated README, configuration authoring guidance, the configuration reference, Python API/CLI guide, and development guide. The fictional example has its own neutral eval set: strict reference validation correctly rejects the project's prompted-PIQA exclusion when used with a catalogue that lacks that eval. No new runtime dependency was added.