Skip to content

Make eval configuration modular with set overrides and explicit relaxed matching - #6

Merged
jonabur merged 4 commits into
mainfrom
feature/per-eval-catalogue
Oct 2, 2026
Merged

jonabur merged 4 commits into
mainfrom
feature/per-eval-catalogue

Conversation

@jonabur

@jonabur jonabur commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Eval configuration is editable one benchmark at a time, and named sets select whole eval groups with concise overrides and exclusions. Python and the standalone dashboard resolve the same rules, with explicit strict/relaxed matching and warnings when differing few-shot settings contribute.

version: 1
name: flagship-1
mode: fixed
exclude_languages: [kat_Geor]
evals:
  - name: ARC Challenge
    shots: 10
  - name: SIB-200
    metric: acc
  - name: X-CSQA
    exclude_languages: [eng_Latn]

This is an abbreviated illustration; the shipped set lists every required eval. Omitted settings inherit from the catalogue. This PR remains draft for review.

Configuration authoring

  • Each configs/evals/*.yaml file contains the eval's interpretation, normalization, warnings, component weights, and task-language mappings. configs/catalogue.yaml is a small directory manifest; adding an eval file needs no second registration list.
  • Per-eval language_defaults removes repeated evidence/notes and allows local exceptions. Scope is inferred from explicit language fields, with explicit overrides supported. Clipping defaults to true and is omitted from ordinary YAML.
  • The config field is metric_filter; CSV exports still use filter. Blank and literal none are distinct exact matches. Existing imported catalogues using the old filter key need that field renamed.
  • Named sets may override metric, metric_filter, and shots for an entire eval group. The UI identifies overridden fields. Catalogue exports retain the global defaults; set exports retain their overrides/exclusions.
  • Whole-eval requirements expand to known eligible catalogue tasks, independent of loaded results. Missing languages warn even when another language is present. The shipped set no longer repeats the task inventory. Explicit task membership lists remain supported; no variant-setting selectors or general task-exclusion regexes are added.
  • Set-wide/per-eval exclude_languages combine, using explicit language mappings. Either translation endpoint excludes the pair. Whole component language groups may be excluded; partial component selection remains an error.
  • Unknown eval names, task references, and language exclusions without a matching assignment are immediate config errors. Present excluded data warns; absent excluded data is not missing coverage. Valid requirements missing from a CSV warn and are excluded from scoring.
  • Builds and Python accept either a manifest or a complete catalogue. Browser import/export remains one self-contained catalogue file. Invalid imports preserve models and settings.

Matching and intentional behavior changes

Strict matching is the default. Relaxed matching prefers the expected shot count, then a uniquely closest alternative; ties are excluded with warnings. It never chooses by score, substitutes metrics/filters, or pairs different harnesses/backends. Component groups still require one consistent actual protocol. A prominent inconsistency notice and structured expected/actual-shot diagnostics appear only for included relaxed mismatches.

Strict shot mismatches are grouped into one warning per eval, model, and expected/actual shot count, with an affected-task count and expandable list. Duplicate missing-setting/no-selected-score/missing-suite warnings for the same cause are suppressed; coverage remains incomplete, and genuinely missing tasks or other scoring failures still warn. Shared tests and browser checks cover the grouping, counts, task highlights, and unchanged exclusion policy.

Working catalogue expectations are ARC Challenge 10 shots, PIQA 10, and MGSM 5. Their 0-shot translated/global variants need relaxed matching unless the set overrides the expectation. These defaults are explicit review points.

The following sample comparison uses Original weights and standard aggregation; it compares against the configuration at c1659fa:

Configuration Selected measurements Score / 100
Before this refactor 403 43.344832
Any available, relaxed 403 43.344832
Any available, strict 338 45.049944
flagship-1, relaxed 396 / 396 required 43.359673
flagship-1, strict 332 / 396 required 45.047417

Relaxed freeform coverage and scores are unchanged. Relaxed flagship coverage removes exactly six Georgian tasks (including both FLORES200 directions) and English X-CSQA. All retained measurements have unchanged normalized scores. Strict flagship is visibly incomplete; its different score reflects the remaining shared subset, not improved model performance. Counts are derived from configuration, not hardcoded.

Validation

  • 55 Python tests passed, including shared Python/JavaScript edge cases and all 72 current combinations of shipped profiles, sets, aggregation modes, matching modes, and independent/complete/missing-coverage sample analyses.
  • 101 JavaScript tests passed.
  • Public Chrome checks passed: modular/portable catalogue import, effective overrides, strict/relaxed transitions, detailed warnings, language exclusions, catalogue/set export separation, rollback on invalid config/data, synthetic regeneration, and no network requests for local imports.
  • Shared tests cover default inheritance, metric/filter/shot overrides, absent/present excluded languages, both translation directions, empty selection, invalid references, component completeness, Unicode task ordering, ambiguity, missing metrics, and preserved actual measurement identities.
  • Standalone sample and fictional builds passed. The installed CLI example reports the expected −2.5-point delta.
  • README and affected commands reviewed; documentation file/heading links, Python 3.8 syntax, formatting/lint, and whitespace checks passed.
  • CI also tests an installed wheel outside the checkout with Node absent from PATH.

Documentation impact

Updated README, configuration authoring guidance, the configuration reference, Python API/CLI guide, and development guide. The fictional example has its own neutral eval set: strict reference validation correctly rejects the project's prompted-PIQA exclusion when used with a catalogue that lacks that eval. No new runtime dependency was added.

@jonabur jonabur changed the title Split the catalogue into self-contained per-eval YAML files Make eval configuration modular with set overrides and explicit relaxed matching Oct 2, 2026
@jonabur
jonabur marked this pull request as ready for review October 2, 2026 08:25
@jonabur
jonabur merged commit 29f99b1 into main Oct 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant