Skip to content

Coalesce intentional exclusions into information instead of warnings - #17

Merged
jonabur merged 2 commits into
mainfrom
feat/intentional-exclusions
Oct 9, 2026
Merged

jonabur merged 2 commits into
mainfrom
feat/intentional-exclusions

Conversation

@jonabur

@jonabur jonabur commented Oct 9, 2026

Copy link
Copy Markdown
Member

flagship-1 produces a separate Not used warning for each model/eval containing deliberately excluded data. Replaced evaluation protocols and excluded languages account for 34 of the default comparison's 37 warnings.

Allow the existing exclude list in fixed sets as well as available-mode sets. Explicit eval and language exclusions now form one Intentional exclusions informational record across models and evals, with expandable task lists. The warning badge, Python warning emission, and strict diagnostic failures count only warnings. Missing required results, unexpected unused data, metric/shot mismatches, and scoring caveats still warn.

The flagship-1 additions explicitly name the original protocols already omitted from the set, plus prompted Global PIQA. Inline comments explain each exclusion, for example:

exclude:
  - AIME25 # Replaced by aime25_cot with corrected CoT evaluation settings.

Excluded evals remain out of scope if results arrive later. An eval cannot be both required and excluded, and unknown names remain errors. Comments document the authored YAML; imports/exports preserve the exclusion list but do not preserve YAML comments.

The current startup comparison has 3 warnings and 1 informational summary: two Global PIQA shot mismatches and the existing MultiBLiMP grouping caveat remain warnings. Scores, coverage, complete measurement audits, effective weights, and contributions are identical before and after this change. The existing CoT selections and Sampo's latest config revert are preserved.

Validation:

  • python3 -m tests.check passes, including both scoring engines and Chrome.
  • Shared regressions cover fixed/available exclusions, both translation endpoints, overlapping exclusions, missing requirements, validation failures, and information delivery through Python and CLI strict modes.
  • Browser regressions verify the warning badge, a single information row, YAML export/import, unchanged scores, and rollback after an invalid exclusion.
  • Compared full reports for the real default pair and reviewed a local render.

Fixes #8.

Leave unmerged for review, including the annotated flagship-1 exclusions.

@github-actions

github-actions Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Preview closed

This PR is closed and its preview has been removed.

@jonabur
jonabur merged commit 8871d5b into main Oct 9, 2026
2 checks passed

This branch was successfully deployed

1 active deployment
github-pages — b9703179 Deployed Oct 9, 2026 by jonabur via publish #25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Coalesce warnings from language exclusion

1 participant