Repository navigation
Coalesce intentional exclusions into information instead of warnings - #17
Merged
Merged
Conversation
|
Preview closed This PR is closed and its preview has been removed. |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
flagship-1produces a separate Not used warning for each model/eval containing deliberately excluded data. Replaced evaluation protocols and excluded languages account for 34 of the default comparison's 37 warnings.Allow the existing
excludelist in fixed sets as well as available-mode sets. Explicit eval and language exclusions now form one Intentional exclusions informational record across models and evals, with expandable task lists. The warning badge, Python warning emission, and strict diagnostic failures count only warnings. Missing required results, unexpected unused data, metric/shot mismatches, and scoring caveats still warn.The
flagship-1additions explicitly name the original protocols already omitted from the set, plus prompted Global PIQA. Inline comments explain each exclusion, for example:Excluded evals remain out of scope if results arrive later. An eval cannot be both required and excluded, and unknown names remain errors. Comments document the authored YAML; imports/exports preserve the exclusion list but do not preserve YAML comments.
The current startup comparison has 3 warnings and 1 informational summary: two Global PIQA shot mismatches and the existing MultiBLiMP grouping caveat remain warnings. Scores, coverage, complete measurement audits, effective weights, and contributions are identical before and after this change. The existing CoT selections and Sampo's latest config revert are preserved.
Validation:
python3 -m tests.checkpasses, including both scoring engines and Chrome.Fixes #8.
Leave unmerged for review, including the annotated
flagship-1exclusions.