Repository navigation
Refresh LUMI results and use corrected CoT evals by default - #11
Merged
Merged
Conversation
This was referenced Oct 9, 2026
jonabur
marked this pull request as ready for review
October 9, 2026 09:42
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The new LUMI exports contain corrected CoT reasoning and code-continuation runs under distinct task names, while
flagship-1still selects the original runs. Refresh the shared exports, add five merge3 checkpoints, and makeflagship-1select the corrected protocols.The new models are
v1anneal_merge3_116k_120k,v1annealC_merge3_116k_120k,v1annealCT_merge3_116k_120k,v1annealTT_merge3_116k_120k, andv2anneal_merge3_116k_120k. Checkpoint labels and CSV bytes are preserved exactly as exported.Each of the eight combined
flag-evals-471files has 2,123 rows, including 101 added rows across 33 corrected task names. Comparing the three refreshed files with their previous versions:v1anneal_120kv1annealC_120k_l0fixv2anneal_120kThe corrected results have separate catalogue entries. The default set replaces ten original entries: AIME24/25, AMC23, GPQA Diamond, JEEBench, MATH500, PolyMath, HumanEval, LiveCodeBench, and MBPP. It selects
pass@1, filterall, at 0 shots except MBPP continuation at 3 shots. PolyMath includes all four difficulty levels in each of six languages, retaining relative weights 1, 2, 4, and 8. Its original task matcher is narrowed so it cannot also match CoT tasks.Original protocols and alternate metrics remain inspectable. Missing corrected results produce incomplete coverage instead of falling back to original runs, including under relaxed matching.
flagship-1counts only the selected protocol; the broadAny availableset can include both. The startup model pair, weighting profile, category assignments, and normalization floors are retained. The inherited JEEBench floor of 0.1055 remains a shared scoring convention rather than a newly established CoT baseline.With the Original profile and strict matching, the current startup comparison changes from v1C 48.96 / v2 49.27 to v1C 53.02 / v2 52.27. Both before and after have 365/396 shared requirements; the existing 31 Global PIQA shot mismatches remain. All 33 corrected tasks contribute for every model without original-protocol duplicates.
Validation:
/scratch/project_465002530/poppelko/flag-evals/results/on LUMI, retrieved 2026-10-09. The combined exports already include the separateflag-cotresults, so one CSV per checkpoint is committed.python3 -m tests.checkpasses, including Python/JavaScript contract tests and public browser checks.