Skip to content

Refresh LUMI results and use corrected CoT evals by default - #11

Merged
jonabur merged 2 commits into
mainfrom
feat/refresh-lumi-merge3-results
Oct 9, 2026
Merged

jonabur merged 2 commits into
mainfrom
feat/refresh-lumi-merge3-results

Conversation

@jonabur

@jonabur jonabur commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

The new LUMI exports contain corrected CoT reasoning and code-continuation runs under distinct task names, while flagship-1 still selects the original runs. Refresh the shared exports, add five merge3 checkpoints, and make flagship-1 select the corrected protocols.

The new models are v1anneal_merge3_116k_120k, v1annealC_merge3_116k_120k, v1annealCT_merge3_116k_120k, v1annealTT_merge3_116k_120k, and v2anneal_merge3_116k_120k. Checkpoint labels and CSV bytes are preserved exactly as exported.

Each of the eight combined flag-evals-471 files has 2,123 rows, including 101 added rows across 33 corrected task names. Comparing the three refreshed files with their previous versions:

Existing model Added rows Removed rows Changes to existing measurements
v1anneal_120k 101 0 None
v1annealC_120k_l0fix 101 0 None
v2anneal_120k 101 0 Source-path metadata only; scores and settings unchanged

The corrected results have separate catalogue entries. The default set replaces ten original entries: AIME24/25, AMC23, GPQA Diamond, JEEBench, MATH500, PolyMath, HumanEval, LiveCodeBench, and MBPP. It selects pass@1, filter all, at 0 shots except MBPP continuation at 3 shots. PolyMath includes all four difficulty levels in each of six languages, retaining relative weights 1, 2, 4, and 8. Its original task matcher is narrowed so it cannot also match CoT tasks.

Original protocols and alternate metrics remain inspectable. Missing corrected results produce incomplete coverage instead of falling back to original runs, including under relaxed matching. flagship-1 counts only the selected protocol; the broad Any available set can include both. The startup model pair, weighting profile, category assignments, and normalization floors are retained. The inherited JEEBench floor of 0.1055 remains a shared scoring convention rather than a newly established CoT baseline.

With the Original profile and strict matching, the current startup comparison changes from v1C 48.96 / v2 49.27 to v1C 53.02 / v2 52.27. Both before and after have 365/396 shared requirements; the existing 31 Global PIQA shot mismatches remain. All 33 corrected tasks contribute for every model without original-protocol duplicates.

Validation:

  • Verified imported CSVs against source SHA-256 checksums. Source: /scratch/project_465002530/poppelko/flag-evals/results/ on LUMI, retrieved 2026-10-09. The combined exports already include the separate flag-cot results, so one CSV per checkpoint is committed.
  • Added a regression shared by Python and JavaScript for corrected task/metric selection, missing-result behavior, and complete PolyMath components. Behavioral expectations use invented scores.
  • Added a controlled browser fixture checking corrected scores, inspectable original runs, and incomplete coverage without fallback in strict and relaxed modes.
  • python3 -m tests.check passes, including Python/JavaScript contract tests and public browser checks.
  • Built the real shared dataset and passed browser smoke across all eight models, both sets, and all six views: 16 comparisons, with the startup pair verified.

@jonabur jonabur changed the title Refresh LUMI exports and add merge3 checkpoints Refresh LUMI results and use corrected CoT evals by default Oct 9, 2026
@jonabur
jonabur marked this pull request as ready for review October 9, 2026 09:42
@jonabur
jonabur merged commit 0a50bb1 into main Oct 9, 2026
3 checks passed

This branch was successfully deployed

1 active deployment
github-pages — 389476aa Deployed Oct 9, 2026 by jonabur via publish #5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant