Skip to content

Separate comparison inputs and standardize scoring and warnings - #1

Merged
jonabur merged 5 commits into
mainfrom
config/sib200-jeebench-baselines
Oct 1, 2026
Merged

jonabur merged 5 commits into
mainfrom
config/sib200-jeebench-baselines

Conversation

@jonabur

@jonabur jonabur commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Compare new model results without redefining how each eval is interpreted. Separate the global catalogue, independent weighting profiles, and optional named eval sets; align scoring choices and make exclusions visible.

  • Separate inputs: configs/catalogue.yaml defines task matching, categories, metrics, normalization, languages, and caveats. configs/weights/ offers Original (startup default) and Code & math emphasis (Code/Math 0.2 each, Reasoning 0.1, Knowledge/Commonsense 0.125 each), including English shares and the default calculation. Reading and the final three category weights are identical between profiles. configs/sets/ defines optional expected coverage. Separate selectors/imports/exports preserve edited weights when changing eval sets.
  • Coverage: Any available compares shared recognized measurements without requiring an eval list. Its explicit exclusion list omits unvalidated prompted Global PIQA. Flagship-1 requires 45 evals and 403 task/shot requirements. Missing requirements mark scores incomplete; extra data remains inspectable. Comparisons exclude unmatched measurements from both scores and redistribute remaining weights.
  • Scoring: keep SIB-200 on acc with a comment instead of a warning; use the shared JEEBench baseline of 0.1055; keep AMC23 at 0; set ARC Easy to 0.25; select FLORES200 chrf++ and OpenSubtitles chrf. Translation retains native 0–100 scores without chance correction, documented as the accepted policy. SQuAD v2 uses f1.
  • Warnings: show attached scoring caveats for evals with shared comparison data, and always in their configuration details. Excluded eval data gets Not used, even when only an alternate metric exists. Required missing evals, missing configured metrics, freeform A/B mismatches, and unconfigured tasks warn and are excluded appropriately. Unused global catalogue rules do not warn. Prompted Global PIQA has a separate warning that metric selection and normalization have not been validated; the default included-eval caveat is MultiBlimp’s Croatian/Serbian pooling. Runtime consistency warnings remain active.

The HTML remains self-contained and works offline. CLI inputs are --catalogue, --weights / --weights-dir, and --eval-set / --sets-dir; Pages and fictional-demo builds use these inputs. A missing configured metric never silently falls back to another metric. Profiles can omit catalogue categories, which then contribute zero with a warning when shared data uses them.

Follow-up: the discussion checklist records the scoring decisions. All 49 entries in Sampo’s updated table now match our metrics and baselines, including JEEBench at 0.1055. Grouping and double counting are deferred to issue #2, as agreed before merging. Current behavior: we combine ARC Challenge/translated ARC, the two MGSM prefixes, and PIQA/global completions into three evals. Filters, shots, scaling, clipping, membership, and aggregation cannot be established from that baseline table alone.

Validation: 35 Python tests, 87 Node tests, and both public-fixture/full-export browser suites passed locally. Tests cover all six warning cases, exclusion arithmetic in every aggregate, absent data on both sides, incompatible protocols, schema rejection, independent selection, import/export, rollback, direct startup with a restricted set, contribution reconciliation, and no network requests during local comparison. Shared/demo builds and documentation links passed. CI exercises public fixtures without private exports.

Documentation: updated the README, configuration reference, development/Pages commands, and config contribution guide for separate inputs, explicit freeform exclusions, metric provenance, and warning behavior.

@jonabur

jonabur commented Oct 1, 2026 •

Copy link
Copy Markdown
Member Author

Scoring alignment with Sampo’s updated table

Compared all 49 task/metric entries, covering all 46 catalogue evals. All selected metrics and baselines now agree, including JEEBench at 0.1055. This does not establish identical grouping or aggregate weights.

  • AMC23: keep zero for generated answers, as confirmed. The pinned evaluator removes the contest’s answer options.
  • ARC Easy: use 0.25 for consistency; the tiny variable-option-count difference is documented in a note.
  • Translation: prefer chrF++ over chrF over BLEU. FLORES200 selects the explicitly exported chrf++ field (rescored); OpenSubtitles selects chrf (none). Both use native 0–100 scores and zero chance correction. These choices match the updated table. The OpenSubtitles harness uses plain chrF (word_order=0); chrF++ uses word_order=2. FLORES has separate chrf and chrf++ export fields, though the CSV does not record a rescoring signature.
  • SQuAD v2: both use f1, baseline 0, per the updated table and discussion.
  • Prompted Global PIQA policy: exclude it from the supplied comparisons. Keep a separate catalogue rule with provisional exact_match / baseline 0 and the warning “Normalization and metric selection have not been validated.” Present excluded data produces a Not used notice; explicitly including it surfaces its scoring caveat. Completion-based PIQA retains acc_norm / 0.5. The baseline dictionary alone does not establish whether Sampo’s aggregate includes the prompted entry.
  • JEEBench: use 0.1055 (10.55%), matching Sampo, as agreed. The config note distinguishes this setting from the paper’s approximate 10.5% figure.
  • Confirm grouping and aggregate weights — tracked separately in issue #2. Our catalogue has 46 evals because it combines arc_challenge + arc_challenge_mt, global_mgsm + mgsm_native_cot, and piqa + global_piqa_completions. Excluding prompted PIQA leaves 45 active evals. Sampo’s dictionary has 49 entries; it does not establish whether these are separate aggregation units. If he gives them separate equal eval weights within categories, the composites differ despite matching metrics/baselines. No grouping changes have been made.

This comparison establishes metric/baseline agreement only. The dictionary does not specify extraction filters, shots, selected task summaries, source-scale conversion, clipping, category/language weights, or aggregation. Those cannot be verified from this table.

The default comparison’s only configured scoring caveat is MultiBlimp’s Croatian/Serbian pooling. Unused prompted PIQA data and runtime data issues still generate notices/warnings. Tests cover all six requested cases: active caveats, unused eval data (including alternate-only metrics), missing required evals, missing scoring fields, freeform mismatches, and missing interpretation rules. Extra catalogue rules absent from both models do not warn in freeform mode.

The draft also separates model results, independent weighting profiles, optional named eval sets, and the global catalogue. Changing the selected eval set preserves edited weights. The user approved merging with grouping and double-counting policy deferred to issue #2.

Sources:

@jonabur jonabur changed the title Refine SIB-200 scoring notes and JEEBench chance baseline Separate comparison inputs and refine SIB-200/JEEBench scoring Oct 1, 2026
@jonabur jonabur changed the title Separate comparison inputs and refine SIB-200/JEEBench scoring Separate comparison inputs and standardize scoring and warnings Oct 1, 2026
@jonabur
jonabur marked this pull request as ready for review October 1, 2026 12:22
@jonabur
jonabur merged commit 025d9f2 into main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant