Separate comparison inputs and standardize scoring and warnings - #1
Conversation
|
Scoring alignment with Sampo’s updated table Compared all 49 task/metric entries, covering all 46 catalogue evals. All selected metrics and baselines now agree, including JEEBench at 0.1055. This does not establish identical grouping or aggregate weights.
This comparison establishes metric/baseline agreement only. The dictionary does not specify extraction filters, shots, selected task summaries, source-scale conversion, clipping, category/language weights, or aggregation. Those cannot be verified from this table. The default comparison’s only configured scoring caveat is MultiBlimp’s Croatian/Serbian pooling. Unused prompted PIQA data and runtime data issues still generate notices/warnings. Tests cover all six requested cases: active caveats, unused eval data (including alternate-only metrics), missing required evals, missing scoring fields, freeform mismatches, and missing interpretation rules. Extra catalogue rules absent from both models do not warn in freeform mode. The draft also separates model results, independent weighting profiles, optional named eval sets, and the global catalogue. Changing the selected eval set preserves edited weights. The user approved merging with grouping and double-counting policy deferred to issue #2. Sources: |
Compare new model results without redefining how each eval is interpreted. Separate the global catalogue, independent weighting profiles, and optional named eval sets; align scoring choices and make exclusions visible.
configs/catalogue.yamldefines task matching, categories, metrics, normalization, languages, and caveats.configs/weights/offers Original (startup default) and Code & math emphasis (Code/Math 0.2 each, Reasoning 0.1, Knowledge/Commonsense 0.125 each), including English shares and the default calculation. Reading and the final three category weights are identical between profiles.configs/sets/defines optional expected coverage. Separate selectors/imports/exports preserve edited weights when changing eval sets.accwith a comment instead of a warning; use the shared JEEBench baseline of 0.1055; keep AMC23 at 0; set ARC Easy to 0.25; select FLORES200chrf++and OpenSubtitleschrf. Translation retains native 0–100 scores without chance correction, documented as the accepted policy. SQuAD v2 usesf1.The HTML remains self-contained and works offline. CLI inputs are
--catalogue,--weights/--weights-dir, and--eval-set/--sets-dir; Pages and fictional-demo builds use these inputs. A missing configured metric never silently falls back to another metric. Profiles can omit catalogue categories, which then contribute zero with a warning when shared data uses them.Follow-up: the discussion checklist records the scoring decisions. All 49 entries in Sampo’s updated table now match our metrics and baselines, including JEEBench at 0.1055. Grouping and double counting are deferred to issue #2, as agreed before merging. Current behavior: we combine ARC Challenge/translated ARC, the two MGSM prefixes, and PIQA/global completions into three evals. Filters, shots, scaling, clipping, membership, and aggregation cannot be established from that baseline table alone.
Validation: 35 Python tests, 87 Node tests, and both public-fixture/full-export browser suites passed locally. Tests cover all six warning cases, exclusion arithmetic in every aggregate, absent data on both sides, incompatible protocols, schema rejection, independent selection, import/export, rollback, direct startup with a restricted set, contribution reconciliation, and no network requests during local comparison. Shared/demo builds and documentation links passed. CI exercises public fixtures without private exports.
Documentation: updated the README, configuration reference, development/Pages commands, and config contribution guide for separate inputs, explicit freeform exclusions, metric provenance, and warning behavior.