You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Resolve how original and multilingual eval variants contribute to the same comparison, including repeated English measurements. Follow-up to #1; this work is deferred from that PR and does not block its merge.
Status: open follow-up. Ownership: Quickdash maintainer (unassigned). Next action: agree on scoring groups and language exclusions, then implement and test the policy. No implementation dependency on an external service.
Current behavior and evidence:
MGSM combines global_mgsm and mgsm_native_cot. Both can contribute English; the inspected export also has German, Spanish, and French in both protocols. The English entries use different few-shot settings (0 versus 5).
PIQA combines piqa and global_piqa_completions. Both can contribute English, with different few-shot settings (10 versus 0) and sample counts.
ARC Challenge combines arc_challenge and arc_challenge_mt. The inspected export has English only in the original task, so there is no repeated English measurement there. Other exports still need a defined policy.
MMLU and Global MMLU are intentionally separate evals and both include English. Review overlap without silently undoing that decision.
These observations establish repeated language measurements, not identical underlying question sets. Investigate task implementations/data before claiming item-level duplication. No private scores or model identities are needed in this issue.
Sampo proposed keeping translated variants separately named and dropping English from multilingual results where the original supplies English. Decide whether this is the desired policy and how to handle overlapping non-English MGSM variants and missing original-language coverage.
The dashboard already has two naming levels: a canonical eval/scoring group, and exact source task names visible in expanded views. There is no intermediate protocol subgroup. Splitting canonical evals changes their equal shares within a category; it is not merely a label change. English-balanced modes retain the configured English share but average multiple English measurements within it.
Acceptance checks:
Document the chosen grouping and overlap policy, including original/multilingual naming, English exclusion, non-English overlap, and MMLU/Global MMLU treatment.
Keep global interpretation rules distinct from comparison membership; express exclusions explicitly and make unused data visible through the existing warnings/inspection views.
Verify the policy against representative A/B inputs, including multilingual-only data and asymmetric coverage. Missing measurements must not silently enter one side's score.
Test standard, English-per-eval, and English-per-category aggregation, contribution reconciliation, and any intended change to effective eval weights.
Ensure breakdowns expose protocol/task identity, and update configuration documentation with a concrete example.
Explain remaining differences from Sampo's aggregation once his grouping policy is established; matching metric/baseline tables alone is insufficient.
Resolve how original and multilingual eval variants contribute to the same comparison, including repeated English measurements. Follow-up to #1; this work is deferred from that PR and does not block its merge.
Status: open follow-up. Ownership: Quickdash maintainer (unassigned). Next action: agree on scoring groups and language exclusions, then implement and test the policy. No implementation dependency on an external service.
Current behavior and evidence:
global_mgsmandmgsm_native_cot. Both can contribute English; the inspected export also has German, Spanish, and French in both protocols. The English entries use different few-shot settings (0 versus 5).piqaandglobal_piqa_completions. Both can contribute English, with different few-shot settings (10 versus 0) and sample counts.arc_challengeandarc_challenge_mt. The inspected export has English only in the original task, so there is no repeated English measurement there. Other exports still need a defined policy.These observations establish repeated language measurements, not identical underlying question sets. Investigate task implementations/data before claiming item-level duplication. No private scores or model identities are needed in this issue.
Sampo proposed keeping translated variants separately named and dropping English from multilingual results where the original supplies English. Decide whether this is the desired policy and how to handle overlapping non-English MGSM variants and missing original-language coverage.
The dashboard already has two naming levels: a canonical eval/scoring group, and exact source task names visible in expanded views. There is no intermediate protocol subgroup. Splitting canonical evals changes their equal shares within a category; it is not merely a label change. English-balanced modes retain the configured English share but average multiple English measurements within it.
Acceptance checks: