Skip to content

Resolve double counting across original and multilingual evals #2

Description

@jonabur

Resolve how original and multilingual eval variants contribute to the same comparison, including repeated English measurements. Follow-up to #1; this work is deferred from that PR and does not block its merge.

Status: open follow-up. Ownership: Quickdash maintainer (unassigned). Next action: agree on scoring groups and language exclusions, then implement and test the policy. No implementation dependency on an external service.

Current behavior and evidence:

  • MGSM combines global_mgsm and mgsm_native_cot. Both can contribute English; the inspected export also has German, Spanish, and French in both protocols. The English entries use different few-shot settings (0 versus 5).
  • PIQA combines piqa and global_piqa_completions. Both can contribute English, with different few-shot settings (10 versus 0) and sample counts.
  • ARC Challenge combines arc_challenge and arc_challenge_mt. The inspected export has English only in the original task, so there is no repeated English measurement there. Other exports still need a defined policy.
  • MMLU and Global MMLU are intentionally separate evals and both include English. Review overlap without silently undoing that decision.

These observations establish repeated language measurements, not identical underlying question sets. Investigate task implementations/data before claiming item-level duplication. No private scores or model identities are needed in this issue.

Sampo proposed keeping translated variants separately named and dropping English from multilingual results where the original supplies English. Decide whether this is the desired policy and how to handle overlapping non-English MGSM variants and missing original-language coverage.

The dashboard already has two naming levels: a canonical eval/scoring group, and exact source task names visible in expanded views. There is no intermediate protocol subgroup. Splitting canonical evals changes their equal shares within a category; it is not merely a label change. English-balanced modes retain the configured English share but average multiple English measurements within it.

Acceptance checks:

  • Document the chosen grouping and overlap policy, including original/multilingual naming, English exclusion, non-English overlap, and MMLU/Global MMLU treatment.
  • Keep global interpretation rules distinct from comparison membership; express exclusions explicitly and make unused data visible through the existing warnings/inspection views.
  • Verify the policy against representative A/B inputs, including multilingual-only data and asymmetric coverage. Missing measurements must not silently enter one side's score.
  • Test standard, English-per-eval, and English-per-category aggregation, contribution reconciliation, and any intended change to effective eval weights.
  • Ensure breakdowns expose protocol/task identity, and update configuration documentation with a concrete example.
  • Explain remaining differences from Sampo's aggregation once his grouping policy is established; matching metric/baseline tables alone is insufficient.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions