Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
-
Updated
Aug 8, 2026 - Python
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
A task-level audit of the grading rubrics in OpenAI's GDPval. 10,453 criteria opened: score scales vary tenfold across tasks, and one untestable line carries up to 21.7% of a task's score.
Add a description, image, and links to the gdpval topic page so that developers can more easily learn about it.
To associate your repository with the gdpval topic, visit your repo's landing page and select "manage topics."