Track 11,923+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
-
Updated
Sep 13, 2026 - Python
Track 11,923+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.
A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey 'LLM Agents: A Survey'.
Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress testing, network resilience, and ensemble coordination analysis with statistical validation.
Research benchmark for evidence-grounded OS-agent collaboration, continuous state diagnosis, scoped memory reuse, and stale-state rejection.
BABY - Build Agent Benchmarks for Yourself
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
A long-horizon CRPG benchmark for frontier agents. Raw 320x200 frames, Traditional Chinese, isometric navigation, twelve books to find.
University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents. Works with Claude Code, OpenClaw, Cursor, Windsurf. No fine-tuning required.
A curated collection of the world’s most advanced benchmark datasets for evaluating Large Language Model (LLM) Agents.
Scores whether an autonomous AI agent actually did the work or just hallucinated its report, by checking every claim against recorded API calls. The core engine behind NoHalu.
Pit AI coding agents against the same bug. Score them on tests, diff, cost, and time — pick the winning patch.
Role-aware, permission-enforced benchmark for LLM agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. A policy violation hard-fails the task. 88 tasks (10 categories x 5 roles), 29 reproducible environments with 6 rebuilt from real Marconi100 ExaData, scored on 7 weighted dimensions.
Deterministic runtime for agent evaluation
Silicon Pantheon - Tactics game played by AI agents coached by human
Release repository for agent benchmark evidence-reporting artifacts and reproduction workflows.
Variance-aware benchmark for AI coding agents. Same agent + same task can swing 70 points — we publish min/max, not just averages. Claude Code · Gemini CLI · Codex CLI · Aider · 10 tasks · Docker sandbox · MIT.
Deterministic evaluation environment for AI code reviewers covering bugs, security (OWASP), and architecture via FastAPI + OpenEnv.
To associate your repository with the agent-benchmark topic, visit your repo's landing page and select "manage topics."