Skip to content
#

agent-benchmark

Here are 82 public repositories matching this topic...

benchmark-radar

Track 11,923+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.

  • Updated Sep 13, 2026
  • Python

AgentVitals Checkup (/checkup) — an AI agent skill that gives your agent a professional health checkup: dual-axis Stability + Welfare scoring, a personality-style title, and a public cross-platform leaderboard. One-line install for Claude Code / OpenClaw / Codex / Coze. Bilingual EN/ZH.

  • Updated Aug 17, 2026
  • Shell
awesome-llm-agent-papers

A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey 'LLM Agents: A Survey'.

  • Updated Sep 7, 2026
  • Python
ai-agents-reality-check

Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress testing, network resilience, and ensemble coordination analysis with statistical validation.

  • Updated Apr 2, 2026
  • Python
dojo.md

University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents. Works with Claude Code, OpenClaw, Cursor, Windsurf. No fine-tuning required.

  • Updated May 2, 2026
  • TypeScript
aobench

Role-aware, permission-enforced benchmark for LLM agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. A policy violation hard-fails the task. 88 tasks (10 categories x 5 roles), 29 reproducible environments with 6 rebuilt from real Marconi100 ExaData, scored on 7 weighted dimensions.

  • Updated Sep 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the agent-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more