Track 11,923+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
-
Updated
Sep 14, 2026 - Python
Track 11,923+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
MyPhoneBench: Do Phone-Use Agents Respect Your Privacy?
20 runnable LLM agent design patterns in Python, benchmarked, traced, offline-compatible. ReAct, multi-agent, constitutional AI, and more.
The same AI agent pipeline built in Mastra and LangChain. Runs in parallel, measures everything.
splyntra
A Claude Agent SDK security benchmark project
Correctness-gated measurement of mutation, edit payload, and semantic revisit across agent trajectories.
To associate your repository with the agent-benchmarking topic, visit your repo's landing page and select "manage topics."