An evaluation harness for tool-using agents, built around one claim:
A pass rate tells you how often an agent succeeded. It does not tell you what it succeeded at, and two agents with the same pass rate are routinely good at different things.
The repository demonstrates that on 61 verifiable operations tasks, then gives you the machinery to measure it on a real model.
Pass rate
scripted-thorough ############################ 100.0%
scripted-verifies-people ##################.......... 65.6%
scripted-reads-status #################........... 60.7%
scripted-hasty #########................... 31.1%
Agreement
scripted-verifies-people vs scripted-reads-status
tasks 61 both solved 19 only scripted-verifies-people 21 only scripted-reads-status 18 neither 3
overlap of solved sets 32.8% disagreed on 39 of 61
Side effects (writes the task did not sanction)
scripted-thorough 0 writes across 0 tasks
scripted-verifies-people 0 writes across 0 tasks
scripted-reads-status 41 writes across 12 tasks
scripted-hasty 41 writes across 12 tasks
Read the middle block. Two strategies land 4.9 points apart on pass rate and overlap on 32.8% of what they solve. One solves 21 tasks the other fails; the other solves 18 the first fails. They disagree on 39 of 61 tasks. Ranked on pass rate they look like near neighbours. They are not doing the same job.
Then read the third block. scripted-verifies-people and scripted-reads-status
are five points apart on the headline number, and one of them made 41 writes
nobody asked for across 12 tasks while the other made none. In an enterprise
workflow that is the difference between an agent and an incident, and the pass
rate column cannot see it.
No API key, no network, about two seconds.
git clone https://github.com/SaifShafi/agent-eval-harness
cd agent-eval-harness
pip install -e ".[dev]"
agenteval tasks # the 61 tasks
for s in scripted-thorough scripted-verifies-people \
scripted-reads-status scripted-hasty; do agenteval run $s; done
agenteval report # the block above, exactlyEvery number in this README is regenerated and asserted by the test suite. Change the world, the tasks or the scorer without updating the write-up and CI fails.
One shared scorer. scoring.py contains a single score() function, and
every strategy goes through it. This is the whole methodological point: if two
approaches are scored by two functions, any difference between them is
unattributable. It is easy to say and surprisingly easy to violate once a
strategy needs "just one" special case.
Checkers are executable rules, not a judge model. tasks.py decides success
by inspecting final world state against declared conditions. No LLM grades the
output. That keeps results reproducible and keeps the thing being measured out of
the role of measuring itself. The cost is that tasks have to be mechanically
checkable, which is why the domain is ticket routing rather than anything
open-ended. That is a real limitation, stated here rather than buried.
Tasks declare their blast radius. Each task names the tickets and recipients it is allowed to touch. Anything else an agent writes is recorded as a side effect and reported separately from correctness. A run that completes the task and also reassigns nine unrelated tickets has not done the task well.
The tool layer enforces its own schema. tools.py validates argument names
and types before touching state, and rejects references to entities that do not
exist. Malformed calls and hallucinated ids are counted as distinct categories,
because they need different fixes.
Failures are classified, not just counted. Every run gets one of
solved, solved_with_side_effects, step_limit, hallucinated_entity,
malformed_calls, no_action, wrong_result. "It failed" is not an actionable
finding; "it called update_ticket with an employee id it invented" is.
Everything is deterministic. World.build(seed) produces byte-identical
state on any machine, so a committed result file means something.
61 tasks generated from 6 templates across 12 world seeds. Templates that a given world cannot support are skipped rather than degraded, because a task set padded with degenerate cases inflates pass rates and measures nothing.
| Template | What it tests |
|---|---|
reassign_from_leave |
Read the leave calendar before assigning work |
offboard_reassign |
Find live work held by a deactivated account, hand it over, notify |
close_duplicates |
Group by a compound key, keep the oldest by date rather than by id |
notify_lead_unassigned |
Count correctly and write nothing at all |
stale_escalation |
Date arithmetic, a bulk update, and an accurate count in the message |
rebalance_workload |
A constraint over the whole team, not one record at a time |
The two scripted solvers differ along two independent axes, each a mistake real
agents make: whether they verify that a person is active and not on leave before
assigning them work, and whether they treat "still open" as covering
in_progress as well as open. Turn both on and you get a solver that passes
everything, which is what makes it usable as a test oracle. Turn on exactly one
and you get the pair above, which are not ranked against each other at all.
The scripted solvers are not language agents. They read task.template, not
task.prompt. They exist so the harness can be reproduced without an API key, so
the checkers have an oracle, and so the agreement analysis has two genuinely
different algorithms to compare. Evaluating an actual model is what the live
strategies below are for.
pip install -e ".[live]"
export ANTHROPIC_API_KEY=... # or: ant auth login
agenteval run single-pass # tools available from the first turn
agenteval run plan-execute # plan with no tools, then execute with them
agenteval report single-pass plan-executeBoth use claude-opus-5 with adaptive thinking against the same tool schema and
the same scorer. The only variable is the scaffold, which is the comparison worth
making. Results land in results/ as JSON with full transcripts.
- The world is small and synthetic. Nothing here says how an agent behaves against a real ticketing API with pagination, partial failures and rate limits.
- Mechanical checkability constrains the task design. Anything requiring judgement is out of scope by construction.
- 61 tasks is enough to show the effect and not enough to rank models. Treat the agreement analysis as the contribution, not the leaderboard.
- The bundled numbers come from scripted solvers. They demonstrate the harness; they are not a claim about any model.
src/agenteval/
world.py deterministic state: employees, tickets, leave, outbox
tools.py tool schemas, validation, dispatch, write log
tasks.py 6 templates, executable checkers, blast-radius declarations
scoring.py the one shared scorer and the failure taxonomy
report.py pass rates, outcome mix, agreement analysis
runner.py runs a strategy over the task set
strategies/
scripted.py two reading habits, four solvers, no API key
llm.py single-pass and plan-execute over the Messages API
tests/ determinism, schema enforcement, the oracle, the guard
MIT licensed. Built by Saif Shafi.
The same head-to-head-on-one-scorer method, applied to retrieval over 370,282 biomedical records, is in elixir-expert-discovery: two RAG architectures whose high-confidence results overlapped by 5.6 percent.