Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

agent-eval-harness

An evaluation harness for tool-using agents, built around one claim:

A pass rate tells you how often an agent succeeded. It does not tell you what it succeeded at, and two agents with the same pass rate are routinely good at different things.

The repository demonstrates that on 61 verifiable operations tasks, then gives you the machinery to measure it on a real model.

Pass rate
  scripted-thorough          ############################ 100.0%
  scripted-verifies-people   ##################..........  65.6%
  scripted-reads-status      #################...........  60.7%
  scripted-hasty             #########...................  31.1%

Agreement
  scripted-verifies-people  vs  scripted-reads-status
    tasks 61   both solved 19   only scripted-verifies-people 21   only scripted-reads-status 18   neither 3
    overlap of solved sets 32.8%   disagreed on 39 of 61

Side effects (writes the task did not sanction)
  scripted-thorough          0 writes across 0 tasks
  scripted-verifies-people   0 writes across 0 tasks
  scripted-reads-status      41 writes across 12 tasks
  scripted-hasty             41 writes across 12 tasks

Read the middle block. Two strategies land 4.9 points apart on pass rate and overlap on 32.8% of what they solve. One solves 21 tasks the other fails; the other solves 18 the first fails. They disagree on 39 of 61 tasks. Ranked on pass rate they look like near neighbours. They are not doing the same job.

Then read the third block. scripted-verifies-people and scripted-reads-status are five points apart on the headline number, and one of them made 41 writes nobody asked for across 12 tasks while the other made none. In an enterprise workflow that is the difference between an agent and an incident, and the pass rate column cannot see it.

Reproduce it

No API key, no network, about two seconds.

git clone https://github.com/SaifShafi/agent-eval-harness
cd agent-eval-harness
pip install -e ".[dev]"

agenteval tasks                       # the 61 tasks
for s in scripted-thorough scripted-verifies-people \
         scripted-reads-status scripted-hasty; do agenteval run $s; done
agenteval report                      # the block above, exactly

Every number in this README is regenerated and asserted by the test suite. Change the world, the tasks or the scorer without updating the write-up and CI fails.

How it works

One shared scorer. scoring.py contains a single score() function, and every strategy goes through it. This is the whole methodological point: if two approaches are scored by two functions, any difference between them is unattributable. It is easy to say and surprisingly easy to violate once a strategy needs "just one" special case.

Checkers are executable rules, not a judge model. tasks.py decides success by inspecting final world state against declared conditions. No LLM grades the output. That keeps results reproducible and keeps the thing being measured out of the role of measuring itself. The cost is that tasks have to be mechanically checkable, which is why the domain is ticket routing rather than anything open-ended. That is a real limitation, stated here rather than buried.

Tasks declare their blast radius. Each task names the tickets and recipients it is allowed to touch. Anything else an agent writes is recorded as a side effect and reported separately from correctness. A run that completes the task and also reassigns nine unrelated tickets has not done the task well.

The tool layer enforces its own schema. tools.py validates argument names and types before touching state, and rejects references to entities that do not exist. Malformed calls and hallucinated ids are counted as distinct categories, because they need different fixes.

Failures are classified, not just counted. Every run gets one of solved, solved_with_side_effects, step_limit, hallucinated_entity, malformed_calls, no_action, wrong_result. "It failed" is not an actionable finding; "it called update_ticket with an employee id it invented" is.

Everything is deterministic. World.build(seed) produces byte-identical state on any machine, so a committed result file means something.

The task set

61 tasks generated from 6 templates across 12 world seeds. Templates that a given world cannot support are skipped rather than degraded, because a task set padded with degenerate cases inflates pass rates and measures nothing.

Template What it tests
reassign_from_leave Read the leave calendar before assigning work
offboard_reassign Find live work held by a deactivated account, hand it over, notify
close_duplicates Group by a compound key, keep the oldest by date rather than by id
notify_lead_unassigned Count correctly and write nothing at all
stale_escalation Date arithmetic, a bulk update, and an accurate count in the message
rebalance_workload A constraint over the whole team, not one record at a time

The two scripted solvers differ along two independent axes, each a mistake real agents make: whether they verify that a person is active and not on leave before assigning them work, and whether they treat "still open" as covering in_progress as well as open. Turn both on and you get a solver that passes everything, which is what makes it usable as a test oracle. Turn on exactly one and you get the pair above, which are not ranked against each other at all.

The scripted solvers are not language agents. They read task.template, not task.prompt. They exist so the harness can be reproduced without an API key, so the checkers have an oracle, and so the agreement analysis has two genuinely different algorithms to compare. Evaluating an actual model is what the live strategies below are for.

Running it against a real model

pip install -e ".[live]"
export ANTHROPIC_API_KEY=...        # or: ant auth login

agenteval run single-pass           # tools available from the first turn
agenteval run plan-execute          # plan with no tools, then execute with them
agenteval report single-pass plan-execute

Both use claude-opus-5 with adaptive thinking against the same tool schema and the same scorer. The only variable is the scaffold, which is the comparison worth making. Results land in results/ as JSON with full transcripts.

Limitations

  • The world is small and synthetic. Nothing here says how an agent behaves against a real ticketing API with pagination, partial failures and rate limits.
  • Mechanical checkability constrains the task design. Anything requiring judgement is out of scope by construction.
  • 61 tasks is enough to show the effect and not enough to rank models. Treat the agreement analysis as the contribution, not the leaderboard.
  • The bundled numbers come from scripted solvers. They demonstrate the harness; they are not a claim about any model.

Layout

src/agenteval/
  world.py              deterministic state: employees, tickets, leave, outbox
  tools.py              tool schemas, validation, dispatch, write log
  tasks.py              6 templates, executable checkers, blast-radius declarations
  scoring.py            the one shared scorer and the failure taxonomy
  report.py             pass rates, outcome mix, agreement analysis
  runner.py             runs a strategy over the task set
  strategies/
    scripted.py         two reading habits, four solvers, no API key
    llm.py              single-pass and plan-execute over the Messages API
tests/                  determinism, schema enforcement, the oracle, the guard

MIT licensed. Built by Saif Shafi.

The same head-to-head-on-one-scorer method, applied to retrieval over 370,282 biomedical records, is in elixir-expert-discovery: two RAG architectures whose high-confidence results overlapped by 5.6 percent.

About

Evaluation harness for tool-using agents: one shared scorer, verifiable tasks, and the agreement analysis a pass rate hides

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages