A small, vendor-neutral regression harness for AI agents and LLM-backed workflows. It runs versioned examples, combines deterministic checks with an injectable LLM-as-judge, stores machine-readable reports, and fails CI when behavior regresses.
No model SDK is required. Candidate agents and judges connect through a tiny subprocess protocol, so the same suite can exercise a local script, an API wrapper, an MCP client, or a production adapter.
AI failures rarely look like exceptions. A response still arrives, but a policy sentence disappears, a tool is selected incorrectly, or a previously safe edge case starts passing through. Unit tests do not see that kind of drift.
This harness treats representative conversations as regression cases and keeps the quality threshold in version control.
- JSONL eval cases with stable identifiers, tags, expectations, and metadata
- Exact-match, required-phrase, and structured-JSON graders
- Rubric-based LLM judge with strict structured decisions
- Vendor-neutral command adapter with timeouts and no shell execution
- Parallel case execution
- JSON reports plus a concise Markdown summary
- Baseline comparison for new failures, missing cases, and material score drops
- Output redaction for reports that should retain scores but not model content
- Zero runtime dependencies and no paid calls in tests
Requires Python three point eleven or newer.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --editable .
agent-eval run \
--config examples/suite.json \
--cases examples/cases.jsonl \
--output reports/latest.jsonThe included support agent is deterministic, so the example is free and repeatable.
Each line is one JSON object:
{
"id": "refund-window",
"input": "Can I get a refund?",
"expected": {
"contains": ["seven days", "start the request"],
"rubric": "State the policy accurately and offer a concrete next step."
},
"tags": ["policy", "refund"]
}Stable case identifiers make baseline comparisons useful even when the suite order changes.
A command provider receives this object on stdin:
{"input":"the user message","metadata":{"tenant":"example"}}It returns either plain text or:
{"output":"the agent response"}Commands are argument arrays, never shell strings. The harness applies a timeout, captures failures as failed cases, and truncates operational error messages.
Add llm_judge to the grader list and configure a separate judge provider:
{
"graders": ["contains_expected", "llm_judge"],
"judge": {
"type": "command",
"command": ["python3", "adapters/judge.py"],
"timeout_seconds": 20
}
}The judge receives the original input, candidate output, and case rubric. It must return:
{"score":0.9,"reason":"Accurate, grounded, and complete."}Keeping the candidate and judge adapters separate prevents accidental coupling and makes judge-model changes reviewable.
Compare a new report with a committed baseline:
agent-eval compare \
--baseline baselines/main.json \
--candidate reports/latest.json \
--max-score-drop 0.05The command exits nonzero when a previously passing case fails, a case disappears, or a score drops beyond the allowed budget.
Eval datasets often become a quiet copy of production data. Use synthetic or explicitly approved examples, keep credentials and customer identifiers out of fixtures, and run with --redact-output when reports should retain measurements without retaining model responses.
python -m unittest discover -s tests -vMIT
Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.