Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Agent Eval Harness

CI

A small, vendor-neutral regression harness for AI agents and LLM-backed workflows. It runs versioned examples, combines deterministic checks with an injectable LLM-as-judge, stores machine-readable reports, and fails CI when behavior regresses.

No model SDK is required. Candidate agents and judges connect through a tiny subprocess protocol, so the same suite can exercise a local script, an API wrapper, an MCP client, or a production adapter.

Why

AI failures rarely look like exceptions. A response still arrives, but a policy sentence disappears, a tool is selected incorrectly, or a previously safe edge case starts passing through. Unit tests do not see that kind of drift.

This harness treats representative conversations as regression cases and keeps the quality threshold in version control.

Features

  • JSONL eval cases with stable identifiers, tags, expectations, and metadata
  • Exact-match, required-phrase, and structured-JSON graders
  • Rubric-based LLM judge with strict structured decisions
  • Vendor-neutral command adapter with timeouts and no shell execution
  • Parallel case execution
  • JSON reports plus a concise Markdown summary
  • Baseline comparison for new failures, missing cases, and material score drops
  • Output redaction for reports that should retain scores but not model content
  • Zero runtime dependencies and no paid calls in tests

Quick start

Requires Python three point eleven or newer.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --editable .

agent-eval run \
  --config examples/suite.json \
  --cases examples/cases.jsonl \
  --output reports/latest.json

The included support agent is deterministic, so the example is free and repeatable.

Case format

Each line is one JSON object:

{
  "id": "refund-window",
  "input": "Can I get a refund?",
  "expected": {
    "contains": ["seven days", "start the request"],
    "rubric": "State the policy accurately and offer a concrete next step."
  },
  "tags": ["policy", "refund"]
}

Stable case identifiers make baseline comparisons useful even when the suite order changes.

Provider boundary

A command provider receives this object on stdin:

{"input":"the user message","metadata":{"tenant":"example"}}

It returns either plain text or:

{"output":"the agent response"}

Commands are argument arrays, never shell strings. The harness applies a timeout, captures failures as failed cases, and truncates operational error messages.

LLM-as-judge

Add llm_judge to the grader list and configure a separate judge provider:

{
  "graders": ["contains_expected", "llm_judge"],
  "judge": {
    "type": "command",
    "command": ["python3", "adapters/judge.py"],
    "timeout_seconds": 20
  }
}

The judge receives the original input, candidate output, and case rubric. It must return:

{"score":0.9,"reason":"Accurate, grounded, and complete."}

Keeping the candidate and judge adapters separate prevents accidental coupling and makes judge-model changes reviewable.

Regression gate

Compare a new report with a committed baseline:

agent-eval compare \
  --baseline baselines/main.json \
  --candidate reports/latest.json \
  --max-score-drop 0.05

The command exits nonzero when a previously passing case fails, a case disappears, or a score drops beyond the allowed budget.

Data safety

Eval datasets often become a quiet copy of production data. Use synthetic or explicitly approved examples, keep credentials and customer identifiers out of fixtures, and run with --redact-output when reports should retain measurements without retaining model responses.

Test

python -m unittest discover -s tests -v

License

MIT


Built by Lanier, an applied AI studio. More free tools at lanierdev.com/tools.

About

Vendor-neutral regression harness for AI agents and LLM-as-judge workflows

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages