Prototype LLM-as-judge evaluator for scoring AI agent trajectories (thought → tool call → observation sequences) against a rubric: task completion, tool correctness, and efficiency. Uses Groq (free tier) for judging, with a deterministic exact-match shortcut when an expected output is known.
Single input/output evals (like RAG faithfulness scoring) don't work for agents — the thing that matters is the full trajectory: which tools were called, with what args, in what order, and whether the end goal was actually reached. This scores that trajectory as a whole.
pip install -r requirements.txt
export GROQ_API_KEY=your_key_herepython run_eval.py # hardcoded demo trajectory
python run_eval.py --input my_test.json # your own trajectory{
"task": "Find the top 3 open issues labeled 'bug' and summarize them",
"expected_output": null,
"steps": [
{
"thought": "Search issues with label bug",
"tool_name": "github_search_issues",
"tool_input": {"label": "bug", "state": "open"},
"observation": "Found 5 issues: #101, #103, #108, #112, #115"
}
],
"final_output": "Summarized #101 only."
}{
"task_completion": {"score": 2, "reason": "Only one issue was summarized"},
"tool_correctness": {"score": 5, "reason": "Correct tool calls with necessary args"},
"efficiency": {"score": 2, "reason": "Did not fetch all required issues"},
"overall": 2,
"method": "llm_judge"
}trajectory.py— trajectory data model + transcript formattingjudge.py— Groq-based LLM-as-judge scoring + deterministic shortcutrun_eval.py— CLI runner, hardcoded demo or--inputcustom JSONsample_trajectory.json— example input
Prototype-stage. Not wired to a live agent — trajectories are currently authored by hand or exported manually from an agent run. Natural next step: extract trajectories directly from a LangGraph messages state and log judge scores alongside LangSmith traces.