Skip to content

About

LLM-as-judge evaluator for scoring AI agent trajectories on task completion, tool correctness, and efficiency

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

agent-trajectory-evaluator

Prototype LLM-as-judge evaluator for scoring AI agent trajectories (thought → tool call → observation sequences) against a rubric: task completion, tool correctness, and efficiency. Uses Groq (free tier) for judging, with a deterministic exact-match shortcut when an expected output is known.

Why

Single input/output evals (like RAG faithfulness scoring) don't work for agents — the thing that matters is the full trajectory: which tools were called, with what args, in what order, and whether the end goal was actually reached. This scores that trajectory as a whole.

Setup

pip install -r requirements.txt
export GROQ_API_KEY=your_key_here

Usage

python run_eval.py                          # hardcoded demo trajectory
python run_eval.py --input my_test.json      # your own trajectory

Trajectory format

{
  "task": "Find the top 3 open issues labeled 'bug' and summarize them",
  "expected_output": null,
  "steps": [
    {
      "thought": "Search issues with label bug",
      "tool_name": "github_search_issues",
      "tool_input": {"label": "bug", "state": "open"},
      "observation": "Found 5 issues: #101, #103, #108, #112, #115"
    }
  ],
  "final_output": "Summarized #101 only."
}

Example output

{
  "task_completion": {"score": 2, "reason": "Only one issue was summarized"},
  "tool_correctness": {"score": 5, "reason": "Correct tool calls with necessary args"},
  "efficiency": {"score": 2, "reason": "Did not fetch all required issues"},
  "overall": 2,
  "method": "llm_judge"
}

Files

  • trajectory.py — trajectory data model + transcript formatting
  • judge.py — Groq-based LLM-as-judge scoring + deterministic shortcut
  • run_eval.py — CLI runner, hardcoded demo or --input custom JSON
  • sample_trajectory.json — example input

Notes

Prototype-stage. Not wired to a live agent — trajectories are currently authored by hand or exported manually from an agent run. Natural next step: extract trajectories directly from a LangGraph messages state and log judge scores alongside LangSmith traces.

About

LLM-as-judge evaluator for scoring AI agent trajectories on task completion, tool correctness, and efficiency

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages