Skip to content

Repository files navigation

Agent Replay πŸ”„

License: MIT Python 3.8+ GitHub Stars

Local-first agent debugging for the "it worked yesterday" moments

Stop debugging non-deterministic AI agent failures in the dark. Agent Replay logs every agent run locally, shows you what changed, and lets you replay the exact same conversation to catch regressions.

Why?

  • Zero cloud lock-in β€” Everything stored in SQLite, runs offline
  • Cost transparency β€” See exactly what each feature costs across OpenAI/Anthropic/etc
  • Replay any run β€” Re-run the exact same conversation to catch non-deterministic failures
  • Hallucination flags β€” Auto-detect when responses contradict previous runs
  • Compare runs β€” Diff two runs side-by-side to see what changed

Problem

You ship an agent. It works. Next day, same input β†’ different (broken) output. "Nothing changed" but it's failing. Was it:

  • A prompt regression?
  • Model update from the provider?
  • Temperature randomness?
  • Context window differences?

Agent Replay records everything so you can replay, compare, and debug.

Quick Start

pip install agent-replay

1. Wrap your agent calls

from agent_replay import AgentLogger

logger = AgentLogger(db_path="./runs.db", feature="user-onboarding")

# Wrap any LLM call
response = logger.log_call(
    provider="openai",
    model="gpt-4",
    messages=[{"role": "user", "content": "Hello"}],
    response="Hi there!",
    metadata={"user_id": "123"}
)

2. View cost breakdown

agent-replay report --feature user-onboarding
Feature: user-onboarding
Total cost: $2.34
Runs: 145

By model:
  gpt-4:      $1.82 (78%)
  claude-3:   $0.52 (22%)

3. Replay a specific run

agent-replay replay --run-id abc123

Shows the full conversation, cost, and flags any hallucinations detected.

4. Compare two runs

agent-replay diff --run1 abc123 --run2 def456

Side-by-side diff of inputs, outputs, costs, and detected issues.

Features

  • βœ… Multi-provider support β€” OpenAI, Anthropic, Cohere, Ollama
  • βœ… Token & cost tracking β€” Automatic calculation per model
  • βœ… SQLite storage β€” No external services, fully offline
  • βœ… Budget alerts β€” Set per-feature cost limits
  • βœ… Run replay β€” Re-execute past conversations
  • βœ… Hallucination detection β€” Compare responses across runs
  • βœ… Export to CSV β€” For external analysis

Installation

pip install agent-replay

Or install from source:

git clone https://github.com/thehypewipe/agent-replay.git
cd agent-replay
pip install -e .

Usage

Basic logging

from agent_replay import AgentLogger

logger = AgentLogger(
    db_path="./agent_runs.db",
    feature="email-classifier"
)

# Log a call
result = logger.log_call(
    provider="anthropic",
    model="claude-sonnet-4",
    messages=[
        {"role": "user", "content": "Classify this email: ..."}
    ],
    response="Category: Support",
    tokens_in=150,
    tokens_out=10
)

print(f"Cost: ${result['cost']:.4f}")
print(f"Run ID: {result['run_id']}")

Budget alerts

logger = AgentLogger(
    db_path="./runs.db",
    feature="chatbot",
    budget_limit=10.0  # Alert if feature exceeds $10
)

# Raises warning if budget exceeded
logger.log_call(...)

CLI Commands

# View all runs
agent-replay list

# Filter by feature
agent-replay list --feature email-classifier

# Show cost report
agent-replay report

# Replay a specific run
agent-replay replay --run-id abc123

# Compare two runs
agent-replay diff --run1 abc123 --run2 def456

# Export to CSV
agent-replay export --output runs.csv

# Watch for budget violations
agent-replay watch --budget 50

Architecture

agent-replay/
β”œβ”€β”€ agent_replay/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ logger.py          # Core logging logic
β”‚   β”œβ”€β”€ models.py          # SQLite schema
β”‚   β”œβ”€β”€ pricing.py         # Model pricing data
β”‚   β”œβ”€β”€ replay.py          # Run replay logic
β”‚   β”œβ”€β”€ diff.py            # Run comparison
β”‚   β”œβ”€β”€ detector.py        # Hallucination detection
β”‚   └── cli.py             # CLI commands
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_logger.py
β”‚   β”œβ”€β”€ test_replay.py
β”‚   └── test_detector.py
β”œβ”€β”€ README.md
β”œβ”€β”€ setup.py
└── requirements.txt

Pricing Data

Agent Replay includes up-to-date pricing for:

  • OpenAI (GPT-4, GPT-3.5, etc.)
  • Anthropic (Claude Opus, Sonnet, Haiku)
  • Cohere
  • Custom models (configurable)

Contributing

Contributions welcome! Please:

  1. Fork the repo
  2. Create a feature branch
  3. Add tests
  4. Submit a PR

License

MIT License - see LICENSE file

Roadmap

  • Web UI for viewing runs
  • Integration with LangChain/LlamaIndex
  • Automatic regression testing
  • Team collaboration features
  • Slack/Discord notifications for budget alerts

Built for developers who are tired of "it worked yesterday" debugging.

⭐ Star this repo if Agent Replay saves you debugging time!

About

Local-first agent debugging for the "it worked yesterday" moments. Track AI costs, replay runs, and catch non-deterministic failures.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages