Local-first agent debugging for the "it worked yesterday" moments
Stop debugging non-deterministic AI agent failures in the dark. Agent Replay logs every agent run locally, shows you what changed, and lets you replay the exact same conversation to catch regressions.
- Zero cloud lock-in β Everything stored in SQLite, runs offline
- Cost transparency β See exactly what each feature costs across OpenAI/Anthropic/etc
- Replay any run β Re-run the exact same conversation to catch non-deterministic failures
- Hallucination flags β Auto-detect when responses contradict previous runs
- Compare runs β Diff two runs side-by-side to see what changed
You ship an agent. It works. Next day, same input β different (broken) output. "Nothing changed" but it's failing. Was it:
- A prompt regression?
- Model update from the provider?
- Temperature randomness?
- Context window differences?
Agent Replay records everything so you can replay, compare, and debug.
pip install agent-replayfrom agent_replay import AgentLogger
logger = AgentLogger(db_path="./runs.db", feature="user-onboarding")
# Wrap any LLM call
response = logger.log_call(
provider="openai",
model="gpt-4",
messages=[{"role": "user", "content": "Hello"}],
response="Hi there!",
metadata={"user_id": "123"}
)agent-replay report --feature user-onboardingFeature: user-onboarding
Total cost: $2.34
Runs: 145
By model:
gpt-4: $1.82 (78%)
claude-3: $0.52 (22%)
agent-replay replay --run-id abc123Shows the full conversation, cost, and flags any hallucinations detected.
agent-replay diff --run1 abc123 --run2 def456Side-by-side diff of inputs, outputs, costs, and detected issues.
- β Multi-provider support β OpenAI, Anthropic, Cohere, Ollama
- β Token & cost tracking β Automatic calculation per model
- β SQLite storage β No external services, fully offline
- β Budget alerts β Set per-feature cost limits
- β Run replay β Re-execute past conversations
- β Hallucination detection β Compare responses across runs
- β Export to CSV β For external analysis
pip install agent-replayOr install from source:
git clone https://github.com/thehypewipe/agent-replay.git
cd agent-replay
pip install -e .from agent_replay import AgentLogger
logger = AgentLogger(
db_path="./agent_runs.db",
feature="email-classifier"
)
# Log a call
result = logger.log_call(
provider="anthropic",
model="claude-sonnet-4",
messages=[
{"role": "user", "content": "Classify this email: ..."}
],
response="Category: Support",
tokens_in=150,
tokens_out=10
)
print(f"Cost: ${result['cost']:.4f}")
print(f"Run ID: {result['run_id']}")logger = AgentLogger(
db_path="./runs.db",
feature="chatbot",
budget_limit=10.0 # Alert if feature exceeds $10
)
# Raises warning if budget exceeded
logger.log_call(...)# View all runs
agent-replay list
# Filter by feature
agent-replay list --feature email-classifier
# Show cost report
agent-replay report
# Replay a specific run
agent-replay replay --run-id abc123
# Compare two runs
agent-replay diff --run1 abc123 --run2 def456
# Export to CSV
agent-replay export --output runs.csv
# Watch for budget violations
agent-replay watch --budget 50agent-replay/
βββ agent_replay/
β βββ __init__.py
β βββ logger.py # Core logging logic
β βββ models.py # SQLite schema
β βββ pricing.py # Model pricing data
β βββ replay.py # Run replay logic
β βββ diff.py # Run comparison
β βββ detector.py # Hallucination detection
β βββ cli.py # CLI commands
βββ tests/
β βββ test_logger.py
β βββ test_replay.py
β βββ test_detector.py
βββ README.md
βββ setup.py
βββ requirements.txt
Agent Replay includes up-to-date pricing for:
- OpenAI (GPT-4, GPT-3.5, etc.)
- Anthropic (Claude Opus, Sonnet, Haiku)
- Cohere
- Custom models (configurable)
Contributions welcome! Please:
- Fork the repo
- Create a feature branch
- Add tests
- Submit a PR
MIT License - see LICENSE file
- Web UI for viewing runs
- Integration with LangChain/LlamaIndex
- Automatic regression testing
- Team collaboration features
- Slack/Discord notifications for budget alerts
Built for developers who are tired of "it worked yesterday" debugging.
β Star this repo if Agent Replay saves you debugging time!