ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.
-
Updated
Oct 10, 2026 - Python
ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.
[EMNLP 2026] Messier is a cross-benchmark corpus of AI agent evaluations, linking tasks and results with models, scaffolds, verifiers, scoring rules, and execution trajectories.
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
Copy inflation in multi-turn search agents: 78-92% of generated tokens are copied from retrieved documents and carry inflated log-probabilities, breaking confidence-based methods. Diagnostic toolkit + Retrieval-Grounded Voting. Findings of EMNLP 2026.
VCR cassettes for agent trajectories: record agent runs as DAGs, replay the canonical path, only call the LLM for net-new paths
Local-first trajectory learning for coding agents. Every run makes the next one better.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Export Claude Code, Codex and Cursor sessions as ATIF trajectories. Rust core, Python CLI, installable with uv.
Discovering named behaviours in LLM agent trajectories and testing if they flag failing runs early.
Independent research on trajectory-aware AI-agent evaluation, delegated authority, control integrity, and failure-preserving reproducibility.
Capture coding-agent sessions (Claude Code / Codex) at the source — trajectory + verifiable git environment — and score each session's training value (grounded × rich × focused). Local, deterministic, no model at runtime.
To associate your repository with the agent-trajectories topic, visit your repo's landing page and select "manage topics."