Skip to content

feat: show agent setup, teardown, and scoring as Braintrust spans - #360

Merged
Rodriguespn merged 9 commits into
mainfrom
Rodriguespn/ai-1279-setup-scoring-spans
Oct 5, 2026
Merged

Rodriguespn merged 9 commits into
mainfrom
Rodriguespn/ai-1279-setup-scoring-spans

Conversation

@Rodriguespn

@Rodriguespn Rodriguespn commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #351. Adds setup, teardown, and a timed score span around task, so the whole run shows up in each Braintrust trace. Before, the root span ended at the result file's mtime with ~50s unexplained after task (Slack), and the first LLM span shrank to ~0s.

eval        run start → scoring end
├─ setup    CLI boot until it records the prompt
├─ task     prompt → last transcript event
├─ teardown CLI exit until agent.run() returns
└─ passed   workspace export + checks + judge calls

Summary

  • Results record host epoch ms agentRunStartedAt, agentPromptAt, agentRunEndedAt, scoringEndedAt (all optional; older results keep the mtime fallback).
  • agentRunDurationMs now measures only agent time, the task span (prompt → last transcript event), falling back to CLI process time without transcript timestamps. The new timing fields stay out of the exported web snapshot.
  • Each harness reads the prompt time from its session files: Codex rollout user message, Claude Code first user record, Grok first user_message_chunk, OpenCode opencode.db via node:sqlite (new enrichEvents). It's clamped between run start and the first transcript event, since it comes from the sandbox clock.

Review guide

  • Start with logTranscript in apps/framework/scripts/upload-braintrust.ts and its test.
  • Then run-eval.ts (where the times are taken) and the four enrichEvents/parser changes.

Verification

  • upload-braintrust + core tests (195), pnpm typecheck, biome on changed files.
  • Eval refresh run: 8/8 passed (4 harnesses × 2 evals). Every trace has setup → task → teardown → passed:
Experiment setup teardown passed
codex-gpt-6-luna 1.5–2.3s 0.2–0.3s 2.9–3.0s
claude-code-sonnet-5 3.6–3.8s 0.4s 2.1–4.1s
grok-4.7 0.7–0.9s 0.3s 3.4–4.0s
opencode-kimi-k3 1.8–2.5s 0.4s 2.6–4.3s

Compare with #351's codex "After" trace, or browse all experiments in this run.

Risk

Sandbox/stack boot before the agent and stack teardown after scoring are still outside the trace. Judge calls inside passed get their own spans in AI-1271.

Linear: AI-1279

@vercel

vercel Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
evals Ignored Ignored Preview Oct 5, 2026 4:55pm UTC

Request Review

@mattrossman

mattrossman commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

I realize that our existing usage.agentRunDurationMs doesn't map cleanly to the setup/task/teardown spans, instead it roughly maps to setup + task + part of teardown. We might want to make it map more cleanly to the task span only since that's the part I expect to vary per agent.

Comment thread apps/framework/scripts/export-results.ts Outdated
@mattrossman

Copy link
Copy Markdown
Collaborator

Traces look good, thanks for adding this!

@Rodriguespn
Rodriguespn marked this pull request as ready for review October 5, 2026 15:54
@Rodriguespn
Rodriguespn requested a review from a team as a code owner October 5, 2026 15:54
@Rodriguespn

Rodriguespn commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Agreed. agentRunDurationMs is now the task span's window, prompt to last transcript event (both on the sandbox clock). On #351's CI Codex artifacts it comes out 1.6–2.5s lower than before

Rodriguespn and others added 9 commits October 5, 2026 09:51
Record host epoch-ms agent run start/end and scoring end in results so the
uploader anchors task at the real start and times teardown and score spans.
Older results keep the mtime fallback.

Refs AI-1279

Co-Authored-By: Claude <noreply@anthropic.com>
Each harness reads when the CLI recorded the user prompt from its session
files (Codex rollout, Claude Code session JSONL, Grok updates.jsonl,
OpenCode opencode.db) and results carry it as agentPromptAt. The uploader
ends setup there and starts task and the first LLM span at it, clamped to
the run start and the first transcript event.

Refs AI-1279

Co-Authored-By: Claude <noreply@anthropic.com>
@Rodriguespn
Rodriguespn force-pushed the Rodriguespn/ai-1279-setup-scoring-spans branch from 9ed618d to adbcd41 Compare October 5, 2026 16:55
@Rodriguespn
Rodriguespn changed the base branch from mattrossman/ai-1264-align-eval-trace-structure-with-braintrusts-eval-span-spec to main October 5, 2026 16:55
@Rodriguespn
Rodriguespn merged commit 6513a88 into main Oct 5, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants