feat: show agent setup, teardown, and scoring as Braintrust spans - #360
Merged
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
Collaborator
|
I realize that our existing |
mattrossman
reviewed
Oct 2, 2026
mattrossman
approved these changes
Oct 2, 2026
Collaborator
|
Traces look good, thanks for adding this! |
Rodriguespn
marked this pull request as ready for review
October 5, 2026 15:54
Contributor
Author
|
Agreed. |
Record host epoch-ms agent run start/end and scoring end in results so the uploader anchors task at the real start and times teardown and score spans. Older results keep the mtime fallback. Refs AI-1279 Co-Authored-By: Claude <noreply@anthropic.com>
Each harness reads when the CLI recorded the user prompt from its session files (Codex rollout, Claude Code session JSONL, Grok updates.jsonl, OpenCode opencode.db) and results carry it as agentPromptAt. The uploader ends setup there and starts task and the first LLM span at it, clamped to the run start and the first transcript event. Refs AI-1279 Co-Authored-By: Claude <noreply@anthropic.com>
This reverts commit 115caf7.
This reverts commit 23b381e.
Rodriguespn
force-pushed
the
Rodriguespn/ai-1279-setup-scoring-spans
branch
from
October 5, 2026 16:55
9ed618d to
adbcd41
Compare
Rodriguespn
changed the base branch from
mattrossman/ai-1264-align-eval-trace-structure-with-braintrusts-eval-span-spec
to
main
October 5, 2026 16:55
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #351. Adds
setup,teardown, and a timed score span aroundtask, so the whole run shows up in each Braintrust trace. Before, the root span ended at the result file's mtime with ~50s unexplained aftertask(Slack), and the first LLM span shrank to ~0s.Summary
agentRunStartedAt,agentPromptAt,agentRunEndedAt,scoringEndedAt(all optional; older results keep the mtime fallback).agentRunDurationMsnow measures only agent time, thetaskspan (prompt → last transcript event), falling back to CLI process time without transcript timestamps. The new timing fields stay out of the exported web snapshot.userrecord, Grok firstuser_message_chunk, OpenCodeopencode.dbvianode:sqlite(newenrichEvents). It's clamped between run start and the first transcript event, since it comes from the sandbox clock.Review guide
logTranscriptinapps/framework/scripts/upload-braintrust.tsand its test.run-eval.ts(where the times are taken) and the fourenrichEvents/parser changes.Verification
upload-braintrust+ core tests (195),pnpm typecheck, biome on changed files.setup → task → teardown → passed:Compare with #351's codex "After" trace, or browse all experiments in this run.
Risk
Sandbox/stack boot before the agent and stack teardown after scoring are still outside the trace. Judge calls inside
passedget their own spans in AI-1271.Linear: AI-1279