Skip to content

Report the model each role actually ran, not the configured one - #39

Merged
sam-bretz merged 1 commit into
mainfrom
issue-38-resolved-models
Sep 17, 2026
Merged

sam-bretz merged 1 commit into
mainfrom
issue-38-resolved-models

Conversation

@sam-bretz

Copy link
Copy Markdown
Owner

Closes #38.

What changed

Every surface named a stage's models from configuration. A role with no configured model is invoked without --model and the harness picks for itself, so harness default was as much as any surface could say — and an alias like sonnet never showed the version behind it. After a run finished, "which model did this use?" had no answer.

Each attempt now records what its harness reports running, and the surfaces prefer that over the configured value.

Configured Reported Shown
unset claude-sonnet-5-20260115 claude-sonnet-5-20260115 (harness default)
sonnet claude-sonnet-5-20260115 claude-sonnet-5-20260115 (configured sonnet)
opus opus opus
opus nothing reported opus

How

Claude names its model on the init event of its stream — the same event the session identity already comes from — so the adapter reads both in one pass (claudeInit). agent.Harness gains Model(stream) string; the backend records it per role beside the session; engine.Observation carries it; the engine writes it onto Attempt.Models, preferring a reported model because it is the more precise of the two.

Display: internal/tui/model.go (stage view), internal/web/view.go + app.js (stage panel and timeline).

Acceptance criteria

  • TUI stage view shows the model each role actually ran
  • Web stage panel does the same
  • A harness-default role shows which model that resolved to
  • Timeline defect fixed (see below)
  • Recorded, not only displayed — answerable after the run finishes

The timeline defect

The issue noted this line:

const modelText = models.worker || models.supervisor ? ` with ${model(models.worker)}` : '';

It tested whether either role had a model and then always printed the worker's, so a stage with a supervisor model and no worker model announced with harness default — naming the wrong role. It now names each role that reported one, and the supervisor appears in the timeline at all for the first time.

A limitation worth knowing

Codex reports nothing. Its documented events carry no model, and I did not want to invent a schema, so Codex.Model reads a top-level model key only if a stream happens to provide one. For Codex runs these surfaces behave exactly as before, falling back to the configured value. If someone knows the real Codex field, that adapter is a two-line change.

Claude's init.model field is likewise read defensively: a stream that says nothing leaves the configured value showing rather than a guess.

Testing

gofmt, go vet ./..., full go test ./... pass. New tests cover the adapter parse (including a stream that says nothing), the engine merge (unset role, alias→version, and that an unchanged attempt is not rewritten), the TUI formatter's five cases, and the web view reporting ran_worker without inventing a supervisor.

Not exercised against a live harness — the init.model field is read from Claude's documented stream shape, so that is worth one real run before merge.

🤖 Generated with Claude Code

Every surface named a stage's models from configuration. A role with no
configured model is invoked without --model and the harness chooses for
itself, so "harness default" was as much as any of them could say, and an
alias like sonnet never showed the version behind it. After a run
finished there was no way to answer which model it had used.

Each attempt now records what its harness reports running. Claude names
its model on the init event of its stream, the same event the session
identity already comes from, so the adapter reads both in one pass. The
backend records it per role beside the session, the observation carries
it, and the engine writes it onto the attempt, preferring a reported
model over a configured one because it is the more precise of the two: it
names what an unset role resolved to and the version behind an alias.

The TUI stage view and the dashboard stage panel show what ran, falling
back to the configured value when a harness reports nothing rather than
guessing. Codex documents no model on its events, so its adapter reads
one only if a stream provides it under the usual key and otherwise
reports nothing, which leaves those surfaces exactly as they were.

Also fixes the dashboard timeline, which tested whether either role had a
model and then always printed the worker's: a stage with a supervisor
model and no worker model announced "with harness default", naming the
wrong role. It now names each role that reported one, and the supervisor
appears there at all for the first time.

Refs #38

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sam-bretz

Copy link
Copy Markdown
Owner Author

Rebased onto main after #40, #41 and #42 merged. One conflict: #42 and this PR each appended a test to the end of internal/web/server_test.go; both are kept. Everything else auto-merged, and the full go test ./... passes.

Heads-up for merge order: this PR and #44 both rewrite the same line in the dashboard's Conversation/Chat panel — this one to show the model that actually ran, #44 to lead the panel with blocking diagnostics. They don't conflict with main today, but whichever merges second will conflict with the first. The resolution is to keep both: blocking(rev, node) first, then the stageModelTUI model line.

@sam-bretz
sam-bretz force-pushed the issue-38-resolved-models branch from 413c370 to 8479a9b Compare September 16, 2026 17:00
@sam-bretz
sam-bretz merged commit cfef4cf into main Sep 17, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Show the resolved worker and supervisor model in the TUI and dashboard

1 participant