Skip to content

feat: supervisor Q&A in Chat, Decision Log, and the dashboard tab rework (#24) - #44

Open
sam-bretz wants to merge 3 commits into
mainfrom
issue-24-chat-decision-log
Open

sam-bretz wants to merge 3 commits into
mainfrom
issue-24-chat-decision-log

Conversation

@sam-bretz

@sam-bretz sam-bretz commented Sep 16, 2026 •

Copy link
Copy Markdown
Owner

Closes #24.

This PR was first opened as partial. It now covers the whole issue: a real back-and-forth with the supervisor, the complete tab removal, and every Decision Log entry the issue lists.

Asking the supervisor

Keys and commands: ? in the dashboard (tab in the composer switches between steer and ask), envctl run ask <run> --node <stage> --text …, and MCP envctl_action with action ask.

It can't change the work. A question is stored on the revision next to messages, not inside any attempt, so answering one can't touch an attempt's result, review, digest or approval. TestQuestionsLeaveTheWorkAndItsApprovalUnchanged checks this.

How it's answered. The coordinator runs a separate, read-only invocation of the stage's supervisor, using the stage's supervisor model:

  • Claude gets only Read,Glob,Grep. Bash is excluded because it can write as easily as it reads.
  • Codex runs with --sandbox read-only.
  • A running worker keeps going, because the answer runs as a separate job.

Grounding. With Claude, the answer branches the supervisor's own session using --fork-session. It knows what it reviewed, and the question isn't added to the transcript the supervisor resumes when steered. Codex can't branch a session, and continuing one would add to it. So Codex.Request refuses Fork, and Codex starts fresh from a prompt holding the stage's summaries, review, checks, pull requests, steering and earlier answers.

Durability. The job ID is derived from the question, the same way readiness probes do it. After a restart the coordinator finds the job it already started instead of asking twice, and each answer is recorded exactly once.

Limits.

  • A run at its token ceiling refuses new questions.
  • Answers count toward that ceiling and are attributed to the stage.
  • A question whose VM has been released (closed, cancelled or superseded run) fails and says to rewind.

Answering after acceptance. Questions are scheduled before the engine's completed-revision skip, so a finished stage can still be asked about while the run keeps its VM.

The rest of the issue

  • Checkpoint is gone as a tab. Tabs are now Chat, Changes, Tests and Decision Log.
  • Chat reads as one conversation in time order. It shows attempts, worker summaries, supervisor accepts and rejections, steering, questions with a "thinking" state and then the answer, and approvals with who approved and when. It also lists the stage's artifacts, so , . o still work.
  • Blockers are in the run summary, visible from every tab. Service states show when a preview is configured or a service is down.
  • The Decision Log adds questions, Plan's scope, retries and stall nudges, token-ceiling stops, and publication results (the PR URL, or why publishing failed).

New durable state: revision notes

Token-ceiling stops, stall nudges and publication failures weren't recorded anywhere the dashboard could read back reliably:

  • Recovery keeps only the latest cause, and it's cleared when a runtime is released.
  • The captain's log exists only when a tracker is configured.

Revisions now keep Notes. Each note is appended inside the same mutation that makes the decision, so it's written once even though the engine reaches those points on every tick. Stall nudges come up from the backend, which is the only part that knows when it sent one.

Two regressions tests caught in my own changes

  • Inherited checkpoints went invisible. A rewind copies earlier checkpoints into the new revision without the attempts that produced them. My first Chat thread was built only from attempts, so an inherited stage showed nothing, and with the Checkpoint tab removed there was nowhere else to see it. The existing history test caught it. TestAStageCarriedOverByARewindStillShowsItsResult now guards it, and I confirmed it fails without the fix.
  • "worker: worker (resume 1)". Removing the "Live:" prefix made a worker whose phase is worker read that way. Now it reads live: worker (resume 1).

Testing

gofmt -l . prints nothing, go vet ./... is clean, and the full go test ./... passes. New tests, by layer:

  • workflow: attaching a question to the latest attempt, every refusal case, the token ceiling, usage accounting, and work and approval staying unchanged.
  • agent: read-only tools, forking, and Codex refusing to fork.
  • engine: answering while running, waiting and accepted; exactly one harness call per answer; every failure path; and each note recorded once.
  • localexec: the answer invocation (fork versus fresh, the stage's model) and the prompt's grounding, including "is there a PR open?"
  • daemon: the ask action.
  • CLI: waiting, failure, timeout, no-wait, and run show.
  • dashboard: conversation order, approval actor and time, inherited checkpoints, the ? composer, and tab toggling.

Docs updated: the first-workflow key table, "Watch it" and a new "Ask instead of steering" section; the CLI reference; and the agent guide.

Not verified against a live harness. The guest-job path, the actual --fork-session behaviour and Codex's --sandbox read-only are covered by tests up to the invocation boundary but haven't been run against real agents in a VM. Asking one real question on a running stage before merge is worth doing.

Merge-order note (unchanged): #39 also edits Chat's model line, so whichever of #39 and this PR merges second will conflict. The fix is to keep both changes.

🤖 Generated with Claude Code

…ision Log

The detail area had eight tabs, and on envctl's own runs most of them did
not help follow or steer one. Services duplicated the preview URL already
in the run summary. Graph repeated the stage strip for linear workflows.
History listed revision and checkpoint IDs, which say what exists but not
what happened or why. And the problems that actually stop a run sat on
Readiness, two keypresses from where you land.

The tabs are now Chat, Checkpoint, Changes, Tests and Decision Log.
Nothing that can block a run was dropped with them: Chat now leads with
the readiness problems Plan must resolve, any recovery in progress, and
any service that is down in the VM, one line each so a small terminal
keeps room for the stage itself. Healthy services, which were most of the
Services tab, are left out on purpose. A stage's dependencies, the useful
part of Graph, are on its Checkpoint.

The Decision Log is derived from state the run already keeps, with no new
persisted fields. Each entry has a time, stage, actor and a one-line
reason: Plan's discovered requirements, supervisor accepts and rejections
with the correction, steering with where it went, approvals, rewinds, and
past needs-attention causes from the captain's log when a tracker is
configured. A supervisor rejection is recorded on the attempt as a failure
prefixed "supervisor correction:", so it is read back and attributed to
the supervisor rather than reported as an ordinary failure.

A rewind is reported as what it reran rather than the stage that was
picked: that choice is not stored on the revision, and changing the
objective reruns every stage whatever was picked, so naming it would
sometimes be false.

The composer now says it steers and that this changes the work, so an ask
mode can sit beside it without the two being confused. Asking the
supervisor a question and getting an answer is not part of this change;
it needs a new agent invocation path and is tracked separately.

Refs #24

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sam-bretz
sam-bretz force-pushed the issue-24-chat-decision-log branch from ca539d7 to 8a23b71 Compare September 16, 2026 16:54
@sam-bretz

Copy link
Copy Markdown
Owner Author

Rebased onto main after #40, #41 and #42 merged. One conflict, in the dashboard footer: #41 added c close while this PR renamed To/i chat to Steer/i steer. Both belong, so the footer now reads Steer <recipient> · i steer · … · c close · … in both the full and compact layouts. No semantic overlap with #41's finished-runs grouping or VM visibility. Full go test ./... passes.

Chat could send instructions but never get an answer. "Is there a PR
open for this?" sat queued for the next attempt and was never answered,
because a message steers the work and there was no way to ask about it.

A question is now its own thing, recorded on the revision beside
messages rather than inside any attempt, so answering one cannot change
an attempt's result, review, digest or approval. The coordinator answers
it with a separate, read-only invocation of the stage's supervisor, on
the stage's supervisor model: Claude gets only Read, Glob and Grep, Codex
runs in its read-only sandbox, and neither can edit or run commands. A
running worker keeps going, since the answer is a different job.

With Claude the answer branches the supervisor's own session with
--fork-session, so it knows what it reviewed without appending the
question to the transcript it resumes when steered. Codex cannot branch
a session and continuing one would pollute it, so it refuses to fork and
starts fresh from the stage's summaries, review, checks, pull requests
and steering; follow-ups see earlier answers either way. The job ID is
derived from the question, like a readiness probe's, so a restart finds
the job it started instead of asking twice, and an answer is recorded
once. A run at its token ceiling refuses questions, answers count toward
that ceiling attributed to the stage, and a question whose VM has been
released fails saying to rewind. `envctl run ask` waits for and prints
the answer, `run show` lists questions, and MCP can ask too.

The rest of the issue: Checkpoint is gone as a tab, so the tabs are Chat,
Changes, Tests and Decision Log. Chat is now one conversation in the
order it happened, with the stage's artifacts for , . and o, and
approvals with who and when. It also reads checkpoints a rewind carried
over without their attempts, which would otherwise show nothing. What
blocks a run moved to the run summary, visible from every tab, with
service states shown when a preview is configured or one is down.

The Decision Log now covers questions, Plan's scope, retries and stall
nudges, token-ceiling stops, and what publishing did. Those coordinator
decisions had no durable record: Recovery keeps only the latest cause
and is cleared when a runtime is released, and the captain's log exists
only with a tracker. Revisions now keep notes, appended inside the same
mutation that makes each decision so each is written once, with stall
nudges carried up from the backend, which alone knows when it sent one.

Refs #24

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sam-bretz sam-bretz changed the title feat: cut the dashboard to Chat, Checkpoint, Changes, Tests and a Decision Log (#24, partial) feat: supervisor Q&A in Chat, Decision Log, and the dashboard tab rework (#24) Sep 16, 2026
Removing the Checkpoint tab shifted every tab after it, and the diff key
still set a fixed panel index, so pressing d to see a change opened Tests
instead. Nothing asserted where d lands, so it passed the suite. The tab
is now looked up by name.

Refs #24

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sam-bretz
sam-bretz added this pull request to stack #46 September 17, 2026 02:59
@sam-bretz
sam-bretz marked this pull request as ready for review September 17, 2026 02:59

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Rework dashboard tabs: real supervisor Chat, Decision Log, fewer panels

1 participant