Conversation
…ision Log The detail area had eight tabs, and on envctl's own runs most of them did not help follow or steer one. Services duplicated the preview URL already in the run summary. Graph repeated the stage strip for linear workflows. History listed revision and checkpoint IDs, which say what exists but not what happened or why. And the problems that actually stop a run sat on Readiness, two keypresses from where you land. The tabs are now Chat, Checkpoint, Changes, Tests and Decision Log. Nothing that can block a run was dropped with them: Chat now leads with the readiness problems Plan must resolve, any recovery in progress, and any service that is down in the VM, one line each so a small terminal keeps room for the stage itself. Healthy services, which were most of the Services tab, are left out on purpose. A stage's dependencies, the useful part of Graph, are on its Checkpoint. The Decision Log is derived from state the run already keeps, with no new persisted fields. Each entry has a time, stage, actor and a one-line reason: Plan's discovered requirements, supervisor accepts and rejections with the correction, steering with where it went, approvals, rewinds, and past needs-attention causes from the captain's log when a tracker is configured. A supervisor rejection is recorded on the attempt as a failure prefixed "supervisor correction:", so it is read back and attributed to the supervisor rather than reported as an ordinary failure. A rewind is reported as what it reran rather than the stage that was picked: that choice is not stored on the revision, and changing the objective reruns every stage whatever was picked, so naming it would sometimes be false. The composer now says it steers and that this changes the work, so an ask mode can sit beside it without the two being confused. Asking the supervisor a question and getting an answer is not part of this change; it needs a new agent invocation path and is tracked separately. Refs #24 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sam-bretz
force-pushed
the
issue-24-chat-decision-log
branch
from
September 16, 2026 16:54
ca539d7 to
8a23b71
Compare
Owner
Author
|
Rebased onto |
5 tasks
Chat could send instructions but never get an answer. "Is there a PR open for this?" sat queued for the next attempt and was never answered, because a message steers the work and there was no way to ask about it. A question is now its own thing, recorded on the revision beside messages rather than inside any attempt, so answering one cannot change an attempt's result, review, digest or approval. The coordinator answers it with a separate, read-only invocation of the stage's supervisor, on the stage's supervisor model: Claude gets only Read, Glob and Grep, Codex runs in its read-only sandbox, and neither can edit or run commands. A running worker keeps going, since the answer is a different job. With Claude the answer branches the supervisor's own session with --fork-session, so it knows what it reviewed without appending the question to the transcript it resumes when steered. Codex cannot branch a session and continuing one would pollute it, so it refuses to fork and starts fresh from the stage's summaries, review, checks, pull requests and steering; follow-ups see earlier answers either way. The job ID is derived from the question, like a readiness probe's, so a restart finds the job it started instead of asking twice, and an answer is recorded once. A run at its token ceiling refuses questions, answers count toward that ceiling attributed to the stage, and a question whose VM has been released fails saying to rewind. `envctl run ask` waits for and prints the answer, `run show` lists questions, and MCP can ask too. The rest of the issue: Checkpoint is gone as a tab, so the tabs are Chat, Changes, Tests and Decision Log. Chat is now one conversation in the order it happened, with the stage's artifacts for , . and o, and approvals with who and when. It also reads checkpoints a rewind carried over without their attempts, which would otherwise show nothing. What blocks a run moved to the run summary, visible from every tab, with service states shown when a preview is configured or one is down. The Decision Log now covers questions, Plan's scope, retries and stall nudges, token-ceiling stops, and what publishing did. Those coordinator decisions had no durable record: Recovery keeps only the latest cause and is cleared when a runtime is released, and the captain's log exists only with a tracker. Revisions now keep notes, appended inside the same mutation that makes each decision so each is written once, with stall nudges carried up from the backend, which alone knows when it sent one. Refs #24 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Removing the Checkpoint tab shifted every tab after it, and the diff key still set a fixed panel index, so pressing d to see a change opened Tests instead. Nothing asserted where d lands, so it passed the suite. The tab is now looked up by name. Refs #24 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7 tasks
sam-bretz
added this pull request to stack #46
September 17, 2026 02:59
sam-bretz
marked this pull request as ready for review
September 17, 2026 02:59
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #24.
This PR was first opened as partial. It now covers the whole issue: a real back-and-forth with the supervisor, the complete tab removal, and every Decision Log entry the issue lists.
Asking the supervisor
Keys and commands:
?in the dashboard (tabin the composer switches between steer and ask),envctl run ask <run> --node <stage> --text …, and MCPenvctl_actionwith actionask.It can't change the work. A question is stored on the revision next to messages, not inside any attempt, so answering one can't touch an attempt's result, review, digest or approval.
TestQuestionsLeaveTheWorkAndItsApprovalUnchangedchecks this.How it's answered. The coordinator runs a separate, read-only invocation of the stage's supervisor, using the stage's supervisor model:
Read,Glob,Grep. Bash is excluded because it can write as easily as it reads.--sandbox read-only.Grounding. With Claude, the answer branches the supervisor's own session using
--fork-session. It knows what it reviewed, and the question isn't added to the transcript the supervisor resumes when steered. Codex can't branch a session, and continuing one would add to it. SoCodex.RequestrefusesFork, and Codex starts fresh from a prompt holding the stage's summaries, review, checks, pull requests, steering and earlier answers.Durability. The job ID is derived from the question, the same way readiness probes do it. After a restart the coordinator finds the job it already started instead of asking twice, and each answer is recorded exactly once.
Limits.
Answering after acceptance. Questions are scheduled before the engine's completed-revision skip, so a finished stage can still be asked about while the run keeps its VM.
The rest of the issue
,.ostill work.New durable state: revision notes
Token-ceiling stops, stall nudges and publication failures weren't recorded anywhere the dashboard could read back reliably:
Recoverykeeps only the latest cause, and it's cleared when a runtime is released.Revisions now keep
Notes. Each note is appended inside the same mutation that makes the decision, so it's written once even though the engine reaches those points on every tick. Stall nudges come up from the backend, which is the only part that knows when it sent one.Two regressions tests caught in my own changes
TestAStageCarriedOverByARewindStillShowsItsResultnow guards it, and I confirmed it fails without the fix.workerread that way. Now it readslive: worker (resume 1).Testing
gofmt -l .prints nothing,go vet ./...is clean, and the fullgo test ./...passes. New tests, by layer:run show.?composer, and tab toggling.Docs updated: the first-workflow key table, "Watch it" and a new "Ask instead of steering" section; the CLI reference; and the agent guide.
Not verified against a live harness. The guest-job path, the actual
--fork-sessionbehaviour and Codex's--sandbox read-onlyare covered by tests up to the invocation boundary but haven't been run against real agents in a VM. Asking one real question on a running stage before merge is worth doing.Merge-order note (unchanged): #39 also edits Chat's model line, so whichever of #39 and this PR merges second will conflict. The fix is to keep both changes.
🤖 Generated with Claude Code