✨ Slice E of #854: run the packaged Plan in xmd repl, under one immutable profile - #864
Merged
Merged
Conversation
taras
force-pushed
the
agent/issue-854-journey
branch
from
October 1, 2026 00:46
47fb629 to
b65f57c
Compare
`xmd repl` took a location and nothing else, so an Agent had no way to reach a
REPL entry and `<Plan>` had no way to run in one. It now reads the same five
Agent options `xmd run` reads, as a first-class Plugin command of its own —
selected Plugins are told `repl`, and the bundled Git Plugin declares for it,
because a REPL entry is a document run under the ordinary profile — one declaration, so the two commands cannot come
to mean different things by the same spelling — and settles them in a fixed
order: the line, then the Agent configuration, then one profile, then the
terminal. Everything a wrong invocation can be refused for is refused before a
per-user directory is formed, a history file exists or the terminal's modes are
touched.
What a REPL execution runs under is one immutable `ReplExecutionProfile`,
assembled once in the command's own scope: the selected Plugins, the Agent
identity components, the real packaged `<Plan>` declaration and the ordinary
evaluation ceiling, with the settled permission mode beside them. The program
takes that one value and nothing else, so what an entry may resolve, what ceiling
a generated fragment runs under and how a permission request is answered are
facts about the command a person invoked rather than about the entry they typed.
It carries data and installations and never authority: no stack, provider, Plan
writer, live request or scope reaches the model, the route or the Journal through
it. The REPL installs no readline permission handling, no foreground launcher and
no browser form — a question is answered in the drawer in front of the person,
and `<Session.Launch>` refuses through the established missing-launcher contract
because this command is using the terminal it would be given.
A returned Plan is a generated fragment, and a fragment that may write may now
also ask: canonical `<Elicit>` joins the ordinary write table at core's own
origin, key and revision, paired only, as a *pinned capability* rather than a
name the fragment resolves. Preflight selects it where it finds an occurrence
under a `write` selection, and the body it selects is core's own `<Elicit>`,
closed over before any document code runs — exactly as the body behind `<File>`
is. Nothing resolves the name: not before the root import, not at invocation,
and middleware answering a generated import with a body of its own is refused,
because only canonical execution answers one. So a same-name repository,
registered, declared, middleware or separately loaded `Elicit` receives no grant
and never runs, including the workflow host's own suspension replacement, which
goes on answering the *authored* element in a run exactly as it did. An
execution whose fragment never writes the element resolves nothing at all, which
is why admitting the entry cannot break a document that only writes a `<File>`.
`allow={["read"]}` still refuses it, because a read selection promises nobody
will be interrupted. What stays contextual is the interaction and only the
interaction: the pinned body asks through the Elicitation Api lexically in scope,
so the REPL drawer answers a generated question and an enclosing `<Answers>`
region answers it without anybody being asked. Its answer is retained by the
ordinary `elicit` operation, and replay restores it without asking anyone. The
`generated_xmd` record keeps the identity and the form and nothing about the
answer.
One defect the journey found is fixed here. The model owned nothing the
packaged Plan did: an
import's retained selection says where its scope's source came from, and for a
component the host *declared* that is the retained origin and the retained bytes
rather than a path, which the reader did not know. Every turn and question inside
`Plan.md` therefore belonged to nothing, so a live screen quietly kept its last
good model and a cold one refused the whole history. Both are read from the Journal — no declaration
lookup, no file read, no fallback owner — and a source this prefix never
admitted, or admitted twice, is still refused rather than attached to a guess.
The program a Plan returns has no file at all, so a position gains a second
source it can name. `SourcePosition` takes an optional `generatedSource`: the id
the fragment was admitted under, which whole-fragment preflight stamps on every
executable position in the candidate text it scans, before the fragment performs
anything. A position names one source — a path, that id, or neither for dynamic
text — and both together is malformed wherever a reader parses one. The member
travels by value through the scanner, expansion snapshots and element sites, the
durable source description, component-resolution copies and Workflow history.
That identity is what owns the work inside a fragment. The effect names its
fragment; journal order and coroutine ancestry only say whether the admission it
names could be its — one that happened afterwards, or on work the effect is not
part of, owns nothing however recently it ran, and ancestry is by coroutine
segment, so `root.1` encloses `root.1.0` and has nothing to do with `root.10`.
Two sequential fragments are sibling scopes even though their effects collide in
name, line and column; a fragment's spawned descendants belong to it; concurrent
siblings are independent of which settled first; and every way a history names no
one fragment — never admitted, admitted only afterwards, admitted twice, admitted
elsewhere, refused, unreadable, or naming no source at all — refuses whole, with
no partial model.
Evidence is `packages/cli/tests/repl-agent-journey.test.ts`, which drives
`runReplProgram()` over a real Journal, a mounted Freedom tree, the real renderer
and a terminal the suite writes bytes to. J1 walks the published Story: the real
packaged Plan, its review in the REPL's own drawer, a refused revision that
appends no answer and starts no turn, feedback resuming the same conversation,
and an approval that admits the returned source byte for byte — then the
generated program's two questions, the preview it showed, the one README it
wrote after confirmation and the decline path that writes nothing. J2 runs three
`<Spawn>` conversations and holds them complete, streaming and queued in one
frame. J3 reruns that whole journey and reopens it in a fresh command
scope — the Plan's scope, the program it returned, both retained turns, both
reviews, both of the program's own questions, the README result and the same
History positions — with zero provider calls and a byte-identical history, and
restores a failed and a cancelled turn with the text each had. C1 is the command line, and X1 is EOF, a
lost renderer, a lost terminal and a cancelled scope — each joining what it
owned, appending nothing after, and giving the modes back once.
One measured limit is reported rather than papered over: while a turn is still
queued it has no provider conversation key yet, so the Sessions surface offers
fewer selectors than there are children until it starts.
taras
force-pushed
the
agent/issue-854-journey
branch
from
October 1, 2026 00:58
b65f57c to
e22b67a
Compare
Owner
Author
|
Planner PASS is recorded for implementation head Fresh test-weight measurement is now running on that exact SHA: run 36800959452. The PR remains draft until the run succeeds, its artifact is committed byte-for-byte with provenance, the body gains |
taras
marked this pull request as ready for review
October 1, 2026 01:25
This was referenced Oct 1, 2026
taras
added a commit
that referenced
this pull request
Oct 1, 2026
`test-weights.json` was provenanced to `084397e1`, which predates the end of Slice B of #854. Nothing measured after it: Slices B, C, D and E each merged without a weights commit, four measurement attempts were cancelled along the way, and the two REPL suites #861 and #864 added — plus the Elicit form suite — had no recorded weight at all. A file without one is charged the heaviest weight the current corpus recorded, so a new test is never treated as free. Three of them were being charged `cli-npm-bin.test.ts`'s 238s each under Deno: 715s of predicted work against 31s of real work. The partition was packing shards against a number that was wrong by eleven minutes per runner. Measured on the runner, at main's exact head, with the provenance that run supplied: the commit, the run URL, the attempt, the runner label and the three runtime versions all come from the environment and none of them is a default. The file is the artifact byte for byte; no millisecond here was typed. Shard counts are unchanged and remain measured rather than chosen. The floors this measurement implies are 14 for Deno, 8 for Node and 5 for Bun, and the installed 15, 10 and 5 each satisfy their own — so nothing here forces a recalibration, which would take its own five consecutive runs on one fixed head.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #854. Slice E of five — the last layer.
Stacked on four merged slices, all on
main: #857 (A, retained truth androute), #856 (
<All>/<Spawn>), #859 (B, live Agent turns andpermission waits), #860 (C, bounded Elicit forms) and #861 (D, Sessions
and permission presentation).
Why
xmd repltook a location and nothing else. An Agent had no way to reach a REPLentry, and
<Plan>had no way to run in one — so the Story this quest is about,reviewing a packaged Plan in the REPL and running the program it returns, could
not be walked at all.
What changes
Before:
xmd replaccepted one optional location. No Agent options, no packaged Plan,and a generated fragment could write but not ask.
After:
xmd replreads the same five Agent optionsxmd runreads, settles them in afixed order, and runs the real packaged
<Plan>under one immutable profile.A returned program can ask its own questions in the drawer in front of the
person, and a cold command reconstructs the whole journey from the Journal.
How it works
One declaration for two commands. The five Agent options are declared once, so
xmd runandxmd replcannot come to mean different things by the samespelling. They settle in a fixed order — the line, the Agent configuration, one
profile, then the terminal — so everything a wrong invocation can be refused for
is refused before a per-user directory is formed, a history file exists, or the
terminal's modes are touched.
What an execution runs under is one immutable
ReplExecutionProfile,assembled once in the command's own scope: the selected Plugins, the Agent
identity components, the real packaged
<Plan>declaration, the ordinaryevaluation ceiling, and the settled permission mode.
runReplProgram()takesthat one value, so what an entry may resolve, the ceiling a generated fragment
runs under, and how a permission request is answered are facts about the command
a person invoked rather than about the entry they typed. It carries data and
installations and never authority: no stack, provider, Plan writer, live request
or scope reaches the model, the route or the Journal through it.
A generated fragment that may write may now also ask. Canonical
<Elicit>joins the ordinary write table at core's own origin, key and revision, paired
only — as a pinned capability, not a name the fragment resolves. Preflight
selects it where it finds an occurrence under a
writeselection, and the bodyit selects is core's own
<Elicit>, closed over before any document code runs,exactly as the body behind
<File>is.Nothing resolves the name: not before the root import, not at invocation, and
middleware answering a generated import with a body of its own is refused —
only canonical execution answers one. So a same-name repository, registered,
declared, middleware or separately loaded
Elicitreceives no grant and neverruns, including the workflow host's own suspension replacement, which goes on
answering the authored element in a run exactly as it did. An execution whose
fragment never writes the element resolves nothing at all, which is why
admitting the entry cannot break a document that only writes a
<File>.allow={["read"]}still refuses it, because a read selection promises nobodywill be interrupted. What stays contextual is the interaction and only the
interaction: the pinned body asks through the Elicitation Api lexically in
scope, so the REPL drawer answers a generated question and an enclosing
<Answers>region answers it without anybody being asked. Its answer isretained by the ordinary
elicitoperation, and replay restores it withoutasking anyone. The
generated_xmdrecord keeps the identity and the form andnothing about the answer.
replis a first-class Plugin command. Selected Plugins are toldreplrather than having a REPL invocation normalize to
runwithreplleft as apositional, and the bundled Git Plugin declares for it — a REPL entry is a
document run under the ordinary profile, so the Git vocabulary is available
there as it is to
run. The Plugin contract and PL4/PL17 say so.A position can name a source that has no file.
SourcePositiontakes anoptional
generatedSource— the id the fragment was admitted under, stamped onevery executable position by whole-fragment preflight before the fragment
performs anything. A position names one source: a path, that id, or neither for
dynamic text; both together is malformed wherever a reader parses one. The member
travels by value through the scanner, expansion snapshots, element sites, the
durable source description, component-resolution copies and Workflow history.
That identity is what owns work inside a fragment. The effect names its
fragment; journal order and coroutine ancestry only say whether the admission it
names could be its. Ancestry is by coroutine segment, so
root.1enclosesroot.1.0and has nothing to do withroot.10. Two sequential fragments aresibling scopes even where their effects collide in name, line and column; a
fragment's spawned descendants belong to it; concurrent siblings are independent
of which settled first. Every way a history names no one fragment — never
admitted, admitted only afterwards, admitted twice, admitted elsewhere, refused,
unreadable, or naming no source at all — refuses whole, with no partial model.
This also fixes what the model used to own for a packaged Plan: an import's
retained selection says where its scope's source came from, and for a component
the host declared, that is the retained origin and the retained bytes rather
than a path. Every turn and question inside
Plan.mdpreviously belonged tonothing, so a live screen quietly kept its last good model and a cold one refused
the whole history. Both are now read from the Journal — no declaration lookup, no
file read, no fallback owner.
Review guide
Start with:
packages/cli/tests/repl-agent-journey.test.tsThen review:
packages/cli/src/repl-profile.tsandsrc/cli.ts— what the command settles,and in what order.
packages/core/src/fragment-capabilities.ts—elicit:ask: the pinned body,and that it is the one capability whose interaction is contextual.
packages/core/src/evaluation-profile.ts—elicitWriteEntry()as a corecapability, and that
Elicitis absent from the eagerly resolved set.packages/core/src/source-position.tsandsrc/generated-xmd.ts— the secondkind of source, and that a position never names two.
packages/cli/src/repl/model.ts— retained ownership for declared sources.Look carefully at:
whole rather than producing a partial model;
EL5/EL7— that no component name is resolved for a generated question, andthat an admitted entry nothing writes resolves nothing.
How to verify it
Two frozen gates, both green at this head:
deno task check,deno task lintandgit diff --checkall exit 0.repl-agent-journey.test.tsdrivesrunReplProgram()over a real Journal, amounted Freedom tree, the real renderer and a terminal the suite writes bytes to:
REPL's own drawer, a refused revision that appends no answer and starts no
turn, feedback resuming the same conversation, an approval that admits the
returned source byte for byte — then the generated program's two questions, the
preview it showed, the one README it wrote after confirmation, and the decline
path that writes nothing.
<Spawn>conversations and holds them complete, streamingand queued in one frame.
Plan's scope, the returned program, both retained turns, both reviews, both of
the program's own questions, the README result and the same History positions —
with zero provider calls and a byte-identical history, and restores a
failed and a cancelled turn with the text each had.
a cancelled scope — each joining what it owned, appending nothing after, and
giving the terminal's modes back once.
The generated-question contract is pinned in
evaluate-component.test.tsandevaluation-profile.test.ts: EL1 the capability arm and thatElicitisabsent from the eagerly resolved set; EL2 the provider answers and the
answer writes; EL3
readrefuses before the provider or a write; EL5 asame-name definition never runs and middleware cannot answer a generated
import; EL6 replay asks nobody; EL7 an admitted entry nothing writes
resolves nothing; EL8 an enclosing
<Answers>answers it; and WGAC18that a workflow host admits it only by stating the entry itself.
Two deliberate defects hold those rows honest, each compiled and run checked:
eager-ordinary-elicit-lookup— admittingelicit:askalso resolves anordinary
Elicitowner during capture and requires it to be core's. It turnsthe repository-shadow
<File>-only case red, with the exact refusal thisdesign removed, and
PRR27red. Nothing else moves, so it reddens theboundary and nothing incidental.
EL7is not among them: a capture-timeresolution is invisible to its middleware probe, so what pins "resolves
nothing" is those two rows, while
EL7pins "breaks nothing".route-through-the-import-chain— the pinned body resolves the name atinvocation instead. It turns
EL2,EL8, bothEL5rows andEL6red.Rebase notes
This is the accepted implementation replayed onto green
main(4af3009d) — onecommit, not its thirteen-commit history. Three resolutions are worth a reviewer's
eye:
session.agentChangeswake with the subscription inside the spawn; ✨ Slice D of #854: present Agent conversations, turns and permissions in the REPL #861landed the same wake with the subscription acquired in
watch()'s enclosingscope, because a spawned body starts a turn later and a turn is long enough to
miss the first change. Git's merge kept both. Main's ownership-correct one is
what remains, carrying both comments' reasoning, and the commit message says so.
specs/repl-spec.mdhad this slice's Sessions and Permission sectionsbeside merged Slice D's section on the same subject — and this slice's text,
written against pre-correction Slice D, described a request as presented as a
drawer, which is the auto-open behavior D removed. The facts only D's section
carried (arriving opens nothing,
alwaysscoped to this Agent session, bothreadings windowed) moved into these sections, and the duplicate went.
repl-agent-interface.test.tshad a thirdrunReplProgramcall site addedby D's correction, which this slice's two-call-site conversion missed. It now
uses the profile form.
Every count change from the accepted pre-rebase reference (61 tests / 335 steps →
68 / 360) is attributable to merged work, itemized: Slice D's U4–U8 (+5 tests,
+13 steps), Slice B's P3/P6 rows (+1, +7), Slice C's F1-conditional describe and
its two new
refusalsentries (+1, +4), and therepl-modelduplicate-sequenceguard (+1 step).
Scope
Included
<Elicit>admission for generated fragments, as a pinned capability.replas a Plugin command, and bundled Git declaring for it.repl-agent-journey.test.ts(J1–J3, C1, X1) and the boundary inventory.Intentionally unchanged
test-weights.json, dependencies,deno.lockand vendor sources.Risks and limitations
still queued it has no provider conversation key, so the Sessions surface
offers fewer selectors than there are children until it starts.
mainand two REPL suites are unmeasured;the measurement at this head repairs that and is the one artifact still to come.
Scope confirmation