Skip to content

✨ Slice E of #854: run the packaged Plan in xmd repl, under one immutable profile - #864

Merged
taras merged 1 commit into
mainfrom
agent/issue-854-journey
Oct 1, 2026
Merged

taras merged 1 commit into
mainfrom
agent/issue-854-journey

Conversation

@taras

@taras taras commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Part of #854. Slice E of five — the last layer.

Stacked on four merged slices, all on main: #857 (A, retained truth and
route), #856 (<All>/<Spawn>), #859 (B, live Agent turns and
permission waits), #860 (C, bounded Elicit forms) and #861 (D, Sessions
and permission presentation).

Draft: test-weights.json is outstanding, and this PR does not yet carry
Closes #854. Weights are measured against this exact head once the
implementation is accepted; the closing keyword and the artifact arrive
together. The previous Slice E measurement (36568888172) failed in
pull-request-read-fork.test.ts and wrote nothing, so it is stale regardless.

Why

xmd repl took a location and nothing else. An Agent had no way to reach a REPL
entry, and <Plan> had no way to run in one — so the Story this quest is about,
reviewing a packaged Plan in the REPL and running the program it returns, could
not be walked at all.

What changes

Before:

  • xmd repl accepted one optional location. No Agent options, no packaged Plan,
    and a generated fragment could write but not ask.

After:

  • xmd repl reads the same five Agent options xmd run reads, settles them in a
    fixed order, and runs the real packaged <Plan> under one immutable profile.
    A returned program can ask its own questions in the drawer in front of the
    person, and a cold command reconstructs the whole journey from the Journal.

How it works

One declaration for two commands. The five Agent options are declared once, so
xmd run and xmd repl cannot come to mean different things by the same
spelling. They settle in a fixed order — the line, the Agent configuration, one
profile, then the terminal — so everything a wrong invocation can be refused for
is refused before a per-user directory is formed, a history file exists, or the
terminal's modes are touched.

What an execution runs under is one immutable ReplExecutionProfile,
assembled once in the command's own scope: the selected Plugins, the Agent
identity components, the real packaged <Plan> declaration, the ordinary
evaluation ceiling, and the settled permission mode. runReplProgram() takes
that one value, so what an entry may resolve, the ceiling a generated fragment
runs under, and how a permission request is answered are facts about the command
a person invoked rather than about the entry they typed. It carries data and
installations and never authority: no stack, provider, Plan writer, live request
or scope reaches the model, the route or the Journal through it.

A generated fragment that may write may now also ask. Canonical <Elicit>
joins the ordinary write table at core's own origin, key and revision, paired
only — as a pinned capability, not a name the fragment resolves. Preflight
selects it where it finds an occurrence under a write selection, and the body
it selects is core's own <Elicit>, closed over before any document code runs,
exactly as the body behind <File> is.

Nothing resolves the name: not before the root import, not at invocation, and
middleware answering a generated import with a body of its own is refused —
only canonical execution answers one. So a same-name repository, registered,
declared, middleware or separately loaded Elicit receives no grant and never
runs, including the workflow host's own suspension replacement, which goes on
answering the authored element in a run exactly as it did. An execution whose
fragment never writes the element resolves nothing at all, which is why
admitting the entry cannot break a document that only writes a <File>.

allow={["read"]} still refuses it, because a read selection promises nobody
will be interrupted. What stays contextual is the interaction and only the
interaction: the pinned body asks through the Elicitation Api lexically in
scope, so the REPL drawer answers a generated question and an enclosing
<Answers> region answers it without anybody being asked. Its answer is
retained by the ordinary elicit operation, and replay restores it without
asking anyone. The generated_xmd record keeps the identity and the form and
nothing about the answer.

repl is a first-class Plugin command. Selected Plugins are told repl
rather than having a REPL invocation normalize to run with repl left as a
positional, and the bundled Git Plugin declares for it — a REPL entry is a
document run under the ordinary profile, so the Git vocabulary is available
there as it is to run. The Plugin contract and PL4/PL17 say so.

A position can name a source that has no file. SourcePosition takes an
optional generatedSource — the id the fragment was admitted under, stamped on
every executable position by whole-fragment preflight before the fragment
performs anything. A position names one source: a path, that id, or neither for
dynamic text; both together is malformed wherever a reader parses one. The member
travels by value through the scanner, expansion snapshots, element sites, the
durable source description, component-resolution copies and Workflow history.

That identity is what owns work inside a fragment. The effect names its
fragment; journal order and coroutine ancestry only say whether the admission it
names could be its. Ancestry is by coroutine segment, so root.1 encloses
root.1.0 and has nothing to do with root.10. Two sequential fragments are
sibling scopes even where their effects collide in name, line and column; a
fragment's spawned descendants belong to it; concurrent siblings are independent
of which settled first. Every way a history names no one fragment — never
admitted, admitted only afterwards, admitted twice, admitted elsewhere, refused,
unreadable, or naming no source at all — refuses whole, with no partial model.

This also fixes what the model used to own for a packaged Plan: an import's
retained selection says where its scope's source came from, and for a component
the host declared, that is the retained origin and the retained bytes rather
than a path. Every turn and question inside Plan.md previously belonged to
nothing, so a live screen quietly kept its last good model and a cold one refused
the whole history. Both are now read from the Journal — no declaration lookup, no
file read, no fallback owner.

Review guide

Start with: packages/cli/tests/repl-agent-journey.test.ts

Then review:

  1. packages/cli/src/repl-profile.ts and src/cli.ts — what the command settles,
    and in what order.
  2. packages/core/src/fragment-capabilities.ts — elicit:ask: the pinned body,
    and that it is the one capability whose interaction is contextual.
  3. packages/core/src/evaluation-profile.ts — elicitWriteEntry() as a core
    capability, and that Elicit is absent from the eagerly resolved set.
  4. packages/core/src/source-position.ts and src/generated-xmd.ts — the second
    kind of source, and that a position never names two.
  5. packages/cli/src/repl/model.ts — retained ownership for declared sources.

Look carefully at:

  • that no authority travels on the profile;
  • the refusal paths for a history that names no one fragment — each refuses
    whole rather than producing a partial model;
  • EL5/EL7 — that no component name is resolved for a generated question, and
    that an admitted entry nothing writes resolves nothing.

How to verify it

Two frozen gates, both green at this head:

deno task test \
  packages/core/tests/source-position.test.ts \
  packages/core/tests/journal-source-position.test.ts \
  packages/core/tests/expansion-identity.test.ts \
  packages/core/tests/generated-xmd.test.ts \
  packages/core/tests/evaluation-profile.test.ts \
  packages/core/tests/evaluate-component.test.ts \
  packages/cli/tests/plan-command-document.test.ts \
  packages/workflow/tests/workflow-lifecycle-inspection.test.ts
# 46 passed (378 steps)

deno task test \
  packages/cli/tests/repl-agent-execution.test.ts \
  packages/cli/tests/repl-agent-interface.test.ts \
  packages/cli/tests/repl-agent-journey.test.ts \
  packages/cli/tests/repl-forms.test.ts \
  packages/cli/tests/repl-model.test.ts \
  packages/cli/tests/repl-route.test.ts \
  packages/cli/tests/repl-execution.test.ts \
  packages/cli/tests/repl-composition.test.ts \
  packages/cli/tests/repl-terminal.test.ts \
  packages/cli/tests/repl-journey.test.ts \
  packages/cli/tests/repl-boundaries.test.ts \
  packages/cli/tests/cli-help.test.ts
# 68 passed (360 steps)

deno task check, deno task lint and git diff --check all exit 0.

repl-agent-journey.test.ts drives runReplProgram() over a real Journal, a
mounted Freedom tree, the real renderer and a terminal the suite writes bytes to:

  • J1 walks the published Story: the real packaged Plan, its review in the
    REPL's own drawer, a refused revision that appends no answer and starts no
    turn, feedback resuming the same conversation, an approval that admits the
    returned source byte for byte — then the generated program's two questions, the
    preview it showed, the one README it wrote after confirmation, and the decline
    path that writes nothing.
  • J2 runs three <Spawn> conversations and holds them complete, streaming
    and queued in one frame.
  • J3 reruns the whole journey and reopens it in a fresh command scope — the
    Plan's scope, the returned program, both retained turns, both reviews, both of
    the program's own questions, the README result and the same History positions —
    with zero provider calls and a byte-identical history, and restores a
    failed and a cancelled turn with the text each had.
  • C1 is the command line; X1 is EOF, a lost renderer, a lost terminal and
    a cancelled scope — each joining what it owned, appending nothing after, and
    giving the terminal's modes back once.

The generated-question contract is pinned in evaluate-component.test.ts and
evaluation-profile.test.ts: EL1 the capability arm and that Elicit is
absent from the eagerly resolved set; EL2 the provider answers and the
answer writes; EL3 read refuses before the provider or a write; EL5 a
same-name definition never runs and middleware cannot answer a generated
import; EL6 replay asks nobody; EL7 an admitted entry nothing writes
resolves nothing; EL8 an enclosing <Answers> answers it; and WGAC18
that a workflow host admits it only by stating the entry itself.

Two deliberate defects hold those rows honest, each compiled and run checked:

  • eager-ordinary-elicit-lookup — admitting elicit:ask also resolves an
    ordinary Elicit owner during capture and requires it to be core's. It turns
    the repository-shadow <File>-only case red, with the exact refusal this
    design removed, and PRR27 red. Nothing else moves, so it reddens the
    boundary and nothing incidental. EL7 is not among them: a capture-time
    resolution is invisible to its middleware probe, so what pins "resolves
    nothing" is those two rows, while EL7 pins "breaks nothing".
  • route-through-the-import-chain — the pinned body resolves the name at
    invocation instead. It turns EL2, EL8, both EL5 rows and EL6 red.

Rebase notes

This is the accepted implementation replayed onto green main (4af3009d) — one
commit, not its thirteen-commit history. Three resolutions are worth a reviewer's
eye:

  • The Agent wake belongs to Slice D. This slice originally added its own
    session.agentChanges wake with the subscription inside the spawn; ✨ Slice D of #854: present Agent conversations, turns and permissions in the REPL #861
    landed the same wake with the subscription acquired in watch()'s enclosing
    scope, because a spawned body starts a turn later and a turn is long enough to
    miss the first change. Git's merge kept both. Main's ownership-correct one is
    what remains, carrying both comments' reasoning, and the commit message says so.
  • specs/repl-spec.md had this slice's Sessions and Permission sections
    beside merged Slice D's section on the same subject — and this slice's text,
    written against pre-correction Slice D, described a request as presented as a
    drawer, which is the auto-open behavior D removed. The facts only D's section
    carried (arriving opens nothing, always scoped to this Agent session, both
    readings windowed) moved into these sections, and the duplicate went.
  • repl-agent-interface.test.ts had a third runReplProgram call site added
    by D's correction, which this slice's two-call-site conversion missed. It now
    uses the profile form.

Every count change from the accepted pre-rebase reference (61 tests / 335 steps →
68 / 360) is attributable to merged work, itemized: Slice D's U4–U8 (+5 tests,
+13 steps), Slice B's P3/P6 rows (+1, +7), Slice C's F1-conditional describe and
its two new refusals entries (+1, +4), and the repl-model duplicate-sequence
guard (+1 step).

Scope

Included

  • The REPL command line, the immutable execution profile, and the packaged Plan.
  • Canonical <Elicit> admission for generated fragments, as a pinned capability.
  • repl as a Plugin command, and bundled Git declaring for it.
  • Generated-source identity and the ownership it gives retained work.
  • repl-agent-journey.test.ts (J1–J3, C1, X1) and the boundary inventory.

Intentionally unchanged

  • Everything Slices A–D delivered, including every merged Slice D correction.
  • test-weights.json, dependencies, deno.lock and vendor sources.

Risks and limitations

  • One measured limit, reported rather than papered over: while a turn is
    still queued it has no provider conversation key, so the Sessions surface
    offers fewer selectors than there are children until it starts.
  • Weights are three slices stale on main and two REPL suites are unmeasured;
    the measurement at this head repairs that and is the one artifact still to come.

Scope confirmation

  • Every changed file supports the purpose described above.
  • Unrelated cleanup and formatting changes are excluded.
  • Generated or mechanical changes are clearly identified.
  • The description matches the final diff and test results.

@taras
taras force-pushed the agent/issue-854-journey branch from 47fb629 to b65f57c Compare October 1, 2026 00:46
`xmd repl` took a location and nothing else, so an Agent had no way to reach a
REPL entry and `<Plan>` had no way to run in one. It now reads the same five
Agent options `xmd run` reads, as a first-class Plugin command of its own —
selected Plugins are told `repl`, and the bundled Git Plugin declares for it,
because a REPL entry is a document run under the ordinary profile — one declaration, so the two commands cannot come
to mean different things by the same spelling — and settles them in a fixed
order: the line, then the Agent configuration, then one profile, then the
terminal. Everything a wrong invocation can be refused for is refused before a
per-user directory is formed, a history file exists or the terminal's modes are
touched.

What a REPL execution runs under is one immutable `ReplExecutionProfile`,
assembled once in the command's own scope: the selected Plugins, the Agent
identity components, the real packaged `<Plan>` declaration and the ordinary
evaluation ceiling, with the settled permission mode beside them. The program
takes that one value and nothing else, so what an entry may resolve, what ceiling
a generated fragment runs under and how a permission request is answered are
facts about the command a person invoked rather than about the entry they typed.
It carries data and installations and never authority: no stack, provider, Plan
writer, live request or scope reaches the model, the route or the Journal through
it. The REPL installs no readline permission handling, no foreground launcher and
no browser form — a question is answered in the drawer in front of the person,
and `<Session.Launch>` refuses through the established missing-launcher contract
because this command is using the terminal it would be given.

A returned Plan is a generated fragment, and a fragment that may write may now
also ask: canonical `<Elicit>` joins the ordinary write table at core's own
origin, key and revision, paired only, as a *pinned capability* rather than a
name the fragment resolves. Preflight selects it where it finds an occurrence
under a `write` selection, and the body it selects is core's own `<Elicit>`,
closed over before any document code runs — exactly as the body behind `<File>`
is. Nothing resolves the name: not before the root import, not at invocation,
and middleware answering a generated import with a body of its own is refused,
because only canonical execution answers one. So a same-name repository,
registered, declared, middleware or separately loaded `Elicit` receives no grant
and never runs, including the workflow host's own suspension replacement, which
goes on answering the *authored* element in a run exactly as it did. An
execution whose fragment never writes the element resolves nothing at all, which
is why admitting the entry cannot break a document that only writes a `<File>`.

`allow={["read"]}` still refuses it, because a read selection promises nobody
will be interrupted. What stays contextual is the interaction and only the
interaction: the pinned body asks through the Elicitation Api lexically in scope,
so the REPL drawer answers a generated question and an enclosing `<Answers>`
region answers it without anybody being asked. Its answer is retained by the
ordinary `elicit` operation, and replay restores it without asking anyone. The
`generated_xmd` record keeps the identity and the form and nothing about the
answer.

One defect the journey found is fixed here. The model owned nothing the
packaged Plan did: an
import's retained selection says where its scope's source came from, and for a
component the host *declared* that is the retained origin and the retained bytes
rather than a path, which the reader did not know. Every turn and question inside
`Plan.md` therefore belonged to nothing, so a live screen quietly kept its last
good model and a cold one refused the whole history. Both are read from the Journal — no declaration
lookup, no file read, no fallback owner — and a source this prefix never
admitted, or admitted twice, is still refused rather than attached to a guess.

The program a Plan returns has no file at all, so a position gains a second
source it can name. `SourcePosition` takes an optional `generatedSource`: the id
the fragment was admitted under, which whole-fragment preflight stamps on every
executable position in the candidate text it scans, before the fragment performs
anything. A position names one source — a path, that id, or neither for dynamic
text — and both together is malformed wherever a reader parses one. The member
travels by value through the scanner, expansion snapshots and element sites, the
durable source description, component-resolution copies and Workflow history.

That identity is what owns the work inside a fragment. The effect names its
fragment; journal order and coroutine ancestry only say whether the admission it
names could be its — one that happened afterwards, or on work the effect is not
part of, owns nothing however recently it ran, and ancestry is by coroutine
segment, so `root.1` encloses `root.1.0` and has nothing to do with `root.10`.
Two sequential fragments are sibling scopes even though their effects collide in
name, line and column; a fragment's spawned descendants belong to it; concurrent
siblings are independent of which settled first; and every way a history names no
one fragment — never admitted, admitted only afterwards, admitted twice, admitted
elsewhere, refused, unreadable, or naming no source at all — refuses whole, with
no partial model.

Evidence is `packages/cli/tests/repl-agent-journey.test.ts`, which drives
`runReplProgram()` over a real Journal, a mounted Freedom tree, the real renderer
and a terminal the suite writes bytes to. J1 walks the published Story: the real
packaged Plan, its review in the REPL's own drawer, a refused revision that
appends no answer and starts no turn, feedback resuming the same conversation,
and an approval that admits the returned source byte for byte — then the
generated program's two questions, the preview it showed, the one README it
wrote after confirmation and the decline path that writes nothing. J2 runs three
`<Spawn>` conversations and holds them complete, streaming and queued in one
frame. J3 reruns that whole journey and reopens it in a fresh command
scope — the Plan's scope, the program it returned, both retained turns, both
reviews, both of the program's own questions, the README result and the same
History positions — with zero provider calls and a byte-identical history, and
restores a failed and a cancelled turn with the text each had. C1 is the command line, and X1 is EOF, a
lost renderer, a lost terminal and a cancelled scope — each joining what it
owned, appending nothing after, and giving the modes back once.

One measured limit is reported rather than papered over: while a turn is still
queued it has no provider conversation key yet, so the Sessions surface offers
fewer selectors than there are children until it starts.
@taras
taras force-pushed the agent/issue-854-journey branch from b65f57c to e22b67a Compare October 1, 2026 00:58
@taras

taras commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Planner PASS is recorded for implementation head e22b67a893b264a33df17ab32537673412a669f4.

Fresh test-weight measurement is now running on that exact SHA: run 36800959452.

The PR remains draft until the run succeeds, its artifact is committed byte-for-byte with provenance, the body gains Closes #854, and CI passes on the weights commit.

@taras
taras marked this pull request as ready for review October 1, 2026 01:25
@taras
taras merged commit f7a1063 into main Oct 1, 2026
44 of 45 checks passed
@taras
taras deleted the agent/issue-854-journey branch October 1, 2026 01:25
taras added a commit that referenced this pull request Oct 1, 2026
`test-weights.json` was provenanced to `084397e1`, which predates the end of
Slice B of #854. Nothing measured after it: Slices B, C, D and E each merged
without a weights commit, four measurement attempts were cancelled along the
way, and the two REPL suites #861 and #864 added — plus the Elicit form suite —
had no recorded weight at all.

A file without one is charged the heaviest weight the current corpus recorded,
so a new test is never treated as free. Three of them were being charged
`cli-npm-bin.test.ts`'s 238s each under Deno: 715s of predicted work against
31s of real work. The partition was packing shards against a number that was
wrong by eleven minutes per runner.

Measured on the runner, at main's exact head, with the provenance that run
supplied: the commit, the run URL, the attempt, the runner label and the three
runtime versions all come from the environment and none of them is a default.
The file is the artifact byte for byte; no millisecond here was typed.

Shard counts are unchanged and remain measured rather than chosen. The floors
this measurement implies are 14 for Deno, 8 for Node and 5 for Bun, and the
installed 15, 10 and 5 each satisfy their own — so nothing here forces a
recalibration, which would take its own five consecutive runs on one fixed head.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant