Skip to content

Commit 53810b6

Browse files
committed
📝 Synchronize the living end-goal target with current XMD contracts
The adversarial implementation workflow target was written against a main that predates the error-model rules (#315), the error vocabulary rename (#317), error-model semantics (#319), and operation-scoped state (#325). It asserted `<CollectFailures>` as structural syntax, "the unresolved diagnostic", and a durable layer that "replays recorded results" — none of which describe main. This re-derives it on 7d7bdf2. The error model reaches the documents. A stage component is split by its `<Output>` boundary: the region inside runs under the `output` error mode, everything outside is documentation and runs under `throw`, which no `<PrintErrors>` region replaces. So a stage returns a complete validated result or it fails, keeping only what it had already rendered — the final `<Parse>` in each repair loop is a real gate. `throwOnError` is load-bearing for the same reason: without it a failed prompt records its failure and returns its text, raising nothing to decide. The markup did not run. Every stage passed props through expression props as `agent={props.planner}`, which fails on main with `props is not defined` — an expression prop reads the bare binding while text interpolation reads the namespace. Unifying them is #305, whose acceptance includes expression props reading `props.name`. 22 sites are corrected to the spelling main supports, and the asymmetry is recorded with the issue that removes it. Vocabulary is collapsed onto the concepts #289, #291, and #298 authorize: artifact ledger, artifact version, run identity, pinned source revision, stop reason, terminal record, stage boundary, declared inputs, and cross-process continuation, in place of the four names these files used for a ledger and the three for a run. Missing capabilities now cite the issue that supplies them rather than saying only "not implemented", and replay is described as reaching the state execution resumes from, never as the continuation itself. Planning-loop exhaustion stays open. It is recorded against #290, which pins the behavior; this change reports `verdict.passed` and does not call an exhausted loop converged. Evidence: `inspectDocument` parses all 9 frontmatters and compiles both schema kinds; `compileParseSchema` compiles all 5 embedded draft-07 schemas; `inspectComponent` resolves 21 shipped and 5 repository names and confirms 9 missing ones unresolved; `InstructionFiles` runs end to end against the repository's own AGENTS.md.
1 parent 9010557 commit 53810b6

11 files changed

Lines changed: 2135 additions & 0 deletions

‎specs/adversarial-implementation-workflow.md‎

Lines changed: 574 additions & 0 deletions
Large diffs are not rendered by default.

‎specs/markdown-agents-vision.md‎

Lines changed: 78 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -178,6 +178,80 @@ Structured distillation is preferred when later control flow depends on the
178178
result. Free-form summaries are useful context, but they are not substitutes for
179179
validated workflow state.
180180

181+
## Workflow-owned development artifacts
182+
183+
A development workflow owns its material environment and logical run state.
184+
Worktrees, working directories, captured handoffs, implementation plans,
185+
feedback, decisions, branches, and pull requests do not belong to whichever
186+
agent happened to create them. The document captures or resolves those assets
187+
deterministically and passes required content into agent prompts explicitly.
188+
189+
`<File>` and `<Glob>` already perform such operations. `<Workflow>` (#289),
190+
`<Worktree>` (#293), and `<PullRequest>` (#295) describe the rest and are not
191+
built. Together they cover the work that should not depend on model judgment:
192+
193+
- record artifact versions of handoffs, plans, reviews, and decisions;
194+
- create or resolve a named workspace, branch, file, or pull request only when
195+
that environmental asset is needed;
196+
- establish the working directory inherited by child operations;
197+
- read and write exact artifact content;
198+
- return paths, commit identities, pull-request numbers, and URLs as workflow
199+
data;
200+
- reconcile existing state when an execution resumes; and
201+
- record the inputs, observed state, effects, and outputs of each operation.
202+
203+
Agent calls analyze evidence and propose changes. Deterministic components apply
204+
approved environmental changes and provide exact required content to the next
205+
call. Generated files are optional exports rather than the handoff protocol.
206+
This removes manual copying between agent-owned transcripts, plan files, and
207+
working directories.
208+
209+
That run state is scoped to the operation that owns it: created inside the run
210+
it describes, provided contextually, and torn down with it. Nothing accumulates
211+
runs in a module-scoped registry, so concurrent runs cannot observe each other.
212+
213+
Resources clean up with their enclosing execution by default. Agent sessions,
214+
processes, streams, and other ongoing effects always stop. An execution may
215+
explicitly retain a workspace for inspection. A failed or cancelled execution
216+
also retains a workspace when cleanup would discard uncommitted or unpushed
217+
changes, and reports the path, branch, and reason that recovery is required.
218+
Durable published results such as commits, issues, and pull requests remain
219+
addressable after scoped resources close.
220+
221+
## Living workflows with `xmd play`
222+
223+
`xmd run` executes a fixed document. `xmd play` treats the document as a living
224+
collaborative workspace:
225+
226+
```sh
227+
xmd play workflow.md
228+
```
229+
230+
The executable document, rather than a hidden conversation, is the shared source
231+
of workflow intent and progress. Agents propose visible document changes or new
232+
executions. The runtime validates proposals, enforces policy, and performs
233+
deterministic effects. The user approves material changes and remains the final
234+
authority for product behavior, scope, architecture, risk, and lasting
235+
constraints.
236+
237+
An accepted proposal becomes an inspectable document revision. Rejected
238+
proposals, failed executions, reviewer rejections, and later successful attempts
239+
retain their provenance so the engineering history explains how the workflow
240+
changed. Hidden session history may help an agent reason, but it is never the
241+
only source of consequential workflow state.
242+
243+
Named agent sessions remain scope-owned while Play is active. Each invocation
244+
receives explicit workflow context and references to workflow-owned artifacts.
245+
The document and execution record identify what each agent received, what it
246+
proposed, what the runtime applied, and which user decision authorized a
247+
material transition.
248+
249+
Play rests on the same deterministic asset and agent orchestration needed by an
250+
automated implementation loop. The loop is the proving ground for worktree,
251+
file, pull-request, review, decision, cleanup, and recovery semantics. Play adds
252+
collaborative document evolution after those operations are reliable; it does
253+
not replace them with agent-managed shell work.
254+
181255
## Foundation and agent layer
182256

183257
Executable.md separates two concerns:
@@ -222,6 +296,10 @@ its declared props:
222296
6. What happens when it fails or returns invalid output?
223297
7. Why did the workflow take a branch or stop?
224298
8. What evidence in the execution record supports those answers?
299+
9. Which environmental assets did the workflow create or resolve, and who owns
300+
their cleanup or retention?
301+
10. In Play, what document change was proposed, what effects were validated, and
302+
which user decision accepted it?
225303

226304
If those answers depend on hidden host behavior, implicit transcript sharing, or
227305
an agent's own account of what it did, the design does not satisfy the product
Lines changed: 63 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,63 @@
1+
---
2+
required: [instructions, planner]
3+
4+
props:
5+
instructions: { type: string }
6+
planner: { type: string }
7+
---
8+
9+
# Discovery
10+
11+
The workflow enters through design discovery or a bounded deferred issue.
12+
Discovery includes a user-planner interview. A sufficiently specified deferred
13+
issue may enter directly at implementor planning.
14+
15+
## Target shape
16+
17+
<Agent name={planner}>
18+
<Session name="planner">
19+
<Prompt as="handoff" throwOnError>
20+
Repository instructions:
21+
22+
{props.instructions}
23+
24+
User request:
25+
26+
<Content />
27+
28+
Produce a user-validated design handoff and a falsifiable implementation
29+
theory. Distinguish user decisions from hypotheses the implementor must
30+
test.
31+
</Prompt>
32+
</Session>
33+
</Agent>
34+
35+
<Output>{handoff}</Output>
36+
37+
## Handoff contents
38+
39+
- Purpose and observable behavior
40+
- User decisions, constraints, non-goals, and accepted risks
41+
- Repository and architectural context
42+
- Falsifiable implementation theory
43+
- Assumptions to confirm or refute
44+
- Required evidence and validation
45+
- Likely pull-request topology
46+
- Decisions that remain with the user
47+
48+
The component declares no `returns`, so its `<Output>` region is its return
49+
value and a caller's `as` binds that rendered text. `<Agent name={planner}>`
50+
selects the agent from a validated prop rather than a literal: the agent
51+
components take their props from a literal or from an expression that resolves
52+
to a string. An expression prop reads the bare binding, while the prompt body
53+
above interpolates `{props.instructions}` — the two spellings that #305 will
54+
unify. Its caller supplies the request as content and decides whether and where
55+
to persist the result. The handoff is a theory for investigation, not an
56+
implementation plan that the implementor follows unquestioningly.
57+
58+
The prompt sits outside `<Output>`, so it runs under the `throw` error mode:
59+
`throwOnError` turns a failed prompt into a failure the mode then ends the
60+
stage on. Without it a failed prompt records its failure and returns its text,
61+
and the stage would hand its caller an empty handoff.
62+
63+
This component runs today.

0 commit comments

Comments
 (0)