Quest outcome
Run repository code review as an explicit root README operation. XMD invokes
Codex through its ACPX Agent integration, and GitHub Actions presents the
result in the review job's Markdown summary instead of creating or updating a
pull-request issue comment.
A contributor can inspect the README to understand the review and can run the
same entrypoint without knowing about hidden .reviews/ documents:
./dist/xmd run --plugin ./packages/code-review-agent/mod.ts README.md#Review --approve-reads
In CI, the PR Review job is the GitHub Actions check. Its summary contains the
review report. The Quest does not create a second custom Check Run.
The bundled Git Plugin, including its GitHub adapter, is active by default in the run/workflow profile established by #822. The supported local and CI command explicitly selects the trusted code-review Plugin before importing the README; merely having that optional package or its component names in the checkout does not activate it. The GitHub adapter remains inert until a GitHub-backed operation invokes it.
Current gap
The review workflow currently invokes hidden executable documents, sends their
prompts to hosted models selected outside the README, and publishes the result
as a marked pull-request comment and separate inline review comments. Most of
the local and hosted documents are duplicated.
CI also loads the reviewer from the pull request checkout while supplying the
credentials used by the reviewer. A pull request can therefore alter the code
that reviews it before that code is trusted.
The current composition gives separate policies broad pull-request and
diagnostic objects, lets some of them invoke the model independently, and has no
documented, bounded way for the model to obtain evidence omitted from its
prompt or correct a malformed final response.
End-to-end contract
The root README visibly composes one review pipeline:
exact pull-request subject
→ deterministic environment, change and lint evidence
→ independently testable review-check components
→ one review request plus the available read-only XMD syntax
→ at most four Codex Prompt turns through ACPX
→ validated structured findings
→ rendered Markdown report
→ GitHub Actions job summary
Each review check is its own component. It receives the immutable pull-request
subject and only the evidence relevant to its question. It returns plain
structural data and receives no credential, provider handle, publication capability or mutable shared review state. A deterministic check may return its
result directly. Agent-assisted checks contribute typed requests to one logical
review so Codex can apply shared context and reconcile overlapping findings.
Separate report sections do not create separate Agent turns.
The review uses <Agent name="codex">, one named review session and ordinary
<Prompt> turns. It does not use the native interactive <Session.Launch>
surface or the standalone OpenAI Codex GitHub Action.
Before the first Prompt, the document obtains a compact syntax catalog from the
same read-only evaluation profile that will answer later information requests.
That generated catalog appears in the initial Prompt exactly once; it is not a
hand-maintained list. Codex may use bare <Syntax /> to inspect that available
surface or <Syntax names={...} /> to read selected component documentation.
Documentation never grants execution control.
Codex receives the repository policy, the applicable check requests, selected
evidence, and access to inspect the exact pull-request subject. Its response is
parsed against a fixed schema as either:
- a complete review for the supplied head commit; or
- an information request carrying a reason and a read-only XMD program.
An information request is a short circuit. It contains no provisional
findings, is not rendered as review output, and runs through <Evaluate allow={["read"]}> under the trusted host's review-read profile. The profile
includes File, Glob, Syntax and Json, exact pull-request changes, and
pull-request comments, reviews and checks. A request may compose those reads and
render only the observations it wants returned to the next Prompt. It cannot
write, execute a command, reach an undeclared network surface or obtain a
credential.
The README owns one bound of four total Prompt turns for the logical review.
Information continuations and schema-repair continuations share that bound. A
malformed response receives its schema diagnostics while another turn remains;
a refused information request receives its safe normalized reason. No fifth
Prompt starts, and a fourth response that is not a valid complete review ends
the job without publishing findings.
Only a complete, schema-valid response becomes the report. Deterministic
sensors remain evidence for the review; Codex supplies judgment rather than
replacing facts the repository can compute directly. The renderer organizes
the final result by check and does not need another Agent turn.
Review findings are advisory. A successful review with findings leaves the job
successful and makes the findings visible in its summary. Failure to collect
required evidence, obtain the syntax catalog, start or complete a Codex turn,
execute an authorized information request, obtain a valid complete review
within four turns, or publish the summary fails the job.
Trust boundary
CI executes the xmd binary, bundled default Git Plugin with its GitHub adapter, explicitly selected code-review Plugin, README entrypoint, review components, read profile, policy, prompt and response schema from the trusted base revision. The
pull request is supplied as subject data and may not replace those sources
during its own review.
The Codex session uses ACPX's approve-reads permission policy and receives no
credential capable of changing repository collaboration state. Information
requests receive only the exact component forms and host operations declared by
the review-read profile. GitHub's job summary channel needs no issue-comment or
Checks API write permission.
This is a trusted CI-agent profile, not an operating-system sandbox. The Quest does not
claim that ACPX prevents every direct filesystem or network action an agent
process could attempt. The ephemeral runner, withheld mutation credentials,
trusted reviewer source, declared permission policy and bounded XMD evaluation
profile define the delivered boundary. Stronger portable Agent enforcement
remains separate work.
Child map
Completion
The Quest is complete when:
- README help identifies the exact review target without executing it;
- the bundled Git Plugin, including its GitHub adapter, is active once, while the supported local and CI command explicitly selects the trusted code-review Plugin and invokes the README target through XMD;
- each review check owns one typed evidence-and-question boundary while one
logical Codex review judges all applicable Agent-assisted checks;
- the initial Prompt receives the catalog produced by the exact read profile,
and information requests can inspect that profile through <Syntax />;
- XMD selects Codex through ACPX and permits at most four Prompt turns shared by
information requests and schema repairs;
- only a complete review for the exact head commit passes structured response
validation and reaches the report;
- the Actions review job shows the rendered report in its Markdown summary and
creates no pull-request or inline review comment;
- findings remain advisory while incomplete review and infrastructure failures
fail the job;
- CI runs trusted reviewer sources and treats the pull request as subject data;
- the Agent and generated information requests receive no GitHub
collaboration-state write permission; and
- focused positive and negative controls distinguish each of those behaviors.
Related work
Out of scope
- Creating a separate custom Check Run or line annotations through the Checks
API.
- Making review findings block delivery.
- Launching Codex's native interactive UI.
- Claiming a general Agent, process, filesystem or network sandbox.
- Moving the repository-analysis workflow into the README.
Quest outcome
Run repository code review as an explicit root README operation. XMD invokes
Codex through its ACPX Agent integration, and GitHub Actions presents the
result in the review job's Markdown summary instead of creating or updating a
pull-request issue comment.
A contributor can inspect the README to understand the review and can run the
same entrypoint without knowing about hidden
.reviews/documents:./dist/xmd run --plugin ./packages/code-review-agent/mod.ts README.md#Review --approve-readsIn CI, the
PR Reviewjob is the GitHub Actions check. Its summary contains thereview report. The Quest does not create a second custom Check Run.
The bundled Git Plugin, including its GitHub adapter, is active by default in the run/workflow profile established by #822. The supported local and CI command explicitly selects the trusted code-review Plugin before importing the README; merely having that optional package or its component names in the checkout does not activate it. The GitHub adapter remains inert until a GitHub-backed operation invokes it.
Current gap
The review workflow currently invokes hidden executable documents, sends their
prompts to hosted models selected outside the README, and publishes the result
as a marked pull-request comment and separate inline review comments. Most of
the local and hosted documents are duplicated.
CI also loads the reviewer from the pull request checkout while supplying the
credentials used by the reviewer. A pull request can therefore alter the code
that reviews it before that code is trusted.
The current composition gives separate policies broad pull-request and
diagnostic objects, lets some of them invoke the model independently, and has no
documented, bounded way for the model to obtain evidence omitted from its
prompt or correct a malformed final response.
End-to-end contract
The root README visibly composes one review pipeline:
Each review check is its own component. It receives the immutable pull-request
subject and only the evidence relevant to its question. It returns plain
structural data and receives no credential, provider handle, publication capability or mutable shared review state. A deterministic check may return its
result directly. Agent-assisted checks contribute typed requests to one logical
review so Codex can apply shared context and reconcile overlapping findings.
Separate report sections do not create separate Agent turns.
The review uses
<Agent name="codex">, one named review session and ordinary<Prompt>turns. It does not use the native interactive<Session.Launch>surface or the standalone OpenAI Codex GitHub Action.
Before the first Prompt, the document obtains a compact syntax catalog from the
same read-only evaluation profile that will answer later information requests.
That generated catalog appears in the initial Prompt exactly once; it is not a
hand-maintained list. Codex may use bare
<Syntax />to inspect that availablesurface or
<Syntax names={...} />to read selected component documentation.Documentation never grants execution control.
Codex receives the repository policy, the applicable check requests, selected
evidence, and access to inspect the exact pull-request subject. Its response is
parsed against a fixed schema as either:
An information request is a short circuit. It contains no provisional
findings, is not rendered as review output, and runs through
<Evaluate allow={["read"]}>under the trusted host's review-read profile. The profileincludes
File,Glob,SyntaxandJson, exact pull-request changes, andpull-request comments, reviews and checks. A request may compose those reads and
render only the observations it wants returned to the next Prompt. It cannot
write, execute a command, reach an undeclared network surface or obtain a
credential.
The README owns one bound of four total Prompt turns for the logical review.
Information continuations and schema-repair continuations share that bound. A
malformed response receives its schema diagnostics while another turn remains;
a refused information request receives its safe normalized reason. No fifth
Prompt starts, and a fourth response that is not a valid complete review ends
the job without publishing findings.
Only a complete, schema-valid response becomes the report. Deterministic
sensors remain evidence for the review; Codex supplies judgment rather than
replacing facts the repository can compute directly. The renderer organizes
the final result by check and does not need another Agent turn.
Review findings are advisory. A successful review with findings leaves the job
successful and makes the findings visible in its summary. Failure to collect
required evidence, obtain the syntax catalog, start or complete a Codex turn,
execute an authorized information request, obtain a valid complete review
within four turns, or publish the summary fails the job.
Trust boundary
CI executes the
xmdbinary, bundled default Git Plugin with its GitHub adapter, explicitly selected code-review Plugin, README entrypoint, review components, read profile, policy, prompt and response schema from the trusted base revision. Thepull request is supplied as subject data and may not replace those sources
during its own review.
The Codex session uses ACPX's approve-reads permission policy and receives no
credential capable of changing repository collaboration state. Information
requests receive only the exact component forms and host operations declared by
the review-read profile. GitHub's job summary channel needs no issue-comment or
Checks API write permission.
This is a trusted CI-agent profile, not an operating-system sandbox. The Quest does not
claim that ACPX prevents every direct filesystem or network action an agent
process could attempt. The ephemeral runner, withheld mutation credentials,
trusted reviewer source, declared permission policy and bounded XMD evaluation
profile define the delivered boundary. Stronger portable Agent enforcement
remains separate work.
Child map
Completion
The Quest is complete when:
logical Codex review judges all applicable Agent-assisted checks;
and information requests can inspect that profile through
<Syntax />;information requests and schema repairs;
validation and reaches the report;
creates no pull-request or inline review comment;
fail the job;
collaboration-state write permission; and
Related work
is independent of this Quest because review starts from the current pull
request's known URL.
quality follow-ups, not prerequisites for moving the review pipeline.
xmd run#743 owns opt-in qualification of ACP Agents against their real backends; itis not part of ordinary review CI.
Out of scope
API.