From 457c74e38671343d89f934d74b17c7215327b317 Mon Sep 17 00:00:00 2001 From: Francis Eytan Dortort Date: Sat, 8 Aug 2026 18:07:09 -0400 Subject: [PATCH 1/3] docs(readme): reflect the shipped Codex/Gemini/Antigravity drivers Status, prerequisites, the What-it-is intro, the design-decisions substrate note, the project-layout tree, and the agent-selection note now cover all four drivers; drops Codex/Gemini from the roadmap backlog (Cursor remains). Co-Authored-By: Claude Opus 4.8 --- README.md | 23 ++++++++++++++--------- 1 file changed, 14 insertions(+), 9 deletions(-) diff --git a/README.md b/README.md index 1ee5df6..cc95348 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,7 @@ > **Playwright for AI Agents** — end-to-end testing for the infrastructure AI agents interact with. -**Status:** Early MVP — Claude driver, structural assertions, record/replay. No published npm package yet. See [Roadmap](#roadmap). +**Status:** Early MVP — Claude, Codex, Gemini, and Antigravity drivers; structural assertions; record/replay. No published npm package yet. See [Roadmap](#roadmap). --- @@ -14,7 +14,7 @@ Playwright drives a real browser and asserts on the DOM. Agentry drives a real A You write TypeScript tests that: -1. **Drive a real agent CLI** (Claude Code, headless) with a prompt and a sandbox workspace. +1. **Drive a real agent CLI** (Claude Code, Codex, Gemini, or Antigravity — all headless) with a prompt and a sandbox workspace. 2. **Observe a normalized event stream** — assistant turns, tool calls, MCP requests, filesystem side-effects, token usage. 3. **Assert on structure, not text** — `toHaveToolCall`, `toHaveFile`, `toFinishWithin`. Exact matchers on what the agent did; no fragile string matching on free-form output. 4. **Replay deterministically** in CI via recorded transcripts — fast (<2 s), free, no agent required. Record once live; replay forever. @@ -27,7 +27,7 @@ The target is the agent's surrounding ecosystem: skills, MCP gateways, plugins There is no published npm package yet. Run from this monorepo directly. -**Prerequisites:** Node >= 20, pnpm >= 10, `claude` CLI in your PATH, `ANTHROPIC_API_KEY` set. +**Prerequisites:** Node >= 20, pnpm >= 10, and the CLI for whichever driver you use in your PATH — `claude` (with `ANTHROPIC_API_KEY`), `codex`, `gemini`, or `agy`, each authenticated per its own vendor. Replay-only runs need none of these. ```bash git clone @@ -68,6 +68,8 @@ export default defineConfig({ }); ``` +`use.agent` selects the driver — `'claude'` (default), `'codex'`, `'gemini'`, or `'antigravity'` — and `use.model` must be a pinned snapshot id for that agent. + ### 2. Test: `examples/basic/tests/todo.agentry.ts` ```ts @@ -214,7 +216,7 @@ Test file → runner → AgentHandle.run(prompt) Key design decisions: -- **Structured/headless substrate.** The Claude driver uses `claude -p --output-format stream-json --verbose` — no TTY, no screen scraping. The JSONL stream is the reliable source of truth (the CDP analog). +- **Structured/headless substrate.** Every driver uses its CLI's structured headless mode — `claude -p --output-format stream-json --verbose`, `codex exec --json`, `gemini -p --output-format stream-json`, `agy -p --output-format stream-json` — no TTY, no screen scraping. The JSONL stream is the reliable source of truth (the CDP analog). - **Normalized event model.** Every driver output is translated into the same `AgentEvent` stream (`tool_use`, `tool_result`, `mcp_request`, `message`, `usage`, `run.end`, …). Assertions read only this model; drivers are pluggable. - **Sandbox isolation.** Each scenario runs in a fresh temp directory. File side-effects are detected by content-hash snapshotting before/after the run, producing `fs` events. - **Budget cap.** `budget.perTest.usd` is enforced post-run in the MVP; `--max-budget-usd` is passed to the Claude CLI as a native backstop. @@ -229,10 +231,13 @@ See [docs/SPEC.md](./docs/SPEC.md) for the full design. ``` agentry/ ├── packages/ -│ ├── core/ @agentry/core — events, config, sandbox, runner, assertions, transcript -│ ├── claude/ @agentry/claude — Claude Code driver -│ ├── mcp/ @agentry/mcp — MockMcpServer + MCP matchers -│ └── cli/ agentry — CLI binary (test, record, init, doctor) +│ ├── core/ @agentry/core — events, config, sandbox, runner, assertions, transcript +│ ├── claude/ @agentry/claude — Claude Code driver +│ ├── codex/ @agentry/codex — Codex CLI driver +│ ├── gemini/ @agentry/gemini — Gemini CLI driver +│ ├── antigravity/ @agentry/antigravity — Antigravity (agy) driver +│ ├── mcp/ @agentry/mcp — MockMcpServer + MCP matchers +│ └── cli/ agentry — CLI binary (test, record, init, doctor) ├── examples/ │ └── basic/ Working example with committed transcript └── docs/ @@ -267,7 +272,7 @@ The SPEC describes the full vision. What is implemented vs. planned: | Skills/plugins CH6 **differential** harness (`toChangeBehaviorVs`) | Phase 2 | | MCP **live**: agent→mock fixture wiring + protocol-compliance matchers (`toBeValidMcpProtocol`, error codes) | Phase 3 | | LLM **gateway** matchers (routing/fallback/cache/rate-limit) + provider-pool resolver (foundation shipped) | Phase 4 | -| Codex, Gemini, Cursor drivers | Phase 5 (spike-gated) | +| Cursor driver | Phase 5 (spike-gated) | | Semantic/LLM-as-judge assertions (Tier 4: `toSatisfyRubric`) | Phase 6 | | HTML/JUnit/JSON reporters; trace bundles | Phase 6 | | HOME-remap / container sandbox; commit wire cassettes (host config leakage today) | Phase 7+ | From 0fe3550bf3ac46d7f80a56b9ac6a161d57d794be Mon Sep 17 00:00:00 2001 From: Francis Eytan Dortort Date: Sat, 8 Aug 2026 18:08:15 -0400 Subject: [PATCH 2/3] docs(roadmap): mark the Codex/Gemini/Antigravity spikes done and reframe Phase 5 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phase 0 cross-agent spikes 0.11 (Codex) and 0.12 (Gemini) are shipped, plus a new 0.14 (Antigravity); Cursor (0.13) is the only remaining adapter. Updates the agents summary, Phase 5, and the status checklist (incl. the 79→108 test count). Co-Authored-By: Claude Opus 4.8 --- docs/ROADMAP.md | 27 +++++++++++++++++---------- 1 file changed, 17 insertions(+), 10 deletions(-) diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 2def59e..cf67ea4 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -16,7 +16,7 @@ architecture, and the decisions still open. **Revised v2** after two pre-impleme | Determinism | Record/replay cassettes (ordered, session-positional); **replay default**, live periodic | | Authoring | Code-first TypeScript v1; declarative YAML deferred | | **Target types in v1** | **All four — skills, plugins, MCP gateways, LLM gateways** (skills & plugins = one workstream → "three workstreams"). **Committed GA.** | -| **Agents** | **Claude** in MVP; **Codex + Gemini + Cursor** committed v1 objectives, spike-gated fast-follow | +| **Agents** | **Claude, Codex, Gemini, and Antigravity** drivers shipped; **Cursor** remains a spike-gated fast-follow | | North star | Follow Playwright; diverge only for LLM non-determinism (3 divergences — SPEC §1.4) | --- @@ -79,9 +79,10 @@ these are answered and written up in `docs/research/phase0-findings.md` (local). ### Cross-agent feasibility (per SPEC §4.3 matrix — gate each adapter) | # | Question | Why it matters | |---|---|---| -| 0.11 | **Codex:** `codex exec --json`; LLM proxy via `model_providers..base_url` under `CODEX_HOME`; MCP support | Codex adapter feasibility | -| 0.12 | **Gemini:** structured stream; base-URL interception (unproven); MCP config | Gemini adapter feasibility | +| 0.11 | **Codex:** `codex exec --json`; LLM proxy via `model_providers..base_url` under `CODEX_HOME`; MCP support | Codex adapter feasibility — **✅ shipped** (`@agentry/codex`; stream verified against codex-cli 0.147; interception is `provider-config`, not wired to the Anthropic proxy) | +| 0.12 | **Gemini:** structured stream; base-URL interception (unproven); MCP config | Gemini adapter feasibility — **✅ shipped** (`@agentry/gemini`; `--output-format stream-json` verified against gemini-cli 0.54; assistant text coalesced from deltas; base-URL interception still unproven → `none`) | | 0.13 | **Cursor (highest risk):** does `cursor-agent` run headless/structured at all (probing previously hung)? LLM/MCP interception? | Cursor adapter feasibility — may demote to fast-follow if it fails | +| 0.14 | **Antigravity (`agy`):** `agy -p --output-format stream-json`; `event`-discriminated stream; permissions via `--dangerously-skip-permissions` | Antigravity adapter feasibility — **✅ shipped** (`@agentry/antigravity`; usage + `run.end` from the terminal `result`; sandbox fs-diff best-effort since agy uses its own scratch dir) | **Exit criteria:** yes/no + evidence per spike; architecture deltas folded into `SPEC.md`. A failed cross-agent spike (esp. 0.13) demotes that agent to post-v1 without affecting the Claude GA path. @@ -121,8 +122,10 @@ Routing, fallback, cache, rate-limit, transformation, token-accuracy, latency-ov propagation matchers over the existing LLM interceptor. ### Phase 5 — Cross-agent adapters (committed v1 objective, spike-gated) -Codex, Gemini, Cursor drivers against the proven interface; per-agent capability gating; shared -suites run across the matrix. Each gated on its Phase 0 spike (0.11–0.13). +The **Codex, Gemini, and Antigravity** drivers ship against the proven interface (`@agentry/codex`, +`@agentry/gemini`, `@agentry/antigravity`), each mirroring the Claude reference (pure parser + +`buildArgs` + honest `capabilities()`), with per-agent capability gating. **Cursor** remains, gated on +spike 0.13; shared suites run across the full matrix once it lands. ### Phase 6 — Tier-3/4 assertions + JUnit/JSON/HTML reporters + parallelism hardening Structured-output (schema/AST) + LLM-as-judge (recorded for free replay, §8.5); rate-limit-aware @@ -201,11 +204,12 @@ firm). Still open: - [x] Spec drafted + revised post-review (`SPEC.md` v2) - [x] Roadmap drafted + revised (this doc v2) - [x] Independent review pass (Claude critic + Codex; in `docs/research/`) -- [~] Phase 0 spikes — **in progress**: 0.1 / 0.2 / 0.4 ✅, 0.6 / 0.7 🟡, config isolation ✅; +- [~] Phase 0 spikes — **in progress**: 0.1 / 0.2 / 0.4 ✅, 0.6 / 0.7 🟡, config isolation ✅, + cross-agent Codex (0.11) / Gemini (0.12) / Antigravity (0.14) ✅; remaining: MCP shim (0.3), HOME-remap (0.5), cassette canonicalization (0.9), mcp-live (0.10), - cross-agent (0.11–0.13) + Cursor (0.13) - [x] Repo shape = monorepo; Phase 0 ownership = Agentry-runs (§7) -- [~] **Implementation — MVP spine + LLM-proxy phase SHIPPED** (79 unit tests, CI green): +- [~] **Implementation — MVP spine + LLM-proxy phase + cross-agent drivers SHIPPED** (108 unit tests, CI green): - [x] **Phase 1 (spine):** monorepo · event model · RunRecord · Sandbox (fs diff) · config (model-pin) · assertion engine (Tiers 1–3) · cassette engine · Claude driver (live) · transcript record/replay · runner + fixtures · console reporter · CLI (init/test/record/doctor). MVP §5 @@ -221,8 +225,11 @@ firm). Still open: - [~] **Phase 4 (LLM gateways):** foundation laid (observable `llm_request`/`llm_response` + positionable proxy + provider-impersonation seam, SPEC §9.3); remaining: provider-pool resolver + routing/fallback/cache matchers. - - [ ] **Phase 5** (Codex/Gemini/Cursor) · **Phase 6** (Tier 4 LLM-as-judge, HTML/JUnit reporters) - · **Phase 7+** (trace viewer, codegen, PTY driver) — not started. + - [~] **Phase 5 (cross-agent adapters):** Claude + **Codex + Gemini + Antigravity** drivers shipped + (each mirrors the reference: pure parser + `buildArgs` + `capabilities()`); `agentry doctor` probes + all four CLIs. Remaining: **Cursor** (spike 0.13). + - [ ] **Phase 6** (Tier 4 LLM-as-judge, HTML/JUnit reporters) · **Phase 7+** (trace viewer, codegen, + PTY driver) — not started. *Known gaps:* wire cassettes capture host config until HOME-remap sandboxing (spike 0.5) lands, so they're not committed for the example yet; `wire-replay` cost display shows recorded (not $0) spend. From 547bde94c4fbcb11840f32149643e0e9d51a55ba Mon Sep 17 00:00:00 2001 From: Francis Eytan Dortort Date: Sat, 8 Aug 2026 18:09:59 -0400 Subject: [PATCH 3/3] docs(spec): document the shipped drivers and rewrite the capability matrix MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The §4.3 matrix now reflects the real capabilities() of the shipped Claude/Codex/Gemini/Antigravity drivers (with an Antigravity column) instead of Phase-0 'verify' placeholders; adds §4.4 documenting the three new drivers' invocations, event schemas, and caveats (PTY driver → §4.5); adds antigravity to the driver id union and the header. §4.3 kept as the matrix (referenced elsewhere); Cursor stays marked unvalidated. Co-Authored-By: Claude Opus 4.8 --- docs/SPEC.md | 63 +++++++++++++++++++++++++++++++++++++++------------- 1 file changed, 47 insertions(+), 16 deletions(-) diff --git a/docs/SPEC.md b/docs/SPEC.md index 867648e..cfba491 100644 --- a/docs/SPEC.md +++ b/docs/SPEC.md @@ -2,7 +2,7 @@ > **Playwright for AI Agents.** End-to-end testing for the things AI agents interact with — > **skills, plugins, MCP gateways, and LLM gateways** — by driving real agent CLIs -> (Claude, Codex, Gemini, Cursor) through scripted scenarios and asserting on what they *do*. +> (Claude, Codex, Gemini, Antigravity, and Cursor) through scripted scenarios and asserting on what they *do*. **Status:** Draft v2 (post-review) · **Owner:** dortort · **Last updated:** 2026-06-27 @@ -157,7 +157,7 @@ roadmap (Phase 0) validates each before committing. ```ts interface AgentDriver { - readonly id: 'claude' | 'codex' | 'gemini' | 'cursor' | string; + readonly id: 'claude' | 'codex' | 'gemini' | 'antigravity' | 'cursor' | string; capabilities(): DriverCapabilities; launch(opts: LaunchOptions): Promise; } @@ -227,23 +227,54 @@ claude -p "" \ `--mcp-config`. - **Skill/plugin observation** via the channels in §6.1 / §9. -### 4.3 Per-agent capability & interception matrix (validate in Phase 0 before freezing the interface) +### 4.3 Per-agent capability & interception matrix -| Capability | Claude | Codex | Gemini | Cursor | -|---|---|---|---|---| -| Machine-readable stream | `stream-json` ✓ | `codex exec --json` ✓ | `--output-format stream-json` ✓ (verify) | `cursor-agent` print/stream ⚠️ unverified | -| LLM interception | `ANTHROPIC_BASE_URL` | `model_providers..base_url` (under `CODEX_HOME`; built-in provider ids reserved) | base-URL **unproven** | **unproven / high-risk** | -| MCP transports | stdio + http/sse | verify | stdio + http/sse (verify) | verify | -| Config isolation | `--strict-mcp-config`, env/HOME | **`CODEX_HOME`** (not project-local) | env/HOME (verify) | verify | -| Session persistence off | `--no-session-persistence` | verify | verify | verify | -| Native budget control | `--max-budget-usd` | verify | verify | verify | -| Known risk | low | medium (config model differs) | medium (base-URL unproven) | **highest** (`cursor-agent --help` hung during probing) | +Claude, Codex, Gemini, and Antigravity are shipped; the cells below reflect each driver's actual +`capabilities()` and invocation. Cursor is not yet built (ROADMAP spike 0.13). -> These are **not interchangeable adapters.** Each driver's `capabilities()` drives capability -> gating (`test.skip(!caps.mcp)`), and each row above is a Phase 0 spike (ROADMAP §3) before that -> driver is built. Claude is the reference; the rest are committed v1 objectives, spike-gated. +| Capability | Claude | Codex | Gemini | Antigravity | Cursor | +|---|---|---|---|---|---| +| Machine-readable stream | `stream-json` ✓ | `codex exec --json` ✓ | `--output-format stream-json` ✓ | `agy -o stream-json` ✓ (`event`-discriminated) | `cursor-agent` ⚠️ unverified | +| `llmInterception` | `base-url` (`ANTHROPIC_BASE_URL`) | `provider-config` (`model_providers..base_url`; **not** wired to Agentry's Anthropic proxy) | `none` (base-URL unproven) | `none` (routed through Antigravity's backend) | unverified | +| `mcpTransports` | stdio + http/sse | stdio | stdio + http/sse | none (unverified) | unverified | +| Config isolation / hermeticity | `--strict-mcp-config`, `--no-session-persistence` | `--ephemeral`, `--skip-git-repo-check` | `--skip-trust` | `--add-dir` (agent still writes to its own scratch dir) | unverified | +| `toolPermissionControl` | ✓ (`--permission-mode`) | ✓ (`--sandbox` / `--dangerously-bypass-approvals-and-sandbox`) | ✓ (`--approval-mode`, `--allowed-tools`) | ✓ (`--dangerously-skip-permissions`) | unverified | +| `nativeBudgetControl` | ✓ (`--max-budget-usd`) | ✗ | ✗ | ✗ | unverified | +| Known risk / caveat | low | no cost reported; no terminal event (`run.end` synthesized on exit) | assistant text streams as deltas (coalesced); auth-tier changes across CLI versions | writes to a scratch dir → sandbox fs-diff best-effort; MCP unverified | **highest** (`cursor-agent --help` hung during probing) | -### 4.4 Secondary PTY/TUI driver +> These are **not interchangeable adapters.** Each driver's `capabilities()` drives capability +> gating (`test.skip(!caps.mcp)`). Claude is the reference; Codex, Gemini, and Antigravity are shipped +> and their rows reflect real, verified behavior. **Cursor** remains an unvalidated Phase 0 spike +> (ROADMAP 0.13). Only Claude wires the LLM proxy today, so **wire cassettes apply to Claude alone** — +> transcript record/replay works for every driver. + +### 4.4 Codex, Gemini, and Antigravity drivers (shipped) + +Each mirrors the Claude reference — a pure native-event → `AgentEvent` mapper, a pure `buildArgs`, and +a `Driver` class that spawns its CLI with stdin ignored (so the process never blocks on stdin). Event +schemas were captured live from the CLIs (codex-cli 0.147, gemini-cli 0.54, agy 1.1) and drive the +unit tests. + +- **Codex** (`@agentry/codex`) — `codex exec --json -C -m --ephemeral --skip-git-repo-check`, + plus `--dangerously-bypass-approvals-and-sandbox` (bypass) or `--sandbox workspace-write`. The stream is + `type`-discriminated (`thread.started`, `item.started/completed{agent_message|command_execution|file_change}`, + `turn.completed`); shell runs → tool `shell`, patches → `apply_patch`; usage from `turn.completed`. Codex + emits no terminal event, so `run.end` is synthesized on exit, and it reports no per-run cost. Interception + is `provider-config`, so the Anthropic LLM proxy is not wired (no wire cassettes). +- **Gemini** (`@agentry/gemini`) — `gemini -p --output-format stream-json -m --skip-trust` + (`--skip-trust` is required for headless untrusted workspaces), plus `--approval-mode yolo|default`. + `type`-discriminated (`init`, `message`, `tool_use`, `tool_result`, `error`, `result`); assistant text + arrives as `delta` chunks that are **coalesced into one message** per turn; usage + `run.end` from the + terminal `result`. Base-URL interception is unproven → `none`. +- **Antigravity** (`@agentry/antigravity`) — `agy -p --output-format stream-json --model --add-dir `, + plus `--dangerously-skip-permissions`. The stream is **`event`-discriminated** (`init`, `step_update`, + `result`); `agent_response` text is coalesced per `step_index`, and usage + `run.end` come from the + authoritative terminal `result`. agy writes to its own project/scratch dir by default, so sandbox fs-diff + capture is best-effort; MCP support is unverified (`mcpTransports: []`). + +`agentry doctor` probes all four CLIs and prints each driver's `capabilities()`. + +### 4.5 Secondary PTY/TUI driver For interactive UX tests, a `node-pty` + `@xterm/headless` driver drives the real TUI and parses screen state. Inherits the harder problems (idle detection, ANSI noise); **opt-in, off the MVP