Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 14 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

> **Playwright for AI Agents** — end-to-end testing for the infrastructure AI agents interact with.

**Status:** Early MVP — Claude driver, structural assertions, record/replay. No published npm package yet. See [Roadmap](#roadmap).
**Status:** Early MVP — Claude, Codex, Gemini, and Antigravity drivers; structural assertions; record/replay. No published npm package yet. See [Roadmap](#roadmap).

---

Expand All @@ -14,7 +14,7 @@ Playwright drives a real browser and asserts on the DOM. Agentry drives a real A

You write TypeScript tests that:

1. **Drive a real agent CLI** (Claude Code, headless) with a prompt and a sandbox workspace.
1. **Drive a real agent CLI** (Claude Code, Codex, Gemini, or Antigravity — all headless) with a prompt and a sandbox workspace.
2. **Observe a normalized event stream** — assistant turns, tool calls, MCP requests, filesystem side-effects, token usage.
3. **Assert on structure, not text** — `toHaveToolCall`, `toHaveFile`, `toFinishWithin`. Exact matchers on what the agent did; no fragile string matching on free-form output.
4. **Replay deterministically** in CI via recorded transcripts — fast (<2 s), free, no agent required. Record once live; replay forever.
Expand All @@ -27,7 +27,7 @@ The target is the agent's surrounding ecosystem: skills, MCP gateways, plugins

There is no published npm package yet. Run from this monorepo directly.

**Prerequisites:** Node >= 20, pnpm >= 10, `claude` CLI in your PATH, `ANTHROPIC_API_KEY` set.
**Prerequisites:** Node >= 20, pnpm >= 10, and the CLI for whichever driver you use in your PATH — `claude` (with `ANTHROPIC_API_KEY`), `codex`, `gemini`, or `agy`, each authenticated per its own vendor. Replay-only runs need none of these.

```bash
git clone <this repo>
Expand Down Expand Up @@ -68,6 +68,8 @@ export default defineConfig({
});
```

`use.agent` selects the driver — `'claude'` (default), `'codex'`, `'gemini'`, or `'antigravity'` — and `use.model` must be a pinned snapshot id for that agent.

### 2. Test: `examples/basic/tests/todo.agentry.ts`

```ts
Expand Down Expand Up @@ -214,7 +216,7 @@ Test file → runner → AgentHandle.run(prompt)

Key design decisions:

- **Structured/headless substrate.** The Claude driver uses `claude -p --output-format stream-json --verbose` — no TTY, no screen scraping. The JSONL stream is the reliable source of truth (the CDP analog).
- **Structured/headless substrate.** Every driver uses its CLI's structured headless mode — `claude -p --output-format stream-json --verbose`, `codex exec --json`, `gemini -p --output-format stream-json`, `agy -p --output-format stream-json` — no TTY, no screen scraping. The JSONL stream is the reliable source of truth (the CDP analog).
- **Normalized event model.** Every driver output is translated into the same `AgentEvent` stream (`tool_use`, `tool_result`, `mcp_request`, `message`, `usage`, `run.end`, …). Assertions read only this model; drivers are pluggable.
- **Sandbox isolation.** Each scenario runs in a fresh temp directory. File side-effects are detected by content-hash snapshotting before/after the run, producing `fs` events.
- **Budget cap.** `budget.perTest.usd` is enforced post-run in the MVP; `--max-budget-usd` is passed to the Claude CLI as a native backstop.
Expand All @@ -229,10 +231,13 @@ See [docs/SPEC.md](./docs/SPEC.md) for the full design.
```
agentry/
├── packages/
│ ├── core/ @agentry/core — events, config, sandbox, runner, assertions, transcript
│ ├── claude/ @agentry/claude — Claude Code driver
│ ├── mcp/ @agentry/mcp — MockMcpServer + MCP matchers
│ └── cli/ agentry — CLI binary (test, record, init, doctor)
│ ├── core/ @agentry/core — events, config, sandbox, runner, assertions, transcript
│ ├── claude/ @agentry/claude — Claude Code driver
│ ├── codex/ @agentry/codex — Codex CLI driver
│ ├── gemini/ @agentry/gemini — Gemini CLI driver
│ ├── antigravity/ @agentry/antigravity — Antigravity (agy) driver
│ ├── mcp/ @agentry/mcp — MockMcpServer + MCP matchers
│ └── cli/ agentry — CLI binary (test, record, init, doctor)
├── examples/
│ └── basic/ Working example with committed transcript
└── docs/
Expand Down Expand Up @@ -267,7 +272,7 @@ The SPEC describes the full vision. What is implemented vs. planned:
| Skills/plugins CH6 **differential** harness (`toChangeBehaviorVs`) | Phase 2 |
| MCP **live**: agent→mock fixture wiring + protocol-compliance matchers (`toBeValidMcpProtocol`, error codes) | Phase 3 |
| LLM **gateway** matchers (routing/fallback/cache/rate-limit) + provider-pool resolver (foundation shipped) | Phase 4 |
| Codex, Gemini, Cursor drivers | Phase 5 (spike-gated) |
| Cursor driver | Phase 5 (spike-gated) |
| Semantic/LLM-as-judge assertions (Tier 4: `toSatisfyRubric`) | Phase 6 |
| HTML/JUnit/JSON reporters; trace bundles | Phase 6 |
| HOME-remap / container sandbox; commit wire cassettes (host config leakage today) | Phase 7+ |
Expand Down
27 changes: 17 additions & 10 deletions docs/ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ architecture, and the decisions still open. **Revised v2** after two pre-impleme
| Determinism | Record/replay cassettes (ordered, session-positional); **replay default**, live periodic |
| Authoring | Code-first TypeScript v1; declarative YAML deferred |
| **Target types in v1** | **All four — skills, plugins, MCP gateways, LLM gateways** (skills & plugins = one workstream → "three workstreams"). **Committed GA.** |
| **Agents** | **Claude** in MVP; **Codex + Gemini + Cursor** committed v1 objectives, spike-gated fast-follow |
| **Agents** | **Claude, Codex, Gemini, and Antigravity** drivers shipped; **Cursor** remains a spike-gated fast-follow |
| North star | Follow Playwright; diverge only for LLM non-determinism (3 divergences — SPEC §1.4) |

---
Expand Down Expand Up @@ -79,9 +79,10 @@ these are answered and written up in `docs/research/phase0-findings.md` (local).
### Cross-agent feasibility (per SPEC §4.3 matrix — gate each adapter)
| # | Question | Why it matters |
|---|---|---|
| 0.11 | **Codex:** `codex exec --json`; LLM proxy via `model_providers.<id>.base_url` under `CODEX_HOME`; MCP support | Codex adapter feasibility |
| 0.12 | **Gemini:** structured stream; base-URL interception (unproven); MCP config | Gemini adapter feasibility |
| 0.11 | **Codex:** `codex exec --json`; LLM proxy via `model_providers.<id>.base_url` under `CODEX_HOME`; MCP support | Codex adapter feasibility — **✅ shipped** (`@agentry/codex`; stream verified against codex-cli 0.147; interception is `provider-config`, not wired to the Anthropic proxy) |
| 0.12 | **Gemini:** structured stream; base-URL interception (unproven); MCP config | Gemini adapter feasibility — **✅ shipped** (`@agentry/gemini`; `--output-format stream-json` verified against gemini-cli 0.54; assistant text coalesced from deltas; base-URL interception still unproven → `none`) |
| 0.13 | **Cursor (highest risk):** does `cursor-agent` run headless/structured at all (probing previously hung)? LLM/MCP interception? | Cursor adapter feasibility — may demote to fast-follow if it fails |
| 0.14 | **Antigravity (`agy`):** `agy -p --output-format stream-json`; `event`-discriminated stream; permissions via `--dangerously-skip-permissions` | Antigravity adapter feasibility — **✅ shipped** (`@agentry/antigravity`; usage + `run.end` from the terminal `result`; sandbox fs-diff best-effort since agy uses its own scratch dir) |

**Exit criteria:** yes/no + evidence per spike; architecture deltas folded into `SPEC.md`. A failed
cross-agent spike (esp. 0.13) demotes that agent to post-v1 without affecting the Claude GA path.
Expand Down Expand Up @@ -121,8 +122,10 @@ Routing, fallback, cache, rate-limit, transformation, token-accuracy, latency-ov
propagation matchers over the existing LLM interceptor.

### Phase 5 — Cross-agent adapters (committed v1 objective, spike-gated)
Codex, Gemini, Cursor drivers against the proven interface; per-agent capability gating; shared
suites run across the matrix. Each gated on its Phase 0 spike (0.11–0.13).
The **Codex, Gemini, and Antigravity** drivers ship against the proven interface (`@agentry/codex`,
`@agentry/gemini`, `@agentry/antigravity`), each mirroring the Claude reference (pure parser +
`buildArgs` + honest `capabilities()`), with per-agent capability gating. **Cursor** remains, gated on
spike 0.13; shared suites run across the full matrix once it lands.

### Phase 6 — Tier-3/4 assertions + JUnit/JSON/HTML reporters + parallelism hardening
Structured-output (schema/AST) + LLM-as-judge (recorded for free replay, §8.5); rate-limit-aware
Expand Down Expand Up @@ -201,11 +204,12 @@ firm). Still open:
- [x] Spec drafted + revised post-review (`SPEC.md` v2)
- [x] Roadmap drafted + revised (this doc v2)
- [x] Independent review pass (Claude critic + Codex; in `docs/research/`)
- [~] Phase 0 spikes — **in progress**: 0.1 / 0.2 / 0.4 ✅, 0.6 / 0.7 🟡, config isolation ✅;
- [~] Phase 0 spikes — **in progress**: 0.1 / 0.2 / 0.4 ✅, 0.6 / 0.7 🟡, config isolation ✅,
cross-agent Codex (0.11) / Gemini (0.12) / Antigravity (0.14) ✅;
remaining: MCP shim (0.3), HOME-remap (0.5), cassette canonicalization (0.9), mcp-live (0.10),
cross-agent (0.11–0.13)
Cursor (0.13)
- [x] Repo shape = monorepo; Phase 0 ownership = Agentry-runs (§7)
- [~] **Implementation — MVP spine + LLM-proxy phase SHIPPED** (79 unit tests, CI green):
- [~] **Implementation — MVP spine + LLM-proxy phase + cross-agent drivers SHIPPED** (108 unit tests, CI green):
- [x] **Phase 1 (spine):** monorepo · event model · RunRecord · Sandbox (fs diff) · config
(model-pin) · assertion engine (Tiers 1–3) · cassette engine · Claude driver (live) · transcript
record/replay · runner + fixtures · console reporter · CLI (init/test/record/doctor). MVP §5
Expand All @@ -221,8 +225,11 @@ firm). Still open:
- [~] **Phase 4 (LLM gateways):** foundation laid (observable `llm_request`/`llm_response` +
positionable proxy + provider-impersonation seam, SPEC §9.3); remaining: provider-pool resolver
+ routing/fallback/cache matchers.
- [ ] **Phase 5** (Codex/Gemini/Cursor) · **Phase 6** (Tier 4 LLM-as-judge, HTML/JUnit reporters)
· **Phase 7+** (trace viewer, codegen, PTY driver) — not started.
- [~] **Phase 5 (cross-agent adapters):** Claude + **Codex + Gemini + Antigravity** drivers shipped
(each mirrors the reference: pure parser + `buildArgs` + `capabilities()`); `agentry doctor` probes
all four CLIs. Remaining: **Cursor** (spike 0.13).
- [ ] **Phase 6** (Tier 4 LLM-as-judge, HTML/JUnit reporters) · **Phase 7+** (trace viewer, codegen,
PTY driver) — not started.

*Known gaps:* wire cassettes capture host config until HOME-remap sandboxing (spike 0.5) lands, so
they're not committed for the example yet; `wire-replay` cost display shows recorded (not $0) spend.
63 changes: 47 additions & 16 deletions docs/SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

> **Playwright for AI Agents.** End-to-end testing for the things AI agents interact with —
> **skills, plugins, MCP gateways, and LLM gateways** — by driving real agent CLIs
> (Claude, Codex, Gemini, Cursor) through scripted scenarios and asserting on what they *do*.
> (Claude, Codex, Gemini, Antigravity, and Cursor) through scripted scenarios and asserting on what they *do*.

**Status:** Draft v2 (post-review) · **Owner:** dortort · **Last updated:** 2026-06-27

Expand Down Expand Up @@ -157,7 +157,7 @@ roadmap (Phase 0) validates each before committing.

```ts
interface AgentDriver {
readonly id: 'claude' | 'codex' | 'gemini' | 'cursor' | string;
readonly id: 'claude' | 'codex' | 'gemini' | 'antigravity' | 'cursor' | string;
capabilities(): DriverCapabilities;
launch(opts: LaunchOptions): Promise<AgentSession>;
}
Expand Down Expand Up @@ -227,23 +227,54 @@ claude -p "<prompt>" \
`--mcp-config`.
- **Skill/plugin observation** via the channels in §6.1 / §9.

### 4.3 Per-agent capability & interception matrix (validate in Phase 0 before freezing the interface)
### 4.3 Per-agent capability & interception matrix

| Capability | Claude | Codex | Gemini | Cursor |
|---|---|---|---|---|
| Machine-readable stream | `stream-json` ✓ | `codex exec --json` ✓ | `--output-format stream-json` ✓ (verify) | `cursor-agent` print/stream ⚠️ unverified |
| LLM interception | `ANTHROPIC_BASE_URL` | `model_providers.<id>.base_url` (under `CODEX_HOME`; built-in provider ids reserved) | base-URL **unproven** | **unproven / high-risk** |
| MCP transports | stdio + http/sse | verify | stdio + http/sse (verify) | verify |
| Config isolation | `--strict-mcp-config`, env/HOME | **`CODEX_HOME`** (not project-local) | env/HOME (verify) | verify |
| Session persistence off | `--no-session-persistence` | verify | verify | verify |
| Native budget control | `--max-budget-usd` | verify | verify | verify |
| Known risk | low | medium (config model differs) | medium (base-URL unproven) | **highest** (`cursor-agent --help` hung during probing) |
Claude, Codex, Gemini, and Antigravity are shipped; the cells below reflect each driver's actual
`capabilities()` and invocation. Cursor is not yet built (ROADMAP spike 0.13).

> These are **not interchangeable adapters.** Each driver's `capabilities()` drives capability
> gating (`test.skip(!caps.mcp)`), and each row above is a Phase 0 spike (ROADMAP §3) before that
> driver is built. Claude is the reference; the rest are committed v1 objectives, spike-gated.
| Capability | Claude | Codex | Gemini | Antigravity | Cursor |
|---|---|---|---|---|---|
| Machine-readable stream | `stream-json` ✓ | `codex exec --json` ✓ | `--output-format stream-json` ✓ | `agy -o stream-json` ✓ (`event`-discriminated) | `cursor-agent` ⚠️ unverified |
| `llmInterception` | `base-url` (`ANTHROPIC_BASE_URL`) | `provider-config` (`model_providers.<id>.base_url`; **not** wired to Agentry's Anthropic proxy) | `none` (base-URL unproven) | `none` (routed through Antigravity's backend) | unverified |
| `mcpTransports` | stdio + http/sse | stdio | stdio + http/sse | none (unverified) | unverified |
| Config isolation / hermeticity | `--strict-mcp-config`, `--no-session-persistence` | `--ephemeral`, `--skip-git-repo-check` | `--skip-trust` | `--add-dir` (agent still writes to its own scratch dir) | unverified |
| `toolPermissionControl` | ✓ (`--permission-mode`) | ✓ (`--sandbox` / `--dangerously-bypass-approvals-and-sandbox`) | ✓ (`--approval-mode`, `--allowed-tools`) | ✓ (`--dangerously-skip-permissions`) | unverified |
| `nativeBudgetControl` | ✓ (`--max-budget-usd`) | ✗ | ✗ | ✗ | unverified |
| Known risk / caveat | low | no cost reported; no terminal event (`run.end` synthesized on exit) | assistant text streams as deltas (coalesced); auth-tier changes across CLI versions | writes to a scratch dir → sandbox fs-diff best-effort; MCP unverified | **highest** (`cursor-agent --help` hung during probing) |

### 4.4 Secondary PTY/TUI driver
> These are **not interchangeable adapters.** Each driver's `capabilities()` drives capability
> gating (`test.skip(!caps.mcp)`). Claude is the reference; Codex, Gemini, and Antigravity are shipped
> and their rows reflect real, verified behavior. **Cursor** remains an unvalidated Phase 0 spike
> (ROADMAP 0.13). Only Claude wires the LLM proxy today, so **wire cassettes apply to Claude alone** —
> transcript record/replay works for every driver.

### 4.4 Codex, Gemini, and Antigravity drivers (shipped)

Each mirrors the Claude reference — a pure native-event → `AgentEvent` mapper, a pure `buildArgs`, and
a `Driver` class that spawns its CLI with stdin ignored (so the process never blocks on stdin). Event
schemas were captured live from the CLIs (codex-cli 0.147, gemini-cli 0.54, agy 1.1) and drive the
unit tests.

- **Codex** (`@agentry/codex`) — `codex exec --json -C <cwd> -m <model> --ephemeral --skip-git-repo-check`,
plus `--dangerously-bypass-approvals-and-sandbox` (bypass) or `--sandbox workspace-write`. The stream is
`type`-discriminated (`thread.started`, `item.started/completed{agent_message|command_execution|file_change}`,
`turn.completed`); shell runs → tool `shell`, patches → `apply_patch`; usage from `turn.completed`. Codex
emits no terminal event, so `run.end` is synthesized on exit, and it reports no per-run cost. Interception
is `provider-config`, so the Anthropic LLM proxy is not wired (no wire cassettes).
- **Gemini** (`@agentry/gemini`) — `gemini -p --output-format stream-json -m <model> --skip-trust`
(`--skip-trust` is required for headless untrusted workspaces), plus `--approval-mode yolo|default`.
`type`-discriminated (`init`, `message`, `tool_use`, `tool_result`, `error`, `result`); assistant text
arrives as `delta` chunks that are **coalesced into one message** per turn; usage + `run.end` from the
terminal `result`. Base-URL interception is unproven → `none`.
- **Antigravity** (`@agentry/antigravity`) — `agy -p --output-format stream-json --model <name> --add-dir <cwd>`,
plus `--dangerously-skip-permissions`. The stream is **`event`-discriminated** (`init`, `step_update`,
`result`); `agent_response` text is coalesced per `step_index`, and usage + `run.end` come from the
authoritative terminal `result`. agy writes to its own project/scratch dir by default, so sandbox fs-diff
capture is best-effort; MCP support is unverified (`mcpTransports: []`).

`agentry doctor` probes all four CLIs and prints each driver's `capabilities()`.

### 4.5 Secondary PTY/TUI driver

For interactive UX tests, a `node-pty` + `@xterm/headless` driver drives the real TUI and parses
screen state. Inherits the harder problems (idle detection, ANSI noise); **opt-in, off the MVP
Expand Down
Loading