diff --git a/README.md b/README.md index bcbc9ad..28717eb 100644 --- a/README.md +++ b/README.md @@ -2,33 +2,242 @@ ### From breaking change to verified PR. +**Live showcase:** [upgradepilot.vercel.app](https://upgradepilot.vercel.app) (real completed mission, static snapshot) · **Proof PR:** [briefbot#2](https://github.com/oussamaelfig/briefbot/pull/2) · **Demo video:** _linked on the submission_ + UpgradePilot is an autonomous dependency-migration agent built on the -[TrueForge](https://github.com/truefoundry/trueforge) agent harness. +[TrueForge](https://github.com/truefoundry/trueforge) agent harness. Point it at a repository +and an upgrade target: + +> "Upgrade this project to the current OpenAI SDK. Read the official migration documentation, +> reproduce what breaks, migrate the code, prove the tests pass, and ask me before touching the +> real repository." -Point it at a repository and an upgrade target. It reads the live migration documentation, -finds the affected code, reproduces the failures in an isolated sandbox, performs the -migration, proves the upgraded software works with real test execution — and then asks for -permission before opening the real pull request. +It does not produce a migration checklist. It performs the migration, proves the upgraded +software works with executed tests, and asks permission before opening the real PR. ```text Docs change ↓ -UpgradePilot understands the migration (Bright Data MCP → live official docs) +UpgradePilot understands the migration Bright Data MCP → live official docs, schema-validated ↓ -Finds affected code (repository analysis via subagents) +Finds affected code parallel subagents map breaking changes to call sites ↓ -Migrates in a sandbox (Daytona — credential-free isolation) +Migrates in a sandbox Daytona — credential-free isolation ↓ -Proves the new version works (pytest before/after, deterministic checks) +Proves the new version works pytest before/after + deterministic legacy-pattern scan ↓ -Asks permission (human approval — harness-enforced) +Asks permission human approval, enforced at the harness tool boundary ↓ -Opens the PR (GitHub MCP, approval-gated write tools) +Opens the PR GitHub MCP — the only path to the real world +``` + +![Mission Control dashboard](docs/screenshots/dashboard.png) + +![Approval screen](docs/screenshots/approval-modal.png) + +## The demo migration + +**`openai` 0.28.1 → current SDK (v1 client interface)** on +[oussamaelfig/briefbot](https://github.com/oussamaelfig/briefbot), a small meeting-notes app +frozen in 2023: module-level `openai.ChatCompletion/Embedding/Moderation` calls, `openai.api_base` +config, dict-style response access, and the removed `openai.error` taxonomy. + +Its test suite is fully offline — a local stub implements the OpenAI REST endpoints and the app +targets it via base-URL override — so the *same* tests grade both sides of the migration: + +| State | Result (validated in a real Daytona sandbox) | +| --- | --- | +| `openai==0.28.1` + original code | **13 passed** | +| current SDK + original code | **9 failed, 1 passed, 1 error** — `APIRemovedInV1` at every call site | +| current SDK + correct migration | **13 passed**, legacy-pattern scan: **0 matches** | + +Verification is never an LLM claim: it is pytest exit codes and a deterministic pattern scan, +executed in the sandbox, reported with raw log excerpts. + +## Architecture + +```mermaid +flowchart TD + subgraph tf [TrueForge harness] + Agent[UpgradePilot agent - OpenAI model] + SubRI[Release Intelligence subagent] + SubRepo[Repo Investigator subagent] + Approval[Native tool approval on GitHub write tools] + end + subgraph mcp [MCP connectors] + BD[brightdata - search_engine, scrape_as_markdown] + GH[github - mutations, approval-gated] + MC[mission-control - ours] + end + Daytona[Daytona sandbox - clone, install, pytest, migrate, verify. Credential-free] + subgraph mission [Mission Control - this repo] + Server[MCP server + approval state machine + zod trust boundary] + UI[Live dashboard - timeline, evidence, approval screen] + end + Target[Target repository on GitHub] + + Agent --> SubRI --> BD + Agent --> SubRepo + Agent --> Daytona + Agent --> MC + MC --- Server --> UI + Agent --> Approval --> GH --> Target +``` + +### How each sponsor capability does real work + +- **TrueForge** runs the whole loop: the agent, two parallel **dynamic subagents** with isolated + contexts (Release Intelligence, Repo Investigator), **sandbox-as-tool** execution on Daytona, + git-backed **skills** with progressive disclosure, automatic context engineering, persistent + sessions, and **native tool approval** on GitHub write tools — the harness itself pauses + before anything irreversible. The capability map below makes each of these provable. +- **Bright Data MCP** feeds the agent live official migration documentation + (`scrape_as_markdown`, `search_engine`). The pipeline is declarative and version-controlled: + [`skills/release-intel/sources.yaml`](skills/release-intel/sources.yaml) registers sources, a + **publisher allowlist** (provenance, not just domain), a recovery query, and the extraction + schema; [`.cursor/rules/brightdata-pipeline.mdc`](.cursor/rules/brightdata-pipeline.mdc) + mirrors it as a project rule. When a source drifts, the skill's fallback chain re-acquires the + docs from the next allowed source and the dashboard shows a `RECOVERED SOURCE` badge — data + provenance stays visible. +- **Daytona** executes everything that touches code: clone, installs, baseline reproduction, + migration edits, verification. The sandbox is **credential-free** — the only path from agent to + the real world is the approval-gated GitHub MCP. +- **OpenAI** is the model layer for the agent and both subagents. +- **Qodo** reviewed every PR in this repository — see the evidence below. + +### TrueForge capability map + +There is no orchestration script in this repository — TrueForge *is* the orchestrator. One row +per harness capability: what it does in this project, and where a judge can verify it. + +| Harness capability | How UpgradePilot uses it | Where to see the proof | +| --- | --- | --- | +| Agent loop + model | The entire mission — discovery → sandbox → approval → PR — is a single TrueForge run of the `upgradepilot` agent on `openai/gpt-5.2` | [`agent/upgradepilot.agent.md`](agent/upgradepilot.agent.md) (the full system prompt), [`agent/setup.md`](agent/setup.md); the session header in TrueForge during the run | +| Three MCP connectors | `brightdata` (live official migration docs), `github` (44 tools — the only path to the real world), `mission-control` (our own 11-tool MCP server; the dashboard is harness-native, no internals scraped) | Connector table in [`agent/setup.md`](agent/setup.md); tool registry in [`mission-control/server/src/mcp.ts`](mission-control/server/src/mcp.ts); the dashboard filling itself in real time | +| Sandbox-as-tool (Daytona) | Clone, venv, installs, baseline reproduction (**9 failed / 1 error**, `APIRemovedInV1`), migration edits, verification (**13/13** + legacy scan **0**) — all in a credential-free sandbox | Sandbox protocol in [`skills/openai-v1-migration/SKILL.md`](skills/openai-v1-migration/SKILL.md); baseline/verification cards on the dashboard; the [produced PR](https://github.com/oussamaelfig/briefbot/pull/2) body | +| Dynamic subagents | Release Intelligence + Repo Investigator fan out in parallel with isolated contexts; only their results return to the root agent's context | Mission protocol step 2 in [`agent/upgradepilot.agent.md`](agent/upgradepilot.agent.md); both subagents in the dashboard activity feed and the TrueForge session view | +| Git-backed skills, progressive disclosure | `release-intel` and `openai-v1-migration` load from this repository (ref `main`); only name + description sit in the input context, bodies are read from the sandbox on demand | Frontmatter of [`skills/release-intel/SKILL.md`](skills/release-intel/SKILL.md) and [`skills/openai-v1-migration/SKILL.md`](skills/openai-v1-migration/SKILL.md); skill registration in [`agent/setup.md`](agent/setup.md) §4 | +| Native tool approval | GitHub write tools carry `require_approval_for_tools: ["@write", "@destructive"]` — the harness pauses before any mutation, *on top of* Mission Control's approval state machine (single decision, action matching) | [`agent/setup.md`](agent/setup.md) §3; guards in [`mission-control/server/src/mission.ts`](mission-control/server/src/mission.ts); regression tests in [`mission-control/server/test/mission.test.ts`](mission-control/server/test/mission.test.ts); both approval moments in the demo | +| Context engineering, automatic | Deferred tool loading (`preload: false` default) keeps all 44 GitHub tool schemas out of context until first use; large scrape responses are offloaded to sandbox files with a preview in context; compaction enabled | TrueForge terminal log during the run — tool-loading and offload lines (see "What to watch" below) | +| Persistent sessions | TrueForge sessions are SQLite-backed and survive page reloads/reconnects; Mission Control state survives server restarts (atomic-write JSON snapshot) and the dashboard reconnects via SSE seq replay with an explicit `replay_gap` protocol | The mid-run refresh in the demo; `save`/`load`/`replayGap` in [`mission-control/server/src/mission.ts`](mission-control/server/src/mission.ts); SSE replay-gap tests in [`mission-control/server/test/routes.test.ts`](mission-control/server/test/routes.test.ts) | + +**Generality:** a second registry entry + playbook +([Flask 2→3, PR #11](https://github.com/oussamaelfig/UpgradePilot/pull/11)) is in review — the +product needed zero code changes. + +#### What to watch during the demo + +Five observable harness moments, in run order: + +1. **Parallel fan-out** — right after the mission starts, *two* subagents appear at once; + the dashboard activity feed and the TrueForge session show Release Intelligence and the + Repo Investigator running concurrently. +2. **Deferred tools** — the GitHub connector exposes 44 tools, but TrueForge keeps their + schemas out of the context window until first use; watch the tool-loading lines in the + TrueForge terminal log. +3. **Response offloading** — the scraped migration guide is large; TrueForge writes it to a + sandbox file and keeps only a preview in context. Also visible in the terminal log. +4. **The mid-run refresh** — hard-reload the TrueForge tab *and* the dashboard tab while the + run is live: the session comes back server-side and the dashboard replays its event log. + Nothing is lost. +5. **Defense in depth at the approval** — after the human approves on Mission Control, + TrueForge's *native* approval card appears for the GitHub write tools. Two independent + gates, both real; the agent cannot talk its way around the harness one. + +### Mission Control (this repository's code) + +The dashboard is itself an MCP server. The agent reports structured progress and evidence +through ordinary tool calls (`start_mission`, `report_breaking_changes`, `report_baseline`, +`report_verification`, `request_approval`, `await_approval`, …) — no harness internals scraped. +Two properties matter: + +1. **Model output is untrusted input.** Every payload crosses a zod trust boundary; malformed + reports are rejected with structured errors the agent can self-correct from. +2. **The approval state machine is deterministic.** Approvals resolve exactly once; + `report_pr_opened` is rejected unless the recorded PR matches the approved action (branch + + repository), so approval for one PR can never launder a different one into the audit state; + an approval from mission A can never satisfy mission B. All of it regression-tested. + +## Human approval model + +Defense in depth, both layers real: + +- **Product gate** — the agent calls `request_approval` with the exact external action and the + before/after evidence; Mission Control renders the approval screen; `await_approval` blocks + until the human decides. Rejection ends the mission. +- **Harness gate** — GitHub MCP write tools carry TrueForge's native + `require_approval_for_tools: ["@write", "@destructive"]`, so even a misbehaving agent cannot + mutate the repository without a human clicking through the harness approval. + +## Running it + +The public site ([upgradepilot.vercel.app](https://upgradepilot.vercel.app)) is a **static +showcase of a real completed run** — the SSE indicator shows reconnecting and the approval +buttons are inert by design. The full live system runs locally per +[`agent/setup.md`](agent/setup.md). + +Prereqs: Node 22+, a TrueForge instance (`npx @truefoundry/trueforge`), and API keys for +OpenAI, Daytona, Bright Data, plus a fine-grained GitHub PAT (Contents + Pull requests, +read/write) for the target repository. + +```bash +# 1. Mission Control (dashboard + MCP server on http://127.0.0.1:4100) +cd mission-control && npm run setup && npm start + +# 2. Wire TrueForge (model, Daytona, the three MCP connectors, both skills, the agent) +# — full click-by-click / API instructions: +open agent/setup.md + +# 3. Open a session with the upgradepilot agent and ask for an upgrade; watch +# http://127.0.0.1:4100 light up. +``` + +Tests: + +```bash +cd mission-control && npm test # server (38) + web (26) +cd demo-target && pip install -r requirements.txt && python -m pytest # fixture (13) ``` -> Built for the Agent Harness Hackathon (Aug 2026). Full architecture, demo evidence, and the -> Qodo Code Review Evidence section land as the build progresses — development itself happens -> through reviewed PRs on this repository. +## Qodo Code Review Evidence + +Every meaningful change in this repository went through a Qodo-reviewed PR before merge — +implement → review → respond point-by-point → fix with regression tests → re-review → merge. +Qodo was not a checkbox; it materially changed this codebase. + +Representative merged PRs (full review trails preserved): + +- [#3 Mission Control server](https://github.com/oussamaelfig/UpgradePilot/pull/3) — **the + strongest cycle**: two full review rounds, 7 findings, 7 applied, 7 regression tests. Best + catch: `report_pr_opened` accepted *any* approved approval, so approval of one PR could have + laundered a different PR into the audit state — exactly the class of bug this component exists + to prevent. Qodo then re-reviewed the fixes and found two follow-up bugs *in them* (half-loaded + corrupt snapshots; ahead-of-store replay gaps), both fixed and tested. +- [#1 Engineering standards](https://github.com/oussamaelfig/UpgradePilot/pull/1) — Qodo reviewed + its own review rules and found a High-severity supply-chain hole: the docs-recovery path + accepted any `site:github.com` result as "official". Fixed with an explicit publisher + allowlist; one finding declined with live-endpoint evidence and verified stale on re-review. +- [#2 briefbot demo fixture](https://github.com/oussamaelfig/UpgradePilot/pull/2) — retry-policy + gaps (transport errors escaping retries) fixed with regression tests; re-review clean. +- [#4 Mission Control dashboard](https://github.com/oussamaelfig/UpgradePilot/pull/4) — a real + stale-response race that could hide the approval modal, caught before it ever hit a demo; + status derivation moved server-side because Qodo enforced *this repo's own* standards file. +- [#5 Agent skills](https://github.com/oussamaelfig/UpgradePilot/pull/5) — baseline-pass + short-circuit bypassing the legacy scan; requested-target-version handling; dict-style response + access added to the scan; a schema-drift contract test added on request. +- [#7 Agent hardening](https://github.com/oussamaelfig/UpgradePilot/pull/7) — 7 findings from the + first full end-to-end mission, all applied with regression tests across two review cycles. +- [#8 UI redesign](https://github.com/oussamaelfig/UpgradePilot/pull/8) — 18 findings across the + landing page and Mission Control restyle, including a focus leak Qodo caught inside one of its + own suggested fixes. +- [#9 Transport contract](https://github.com/oussamaelfig/UpgradePilot/pull/9) — GET/DELETE on + `/mcp` now answer 405 per the stateless streamable-HTTP convention. +- [#10 Daytona console restyle](https://github.com/oussamaelfig/UpgradePilot/pull/10) — 12 + findings, including empty-state evidence chips that fabricated data, caught before any demo; + reviewed on its own stacked PR, merged to main via #8. + +Each PR ends with a **Qodo engagement log** table classifying every finding +(APPLIED / DECLINED WITH REASONING / FOLLOW-UP / STALE) with commits and regression tests. ## Repository layout diff --git a/docs/blog-assets/briefbot-pr.png b/docs/blog-assets/briefbot-pr.png new file mode 100644 index 0000000..ab25aa2 Binary files /dev/null and b/docs/blog-assets/briefbot-pr.png differ diff --git a/docs/blog-assets/daytona-approval-modal.png b/docs/blog-assets/daytona-approval-modal.png new file mode 100644 index 0000000..5abec89 Binary files /dev/null and b/docs/blog-assets/daytona-approval-modal.png differ diff --git a/docs/blog-assets/daytona-dashboard-mobile.png b/docs/blog-assets/daytona-dashboard-mobile.png new file mode 100644 index 0000000..573bebf Binary files /dev/null and b/docs/blog-assets/daytona-dashboard-mobile.png differ diff --git a/docs/blog-assets/daytona-dashboard.png b/docs/blog-assets/daytona-dashboard.png new file mode 100644 index 0000000..21d32d0 Binary files /dev/null and b/docs/blog-assets/daytona-dashboard.png differ diff --git a/docs/blog-assets/daytona-landing.png b/docs/blog-assets/daytona-landing.png new file mode 100644 index 0000000..050b3b3 Binary files /dev/null and b/docs/blog-assets/daytona-landing.png differ diff --git a/docs/blog-assets/mission-pr-opened.png b/docs/blog-assets/mission-pr-opened.png new file mode 100644 index 0000000..5359c2d Binary files /dev/null and b/docs/blog-assets/mission-pr-opened.png differ diff --git a/docs/blog.md b/docs/blog.md new file mode 100644 index 0000000..3ddd738 --- /dev/null +++ b/docs/blog.md @@ -0,0 +1,119 @@ +# From breaking change to verified PR: building UpgradePilot in one hackathon day + +*Field report from the Agent Harness Hackathon, San Francisco — August 29, 2026. Built in one day on TrueFoundry's TrueForge harness, with Bright Data, Daytona, Qodo, and OpenAI doing load-bearing work.* + +![UpgradePilot landing page — "From breaking change to verified PR."](./blog-assets/daytona-landing.png) + +## The shaped hole + +Every AI engineer in the room lived through the same week in 2023: `openai>=1.0` shipped, and every `openai.ChatCompletion.create` in production started throwing `APIRemovedInV1`. The fix was never intellectually hard — the migration guide existed. The cost was toil: read the docs, find every call site, upgrade, run the tests, open the PR, convince a reviewer you didn't miss anything. + +We picked that exact migration for three reasons. + +First, the hackathon's bar was explicit: *perform the work, don't explain it*. Nobody wants a chatbot that summarizes a migration guide; they want the verified PR. + +Second, honesty by construction. The v1 breakage is deterministic — install the new SDK against old code and it fails at every call site, offline, with zero API keys. Our demo target ([briefbot](https://github.com/oussamaelfig/briefbot), a meeting-notes app frozen in 2023) runs its whole suite against a local stub server, so the *same* tests grade both sides of the migration. The evidence is pytest exit codes; it can't be faked. + +Third, the kicker: the demo migrates a sponsor's own SDK. + +## What UpgradePilot does + +Input is one sentence: + +> "Upgrade this project to the current OpenAI SDK. Read the official migration documentation, reproduce what breaks, migrate the code, prove the tests pass, and ask me before touching the real repository." + +From there: + +1. **Understand** — a Release Intelligence subagent pulls the official migration guide live through Bright Data and extracts breaking changes into a schema-validated contract, every entry citing its official source URL. +2. **Locate** — a parallel Repo Investigator subagent maps those changes to actual call sites. +3. **Reproduce** — in a Daytona sandbox: install the new SDK against the unmodified code and watch the suite burn. **9 failed, 1 error** — `APIRemovedInV1` at every call site. +4. **Migrate** — the smallest correct change set (seven files on briefbot). +5. **Verify** — full suite re-run (**13/13 passed**) plus a deterministic scan for leftover legacy patterns (**0 matches**). Exit codes, not model claims — the raw command output is quoted in the [produced PR's body](https://github.com/oussamaelfig/briefbot/pull/2). +6. **Ask** — stop. Show a human the exact GitHub action and the before/after evidence. Wait. +7. **Act** — only after approval: branch, commit, real pull request through the GitHub MCP. + +![Mission Control mid-run: execution timeline, deterministic before/after test evidence, breaking changes with publisher provenance](./blog-assets/daytona-dashboard.png) + +*Mission Control mid-run, paused at the human gate. (Dashboard shots in this post come from a seeded preview mission used during UI work; the live run's numbers are quoted in the PR below.)* + +## TrueForge: the agent loop we didn't have to build + +The honest review of the harness is a list of things we never wrote: the agent loop, subagent orchestration, context management, approval UX, session persistence. TrueForge ran all of it. Three MCP connectors (Bright Data, GitHub, our own Mission Control), Daytona as sandbox-as-tool, parallel dynamic subagents with isolated contexts, and git-backed skills with progressive disclosure — only name and description in context, bodies read on demand. + +Context economics paid rent too: deferred tool loading kept all 44 GitHub tool schemas out of the context window until first use, and the large scraped migration guide was offloaded to a sandbox file with a preview left in context. Sessions persist server-side, which became a demo beat: hard-refresh both tabs mid-run and everything comes back, the dashboard replaying its event log. + +The design decision we're proudest of: **the dashboard is itself an MCP server.** Mission Control exposes eleven tools (`start_mission`, `report_baseline`, `request_approval`, `await_approval`, …), and the agent drives the UI through ordinary tool calls. No harness internals scraped, no narrator LLM. Model output is untrusted input, so every payload crosses a zod trust boundary; malformed reports bounce with structured errors the agent can self-correct from. + +Approval is defense in depth, and both layers are real. Our state machine guarantees an approval resolves exactly once, is scoped to its mission, and that a recorded PR must *match* the approved action (branch and repository). Underneath it, TrueForge's native tool gate holds GitHub write tools behind `require_approval_for_tools: ["@write", "@destructive"]` — even a misbehaving agent can't talk its way past the harness. The only path from agent to the real world runs through a human click. + +![The approval sheet: the exact external action, before/after evidence, and a human decision](./blog-assets/daytona-approval-modal.png) + +*The approval sheet. The payload is fixed — approval authorizes this action only.* + +![Mission Control on a phone](./blog-assets/daytona-dashboard-mobile.png) + +*The console holds up at phone width.* + +## Bright Data: live docs, provenance enforced + +Migration guidance comes from live official documentation via the Bright Data MCP (`scrape_as_markdown` for known URLs, `search_engine` for recovery). What makes it an engineering artifact rather than a scrape script is the registry: `skills/release-intel/sources.yaml` version-controls the source list in priority order, a **publisher allowlist**, a recovery query, and the extraction schema. + +The allowlist is the point: `github.com` is shared hosting, so a domain match proves nothing. Authenticity requires the owner segment (`github.com/openai/...`), and recovery results that don't match are discarded regardless of ranking. When a source drifts — failed scrape, thin content, failed schema validation — the skill falls through the registered fallbacks, then recovery search, and the dashboard shows a `RECOVERED SOURCE` badge. The pipeline lives inside the agentic workflow as a skill and a project rule, not beside it as a script. + +## Daytona: every claim is an executed command + +Everything that touches code happens in a Daytona sandbox driven through TrueForge's sandbox-as-tool: clone, virtualenv, installs, baseline reproduction, migration edits, verification. The evidence rule is absolute: every number on the dashboard comes from executed command output, raw log excerpts attached. + +The security property is just as load-bearing: the sandbox is credential-free. No GitHub token, no API keys (the fixture's tests run against a local stub). Prompt injection from scraped docs or a confused agent can't reach the real world from inside the sandbox — the only mutation path is the approval-gated GitHub MCP. + +The sandbox was also our best regression tool during development — twice it caught us about to quietly break the demo. More below. + +## Qodo: a second engineer, not a checkbox + +Every one of the ten PRs in this repo went through review → point-by-point challenge → fix with regression tests → re-review. Three catches worth naming: + +- **A supply-chain hole in our own rules (PR #1, High).** Our docs-recovery search accepted any `site:github.com` result as "official." Schema validation checks shape, not provenance. Fixed with the publisher allowlist — before any feature code existed. +- **Approval laundering (PR #3).** `report_pr_opened` originally required *an* approved approval — not that the recorded PR *was the approved one*. Approve PR A, record PR B. The guard now matches branch and repository, with regression tests. That cycle ran two full rounds — 7 findings, 7 applied, 7 regression tests — and the re-review found two further bugs *in our fixes*. +- **The focus leak inside its own fix (PR #8).** An earlier round had us harden the approval modal's focus trap. One cycle later Qodo flagged that while a decision is in flight both buttons are disabled, the focusables query returns empty, and Tab escapes the modal — a correctness bug inside the very behavior it had asked for. Fixed and re-verified. + +We also declined findings with evidence, on the record: when Qodo claimed our batch scrape tools were unavailable, we ran `tools/list` against the live connector, posted the output, and merged with the decline recorded. + +## OpenAI: the model, and the punchline + +The root agent and both subagents run on `gpt-5.2`. The work is structure-heavy — schema-valid extraction, exact sandbox command sequences, tool discipline across a seven-stage protocol — and the model held it across a full autonomous mission. And the poetry writes itself: OpenAI models migrating OpenAI's own SDK from OpenAI's own migration guide. + +## What broke along the way + +- **The SDK moved under us.** `pip install --upgrade openai` in the sandbox resolved 3.6.0, which depends on `httpx2` — not `httpx` — and our reference migration's tests, which constructed v1-style exceptions with an `httpx.Request`, broke. The durable fix: construct SDK exceptions with `request=None` so tests don't care which HTTP library the SDK drags in. Validated in a live sandbox before the demo depended on it. +- **Our own fix broke the demo's soul.** Responding to a Qodo finding about `api_base` restore semantics, our first patch read `openai.api_base` at import time — which turned the iconic call-time `APIRemovedInV1` failures into a wall of import-time `AttributeError`s. The demo's *failure signature* is a feature. Caught by re-running the sandbox validation protocol; `getattr` with a default preserved both behaviors. +- **GitHub closed our stacked PR for us.** PR #4 was stacked on PR #3's branch. Merging #3 deleted that branch, and GitHub closed #4 one second later (`base_ref_deleted` → `closed`, per the event log). Reopen, retarget to `main`, merge. Lesson: retarget stacked PRs before deleting their base. +- **The read-only PAT.** The fine-grained GitHub token behind the connector was minted read-only — fine-grained PATs default that way per permission — so every discovery call worked and the first write didn't. Re-minted with Contents + Pull requests read/write; the README now spells out the exact scopes. +- **Subagents hijacked the dashboard.** In the first end-to-end attempt, subagents called `start_mission` themselves and littered Mission Control with junk missions. The agent definition now reserves mission lifecycle for the root agent; subagents receive the `mission_id` and report only their own evidence. +- **The agent gave up at the finish line.** `await_approval` returned `pending`; the agent treated that as terminal and ended its turn instead of polling — a completed migration, stalled at the gate. Instructions now mandate the poll loop until a human decision, with no external actions while pending. + +Every one of these was a real failure observed in live runs or event logs, and each fix is a commit you can read. + +## The numbers + +- **Ten PRs** ([#1–#10](https://github.com/oussamaelfig/UpgradePilot/pulls?q=is%3Apr)): seven merged, three open at time of writing (docs and two UI restyles). Every PR carries a **Qodo engagement log** classifying each finding — applied, declined with reasoning, or verified stale — with commits and regression tests. +- **About fifty findings** evaluated on the record across those logs; most applied with regression tests, a handful declined with evidence. +- **Test growth as a side effect of review:** mission-control server suite 21 → 38 (PR #3's two review rounds alone took it 21 → 30); web suite 6 → 26 across the UI and hardening PRs; the demo fixture holds 13. (Executed at merge time: server 38 passed, web 26 passed.) +- **The artifact:** [briefbot PR #2](https://github.com/oussamaelfig/briefbot/pull/2), opened autonomously after human approval — and since merged. Its body quotes the executed evidence: baseline exit 1 (9 failed, 1 error, 1 passed on `openai==3.6.0`), post-migration exit 0 (13 passed), legacy-pattern scan 0 matches, all inside a credential-free sandbox. + +![The real pull request on GitHub: merged, with the evidence in the body and Qodo as reviewer](./blog-assets/briefbot-pr.png) + +*The proof: a real PR with baseline and verification evidence in the body — reviewed, then merged.* + +![Mission complete: all nine stages green, PR opened](./blog-assets/mission-pr-opened.png) + +*Mission Control at the end of a run: nine of nine stages, approval recorded, pull request opened.* + +## What we'd build next + +The architecture already generalizes: `sources.yaml` registers documentation sources and allowlists *per package*, and each migration is a playbook skill. Supporting the next dependency means a registry entry and a playbook — the agent, mission protocol, approval model, and dashboard don't change. Beyond that: triggering missions from release feeds instead of prompts, and letting the deterministic legacy scan gate CI so a migration PR can't merge with a stale call site. + +The software didn't say it worked. It proved it, and then it asked. + +--- + +*UpgradePilot: [github.com/oussamaelfig/UpgradePilot](https://github.com/oussamaelfig/UpgradePilot) · the autonomous run's PR: [briefbot#2](https://github.com/oussamaelfig/briefbot/pull/2) · every PR's Qodo review → response → re-review trail is public in the repo, engagement log included.* diff --git a/docs/dashboard-walkthrough.webm b/docs/dashboard-walkthrough.webm new file mode 100644 index 0000000..4fed1db Binary files /dev/null and b/docs/dashboard-walkthrough.webm differ diff --git a/docs/demo.md b/docs/demo.md new file mode 100644 index 0000000..14a935e --- /dev/null +++ b/docs/demo.md @@ -0,0 +1,95 @@ +# Demo runbook — 2.5 minutes + +## Pre-flight (10 minutes before judging) + +1. Mission Control running: `cd mission-control && npm start` → http://127.0.0.1:4100 shows the + empty state ("Waiting for a mission"). +2. TrueForge running (`npx @truefoundry/trueforge`), agent `upgradepilot` in the library; + connectors green: `brightdata`, `github`, `mission-control`; Daytona sandbox provider `ready`. +3. Reset the demo target if a previous run left a PR/branch: + `gh pr close --repo oussamaelfig/briefbot --delete-branch` (keep `main` pinned to + `openai==0.28.1`). +4. Fresh mission state: stop the server, delete `mission-control/server/data/`, restart. +5. Two windows side by side: left = TrueForge chat, right = Mission Control dashboard. +6. Keep the TrueForge terminal log visible (a strip under the chat works): it is where the + deferred-tools moment (44 GitHub tool schemas loaded only on first use, not up front) and + the offloading moment (large scraped docs written to sandbox files, preview in context) + actually show up. + +## The run + +Say: *"briefbot is a real app frozen in 2023 on openai 0.28. Watch UpgradePilot ship the +upgrade — not explain it."* + +Paste into a new `upgradepilot` session: + +> Upgrade https://github.com/oussamaelfig/briefbot from openai 0.28.1 to the current OpenAI +> Python SDK. Read the official migration documentation, determine what breaks, reproduce the +> failures in the sandbox, migrate the code, prove the migration works, then request approval +> before opening the real PR on GitHub. + +Narration beats (the dashboard drives itself): + +1. **Timeline lights up** — "Two subagents in parallel: Release Intelligence is pulling the + official migration guide live through Bright Data; the Repo Investigator maps it to actual + call sites." +2. **Breaking-changes table fills** — "Six breaking changes, each linked to the official doc it + came from. Provenance is allowlisted to the publisher — a random GitHub repo can't feed this." +3. **Baseline goes red** — "It installed the new SDK in a Daytona sandbox and *reproduced* the + damage: 9 failed, 1 error — `APIRemovedInV1`. No credentials in that sandbox, ever." +4. **Verification goes green** — "Migration applied, full test suite re-run, plus a + deterministic scan for leftover legacy calls: 13/13 pass, zero legacy sites. No LLM claims — + exit codes." +5. **The refresh** (10 seconds) — hard-refresh the TrueForge tab AND the dashboard tab, mid-run: + both come back exactly where they were. "The harness keeps the run alive server-side; the + dashboard replays from its event log." +6. **Approval modal** — "And here it stops. The exact action, the evidence, and a human + decision. The harness enforces this too: GitHub write tools are approval-gated at the tool + boundary." → Click **Approve** (+ approve the harness card in TrueForge). +7. **PR card** — open the real PR on GitHub: verified changes, evidence, doc links. "From + breaking change to verified PR." + +## Drift-recovery encore (30s, if asked about Bright Data) + +New session: same prompt + *"Simulate documentation drift on the primary source."* The +dashboard shows the drift warning and the `RECOVERED SOURCE` badge — the fallback fetch is real, +only the trigger is simulated. + +## Fallbacks + +- Live run slow → a pre-run mission's populated dashboard stays on screen (state persists); + walk the timeline while the live one progresses. +- Catastrophic (venue Wi-Fi, provider outage) → `docs/dashboard-walkthrough.webm` + the merged + PRs and README evidence carry the story. + +## Reset between runs + +Stop server → `rm -rf mission-control/server/data` → start server → close/delete the briefbot +PR and branch → new TrueForge session. + +## Judge-question cheat sheet + +- **"Does it only do OpenAI migrations?"** — The architecture is registry + playbook. + `skills/release-intel/sources.yaml` registers documentation sources, a publisher allowlist, + and a recovery query *per package*; `openai-v1-migration` is one migration playbook. + Supporting a new dependency means registering its sources and adding its playbook skill — + the agent, mission protocol, approval model, and dashboard are unchanged. +- **"What happens if the human rejects?"** — `await_approval` returns `rejected`; the agent + reports the stage failed, summarizes state, and stops — no retry, by instruction. Belt and + braces: a rejected approval can never record a PR (regression-tested in + `mission-control/server/test/mission.test.ts`), and GitHub write tools are still gated by + TrueForge's native approval even for a misbehaving agent. +- **"What if the docs site breaks?"** — Drift recovery is designed in: fall through the ordered + sources in `sources.yaml`, then `search_engine` with the registered recovery query, keeping + only results whose URL matches the publisher allowlist. The dashboard shows the drift warning + and a `RECOVERED SOURCE` badge — run the encore ("simulate documentation drift"); the + recovery fetch is real, only the trigger is simulated. +- **"Why should I trust the test numbers?"** — They are not model claims. Every reported number + must come from executed command output (pytest exit codes plus a deterministic legacy-pattern + scan), log excerpts are copied verbatim, and every report crosses Mission Control's zod trust + boundary, which rejects malformed payloads with structured errors. +- **"Why is the dashboard an MCP server?"** — It makes the UI harness-native: the agent reports + evidence through ordinary tool calls, so it works with any MCP client and scrapes no harness + internals. It also puts the approval state machine behind a tested trust boundary — one + decision per approval, the recorded PR must match the approved action (branch + repository), + and an approval from one mission can never satisfy another.