Verify that what you built is what you designed.
walkdown is the integration layer between design prototypes (Claude Design, Figma) and AI-built applications (Claude Code). It extracts a product's intent into a blueprint — a versioned, file-based specification of features, stories, and rules — then tracks whether the built application actually satisfies that intent, through the project's own test suite and through judgment-based verification sessions called walkdowns — performed by AI agents as a screening tier, and by humans as the final word.
The name comes from engineering practice: a walkdown is when engineers physically walk a job site to verify that construction matches the drawings.
The blueprint is the hub. Everything else is a projection of it:
- The PRD (Notion) and prototype (Claude Design / Figma) are inputs — extracted from, never synced to.
- The built app is verified against it.
- The test suite (the project's own — RSpec workflow specs, Playwright, anything) is linked to it via lightweight tags.
- The panel renders it beside the page under review, and the embed lets humans and agents pin feedback and questions to real elements on real screens.
Both audiences work in the same place: agents read and write blueprint files with ordinary file tools inside Claude Code; humans review in the browser, with walkdown's panel riding beside the running app and rendering those same files.
| Doc | Contents |
|---|---|
| 00-vision.md | The problems walkdown takes on, thesis, principles, v1 scope and non-goals |
| 01-glossary.md | The vocabulary — every term, one meaning |
| 02-blueprint-schema.md | Entities, file layout, IDs, dual representation |
| 03-runner-contract.md | How any test framework plugs in (linkage, execution, results) |
| 04-embed-and-anchors.md | The anchor convention (test-id reuse), embed script, the docked and framed panel layouts |
| 05-runs-ledger.md | Verification runs, walkdowns, derived status |
| 06-prototype-contract.md | What a prototype must carry, and who owns it |
| 07-roadmap.md | Which problems are solved, what is next, what we decided not to do — a dated snapshot |
| 08-locations.md | Where a project's spec and records live, and why the default is outside the repository |
| 09-delivery.md | How walkdown arrives: the clone, the one dependency, the skills, the extension |
| 10-house-style.md | How code here is written — what the repo already does, written down |
| 11-architecture.md | The 5000-foot review: what is healthy, what to refactor, in what order |
walkdown is not on a registry, and does not need one. The clone is the install:
git clone --branch v0.3.0 https://github.com/profoundry-us/walkdown.git ~/.walkdown/walkdown
node ~/.walkdown/walkdown/bin/walkdown.js --help
v0.3.0 is the latest tagged version, and a tag is the stable copy to grab. main
moves daily; leave --branch off to follow it. walkdown is pre-1.0, so a minor version
may still change the ledger, thread or registry formats.
Git may warn that the tag "is not a commit"; that is how it announces an annotated tag,
and the clone is fine. What changed between versions is in CHANGELOG.md,
and UPGRADING.md says how to move an existing install to the next one.
No npm install, no build, no network. The panel bundle, the stylesheet and walkdown's
one dependency (vendor/yaml.js) are all committed; rollup, tailwind
and playwright are build-time only. Then:
walkdown init # this machine: home, registry, you, the skills
walkdown blueprints new --dir <your-project> # a spec and its ledger — outside your repo by default
Or hand the whole thing to an agent: "visit https://walkdown.dev/setup and set walkdown
up for this project" — site/setup.md is written to be followed, and the
setup skill (/walkdown:setup) takes it from there. docs/09-delivery.md has
the reasoning, including what it would take to remove the last package.
A command is a noun and then a verb; the noun alone lists, and walkdown <noun> help
lists its verbs (ADR 0012).
walkdown help has them all. The ones you reach for first:
walkdown init [--force]
walkdown blueprints new [<name>] [--dir <project-root>] [--commit none|spec|all]
walkdown blueprints import <path> [--all|--only <folders>]
walkdown upgrade [--dry-run]
walkdown skills [--into <dir>] [--project] [--force]
walkdown where [<kind>] [--blueprint <id>] [--json]
walkdown pointer [--dir <project-root>] [--into <file>]
walkdown run [--target <name>] [--rule <id>]
walkdown status [<rule-id>] [--blueprint <id>] [--target <name>] [--json]
walkdown lint [--blueprint <id>] [--no-checks] [--json]
walkdown hash [--blueprint <id>] [--write]
walkdown threads [--rule <id>] [--all] [--json]
walkdown threads show <id> [--json]
blueprints new scaffolds a home for a project — one flat folder holding a spec.yml
and storyboard template, a feature template, and AGENTS.md: the conventions any AI agent working
in the repo follows (read the blueprint first, carry anchors, tag checks with rule
ids, work the agent queue, claim-never-accept, never touch prototype/);
walkdown pointer --into CLAUDE.md tells agents where the specs are, in a fixed paragraph
that names no blueprint. init links the clone into Claude Code as the
walkdown plugin (~/.claude/skills/walkdown, one link, so updating the
clone updates it): /walkdown:judge (the agent-walkdown ritual — evidence
screenshots, judgment, run record, fail threads), /walkdown:incorporate (fold
answered questions into the blueprint; address notes), /walkdown:formulate
(turn a design/PRD into storyboard + rules + checks), /walkdown:setup and
/walkdown:backlog, plus /walkdown:lint and /walkdown:status. AGENTS.md carries the knowledge; skills carry the
procedures. run executes the project's checks through the runner
contract — run_all, or run_for_rule with --rule — injecting the target's env
and WALKDOWN_TARGET, and confirms which run record the reporter appended.
threads lists active questions and notes (newest first, with anchors and a body
preview); --all includes resolved ones. threads show <id> shows one in full — anchor,
body, and replies; threads reply and threads set answer and move one. The status table shows at most two thread refs per rule and
truncates with +N; threads --rule <id> is the full list.
status renders the derived per-rule verification table straight from the runs ledger:
latest checks per target, latest agent and human walkdown verdicts, open threads, and a
per-rule verdict (verified / pending / failing — exit 1 on any failure). A pass recorded
against an older statement (per-result statement_hash) renders as ~ stale, never as
passing. The table is followed by an active-threads digest (truncated bodies, capped at
six). With a rule id (walkdown status waitlist.join.visual-match), shows that rule in
full: statement, each evidence source's latest result with run provenance and evidence
paths, and every thread anchored to it. --json is the agent-facing form.
lint validates the blueprint end to end: schema and duplicate IDs, storyboard
screen/anchor references (including anchors mentioned in steps), statement-hash staleness,
check coverage from the rule tags in your test files (or a runner.list command, if
you name one), stale check comments, thread
lifecycles (answered-but-not-incorporated, waived without waived_by), and run
records. Errors exit 1; warnings don't.
hash reports statement_hash status per rule; --write repairs missing/stale hashes
in place, preserving YAML formatting. Hashes are sha256 of the whitespace-normalized
statement, stored truncated (sha256: + 12 hex).
walkdown serve starts the local server (default port 4700, 127.0.0.1 only). One server
answers for every blueprint registered on this machine, so start it from anywhere: the
panel finds a page's blueprint from its address, and asks when none or several claim it.
Started inside a registered project, or with --blueprint <id>, that blueprint is also
what a request naming none opens. A blueprint registered, or a storyboard edited, while it
runs is picked up without a restart.
- The review page —
/is walkdown's own page: the panel, with the blueprint's front door framed inside it. Put a URL in the fragment (/#http://localhost:3000/) to review a page the storyboard has never heard of. The browser extension (extension/) opens the same page over any tab, loading the same/panel.js. The prototype is mounted at/prototype/, and picking a rule takes the frame to that rule's screen via the storyboard. - The stand-in app — for screens with no running page to point at (the thing being
built is chrome, or is not built yet), an app path of
/stand-in/<screen-id>serves that screen's own design back as the app — different theme, ringed edge, corner label — so fading, ghosting and pinning on the app surface all work. It is not evidence: what it shows is the design. - Pin mode — with the embed snippet in a page (
<script src="http://localhost:4700/embed.js" data-walkdown data-bp="example">), clicking a real element pins a note or question to its anchor; the thread lands in the home'sthreads/with rule, screen, element, and author attached. Escape leaves pin mode (closing an open form first), as does clicking the badge again. Standalone pages (opened without the panel) resolve their screen from the URL and post to the same server — including HTTPS staging pages, via the Private-Network-Access preflight.data-bpnames the blueprint; omit it and pins file against the blueprintwalkdown servewas started in, which a server started outside any project does not have. - Human walkdowns — Start walkdown (your name arrives from git identity), judge
rules, Finish: the session is appended to the home's
runs/as a hash-stampedkind: walkdownrecord, satisfyinghumanverify requirements. A feedback box rides above the verdict buttons — anything written is filed as a note thread on the rule and linked into the run result, and a fail is refused until it has a why (the note, or a pin dropped during the session; failing arms pin mode so the note lands where the problem is). Rules with no build evidence show Approve / Refine instead of Pass / Fail — sign-off on the spec, recorded asapproved/refining, never counted as verification. Each verdict is written to the home'sdrafts/the moment it is given, so an unfinished sitting survives a reload or a closed browser and shows up inwalkdown statusas in progress; the ledger still gains exactly one record, at Finish, which deletes the draft. Drafts are working state — the directory ignores its own contents. - Threads read as conversations — a thread is one stream of messages (the opening note is simply the first), with initials, times, consecutive messages grouped under one name, a line marking what arrived since you last looked, and a standing composer where Enter sends and the reply appears before the server answers. Append-only by design: no editing, no deleting, and no reactions standing in for a verdict.
- Thread lifecycle — clicking a pin or a thread opens it: body, replies, and the
actions its state allows (Reply / Mark addressed / ✓ Verify / Reopen / Waive; Answer
and Mark incorporated for questions). Transitions are validated server-side;
verifiedandwaivedrequire a named human — agents claim work (addressed), never accept it. The same mutations are available from the terminal:walkdown threads reply n-0002 --as-agent "fixed in run …" --status addressed, thenwalkdown threads set n-0002 --verifyas yourself. Who a mutation records under is never an argument — it is theidentity:in your~/.walkdown/profile.yml. An agent working for you records under your name, because it is your instruction;--as-agentadds the provenance beside it, and refusesverifiedandwaivedoutright.
Every surface walkdown draws (the panel, the embed's pins and ghost, the prototype
wireframes, the example app) is styled with Tailwind CSS 4 + daisyUI 5, built ahead of time
into a single lib/viewer/walkdown.css that is committed and shipped in the package.
Installing walkdown still runs no build and pulls no CSS toolchain: tailwindcss and
daisyui are devDependencies, and yaml remains the only runtime dependency.
npm run build:css # rebuild lib/viewer/walkdown.css from styles/walkdown.css
npm run watch:css # …and keep rebuilding while you work on the UIbuild:css also drops a copy at example/app/walkdown.css: the example is meant to
behave like a real application, and a real application ships its own stylesheet rather
than fetching one from a dev tool. Prototype screens do borrow walkdown's copy over
http://localhost:4700/walkdown.css, right beside the embed.js tag they already
depend on.
Two themes ship in that stylesheet. light is the default and dresses the panel and
example apps. wireframe dresses the prototype screens — a mockup wears
<html data-theme="wireframe"> so it reads as a drawing rather than as a build, which
is exactly the distinction a walkdown is judging. Any element can pin either theme with
data-theme, so a prototype can be previewed in the real skin without editing it.
The docked panel is the one surface that needed care: it is injected into somebody
else's running application, where Tailwind's preflight would restyle their buttons
and their stylesheet would restyle ours. So the panel renders inside a shadow root
with the stylesheet scoped to it. The single exception is the sheet's @property
rules, which the CSS Properties API only registers at document level; the panel copies
just those into the host page, where they declare custom-property types and paint
nothing. embed.js keeps its own dozen lines of scoped CSS — pin markers are
positioned against host elements and have no business carrying a design system.
walkdown is specified with itself: .walkdown/blueprints/ holds
the tool's own rules (starting with thread-lifecycle governance) in two blueprints, walkdown
for the panel and the embed and cli for the rest, verified by the repo's node:test suite — tests tagged by name (... @rule:<id>), recorded by the node:test
reporter (walkdown/node-reporter, the third emitter alongside Playwright and
RSpec). walkdown run here runs the tool's checks through its own runner contract.
docs/ remains the design (why); the blueprints are the verifiable what.
The write side of the loop: add ['walkdown/reporter'] to a project's Playwright
reporter array and every npx playwright test run appends a run record to the
home's runs/ — per-rule results aggregated from @rule: tags, current
statement_hash stamped (so future staleness is detectable), failure screenshots as
evidence, and git provenance (<sha>-dirty on an unclean tree). WALKDOWN_TARGET sets the target
(default local); who it is recorded under is ci under CI and your configured
identity otherwise. The reporter never fails a test run.
Run npm test for the suite. The example is both the schema demo and the
integration fixture: cd example && node ../bin/walkdown.js lint.
Design docs + first milestone (a hand-run of the schema in example/) +
lint/hash/status/threads tooling + run-record emitters for both ecosystems:
the Playwright reporter (walkdown/reporter) and the RSpec formatter
(adapters/rspec/, with its own fixture suite) + walkdown serve
(panel, embed, pins, human walkdowns). The v1 scope from
00-vision.md is complete except extract (PRD/prototype →
blueprint merge). Publishing (npm walkdown + @profoundry scope, RubyGems
walkdown-rspec) is the next step before sharing.