A full-cycle delivery pipeline for coding agents. One skill takes a substantial task, interrogates it into a complete brief, then walks it through ten gated stages — and refuses to advance until each gate passes.
Agents write code well and judge when to stop asking you things badly. A
substantial task becomes twenty interruptions, or a confident build that skipped
the tests and quietly delivered two thirds of what you asked for. task-pipeline
front-loads every decision into one intake conversation, then runs to the end
without checking in — and closes by accounting for every requirement, from a list
rather than from memory.
Built for Claude Code, and installable into any agent that reads skills (Cursor, Codex, OpenCode, …). Every stage's doctrine ships inside the skill — no companion plugin, nothing to resolve, nothing that breaks when a dependency is missing.
intake grill → docs study → brainstorm + decompose → spec → plan → subagent build
→ tests → lint/deploy → post-deploy log check → docs/wiki sync → acceptance
flowchart TD
S0["0 · Harvest + intake grill<br/>brief · REQ table · source ledger"]
S1["1 · Docs study"]
S2["2 · Brainstorm + decompose"]
S3["3 · Spec — UX track first, if UI"]
S4["4 · Plan"]
S5["5 · Dev — worktree, subagents, TDD"]
S6["6 · Tests"]
S7["7 · Lint + deploy"]
S8["8 · Post-deploy"]
S9["9 · Docs + wiki"]
S10["10 · Acceptance"]
S0 --> S1 --> S2 --> S3 --> S4 --> S5 --> S6 --> S7 --> S8 --> S9 --> S10
S10 -. "platform: next module" .-> S3
S10 -. "accounts for every REQ in the brief" .-> S0
classDef manual fill:#fde68a,stroke:#b45309,color:#111827
classDef auto fill:#dbeafe,stroke:#1d4ed8,color:#111827
class S0,S2,S3,S7,S10 manual
class S1,S4,S5,S6,S8,S9 auto
Every gate is typed: auto — the orchestrator verifies it itself, pass/fail
(blue); manual — it waits for your explicit go (amber).
| # | Stage | Gate | Type |
|---|---|---|---|
| 0 | Harvest + intake grill — mandatory | source ledger written; shared understanding + autonomy sweep; brief locked | manual |
| 1 | Docs study | contracts grounded on current docs | auto |
| 2 | Brainstorm + decompose | design approved; UI verdict recorded; every REQ answered; platform: module map approved | manual |
| 3 | Spec | committed + reviewed; UI: super-ux chain validated, linter green | manual |
| 4 | Plan | parallel-ready, DoD per task | auto |
| 5 | Dev | tasks DONE (three review verdicts each), TDD green per task | auto |
| 6 | Tests | full suite green, new code covered | auto |
| 7 | Lint + deploy | lint clean + suite green before deploy | manual |
| 8 | Post-deploy | clean boot / honest degradation | auto |
| 9 | Docs + wiki | every stale source-ledger row updated; docs + wiki synced; the code graph refreshed and checked against the docs | auto |
| 10 | Acceptance | every REQ accounted for with evidence; operator signs off; the retro written — pruned before anything was added | manual |
- The intake grill asks what a senior engineer would ask before anything is touched — scope, edge cases, failure modes, rollback, who the user is — so the build does not stall halfway through.
- Every stage has a gate. No code before a spec. No deploy before tests. No "done" before the post-deploy logs have been read.
- Nothing falls out the back. The request becomes a frozen, addressable list of requirements, and the last stage accounts for every one with evidence — then walks the ladder for what should have been on the list and never was.
- Team discipline without a team. ADRs, a written plan, a real test suite, a wiki entry — produced as part of the work, not promised for later.
- It adapts to your repo, not the reverse. Deploy, docs and wiki conventions are read from the host project, so nothing is imposed.
- It gets better at your project, without getting longer. Each run ends with a retrospective, and the next run reads it — but the standing-instruction list is capped at ten and pruned before anything is added, so what you inherit is the rules that still fire, not an archive.
/plugin marketplace add ssheleg/task-pipeline
/plugin install task-pipeline@task-pipeline
Then say "run this through the pipeline", "the full cycle", or invoke
/task-pipeline <one-line task>. Russian phrasings ("полный цикл", "прогони по
конвейеру") route the same way. The skill creates a TaskList with one entry per
stage and walks the gates. See Install for the other channels.
The doctrine each stage runs on ships inside the skill. Nothing to install for it, nothing to resolve at preflight, no version skew with someone else's repo, and no stage that can fail because a plugin is missing:
| Stage | Built-in doctrine |
|---|---|
| 0 Knowledge harvest | knowledge-sources.md — source list, the wiki, the ledger, the stage-9 loop-back |
| 0 Intake grill | grill.md — interview loop, domain awareness, autonomy sweep |
| 2 Brainstorm | brainstorm.md — approaches, YAGNI, the no-code-before-approval gate |
| 2 Decompose | decomposition.md — platforms only: brick criteria, module map, build order |
| 3 Spec | spec.md — UX-track order, locked contracts, global constraints, self-review |
| 4 Plan | planning.md — zero-context tasks, parallel groups, no placeholders |
| 5 Build | build.md + review.md — isolation, ledger, subagent loop, review rubric, fix loop |
| 5–6 TDD | tdd.md — the iron law, red/green/refactor, the suite gate |
| 10 Acceptance | acceptance.md — REQ coverage table, evidence rules, the closing question |
| 10 + any audit | audit.md — the L0→L7 ladder and its seams, axis rotation, ratchets, proven checks |
| any loop | loop-guard.md — churn detection, caps, the break protocol |
Ported, not depended on. Stage 0 is adapted from
Matt Pocock's grilling / grill-with-docs
and stages 2–6 from the corresponding skills in
obra/superpowers — both MIT, both credited in
LICENSE → Third-party. Nothing at runtime reaches for either.
Optional bridge: if you already run an equivalent skill set, map it onto stages
2/4/5/6 in your pipeline.json → skills[]. That's a substitution, never a
requirement — the gates still govern, and nothing detects, recommends or waits for
an external provider.
Before any technical work, task-pipeline interviews you relentlessly — one question per turn, each with a recommended answer, exploring the codebase before asking — until every decision branch is resolved and locked into a task brief. There is no "clear enough task" exemption: no stage-1 work starts without a committed, confirmed brief.
Domain awareness. While exploring, the grill reads the project's own
CONTEXT.md / docs/adr/ and holds you to them — calling out terms that conflict
with the glossary, replacing overloaded words with a canonical one, stress-testing
relationships against concrete edge cases, and surfacing where the code contradicts
what you just said. Resolved terms are written into CONTEXT.md as they land;
decisions that are hard to reverse, surprising without context and the result of
a real trade-off get an ADR. Both files are created lazily.
Autonomy comes from the sweep. Beyond the task itself, the grill pre-resolves everything that would otherwise interrupt stages 1→10: which external libs need docs, branch and task-tracker policy, the test command and what "green" means, the lint command, the deploy target and its authorization, where logs and health live, which docs and runbooks to update, and the model. Each gets an answer or an explicit "stop and ask me here" — an unasked question is a scheduled interruption. Deploy authorization has a hard floor: a standing go counts only if it names the target and the preconditions.
Stage 0 doesn't open with a question. It opens by finding what the project already
knows about this task
(knowledge-sources.md):
the code, CLAUDE.md, CONTEXT.md and the ADRs, docs/ and docs/ux/, previous
pipeline briefs and their carry-over ledgers, the retro's standing instructions
(read in full — they bind the run; see below), the knowledge wiki if you have one,
and any other repository or hosted doc system your project names as its docs. It's
retrieval scoped by the task's own nouns, not a read of everything, and it ends with a
source ledger written into the brief — one row per source, what it says, how
fresh, and whether this run makes it stale.
That buys two things. The cheap one: you don't get asked what an ADR already answers. The one that matters: an answer nobody can check is a recollection. People answer from memory about systems they wrote a year ago, and without the document in hand there is no way to tell a decision from a misremembering — so the run builds on it and every later gate passes honestly on a false premise. With the harvest in hand the grill quotes the source instead: "the March ADR says orders go through the command handler, you just described a direct write — has that changed?" You outrank every document, but only out loud: an override quoted against its source is a recorded decision, an unquoted one is an undetected divergence. When two sources disagree, precedence is code > host docs/ADRs > wiki > memory.
Then the loop closes: stage 9 updates exactly what stage 0 read. Every doc the run proved stale is already in the ledger with what's wrong, so "docs updated" means the sources the next run will trust — not just the files this change happened to touch.
The wiki is obsidian-wiki (Karpathy's
LLM-wiki pattern), and it's the one source that carries why across projects and
across months. Detected via ~/.obsidian-wiki/config or a resolving wiki-query.
Installed → queried at stage 0, synced with wiki-update at stage 9. Not installed →
recommended once, with the line, and the run continues:
pip install obsidian-wiki
obsidian-wiki setup --vault /path/to/your/vaultIt is a recommendation, never a gate — no stage blocks on a missing wiki, and nothing asks twice in one run.
A grep finds a name. A graph finds reach: what actually calls this, what breaks if it moves, what every change passes through. That is the question stage 0 needs answered before it asks you anything, and the one documents answer least reliably — a document records the reach its author remembered.
So where a code graph exists, the pipeline uses it
(knowledge-graph.md).
The tool is graphify; detected via
graphify-out/graph.json. Not installed → recommended once, in the preflight block,
with the lines — then the run continues:
uv tool install graphifyy # the CLI
graphify install # add the /graphify skill to this agentthen, in the project root:
/graphify .
Stage 0 asks it what grep can't — graphify query "how does session reach the API layer", graphify affected "AuthModule", graphify god-nodes — and records it
in the source ledger with its build date, because a graph goes stale exactly like
a doc. It points; the code decides.
Stage 9 closes three artifacts, not two. Docs, wiki, and the graph — in the agent, so the documents this stage just edited are re-extracted too:
/graphify . --update
There is a CLI shortcut, graphify update ., which is structural, model-free and
code-only — the wrong default at the one stage whose job was changing the docs,
because it produces the most expensive kind of stale graph: one that was refreshed.
And the reason the graph is a peer of the docs rather than an afterthought: the
next run's harvest queries it first, so a stale graph is a false premise delivered
with the authority of a machine. A wrong doc gets argued with. A wrong graph gets
believed.
Then the divergence check — two independent statements of the same system. This is the part a doc linter cannot do, because it compares your docs against the code's actual shape rather than against itself:
| Ask the graph | A disagreement means |
|---|---|
graphify god-nodes |
a hub no document names — an undocumented seam: the thing every change passes through and nothing explains |
graphify path "A" "B" |
an edge the docs deny — either a leak in the code or a lie in the docs, and which one is a decision, not a guess |
graphify affected "X" |
callers the docs never mention — the documented blast radius is smaller than the real one |
| a doc naming a module the graph has no node for | the doc describes something that no longer exists |
Doc-side findings are fixed at stage 9. Absences go to stage 10's ladder walk and
become REQ rows with their checks — the graph is the fourth audit axis, and the
only one that finds an absence without reading for it. The graph is derived, so it
is never hand-edited and graphify-out/ is git-ignored by default: you fix the code
or the doc and re-extract.
Cadence: refresh every close-out, sweep periodically (stage 10, or when another axis goes quiet). Like the wiki, it is a recommendation, never a gate.
Every gate before the last one asks "is this artifact good?" — none asks "does this still contain everything that was asked for?" Scope doesn't leak inside a stage; it leaks on the seams, because brief → spec → plan → task briefs is four rewrites and nothing compares the lists.
So the grill's second hard output is an addressable requirement table: one row per independently verifiable deliverable, each naming how it will be verified. A requirement you can't say how to verify is a badly-stated requirement — it gets split during the grill, not discovered at the end.
From there the ids thread through everything:
| Where | What it does |
|---|---|
| Spec | every section carries covers: REQ-… |
| Plan | every task carries Implements: REQ-…; the gate is set equality against the brief — a difference is printed as the explicit list of dropped requirements |
| Build | the implementer's brief quotes the REQ statement verbatim, so it optimises the requirement and not just the instruction |
| Review | a third verdict beside spec-compliance and code-quality: does this satisfy its REQ? |
| Deploy | no REQ may still be open; a partial ships only with explicit acceptance |
| Acceptance | every REQ gets verified / partial / deferred / dropped — and verified requires evidence: a passing test name, a file:line, a command and its output |
Two rules keep it honest. The list is frozen — adding mid-run is free, removing or narrowing needs your explicit agreement, because silently restating the task smaller makes every later gate pass honestly on a shrunken task. And deferred out loud is forgotten — anything postponed, dropped or half-done goes into an append-only carry-over ledger the moment it's said, including implementer concerns and non-blocking review findings.
Stage 10 closes the circle with the question the pipeline exists to be able to answer from a list rather than from memory: here's what you asked for, here's what shipped, here's what's deferred and where it lives — what's missing?
A one-feature task runs the pipeline once. A platform — several independent
capabilities, several separately shippable surfaces, requirements no single
deliverable satisfies — gets cut into modules at stage 2, before any spec is
written (decomposition.md).
Modules are cut by capability, never by layer ("Ordering", "Billing" — not "Controllers", "Services"), and a candidate is only a brick when it is independently specifiable, buildable and testable, owns its own entities, talks to its neighbours through declared contracts only, and can land while leaving the system working. The committed module map fixes the build order — walking skeleton first, then topological, no cycles — and every requirement maps to exactly one module.
Then stages 3→10 run per module: dossier → plan → build → tests → deploy → post-deploy → docs → acceptance → next brick. Stages 0–2 run once for the platform, and the map's status column is what a resumed session reads to know where it stopped. Each module's spec is a full dossier: architecture, entities and ownership, contracts in and out with their failure behavior, business rules, edge and failure cases, UI/Figma chain, limits, open questions.
Any repeating pass can start undoing the previous one: two shapes alternating, the same file rewritten round after round, a finding that was closed coming back. That looks like progress and consumes a run, so it is detected mechanically: every repeat pass logs one line per touched file with the reason that forced it — a finding id, a failed gate item. "Cleanup" is not a reason.
It trips on revert-oscillation, a file edited twice for the same reason, a resurrected finding, a third entry into one stage, or two loops editing one file — plus hard caps (5 fix rounds per task, 2 re-entries per stage, 3 passes per module). On a trip the run stops editing, names shape A and shape B with their evidence, escalates to the layer that owns the conflict (rubric → operator → plan → spec → module map), re-plans the check as an ordered checklist with one verification command per item, and goes through it one at a time. A higher-layer conflict is never settled inside a lower loop.
The REQ spine catches a requirement that was named and lost. It cannot catch one that was never named — because a comparison needs two sides, and an absence has one. Nothing in a diff between spec and plan reveals the error path nobody specified, the entity nobody gave an owner, the failure mode nobody thought of.
So stage 10 opens with a ladder walk
(audit.md), not
with the coverage table. Each requirement is walked bottom-up through its rungs
— recorded decision → spec section → contract and its failure behavior → plan
task → the change in the tree → an executed named assertion → the surface a
user reaches and its docs — and the work is the seam between each pair: did the
decision reach the spec, does every contract have a task, did the DoD land in the
diff, would that test still pass with the production code deleted, does what
shipped satisfy the requirement's own statement rather than the task's
instructions. Findings are ordered by seam, never by file — the seam names
which layer of your process leaks. Every absence becomes a new REQ row with its
check before the table is written.
Bottom-up is not taste: a missing artefact low on the ladder makes everything above it meaningless, so top-down you spend the pass polishing a surface for a contract that does not exist.
The frame is a rung too, where the project designs visually. super-ux owns the
frame completely — the Figma on/off choice, the MCP preflight, the
SCR-NN/<Screen>/<state> naming, and a linter that catches a missing, misnamed or
stale link. What no linter can check is what the frame says. A link can be
present, correctly named and fresh while the picture behind it promises a retention
window, a credit meter or a pricing tier the spec never described and the code
never built — a rendered claim about the product, seen by more people than the
spec, and often the version stakeholders believe. Compare frames to frames and they
agree; compare specs to specs and they agree; the defect lives in the seam. So the
walk adds two questions on UI work: does the frame render what the spec says, and
does what shipped still match the frame. The spec is the contract — say which
document you propose to move, and remember that editing a shared design file is
outward, like a PR or a deploy.
Three rules keep the audit from becoming another loop:
- Every pass changes the axis, not the effort. A searching pass doesn't oscillate, it converges: each pass edits the corpus the next one reads, so the newest edits are always the least-reviewed text and are what the next pass finds. Measured over seven passes on a production repository, by pass six the audit was mostly repairing its own previous pass — while the finding count still looked healthy. So count both numbers every pass (new findings vs. self-inflicted ones); when the second overtakes the first, rotate the axis — seams down one deliverable, then invariants across deliverables, then one class swept end to end.
- A class that repeats twice becomes a gate, not a note. Once is an incident; twice is a category, and a category belongs in lint or CI where nobody has to remember it. The third instance in a ledger is how a mechanical defect becomes permanent.
- What can't be fixed now becomes a ratchet, never a TODO — a named, counted
set that may only shrink, printed beside every gate verdict
(
carry-over: 4 open (was 6) · unresolved: 0). A TODO is invisible until someone opens the file; a ratchet makes "green" never read as "verified" — it reads as "green, and here is exactly what was not looked at".
And the exit criterion that is usually skipped: a deliverable is audited when every rung has its artefact and every check you are relying on has been seen failing once against a planted defect. That is the TDD iron law — if you didn't watch it fail, you don't know it tests the right thing — raised from one test to every gate, linter and script in the run. A green result from an unproven check is worth nothing.
Every gate in this flow is good at this run and blind across runs. So the same class of failure gets caught, fixed and forgotten five times, and nothing in the pipeline notices it is the same one.
The last act of stage 10 is therefore a retrospective, written to
docs/superpowers/retro.md — one file per project, not per run
(retrospective.md).
Every run prunes and stamps; only a run that diverged writes an entry:
symptom with evidence, the stage it surfaced at, the stage that owned it, the
root cause, the fix, and the check that catches it the first time from now on.
Fixes come in three grades, and you take the highest one that works:
| Grade | What it is | What it costs later |
|---|---|---|
| 1 — mechanical | a test, a lint rule, a gate criterion, a hook | nothing: the check is the memory |
| 2 — standing instruction | a rule agents read, for what no check can decide | one of ten slots, and its retirement trigger must be written at birth |
| 3 — a note | something still being understood | expires in two runs, then it is promoted or deleted |
The prune is mandatory and runs before anything is added. Every standing instruction is checked against three retirement triggers — it became a check · every path or command it names is gone · it has not fired in the last five run stamps — and the list is held to a hard cap of ten. At eleven, the oldest never-fired rule goes; "but they all matter" is exactly the state in which the list stopped being read, and the ninth stale rule is what discredits the two that are load-bearing.
Nothing is deleted silently: every retirement writes one line in the log, so the incident survives and only the instruction leaves. And the counts print beside the gate verdict, like the carry-over ledger's, so a list that quietly grew back is visible where it happened:
GATE 10 acceptance: PASS — 14/14 REQ verified
carry-over: 0 unresolved · retro: 7 standing (was 9) · retired 3 · added 1
Stage 0 reads those standing instructions in full on the next run — which is the whole reason the cap exists and the prune is a gate criterion instead of a good intention. A rule nobody reads to the end is worse than no rule: everyone believes it is covered.
The moment a task touches any user-facing surface (web / mobile / CLI / TUI — a
screen, command, or visible behavior), super-ux
is the recommended workflow, detected early in the stage-0 grill. If it's
installed, task-pipeline uses it; if not, it gives you the install line on the spot.
The spec stage runs it before any plan is written: /ux (setup check) →
ux-foundation (personas, JTBD, customer journey maps, user stories) →
ux-flows (user flows + screens.md UI map, Figma frames) → ux-scenarios
(usage scenarios validated against super-ux's own scenario-format contract) →
/ux-lint (must pass). The spec then embeds the UX layer — scenario IDs, CJM
stages served, applicable UX patterns — and the plan's UI tasks carry scenario IDs
in their DoD. Scenarios come before interface.
/plugin marketplace add ssheleg/super-ux
/plugin install super-ux@super-ux
Figma is super-ux's, and the decision about it is stage 0's. super-ux mirrors
every SCR- screen and state into a frame when the project designs visually, and
it handles all of it: the on/off choice, the MCP preflight, the naming contract,
the drift linter. task-pipeline only settles the part that would otherwise
interrupt a run — is Figma on, is the MCP connected, and if it isn't, do we ship
text-only or stop and connect it? That last clause matters: super-ux recommends
the MCP and then continues text-only on its own, never blocking, so an unasked
question quietly narrows the delivery from "designed" to "described". The stage-0
sweep decides it, and the preflight block flags the missing MCP in the same
exchange as everything else.
One file, in a named team, decided before anything is drawn. Left to drawing
time, "where do I put this?" gets answered by whichever agent is holding the brush,
and the answer is usually create a new file — which is how a project acquires
three files called some variation of "Design", each with real work in it. So the
sweep settles the team or organization by name (a file URL says which file, not
whose workspace — a design that lands in someone's personal drafts is invisible to
everyone who needs it) and the file: the one already recorded, a URL you supply,
or creation in that named team, explicitly authorized the same way a deploy target
is. Two rules make it stick: never create while a recorded file resolves, and
if the recorded file doesn't resolve, stop and ask — never create a replacement,
because "I couldn't open it so I made a new one" is both the duplicate and a hidden
permissions problem. The URL is written to the project's canonical record —
docs/ux/foundation.md → Design tooling, or the repo's own docs when there's no
UX chain — before the first frame, and the audit's F rung then checks it
mechanically: every screens.md deep link is figma.com/design/:fileKey/…, so a
key that differs from the recorded one is a second file, caught by a string match.
The default recommendation is the most capable reasoning model the environment
offers — currently the latest Opus generation, but that's a tier, not a string.
Model ids go stale as generations ship, and you may be on another provider entirely,
so nothing is hardcoded: the pipeline resolves the top tier available at runtime and
stage configs use provider-agnostic tokens (default / inherit).
You confirm or override it (per-stage overrides welcome) before stage 0 — then it
stops asking. A skill can't switch the main-loop model; /model is yours.
Stage-5 subagents are pinned to the confirmed model automatically. If the
recommended tier isn't available, the pipeline says which one it's using and
continues — a reminder, never a block.
Stages 0→10 above are the plugin's example flow. It is a machine-readable config
(pipeline.example.json)
written against a universal contract
(pipeline.schema.json):
copy the example to pipeline.json in your repo and rewrite it with your own stages
(any count), your own skills[], and your own auto/manual gate types. The
framework bakes in no fixed stage count and no opinion on which gates are manual.
A pipeline config may declare an optional release block: a master enabled
toggle, a trigger, project-defined steps, and verify smoke-checks. It's off
unless a project turns it on, and every project configures its own. This repo's
own instance is .github/workflows/release.yml —
armed per repo by the RELEASE_ENABLED variable (unset = off), it validates the tag
against the manifests, cuts a GitHub release from the CHANGELOG, and smoke-tests
npx from a clean checkout. Copy and adapt it; nothing is hardcoded.
Stages 6–10 read the host project's CLAUDE.md conventions (tests / lint / deploy /
docs / wiki) with detection fallbacks, so the skill works in any repo. The canonical
artifact layout each stage writes to is fixed in
artifacts.md.
Claude Code plugin (recommended):
/plugin marketplace add ssheleg/task-pipeline
/plugin install task-pipeline@task-pipeline
Any agent via the skills CLI (Cursor, Codex, OpenCode, 70+ — not Claude Code, use the plugin above):
npx skills add ssheleg/task-pipeline --agent cursor --agent codex --global(one repeated --agent per agent; never include claude-code while the plugin is
installed — the plain copy shadows it)
npm installer (no clone needed):
npx github:ssheleg/task-pipeline # straight from GitHub
npx task-pipeline-skill # from the npm registry(package is task-pipeline-skill — the unscoped task-pipeline name is taken
on npm; installs the same skill + /task-pipeline command into ~/.claude,
idempotent, --force to overwrite)
Cursor: the skills CLI above with --agent cursor, or per project copy
cursor/rules/task-pipeline.mdc into the repo's
.cursor/rules/. Cursor has no global rules directory — use the skills CLI for a
global install, the .mdc for per-project, or paste it into Cursor Settings →
Rules. The rule is self-contained (no external links), so it works copied anywhere.
Plain skill:
git clone https://github.com/ssheleg/task-pipeline
cd task-pipeline && ./install.sh(copies the skill into ~/.claude/skills/task-pipeline and the /task-pipeline
command into ~/.claude/commands/; idempotent — rerun skips existing installs,
./install.sh --force overwrites)
The family updates as one package — a bundle with one member current and the rest stale is a combination nobody tested:
npx sshlg-skills update # installed but behind — updates everything
npx sshlg-skills install # nothing installed yet
npx --yes sshlg-skills@latest list # what the current release of each member isRestart your agent afterwards: skills and hooks load at session start.
Per-channel, when you are updating this one member only:
Pick one channel per agent — running the plugin and the plain/skills-CLI copy on the same Claude Code install yields a duplicate, shadowing skill.
| Agent / channel | Update |
|---|---|
| Claude Code (plugin) | claude plugin marketplace update task-pipeline → claude plugin update task-pipeline@task-pipeline → restart |
| Any agent (skills CLI) | npx skills update task-pipeline --global --yes; to add: repeated --agent <name> (never claude-code when the plugin is installed) |
| Cursor | skills CLI (above) with --agent cursor, or re-copy the .mdc per project |
| npm | npx task-pipeline-skill@latest / npx github:ssheleg/task-pipeline (ephemeral — always latest) |
| Plain skill | git pull && ./install.sh --force |
None for the pipeline itself — the doctrine for every stage ships inside the skill. Four optional companions make individual stages better:
| Companion | For | Required? |
|---|---|---|
| super-ux | the stage-3 UX track | only for user-facing tasks |
| context7 (MCP) | stage-1 docs study | recommended — web-search fallback |
| obsidian-wiki | stage-0 harvest + stage-9 sync | recommended — never a gate |
| graphify | stage-0 reach queries + stage-9 refresh + the graph↔docs divergence check | recommended — never a gate |
A single preflight block prints which are ready, which to install, and the model
recommendation, so you arm the whole run in one exchange. Detail:
companion-skills.md.
| File | What's in it |
|---|---|
SKILL.md |
the orchestrator: how to run, the stage table, the model decision |
references/stages.md |
per-stage detail and the exact gate criteria |
references/artifacts.md |
the canonical document layout each stage writes to |
references/conventions.md |
how stages 6–10 read the host project's CLAUDE.md |
references/knowledge-graph.md |
the code graph: install line, stage-0 reach queries, the stage-9 refresh, the graph↔docs divergence check |
references/retrospective.md |
the project retro: the three grades of fix, the mandatory prune, the cap of ten |
references/model-tiering.md |
model policy, the /model reminder, overrides |
templates/ |
brief, carry-over ledger, CONTEXT.md and ADR skeletons |
CHANGELOG.md |
every release, with the reasoning behind it |
CONTRIBUTING.md |
dev setup, the validator, the version-sync rule, release flow |
Issues and pull requests are welcome — see CONTRIBUTING.md for the repo's invariants (the structural validator, four-way version sync, and the surfaces that must never drift apart). Security reports: SECURITY.md. Everyone participating is expected to follow the Code of Conduct.
npm test # python3 test/validate.py — the structural validatorBuilt by ssheleg — sshlg.me
- X / Twitter — @fuck_this_year
- Telegram — @sshlg
Part of the ssheleg skill family:
super-ux, task-pipeline, agent-sync, make-skill, sheleg-design, seo-aeo-audit.
The family installs and updates as one package, for every agent you use — a bundle with one
member current and the rest stale is a combination nobody tested:
npx sshlg-skills install # nothing installed yet — the whole family, any agent
npx sshlg-skills update # installed but behind — updates everything
npx --yes sshlg-skills@latest list # what the current release of each member isRestart your agent afterwards: skills and hooks load at session start, so the session that updates is not the session that gets the new ones.
MIT © 2026 ssheleg. Third-party portions (the ported stage doctrine) are credited and licensed in LICENSE → Third-party.
{ "version": 1, "stages": [ { "id": 1, "state": "spec", "name": "Spec", "model": "default", // 'default' = the run's confirmed model "skills": ["your-team:spec"], // whatever your environment resolves "gate": { "type": "manual", "check": "spec committed and reviewed" } } ] }