Document Date: May 2026 (revised) Purpose: Guide the implementation of automated agentic development pipeline enhancements for the NAAS project. Audience: Project architect and Claude Code agents operating on the NAAS codebase.
This document defines three prioritized enhancements to the NAAS project's agentic development workflow. The goal is to transform the current manual, sequential subagent invocation pattern into an automated pipeline with quality gates, structured observability, and pipeline recovery.
Five worker subagents exist in .claude/agents/:
technical-architect— creates implementation plans from functional specsfeature-implementer— implements plans chunk by chunkcode-security-reviewer— reviews code for security concernstest-suite-generator— generates test suitesintegration-validator— runs integration tests
Current workflow: Manual sequential invocation. Developer triggers each agent, waits for output, manually invokes the next.
An automated pipeline where a single invocation (e.g., "Implement Spec 3") to a dedicated pipeline-orchestrator skill triggers the full chain: SCM initialization → architecture + plan decomposition → per-chunk TDD implementation with quality gate loops → final integration validation → SCM finalization (branch push, draft PR). Human intervention only required for escalated issues and the final PR merge.
| Component | Role | Category | Surface |
|---|---|---|---|
pipeline-orchestrator |
Pipeline entry point, lifecycle manager, and coordination loop | Orchestration | Skill (.claude/skills/) |
technical-architect |
Analyzes specs, produces implementation plans and chunk decompositions | Worker | Subagent (.claude/agents/) |
feature-implementer |
Implements code chunk by chunk, makes tests pass, fixes security issues | Worker | Subagent (.claude/agents/) |
code-security-reviewer |
Reviews code for security and quality | Worker | Subagent (.claude/agents/) |
test-suite-generator |
Generates test suites (TDD-first and post-implementation) | Worker | Subagent (.claude/agents/) |
integration-validator |
Tests cross-service integration | Worker | Subagent (.claude/agents/) |
pipeline-simulator |
Validates the orchestrator state machine by simulating runs without invoking real workers | Validation | Subagent (.claude/agents/) |
The pipeline-orchestrator is invoked directly by the developer (via /pipeline-orchestrator <spec>) and runs in the main Claude Code session. All five worker agents are invoked by the orchestrator via the Agent tool, never by the developer. (The pipeline-simulator is a separate validation harness — see "Pipeline Simulation Harness" under Observability — and is not part of a production pipeline run; it is invoked via the /pipeline-simulator-run skill.) The orchestrator MUST run in the main session because subagents cannot themselves invoke other subagents — placing the orchestrator in .claude/agents/ would prevent it from delegating to the worker pool.
The orchestrator manages the entire pipeline execution loop — not just pre/post phases. It invokes each worker via the Agent tool, reads results from Agent tool responses and artifact files, updates pipeline state (state.json) and the execution log, and decides the next step.
Workers are stateless specialists. They receive their context via the orchestrator's Agent prompt, do their work, produce artifact files, and return a summary. They never read or write pipeline state files (state.json, chunks.json).
Phase decomposition: Rather than encoding all pipeline logic in a single monolithic system prompt, the orchestrator's detailed per-phase instructions live in separate files under .claude/pipeline/phases/. The orchestrator prompt defines the phase sequence and maps each state.json phase value to an instruction file. When entering a phase, the orchestrator reads the corresponding file for detailed guidance.
This design:
- Reduces the orchestrator's active instruction set at any given time (one phase file vs. the entire prompt)
- Makes each phase independently reviewable and editable
- Eliminates fragile step-number cross-references — transitions reference phase names
- Centralizes the human-review escalation protocol in a single shared file
.claude/pipeline/phases/
├── pre-pipeline.md # Branch creation, state init, log init
├── architecture.md # Architect invocation, chunks.json validation
├── per-chunk.md # Test gen, implementation, security review, commit
├── integration.md # Integration validator invocation
├── post-pipeline.md # Push, PR creation, finalization
└── human-review.md # Shared escalation/resume protocol
The pipeline incorporates several layers of guardrails — iteration caps on the implementer (3 attempts to make tests pass), iteration caps on the security review reflection loop (3 attempts before escalation), an invocation-count budget guard (pause at 30 invocations), regression checks after every security fix, and explicit HUMAN_REVIEW escalation paths. With current-generation frontier models (Claude Opus 4.8, Claude Sonnet 5), some of this scaffolding is heavier than what is strictly required for the pipeline to produce correct output. Modern models exhibit stronger task persistence, better self-verification, and more reliable tool use than earlier generations, and many runs would succeed without any of these guardrails firing.
These layers are retained deliberately for three reasons:
-
Demonstration value. A pipeline that includes explicit quality gates, escalation paths, and budget controls visibly demonstrates agentic engineering discipline — exactly the discipline that distinguishes a production-ready agentic system from a prototype. The receipts (iteration counts, security fixes, escalations) tell a verifiable story.
-
Robustness across model generations. The pipeline is designed to remain reliable if a less capable model is substituted for cost reasons (e.g., Sonnet 5 in place of Opus 4.8), or if a future model exhibits regression on a particular workflow. The guardrails are calibrated for the minimum trustworthy behavior, not the typical case.
-
Catching the long tail. Even with a well-behaved model, edge cases — flaky test environments, ambiguous spec requirements, intricate security findings — can produce a runaway loop or a confidently wrong output. The guardrails catch these without requiring the developer to babysit every run.
The cost of this defense-in-depth is mostly cognitive surface area, not runtime overhead — the guards rarely fire on a healthy run, but their presence makes the pipeline trustworthy enough to leave unattended for spans of 30+ minutes. The per-spec pipeline-quality-report.md artifact (see Post-Pipeline Phase) makes these guardrails visible by recording when and how often each fired during a run.
Update all five existing worker subagent files and create the new pipeline-orchestrator skill using the current Claude Code formats — subagents with YAML frontmatter fields for tool scoping, model selection, and memory; the orchestrator skill with its own frontmatter (allowed-tools, model, effort).
.claude/skills/pipeline-orchestrator/SKILL.md
The pipeline-orchestrator is a thick orchestrator that runs in the main Claude Code session and manages the entire pipeline lifecycle:
- Pre-pipeline phase: Parse the spec identifier from the developer's invocation, create the feature branch, initialize pipeline state (
state.json), and create the pipeline execution log. - Architecture phase: Invoke the
technical-architectvia theAgenttool. Read the resulting plan file andchunks.json. - Per-chunk loop: For each chunk, invoke the
test-suite-generator,feature-implementer, andcode-security-reviewervia theAgenttool, handling reflection loops when quality gates fail. Commit each chunk after it passes. - Integration phase: Invoke the
integration-validatorvia theAgenttool. - Post-pipeline phase: Push the feature branch, generate the draft PR from the execution log, write the final pipeline summary, and emit the per-spec quality report.
- Resume: If
state.jsonalready exists, resume from the last recorded state instead of starting fresh. - Cleanup: On "Clean up pipeline for Spec X", remove transient pipeline artifacts.
After every Agent tool completion, the orchestrator performs a three-step update:
- Extract key data from the worker's response and any artifact files
- Update
state.jsonwith structured data - Append a summary line to the pipeline execution log
Why a skill, not a subagent: Claude Code subagents cannot themselves invoke other subagents. Placing the orchestrator in .claude/agents/ would prevent it from delegating to the worker pool — the Agent tool is unavailable inside a subagent context. Skills, by contrast, run in the main Claude Code session, which retains full Agent tool access. The skill body becomes the orchestrator's operating instructions when invoked.
Recommended configuration:
| Frontmatter field | Value | Rationale |
|---|---|---|
name |
pipeline-orchestrator |
Invoked as /pipeline-orchestrator <spec> |
description |
"Pipeline entry point and lifecycle manager. Invoke as /pipeline-orchestrator <spec> to run the full automated pipeline, /pipeline-orchestrator resume <spec> to continue an interrupted run, or /pipeline-orchestrator cleanup <spec> to remove transient artifacts." |
Loaded into context so Claude knows when to apply the skill |
argument-hint |
[spec-id] |
Autocomplete hint when typing the slash command |
disable-model-invocation |
true |
Pipeline runs are explicit developer actions, never auto-triggered by Claude pattern-matching chat |
allowed-tools |
Bash Read Write Edit Agent Grep Glob AskUserQuestion TaskCreate TaskGet TaskList TaskUpdate |
Pre-approves the toolset the orchestrator needs, eliminating per-call permission prompts during a run |
model |
claude-opus-4-8[1m] |
Complex multi-step coordination with significant accumulated context. Benefits from 1M context window, improved long-horizon focus, and stronger file-system-based memory |
effort |
high |
The orchestrator is a coordinator, not a reasoner — it dispatches workers, extracts their results, and applies deterministic state-machine and contract rules (phase-to-file mapping, CONTRACTS.md §5 log lines, the three-step state.json update). The deep cognitive work (architecture, implementation, security review) is delegated to worker subagents, each of which gets its own reasoning budget when invoked. high gives enough headroom for the careful state/sequencing work (resume logic, budget-guard-before-escalation ordering) while avoiding the per-turn cost xhigh would impose across the dozens of coordinator turns in a full run |
System prompt structure (skill body):
The skill body is intentionally concise (~70 lines). It defines:
- Identity and role
- First-action document loading (CLAUDE.md, AI-AGENT-PRINCIPLES.md, CONTRACTS.md)
- Three entry modes (fresh start, resume, cleanup) and how arguments select between them
- The state machine diagram and phase-to-file mapping table
- Budget guard (pause at 30 invocations) and how to surface it to the developer
- Critical rules (sole
state.jsonwriter, three-step update, targeted staging, pipeline mode instruction) - The post-pipeline obligation to emit
pipeline-quality-report.md
Detailed per-phase instructions remain at .claude/pipeline/phases/*.md (unchanged from prior layout). The skill reads the relevant phase file when entering each phase. This keeps the skill body focused on structure and constraints while phase files provide natural-language guidance for execution. Phase files remain at .claude/pipeline/ rather than moving inside the skill directory in order to keep their pipeline-aligned behavior more generally consumable.
Token budget guard note: The Anthropic API beta task_budget feature is set via API headers and is not exposed in terminal Claude Code. The pipeline therefore retains its existing invocation_count field in state.json as the budget control mechanism. If the orchestrator is ever migrated to the Agent SDK, task_budget becomes available as a refinement.
The technical-architect absorbs the plan decomposition responsibility that was previously handled by the chunking skill (.claude/skills/chunk-plans/SKILL.md). In a single invocation, it produces both:
- An implementation plan at
.claude/pipeline/plans/<slug>-plan.md(human-readable) - A
chunks.jsonat.claude/pipeline/chunks.json(machine-readable, see CONTRACTS.md Section 2)
Why merge decomposition into the architect: The architect already has full context — the spec, the codebase, the architectural decisions. A separate decomposer would re-read this context with loss of nuance, adding a failure point and a context-loading cycle for no benefit.
Chunking rules to absorb from the skill:
- Each chunk: ~200-500 lines of new code, ~30-45 min to implement
- Each chunk has standalone "Done When" verification criteria
- Chunks are ordered sequentially (dependencies only reference earlier chunks)
- First chunk: scaffold (directories, Dockerfile, docker-compose entry, FastAPI skeleton, health endpoint)
- Last chunk: integration-facing (connects to upstream/downstream services)
- Shared library changes get their own chunk when significant
scope_boundaryfiles do not overlap across chunks;shared_files(e.g., docker-compose.yml) may overlapdo_not_touchenforces hard boundaries between chunks
Recommended configuration:
| Field | Value | Rationale |
|---|---|---|
tools |
Read, Write, Grep, Glob, AskUserQuestion |
Write for plan files and chunks.json. AskUserQuestion for manual invocation mode. No Bash (doesn't run code). No Agent (doesn't invoke other agents — the orchestrator invokes it). |
model |
claude-opus-4-8[1m] |
Deep reasoning for architectural analysis and plan decomposition. |
effort |
xhigh |
Most reasoning-intensive role; decomposition quality scales with reasoning depth, and a decomposition error cascades through every downstream chunk. Deviates from the high baseline most components use — see "Effort levels" below. |
memory |
project |
Remembers architectural decisions across spec implementations. |
Delete the chunking skill: Remove .claude/skills/chunk-plans/SKILL.md — its functionality is now part of the architect's standard responsibilities.
For each .claude/agents/<agent-name>.md file:
1. Verify tools: field restricts capabilities appropriately.
| Component | Tools | Rationale |
|---|---|---|
pipeline-orchestrator (skill allowed-tools) |
Bash Read Write Edit Agent Grep Glob AskUserQuestion TaskCreate TaskGet TaskList TaskUpdate |
Full pipeline management + developer escalation. Agent invokes worker subagents (formerly named Task). TaskCreate/TaskGet/TaskList/TaskUpdate populate the Claude Code task UI for visual progress tracking |
technical-architect (subagent tools) |
Read, Write, Grep, Glob, AskUserQuestion |
Plan + chunks.json production |
feature-implementer (subagent tools) |
Read, Write, Edit, Bash, Grep, Glob, LSP, AskUserQuestion |
Full implementation toolset. LSP provides live type errors after edits when a code-intelligence plugin is installed for the language |
code-security-reviewer (subagent tools) |
Read, Grep, Glob, LSP |
Read-only by design. LSP enables call-hierarchy and reference-finding for vulnerability analysis |
test-suite-generator (subagent tools) |
Read, Write, Edit, Bash, Grep, Glob, LSP, AskUserQuestion |
Test file creation + verification. LSP flags type errors in generated test code |
integration-validator (subagent tools) |
Read, Bash, Grep, Glob, AskUserQuestion |
Test execution + diagnostics |
No worker subagent has access to Bash for git operations. The feature-implementer, test-suite-generator, and integration-validator use Bash for running code and tests, not for SCM. Git operations are exclusively the pipeline-orchestrator's responsibility.
LSP activation note: The LSP tool is inactive until a Claude Code code-intelligence plugin is installed for the relevant language (Python, TypeScript). The agents declare LSP in their tool lists so the capability is available when plugins are installed; absent a plugin the tool entry is harmless. Installing language-specific code-intelligence plugins is recommended but optional.
Frontmatter field naming: Skill frontmatter uses allowed-tools (hyphenated, space-separated) while subagent frontmatter uses tools (comma-separated). The set of valid tool names is identical between the two surfaces.
2. Verify model: and effort: fields.
Use claude-opus-4-8[1m] for components that benefit from deeper reasoning (pipeline-orchestrator, technical-architect, code-security-reviewer, integration-validator). Use claude-sonnet-5 for components that benefit from speed (feature-implementer, test-suite-generator).
Opus 4.8 brings three improvements that disproportionately benefit the orchestration and review roles: stronger long-horizon task persistence (relevant to multi-chunk pipeline runs), better file-system-based memory (relevant to the orchestrator's state.json and execution-log discipline), and proactive output self-verification (relevant to architect chunks.json validation and security reviewer verdicts). The [1m] model id selects the 1M-token context window, and standard Opus pricing applies.
Effort levels. Reasoning effort is configurable per agent via the effort: frontmatter field. Both subagents (.claude/agents/*.md) and skills (.claude/skills/*/SKILL.md) support it; when set it overrides the session effort level, and when omitted it inherits whatever effort the invoking session is currently using. Valid values are low, medium, high, xhigh, max; both Opus 4.8 and Sonnet 5 support the full range, including xhigh and max.
Convention: every agent and skill sets effort explicitly. Each component's effort level is chosen to match the reasoning demands of its task, and pinning it in the frontmatter keeps that choice stable. An omitted effort would inherit the session's level and drift with it — silently down-leveling a component below the depth its task needs, or up-leveling a mechanical one into needless latency and cost. Setting the value explicitly on every component makes effort a deliberate, self-documenting per-component decision rather than an artifact of the ambient session, and prevents inadvertent drift in either direction. Rationale that cannot live in the frontmatter is captured here.
| Component | Model | Effort | Rationale |
|---|---|---|---|
pipeline-orchestrator |
Opus 4.8 | high (explicit) |
Coordinator/dispatcher — the deep cognitive work is delegated to workers, each of which gets its own reasoning budget. The explicit line records the deliberate decision not to use xhigh here despite the long-horizon role, and avoids paying the xhigh premium across the dozens of coordinator turns in a full run. |
technical-architect |
Opus 4.8 | xhigh |
Most reasoning-intensive role; decomposition quality scales with reasoning depth, and a decomposition error cascades through every downstream chunk. Low invocation count per run, so the cost is bounded. |
code-security-reviewer |
Opus 4.8 | xhigh |
Subtle vulnerability detection in IAM code benefits from maximum reasoning depth — a missed flaw is a product failure. Also fires only a few times per run. |
integration-validator |
Opus 4.8 | high (explicit) |
Mostly execution and diagnosis against running services; high matches the task, and pinning it keeps the level from drifting with the session. |
feature-implementer |
Sonnet 5 | high (explicit) |
Chosen for speed; effort: high is pinned so implementation reasoning can't silently drop below high when the session effort is lowered. Sonnet 5 offers xhigh/max, but mechanical implementation work doesn't warrant the added latency. |
test-suite-generator |
Sonnet 5 | high (explicit) |
Chosen for speed; effort: high is pinned so test-generation reasoning is fixed at the intended baseline rather than inherited from the session. |
technical-architect and code-security-reviewer set effort: xhigh for their deep-reasoning roles; every other pipeline component pins effort: high. No component inherits its effort — each value is chosen for its task and fixed in the frontmatter so it cannot drift with the session.
3. Refactor system prompts to remove duplicated project context.
Agents automatically load CLAUDE.md context. For architectural details, agents should read the source documents directly via the Read tool as their first action (see Document Loading Convention below).
4. Add pipeline mode instructions.
Each worker's system prompt should include instructions for two operating modes:
- Pipeline mode (invoked via the Agent tool by orchestrator): Do not use
AskUserQuestion. If encountering an issue requiring human input, clearly state the problem and what is needed in the response. The orchestrator handles escalation. - Manual mode (invoked directly by developer): Use
AskUserQuestionfreely for ambiguities.
5. Add lint/format gate to the feature-implementer.
The feature-implementer's system prompt should include a self-verification step after all tests pass:
After all tests pass, run:
Python: ruff check + ruff format --check on scope_boundary files
TypeScript: tsc --noEmit + eslint on scope_boundary files
Fix any issues before declaring implementation complete.
6. Add test quality verification to the test-suite-generator.
In TDD mode, the test-suite-generator must run its generated tests and verify they all fail before completing. If any test passes before implementation exists, it is not testing new behavior — rewrite it.
7. Workers do NOT interact with pipeline state files.
Workers never read or write state.json or chunks.json. They receive their context from the orchestrator's Agent prompt and communicate results through their Agent response and artifact files. The orchestrator is the sole writer of state.json.
Each agent's system prompt should instruct it to read the relevant project documents as its first action using the Read tool. The NAAS documentation footprint is modest (~1,800 lines / ~69 KB across architecture doc, functional spec, and behavioral principles) and fits comfortably within context window limits.
Do not create Skills, curated subsets, or caching layers for document loading. Agents read the canonical source files directly.
The standard reading order for agents is:
CLAUDE.md— loaded automatically by Claude Codedocs/AI-AGENT-PRINCIPLES.md— behavioral guidelines (all agents)docs/architecture/SYSTEM_ARCHITECTURE.md— system architecture (agents that need cross-service context)- The relevant functional spec file — passed via the pipeline context or specified in the agent's task prompt
After modernization, verify that:
- The
pipeline-orchestratorcan create a branch, initialize state, and invoke thetechnical-architectvia the Agent tool. - The
technical-architectcan read a spec and produce both a plan file and a validchunks.json. - Each worker agent can still be invoked manually via
/agentsand performs its role correctly. - Tool restrictions are working (e.g.,
code-security-reviewercannot write files, worker agents cannot run git commands). - Agents correctly read project documents via the
Readtool as their first action. - Workers do not attempt to read or write
state.jsonorchunks.json.
Implement the orchestrator's state machine so that completing one phase automatically proceeds to the next, with the orchestrator managing the entire execution loop via explicit Agent-tool invocations.
┌──────────────────────────┐
│ Developer Input │ "Implement Spec X"
│ (invokes orchestrator) │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ pipeline-orchestrator │ PRE-PIPELINE PHASE:
│ (entry point) │ - Parse spec slug
│ │ - Create feature branch
│ │ - Initialize state.json
│ │ - Create pipeline log
│ │ - Invoke technical-architect via the Agent tool
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ technical-architect │ Analyze spec, produce plan + chunks.json
│ (invoked via Agent) │
└────────────┬─────────────┘
│ Agent call returns to orchestrator
│ Orchestrator reads chunks.json, updates state.json
▼
┌──────────────────────────────────────────────────────────────────┐
│ PER-CHUNK LOOP (Priority 3) │
│ Orchestrator invokes each worker via Agent │
│ │
│ ┌───────────────────┐ │
│ │test-suite-generator│ Write failing tests FIRST │
│ │ (invoked via Agent)│ │
│ └─────────┬─────────┘ │
│ │ Agent call returns, orchestrator updates state │
│ ▼ │
│ ┌───────────────────┐ │
│ │ feature-implementer│◄──── fix instructions (from orchestrator)│
│ │ (invoked via Agent,│ │ │
│ │ iterates until │ │ │
│ │ tests pass) │ │ │
│ └─────────┬─────────┘ │ │
│ │ Agent call returns │ │
│ ▼ │ │
│ ┌────────────────────────┐ │ │
│ │code-security-reviewer │────┘ │
│ │ (invoked via Agent) │ FAIL → orchestrator re-invokes │
│ │ │ implementer with fixes │
│ │ PASS → orchestrator │ │
│ │ commits chunk, │ │
│ │ advances to next │ │
│ └────────────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
│ All chunks complete
▼
┌──────────────────────────┐
│ integration-validator │ Full integration validation
│ (invoked via Agent) │
└────────────┬─────────────┘
│ Agent call returns to orchestrator
▼
┌──────────────────────────┐
│ pipeline-orchestrator │ POST-PIPELINE PHASE:
│ (continues its loop) │ - Push feature branch
│ │ - Generate draft PR from execution log
│ │ - Write final pipeline summary
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ HUMAN REVIEW │ Developer reviews and squash-merges PR
└──────────────────────────┘
The orchestrator drives the pipeline through these phases. Each phase's detailed instructions live in a separate file under .claude/pipeline/phases/:
PRE-PIPELINE → ARCHITECTURE → PER-CHUNK LOOP → INTEGRATION → POST-PIPELINE → DONE
↑
Any phase can → HUMAN_REVIEW (ask developer)
state.json Phase |
Instruction File | Summary |
|---|---|---|
starting |
phases/pre-pipeline.md |
Parse spec, create branch, init state + log |
architecture |
phases/architecture.md |
Invoke architect, validate chunks.json |
implementing |
phases/per-chunk.md |
Test gen → implementation → security review → commit (per chunk) |
integration_validation |
phases/integration.md |
Invoke integration validator |
post_pipeline |
phases/post-pipeline.md |
Push branch, create draft PR, finalize |
human_review |
phases/human-review.md |
Shared escalation/resume protocol |
The terminal phases complete (the "DONE" state above — a successful run) and failed (a run the developer aborted from human_review) have no instruction file and so do not appear in the table; they are recorded in state.json only. See CONTRACTS.md Section 3 for the full set of phase values.
Phase files use natural language guidance anchored by formal constraints (retry limits, state.json field names, phase values). They define entry conditions, execution guidance, state updates, success criteria, and escalation paths.
The orchestrator accumulates worker outputs in its conversation context — each Agent response returns to the orchestrator. Over a full pipeline run:
- Each Agent response: ~1-5K tokens
- 15-25 invocations (5 chunks × 3+ agents, with some reflection loops): ~15-125K tokens
- Plus orchestrator reasoning, Agent prompt construction, state/log reads: ~30-60K tokens
- Total estimated: ~50-200K tokens
This is why the orchestrator uses claude-opus-4-8[1m] (1M context) — the accumulated context fits comfortably. Two mechanisms prevent context pressure:
- Claude Code's automatic context compression. Older messages are compressed as context fills. The orchestrator only needs detailed access to the most recent Agent response.
- Persistent ground truth files. After each Agent call, the orchestrator writes state.json and appends to the execution log. If older context is compressed, these files serve as ground truth for any data the orchestrator needs later.
For detailed schemas of state.json, chunks.json, commit messages, and the execution log, see .claude/pipeline/CONTRACTS.md.
Key design principle: The orchestrator is the sole writer of state.json. Workers never read or write it. This eliminates dual-write bugs, simplifies workers, and ensures state consistency.
How workers communicate results to the orchestrator:
- Agent response text — natural language summary of what was done
- Artifact files — plan files, chunks.json, test files, review reports
- Test/lint results — pass/fail counts reported in the Agent response
The orchestrator synthesizes these into state.json updates and log entries.
How the orchestrator passes context to workers:
The orchestrator reads chunks.json and extracts the relevant chunk's fields into each worker's Agent prompt. Workers receive self-contained, unambiguous prompts — they don't need to know about pipeline state files, chunk IDs, or the broader pipeline context.
If the pipeline is interrupted (session crash, network failure), the orchestrator supports resume:
- Developer invokes orchestrator with: "Resume pipeline for Spec X"
- Orchestrator reads existing
state.json - Orchestrator maps the recorded phase and chunk status to its state machine and re-enters the loop at the correct point
- Previously completed chunks are not re-executed
The orchestrator tracks invocation_count in state.json, incrementing after each Agent call. The first time invocation_count reaches the threshold (>= 30) while budget_guard_triggered is false, the orchestrator pauses, reports current pipeline status to the developer, and sets budget_guard_triggered = true. The guard fires exactly once per run — the flag prevents it from re-pausing as the count keeps climbing past 30. This prevents runaway reflection loops from consuming excessive resources.
The orchestrator supports a cleanup command: "Clean up pipeline for Spec X"
- Deletes transient state files (
state.json,chunks.json) - Deletes plan and review files under
.claude/pipeline/ - Asks for confirmation before discarding uncommitted changes or deleting the feature branch
- Invoke the
pipeline-orchestratorwith a small spec (Spec 1: Event Ingestion — simplest, most self-contained). - Verify the orchestrator creates the branch, initializes state, invokes the architect via the Agent tool, reads chunks.json, and begins the per-chunk loop.
- Verify that each worker is invoked via the Agent tool with appropriate context extracted from chunks.json.
- Verify that
state.jsonupdates correctly after every Agent completion. - Verify that the execution log is appended after every Agent completion.
- Verify that the orchestrator commits chunks with targeted staging (not
git add -A). - Verify that HUMAN_REVIEW escalation works (e.g., architect flags an ambiguity).
- Verify that the orchestrator pushes the branch and creates a draft PR after all chunks pass integration.
The per-chunk loop includes conditional feedback loops so that if the code-security-reviewer finds issues, the orchestrator routes back to the feature-implementer with specific fix instructions.
The orchestrator reads the code-security-reviewer's Agent response to determine PASS or FAIL. On FAIL, the orchestrator constructs a new Agent prompt for the feature-implementer that includes the specific issues found, file paths, line numbers, and fix instructions.
-
Maximum iterations per chunk: 3 attempts at the security review stage. If the quality gate still fails after 3 iterations, the orchestrator escalates to HUMAN_REVIEW with a summary of all issues found.
-
Iteration context: Each reflection loop pass, the orchestrator includes in the implementer's Agent prompt:
- The original chunk implementation instructions
- The specific issues found by the reviewer
- The iteration count (so the implementer knows urgency increases)
- File paths and line numbers where issues were found
-
Loop scope: The reflection loop runs between
feature-implementerandcode-security-revieweronly. Architectural issues escalate to HUMAN_REVIEW rather than trying to auto-fix. -
TDD-first pattern: The
test-suite-generatorruns FIRST for each chunk, writing failing tests that define the chunk's success criteria. The implementer iterates on its implementation by running the test suite after each change until all tests pass. Only then does the security review run. -
Test-implementation loop: The
feature-implementeris allowed up to 3 internal iterations to make all tests pass. If tests are still failing after 3 iterations, the orchestrator escalates to HUMAN_REVIEW. This is separate from the security review iteration count. -
Post-security-fix regression check: When the implementer receives fix instructions from a security review, it must re-run the existing test suite after applying fixes to ensure no regressions. If security fixes break tests, the implementer must resolve both issues within its iteration budget.
-
Lint/format gate: After all tests pass, the implementer runs
ruff check+ruff format --check(Python) ortsc --noEmit+eslint(TypeScript) and fixes any issues before declaring implementation complete.
test-suite-generator (write failing tests for chunk N)
│
▼
feature-implementer (implement until tests + lint pass, max 3 iterations)
│
├── tests/lint failing → iterate (fix code, re-run, repeat)
│ │
│ ├── still failing after 3 attempts → HUMAN_REVIEW
│ └── passing → continue
│
└── tests + lint passing → continue
│
▼
code-security-reviewer
│
├── FAIL → feature-implementer (fix security issues)
│ │
│ ▼
│ re-run tests (regression check)
│ │
│ ├── tests broken → fix both, then back to reviewer
│ └── tests pass → back to reviewer
│
└── PASS → orchestrator commits chunk → next chunk
(or integration-validator if last chunk)
When a chunk passes the security review gate, the orchestrator (not a hook script) commits all changes:
- Read
chunks.json→ getscope_boundary+shared_filesfor current chunk git addeach file inscope_boundaryandshared_filesgit addcorresponding test files (derived from scope_boundary paths usingtests/mirror convention)- Never use
git add -Aorgit add . - Commit with structured message (see CONTRACTS.md Section 4)
This produces a commit history that tells the story of iterative, quality-gated development:
feat(spec-3-enrichment/chunk-5): Dashboard integration for enrichment metrics
feat(spec-3-enrichment/chunk-4): Risk score aggregation pipeline
feat(spec-3-enrichment/chunk-3): Geo-location enrichment service
feat(spec-3-enrichment/chunk-2): IP reputation enrichment
feat(spec-3-enrichment/chunk-1): Redis Stream consumer setup
- Intentionally introduce a security issue in a chunk's scope and verify the reflection loop catches it, feeds fix instructions to the implementer, and re-reviews.
- Verify that the max iteration count is respected and escalation to HUMAN_REVIEW works.
- Verify that iteration counts are tracked in
state.json. - Verify that per-chunk commits contain only files from
scope_boundary+shared_files+ test files. - Verify that lint/format checks run as part of the implementation verification loop.
Status: not yet implemented. Priorities 1–3 are fully built; Priority 4 remains a forward-looking design. There is no Agent Teams configuration in the repository (no
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMSflag, no teammate spawn templates). The section below is retained as the intended approach should this optimization be pursued.
Use Claude Code Agent Teams to parallelize test generation for a single module, spawning teammates for unit tests, integration tests, and security tests simultaneously.
- Priorities 1-3 must be working and stable.
- Enable Agent Teams: set
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1in settings.json env.
Apply Agent Teams to ONE well-bounded task only: generating the complete test suite for a single NAAS service after the feature-implementer completes all chunks for a spec.
Spawn three teammates from the test-suite-generator agent's completion:
| Teammate | Responsibility | File Scope |
|---|---|---|
| Unit Test Writer | Unit tests for individual functions/classes | tests/unit/ |
| Integration Test Writer | Service-to-service integration tests | tests/integration/ |
| Security Test Writer | Security-focused tests (injection, auth bypass, etc.) | tests/security/ |
Teammates coordinate via the shared task list. Each teammate owns a non-overlapping file scope to prevent merge conflicts.
Agent Teams run at approximately 15x standard token usage. Only use this for:
- Complete test suite generation after a full spec is implemented
- Not for per-chunk test generation (the sequential
test-suite-generatorhandles that)
- Run Agent Teams on Spec 1 (Event Ingestion) test generation as a proof of concept.
- Compare output quality and coverage against sequentially-generated tests.
- Measure token cost and wall-clock time savings.
Estimated effort: 3-4 hours
- Create the
pipeline-orchestratorskill withallowed-tools,model, andeffortfrontmatter (plusname,description,argument-hint,disable-model-invocation), and a skill body defining the state machine, phase-to-file mapping, the three entry modes, Agent-tool prompt templates, commit logic, resume logic, cleanup logic, and budget guard. - Update the
technical-architectagent to absorb plan decomposition: add chunks.json production, chunking rules, update tool list (Read, Write, Grep, Glob, AskUserQuestion), add pipeline mode instructions. - Remove the old
chunkingskill (.claude/skills/chunk-plans/) — its responsibilities are now owned by thetechnical-architect. - Update all five existing worker agent
.mdfiles: verify tool lists, add pipeline mode instructions (workers do not interact with state.json/chunks.json), add lint gate to implementer, add TDD verification to test-suite-generator. - Audit worker agent system prompts and remove duplicated project context. Ensure each agent's first-action instructions follow the Document Loading Convention.
- Update
.gitignoreto track.claude/pipeline/(exceptstate.jsonandchunks.json). - Verify the
pipeline-orchestratorcan create a branch, initialize state, and invoke thetechnical-architectvia the Agent tool. - Verify the
technical-architectcan read a spec and produce both a plan file and a validchunks.json. - Verify each worker agent still works correctly via manual invocation.
Estimated effort: 3-4 hours
- Implement the orchestrator's full state machine loop (pre-pipeline → architecture → per-chunk → integration → post-pipeline).
- Implement the three-step post-Agent-call update pattern (extract data → update state.json → append to log).
- Implement per-chunk commit logic with targeted staging.
- Implement resume logic (read state.json, determine position, re-enter loop).
- Implement budget guard (pause at 30 invocations).
- Test the full pipeline end-to-end on Spec 1 (Event Ingestion), from orchestrator invocation through draft PR creation.
Estimated effort: 2-3 hours
- Implement the reflection loop in the orchestrator's per-chunk logic: security review FAIL → re-invoke implementer with fix instructions → re-run review.
- Implement max iteration enforcement (3 per quality gate stage).
- Implement HUMAN_REVIEW escalation with detailed context.
- Implement the post-security-fix regression check (implementer re-runs tests after applying security fixes).
- Test the reflection loop by introducing deliberate issues.
Estimated effort: 2 hours
- Enable Agent Teams in settings.
- Create a spawn prompt template for the three test-generation teammates.
- Test on Spec 1's completed implementation.
- Evaluate cost/benefit for continued use.
Worker agents do not interact with git or GitHub directly. All source control operations are owned by the pipeline-orchestrator:
- Pre-pipeline: Branch creation
- Per-chunk: Targeted file staging and commits (after each security review PASS)
- Post-pipeline: Branch push and draft PR creation
This enforces a clean separation of concerns: worker agents reason about code, the orchestrator manages the development lifecycle.
One feature branch per spec, squash-merged to main upon completion.
# Derived from developer prompt: "Implement Spec 3: Enrichment and Evaluation"
SPEC_SLUG="spec-3-enrichment"
git checkout main
git pull origin main
git checkout -b "feature/${SPEC_SLUG}"
mkdir -p .claude/pipeline/logs
cat > .claude/pipeline/state.json << EOF
{
"contract_version": 2,
"spec": "Spec 3: Enrichment and Evaluation",
"spec_slug": "${SPEC_SLUG}",
"branch": "feature/${SPEC_SLUG}",
"phase": "starting",
"current_chunk": 0,
"total_chunks": 0,
"invocation_count": 0,
"budget_guard_triggered": false,
"chunks": [],
"started_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)",
"completed_at": null
}
EOF
echo "# Pipeline Run: Spec 3 — Enrichment and Evaluation" > ".claude/pipeline/logs/${SPEC_SLUG}.md"
echo "# Started: $(date -u +%Y-%m-%dT%H:%M:%SZ)" >> ".claude/pipeline/logs/${SPEC_SLUG}.md"After each chunk passes its security review quality gate, the orchestrator stages and commits with a structured message:
# Read chunk's file list from chunks.json
# Stage ONLY scope_boundary + shared_files + test files
git add services/signal-enrichment/app/consumer.py
git add services/signal-enrichment/app/models.py
git add docker-compose.yml
git add tests/unit/test_consumer.py
git add tests/unit/test_models.py
git commit -m "feat(spec-3-enrichment/chunk-1): Redis Stream consumer setup
Tests: 8 written, all passing
Implementation iterations: 1
Security review iterations: 1
Security issues caught: 0
Pipeline: auto-committed by agentic pipeline"STATE_FILE=".claude/pipeline/state.json"
SPEC_SLUG=$(jq -r '.spec_slug' "$STATE_FILE")
LOG_FILE=".claude/pipeline/logs/${SPEC_SLUG}.md"
# Push feature branch
git push -u origin "feature/${SPEC_SLUG}"
# Create draft PR with body generated from pipeline execution log
gh pr create \
--draft \
--title "Spec 3: Enrichment and Evaluation" \
--body "$(cat <<EOF
## Summary
[Generated from pipeline execution log]
$(cat "${LOG_FILE}")
## Agentic Development Process
This implementation was produced by an automated agentic pipeline:
- **Orchestration:** \`pipeline-orchestrator\` managed the full pipeline lifecycle
- **Architecture:** \`technical-architect\` analyzed the spec and produced the chunked plan
- **TDD:** \`test-suite-generator\` wrote failing tests before each chunk
- **Implementation:** \`feature-implementer\` implemented each chunk iteratively
- **Security:** \`code-security-reviewer\` reviewed each chunk with reflection loops
- **Integration:** \`integration-validator\` verified cross-service behavior
See \`.claude/pipeline/logs/\` for full execution details.
EOF
)" \
--base main \
--head "feature/${SPEC_SLUG}"The developer's only manual git step: review the draft PR and squash-merge to main.
Do not create GitHub Issues. The functional specs serve as the issue tracker. The chunks.json file serves as the task breakdown. The state.json file serves as the progress tracker.
Do not create per-chunk PRs. One PR per spec keeps the PR history clean and each PR represents a coherent, reviewable unit of functionality.
Do not give worker agents direct access to git. The feature-implementer should implement features, not manage version control. SCM operations belong in the orchestrator.
gitmust be configured with credentials that allow push access to the NAAS repository.ghCLI must be installed and authenticated (gh auth login).- Both tools should be verified during Phase 1 setup.
Each pipeline run produces a human-readable summary in .claude/pipeline/logs/. The orchestrator appends to this log after every Agent completion — see CONTRACTS.md Section 5 for the format and the specific entries written at each phase.
This log is a demonstration artifact. It shows that the pipeline ran, caught real issues, and resolved them autonomously. The orchestrator includes it in the PR description during the post-pipeline phase.
After the draft PR is created, the orchestrator emits a per-spec quality report that summarizes the run's defense-in-depth receipts:
# Pseudocode — actual implementation lives in .claude/pipeline/phases/post-pipeline.md
mkdir -p .claude/pipeline/reports
REPORT_FILE=".claude/pipeline/reports/${SPEC_SLUG}-quality-report.md"
# Generate from state.json and the execution log
# Format defined in .claude/pipeline/CONTRACTS.md Section 6The report is a durable, version-controlled artifact that records:
- Per-chunk metrics: tests written, implementation iterations, security review iterations, security issues caught
- Aggregate metrics: total tests, total reflection-loop firings, total HUMAN_REVIEW escalations
- Self-correction events: instances where the security review caught issues the implementer fixed without human intervention
- Defense-in-depth receipts: confirmation that iteration caps, budget guards, and regression checks operated as designed
- Time metrics: pipeline duration
The report serves three audiences: (a) the developer, who can scan it to confirm a clean run; (b) the code reviewer on the resulting PR, who can verify quality without reading the full execution log; (c) any future portfolio reviewer evaluating the agentic engineering discipline of the project. Schema details are in CONTRACTS.md Section 6.
The orchestrator's state machine is validated without invoking real worker agents (and without running shell commands, performing git operations, or modifying any code) by the pipeline-simulator agent. It stands in for the orchestrator, reads CONTRACTS.md and the phase files, and applies the same state-transition and log-formatting rules deterministically against predetermined agent outcomes — producing the full set of real state artifacts (state.json, chunks.json, the execution log, per-step state snapshots, and the various report files) under .claude/pipeline/simulation/runs/<scenario>/.
Three scenarios live in .claude/pipeline/simulation/scenarios/:
| Scenario | Exercises |
|---|---|
happy-path |
Every agent succeeds on the first attempt across all chunks — no retries or escalations. Validates basic phase progression, chunk sequencing, log formatting, and invocation counting. |
max-recovery |
Agents fail up to but not exceeding their iteration thresholds, maximizing automatic retries without triggering human escalation. Exercises every non-escalating retry path (impl/security iterations, the security-fix → regression-check → re-review loop). |
all-failures |
Every failure mode fires with human escalation at every step, plus the budget guard. Validates human_review transitions, chunk-phase retention during escalation, iteration resets on guidance, and the accept-risk commit path. |
The pipeline-simulator-run skill (/pipeline-simulator-run <scenario>, or all to run the three sequentially) is a thin wrapper: it invokes the pipeline-simulator agent for each requested scenario and persists the agent's final report to report.md in that scenario's run directory. All other simulation artifacts are written by the agent itself.
Add a section to the NAAS README describing the agentic development methodology:
- Link to
.claude/skills/pipeline-orchestrator/as the developer-facing entry point, and to.claude/agents/with brief descriptions of each worker agent's role - Link to
.claude/pipeline/CONTRACTS.mdfor the inter-agent communication protocols - Link to a sample pipeline execution log showing the reflection loop in action
-
The orchestrator manages the loop, workers do the work. All pipeline control flow lives in the orchestrator's state machine. Workers are stateless specialists invoked via the Agent tool — they receive context, do their job, and return results.
-
State is orchestrator-owned and inspectable. The
state.jsonfile has a single writer (the orchestrator), is always consistent, and enables resume after interruptions. A developer cancat state.jsonat any time to see exactly where the pipeline stands. -
Human escalation is a feature, not a failure. The pipeline surfaces hard problems to the developer rather than silently making bad decisions. A well-designed escalation path is more valuable than a fully autonomous loop that occasionally produces garbage.
-
Start with the simplest spec. Always test pipeline changes on Spec 1 (Event Ingestion) first. It's the smallest, most self-contained spec and will surface integration issues without wasting time on complex debugging.
-
Each enhancement is independently valuable. If time runs out after Priority 2, you still have a working automated pipeline. Priority 3 makes it smarter. Priority 4 makes one part faster.
-
Workers are self-contained. Workers receive everything they need in their Agent prompt — they don't read pipeline state files, parse other agents' output, or know about the broader pipeline context. This makes them independently testable and reusable outside the pipeline.
-
Let the artifacts tell the story. Structured commit messages, rich PR descriptions auto-generated from pipeline logs, and the execution logs themselves provide all the project management visibility a demonstration project needs.