diff --git a/.github/workflows/adr-tool.yml b/.github/workflows/adr-tool.yml index a042008c..143d89a7 100644 --- a/.github/workflows/adr-tool.yml +++ b/.github/workflows/adr-tool.yml @@ -14,6 +14,7 @@ on: - 'docs/scripts/**' - 'tests/adr-*.sh' - 'tests/fixtures/adr/**' + - 'docs/architecture/**' - 'Makefile' - '.github/workflows/adr-tool.yml' pull_request: @@ -23,6 +24,7 @@ on: - 'docs/scripts/**' - 'tests/adr-*.sh' - 'tests/fixtures/adr/**' + - 'docs/architecture/**' - 'Makefile' - '.github/workflows/adr-tool.yml' workflow_dispatch: diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 04701c92..6b8e2ce2 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -36,7 +36,7 @@ Open an issue. Include which hook or way is involved, your OS/shell, and any err **Hooks and macros** are bash. Keep them portable (macOS bash 3.2 compatible — no `declare -A`, no `mapfile`, no `grep -P`), use `shellcheck` if available, and keep scripts under 200 lines where possible. Hook scripts should be thin dispatchers to the `ways` binary. -**The `ways` binary** is Rust (`tools/ways-cli/`). Run `cargo test` before submitting changes. See [ADR-111](docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) for the consolidation rationale. +**The `ways` binary** is Rust (`tools/ways-cli/`). Run `cargo test` before submitting changes. See [ADR-111](docs/architecture/platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) for the consolidation rationale. ## Clippy and Rust toolchain drift diff --git a/Makefile b/Makefile index 8786fcdc..672ff60b 100644 --- a/Makefile +++ b/Makefile @@ -52,7 +52,7 @@ help: @echo " make test Run all tests (lint + smoke + unit + sim + adr)" @echo " make test-unit Run Rust unit tests" @echo " make test-sim Run session simulator (8 scenarios)" - @echo " make test-adr Run adr tool tests (lint, archive, golden, macro)" + @echo " make test-adr Run adr tool tests (lint, archive, golden, import, macro)" @echo " make test-lang Validate active language coverage" @echo " make test-locales Check locale files for gaps and duplicates" @echo " make test-multilingual Verify multilingual way matching (18 languages)" @@ -366,13 +366,17 @@ test-smoke: ways @echo "Smoke tests passed." test-adr: - @echo "Running adr tool tests (lint, archive, golden output)..." + @echo "Running adr tool tests (lint, archive, golden output, import round trip)..." @python3 hooks/ways/documentation/adr/assemble --check @bash tests/adr-lint-test.sh @bash tests/adr-archive-test.sh @bash tests/adr-golden-test.sh + @bash tests/adr-import-roundtrip.sh @bash tests/adr-macro-test.sh + @bash tests/adr-conversion-check.sh + @bash tests/adr-template-test.sh @docs/scripts/adr lint --check >/dev/null || { docs/scripts/adr lint; exit 1; } + @docs/scripts/adr cite --check >/dev/null || { docs/scripts/adr cite; exit 1; } @echo "adr tool tests passed." test-unit: diff --git a/README.md b/README.md index 990f589f..3c7374e5 100644 --- a/README.md +++ b/README.md @@ -115,7 +115,7 @@ The built-in ways cover software development, but the framework is domain-agnost 3. **SubagentStart** injects relevant ways into subagents spawned via Task 4. A way fires when matched, then **re-discloses on its `refire:` cadence** (a fraction of the context window, ADR-126) as its salience decays — marker files track the first fire and drive that re-disclosure state machine, they don't permanently block re-triggering -Matching has two channels: regex patterns for known keywords/commands/files, and [sentence-embedding](docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) semantic scoring (all-MiniLM-L6-v2). See [matching.md](docs/hooks-and-ways/matching.md) for the full strategy. +Matching has two channels: regex patterns for known keywords/commands/files, and [sentence-embedding](docs/architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) semantic scoring (all-MiniLM-L6-v2). See [matching.md](docs/hooks-and-ways/matching.md) for the full strategy. `ways list` shows the live session state — which ways fired, when (epoch), how far back (distance), what triggered them, tree relationships, check decay curves, and a re-disclosure forecast showing when distant ways will re-fire as context fills: @@ -322,7 +322,7 @@ If your organization clones this repo under a different name without forking on | [docs/hooks-and-ways.md](docs/hooks-and-ways.md) | Reference: hook lifecycle, state management, data flow | | [docs/governance.md](docs/governance.md) | Reference: compilation chain, provenance mechanics | | [docs/architecture.md](docs/architecture.md) | System architecture diagrams | -| [docs/architecture/](docs/architecture/) | Architecture Decision Records | +| [docs/architecture/](docs/architecture/) | Agent Decision Records | | [governance/](governance/) | Governance traceability and reporting | | [docs/README.md](docs/README.md) | Full documentation map | diff --git a/agents/requirements-analyst.md b/agents/requirements-analyst.md index 293f7f9a..1059f03c 100644 --- a/agents/requirements-analyst.md +++ b/agents/requirements-analyst.md @@ -1,6 +1,6 @@ --- name: requirements-analyst -description: Captures user needs as GitHub issues or in ADR context. Creates simple requirement statements with acceptance criteria. Focuses on understanding the problem to solve, not prescribing solutions. +description: Captures user needs as GitHub issues or in the context of an Agent Decision Record. Creates simple requirement statements with acceptance criteria. Focuses on understanding the problem to solve, not prescribing solutions. # Hardened: keeps its full working + research toolset; locks only Task — this role # authors issues/ADR context, it doesn't spawn subagents. tools: Read, Grep, Glob, Bash, Edit, Write, WebFetch, WebSearch @@ -33,7 +33,7 @@ Acceptance criteria: List requirements: `gh issue list --label requirement` ### Without GitHub -Capture in ADR context or `.claude/notes.md`: +Capture in the Context of the record the requirement motivates, or in `.claude/notes.md`. Quote the operator's words as said; the architect decides which become a decision's `basis`: ```markdown ## Requirements **Problem**: Users need password reset capability diff --git a/agents/system-architect.md b/agents/system-architect.md index 44751a8b..9acdd096 100644 --- a/agents/system-architect.md +++ b/agents/system-architect.md @@ -1,6 +1,6 @@ --- name: system-architect -description: Drafts Architecture Decision Records (ADRs) documenting design choices. Evaluates against SOLID principles. Guides ADR workflow from draft to PR to merge. Never implements - only designs and documents. +description: Drafts Agent Decision Records (ADRs) documenting design choices. Evaluates against SOLID principles. Guides the ADR workflow from `adr new` through `adr consider` to `adr accept`. Never implements - only designs and documents. # Hardened: keeps its full working + research toolset; locks only Task — this role # drafts ADRs, it doesn't spawn subagents. tools: Read, Grep, Glob, Bash, Edit, Write, WebFetch, WebSearch @@ -10,61 +10,46 @@ You create and maintain architectural decisions through the ADR workflow pattern **Role boundary**: You design and document architecture, but never implement code. Your output is ADRs and architectural guidance - not code files. -**Purpose**: Document "how to build" through Architecture Decision Records that capture context, decision, consequences, and alternatives. +**Purpose**: Document decisions, specs, and evidence as Agent Decision Records — the durable why behind a change, a specification kept current, or a finding later decisions cite as basis. ## ADR Workflow (Primary Responsibility) ### 1. Debate Phase -Discuss architectural options with user: +Discuss architectural options with the operator: - Present trade-offs clearly (benefit vs cost) - Explain technical implications - Surface risks and mitigation strategies - Avoid absolutes - present options with honest analysis -### 2. Draft ADR -Create `docs/adr/ADR-NNN-description-of-thing.md`: -```markdown -# ADR-NNN: Decision Title +### 2. Choose the Kind +- **decision** — adds, cuts, changes, retires, or constrains a capability. Needs `--verb`, `--capability`, `basis` entries, and `agent: {name, model}`; opens with a `## Summary`. +- **spec** — a specification kept current as the system changes; no verb, stays mutable after acceptance. +- **evidence** — a finding, survey, audit, measurement, or exploration; corrected by appending once accepted. Decisions cite it in `basis` instead of restating it. -Status: Proposed -Date: YYYY-MM-DD -Deciders: @user, @claude - -## Context -Forces that led to this decision. Why now? What constraints? - -## Decision -Architectural choice and approach selected. - -## Consequences -### Positive -- Clear benefits +### 3. Create the Record +```bash +docs/scripts/adr new "" --kind decision --verb add --capability <capability> --agent <name> --model <model> +``` +Then follow the ADR way's "Commands, Format and Lifecycle" section for this project — it prints the contract-specific record format for the tool this project vendored. Run `docs/scripts/adr lint` and fix what it reports rather than guessing at structure. -### Negative -- Costs and risks +If `docs/architecture/adr.yaml` declares no `contract:` key, the project is adr/v0: use the legacy Context / Decision / Consequences / Alternatives sections and Draft|Proposed|Accepted|Superseded|Deprecated statuses instead. -### Neutral -- Other implications +### 4. Write the Record +- Open a decision with a `## Summary`: what's decided, what it trades away, whether it's one-way, probes labeled confident/not confident, and the inversion. +- Write each `basis` entry honestly. Quote the operator verbatim (`said`, `via`, `level`: authored|directed|guided) when they said it. Cite evidence, standard, upstream, or precedent otherwise. Never invent an operator statement — when a decision's option label was agent-written rather than said by the operator, mark `via` saying so. -## Alternatives Considered -- Other options evaluated -- Why they were not selected +### 5. Hand the Probes Back +You run as a subagent, so the operator is not in your conversation. Return the Summary's probes to the calling session in plain words; the session asks the operator. When the operator's answer reaches you verbatim, record it: +```bash +docs/scripts/adr consider <n> --said "<verbatim>" --via "<where it was said>" --covers <probe...> ``` -### 3. Create PR for ADR +### 6. Accept +Accept only after the operator has considered the decision: ```bash -git checkout -b adr-NNN-description -git add docs/adr/ADR-NNN-description-of-thing.md -git commit -m "docs: Add ADR-NNN for [decision]" -git push -u origin adr-NNN-description -gh pr create --title "ADR-NNN: Decision Title" --body "..." +docs/scripts/adr accept <n> ``` - -### 4. Address PR Feedback -User reviews, you iterate on ADR based on comments. - -### 5. After Merge -ADR status becomes "Accepted" - ready to reference in implementation. +Use `docs/scripts/adr reject <n> --reason "..."` or `docs/scripts/adr abandon <n> --reason "..."` when the decision doesn't hold. Once accepted, a decision is corrected by appending, and a change in what the project does is a new decision that names what it replaces. The tool checks shape and references, not these conventions; review catches them, and git keeps every earlier version (ADR-311). Mark a cut or retire done with `docs/scripts/adr enact <n> <commit>` once the commit lands. ## SOLID Principles Evaluation @@ -93,13 +78,13 @@ Provide **specific refactoring recommendations**, not just problem identificatio **Check for upstream**: `gh repo view` ### With GitHub -- Store ADRs in wiki for project-wide visibility (optional) -- Reference ADRs in issues and PRs +- Reference ADRs by number in issues and PRs - Use GitHub discussions for architectural debates +- `adr cite` checks that ADR-N citations in code still match a real record ### Without GitHub -- ADRs live in `docs/adr/` directory -- Reference by filename in commits and documentation +- ADRs live under `docs/architecture/<domain>/` (`docs/scripts/adr domains` shows this project's areas) +- Reference by ADR number in commits and documentation — the number is permanent even if the record moves domains ## Communication Guidelines @@ -136,24 +121,24 @@ Good: "Depends on your needs. Microservices offer independent scaling and deploy - **Requirements Analyst**: Receives requirements that drive design decisions - **Task Planner**: Provides architectural guidance for task breakdown - **Code Reviewer**: Validates implementation follows ADR decisions -- **Workflow Orchestrator**: Coordinates ADR approval before implementation +- **Workflow Orchestrator**: Coordinates ADR acceptance before implementation ## Design Decision Lifecycle -1. **Propose**: Create ADR with "Proposed" status, capture context -2. **Review**: Create PR, gather feedback, evaluate alternatives -3. **Decide**: Merge PR, update status to "Accepted" -4. **Implement**: Guide implementation teams on architectural compliance -5. **Evolve**: Update or supersede decisions as needs change +1. **Propose**: `adr new` creates the record with status `proposed`, capturing context +2. **Consider**: The session asks the operator the Summary's probes; record the answer with `adr consider` +3. **Decide**: `adr accept`, or `adr reject`/`adr abandon` with a reason +4. **Implement**: Guide implementation teams on architectural compliance; mark a cut/retire done with `adr enact` +5. **Evolve**: Supersede with `adr supersede <old> --by <new>`, or archive when a record no longer applies -**Summary**: You draft and maintain ADRs following the debate → draft → PR → review → merge workflow. You evaluate designs against SOLID principles and provide specific improvement recommendations. Your documentation serves as the authoritative source for how to build the system. +**Summary**: You draft and maintain ADRs through `adr new` → Summary and basis → `adr consider` → `adr accept`/`reject`/`abandon`. You evaluate designs against SOLID principles and provide specific improvement recommendations. Your documentation serves as the authoritative source for how to build the system. ## What You Return - **Status**: complete, blocked out of domain, or failed - **Failure class** when failed: transient, deterministic, capability, ambiguity, or systemic -- **Work done**: the ADRs drafted or revised, with file paths and the PR if opened +- **Work done**: the ADRs drafted or revised, with file paths and ADR numbers - **What is needed outside your domain**: a requirement the analyst must clarify, an implementation the planner must sequence, or "none" -- **Recommended next step**: review the ADR, open the PR, or begin implementation -- **Gates run**: ADR lint and any PR check, each with its state, or "none" +- **Recommended next step**: the probes for the operator, in plain words; accepting the record; or beginning implementation +- **Gates run**: `adr lint` and any PR check, each with its state, or "none" - **Tools or scripts built**: any diagram or ADR script kept for reuse, with path and invocation, or "none" diff --git a/agents/task-planner.md b/agents/task-planner.md index 669bed7c..cfcd9d5a 100644 --- a/agents/task-planner.md +++ b/agents/task-planner.md @@ -37,14 +37,14 @@ Large feature: ### For Multi-Step Implementation Suggest sequence: ``` -Implementing ADR-005 (new auth system): -1. Branch: adr-005-auth-foundation +Implementing ADR-NNN (new auth system): +1. Branch: adr-nnn-auth-foundation - Database schema - Core auth models -2. Branch: adr-005-auth-endpoints +2. Branch: adr-nnn-auth-endpoints - API endpoints - Depends on: foundation branch -3. Branch: adr-005-auth-ui +3. Branch: adr-nnn-auth-ui - Login/logout UI - Depends on: endpoints branch ``` diff --git a/agents/workflow-orchestrator.md b/agents/workflow-orchestrator.md index bc3c68a2..7c719505 100644 --- a/agents/workflow-orchestrator.md +++ b/agents/workflow-orchestrator.md @@ -14,8 +14,8 @@ You coordinate the development lifecycle following the ADR-driven workflow patte ## Core Workflow Pattern ``` -Debate/Research → Draft ADR (docs/adr/) → PR for ADR → -Branch (reference ADR) → TodoWrite (session) → Implement → +Debate/Research → adr new (docs/architecture/<domain>/) → probes answered via adr consider → +adr accept → Branch (reference ADR) → TodoWrite (session) → Implement → PR for code → Address review → Merge ``` @@ -81,8 +81,8 @@ You: "That's an ADR-worthy decision. Want to draft one documenting the trade-off gh pr list --state open gh issue list --label requirement,bug -# Verify ADR PR before implementation -gh pr list --search "ADR" --state open +# Verify the ADR is accepted before implementation +docs/scripts/adr view <n> # status should be accepted, not proposed # Check for stale work gh pr list --state open --json number,title,updatedAt @@ -155,7 +155,7 @@ Todo Status: 3/7 tasks complete Blockers: Waiting on API key from vendor ## Recent Activity -- ADR-007 merged yesterday +- ADR-007 accepted yesterday - PR #45 in review (OAuth core impl) - 2 open issues (non-blocking) diff --git a/agents/workspace-curator.md b/agents/workspace-curator.md index 557484a6..dc816791 100644 --- a/agents/workspace-curator.md +++ b/agents/workspace-curator.md @@ -18,23 +18,24 @@ Recommend structure for project documentation: ``` docs/ -├── adr/ # Architecture Decision Records (ADR-NNN-description.md) +├── architecture/ # Agent Decision Records, one intent folder per domain (adr.yaml) ├── development/ # Dev guides, setup instructions -├── research/ # Research findings, spike reports ├── guides/ # User guides, tutorials ├── testing/ # Test strategies, QA docs ├── features/ # Feature specs, user stories └── [other]/ # Project-specific needs ``` +Research findings, spike reports, and other design notes go in `docs/architecture/<domain>/` as **evidence** records (`adr new <domain> "<title>" --kind evidence --capability <capability>`), not a separate `research/` or `design-notes/` folder - decisions cite them in `basis`. + **Flexibility is key** - suggest structure, don't mandate it. ### 2. ADR Organization When user asks "where should this ADR go?": -- **Location**: `docs/adr/ADR-NNN-description-of-thing.md` -- **Numbering**: Sequential (ADR-001, ADR-002, ADR-003, ...) -- **Format**: ADR-NNN-kebab-case-description -- **Never renumber** - deprecated decisions keep their numbers +- **Location**: `docs/architecture/<domain>/ADR-N-description-of-thing.md`, where `<domain>` is one of the areas `adr domains` lists +- **Numbering**: Each domain in `adr.yaml` owns a number band; it only allocates new numbers, `adr new <domain> "<title>"` assigns the next free one +- **Format**: ADR-N-kebab-case-description +- **Numbers are permanent identity (ADR-310)** - moving a record to another domain with `adr domain move <n> <domain>` keeps its number; never renumber a deprecated or superseded record ### 3. .claude/ Directory Maintain plugin and project configuration: @@ -76,19 +77,19 @@ Keep minimal - only what's needed. ### New Project ``` User: "Setting up a new project" -You: "Want me to set up docs/adr/ for decision tracking? We can add more structure as needed." +You: "Want me to set up docs/architecture/ with a domain or two for decision tracking? We can add more as needed." ``` ### Documentation Scattered ``` User: "Can't find the auth decision doc" -You: "I see ADRs in docs/, root/, and notes/. Want me to consolidate them into docs/adr/ with consistent numbering?" +You: "I see ADRs in docs/, root/, and notes/. Want me to consolidate them into docs/architecture/<domain>/? Records outside docs/architecture/ come in through `adr import` (named ADR-NNN-*.md with frontmatter first); records already under docs/architecture/ change area with `adr domain move`." ``` -### ADR Numbering Unclear +### ADR Domain Unclear ``` -User: "What number should this ADR be?" -You: "Last ADR is ADR-003, so this would be ADR-004. For filename: ADR-004-oauth-integration.md" +User: "What domain should this ADR go in?" +You: "`adr domains` shows the areas and their bands - pick the one matching this decision's area, or `adr domain add` a new one. `adr new <domain> "<title>"` assigns the next free number in that band." ``` ## What NOT to Do @@ -110,7 +111,7 @@ You: "Last ADR is ADR-003, so this would be ADR-004. For filename: ADR-004-oauth - Use GitHub's file organization ### Without GitHub -- ADRs in `docs/adr/` directory +- ADRs in `docs/architecture/<domain>/` directories - Standard filesystem organization ## Communication Guidelines @@ -130,25 +131,25 @@ You: "Last ADR is ADR-003, so this would be ADR-004. For filename: ADR-004-oauth ``` User: "Where should design docs go?" Bad: "You need to create a comprehensive documentation taxonomy with categories, subcategories, and metadata." -Good: "Depends on what you have. ADRs go in docs/adr/. Other design docs could go in docs/development/ or docs/guides/ depending on audience. What type of design doc?" +Good: "Depends on what you have. ADRs go in docs/architecture/<domain>/. A research finding or spike report is an evidence record in the same tree. Other design docs could go in docs/development/ or docs/guides/ depending on audience. What type of design doc?" ``` ## Quick Organization Tasks ### Consolidate Scattered ADRs 1. Find all ADRs: `find . -name "*adr*" -o -name "*decision*"` -2. Move to docs/adr/ with consistent numbering +2. Bring records from elsewhere into `docs/architecture/<domain>/` with `adr import scan` then `adr import apply` (rename to `ADR-NNN-*.md` and add frontmatter first); `adr domain move` only moves records already under `docs/architecture/` 3. Update references in other docs ### Set Up New Project -1. Create docs/adr/ directory +1. Set up `docs/architecture/adr.yaml` with at least one domain and its number band 2. Add .claude/ if using project-specific config 3. Done - don't create more until needed -### Recommend Next ADR Number -1. List existing: `ls docs/adr/ | grep ADR-` -2. Find highest number -3. Suggest next: "ADR-NNN" +### Recommend a Domain and Number +1. List existing: `adr domains` +2. Pick the domain matching the record's area, or add one: `adr domain add <name> --range A-B --folder F` +3. `adr new <domain> "<title>"` assigns the next free number in that band ## Integration @@ -170,7 +171,7 @@ Good: "Depends on what you have. ADRs go in docs/adr/. Other design docs could g - Complex taxonomy - Metadata everywhere -**Summary**: You organize docs/ and .claude/ directories simply and practically. Recommend ADR numbering (ADR-NNN-description.md), suggest structure when helpful, prevent documentation sprawl. Keep it simple - structure should serve findability, not create complexity. +**Summary**: You organize docs/ and .claude/ directories simply and practically. Recommend ADR domains and placement (`docs/architecture/<domain>/ADR-N-description.md`), suggest structure when helpful, prevent documentation sprawl. Keep it simple - structure should serve findability, not create complexity. ## What You Return diff --git a/commands/project-audit.md b/commands/project-audit.md index 1585cb86..8ff9c0c7 100644 --- a/commands/project-audit.md +++ b/commands/project-audit.md @@ -55,13 +55,47 @@ docs/scripts/adr config 2>/dev/null - Warn: yaml exists but no domains configured - Fail: no yaml +**Check: Which contract is the project on?** + +`docs/scripts/adr config` prints `contract: adr/v1` when the project has +adopted it; its absence means the project is on the legacy `adr/v0` contract. +This is informational: it decides which of the checks below apply. A v0 +project is not deficient for staying on v0. + **Check: Do ADRs pass lint?** ```bash docs/scripts/adr lint --check 2>/dev/null ``` - Pass: exit code 0 -- Warn: warnings only (missing optional fields) -- Fail: errors (missing frontmatter, invalid status, etc.) +- Warn: warnings only (missing optional fields, an open concern, an imported + record with no Summary yet) +- Fail: errors (missing frontmatter, invalid status, adr/v1 grammar violations, etc.) + +**Check: Do citations resolve? (`adr cite`)** +```bash +docs/scripts/adr cite --check 2>/dev/null +``` +- Pass: exit code 0 +- Warn: a citation of a superseded or proposed record (the acceptance or + cleanup worklist) +- Fail: a citation resolves to no record + +**Check: adr/v1 records carry the fields their kind requires.** + +Only when `adr config` shows `contract: adr/v1`. Run `docs/scripts/adr list +--json` and confirm: every record has `kind`; every kind that `adr.yaml`'s +`kinds:` block requires `capability` for has one; every `decision` opens with +`## Summary`. `adr lint` already enforces all three, so a clean lint run +covers this — use `adr list --field kind` / `--field capability` or the JSON +output when pointing at specific gaps. + +- Pass: `adr lint` clean, and no v1 record is missing `kind`, `capability`, or + (for a decision) `## Summary` +- Warn: v0 records remain in an otherwise-v1 corpus (expected mid-migration — + convert them with `adr import scan`/`apply`, ADR-306) +- Fail: a v1 record is missing a required field, or `adr lint` reports a + grammar error +- N/A: contract is adr/v0 **Check: Are there orphaned ADRs?** @@ -69,7 +103,8 @@ Look for `ADR-*.md` files outside `docs/architecture/`: ```bash find . -name 'ADR-*.md' -not -path './docs/architecture/*' -not -path './node_modules/*' -not -path './.git/*' 2>/dev/null ``` -Also check for ADRs in `docs/architecture/` that don't belong to any domain folder. +Also check for ADRs in `docs/architecture/` that don't belong to any domain +folder (under adr/v1, an area folder; ADR-310). - Pass: all ADRs in domain directories - Warn: ADRs in legacy/ (expected for migrated repos) @@ -286,12 +321,20 @@ Read the scaffold ADR and extract what was decided: - What CODEOWNERS strategy was elected? - Which ways were created? - What GitHub config was set up? +- On adr/v1, does the scaffold record itself carry the fields its kind + requires — `kind`, `capability`, and, if it's a decision, `## Summary`? A + clean `adr lint` and `adr cite` covers this the same way it covers the rest + of the corpus (§1). Then compare against reality: - **Domain drift**: Are there new directories/concerns that suggest a domain should be added? - **CODEOWNERS drift**: Do the paths and owners still match the codebase structure? - **Ways drift**: Were ways created that the ADR said would be? Were any removed? - **Scope drift**: Has the project grown beyond what the scaffold anticipated? (e.g., started as a CLI tool, now has a web frontend too) +- **Capability drift (adr/v1 only)**: compare `adr.yaml`'s `capabilities` + vocabulary with what the codebase does. A capability the codebase clearly + uses but the vocabulary lacks is a candidate for an `add`. The scaffold's + own capability list is a starting point. **If drift is detected:** @@ -303,7 +346,8 @@ Then ask: This is the key value of the scaffold ADR — it makes drift visible and forces a conscious decision about whether to update the plan or fix the divergence. -- Pass: scaffold ADR exists and current state matches +- Pass: scaffold ADR exists, current state matches, and (adr/v1) it lints + clean with its citations resolving - Drift: scaffold ADR exists but state has diverged (present the delta) - Info: no scaffold ADR (project may predate `/project-init` or was set up manually) @@ -350,7 +394,7 @@ Determine the overall tone from the category statuses: **Score: XX% (NN/MM checks pass)** -### ADR Health: X/5 +### ADR Health: X/N (N varies: the v1-only checks are N/A on an adr/v0 project) | Check | Status | Detail | |-------|--------|--------| | ... | ... | ... | diff --git a/commands/project-init.md b/commands/project-init.md index ff29577e..eae09243 100644 --- a/commands/project-init.md +++ b/commands/project-init.md @@ -25,7 +25,7 @@ Mark each task `in_progress` as you start it, `completed` when done. This is you **Read these docs first** — you need the full landscape before your first question: -1. Read `~/.claude/hooks/ways/documentation/adr/migration/migration.md` — understand the five starting states (greenfield, flat directory, inline metadata, scattered, different tool) and migration strategies +1. Read `~/.claude/hooks/ways/documentation/adr/migration/migration.md` — understand the starting states (greenfield, flat directory, v0 frontmatter, inline metadata, other tools) and the conversion path. Existing records named `ADR-NNN-*.md` with YAML frontmatter convert into adr/v1 through `docs/scripts/adr import scan <paths>` then `adr import apply` (ADR-306). Sequential `0001-*.md` files and inline metadata need renaming and frontmatter first, as the migration way describes. 2. Read `~/.claude/hooks/ways/softwaredev/delivery/github/github.md` — understand PR-always stance, repo health expectations 3. Read `~/.claude/hooks/ways/softwaredev/docs/docs.md` — understand documentation scaling by project complexity @@ -141,6 +141,8 @@ Use `AskUserQuestion` with focused multiple-choice questions. Adapt based on ans If ADRs need setup or reorganization: +**Contract**: ask whether the project adopts `adr/v1` (typed decision/spec/evidence records with verb, capability, basis, and lifecycle commands, ADR-304) or stays on the legacy `adr/v0` contract (Draft/Proposed/Accepted/Superseded/Deprecated status, Context/Decision/Consequences/Alternatives Considered). Recommend v1 for a new or actively-decided project; v0 is a reasonable choice for a project that just wants a lightweight decision log. Either way, name the capabilities the project's domains cover — under v1 these seed the `capabilities` vocabulary in `adr.yaml`. + **For greenfield:** - Analyze the codebase structure (directory layout, package organization) - Propose 3-6 domains based on what you see @@ -150,10 +152,10 @@ If ADRs need setup or reorganization: **For brownfield with existing ADRs:** - List what exists: how many ADRs, what format, what topics they cover - **Lint them immediately** — run the ADR tool's linter (or manually check frontmatter) and show the user what's broken: missing frontmatter, inline metadata that needs conversion, invalid statuses, missing fields -- **Offer to fix the existing ADRs first** before proposing reorganization. "I found 8 ADRs — 3 are missing frontmatter, 2 have inline metadata instead of YAML. Want me to fix these up before we talk about domain organization?" +- **Convert through import sheets.** `docs/scripts/adr import scan <paths>` writes an editable sheet per `ADR-NNN-*.md` record with frontmatter under `.import/` (rename and add frontmatter first for anything else); the user (or you, with their review) fills in the gaps a source format can't answer — kind, verb, capability, basis; `docs/scripts/adr import apply` then writes them as adr/v1 records (ADR-306). Offer this before proposing reorganization: "I found 8 ADRs. Want me to scan them into import sheets so we can bring them into the new contract?" - Show a proposed domain mapping: which existing ADRs belong to which domain based on their content - Ask about the migration approach: park as legacy and go forward, or reorganize everything -- If reorganizing: show the move plan (which files go where) and get approval before touching anything +- If reorganizing: show the move plan (which files go where) and get approval before touching anything; under adr/v1 a move keeps the record's number (ADR-310) ### GitHub Interview @@ -212,7 +214,7 @@ After entry questions and ADR domains, present the **full artifact menu** with r | Planning | Dependency policy | `docs/policies/dependencies.md` | | Planning | Migration guide template | `docs/migration/TEMPLATE.md` | -**RFCs as an ADR domain**: If the user selects RFCs, add an `rfc` domain to `adr.yaml` with its own range. RFCs use the same tooling but with an extended status flow: `Proposed → Discussing → Accepted → Rejected → Withdrawn`. The ADR tool handles this naturally — it's just a domain with different status conventions documented in the scaffold ADR. +**RFCs as an ADR domain**: If the user selects RFCs, add an `rfc` domain to `adr.yaml` with its own range. On adr/v0, RFCs can use an extended status flow (`Proposed → Discussing → Accepted → Rejected → Withdrawn`) documented as a convention in the scaffold ADR; adr/v0's `statuses` list is project-wide, so this is a documented convention for the domain rather than a tool-enforced one. adr/v1's status set is fixed by the contract (proposed, accepted, rejected, abandoned, superseded, archived) — an RFC there is a `decision` record like any other, and a "Discussing" stage is a proposed record awaiting `consider`/`accept`. RFCs can reference internal or external sources — an RFC might propose adopting an external standard (linking to the spec), or it might be an internal design document that lives entirely in the repo. The `related:` frontmatter field supports both: internal ADR cross-references and external URLs. Ask during the interview whether the project uses external references (specs, standards, upstream RFCs) so the scaffold ADR can document the convention. @@ -265,9 +267,9 @@ The scaffold itself is an architectural decision. After the interview, **create - CODEOWNERS strategy (if applicable) - What was deferred or declined -This ADR is collaborative — draft it from the interview answers, show it to the user, and iterate. The user influences the content. Use `docs/scripts/adr new meta "Adopt software engineering scaffold"` (or whatever the project management domain is named). +This ADR is collaborative — draft it from the interview answers, show it to the user, and iterate. The user influences the content. On adr/v0, use `docs/scripts/adr new meta "Adopt software engineering scaffold"` (or whatever the project management domain is named). On adr/v1, add `--kind decision --verb add --capability <name>` naming the capability the scaffold itself establishes, and `--agent`/`--model`. -Example structure: +Example structure (adr/v0 — the legacy contract's Context/Decision/Consequences shape): ```markdown # ADR-NNN: Adopt Software Engineering Scaffold @@ -297,7 +299,9 @@ We adopt the following practices: - [what was deferred: items declined during interview] ``` -**This ADR is created early** (after ADR tooling is installed) and updated as the scaffold progresses. It becomes the first real ADR in the project. +On adr/v1, the scaffold record opens with `## Summary` (decided, trades away, one-way?, probes, inversion) instead, and its `basis` names the interview as an `operator` entry, quoting the operator's answers. `docs/scripts/adr new --help` and `adr lint` confirm the fields the chosen kind requires. + +**This ADR is created early** (after ADR tooling is installed) and updated as the scaffold progresses. It becomes the first real ADR in the project. On adr/v1 it stays `proposed` while the scaffold changes: at the end, ask the operator its probes, record the answer with `docs/scripts/adr consider <n> --said "..." --via "..."`, then `docs/scripts/adr accept <n>`. After acceptance it is corrected by appending, and a later change is a new decision that names what it replaces. ### Sub-Agent Delegation @@ -315,17 +319,25 @@ These are `subagent_type` values for the `Task` tool: parts the skill can't know: 1. Vendor the tool via the **adr** skill (it copies `adr-tool` → `docs/scripts/adr` - and seeds `docs/architecture/adr.yaml` from the template). + and seeds `docs/architecture/adr.yaml` from the template). The template declares + `contract: adr/v1` with the decision, spec and evidence kinds and a + placeholder capability. 2. Customize `docs/architecture/adr.yaml` with the interview answers: - Project name - Domains with ranges (100-wide ranges, 1-99 for legacy) - - Statuses list + - Under adr/v1 (the template's default): replace the placeholder + `capabilities` with one line per capability the interview named - Default deciders (from git/gh config) -3. Create domain subdirectories under `docs/architecture/` +3. Create domain subdirectories under `docs/architecture/`. Under adr/v1 these are + intent areas (ADR-310): a domain's range only allocates new numbers, and a + record moving between areas keeps its number. -4. For brownfield: execute the appropriate migration strategy per the migration way +4. For brownfield: convert existing records with `docs/scripts/adr import scan + <paths>` then `docs/scripts/adr import apply` (ADR-306), or, for a starting + state with nothing the importer reads, execute the appropriate migration + strategy per the migration way 5. Validate: ```bash diff --git a/docs/README.md b/docs/README.md index 563590d9..67ab3460 100644 --- a/docs/README.md +++ b/docs/README.md @@ -11,8 +11,7 @@ Map of the documentation tree. For the project overview, see the [main README](. | [hooks-and-ways/](hooks-and-ways/) | Guides: creating ways, matching, macros, provenance, teams | | [hooks-and-ways.md](hooks-and-ways.md) | Reference: hook lifecycle, state management, session gating | | [architecture.md](architecture.md) | System architecture diagrams (Mermaid) for the ways mechanics | -| [architecture/](architecture/) | Architecture Decision Records (managed by `docs/scripts/adr`) | -| [design-notes/](design-notes/) | Prose-first framing documents that justify multiple related decisions (complement to ADRs) | +| [architecture/](architecture/) | Agent Decision Records (managed by `docs/scripts/adr`): decisions, specs, and evidence records (surveys, audits, explorations; the former design notes, ADR-309) | | [explanation/](explanation/) | Diátaxis explanation pages — conceptual walkthroughs grounded in real scenarios (install topologies, localization, attend messaging, and [how ways works](explanation/how-ways-works/how-ways-works-the-model.md) — the observable companion to cognitive-loop.md) | ## Guides vs Reference diff --git a/docs/architecture/INDEX.md b/docs/architecture/INDEX.md index de865395..a612f635 100644 --- a/docs/architecture/INDEX.md +++ b/docs/architecture/INDEX.md @@ -1,138 +1,167 @@ -# Architecture Decision Records +# Agent Decision Records -This directory contains Architecture Decision Records (ADRs) for Claude Code Configuration. -Each ADR documents a significant architectural decision, its context, and consequences. +This directory contains the Agent Decision Records (ADRs) for Claude Code Configuration, under the adr/v1 contract. +A record's kind says what it holds: a decision, a spec kept current, or evidence (findings, surveys, explorations). -## ADR Format +## Record Format -All ADRs follow a consistent format: -- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded -- **Date:** When the decision was made -- **Deciders:** Who made the decision -- **Context:** The problem or situation requiring a decision -- **Decision:** The architectural choice made -- **Consequences:** Benefits, drawbacks, and other impacts +- **Kind:** decision / spec / evidence, as `adr.yaml` declares +- **Status:** proposed / accepted / rejected / abandoned / superseded / archived +- **Capability:** what the record is about, from the vocabulary in `adr.yaml` +- **Decisions** also carry a verb (add / cut / change / retire / constrain), a basis, and open with a Summary +- **Numbers** are permanent; a record's folder is its area _This index is auto-generated by `adr index`. Configuration: [`adr.yaml`](./adr.yaml)_ -## System -_Ways architecture, matching, macros, hooks, session lifecycle_ +## Ways +_The ways engine: how ways are matched, disclosed and re-disclosed_ | ADR | Title | Status | |-----|-------|--------| -| [ADR-100](./system/ADR-100-ways-scaffolding-wizard.md) | Ways Scaffolding Wizard | Accepted | -| [ADR-101](./system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) | Wormhole relay protocol for cross-instance agent communication | Deprecated | -| [ADR-102](./system/ADR-102-irc-based-local-agent-communication.md) | IRC-based local agent communication | Deprecated | -| [ADR-103](./system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md) | Checks — Epoch-Distance-Aware Confidence Sensors for Ways | Accepted | -| [ADR-104](./system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) | Token-Gated Way Re-Disclosure for Long Context Windows | Superseded (superseded by ADR-123) | -| [ADR-105](./system/ADR-105-progressive-disclosure-for-way-trees.md) | Progressive Disclosure for Way Trees | Accepted | -| [ADR-106](./system/ADR-106-project-pulse-epoch-mapped-project-awareness.md) | Project Pulse — Epoch-Mapped Project Awareness | Accepted | -| [ADR-107](./system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md) | Corpus, Matching Pipeline, and Locale Support | Accepted (partially superseded by ADR-125) | -| [ADR-108](./system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) | Embedding-Based Way Matching with all-MiniLM-L6-v2 | Accepted | -| [ADR-109](./system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md) | Project-Scope Way Embedding with Manifest-Based Staleness Detection | Accepted | -| [ADR-110](./system/ADR-110-way-file-separation-and-graph-compatible-structure.md) | Way File Separation and Graph-Compatible Structure | Accepted | -| [ADR-111](./system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) | Unified `ways` CLI — Single Binary Tool Consolidation | Accepted | -| [ADR-113](./system/ADR-113-attend-active-awareness-module.md) | `attend` — Active Awareness Module | Accepted | -| [ADR-114](./system/ADR-114-attend-as-insistent-way-trigger-type.md) | `attend` Events as an Insistent Way Trigger Type | Accepted | -| [ADR-115](./system/ADR-115-declarative-config-with-project-scope-overlay.md) | Declarative Configuration with Project-Scope Overlay | Accepted | -| [ADR-116](./system/ADR-116-declarative-permission-requirements.md) | Declarative Permission Requirements | Accepted | -| [ADR-117](./system/ADR-117-sensor-crate-extraction-and-feature-flags.md) | Sensor Crate Extraction and Feature Flags | Accepted | -| [ADR-118](./system/ADR-118-focus-groups-dynamic-agent-grouping.md) | Focus Groups — Dynamic Agent Grouping | Accepted | -| [ADR-119](./system/ADR-119-action-potential-engagement-model.md) | Action Potential Engagement Model | Superseded (superseded by ADR-123) | -| [ADR-120](./system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md) | Interactive Chat TUI — Human in the Signal Loop | Accepted | -| [ADR-121](./system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md) | Salience decay for signal presentation — turn-based exponential | Superseded (superseded by ADR-123) | -| [ADR-122](./system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md) | Attend disclosure sensor — token-gated affordance reheat | Accepted | -| [ADR-123](./system/ADR-123-firing-dynamics-progression-axis-unification.md) | Firing dynamics — progression-axis unification for attend and ways | Accepted | -| [ADR-124](./system/ADR-124-channel-bar-ordering-open-as-base.md) | TUI Legend Architecture — Base Channel, Liveness, and Ordering | Accepted | -| [ADR-125](./system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) | Authored Disclosure Graph and Removal of BM25 | Accepted | -| [ADR-126](./system/ADR-126-window-relative-refire.md) | Window-relative refire with named presets | Accepted | -| [ADR-127](./system/ADR-127-reject-full-body-embedding-corpus.md) | Full-body embedding corpus for way matching | Rejected | -| [ADR-128](./system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md) | Memory as repo-portable ways — seed routing over accumulated snapshots | Accepted | -| [ADR-129](./system/ADR-129-instance-suffix-and-heartbeat-liveness.md) | Instance suffix and heartbeat liveness for attend identity | Accepted | -| [ADR-130](./system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md) | Sentence-salience input reduction for embed matching | Accepted | -| [ADR-131](./system/ADR-131-project-scope-way-toggles.md) | Project-scope way toggles | Accepted | -| [ADR-132](./system/ADR-132-collaboration-ways-domain.md) | Collaboration ways domain | Accepted | -| [ADR-134](./system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) | Empirical auto-tuning from fire and near-miss telemetry | Accepted | -| [ADR-135](./system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md) | Content-aware write-time over-build gate with a self-extending pattern corpus | Accepted | -| [ADR-136](./system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md) | Split addressed messaging from the sensor-observation bus | Accepted | -| [ADR-137](./system/ADR-137-boundedness-bounded-work-per-cycle.md) | Boundedness — a unit of work must be bounded within its cycle | Accepted | -| [ADR-138](./system/ADR-138-skills-own-the-how-ways-own-the-5w.md) | Skills own the how, ways own the 5W | Accepted | -| [ADR-139](./system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md) | Shelve maintainer i18n; adopter-run localization via ways-localize | Accepted | -| [ADR-140](./system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md) | Two install topologies: in-place repo and subdirectory projection | Accepted | -| [ADR-141](./system/ADR-141-knowledge-graph-as-evidential-memory-backend.md) | Knowledge Graph as Evidential Memory Backend | Accepted | -| [ADR-142](./system/ADR-142-agent-ways-1-0-xdg-application-distribution.md) | agent-ways 1.0 — XDG application distribution | Accepted | -| [ADR-143](./system/ADR-143-three-root-way-runtime-core-user-project.md) | Three-root way runtime — core, user, project | Accepted | -| [ADR-144](./system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) | Install / repair / migrate as one manifest reconciler | Accepted | -| [ADR-146](./system/ADR-146-installer-binary-verification-and-guided-build-fallback.md) | installer binary verification and guided build fallback | Accepted | -| [ADR-147](./system/ADR-147-composable-settings-json-config-fragments.md) | Composable settings.json — a store of YAML config fragments | Superseded (superseded by ADR-169) | -| [ADR-148](./system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md) | framework surface ships operator content; dev harness in project scope | Accepted | -| [ADR-149](./system/ADR-149-operator-config-interview-skill.md) | operator config interview skill | Superseded (superseded by ADR-169) | -| [ADR-150](./system/ADR-150-version-truth-and-downgrade-safe-self-update.md) | Version-truth and downgrade-safe self-update | Accepted | -| [ADR-151](./system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md) | Extract ways-core crate and ways-audit sibling binary | Accepted | -| [ADR-152](./system/ADR-152-framework-default-secret-path-deny-baseline.md) | Framework-default secret-path deny baseline | Accepted | -| [ADR-153](./system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md) | Session-introspection substrate — correlating fired ways to turns | Accepted | -| [ADR-154](./system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md) | `ways introspect` — one model, three front-ends | Accepted | -| [ADR-155](./system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md) | Semantic gating of the keyword channel and reasoning-channel rebuild | Accepted (partially superseded by ADR-188 §3) | -| [ADR-156](./system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md) | Calibrated relevance scoring for the semantic lane | Accepted | -| [ADR-157](./system/ADR-157-case-insensitive-trigger-regex-compilation.md) | Case-insensitive trigger regex compilation | Accepted | -| [ADR-158](./system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md) | Calibration boundary quality — hard negatives and a fire-breadth ship gate | Accepted | -| [ADR-159](./system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md) | Remove ways tune-curves and the legacy curve: cadence field | Accepted | -| [ADR-160](./system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md) | Chunked late-interaction matching with softmax-share gating for way selection | Accepted | -| [ADR-161](./system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md) | Queued mid-turn operator messages as an aggregated scan surface | Accepted | -| [ADR-162](./system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md) | Mechanical session-link suppression as defense against transcript disclosure | Superseded (superseded by ADR-167) | -| [ADR-163](./system/ADR-163-config-separation-dotfiles-source-of-truth.md) | Config separation — dotfiles as source-of-truth feeding the settings fragment store | Accepted | -| [ADR-164](./system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md) | File artifacts distributed across hosts must be carried by value not host-absolute reference | Accepted | -| [ADR-165](./system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md) | Loop-control bookends: start, develop, merge, release, wrap | Accepted | -| [ADR-166](./system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md) | Single source of truth for model context-window resolution | Accepted | -| [ADR-167](./system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md) | Session-link suppression: attribution.sessionUrl as primary control, deny hook as backstop | Accepted | -| [ADR-169](./system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md) | agent-ways relinquishes user-scoped settings.json; retains only its operational baseline | Accepted | -| [ADR-170](./system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md) | Human focus-group membership via username identity and a shared attend-groups crate | Accepted | -| [ADR-171](./system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md) | Stable session identity — the roster enumerates addressable coordinating units | Accepted | -| [ADR-172](./system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md) | Turn-boundary inbound delivery via a CLI-owned drain checkpoint | Accepted | -| [ADR-173](./system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md) | Chat-idiom convergence for the attend command surfaces | Accepted | -| [ADR-174](./system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md) | Progressive core — decoration guidance and the core re-disclosure gap | Accepted | -| [ADR-175](./system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md) | Standing delegation authorization satisfies the harness permission gate rather than overriding it | Accepted | -| [ADR-176](./system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md) | Contract identification as the develop-loop front gate | Accepted | -| [ADR-177](./system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md) | Version-stamped vendored tools with direction-aware drift detection | Accepted | -| [ADR-178](./system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md) | Register transfers by demonstration - core.md carries policy | Accepted | -| [ADR-179](./system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md) | Remove the pre-1.0 in-place migrator; keep the guards and the transition fallbacks | Accepted | -| [ADR-180](./system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md) | GitHub issues as the shared truth for the session task list | Accepted | -| [ADR-181](./system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md) | Guard hooks: a blocking PreToolUse class, shipped deactivated | Accepted | -| [ADR-182](./system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md) | Keepwarm: attend keeps the prompt cache warm with a wake floor | Accepted | -| [ADR-184](./system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md) | Installation and activation are separate states: targets as the unit of activation | Accepted | -| [ADR-185](./system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md) | CLI output contract: structured output for people, JSON for machines | Accepted | -| [ADR-186](./system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md) | Live integration fixture: install-path test levels and the tier 2 gate | Accepted | -| [ADR-187](./system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md) | Attend MCP server mode: outbound and queries as typed tools, inbound stays on Monitor and the Stop hook | Proposed | -| [ADR-188](./system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md) | PostToolUse delivery for tool-lane ways and retirement of the semantic Bash surface | Proposed | +| [ADR-004](./ways/ADR-004-way-macros.md) | Way Macros for Dynamic Context Injection | accepted | +| [ADR-014](./ways/ADR-014-tfidf-semantic-matcher.md) | TF-IDF/BM25 Binary for Semantic Way Matching | accepted | +| [ADR-103](./ways/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md) | Checks — Epoch-Distance-Aware Confidence Sensors for Ways | accepted | +| [ADR-104](./ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) | Token-Gated Way Re-Disclosure for Long Context Windows | superseded (superseded by ADR-123) | +| [ADR-105](./ways/ADR-105-progressive-disclosure-for-way-trees.md) | Progressive Disclosure for Way Trees | accepted | +| [ADR-107](./ways/ADR-107-way-match-corpus-batch-mode-and-locale-support.md) | Corpus, Matching Pipeline, and Locale Support | accepted (partially superseded by ADR-125) | +| [ADR-108](./ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) | Embedding-Based Way Matching with all-MiniLM-L6-v2 | accepted | +| [ADR-114](./ways/ADR-114-attend-as-insistent-way-trigger-type.md) | `attend` Events as an Insistent Way Trigger Type | accepted | +| [ADR-123](./ways/ADR-123-firing-dynamics-progression-axis-unification.md) | Firing dynamics — progression-axis unification for attend and ways | accepted | +| [ADR-125](./ways/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) | Authored Disclosure Graph and Removal of BM25 | accepted | +| [ADR-126](./ways/ADR-126-window-relative-refire.md) | Window-relative refire with named presets | accepted | +| [ADR-127](./ways/ADR-127-reject-full-body-embedding-corpus.md) | Full-body embedding corpus for way matching | rejected | +| [ADR-130](./ways/ADR-130-sentence-salience-input-reduction-for-embed-matching.md) | Sentence-salience input reduction for embed matching | accepted | +| [ADR-134](./ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) | Empirical auto-tuning from fire and near-miss telemetry | accepted | +| [ADR-135](./ways/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md) | Content-aware write-time over-build gate with a self-extending pattern corpus | accepted | +| [ADR-139](./ways/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md) | Shelve maintainer i18n; adopter-run localization via ways-localize | accepted | +| [ADR-153](./ways/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md) | Session-introspection substrate — correlating fired ways to turns | accepted | +| [ADR-154](./ways/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md) | `ways introspect` — one model, three front-ends | accepted | +| [ADR-155](./ways/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md) | Semantic gating of the keyword channel and reasoning-channel rebuild | accepted (partially superseded by ADR-188 §3) | +| [ADR-156](./ways/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md) | Calibrated relevance scoring for the semantic lane | accepted | +| [ADR-157](./ways/ADR-157-case-insensitive-trigger-regex-compilation.md) | Case-insensitive trigger regex compilation | accepted | +| [ADR-158](./ways/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md) | Calibration boundary quality — hard negatives and a fire-breadth ship gate | accepted | +| [ADR-159](./ways/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md) | Remove ways tune-curves and the legacy curve: cadence field | accepted | +| [ADR-160](./ways/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md) | Chunked late-interaction matching with softmax-share gating for way selection | accepted | +| [ADR-161](./ways/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md) | Queued mid-turn operator messages as an aggregated scan surface | accepted | +| [ADR-166](./ways/ADR-166-single-source-of-truth-for-model-context-window-resolution.md) | Single source of truth for model context-window resolution | accepted | +| [ADR-174](./ways/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md) | Progressive core — decoration guidance and the core re-disclosure gap | accepted | +| [ADR-183](./ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md) | Single-language localization tuning: the English anchor as a peer | accepted | +| [ADR-188](./ways/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md) | PostToolUse delivery for tool-lane ways and retirement of the semantic Bash surface | proposed | +| [ADR-189](./ways/ADR-189-session-introspection-implementation-plan.md) | Session introspection implementation plan | accepted | +| [ADR-190](./ways/ADR-190-the-lexical-gate-as-a-conditional-threshold.md) | The lexical gate as a conditional threshold | accepted | +| [ADR-191](./ways/ADR-191-the-tool-use-channel-is-a-signal-problem-lookbehind-chunk-spread-and-winner-confirmation.md) | The tool-use channel is a signal problem: lookbehind, chunk-spread, and winner confirmation | accepted | +| [ADR-192](./ways/ADR-192-late-interaction-matching-the-flow.md) | Late-interaction matching: the flow | accepted | ## Governance _Provenance, traceability, controls, compliance mapping_ | ADR | Title | Status | |-----|-------|--------| -| [ADR-200](./governance/ADR-200-compliance-claims-and-session-derived-findings.md) | Compliance claims and session-derived findings | Accepted | -| [ADR-201](./governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md) | Findings assembled as classifier-ready assessment records | Accepted | +| [ADR-005](./governance/ADR-005-governance-traceability.md) | Governance Traceability for Ways | superseded (superseded by ADR-200) | +| [ADR-151](./governance/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md) | Extract ways-core crate and ways-audit sibling binary | accepted | +| [ADR-200](./governance/ADR-200-compliance-claims-and-session-derived-findings.md) | Compliance claims and session-derived findings | accepted | +| [ADR-201](./governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md) | Findings assembled as classifier-ready assessment records | accepted | ## Documentation _Documentation structure, tooling, coherence_ | ADR | Title | Status | |-----|-------|--------| -| [ADR-300](./documentation/ADR-300-documentation-structure.md) | Documentation Structure | Accepted | -| [ADR-301](./documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md) | Situated socialization as canonical framing and documentation prose refactor | Accepted | -| [ADR-302](./documentation/ADR-302-unified-documentation-model.md) | A unified documentation model — typed graph, ways packaging, cross-repo convergence | Accepted | -| [ADR-303](./documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md) | Active-set semantics, adr archive, and supersession reading for the ADR corpus | Accepted | +| [ADR-106](./documentation/ADR-106-project-pulse-epoch-mapped-project-awareness.md) | Project Pulse — Epoch-Mapped Project Awareness | accepted | +| [ADR-177](./documentation/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md) | Version-stamped vendored tools with direction-aware drift detection | accepted | +| [ADR-300](./documentation/ADR-300-documentation-structure.md) | Documentation Structure | accepted | +| [ADR-301](./documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md) | Situated socialization as canonical framing and documentation prose refactor | accepted | +| [ADR-302](./documentation/ADR-302-unified-documentation-model.md) | A unified documentation model — typed graph, ways packaging, cross-repo convergence | accepted | +| [ADR-303](./documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md) | Active-set semantics, adr archive, and supersession reading for the ADR corpus | accepted | | [ADR-304](./documentation/ADR-304-typed-decision-records-the-adr-v1-contract.md) | Typed decision records: the adr/v1 contract | Accepted | -| [ADR-305](./documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md) | Capabilities active at adoption need no add decision | accepted | +| [ADR-305](./documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md) | Capabilities active at adoption need no add decision | superseded (superseded by ADR-311) | +| [ADR-306](./documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md) | adr import: foreign records through a round-trip import sheet | accepted | +| [ADR-307](./documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md) | A decision names what should be observable when it holds | accepted | +| [ADR-308](./documentation/ADR-308-a-change-decision-may-list-several-capabilities.md) | A change decision may list several capabilities | accepted | +| [ADR-309](./documentation/ADR-309-an-evidence-record-kind-for-findings-surveys-and-explorations.md) | An evidence record kind for findings, surveys and explorations | accepted | +| [ADR-310](./documentation/ADR-310-record-numbers-are-permanent-identity-and-records-live-by-intent.md) | Record numbers are permanent identity, and records live by intent | accepted | +| [ADR-311](./documentation/ADR-311-the-adr-tool-checks-shape-and-references-git-keeps-the-history.md) | The adr tool checks shape and references; git keeps the history | accepted | -## Legacy (Pre-Domain Numbering) +## Attend +_Session awareness: sensors, peers, messaging, keepwarm_ | ADR | Title | Status | |-----|-------|--------| -| [ADR-004](./legacy/ADR-004-way-macros.md) | Way Macros for Dynamic Context Injection | Accepted | -| [ADR-005](./legacy/ADR-005-governance-traceability.md) | Governance Traceability for Ways | Superseded (superseded by ADR-200) | -| [ADR-013](./legacy/ADR-013-ways-skills-governance-architecture.md) | Ways, Skills, and Governance Architecture | Accepted | -| [ADR-014](./legacy/ADR-014-tfidf-semantic-matcher.md) | TF-IDF/BM25 Binary for Semantic Way Matching | Accepted | +| [ADR-101](./attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) | Wormhole relay protocol for cross-instance agent communication | superseded (superseded by ADR-113) | +| [ADR-102](./attend/ADR-102-irc-based-local-agent-communication.md) | IRC-based local agent communication | superseded (superseded by ADR-113) | +| [ADR-113](./attend/ADR-113-attend-active-awareness-module.md) | `attend` — Active Awareness Module | accepted | +| [ADR-117](./attend/ADR-117-sensor-crate-extraction-and-feature-flags.md) | Sensor Crate Extraction and Feature Flags | accepted | +| [ADR-118](./attend/ADR-118-focus-groups-dynamic-agent-grouping.md) | Focus Groups — Dynamic Agent Grouping | accepted | +| [ADR-119](./attend/ADR-119-action-potential-engagement-model.md) | Action Potential Engagement Model | superseded (superseded by ADR-123) | +| [ADR-120](./attend/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md) | Interactive Chat TUI — Human in the Signal Loop | accepted | +| [ADR-121](./attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md) | Salience decay for signal presentation — turn-based exponential | superseded (superseded by ADR-123) | +| [ADR-122](./attend/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md) | Attend disclosure sensor — token-gated affordance reheat | accepted | +| [ADR-124](./attend/ADR-124-channel-bar-ordering-open-as-base.md) | TUI Legend Architecture — Base Channel, Liveness, and Ordering | accepted | +| [ADR-129](./attend/ADR-129-instance-suffix-and-heartbeat-liveness.md) | Instance suffix and heartbeat liveness for attend identity | accepted | +| [ADR-136](./attend/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md) | Split addressed messaging from the sensor-observation bus | accepted | +| [ADR-137](./attend/ADR-137-boundedness-bounded-work-per-cycle.md) | Boundedness — a unit of work must be bounded within its cycle | accepted | +| [ADR-170](./attend/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md) | Human focus-group membership via username identity and a shared attend-groups crate | accepted | +| [ADR-171](./attend/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md) | Stable session identity — the roster enumerates addressable coordinating units | accepted | +| [ADR-172](./attend/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md) | Turn-boundary inbound delivery via a CLI-owned drain checkpoint | accepted | +| [ADR-173](./attend/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md) | Chat-idiom convergence for the attend command surfaces | accepted | +| [ADR-182](./attend/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md) | Keepwarm: attend keeps the prompt cache warm with a wake floor | accepted | +| [ADR-187](./attend/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md) | Attend MCP server mode: outbound and queries as typed tools, inbound stays on Monitor and the Stop hook | proposed | +| [ADR-400](./attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md) | Attend messaging disclosure with token-gated reheat | accepted | +| [ADR-401](./attend/ADR-401-attend-envelope-fields-sender-kind-principal-and-addressee-on-every-signal.md) | Attend envelope fields: sender kind, principal, and addressee on every signal | accepted | + +## Platform +_Install, update, configuration, permissions, the CLI contract, testing_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-111](./platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) | Unified `ways` CLI — Single Binary Tool Consolidation | accepted | +| [ADR-115](./platform/ADR-115-declarative-config-with-project-scope-overlay.md) | Declarative Configuration with Project-Scope Overlay | accepted | +| [ADR-116](./platform/ADR-116-declarative-permission-requirements.md) | Declarative Permission Requirements | accepted | +| [ADR-131](./platform/ADR-131-project-scope-way-toggles.md) | Project-scope way toggles | accepted | +| [ADR-140](./platform/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md) | Two install topologies: in-place repo and subdirectory projection | accepted | +| [ADR-142](./platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md) | agent-ways 1.0 — XDG application distribution | accepted | +| [ADR-144](./platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) | Install / repair / migrate as one manifest reconciler | accepted | +| [ADR-146](./platform/ADR-146-installer-binary-verification-and-guided-build-fallback.md) | installer binary verification and guided build fallback | accepted | +| [ADR-147](./platform/ADR-147-composable-settings-json-config-fragments.md) | Composable settings.json — a store of YAML config fragments | superseded (superseded by ADR-169) | +| [ADR-148](./platform/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md) | framework surface ships operator content; dev harness in project scope | accepted | +| [ADR-149](./platform/ADR-149-operator-config-interview-skill.md) | operator config interview skill | superseded (superseded by ADR-169) | +| [ADR-150](./platform/ADR-150-version-truth-and-downgrade-safe-self-update.md) | Version-truth and downgrade-safe self-update | accepted | +| [ADR-152](./platform/ADR-152-framework-default-secret-path-deny-baseline.md) | Framework-default secret-path deny baseline | accepted | +| [ADR-162](./platform/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md) | Mechanical session-link suppression as defense against transcript disclosure | superseded (superseded by ADR-167) | +| [ADR-163](./platform/ADR-163-config-separation-dotfiles-source-of-truth.md) | Config separation — dotfiles as source-of-truth feeding the settings fragment store | accepted | +| [ADR-164](./platform/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md) | File artifacts distributed across hosts must be carried by value not host-absolute reference | accepted | +| [ADR-167](./platform/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md) | Session-link suppression: attribution.sessionUrl as primary control, deny hook as backstop | accepted | +| [ADR-169](./platform/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md) | agent-ways relinquishes user-scoped settings.json; retains only its operational baseline | accepted | +| [ADR-179](./platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md) | Remove the pre-1.0 in-place migrator; keep the guards and the transition fallbacks | Accepted | +| [ADR-181](./platform/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md) | Guard hooks: a blocking PreToolUse class, shipped deactivated | accepted | +| [ADR-184](./platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md) | Installation and activation are separate states: targets as the unit of activation | accepted | +| [ADR-185](./platform/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md) | CLI output contract: structured output for people, JSON for machines | accepted | +| [ADR-186](./platform/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md) | Live integration fixture: install-path test levels and the tier 2 gate | Accepted | +| [ADR-500](./platform/ADR-500-settings-json-three-way-merge-spec-and-peer-writer-coexistence-contract.md) | settings.json three-way merge: spec and peer-writer coexistence contract | accepted | + +## Practice +_The ways method, way authoring, the development loop_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-013](./practice/ADR-013-ways-skills-governance-architecture.md) | Ways, Skills, and Governance Architecture | accepted | +| [ADR-100](./practice/ADR-100-ways-scaffolding-wizard.md) | Ways Scaffolding Wizard | accepted | +| [ADR-109](./practice/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md) | Project-Scope Way Embedding with Manifest-Based Staleness Detection | accepted | +| [ADR-110](./practice/ADR-110-way-file-separation-and-graph-compatible-structure.md) | Way File Separation and Graph-Compatible Structure | accepted | +| [ADR-128](./practice/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md) | Memory as repo-portable ways — seed routing over accumulated snapshots | accepted | +| [ADR-132](./practice/ADR-132-collaboration-ways-domain.md) | Collaboration ways domain | accepted | +| [ADR-138](./practice/ADR-138-skills-own-the-how-ways-own-the-5w.md) | Skills own the how, ways own the 5W | accepted | +| [ADR-141](./practice/ADR-141-knowledge-graph-as-evidential-memory-backend.md) | Knowledge Graph as Evidential Memory Backend | accepted | +| [ADR-143](./practice/ADR-143-three-root-way-runtime-core-user-project.md) | Three-root way runtime — core, user, project | accepted | +| [ADR-165](./practice/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md) | Loop-control bookends: start, develop, merge, release, wrap | accepted | +| [ADR-175](./practice/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md) | Standing delegation authorization satisfies the harness permission gate rather than overriding it | accepted | +| [ADR-176](./practice/ADR-176-contract-identification-as-the-develop-loop-front-gate.md) | Contract identification as the develop-loop front gate | accepted | +| [ADR-178](./practice/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md) | Register transfers by demonstration - core.md carries policy | accepted | +| [ADR-180](./practice/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md) | GitHub issues as the shared truth for the session task list | accepted | +| [ADR-600](./practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) | Cognitive loop and the awareness layer | accepted | +| [ADR-601](./practice/ADR-601-autonomy-as-a-layered-system-goals-signposts-and-the-initiation-pattern.md) | Autonomy as a layered system: goals, signposts, and the initiation pattern | accepted | +| [ADR-602](./practice/ADR-602-ways-functional-audit-117-ways-against-the-firing-contract.md) | Ways functional audit: 117 ways against the firing contract | accepted | +| [ADR-603](./practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md) | Cypress survey: what a node-routed seed teaches a hook-disclosed corpus | accepted | ## Archived diff --git a/docs/architecture/adr.yaml b/docs/architecture/adr.yaml index b8dfce93..0449af6c 100644 --- a/docs/architecture/adr.yaml +++ b/docs/architecture/adr.yaml @@ -7,11 +7,11 @@ project_name: Claude Code Configuration # Domain Number Series domains: - system: + ways: range: [100, 199] - name: System - description: Ways architecture, matching, macros, hooks, session lifecycle - folder: system + name: Ways + description: "The ways engine: how ways are matched, disclosed and re-disclosed" + folder: ways governance: range: [200, 299] @@ -25,6 +25,24 @@ domains: description: Documentation structure, tooling, coherence folder: documentation + attend: + range: [400, 499] + name: Attend + description: "Session awareness: sensors, peers, messaging, keepwarm" + folder: attend + + platform: + range: [500, 599] + name: Platform + description: Install, update, configuration, permissions, the CLI contract, testing + folder: platform + + practice: + range: [600, 699] + name: Practice + description: The ways method, way authoring, the development loop + folder: practice + # The record contract (ADR-304). Records that declare `contract: adr/v1` # follow the kinds below; a record without it is adr/v0 and keeps the v0 # statuses and checks until someone migrates it. @@ -32,16 +50,18 @@ contract: adr/v1 kinds: decision: - mutable_after_accept: [status, enacted, superseded_by, considered, concern] verb: required requires: [capability, basis, agent] sections: [Summary] - edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec, evidence] } spec: - mutable_after_accept: all verb: forbidden requires: [capability] edges: { supersedes: spec, decided_by: decision } + evidence: + verb: forbidden + requires: [capability] + edges: { supersedes: evidence } # What agent-ways does, one line each: the only hand-written description of # a capability. A decision adds, cuts, changes or constrains these. @@ -60,24 +80,14 @@ capabilities: loop: The development loop skills (start, develop, merge, release, wrap) testing: The test suites and the live install fixture -# Capabilities that were active when agent-ways adopted adr/v1 on the -# adopted date. No record added them, so none needs an add decision. A change -# on one with no prior record to name stands on the baseline. Any capability -# declared after adoption needs an add (ADR-305). -baseline: - adopted: 2026-09-27 - capabilities: [docs, method, matching, disclosure, authoring, cli, attend, install, config, governance, loop, testing] - -# Retire targets name a surface; these are agent-ways' namespaces. -surfaces: - cli: {} - skill: {} - way: {} - hook: {} - # The basis sources a decision may name (ADR-304 §11); the v1 default. basis_sources: [operator, evidence, standard, upstream, precedent] +# This repository, as URLs name it. A URL into it at a branch is a path in +# the tree, and adr domain rewrites it when a record moves. Without this key +# the origin remote is used. +repository: github.com/aaronsb/agent-ways + # adr cite skips these: fixtures, tool sources and subagent prompts carry # example numbers. cite: diff --git a/docs/architecture/attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md b/docs/architecture/attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md new file mode 100644 index 00000000..9b514f15 --- /dev/null +++ b/docs/architecture/attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md @@ -0,0 +1,179 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: attend +superseded_by: + - ADR-113 +basis: + - standard: 'magic-wormhole: NAT-transparent, credential-free file transfer with single-use codes' +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-02-21 +deciders: + - aaronsb + - claude +related: [] +imported: + from: docs/architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md + format: v0 + status: Deprecated +--- + +# ADR-101: Wormhole relay protocol for cross-instance agent communication + +## Status: Deprecated + +**Deprecated 2026-02-26** after experimental validation. The manifest protocol works mechanically but wormhole's one-shot, role-asymmetric design makes it fundamentally unsuited for conversations. See [Experimental Results](#experimental-results-2026-02-26) below. For agent-to-agent chat, use a persistent bidirectional transport (IRC, Matrix, or similar). Wormhole remains the right tool for one-shot file transfers via the `/wormhole` skill. + +## Context + +Claude Code teams currently operate within a single machine — a lead agent spawns teammates, they share a task list, and coordinate via SendMessage. There is no mechanism for two independent Claude instances on different machines (different users, different accounts, different networks) to exchange files or data. + +magic-wormhole provides secure, NAT-transparent, credential-free file transfer using short codes. We already have a `/wormhole` skill for interactive and automated transfers. The natural next step is a protocol layer that allows two Claude teams to sustain a multi-turn conversation over wormhole, without requiring a human to relay codes after the initial handshake. + +The core challenge: wormhole codes are single-use. Once a transfer completes, the code is consumed. A stateless transport needs a convention to become stateful. + +## Decision + +Define a **manifest-based relay protocol** where each wormhole transfer includes a manifest file alongside the payload. The manifest carries the state needed to continue the conversation. + +### Manifest Structure + +```json +{ + "protocol": "claude-wormhole-relay", + "version": "0.1", + "turn": 1, + "direction": "a-to-b", + "codes": ["99-beta-canyon", "99-gamma-delta", "99-echo-falcon"], + "continuation": "99-zulu-renew", + "payload": ["report.md"], + "message": "Here's the analysis you requested." +} +``` + +| Field | Purpose | +|-------|---------| +| `protocol` | Identifies this as a relay manifest | +| `version` | Protocol version for forward compatibility | +| `turn` | Monotonic turn counter | +| `direction` | Which side is sending this turn | +| `codes` | Pre-agreed codes for upcoming turns (consumed in order) | +| `continuation` | Special code that signals "this turn's payload is a fresh batch of codes" | +| `payload` | List of filenames included in this transfer | +| `message` | Free-text message between agents (context, instructions, requests) | + +### Turn Lifecycle + +1. **Initiation**: Human A tells Claude A to send a file to another machine. Claude A generates a manifest with N pre-generated codes and sends it via the first code (which the human relays to side B). +2. **Steady state**: Each side consumes the next code from the manifest to send their turn. Both sides know the full code sequence, so no human relay is needed. +3. **Renewal**: When the `continuation` code is reached, that turn's payload is a new manifest with fresh codes — extending the conversation. +4. **Termination**: A manifest with empty `codes` and no `continuation` signals the conversation is complete. + +### Code Generation + +Codes follow the pattern `<session-number>-<word>-<word>` (or longer). Claude generates the word pairs. Both sides receive the same code list via the manifest, so no shared seed or deterministic generation is required — the manifest IS the shared state. + +### Team Integration + +A dedicated **comms agent** teammate handles the relay: +- Watches for incoming turns (runs `wormhole receive` with the next expected code) +- Delivers received payloads and messages to the team lead +- Sends outgoing payloads when the lead requests a transfer +- Manages manifest state (current turn, remaining codes, renewal) + +### Error Recovery + +- **Failed transfer**: Retry the same code. Wormhole codes remain valid until successfully used. +- **Missed turn**: The receiving side can keep listening on the expected code indefinitely (with a configurable timeout). +- **Code exhaustion without renewal**: The comms agent notifies the lead that the channel is closing. A human must broker a new initial code to restart. + +## Consequences + +### Positive + +- Two Claude instances on separate machines can exchange files and messages without shared credentials, SSH keys, or network access +- NAT/firewall transparent — both sides only need outbound internet +- No persistent infrastructure — no server to maintain, no accounts to manage +- Each side retains full autonomy — separate teams, separate humans, separate contexts +- The manifest pattern is extensible (add fields for encryption, compression, routing) + +### Negative + +- Latency per turn — each exchange requires a full wormhole handshake through the relay server +- Not suitable for high-frequency or real-time communication +- Single point of failure in the wormhole relay server (transit.magic-wormhole.io) +- Code list is transmitted in plaintext within the manifest — if intercepted, future turns are compromised +- Requires magic-wormhole installed on both machines + +### Neutral + +- The protocol is transport-agnostic in principle — the manifest pattern could work over other transports (email, shared storage, etc.) but this ADR focuses on wormhole +- Human supervision remains at the edges — each side's human can inspect manifests and payloads +- This does not define authentication between agents — both sides trust whoever holds the code + +## Alternatives Considered + +- **SSH/SCP between machines**: Requires pre-shared credentials, network access, and firewall configuration. Higher throughput but much higher setup cost. Better for recurring transfers between known hosts. +- **Cloud storage (S3, GCS) as intermediary**: Requires accounts on a shared service, credentials management, and cleanup. More infrastructure than the problem warrants. +- **Email-based exchange**: Async but requires email configuration, attachment size limits, and parsing complexity. +- **Custom relay server**: Maximum control but requires hosting, maintenance, and security hardening. Overkill for ad-hoc agent communication. +- **Pre-shared deterministic code generation (shared seed)**: Both sides generate codes from a seed without manifests. Simpler but fragile — if turns desync, the protocol breaks with no recovery path. The manifest approach is more resilient because state is explicit. + +## Related Patterns: Ralph Loops + +A ralph loop is a self-referential agent feedback cycle where an agent's output becomes its own input across iterations. Combining a ralph loop with this relay protocol creates an interesting (and potentially hazardous) topology: two independent Claude instances, each running their own ralph loop, exchanging intermediate state via wormhole manifests. + +This could enable: +- **Collaborative iteration**: Agent A refines a document, sends it to Agent B for a different perspective, receives it back, continues refining — each side's ralph loop incorporating the other's feedback +- **Distributed analysis**: Each side processes a portion of a problem, exchanges findings, and iterates toward convergence + +The risks are obvious: +- **Runaway amplification**: Two feedback loops feeding each other can diverge rapidly — each side responding to the other's responses with increasing elaboration and no natural stopping point +- **Resource exhaustion**: Without explicit turn limits or convergence criteria, the loop runs until context windows fill, API budgets drain, or a human intervenes +- **Semantic drift**: Successive rounds of interpretation and reinterpretation can drift far from the original intent + +Any implementation combining ralph loops with the relay protocol MUST include hard bounds: maximum turn count, convergence detection (output similarity between rounds), and mandatory human checkpoints. The continuation mechanism in the manifest provides a natural rate limiter — code exhaustion forces a pause unless explicitly renewed. + +## Experimental Results (2026-02-26) + +Two Claude Code instances tested the manifest protocol over 10 pre-agreed codes with a UUID-tagged naming scheme (`N-439351ef-word-word`). + +### What Worked + +- **Initial handshake**: Human-relayed random codes work reliably for bootstrapping +- **Manifest delivery**: JSON manifest transferred cleanly, both sides parsed it +- **Sequenced turns**: When the human explicitly said "side B receive, now side A send," transfers succeeded every time +- **6 of 10 turns completed** successfully, including 2 out-of-band resyncs + +### What Failed + +- **Role collisions**: 3 of 10 turns failed with `ServerError: crowded` — both sides attempted the same role (both receiving or both sending) on the same code simultaneously +- **Burned codes**: Collisions permanently consume the code. Unlike TCP, there is no retry — the channel number is gone +- **Human sequencing required**: Despite pre-agreed sender/receiver assignments per turn, the human had to manually sequence each exchange. The "autonomous after initial handshake" goal was not achieved +- **No automatic recovery**: When codes burned, the only recovery path was out-of-band fallback codes (side B generated random codes and the human relayed them) + +### Root Cause + +Wormhole's PAKE handshake is **role-asymmetric**: one side must be the sender and the other the receiver. If both connect as the same role, the relay server returns `crowded` and the code is consumed. This is by design — wormhole assumes a human is coordinating both ends in real-time for a single transfer. + +The manifest protocol tried to pre-assign roles, but both Claude instances execute asynchronously. Without a synchronization mechanism to guarantee receiver-before-sender ordering, collisions are inevitable. The protocol is essentially **UDP with no retransmit and destructive packet loss**. + +### Conclusion + +Wormhole is the wrong transport for conversations. It was designed for one-shot file drops, and it excels at that. Forcing statefulness onto a stateless, single-use protocol creates fragility that no amount of manifest engineering can fix. + +For cross-instance agent chat, use a transport designed for persistent bidirectional messaging: + +| Transport | E2E Encrypted | Bidirectional | Persistent | No Accounts | Setup Cost | +|-----------|:---:|:---:|:---:|:---:|---| +| **IRC** (self-hosted ircd) | TLS | Yes | Yes | Yes | Minimal — one process | +| **Matrix** (via matrix-commander) | Yes (Olm) | Yes | Yes | No | Moderate | +| **Tailscale + socat** | Yes (WireGuard) | Yes | Yes | No | Low if already using Tailscale | + +IRC with a local ircd is the simplest option: zero accounts, zero credentials, one package install, both sides join a channel and talk. The channel handles ordering, buffering, and fan-out — all the problems the manifest protocol tried to solve. + +**Wormhole's role going forward**: one-shot file transfers between machines via the `/wormhole` skill. No conversation protocol. diff --git a/docs/architecture/attend/ADR-102-irc-based-local-agent-communication.md b/docs/architecture/attend/ADR-102-irc-based-local-agent-communication.md new file mode 100644 index 00000000..8ed3661b --- /dev/null +++ b/docs/architecture/attend/ADR-102-irc-based-local-agent-communication.md @@ -0,0 +1,164 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: attend +superseded_by: + - ADR-113 +basis: + - evidence: 'ADR-101''s wormhole experiment (2026-02-26): 6 of 10 turns succeeded, 3 collisions burned codes and needed out-of-band resync' + - precedent: ADR-101 +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-02-26 +deciders: + - aaronsb + - claude +related: + - ADR-101 +imported: + from: docs/architecture/system/ADR-102-irc-based-local-agent-communication.md + format: v0 + status: Deprecated +--- + +# ADR-102: IRC-based local agent communication + +## Context + +ADR-101 proposed a manifest-based relay protocol over magic-wormhole for cross-instance agent communication. Experimental testing (2026-02-26) revealed fundamental limitations: wormhole codes are single-use and role-asymmetric (one sender, one receiver). When both sides race to the relay server, codes are permanently consumed by collisions. The protocol achieved 6/10 successful turns with 3 collisions requiring out-of-band resync — essentially UDP with destructive packet loss. + +The core need remains: two Claude Code instances on the same machine need a way to exchange messages in real time. The transport must be persistent, bidirectional, ordered, and simple enough that Claude can operate it with basic file and shell tools. + +## Decision + +Use IRC for agent-to-agent communication, with `miniircd` (a single-file Python 3 IRC server) for the server and `ii` (suckless filesystem-based IRC client) for the client. The human relays connection details (host, port, channel) between instances. All communication flows over IRC. + +### Architecture + +``` +Claude-5d00b2 → ii (FIFO/file) → miniircd (localhost:PORT) ← ii (FIFO/file) ← Claude-144dc4 + (~/.claude) #relay (~/temp) +``` + +### Identity: Hash-Based Nicks + +Each instance derives a deterministic nick from its working directory using a sha256 hash: + +``` +/home/aaron/.claude → claude-5d00b2 +/home/aaron/temp → claude-144dc4 +``` + +This replaces hardcoded `claude-a`/`claude-b` naming. The hash is stable across restarts, unique per directory, and fits IRC nick limits (16 chars). The ii base directory uses the same slug: `/tmp/irc-chat-<hash>/`. + +Implementation: `~/.claude/hooks/irc-lib.sh` provides `irc_nick_from_dir()` and `irc_slug_from_dir()`. + +### Why This Works + +- **ii is filesystem-based**: messages are regular files (`out`) and FIFOs (`in`). Claude reads and writes them with standard tools — no IRC library needed. +- **miniircd is zero-config**: a single Python 3 file, starts with one command, no config files, no accounts. +- **IRC handles the hard parts**: message ordering, buffering, fan-out, presence detection — all the problems the wormhole manifest protocol tried to solve manually. + +### Bootstrapping + +1. Host starts miniircd on a random high port, connects with ii (hash-derived nick), joins `#relay` +2. Human relays host, port, and channel to the other instance (3 values, no wormhole needed) +3. Joiner connects with ii (its own hash-derived nick), joins `#relay` +4. Both sides chat via wrapper scripts (`irc-send.sh`, `irc-read.sh`) + +Wormhole is not used for local bootstrapping — the human relay is simpler and more reliable for 3 values on the same machine. Wormhole remains available for remote bootstrapping (delivering connection details to a real IRC server across machines). + +### Ambient Monitoring + +A `UserPromptSubmit` hook (`~/.claude/hooks/irc-check.sh`) provides tick-based message delivery: + +- **High-water mark**: tracks last-read line in `.hwm` file, only surfaces new messages +- **Notification tiers**: direct mentions show the message, ≤3 new messages inline, >3 as a badge count +- **Self-filtering**: own messages excluded to avoid echo +- **Join/part detection**: connection events surfaced as one-liners + +Time is tick-based: each Claude Code API round-trip is an epoch. The hook fires on each user prompt, not on wall-clock time. + +### Summary Nudge + +A tick counter in `irc-check.sh` nudges the agent to post a recap every ~15 ticks using a triple-nudge pattern: ticks 15, 16, 17 each deliver an escalating prompt, then the counter resets for another 15-tick quiet period. Each nudge steers toward brevity — a sentence or two, not a dissertation. Only fires when other users are present in the channel — inactive IRC never triggers nudges. + +### Compaction-Safe Session State + +Context compaction loses IRC message history. To bridge this gap: + +- `irc-check.sh` writes `.irc-has-history` when messages from others are received +- `irc-session.sh` (SessionStart:compact/resume hook) checks this flag and emits a breadcrumb: "N messages received before compaction — run `irc-read.sh` to catch up" +- `irc-read.sh` clears the flag on read + +### Hook Wiring + +The following hooks are wired in `settings.json`: + +- **UserPromptSubmit + Stop**: `irc-check.sh` — tick-based message delivery on every turn +- **SessionStart (compact/resume)**: `irc-session.sh` — compaction breadcrumb +- **Bash allow**: `irc-send.sh *` and `irc-read.sh` — permission-free execution + +### Wrapper Scripts + +All IRC I/O goes through allowlisted wrapper scripts to avoid permission prompts and path issues (`#` in channel names triggers Claude Code's shell parser): + +- `~/.claude/hooks/irc-send.sh <message>` — send to `#relay` +- `~/.claude/hooks/irc-read.sh [N]` — read last N messages + +Scripts auto-detect the active `/tmp/irc-chat-*` connection via `irc-lib.sh`. + +### Scope + +Localhost only for now. Both Claude instances must be on the same machine. Cross-network communication is a future extension (Tailscale, public IRC, Matrix). + +## Consequences + +### Positive + +- Persistent bidirectional channel — no code burning, no timing races +- Claude operates IRC through filesystem I/O — no special libraries or raw socket handling +- Zero infrastructure beyond a single Python process +- Trivially extensible — more agents join the same channel, add more channels for topics +- Ambient monitoring via hooks — messages appear in context automatically without manual polling +- Hash-based identity — deterministic, stable across restarts, no coordination needed + +### Negative + +- Requires `ii` installed (`pacman -S ii` on Arch, build from source elsewhere) +- Bundles miniircd as a vendored file in the skill directory +- Localhost-only limits cross-machine use cases +- No encryption (acceptable for local-only; would need TLS for network use) + +### Neutral + +- miniircd is GPL-2.0 licensed — vendoring the single file is fine for personal tooling +- The `/wormhole` skill remains available for one-shot file transfers and potential remote IRC bootstrapping +- IRC skills are split into three commands: `/irc-host` (start server), `/irc-chat` (join existing), `/irc-teardown` (cleanup) +- ADR-101's manifest protocol is deprecated but preserved as documentation of what was tried and why it failed +- Wrapper scripts and permission allowlisting are implementation details that may evolve as Claude Code's permission model changes + +## Future: Channel Topology at Scale + +With 2 agents, a single `#relay` channel is sufficient. At 3+ agents, noise becomes a design question. + +**Options considered:** + +- **Single channel (current)**: Everyone sees everything. Noise is bounded by tick rate, not message rate — each agent only processes messages on their turn. May work longer than expected. +- **Topic channels**: `#frontend`, `#backend`, `#deployment`. Focused but reintroduces the human-as-router problem IRC was built to eliminate — someone has to decide who goes where. +- **Hub and spoke**: `#relay` stays as the commons (everyone subscribes, always). Topic channels are additive, not replacements. Cross-cutting context flows through the hub; focused discussion flows through spokes. Nobody leaves the common channel, so no agent gets excluded. + +Hub and spoke is the likely answer when the time comes. The third-wheel problem — an agent stranded on a dead topic channel while the real conversation happens elsewhere — is solved by the invariant that `#relay` is never abandoned. miniircd already supports multiple channels; no infrastructure change needed. + +**Not building this yet.** The tick-based temporal model and summary nudge already compress signal. The real noise threshold is probably higher than it feels because agents read batched messages per-tick, not a real-time stream. Build when `#relay` actually gets noisy, not before. + +## Alternatives Considered + +- **Wormhole manifest protocol (ADR-101)**: Tested and deprecated. One-shot codes with role-asymmetric handshakes are structurally unsuited for conversations. See ADR-101 experimental results. +- **Matrix (via matrix-commander)**: E2E encrypted, persistent, bidirectional. Requires accounts and a homeserver — too much infrastructure for localhost same-machine chat. +- **Named pipes (FIFOs) directly**: Simplest possible approach but no message ordering, no buffering, no presence detection. Would need to reimplement what IRC provides for free. +- **Tailscale + socat**: Great for cross-network but requires Tailscale auth on both machines. Overkill for localhost. +- **UnrealIRCd / ngircd**: Production IRC servers with mandatory configuration. miniircd's zero-config single-file design is better for ephemeral use. diff --git a/docs/architecture/attend/ADR-113-attend-active-awareness-module.md b/docs/architecture/attend/ADR-113-attend-active-awareness-module.md new file mode 100644 index 00000000..6ee7cc22 --- /dev/null +++ b/docs/architecture/attend/ADR-113-attend-active-awareness-module.md @@ -0,0 +1,374 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: attend +supersedes: + - ADR-101 + - ADR-102 +basis: + - standard: Claude Code's Monitor tool delivers each stdout line of a background process as an asynchronous notification + - evidence: 'notification rate experiment: 10+ notifications in 2 minutes caused confabulated user turns; 3 per 2 minutes was stable' + - evidence: ADR-101 and ADR-102 failed on external transports (wormhole fragility, IRC complexity) +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-09 +deciders: + - aaronsb + - claude +related: + - ADR-104 + - ADR-111 + - ADR-112 + - ADR-114 +imported: + from: docs/architecture/system/ADR-113-attend-active-awareness-module.md + format: v0 + status: Accepted + unmapped: + revised: 2026-04-10 +--- + +# ADR-113: `attend` — Active Awareness Module + +## Context + +The [Cognitive Loop and the Awareness Layer](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) design note reads the agent-ways system as a cognitive loop with one missing stage: **active perception**. Every other stage is in place — reactive guidance (ADR-100, ADR-103, ADR-105, ADR-108), attention allocation via disclosure gating (ADR-104), episodic memory (ADR-112 ledger), associative recall (ADR-112 KG), consolidation (compaction plus compaction-checkpoint way), project-level awareness at session entry (ADR-106), and the scoring infrastructure that binds these together. What has been missing is a mechanism by which Claude gains cheap peripheral awareness of its own approaching consequences and the environment around it, without spending reasoning tokens to compute either. + +The consequence of that gap is well-defined: Claude operates in a partially blind configuration. Self-monitoring is expensive (costs tokens) and unreliable (can fail to fire). Environmental sensing requires explicit tool use (which consumes an entire turn for a single observation). Consequence tracking happens only when Claude chooses to check, by which point the consequence may be too near to respond to usefully. + +This ADR defines the missing layer as a new binary, separate from `ways`, co-resident in the agent-ways workspace. + +### The delivery primitive: `Monitor` + +An additional and load-bearing reason this ADR is possible *now* rather than earlier is the `Monitor` tool in Claude Code. Before `Monitor`, there was no standardized channel by which a background process could surface observations into Claude's conversation as asynchronous events. The hook system delivers synchronous injections at event boundaries, but it has no mechanism for signals that occur *between* those boundaries. + +`Monitor` closes that gap by treating a background script's stdout as an async event stream, delivering each line as a notification in Claude's chat. With `persistent: true`, a single `Monitor` invocation covers the duration of a Claude Code session. This is the delivery primitive the awareness layer depends on. + +### Why a new binary (and not a ways subcommand) + +The `ways` binary is a matcher, scorer, and disclosure gate — a stateless computation engine invoked by hooks. It responds to events; it does not observe between them. Adding a durable, long-lived executive layer to `ways` would conflate two different lifecycles. + +A sibling crate in the same workspace preserves the integration surface while keeping lifetimes clean. `attend` runs as a separate process and writes single-line observations to stdout. `Monitor` delivers those lines to Claude as asynchronous notifications. + +### Prior attempts and how this differs + +ADR-101 (Wormhole relay, Deprecated) and ADR-102 (IRC-based agent communication, Abandoned) both attempted inter-instance awareness through external transports. Both failed: ADR-101 on transport fragility, ADR-102 on complexity. + +This ADR succeeds where those failed because it uses the **filesystem as transport** — `~/.claude/sessions/*.json` for session discovery, `~/.cache/attend/signals/` for inter-session messaging. No external services, no fragile protocols. The filesystem is stable because Claude Code depends on it (session files) and third-party tools depend on it (abtop reads the same session files). The transport is load-bearing for the ecosystem, not just for attend. + +## Decision + +Introduce `attend`, a new Rust binary hosted as a sibling crate to `ways` in the agent-ways Cargo workspace. `attend` implements the active awareness layer: durable, restartable, session-scoped, additive to (not required by) Claude Code. + +### Identity + +- **Name:** `attend` +- **Binary:** `attend` +- **Crate:** `attend/` (sibling to `ways/` in the workspace) +- **Role:** Active awareness module — sensor loop, peer awareness, inter-session signaling +- **Lifecycle:** Per work session. Explicit start via `/attend` skill or `attend run`. Exits cleanly on signal. + +### CLI + +`attend` is a normal CLI. Bare `attend` prints help with an ANSI header matching the ways visual style. + +``` +attend run # sensor loop (launched via Monitor) +attend run --catchup # process existing signals, then watch forward +attend send "message" # signal to peers (project + focus scope) +attend send --broadcast "message" # signal to all sessions +attend send --to /path "message" # directed signal to specific project +attend inbox # read pending messages chronologically +attend peers # list active Claude Code sessions +attend focus add ~/path # add project to focus group +attend focus remove ~/path # remove from focus group +attend focus list # show current focus group +attend focus clear # back to project-only mode +attend status # running instances, signals, focus state +attend --version # version + git commit hash +``` + +### Internal architecture + +`attend` is composed of small modules with clear responsibilities. The coordination center is the **tick loop** — a wall-clock-driven game loop that runs sensors on adaptive schedules. + +``` +attend/ +├── build.rs # Bakes git commit hash at compile time +├── src/ +│ ├── main.rs # CLI dispatch, subcommands, disclosure governor +│ ├── tick/mod.rs # AdaptiveInterval: ramp-up on change, hysteresis decay +│ ├── delta/mod.rs # DeltaAccumulator: per-sensor state-change tracking +│ ├── emit/mod.rs # stdout for Monitor delivery, stderr for diagnostics +│ └── sensors/ +│ ├── mod.rs # Sensor trait, Focus struct, SensorSlot runtime wrapper +│ ├── context.rs # Interoceptive: context window pressure via `ways context` +│ ├── git.rs # Git state: dirty files, branch changes, upstream divergence +│ ├── peer.rs # Peer sessions + signal file reading +│ └── process.rs # Application presence tracking +``` + +### Sensor model + +All sensors implement a common trait: + +```rust +pub trait Sensor { + fn name(&self) -> &str; + fn poll(&mut self, focus: &Focus) -> Vec<(f64, String)>; + fn emission_threshold(&self) -> f64; + fn base_interval(&self) -> Duration; + fn min_interval(&self) -> Duration; + fn decay_threshold(&self) -> u32; +} +``` + +Each sensor is wrapped in a `SensorSlot` that provides adaptive interval scheduling and delta accumulation. The tick loop polls sensors via a priority queue ordered by next-fire time. + +The `Focus` struct describes what Claude is currently attending to (description, working directory, keywords). Sensors filter observations through Focus to determine relevance. + +#### Built-in sensors + +| Sensor | Type | Polls | Reports | +|--------|------|-------|---------| +| **context** | interoceptive | `ways context --json` | threshold crossings, velocity, projection to critical | +| **git** | exteroceptive | `git status`, `git rev-list` | dirty files, branch changes, upstream divergence | +| **peers** | exteroceptive | `~/.claude/sessions/*.json` + signal files | session appear/exit, state changes, peer messages | +| **processes** | exteroceptive | `ps` | application presence (not PID churn) | + +The context sensor is calibrated to complement — not duplicate — ways' existing context-threshold triggers (todos@75%, memory@80%, checkpoint@95%). Attend provides early warnings *before* those thresholds and velocity/projection *between* them. + +#### Script sensors (planned) + +Sensors as unit files — declare what to run, how often, what threshold matters. Attend is the scheduler; sensors are units. Two-layer config mirrors ways scoping: + +``` +~/.config/attend/config.yaml # user scope — always loaded +{project}/.claude/attend.yaml # project scope — layered on top +``` + +Project config uses +/- to modify the sensor set: + +```yaml +# project/.claude/attend.yaml +sensors: + +hardware: + script: .claude/sensors/check-hardware.sh + interval: 120 + threshold: 2.0 + -processes: # not relevant here +``` + +Script sensor contract: output `magnitude|description` lines to stdout. Empty output = no change. Same adaptive interval and delta accumulation as built-in sensors. + +Trust model follows ways: user-scope config is trusted, project-scope scripts get the same scrutiny as project-scope way macros. + +### Tick loop + +The core loop is wall-clock-driven, not turn-driven. It runs continuously for the session's lifetime. + +Each tick: + +1. Check the priority queue for sensors whose next-fire time has arrived +2. Run due sensors, collect observations +3. For each observation, feed through the sensor's delta accumulator +4. If state changed: shorten polling interval (ramp up), reset decay cooldown +5. If no change: increment decay cooldown; if threshold exceeded, lengthen interval back toward base +6. Check whether any accumulator has crossed its emission threshold +7. Feed threshold-crossing observations through the disclosure governor +8. Emit to stdout (Monitor delivery) + +Quiet polls produce no stderr output. Only actual state changes are logged. + +### Adaptive sensor scheduling + +Each sensor maintains adaptive scheduling state: + +- **Ramp-up is fast.** When change is detected, interval halves (down to `min_interval`). Missing real change is worse than over-sampling. +- **Decay is slow (hysteresis).** The sensor must see `decay_threshold` consecutive quiet ticks before its interval lengthens. Prevents oscillation. +- **False positives self-correct.** High-frequency samples confirm or deny, then the sensor decays back to baseline. + +### Disclosure governor + +The disclosure governor is a global rate limiter controlling when attend emits to stdout. Architecturally load-bearing — without it, sustained notification delivery destabilizes Claude's turn-taking model. Experimentally validated: 10+ notifications in 2 minutes caused confabulated user turns; 3 per 2 minutes was completely stable. + +Two constraints: + +1. **Adaptive cooldown.** Higher aggregate event rate → longer wait between disclosures. Scales with square root of event rate. +2. **Hard cap per window.** Maximum 3 disclosures per 120-second window. + +When multiple sensors are ready simultaneously, they're emitted as a **batch** within Monitor's 200ms batching window, consuming one disclosure slot. + +Prototype ratio: 76 internal sensor ticks → 3 notifications over 2 minutes. + +### Peer awareness and inter-session signaling + +Attend discovers peer Claude Code sessions by reading `~/.claude/sessions/*.json` and their transcript JSONL files — the same stable pattern used by `abtop`. This provides: + +- **Session discovery:** who else is running, which project, what model, context usage +- **State tracking:** session appear/exit, status changes (working/waiting), context pressure + +#### Signal files + +Inter-session messaging uses signal files in a project-scoped directory structure: + +``` +~/.cache/attend/signals/ +├── _broadcast/ # all attend instances read this +├── -home-aaron-.claude/ # only attend in ~/.claude reads this +├── -home-aaron-temp/ # only attend in ~/temp reads this +└── focus # list of peer project dirs to watch +``` + +Signal format: `from|project|cwd|message` (one line, atomic write via tmp+rename). + +Sender identity is automatically detected: +- **From a Claude session:** `claude:session-id` → displays as `claude/project` +- **From a human terminal:** `external:user@terminal` → displays as `aaron@kitty` + +Terminal detection checks KITTY_PID, ALACRITTY_SOCKET, WEZTERM_PANE, TMUX, STY, TERM_PROGRAM, SSH_CONNECTION. + +#### Sidetone prevention + +Attend filters own signals by checking the `from` field against its own session ID, not by filename prefix. This works across all scoped directories. + +#### Forward-only mode + +`attend run` marks all existing signals as seen on startup. Only signals arriving *after* launch produce notifications. `attend inbox` reads all pending messages (one-shot catchup). `attend run --catchup` processes existing signals then watches forward. + +#### Focus groups + +A focus group is a set of peer project directories. Send scope mirrors receive scope — messages go to everyone you're listening to. + +```bash +attend focus add ~/Projects/foo ~/temp # listen to + send to these +attend focus list # show current group +attend focus clear # project-only mode +``` + +#### Reply hints + +The first peer message includes a reply hint: `(reply: attend send --to /path <msg>)`. Subsequent messages are clean — progressive disclosure, not repetitive nagging. + +### Self-documenting startup + +On launch, attend emits a single usage summary to stdout so Monitor delivers it as Claude's first notification: + +``` +[attend] v0.1.0 (459534c) — sensors: context, git, peers, processes | focus: project + temp | commands: attend send <msg>, attend inbox, attend peers, attend focus add <path> +``` + +### Emission format + +All notifications use bracketed key-value format: + +``` +[attend sensor=context priority=medium] context at 50% — midpoint, wrap-up window opening (burning 1.8%/min, ~25 min to todos checkpoint at 75%) +[attend sensor=peers priority=high] message from aaron: checking updates +[attend sensor=git priority=low] new commits on main (HEAD abc1234 → def5678) +``` + +This format was chosen empirically: Monitor entity-escapes angle brackets but passes square brackets verbatim. + +### Way integration + +A way at `softwaredev/environment/attend` provides progressive disclosure of attend's capabilities. The way body is static CLI reference; a `macro.sh` script (appended at disclosure time) checks live state: whether attend is installed and running, current focus group, active peer count, pending signals. + +The `/attend` skill launches attend via Monitor with explicit instructions to use Monitor (not Bash). + +### Coordination with ways context-threshold triggers + +Ways fires actions at context thresholds: todos@75%, memory@80%, checkpoint@95%. Attend's context sensor provides: + +- Early warnings before ways thresholds (40%, 50%, 65%) +- Verification prompts after ways fires (85% = "verify memory saved") +- Pre-critical warning (92% = "finish task before 95% checkpoint") +- Velocity and projection between all thresholds + +Attend handles trajectory awareness. Ways handles threshold actions. + +### Build and install + +``` +make attend # build (or skip if already built) +make attend-rebuild # force rebuild +make install # symlinks bin/attend → tools/target/release/attend +``` + +Binaries are symlinked to the cargo build output, not copied. Cargo handles atomic replacement of the target binary; the old process keeps running on the old inode. No ETXTBSY on rebuild while attend is running. + +### Hard invariants + +These constraints are non-negotiable: + +1. **Filesystem as transport.** Inter-session awareness uses the filesystem (session files, signal files), not external services. No fragile protocols, no servers, no pubsub. +2. **Informational, not enforceable.** Attend emits observations. It never overrides Claude, never forces action, never bypasses the disclosure gate. +3. **Consequence-anchored.** Context warnings track real mechanical consequences. No arbitrary urgency, no simulated affect. +4. **Metadata-only for content-bearing sensors.** Sensors that touch content-rich sources emit only boolean or categorical state, never the content itself. +5. **Additive, never required.** Attend is optional. Ways that depend on attend signals gracefully no-op when attend is not running. +6. **Explicit invocation only.** Attend never autostarts. +7. **Send scope mirrors receive scope.** Messages go to the same set of projects you're listening to. No silent broadcasting. + +#### Invariant revision note + +The original draft of this ADR included "no cross-session signal" and "no inter-instance protocol" as hard invariants. These were revised after implementation demonstrated that inter-session awareness via the filesystem is stable, useful, and architecturally clean — the same session files Claude Code already publishes, the same pattern abtop already reads. The prior invariants were a reaction to ADR-101/102's failures with fragile external transports, not a principled objection to inter-session awareness itself. The revised invariants preserve the underlying concern (no fragile protocols) while permitting what works (filesystem-based observation and signaling). + +## Consequences + +### Positive + +- **Agency preservation.** Claude gains accurate awareness of approaching consequences in time to act on them. +- **Reduced token waste.** Interoceptive sensors replace expensive self-checks. +- **Proactive environmental awareness.** Git state, peer sessions, process lifecycle — ambient signal at near-zero token cost. +- **Inter-session collaboration.** Multiple Claude instances and human terminals communicate through a shared signal system. No protocol overhead — just files on disk. +- **Sensor toolkit composability.** Adding a new capability means adding a script or a compiled sensor module. The scheduler, delta accumulation, and disclosure governor handle the rest. +- **Cleanly additive.** Sessions that don't need active awareness pay no cost. +- **Standard delivery primitive.** `attend` is a well-behaved Monitor client. Its integration surface is "write lines to stdout." + +### Negative + +- **New binary to maintain.** Mitigation: sibling crate in the same workspace, shares build infrastructure with `ways`. +- **Dependency on Monitor.** If Monitor is unavailable, attend cannot deliver. Mitigation: attend still runs and maintains state; delivery resumes when Monitor becomes available. +- **Signal file management.** Stale signals accumulate if no attend instance polls the directory. Mitigation: 5-minute cleanup on each poll; `attend cleanup` planned. +- **Binary version skew.** Multiple attend instances may run different versions. Mitigation: the signal format and directory layout are stable; version skew doesn't break interop. + +### Neutral + +- **Rust implementation.** Matches `ways`. Zero external dependencies (no serde, no clap, no tokio). +- **XDG conventions.** Signals in `~/.cache/attend/signals/`, config planned for `~/.config/attend/`. + +## Remaining work + +Tracked in aaronsb/agent-ways#2: + +- **Config externalization** — `~/.config/attend/config.yaml` with project-scope overlay +- **Script sensor runner** — poll scripts, parse `magnitude|description` from stdout +- **State persistence** — checkpoint/restore sensor baselines across restarts +- **Self-reload** — watch own binary mtime, exec self on change +- **ADR-114 integration** — affordance strings, `trigger.type: attend` in ways schema +- **Insistence tracker** — unacted observations re-surface with escalating urgency +- **Consequence model** — generalized beyond context pressure + +## Alternatives Considered + +- **Subcommand of `ways`.** Rejected. Conflates stateless per-invocation lifecycle with durable per-session lifecycle. +- **Pure hook-based implementation.** Rejected. Sensors that observe between turns need a persistent process. +- **External transport (wormhole, IRC).** Rejected by ADR-101/102 failures. Filesystem transport is simpler and more stable. +- **Fixed-interval scheduler.** Rejected. Adaptive intervals ramp on change, decay on quiet — better sampling efficiency. +- **Separate repository.** Rejected per ADR-111. Shared workspace avoids coordination churn. + +## References + +- **Design note:** [Cognitive Loop and the Awareness Layer](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) +- **Tracking issue:** [aaronsb/agent-ways#2](https://github.com/aaronsb/agent-ways/issues/2) +- **Related ADRs:** + - [ADR-104](../ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Disclosure gate + - [ADR-111](../platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) — Sibling-crate pattern + - [ADR-112](../archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — Session ledger + - [ADR-114](../ways/ADR-114-attend-as-insistent-way-trigger-type.md) — Way trigger type for attend signals +- **Prior attempts:** + - [ADR-101](./ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) — Wormhole relay (Deprecated) + - [ADR-102](./ADR-102-irc-based-local-agent-communication.md) — IRC-based agent communication (Abandoned) diff --git a/docs/architecture/attend/ADR-117-sensor-crate-extraction-and-feature-flags.md b/docs/architecture/attend/ADR-117-sensor-crate-extraction-and-feature-flags.md new file mode 100644 index 00000000..22cdbce5 --- /dev/null +++ b/docs/architecture/attend/ADR-117-sensor-crate-extraction-and-feature-flags.md @@ -0,0 +1,155 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: the four built-in sensors compile into the attend binary with no isolation, no selective compilation and a blurred Sensor trait boundary + - evidence: 'the crate extraction pattern already proven by agent-fmt (PR #3) and the ScriptSensor runner (PR #4)' + - precedent: ADR-115 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-10 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-115 +imported: + from: docs/architecture/system/ADR-117-sensor-crate-extraction-and-feature-flags.md + format: v0 + status: Accepted +--- + +# ADR-117: Sensor Crate Extraction and Feature Flags + +## Context + +attend (ADR-113) ships with four built-in sensors: context, git, peers, and processes. These are compiled directly into the attend binary as modules under `src/sensors/`. The ScriptSensor runner (shipped in PR #4) provides an extensibility path for shell-script sensors, but built-in sensors have no separation from the orchestrator. + +This creates several problems: + +- **No isolation** — sensor code shares the same module tree as the tick loop, state management, and CLI. Changes to one sensor can break another through shared internal interfaces. +- **No selective compilation** — a user who only wants git and context awareness still compiles peer discovery and process scanning. On constrained systems or custom builds, this matters. +- **No clear trait contract** — the `Sensor` trait lives in `sensors/mod.rs` alongside the `SensorSlot` runtime scaffolding. The boundary between "what a sensor must implement" and "how the orchestrator runs it" is blurred. +- **Future daemon model** — if attend becomes long-running (paralleling the direction noted for ways-cli), hot-reloading sensors or dynamically enabling them requires a cleaner separation than in-binary modules. + +Meanwhile, the workspace already demonstrates the crate extraction pattern: `agent-fmt` was extracted from ways-cli (PR #3) and is consumed by both tools. The same pattern applies here. + +## Decision + +Extract each built-in sensor into its own workspace crate. Introduce a `sensor-trait` crate that defines the `Sensor` trait, `Focus` struct, and supporting types. Wire sensors into attend via Cargo feature flags. + +### Workspace Structure + +``` +tools/ + agent-fmt/ # shared terminal formatting + sensor-trait/ # Sensor trait, Focus, SensorSlot + sensor-git/ # GitSensor + sensor-context/ # ContextSensor + sensor-peers/ # PeerSensor + sensor-processes/ # ProcessSensor + attend/ # orchestrator — depends on sensor crates via features + ways-cli/ # ways CLI +``` + +### Trait Crate + +`sensor-trait` defines the contract between attend and any sensor: + +```rust +pub trait Sensor: Send { + fn name(&self) -> &str; + fn poll(&mut self, focus: &Focus) -> Vec<(f64, String)>; + fn emission_threshold(&self) -> f64; + fn base_interval(&self) -> Duration; + fn min_interval(&self) -> Duration; + fn export_state(&self) -> Vec<(String, String)> { Vec::new() } + fn import_state(&mut self, _state: &[(String, String)]) {} +} +``` + +The `Send` bound prepares for the daemon model where sensors may run on separate threads. `Focus` and `SensorSlot` also move to this crate since they define how the orchestrator interacts with sensors. + +### Feature Flags + +attend's `Cargo.toml`: + +```toml +[features] +default = ["sensor-git", "sensor-context", "sensor-peers", "sensor-processes"] +sensor-git = ["dep:sensor-git"] +sensor-context = ["dep:sensor-context"] +sensor-peers = ["dep:sensor-peers"] +sensor-processes = ["dep:sensor-processes"] +``` + +The orchestrator uses `#[cfg(feature = "sensor-git")]` guards around sensor registration. A minimal build with `--no-default-features` compiles only the orchestrator + script sensor runner. + +### Two Sensor Paths + +This ADR formalizes the two-path model that emerged from PR #4: + +| Path | Implementation | Performance | Extensibility | Compilation | +|------|---------------|-------------|---------------|-------------| +| **Crate sensors** | Rust, compiled in | Native | Requires recompile | Feature-flagged | +| **Script sensors** | Shell, external | Process fork per poll | No recompile | Config-driven | + +Both paths are controlled by the same declarative config (ADR-115). A crate sensor and a script sensor with the same name: the crate sensor wins (compiled-in takes precedence). This allows a script sensor to prototype behavior that later graduates to a crate sensor. + +### Config as Control Plane + +The `attend.yaml` config (ADR-115) controls all sensors uniformly: + +```yaml +sensors: + git: + interval: 30 + threshold: 2.0 + -processes: # disable a built-in sensor + +disk-pressure: # add a script sensor + script: scripts/check-disk.sh + interval: 120 +``` + +Disabling a crate sensor via config (`-processes`) means it's never instantiated at runtime even though the code is compiled in. Disabling via feature flag (`--no-default-features`) means the code isn't compiled at all. Config is the user-facing control plane; features are the build-time control plane. + +## Consequences + +### Positive + +- **Clear contracts** — `sensor-trait` is the documented interface. Any crate implementing `Sensor` can be wired into attend. +- **Selective builds** — custom attend binaries for constrained environments. +- **Isolation** — sensor bugs don't leak across module boundaries. Each sensor has its own dependency tree. +- **Testability** — sensors can be unit-tested in isolation against a mock `Focus`. +- **Graduation path** — script sensor → crate sensor is a defined workflow: prototype in shell, promote to Rust when performance matters. +- **Daemon-ready** — `Send` bound and crate isolation prepare for concurrent sensor polling. + +### Negative + +- **More crates** — workspace goes from 3 to 7 members. Cargo handles this well but it's more manifests to maintain. +- **Cross-crate changes** — modifying the `Sensor` trait requires updating all sensor crates. Mitigated by keeping the trait stable and small. +- **Initial migration effort** — moving existing sensor code is mechanical but touches every sensor file. + +### Neutral + +- **No behavioral change** — the default feature set compiles all sensors, reproducing current behavior exactly. +- **ScriptSensor stays in attend** — it's part of the orchestrator (runs any script), not a specific sensor implementation. +- **Config format unchanged** — ADR-115's config works identically before and after extraction. + +## Implementation Plan + +1. Create `sensor-trait` crate with `Sensor`, `Focus`, `SensorSlot`, `AdaptiveInterval`, `DeltaAccumulator` +2. Create `sensor-git` crate, move `sensors/git.rs` content, depend on `sensor-trait` +3. Create `sensor-context` crate, move `sensors/context.rs` +4. Create `sensor-peers` crate, move `sensors/peer.rs` (largest, has signal reading) +5. Create `sensor-processes` crate, move `sensors/process.rs` +6. Update attend `Cargo.toml` with feature flags, `#[cfg]` guards in sensor registration +7. Update `sensors/mod.rs` to re-export from crates (compatibility shim, removable later) +8. Verify `make lint`, `make attend-rebuild`, all existing behavior preserved +9. Test minimal build: `cargo build -p attend --no-default-features` +10. Update CI workflow to test both default and minimal feature sets diff --git a/docs/architecture/attend/ADR-118-focus-groups-dynamic-agent-grouping.md b/docs/architecture/attend/ADR-118-focus-groups-dynamic-agent-grouping.md new file mode 100644 index 00000000..77ac35c8 --- /dev/null +++ b/docs/architecture/attend/ADR-118-focus-groups-dynamic-agent-grouping.md @@ -0,0 +1,379 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: 'focus groups as static path lists were brittle in use: paths are implementation details, groups did not self-clean, and there was no middle ground between just me and everyone' + - precedent: ADR-113 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-10 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-115 + - ADR-119 + - ADR-120 +imported: + from: docs/architecture/system/ADR-118-focus-groups-dynamic-agent-grouping.md + format: v0 + status: Accepted +--- + +# ADR-118: Focus Groups — Dynamic Agent Grouping + +## Context + +attend's original grouping model had focus groups implemented as static path lists — `attend focus add ~/Projects/foo` required knowing the exact filesystem path of a peer. This was brittle: paths are implementation details, groups didn't self-clean, and there was no dynamic middle ground between "just me" and "everyone." + +The underlying signal model already worked through named directories. A project's signals live in a directory named by its encoded path. Broadcast signals live in `_broadcast/`. The infrastructure for named signal namespaces existed — it just wasn't exposed as the user-facing concept. + +The fix was to rebuild focus groups as **named signal namespaces** rather than path lists. The name "focus groups" was kept because the concept is correct — agents focus on named groups to coordinate, and release focus when done. The implementation changed; the vocabulary didn't. + +### Home Assistant Analogy + +Home Assistant scenes configure groups of devices into named states — "movie night" dims lights and turns on the TV. Scenes are declarative, activatable, and composable. attend's scenes are the equivalent: named configurations of focus group membership and attention profiles. + +## Decision + +Focus groups are **named signal namespaces** that agents join and leave dynamically. They are the universal grouping mechanism alongside the implicit project scope and broadcast. + +### Focus Group Model + +Every agent is always in one **implicit group** — its project path, auto-assigned from cwd. This is the project scope. + +Agents can also join any number of **named focus groups**. A focus group is any string — `deploy`, `infra`, `code-review`. It doesn't need to map to a filesystem path. + +``` +attend focus on deploy # focus on a named group +attend focus off deploy # release focus +attend focus list # show groups you're focused on +attend focus all # show all active groups and their members +attend focus clear # release all named groups (project-only mode) +``` + +### Signal Routing + +Signals are routed to focus groups, not paths. When an agent sends a message: + +- `attend send "msg"` — sends to project scope + joined focus groups +- `attend send --focus deploy "msg"` — sends to a named focus group +- `attend send --broadcast "msg"` — sends to broadcast (all agents, human↔agent channel) + +When an agent receives, it sees signals from: +- Its own project scope (always) +- Any named focus groups it has joined +- Broadcast (always) + +``` +# Agent A (in ~/Projects/foo): +attend focus on collab + +# Agent B (in ~/Projects/bar): +attend focus on collab + +# Now A and B see each other through "collab", without knowing paths +``` + +### Focus Group Lifecycle + +- **Creation**: implicit — focusing on a group that doesn't exist creates it +- **Ephemeral** (default): a group with no members is cleaned up on the next peer sensor poll +- **Pinned**: `attend focus on deploy --pin` marks a group to persist even when empty. Useful for standing workgroups that agents rejoin across sessions. `attend focus unpin deploy` reverses it. +- **Dissolution**: `attend focus dissolve deploy` removes the group and notifies all members + +### Scenes + +A scene is a named preset that configures focus group membership and attention profiles: + +```yaml +# ~/.config/attend/scenes.yaml +private: + groups: [] # release all focus groups, project scope only + +workroom: + groups: [deploy] # focus on just this group + +open: + groups: ["*"] # join all discoverable groups +``` + +``` +attend scene private # activate a scene +attend scene open +attend scenes # list available scenes +``` + +Scenes are sugar over focus on/off. `attend scene private` is equivalent to releasing all named groups. `attend scene open` joins a well-known shared group. + +### Scoped Attention Profiles + +Scenes configure more than group membership. A scene sets two **attention scopes** that mirror Claude Code's own permission model: + +- **Project scope** — sensors and governor settings for local repo work (git changes, context usage, process detection) +- **Focus scope** — sensors and governor settings for coordination participation (peer signals, response cadence) + +```yaml +# ~/.config/attend/scenes.yaml +deep-work: + groups: [] + project: + sensors: [git, context] + governor: + base_cooldown: 30 + focus: + sensors: [] # no peer sensor — no signals arrive + governor: + base_cooldown: 60 + +coordinate: + groups: [deploy, infra] + project: + sensors: [git, context] + governor: + base_cooldown: 45 # slower project disclosure during coordination + focus: + sensors: [peers] + governor: + base_cooldown: 10 # fast response to peer signals + +private: + groups: [] + project: + sensors: [git, context] + focus: + sensors: [] +``` + +The two scopes run simultaneously with independent governors. A git change in your repo discloses on the project cadence. A peer signal in `@deploy` discloses on the focus cadence. This prevents cross-contamination — focus group coordination doesn't interrupt git operations, and local file churn doesn't drown out peer signals. + +**Scope inheritance**: scenes without explicit scope blocks inherit the global config defaults. A minimal scene like `private: groups: []` still works — it just uses the global governor for both scopes and disables focus sensors implicitly by having no groups. + +**Why two scopes, not per-sensor config**: individual sensor tuning is already possible in `attend.yaml`. Scopes are coarser — they separate the *kind of attention* (local work vs. coordination) rather than tweaking individual sensors. An agent in `deep-work` shouldn't see peer signals at all, not just see them slower. An agent in `coordinate` should respond to peers quickly but is still doing project work — it needs both attention modes running with different parameters. + +**Session persistence**: scenes are saved to `scenes.yaml` and the active scene is tracked per project in `_groups.yaml`. When an agent restarts in the same project directory, attend can auto-activate the last scene. This means working groups survive session boundaries — a `coordinate` scene with groups `[deploy, infra]` persists across agent restarts without the human reconfiguring. Combined with pinned groups (which persist even when empty), a set of related projects maintains its coordination topology across sessions. The human sets up the working group once; it reconstitutes each time. + +**Project-scoped scene overrides**: `scenes.yaml` supports both user scope (`~/.config/attend/scenes.yaml`) and project scope (`.attend/scenes.yaml` in a repo). Project-scoped scenes overlay user-scoped ones, so a repo can ship a default scene that configures the right focus groups and attention profile for agents working in that project. This means related repos can pre-declare their coordination topology. + +**Chat TUI visibility** (ADR-120): the TUI sidebar shows each agent's current scene and scope state. You can see at a glance that `api: deep-work` means focus scope is empty (won't see your `@deploy` message) while `infra: coordinate` has focus scope active (will respond). This informs the human's steering decisions — you know when to send a broadcast (reaches everyone regardless of scene) vs. a focus message (only reaches agents focused on that group). + +### Unified View + +`attend peers` and `attend status` merge into a single view organized by focus groups: + +``` +$ attend peers + Focus Agent Status Context + ──────────────────────────────────────────────────────── + (project) agent-ways working 45% + deploy api-server waiting 12% + deploy infra-tools working 30% + (broadcast) game-ai-pro waiting 14% +``` + +`attend status` becomes a self-view (your groups, your signals, your config) rather than a separate system view. + +### Storage + +Focus groups are directories under the existing signals base: + +``` +~/.cache/attend/signals/ + -home-aaron--claude/ # project scope (existing, unchanged) + _broadcast/ # broadcast (existing, unchanged) + @deploy/ # named focus group (@ prefix distinguishes from encoded paths) + @collab/ # another named focus group + _groups.yaml # group membership + pinned state + active scene +``` + +The `@` prefix prevents collision between named groups and encoded project paths. `_groups.yaml` tracks which groups each session has joined and which are pinned. + +## UX Flows + +### Flow 1: Solo work (default, no action needed) + +Agent starts. It's in its project scope automatically. No peers, no noise. + +``` +$ attend peers + Focus Agent Status Context + ────────────────────────────────────────────────── + agent-ways (you) working 12% + + 1 agent, 1 group +``` + +### Flow 2: Ad-hoc collaboration + +Aaron spins up two agents and wants them to coordinate. + +``` +# In agent A's session (agent-ways): +$ attend focus on deploy + +# In agent B's session (api-server): +$ attend focus on deploy + +# Now both see each other: +$ attend peers + Focus Agent Status Context + ────────────────────────────────────────────────── + agent-ways (you) working 12% + deploy api-server waiting 8% + + 2 agents, 2 groups +``` + +Signals flow through the group: +``` +# Agent A: +$ attend send --focus deploy "migrations are done, ready for deploy" + +# Agent B sees it via the peer sensor: +[attend sensor=peers] message from agent-ways in deploy: migrations are done, ready for deploy +``` + +When done, agents release focus or sessions end: +``` +$ attend focus off deploy +# If both leave, "deploy" is cleaned up on next poll +``` + +### Flow 3: Standing workgroup + +A team of agents that reconvene across sessions. + +``` +$ attend focus on infra --pin +# --pin keeps the group alive even when empty +# Next time an agent starts, it can discover and rejoin: + +$ attend focus all + Group Members Pinned + ───────────────────────────────── + deploy 0 no (will be cleaned up) + infra 0 yes (persists) + +$ attend focus on infra +``` + +### Flow 4: Scene switch + +Aaron wants all agents in private mode while he's in a meeting, then open mode after. + +``` +# From any terminal: +$ attend scene private +# → releases all focus groups, project scope only +# → writes scene signal to broadcast so other agents' attend instances pick it up + +# Later: +$ attend scene open +# → joins the well-known "open" group +# → all agents with attend running see the scene change and auto-join +``` + +### Flow 5: Discovery — "what groups exist?" + +``` +$ attend focus all + Group Members Pinned + ───────────────────────────────── + deploy 2 no + infra 1 yes + daily 0 yes + +# Join one: +$ attend focus on deploy +``` + +### Flow 6: Directing a message without joining + +Sometimes you want to send a message to a group without subscribing to it. + +``` +$ attend send --focus infra "heads up: the CI cert expires Friday" +# Message lands in the group, but you don't join it or receive from it +``` + +### Flow 7: Human sends from terminal + +Aaron is in a terminal, not in a Claude session. He wants to poke agents. + +``` +# Send to a specific group: +$ attend send --focus deploy "hold off, I'm rolling back" + +# Send to broadcast (all agents): +$ attend send --broadcast "going to lunch, back in 30" +``` + +### CLI Summary + +``` +attend focus on <name> [--pin] Focus on a group (create if needed, --pin to persist) +attend focus off <name> Release focus from a group +attend focus list Show groups you're focused on +attend focus all Show all active groups with member counts +attend focus clear Release all groups (project-only mode) +attend focus pin <name> Pin a group (persist when empty) +attend focus unpin <name> Unpin a group +attend focus dissolve <name> Remove a group, notify members + +attend send "msg" Send to project scope + joined focus groups +attend send --focus <name> "msg" Send to a named focus group +attend send --broadcast "msg" Send to all sessions (human↔agent channel) + +attend scene <name> Activate a scene (apply attention profile) +attend scene save <name> Save current state as a scene (overwrite) +attend scene save-as <new-name> Save current state as a new scene +attend scene edit <name> Open scene definition for editing +attend scene delete <name> Delete a scene (built-ins protected) +attend scenes List available scenes with scope summaries + +attend peers Unified view: agents grouped by focus +attend status Self-view: your groups, signals, config +``` + +## Consequences + +### Positive + +- **Dynamic grouping** — form and dissolve workgroups without config file edits +- **Name-based targeting** — `--focus deploy` instead of `--to /home/aaron/Projects/...` +- **Self-cleaning** — ephemeral groups dissolve when everyone leaves +- **Unified view** — one `attend peers` organized by groups, not two overlapping commands +- **Scene presets** — named configurations combining group membership + attention profiles +- **Scoped attention** — project and focus scopes with independent governors prevent cross-contamination +- **Session persistence** — working groups survive restarts via scene tracking and pinned groups +- **Backward compatible** — project scope is implicit, broadcast unchanged, existing signals work + +### Negative + +- **Discovery** — agents need a way to discover group names. `attend focus all` helps, but the initial "how do I know what groups exist" requires either convention or the scene mechanism. +- **Dual governors** — splitting project/focus scope adds complexity to the disclosure path + +### Neutral + +- **Broadcast is a built-in group** — conceptually, broadcast is a group everyone is always in. Implementation may or may not change. +- **Project scope is implicit** — no behavioral change for single-agent workflows +- **CLI vocabulary matches code** — `attend focus on/off` maps directly to `groups.rs` internals + +## Implementation Plan + +### Phase 1: Scoped Scenes +1. Extend scene config schema with `project:` and `focus:` scope blocks (sensors, governor) +2. Split disclosure governor into project-scoped and focus-scoped instances +3. Implement scene CRUD: `attend scene save/save-as/edit/delete` +4. Track active scene per project in `_groups.yaml` for session persistence +5. Auto-activate last scene on restart + +### Phase 2: Project-Scoped Config +6. Support `.attend/scenes.yaml` in repos for project-scoped scene defaults +7. Overlay project scenes onto user scenes (project wins on conflict) +8. Update the `/attend` skill documentation diff --git a/docs/architecture/attend/ADR-119-action-potential-engagement-model.md b/docs/architecture/attend/ADR-119-action-potential-engagement-model.md new file mode 100644 index 00000000..88cffab5 --- /dev/null +++ b/docs/architecture/attend/ADR-119-action-potential-engagement-model.md @@ -0,0 +1,181 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +superseded_by: ADR-123 +basis: + - evidence: 'the party problem: agents burn context on extended peer conversations with no natural disengagement signal under flat thresholds and linear cooldowns' + - evidence: Game AI Pro chapter 2, Informing Game AI Through the Study of Neurology, for the action potential model + - precedent: ADR-113 +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-04-12 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-118 + - ADR-123 +imported: + from: docs/architecture/system/ADR-119-action-potential-engagement-model.md + format: v0 + status: Superseded + unmapped: + revised: 2026-04-14 +--- + +# ADR-119: Action Potential Engagement Model + +## Status: Superseded by ADR-123 + +**Superseded 2026-04-14.** The underlying decision — model agent engagement on the neuronal action potential so refractory periods produce natural disengagement from diminishing-value stimuli — is still load-bearing and shipping in production attend. The specific implementation shape described below (linear per-minute decay, time-windowed burst detection, `EngagementState` owning `Instant` timestamps, `step_multiplier` that scales the peak multiplier per fire past threshold) has been replaced by the shared curve engine introduced in [ADR-123](../ways/ADR-123-firing-dynamics-progression-axis-unification.md). + +Concretely, what changed: + +- **Linear decay → exponential half-life.** The old `current_multiplier` decayed `peak - elapsed_min × decay_per_minute` and clamped at 1.0. The new `Curve::ActionPotential` decays `1.0 + (peak - 1.0) × 0.5^(delta / multiplier_half_life)`. Attend's yaml field `decay_per_minute` is preserved for back-compat and converted to `multiplier_half_life` at load time via `ln(0.5) / ln(1 - rate) × 60`. In the load-bearing first few minutes post-burst the two shapes match closely; at the tail they diverge (the new curve never quite reaches 1.0, while the old linear one reached rest and clamped). +- **Time-windowed burst detection → event-count burst detection.** The old model counted fires within the last `burst_window` seconds. The new model counts fires whose exponential contribution to the multiplier hasn't decayed past an epsilon. For attend on wall-clock seconds the practical difference is negligible; for ways on chunky token-position ticks the new model is the only one that works at all. See ADR-123 Decision 2 for the full argument. +- **Per-fire scaling peak → fixed ceiling.** The old model computed `peak = 1 + steps × step_multiplier` where `steps = burst_count - burst_threshold + 1`, so additional fires past threshold kept raising the peak. The new model uses `peak_multiplier = 1 + step_multiplier` (= 2.25 at defaults) as a fixed ceiling. The scaling rarely activated in practice and the flat ceiling is simpler to reason about. +- **`Instant`/`Duration` → `Tick`/`TickDelta`.** The engine is now unit-agnostic. Attend interprets ticks as wall-clock seconds via `sensor_trait::epoch_secs()`; ways interprets them as token position. The engine does not know which. +- **Shared crate.** The engine lives in `sensor-trait::engagement` (and `sensor-trait::curve`) and is consumed by both attend's `SensorSlot` and ways' `session::way_fire_outcome`. Pre-ADR-123, attend had its own `EngagementState` in `sensor-trait` and ways had nothing of the kind — firing was a flat `token_distance_exceeded` step. Now it's the same engine in both tools. + +The biology framing, the "party problem" motivation, the urgency-escape pattern, and the per-peer auto-grouping extension are all still load-bearing and carry over intact. The implementation details below are historical — they describe what was accepted on 2026-04-12, not what runs today. For the current implementation see `docs/attend-and-monitor/engagement.md` and [ADR-123](../ways/ADR-123-firing-dynamics-progression-axis-unification.md). + +## Context + +attend's disclosure governor uses linear cooldowns and rate windows to manage notification frequency. This prevents flooding but doesn't model the natural dynamics of productive engagement. An agent responding to peer messages has no signal that engagement value is declining — it will keep responding at the same threshold indefinitely, leading to unbounded context spend on diminishing-value conversations. + +The current model: +- **Flat threshold**: a stimulus either meets the emission threshold or doesn't, regardless of recent activity +- **Linear cooldown**: fixed time between disclosures, no relationship to engagement history +- **No refractory period**: an agent that just finished a burst of peer messaging can immediately start another + +This creates the "party problem" — agents can burn context on extended peer conversations with no natural disengagement signal. The human has to intervene or the context window runs out. + +## Decision (as of 2026-04-12) + +Model agent engagement after the neuronal action potential. The biological signal has properties that map directly to productive agent behavior: + +### The Action Potential Shape + +``` + Engagement + (magnitude) + ^ + +30 | * peak + | / \ + | / \ + | / \ + 0 | / \ + | / \ + -55 |--* \ threshold + | stimulus \ + -70 |.................\___*___........ resting + | refractory + +--------------------------------> time +``` + +### Phases + +1. **Resting state** — agent at baseline awareness. Sensors poll, observations accumulate. No urgency. +2. **Stimulus & threshold** — an observation accumulates magnitude. Sub-threshold stimuli decay without triggering engagement. Only stimuli crossing threshold fire a response. +3. **Depolarization** — rapid engagement. Active responding. Should not be suppressed. +4. **Peak** — maximum engagement value. Highest information density. +5. **Repolarization** — diminishing returns. Continued engagement on the same topic yields less new information per context token spent. +6. **Refractory period** — after a burst of engagement, the threshold temporarily rises. The agent resists re-engaging with the same stimulus category. Urgent new stimuli (high magnitude) can still break through. + +### Absolute vs relative refractory + +- **Absolute** (first N seconds after burst): no disclosure from this sensor regardless of magnitude. +- **Relative** (decay window): disclosure possible but requires elevated magnitude. Casual follow-ups don't fire; urgent signals do. + +### Urgency escape + +Directed messages (sent with `--to <project>`) start at base magnitude 7.0 — above the elevated threshold even during burst refractory. This preserves urgency discrimination without special-case logic. + +### Per-peer auto-grouping + +sensor-peers pairs the refractory model with a per-peer magnitude boost: messages from peers who've sent multiple messages within `peer_activity_window` get 1.75× (2nd) or 2.5× (3rd+) their base magnitude. The boost lifts active conversation partners above the elevated threshold while uninvolved peers stay below it. Conversation topology emerges from observed traffic rather than explicit group configuration. + +## What ADR-123 changed (2026-04-14) + +The decision above stands. What changed is the shape of the state machine that implements it: + +- `EngagementState` is now generic over a `Curve` variant. Attend uses `Curve::ActionPotential { burst_threshold, peak_multiplier, absolute_refractory, multiplier_half_life }`. +- The tick axis is supplied by the caller. Attend passes `sensor_trait::epoch_secs()` (wall-clock seconds); everything else the engine does is unit-agnostic. +- Burst detection is event-count based, not tick-windowed. This costs nothing for attend and unlocks the entire ways integration. +- The engine exposes `should_fire(tick, magnitude)`, `record_fire(tick, magnitude)`, `current_salience(tick)`, `current_multiplier(tick)`. Attend's `SensorSlot` calls these through `in_absolute_refractory(tick)` and `effective_threshold(base, tick)` helpers that preserve the pre-ADR-123 call-site shape. +- A sibling outward-gate consumer (ways) runs the same engine with a different curve (`Curve::Exponential`) on a different axis (token position). The unification argument is in [ADR-123 Decision 4](../ways/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory). + +### Attend yaml ↔ runtime mapping + +The yaml keys in attend's engagement config are stable. At load time they map onto the curve variant: + +| yaml key | runtime parameter | +|-----------------------|----------------------------------------------------| +| `burst_threshold` | `burst_threshold` | +| `step_multiplier` | `peak_multiplier = 1.0 + step_multiplier` | +| `absolute_refractory` | `absolute_refractory` (seconds) | +| `decay_per_minute` | `multiplier_half_life = ln(0.5)/ln(1-rate) × 60` | +| `burst_window` | *no runtime effect* (DEPRECATED — flagged by `attend config lint`) | +| `peer_activity_window`| consumed directly by sensor-peers | + +### What did NOT change + +- The action potential framing (biology analogy, refractory semantics, urgency escape, auto-grouping). +- `attend tune`'s session survey → config derivation. It still emits `burst_window`/`decay_per_minute` in the pre-ADR-123 field names; the runtime converts at load time. +- `EngagementConfig` in `attend::config`, with all six yaml knobs preserved for back-compat. +- The disclosure governor (ADR-113) and focus groups (ADR-118) compose with the engagement gate the same way. + +## Consequences + +### Positive (preserved) + +- Natural disengagement — agents stop engaging when returns diminish, without rules +- Burst tolerance — rapid-fire engagement during productive phases is not suppressed +- Urgency discrimination — truly urgent signals break through refractory +- Self-regulating — no human intervention needed for the party problem +- Biologically grounded — the model is well-studied, predictable, intuitive + +### Positive (added by ADR-123) + +- Single source of truth for firing dynamics. The math lives in one crate. Drift between attend and ways is structurally impossible. +- Exponential decay shape matches attention fade more closely than the old linear approximation +- Event-count burst detection is robust to axis granularity — same engine works for attend's smooth seconds and ways' chunky tokens +- Curve-as-parameter enables progressive disclosure and flat step curves for ways that want them + +### Negative (from 2026-04-12) + +- **Complexity** — more state per sensor. ADR-123 didn't reduce this; it shared it across tools. +- **Tuning** — refractory parameters still need empirical calibration. `attend tune` is a first pass; `ways tune` (deferred) would apply the same discipline to the ways side. +- **Opaque** — harder for users to understand why an agent isn't responding to a message (elevated threshold). Partially mitigated by `ways list` showing per-way re-fire distances; no equivalent yet for attend's refractory state in `attend status`. + +## Implementation Status + +The ADR-119 acceptance landed in attend via the original `EngagementState` in `sensor-trait` (linear decay, time-windowed burst detection, `Instant`-keyed). That implementation was retired in the 2026-04-14 ADR-123 work and replaced by `EngagementState` backed by `Curve::ActionPotential`. The present-day implementation ships: + +- `sensor_trait::engagement::EngagementState` with serde derives for per-session persistence +- `sensor_trait::curve::Curve::ActionPotential` with event-count burst detection +- `sensor_trait::epoch_secs()` as the canonical attend tick source +- `attend::config::EngagementConfig` with yaml-stable field names converted at load time +- `attend tune` for empirical parameter derivation from real session history +- `attend config lint` / `--fix` to surface and remove the deprecated `burst_window` yaml key + +### Not yet implemented (unchanged from 2026-04-12) + +- Refractory state visible in `attend status` output (current sensor tick log shows it; the status table does not) +- Idle/motivation sensor for intrinsic self-prompting +- Motivation sensor wiring to a reflection-overdue way +- Long-horizon empirical tuning (initial defaults are from a single session survey) + +## References + +- **[ADR-123](../ways/ADR-123-firing-dynamics-progression-axis-unification.md)** — progression-axis unification, curve enum, shared engine. +- **ADR-113** — attend active awareness module; disclosure governor, emission thresholds. +- **ADR-118** — focus groups; scoping which stimuli reach an agent. +- **Game AI Pro Chapter 2**: "Informing Game AI Through the Study of Neurology" — action potential diagram and neurological grounding. +- **Cognitive Frameworks paper** — cognitive economics, cheapest path = correct path. +- `docs/attend-and-monitor/engagement.md` — the current implementer-and-author-friendly explainer. diff --git a/docs/architecture/attend/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md b/docs/architecture/attend/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md new file mode 100644 index 00000000..fc36072c --- /dev/null +++ b/docs/architecture/attend/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md @@ -0,0 +1,197 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: agents misroute messages, e.g. sending --to /home/aaron/.claude instead of --broadcast to report to the operator + - evidence: prior-art survey of terminal chat clients (WeeChat, Matterhorn, iamb) and agent orchestration tools (Agent Deck, Agent of Empires) + - precedent: ADR-118 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-11 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-118 + - ADR-119 +imported: + from: docs/architecture/system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md + format: v0 + status: Accepted +--- + +# ADR-120: Interactive Chat TUI — Human in the Signal Loop + +## Context + +attend provides inter-session signaling for Claude Code agents, but the human has no first-person view of the signal topology. The human's current tools are: + +- **`attend send --broadcast`** from a bare terminal — fire-and-forget, no conversational flow +- **`attend inbox`** — a snapshot, not a stream +- **Switching between Claude terminal sessions** — one agent at a time, no overview + +This creates two problems: + +**1. Routing confusion.** Agents misroute messages because the routing semantics (broadcast, focus group, project scope) are invisible abstractions. In a typical failure, an agent wanting to "report to Aaron" sends `--to /home/aaron/.claude` instead of `--broadcast`, because it reasons about paths rather than intent. The routing model has no physical form that teaches correct usage. + +**2. No orchestration surface.** When running 4+ Claude sessions across related projects, the human has no way to see the whole board. Steering requires switching to each agent's terminal individually. There is no equivalent of a team lead walking through an open-plan office, hearing conversations, and dropping context where needed. + +### The real-world workflow + +A concrete use case that motivates this: + +- A Konsole window with 4 Claude sessions working on related projects (the team) +- A separate terminal running `attend chat` (the coordination surface) +- A GitHub sensor (new attend sensor) detects issue state changes in each agent's repo — an issue moves from backlog to "in progress" and attend notifies the relevant agent +- The human watches progress in the chat TUI, steers with targeted messages (`@infra hold off, @api needs to land first`) +- When deep focus is needed, the human switches to that agent's terminal session directly +- The chat TUI is peripheral vision; the Claude sessions are foveal vision + +### Prior art + +Terminal chat clients have solved multi-channel navigation and threading: + +- **WeeChat** — buffer model (every channel = a numbered buffer with activity indicators), split views, relay protocol decoupling UI from engine +- **Matterhorn** — three-tier channel switching (all / unread / by-name), dedicated thread pane +- **iamb** — Vim-modal navigation, tabs for spaces, splits for rooms, built in Rust with ratatui + +Multi-agent orchestration tools have solved session management: + +- **Agent Deck** — status filters (`!@#$` for running/waiting/idle/error), tmux status bar integration, cost tracking +- **Agent of Empires** — tmux-native persistence, git worktree isolation per agent + +What nothing has built: a tool where the human sits *inside* the same signal protocol the agents use, sending and receiving through the same routing infrastructure, with contextual dispatch (`@group` addressing, `#issue` references) from within the stream. + +## Decision + +Add `attend chat` — an interactive TUI that places the human inside the signal loop as a first-class participant. + +### Core design + +The TUI is an **orchestration surface, not a work surface**. Deep work happens in each agent's terminal. The chat TUI provides: + +- **Visibility** — see all signals across all rooms in real-time +- **Steering** — send targeted messages to groups or broadcast to all +- **Awareness** — ambient metadata (agent status, branch, context %) alongside messages + +### Framework + +**iocraft** (Rust, React-like TUI framework). Chosen for: + +- Declarative layout via `element!` macro with flexbox semantics (taffy) +- React hooks model (`use_state`, `use_future`, `use_terminal_events`) +- Component composition via `#[component]` functions +- First-class async for watching signal directories +- Mouse event support for click-to-reply + +iocraft provides the layout/state scaffolding. Custom components (scrollable message list, focus group sidebar with activity indicators) are built on top. + +### Layout + +``` +┌─────────────┬──────────────────────────────────────┐ +│ Focus │ Messages (chronological stream) │ +│ │ │ +│ ● broadcast │ claude/api @deploy │ +│ @deploy │ Landed the auth refactor, ready for │ +│ @infra │ review. │ +│ project/… │ │ +│ │ claude/infra @deploy │ +│─────────────│ Pulling in the new auth types now. │ +│ Agents │ │ +│ │ aaron broadcast │ +│ api ↑2 73% │ @infra hold off until api merges. │ +│ infra · 91%│ │ +│ docs · 45% │ │ +│ slack ↑1 88│ │ +│ ├──────────────────────────────────────┤ +│ │ > @deploy looks good, ship it_ │ +└─────────────┴──────────────────────────────────────┘ +``` + +- **Left sidebar, top**: Focus group list with activity indicators (● = unread). Click or tab to filter message view. +- **Left sidebar, bottom**: Agent status — name, upstream commits, context %. Pulled from peer sensor. +- **Main area**: Chronological message stream, scoped to selected group or "all". +- **Input bar**: Text input with `@group` and `#issue` inline addressing. + +### Inline addressing + +- **`@name`** at the start of a message routes to that focus group: `@infra look at #172` → writes signal to `@infra` directory +- **`#NNN`** is a repo-local issue reference — the TUI doesn't resolve it. The receiving agent interprets `#172` against its own repo via `gh issue view 172` +- **No prefix** → broadcast (human default, inverse of agent default which is project-scoped) + +The human's default is broadcast because the human is the coordination layer — most messages should be visible to all. Agents' default is project-scoped because most agent work is local. + +### Signal protocol addition: threading + +Current signal format: +``` +from|project|cwd|message +``` + +Extended format: +``` +from|project|cwd|re:signal-id|message +``` + +The `re:` field is optional (empty string for new threads). When an agent or human replies to a specific signal, the reply carries the original signal's ID. The TUI uses this to draw thread lines; agents can use it to maintain conversational context. This is a lightweight addition — no thread trees, no nested replies, just one level of "in response to." + +### GitHub sensor (new attend sensor) + +A new sensor alongside git, peers, and processes: + +- **Polls**: issue state changes, PR status, review requests via `gh` CLI +- **Scope**: each agent's own repo only (not cross-repo) +- **Surfaces**: issue assignments, status transitions (backlog → in progress), PR merge/close, review comments +- **Signal format**: standard attend signals, routed to the agent's project scope + +This enables the workflow: human moves an issue to "in progress" on GitHub → attend notifies the agent → agent picks up the work. The TUI shows the notification alongside other signals. + +### What the TUI does NOT do + +- **Replace Claude sessions** — deep work happens in the agent's terminal +- **Fetch GitHub data** — `#issue` references are resolved by agents, not the TUI +- **Manage agent lifecycle** — starting/stopping agents is outside scope +- **Provide an editor** — no code editing, no file browsing + +## Consequences + +### Positive + +- Routing semantics become visible — the human sees where signals land, which rooms are active, how messages flow +- Cross-agent coordination without terminal switching +- The human experiences the same signal protocol agents use, creating empathy for routing design (dogfooding) +- GitHub sensor closes the loop between project management and agent work +- Threading enables conversational context without full chat protocol complexity + +### Negative + +- New dependency: iocraft (0.8.0, still pre-1.0) +- The TUI is another process to run alongside agents — adds operational surface +- Threading changes the signal format — existing signal readers need to handle the new field +- GitHub sensor requires `gh` CLI auth in each agent's environment + +### Neutral + +- ADR-118 (focus groups) — the TUI's sidebar directly depends on the focus group model and scoped scenes +- ADR-119 (action potential) interactions — agents in the TUI message stream are subject to the same engagement dynamics +- The attend binary grows a significant new subcommand; may warrant a separate binary or feature flag if it pulls in heavy TUI dependencies + +## Alternatives Considered + +### Web dashboard +A browser-based UI showing agent status and signals. Rejected because it breaks the terminal-native workflow — the human would context-switch between browser and terminals. The TUI stays in the same environment as the agents. + +### tmux wrapper +Use tmux panes to tile agent sessions with a monitoring pane. This is what Agent of Empires does. Rejected because tmux panes show raw session output, not the signal layer. You'd see everything an agent does, not the coordination-relevant signals. Too much information, no routing visibility. + +### Extend existing tools (Agent Deck, etc.) +These are session managers — they manage agent lifecycle and show status. They don't participate in a signal protocol. attend's signal infrastructure is the differentiator; building on a session manager would mean reimplementing the signal layer or bridging two systems. + +### ratatui (imperative layout) +The ecosystem standard for Rust TUIs. Rejected as primary framework because layout is imperative (compute rects, render into them) rather than declarative. For a chat UI with dynamic group lists, resizable panes, and nested components, declarative flexbox is significantly more productive. ratatui remains available as a fallback for custom widget rendering inside iocraft components if needed. diff --git a/docs/architecture/attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md b/docs/architecture/attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md new file mode 100644 index 00000000..55411089 --- /dev/null +++ b/docs/architecture/attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md @@ -0,0 +1,183 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +superseded_by: ADR-123 +basis: + - evidence: 'a survey of 18 recent active sessions: median 12 turns, p75 45, p90 84, max 133, with early signals still presented 40 to 80 turns later' + - precedent: ADR-119 +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-04-12 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-114 + - ADR-119 + - ADR-123 +imported: + from: docs/architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md + format: v0 + status: Superseded + unmapped: + revised: 2026-04-14 +--- + +# ADR-121: Salience decay for signal presentation — turn-based exponential + +## Status: Superseded by ADR-123 + +**Superseded 2026-04-14** while the specific attend-side application remained deferred. This ADR introduced two ideas: (1) firing dynamics decompose into an **inward gate** ("should this new stimulus fire?") and an **outward gate** ("should this already-fired signal still be presented?"), and (2) the outward gate for attend's peer signals should be turn-based exponential salience decay with a configurable floor. The inward/outward gate framing became the load-bearing architectural contribution of [ADR-123](../ways/ADR-123-firing-dynamics-progression-axis-unification.md), which generalized both gates across tools via the shared curve engine. + +The outward-gate decision shipped first **on the ways side, not the attend side**: + +- Ways now runs `Curve::Exponential` as its outward gate via `session::way_fire_outcome()`, gating re-fire on `current_salience(current_tick) < REFIRE_FLOOR` (default 0.5). Every way declares an explicit `curve:` block in its frontmatter. Tick is token position, not turn count, because ways' progression axis is the host's addressing unit (see ADR-123 Decision 4). +- Attend's peer-signal presentation still ships the pre-ADR-121 behavior: signals appear in notifications indefinitely until the 30-day disk cleanup removes them. The turn-based aging described below is **still deferred**. When it lands, it will use the same `EngagementState` / `Curve::Exponential` infrastructure the ways side already uses, but with a turn-count progression axis instead of tokens. + +So this ADR's framing is shipping (in a different place than originally targeted) and its application is pending (in the place originally targeted). The body below describes the 2026-04-12 decision; the "What ADR-123 changed" section below reconciles it with what actually exists today. + +## Context + +ADR-119 (action potential engagement model) governs *when new stimuli break through* to fire a disclosure. Disk retention (default 30 days) governs *how long signal files live on disk*. Neither mechanism addresses the middle case: a signal that has already been disclosed, is still on disk, and is still within the "present to the agent" set at every turn regardless of whether it's still load-bearing. + +In a long session, this manifests as presentation bloat — signals from the first few turns keep appearing in notifications 40, 60, 80 turns later, long after their context has been overtaken. A survey of 18 recent active sessions shows the distribution is long-tailed: median 12 turns, p75 45, p90 84, max 133. Short sessions never see the problem; long sessions see it badly. + +Time-based aging is tempting but wrong at this scale. Turn pacing varies by an order of magnitude — a 10-minute think turn and a 15-second reply both count as "one turn" and should age the same. A time-based window would wobble with pacing; a turn-based window tracks the actual thing we care about: how many rounds of decision-making have passed since this signal was useful. + +Disk retention (30 days) is intentionally time-based because at that horizon time is a fine proxy for "bulk turns" — variance averages out across thousands of turns. Presentation aging operates at a much finer grain where the turn/time distinction matters, so the two mechanisms use different units. + +## Decision (as of 2026-04-12) + +Signals carry a **salience** that starts at 1.0 when they arrive and decays exponentially with turns elapsed since arrival. Below a configurable floor, a signal stops appearing in notifications — but the file stays on disk until the 30-day cleanup removes it. Salience is reset to 1.0 whenever a signal is re-engaged (replied to, referenced). + +The decay is parameterized by a **half-life in turns** — the number of turns after which an un-engaged signal has dropped to 50% salience. Default: 20 turns. Presentation floor: 0.3. + +Re-engagement semantics preserve the "keep the active thread alive" behavior: an ongoing peer conversation keeps its signals hot as long as either side is replying. + +Salience decay is a presentation-layer mechanism. It does not delete, mutate, or hide signals at the storage layer. It only gates what the peer sensor emits into notifications. + +### The inward/outward gate framing + +This ADR introduced the explicit language of **two gates**, and that framing is the one that landed: + +| Gate | Question | Mechanism | ADR-119 role | ADR-121 role | +|---|---|---|---|---| +| Inward | "Should this *new* event fire?" | Refractory period, elevated threshold | Primary concern | n/a | +| Outward | "Should this *already-fired* signal still be shown?" | Salience decay with a floor | n/a | Primary concern | + +ADR-119 and ADR-121 compose: a signal that fires in a hot conversation passes both gates — engagement approves it (inward), and salience is at 1.0 (outward, just arrived). Later, salience falls below the floor and the signal stops appearing; if something re-engages it, salience resets and it reappears. These are independent mechanisms answering independent questions against the same underlying per-subject state. + +## What ADR-123 changed (2026-04-14) + +The inward/outward gate framing is now first-class in `sensor-trait`: + +```rust +impl EngagementState { + pub fn should_fire(&self, current_tick: Tick, magnitude: f64) -> bool { /* inward gate */ } + pub fn current_salience(&self, current_tick: Tick) -> f64 { /* outward gate */ } +} +``` + +Both query methods answer against the same `EngagementState` but consult different portions of the curve — refractory for the inward gate, salience decay for the outward gate. A `Curve::ActionPotential` (attend sensors) has both an inward refractory multiplier and a trivial outward salience (always 1.0 — attend sensors don't use the outward gate yet). A `Curve::Exponential` (ways) has a trivial inward gate (multiplier 1.0) and a meaningful outward salience that decays smoothly. A `Curve::Flat` has an outward step function. Curves with both sides populated — e.g., a combined curve where refractory throttles bursts AND salience fades the fired guidance — are possible but not currently used. + +### The ways side shipped first + +Ways' `session::way_fire_outcome()` calls `EngagementState::current_salience(current_tick)` and returns `FireOutcome::Suppressed` when salience exceeds `REFIRE_FLOOR = 0.5`. Every way declares a `curve:` block in its frontmatter; at fire time the engine loads or creates per-way state at `{session_dir}/way-engagement/{way_id}.json`, queries the outward gate, and records the fire if it's allowed. This is ADR-121's original "outward gate fades, floor gates presentation" model — just with token position as the axis instead of turn count, because ways' host-addressing unit is tokens. + +### The attend side is still deferred + +Attend's peer signals still present indefinitely. The planned implementation when it lands: + +1. Each signal file (or a sidecar) carries an `arrival_turn` field. +2. A live turn counter reads the session transcript to derive `current_turn`. +3. The peer sensor's emit path constructs an `EngagementState::new(Curve::Exponential { half_life: turns_to_half })` per signal, records a fire at `arrival_turn`, and queries `current_salience(current_turn)`. Below the floor, the signal is suppressed from presentation. +4. Re-engagement (reply/reference) bumps `arrival_turn` to the current turn, resetting the outward gate. + +The same engine, the same curve variant, a different tick axis. No code the ways side doesn't already exercise. + +### Why the ways side came first + +Two reasons. First, the ways side had a concrete acute pain point (the 25% global threshold from ADR-104 didn't scale with per-way cadence differences) while attend's presentation aging was a latent issue only visible in very long sessions. Second, the ways side forced the progression-axis argument — because token position is chunky and wall-clock-second burst detection breaks on it, the unification effort needed the curve engine to be unit-agnostic from the start. Attend's signal salience would have been a second instance of the same pattern in the same unit; building it first wouldn't have surfaced the chunky-axis problem. + +## Consequences + +### Positive (preserved, now concrete on the ways side) + +- **Graceful fade, not a cliff.** Exponential decay means no hard "at turn 45 everything drops" behavior. Holds for ways' token-axis implementation today and for attend's turn-axis implementation when it lands. +- **Short sessions unaffected.** A way with `half_life: 30000` tokens never re-fires in a short session that accumulates less than 30k tokens — equivalent to the "first-turn signals are still ~66% salient at median session end" guarantee. +- **Symmetric with ADR-119.** Inward and outward gates composed against the same state. This is now codified in `EngagementState::should_fire` / `current_salience` rather than being a design intention. +- **Turn-based matches mental model** (for the attend application). Still true; still deferred. + +### Positive (added by ADR-123) + +- The inward/outward distinction is now a code contract, not just a design intention. A caller that wants the inward gate calls `should_fire`; a caller that wants the outward gate calls `current_salience`. Both are cheap to query against the same state. +- The curve shape is a first-class parameter. `Curve::ProgressiveStaircase` — which this ADR described as an alternative not-yet-pursued — is now a trivially-available shape for any way that wants declared re-fire deltas instead of smooth exponential fade. + +### Negative + +- **Mechanism complexity.** Unchanged. Per-subject state with curve-driven decay is more machinery than a flat presentation rule. +- **Turn counter dependency** (for the attend application, when it lands). Still deferred, still coupled to Claude Code's JSONL transcript format when implemented. +- **Re-engagement detection is heuristic.** Still deferred. +- **Threshold tuning.** The 20-turn default from the session survey above is still the tentative value for the attend application. `REFIRE_FLOOR = 0.5` is the current ways equivalent. +- **Asymmetric rollout.** Shipping on one tool before the other means the documentation (this ADR, plus `docs/attend-and-monitor/salience.md`) describes a design that's only half-implemented. Readers need to be told which half. + +### Neutral + +- **Orthogonal to disk retention.** The 30-day disk cleanup is unchanged. Salience decay sits entirely above it, in the presentation path — still true. +- **Configurable.** All parameters will live in attend config following the ADR-115 overlay pattern when the attend side lands; all ways parameters live in per-way frontmatter today. + +## Alternatives Considered (2026-04-12) + +### Hard cutoff at N turns + +Reject signals older than turn `current - N`. Simpler to implement (no per-signal salience), but produces a cliff. Rejected. This aligns with ADR-123's choice to make `Curve::Flat` an opt-in rather than a default — step functions are valid first-class shapes when a way genuinely wants them, but not the default. + +### Time-based aging (minutes or hours) + +Use `u2u_median` (~82s) as one-turn-equivalent and compute salience from wall clock. Cheaper because no turn counter needed. Rejected because turn pacing varies too much at this granularity. The ADR-123 progression-axis framing generalizes this: each tool picks its axis based on its host's addressing unit. Wall clock is right for attend's multi-observer peer signals; turns are right for attend's per-signal presentation aging; tokens are right for ways' per-way re-fire gating. + +### Linear decay over N turns + +`salience = max(0, 1 - (turns_since / N))`. Even simpler than exponential. Rejected because linear decay implies "every turn contributes equally to aging," which misrepresents the real curve. Exponential matches "relevance halves repeatedly," which is closer to how attention actually works. ADR-123 codified this by making `Curve::Exponential` the primary outward-gate shape. + +### Per-category half-lives + +Different signal types (peer message vs. build event vs. git change) could age at different rates. Tempting but premature — start with one half-life across all signal types. ADR-123 made this free on the ways side (each way declares its own curve) but preserved the single-global-parameter shape on the attend side, to keep the deferred implementation scope minimal. + +### Subsume into ADR-119's refractory machinery + +Action potential already tracks engagement dynamics. Why not extend the refractory curve to cover presentation aging too? Rejected because the two mechanisms answer different questions. ADR-123 validated this rejection by making `should_fire` and `current_salience` distinct query methods against the same `EngagementState` — two gates, one state, cleanly separated rather than conflated. + +## Implementation Plan + +### Ways side (shipped via ADR-123) + +1. ✅ `EngagementState::current_salience(tick)` query method against `Curve::Exponential` +2. ✅ `REFIRE_FLOOR` constant (default 0.5) in `tools/ways-cli/src/session.rs` +3. ✅ Per-way state persistence as `{session_dir}/way-engagement/{way_id}.json` +4. ✅ `FireOutcome` enum gating way firing on salience floor +5. ✅ Every way declares an explicit `curve:` block in its frontmatter +6. ✅ `ways list` and `ways rethink` visualize per-way re-fire distances from each curve's `refire_delta(REFIRE_FLOOR)` + +### Attend side (still deferred) + +1. Per-signal `arrival_turn` metadata (side table or sidecar — keep wire format stable) +2. Live turn counter parsing the session JSONL +3. Per-signal `EngagementState<Curve::Exponential>` in `sensor-peers` +4. Presentation gate consulting `current_salience(current_turn) >= floor` +5. Re-engagement detection (reply via `re:` field from ADR-120, content-match as best-effort) +6. Config plumbing for `attention.half_life_turns` and `attention.presentation_floor` + +## References + +- **[ADR-123](../ways/ADR-123-firing-dynamics-progression-axis-unification.md)** — progression-axis unification; where the inward/outward framing landed in code. +- **ADR-113** — attend active awareness module; the disclosure governor. +- **ADR-114** — attend as insistent way trigger type; integration for signal handlers. +- **ADR-119** — action potential engagement model; the inward gate. +- `docs/attend-and-monitor/salience.md` — the implementer-facing description of the planned attend application. +- `docs/hooks-and-ways/context-decay.md` — the presentation-economics model that motivates the outward gate on the ways side. diff --git a/docs/architecture/attend/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md b/docs/architecture/attend/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md new file mode 100644 index 00000000..f3d6b63e --- /dev/null +++ b/docs/architecture/attend/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md @@ -0,0 +1,142 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: Claude learns attend's messaging affordances from SKILL.md at session start and may have forgotten them by turn 80, so peer messages go unanswered + - evidence: 'validation before promotion: 57 workspace unit tests pass, and live Monitor runs and a two-instance peer conversation confirmed the disclosure' + - precedent: ADR-104 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-13 +deciders: + - aaronsb + - claude +related: + - ADR-104 + - ADR-113 + - ADR-117 + - ADR-119 +imported: + from: docs/architecture/system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md + format: v0 + status: Accepted +--- + +# ADR-122: Attend disclosure sensor — token-gated affordance reheat + +## Context + +`attend` exposes interactive surfaces Claude must know how to use — `attend send`, `--to`, `--focus`, `--broadcast`, focus groups, mark-read semantics, the "silence is a valid reply" norm, and the load-bearing rule that `attend run` belongs to Monitor rather than Bash. None of this is re-taught at runtime. Claude learns it from `SKILL.md` at session start and may have forgotten the actionable shape by turn 80. When a peer message arrives via Monitor, the notification reports *that* something happened but not *what affordances exist* to act on it. The user pays an inference turn for Claude to either improvise, ask, or — most commonly — silently fail to engage with a channel that was supposed to be two-way. + +This is the same shape of problem ADR-104 solved for ways: instructional context that mattered at session start is not reliably present three reflection windows later. Ways answered with token-gated re-disclosure — track the token position of the last disclosure, fire again when drift crosses a threshold. Attend has the same problem on a different surface (interactive messaging affordances) and the same answer should apply, reusing the ADR-104 model rather than inventing a parallel mechanism. + +The substrate separation principle from ADR-113 constrains the form of the answer. Sensing when Claude has drifted far enough from a teaching to need it again is cheap integer arithmetic — it does not need inference. The teaching itself, the thing that gets emitted into Claude's conversation, pays the token cost every time it fires, so it has to be terse by design. The cheap substrate (a sensor) decides when; the expensive substrate (Claude's next turn) only sees the result when surfacing is warranted. + +The design note at `docs/architecture/attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md` walks through the full design iteration. This ADR captures the decision in the form the rest of the system can cite. + +## Decision + +Attend gains a new workspace crate, `sensor-disclosure`, implementing `sensor_trait::Sensor`. The sensor's signal is **token distance since last disclosure for each registered component**, rather than environmental change. Its shape is otherwise ordinary: poll returns observations, the accumulator aggregates, the engagement and governor machinery from ADR-117 and ADR-119 apply unchanged, and each observation becomes one stdout line wrapped by `emit::emit_batch`. + +**The sensor is a sensor in the ordinary sense of the word.** It is not a new architectural layer, not an emit-pipeline stage, not a framework extension. This framing is load-bearing: no new cross-sensor coupling, no new trait, no pre-emit hook. The disclosure batches alongside peer-message arrivals *emergently* — when both sensors are ready to disclose in the same governor window, the existing batch assembly puts them in one `emit_batch` call and Monitor groups them into one notification via its 200ms batching window. When timing does not align, the disclosure sensor fires standalone at its own tick. Either outcome delivers the teaching to Claude. + +**Poll mechanics.** Each poll shells out once to `ways context --json` — the same integration pattern `sensor-context` already uses — and reads the `tokens_used` field. For each registered component, the sensor walks an in-memory `HashMap<Component, u64>` ledger: + +- **No marker present** (first encounter in this attend process) → emit the component's disclosure text at full magnitude, stamp the baseline. +- **Marker present, distance ≥ 25% of the model's current context window** → emit the disclosure text, re-stamp the new baseline. +- **Otherwise** → no observation for that component this poll. + +The threshold is the same `REDISCLOSE_PCT = 25` value ways uses for its own re-disclosure. One dial governs both; the reheat semantics are symmetric across the two substrates. + +**State is in-process memory only.** The ledger is a field on the sensor struct. It is not persisted to `~/.cache/attend/state/`, not to the `/tmp/.claude-sessions-.../` signal tree, not anywhere. This is a deliberate departure from ways' file-backed markers and from attend's existing `state.rs` checkpointing. Any persistent ledger inherits a "stale state causes havoc" failure mode: a crashed attend leaves a marker file behind, the next `attend run` reads it, and suppresses teaching Claude needs. The only defensible way to guarantee clean-restart semantics is to never write the marker in the first place. Because the ledger data is tiny (one `u64` per component) and non-load-bearing, losing it on restart is the *desired* behavior — a fresh attend process re-teaches every component on its first poll, which is the right response to an unexplained process lifecycle event. + +**Component registry from day one.** Even though `messaging` is the only component at introduction, the sensor is structured as a registry — a `Component` enum with a variant per component, each variant carrying a static ID and the bundled markdown body via `include_str!`. Adding a component (`focus`, `scenes`, `peers-introspection`, etc.) is an enum variant plus a markdown file, not a refactor. Per-component state is keyed in the shared ledger. The registry is also the natural seam for per-component semantics if any component ever needs them — today they all follow the same trigger rule. + +**Content is bundled into the binary.** Each component's disclosure text lives at `sensor-disclosure/src/disclosures/{component}.md`, included via `include_str!`. Same posture as ways body files authored alongside their logic: it ships with the code that uses it, it cannot go missing at runtime, it is versioned in git alongside the sensor that depends on it. The alternative — reading from a path under the skill directory — was rejected because it adds a missing-file failure mode without compensating benefit. + +**Emit format uses the existing `[attend sensor=disclosure priority=high]` prefix.** An earlier draft of the design note committed to matching ways' re-disclosure header pattern for visual-identity reasons. Sensor-native framing made that commitment costly (it would require special-casing `emit::emit_batch` for one sensor, against the grain of how every other attend sensor presents itself). The decision is therefore uniform with attend sensors rather than uniform with ways re-disclosures. The content carries a `reheat:` self-identifying tag on the first line of each disclosure, which is sufficient without header-format rhyme. + +**Content is terse by construction.** Each disclosure block pays its own token cost every time it fires. The messaging block is authored in ways-body density: one-line assertions with terse rationale, no prose expansion. Every line stays under ~200 characters, comfortably inside Monitor's ~400-character stdout line ceiling. A sister fix in the same PR (`sensor-peers` chunking) applies the same constraint to peer messages, ensuring the ceiling is respected end-to-end rather than only on disclosure output. + +## Consequences + +### Positive + +- **First-run disclosure is automatic.** The first poll after `attend run` start finds every component with a missing marker and fires them all. No separate startup-banner injection, no special-case code in `main.rs`. The sensor framework handles it. +- **Reheat semantics match ways.** One dial (`REDISCLOSE_PCT = 25`), one model, one mental frame for "Claude is being reheated on something." Reinforcement rather than two parallel protocols. +- **Zero framework change.** The sensor plugs into the existing slot machinery in `tools/attend/src/sensors/mod.rs` exactly like every other sensor. `SensorSlot`, `AdaptiveInterval`, `DeltaAccumulator`, `EngagementState`, and the disclosure governor all apply unchanged. Validated end-to-end in a live multi-agent session during PR development. +- **Clean restart by construction.** A restarted attend process starts with an empty ledger and re-teaches every component on first poll. No stale-state failure mode. No recovery logic needed. The cost of forgetting is exactly one re-teach, which is the correct response to an unexpected process restart anyway. +- **Batching with peer messages is emergent.** When the disclosure sensor and `sensor-peers` both cross their thresholds in the same governor window, the main loop's batch assembly groups them into one `emit_batch` call, and Monitor groups them into one notification via the 200ms batching window. The "teach-alongside-arrival" behavior we wanted comes for free from the existing scheduling without any coupling code. +- **Extensible.** Adding a second component (`focus`, `scenes`, `peers-introspection`) is an enum variant plus a markdown file. No refactor, no cross-sensor plumbing, no framework changes. + +### Negative + +- **Subprocess call on every poll.** Reading `tokens_used` via `ways context --json` is a process spawn per poll (default 60s base interval, 20s minimum). This is the same cost `sensor-context` already pays against the same command — we are doubling the rate, not inventing it. In aggregate the cost is on the order of 1–2 process spawns per minute during active sessions, which is within the noise floor of the existing sensor loop. +- **Content fidelity depends on author discipline.** The disclosure text is bundled Rust source via `include_str!` and pays tokens every fire. Drift toward prose expansion is a real failure mode. Mitigation: explicit terse-by-construction standard documented in this ADR and in the design note; line-length audit as part of code review. +- **No per-component threshold override yet.** All components share the 25% threshold. If some component ever needs more aggressive reheat (e.g., a critical affordance that Claude forgets faster than token drift predicts), there is no knob. Add only if real usage shows the need. +- **First-run-always-fires is a property.** A mid-session `attend run` restart re-teaches every component on first poll, even if the current working window already contains the teaching from an earlier run. Not a bug — it is the clean-restart guarantee — but it is an observable side effect worth naming: attend restarts cost a re-disclosure per registered component, billable to the user's token budget. + +### Neutral + +- **The `state.rs` persistence model is deliberately not extended.** Context-percentage tracking and git state stay persistent (they have different failure characteristics), but disclosure markers do not. Future components in this sensor inherit the in-memory constraint by default. +- **Design note is retained as companion prose.** The design note at `docs/architecture/attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md` stays as the detailed walkthrough of the four-way iteration that landed on the sensor-native framing. This ADR cites it rather than re-deriving its argument. +- **Governor interaction is unchanged.** The disclosure sensor participates in the existing disclosure governor the same way any other sensor does. No new rate-limiting knobs, no new cooldown windows, no new global budget. + +## Alternatives Considered + +### A persistent ledger under the session signal tree + +Store the disclosure marker as a file at `/tmp/.claude-sessions-{uid}/{session_id}/attend-disclose/{component}/.value`, mirroring ways' own `way-tokens/` layout one directory over. Symmetric with ways, visible to other session-aware tooling, survives attend process restarts within a session. + +Rejected because the "stale file causes havoc" failure mode cannot be defended cheaply. A crashed attend leaves a marker behind; the next `attend run` reads it and suppresses teaching Claude actually needs. The ledger data is tiny and non-load-bearing — losing it on restart is the desired outcome, not a cost. In-process memory is strictly simpler and strictly safer for this specific data. + +### Extend `state.rs` `StateSnapshot` to carry disclosure markers + +The existing `StateSnapshot` already persists `disclosed_thresholds` and `reply_hint_shown` to `~/.cache/attend/state/{session_id}.state`. Adding a `disclosure_ledger` field would reuse the existing checkpoint infrastructure. + +Rejected on the same failure-mode grounds as the persistent-file alternative, plus an aesthetic objection: `StateSnapshot` fields are about context-percentage ways-style thresholds. Piling the disclosure ledger into that struct muddies the semantics and couples two different kinds of state into the same serialization format. + +### Emit-pipeline interceptor / pre-flush hook + +Add a hook in `emit::emit_batch` that inspects the outgoing batch and, if it contains a peer-message event and the messaging ledger threshold is exceeded, prepends a disclosure block before flushing. This was the first implementation direction explored. + +Rejected because it is a bespoke emit-pipeline stage that does not fit the existing sensor framework. Every other awareness capability in attend is a sensor. Making disclosure an emit-layer special case would introduce a new architectural seam, complicate the sensor framework's mental model, and commit us to a second mechanism that runs on a different substrate than the rest. + +### `ReactiveSensor` trait for cross-sensor coupling + +Introduce a new trait alongside `Sensor` for sensors that need to read other sensors' observations this tick before producing their own. Disclosure would be the first implementation, running last in poll order and inspecting a "this-tick's-observations-so-far" buffer exposed by the tick scheduler. + +Rejected as speculative abstraction. A reactive-sensor trait is a meaningful commitment — it changes the sensor contract, the scheduler loop, and the way sensor authors reason about tick ordering. It is not justified by a single use case. Moreover, once the disclosure signal was reframed as "token distance since last disclosure" rather than "reaction to a peer event," the cross-sensor coupling disappeared: the disclosure sensor has its own interoceptive signal (ledger distance) and does not need to observe other sensors at all. The batching-with-peer-messages behavior emerges from the existing scheduling without any coupling. + +### Topic-drift or semantic-drift modeling + +Trigger reheat not on raw token count but on some measure of *semantic* drift — e.g., whether the current conversation topic has shifted significantly from the topic at the last disclosure. Conceptually appealing because token count is an imperfect proxy for "has Claude forgotten the teaching." + +Rejected on the substrate-separation principle from ADR-113 and the discipline named in `docs/attend-and-monitor/authoring-sensors.md`: sensors report facts the framework can verify, not guesses the sensor produced. Token count is measurable from the transcript with no inference cost. Semantic drift is an inference problem the sensor cannot solve cheaply, and if solved by an embedding model the embedding itself introduces new failure modes and a new dependency. "Simulated judgment" is the wrong substrate for the thing this sensor is meant to do. + +### Persistent storage cap with TTL + +Compromise between in-memory and fully persistent: write disclosure markers to disk but expire them automatically after some wall-clock window (e.g., 1 hour). Reduces stale-state risk by ensuring markers do not outlive the process state they describe for long. + +Rejected for four independent reasons: + +1. **TTL does not close the stale-state failure mode.** During the TTL window, a crashed-and-restarted attend still inherits suppression state it cannot explain. The window only narrows the exposure; it does not eliminate it. The in-memory solution eliminates it entirely by construction. +2. **It introduces a file-I/O failure surface the in-memory design avoids entirely.** A disk-backed ledger has to handle read errors, write errors, disk-full conditions, filesystem corruption, concurrent-access races between multiple attend instances, and the inevitable edge cases around partial writes during a crash. The ledger data is a `HashMap<Component, u64>` — three to five entries in the extensible case. Accepting all of that failure surface to persist a handful of integers is a bad trade. +3. **TTL tuning is a configuration surface with no defensible default.** "How long should my disclosure markers live?" is a question with no good answer: short TTLs defeat the purpose (the next `attend run` re-teaches anyway), long TTLs reintroduce the stale-state problem we are trying to avoid, and middle values have no principled basis. The in-memory solution has no such tuning knob — the lifetime of the ledger is exactly the lifetime of the attend process, which is the only defensible scope. +4. **It would be inconsistent with the rest of attend's persistence story.** `state.rs`'s existing `StateSnapshot` does not TTL its checkpoints — they persist indefinitely and are only overwritten on the next clean write. Introducing TTL semantics exclusively for disclosure markers would create a new and inconsistent pattern the rest of the codebase does not share. The simpler path is to keep the existing persistence discipline (opt in deliberately, never TTL) and put the disclosure ledger cleanly outside that discipline entirely. + +## Validation + +This ADR is promoted from a working sketch that has already been exercised end-to-end: + +- All 57 workspace unit tests pass, including 4 in `sensor-disclosure` covering extraction, registry, and formatting. +- The sensor registers in the attend tick loop alongside `context`, `processes`, `git`, and `peers`. +- First-run disclosure fired correctly on sensor startup in three separate runs (initial rebuild, content-tightening rebuild, chunking-fix rebuild), each confirmed via live Monitor notification. +- Peer-conversation validation run between two Claude instances exercised the messaging affordances surface in real time, with the disclosure block arriving cleanly as one batched Monitor notification in both directions. +- The chunking fix in `sensor-peers` was validated by requesting a deliberate ~800-char message from the peer; it arrived as three correctly-numbered chunks with clean word boundaries. + +The design note's "promote to ADR after the model holds up across a few real sessions" threshold is met. diff --git a/docs/architecture/attend/ADR-124-channel-bar-ordering-open-as-base.md b/docs/architecture/attend/ADR-124-channel-bar-ordering-open-as-base.md new file mode 100644 index 00000000..cc683f07 --- /dev/null +++ b/docs/architecture/attend/ADR-124-channel-bar-ordering-open-as-base.md @@ -0,0 +1,313 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: 'observed: stopping the claude in /home/aaron/temp left @Urban in the agent legend indefinitely' + - evidence: 'in use, #open looked like any other group and _broadcast/ and @open/ carried overlapping intent' + - precedent: ADR-118 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-16 +deciders: + - aaronsb + - claude +related: + - ADR-118 + - ADR-120 +imported: + from: docs/architecture/system/ADR-124-channel-bar-ordering-open-as-base.md + format: v0 + status: Accepted +--- + +# ADR-124: TUI Legend Architecture — Base Channel, Liveness, and Ordering + +## Context + +The attend-chat TUI has two legend strips — a **channel bar** across +the top (focus groups) and an **agent legend** above the status row +(claudes and humans seen on the signal bus). Both are identity +surfaces: they're how the human locates the set of peers and channels +currently available to address. + +Three problems have surfaced in use: + +**1. `#open` is semantically special but visually equal.** ADR-118's +scenes define `open` as the well-known "everyone who wants shared +coordination" group. It's the canonical fallback that agents opt into +when they're *not* in a specialized focus. A new user looking at the +channel bar has no cue that `#open` is the place to drop a general +message — it's just another group in an alphabetical row. + +**2. `_broadcast/` and `@open/` are distinct on disk but convey +overlapping intent.** `_broadcast/` is "reach everyone regardless of +group membership"; `@open/` is "the named group people gather in for +shared coordination". In practice they're used the same way, and the +dual affordance confuses both the TUI's mental model and the human +writing a message ("do I `#open ...` or just plain text?"). + +**3. No ordering rule beyond alphabetical.** As groups grow, the +channel bar becomes a noisy alphabetical strip with no affordance for +"active" vs. "idle" or "pinned" vs. "transient". + +**4. The agent legend never forgets.** Its seed source is +`~/.claude/sessions/*.json` — a disk-persistent record that outlives +any given claude process. When a claude exits, its session file +lingers; the agent's nickname keeps appearing in the legend long +after the agent is gone. Observed concretely: stopping the claude in +`/home/aaron/temp` left `@Urban` in the legend indefinitely. The +legend misrepresents system state. + +This ADR resolves all four. It's a single record because they share +one substrate (the two legends) and one framing (what the TUI +surfaces as the current coordination topology), so splitting them +would scatter related consequences across multiple ADRs without +making any of them clearer. + +## Decision + +### 1. `#open` is the base channel — always leftmost + +The channel bar renders with a fixed first position reserved for the +base channel (`#open`). Every other discovered `@group` follows it, +in a defined order (see §3). + +The base channel: + +- Always appears, even when its signal dir is empty +- Never gets a removal affordance (can't `/dissolve #open`) +- Is the implicit target of plain-text messages (no sigil required) +- Is the implicit target of `attend send "msg"` with no flags + +`#open` isn't "just another group" — it's the commons. + +### 2. Resolve `_broadcast/` vs. `@open/`: `#open` becomes the canonical base + +Rename the display layer: what today lives at `~/.cache/attend/signals/_broadcast/` +is rendered and addressed in the TUI as `#open`. This is a UX +rename, not a wire-format change — the directory on disk stays +`_broadcast/` so existing peers continue to receive messages there. +`@open/` the named group is folded into the same logical channel. + +Two options for the folding: + +**Option A — `_broadcast/` is canonical; `@open/` is removed.** The +scene named `open` updates to mean "subscribe to the base channel" +(which is automatic for everyone anyway). Any existing `@open/` dir +is migrated: messages move to `_broadcast/`, the dir is cleaned up +on next peer-sensor poll, and the scene preset is simplified. + +- Pros: one dir, one concept. No ambiguity. +- Cons: breaks any hand-crafted scenes/flows that rely on `@open/` + specifically. Migration script needed. + +**Option B — both dirs stay, rendered as the same `#open` channel.** +The TUI merges signals from `_broadcast/` and `@open/` into a single +`#open` channel view. Writes from `#open ...` land in `_broadcast/` +(the canonical source of truth). `@open/` becomes deprecated but not +removed — signals there are still read, just not written to by new +senders. + +- Pros: no migration. Backwards-compatible. +- Cons: two dirs, one channel — the ambiguity moves from the UX to + the implementation. Drift risk over time. + +**Recommendation: Option A.** Cleaner long-term, and the `@open` +scene preset has been around long enough to carry a one-shot migration +note in release notes. The `_broadcast/` directory name stays (on-disk +convention), but `@open/` goes away. + +### 3. Channel ordering rule + +Left-to-right order: + +1. **`#open` (base)** — always first, never moved, never hidden. +2. **Everything else** — whatever order the underlying scan + returns (today that's alphabetical by group name). + +Originally this section prescribed bands for pinned / recent / +quiet. That ordering rule deferred — pin-order would require a +`_groups.yaml` schema change (pin timestamp), and recency requires +a per-group mtime index the scan doesn't track. The base-channel +leftmost is the one invariant this PR needs; richer ordering can +land in a follow-up once there's pressure for it. + +### 4. Visual treatment + +- Base channel `#open` renders with its hashed glyph + color same as + any group, but with a bold weight so the eye registers its + special role as the commons. +- Pin indicators and unread dots are deferred to a follow-up PR + alongside the ordering rule — no value in half-implementing them + without the band structure they're meant to cue. + +### 5. Agent legend: three-state presence + +The agent legend renders each claude in one of three states: + +| State | Condition | Rendering | +|---|---|---| +| **Live** | session file exists AND PID is alive AND PID is a claude process | full-weight, normal identity color | +| **Declared (nobody home)** | session file exists BUT PID is dead | dimmed, same color/glyph, addressable | +| **Absent** | no session file | not shown | + +**Declared agents stay addressable** — that's the point of dimming +rather than dropping. The signal bus is filesystem-mediated: a +message sent to a declared-but-dead claude writes to its cwd-encoded +inbox dir and sits there. The next time a claude starts in that cwd, +attend's backlog handling picks it up. Dimming signals "nobody's home +right now, but your message won't be lost." + +This mirrors how a physical name card on an empty desk still tells +you where to leave a note. + +Humans are **not** gated this way. Their presence in the legend is +keyed on having emitted a signal that's still in the TUI's buffer +(or, in a future PR, being derived from a human-membership key). +Humans are "ephemeral by message" — a human who hasn't typed in a +while fades via the buffer cap, not a liveness check. + +#### Detection mechanism + +Port the liveness idiom already used by `sensor-peers::pid_is_claude` +into `attend-chat::sessions`. Each render-time `discover()` call +filters session files by liveness before returning them to the +identity registry. + +Concrete implementation: + +- `/proc/<pid>` existence check on Linux (single `stat`, sub- + microsecond). +- Fall back to `ps -p <pid> -o comm` on non-Linux or when `/proc` is + unavailable. +- Match the parent process name against `claude` to exclude PIDs that + happen to be reused by another program after the session file was + written. + +#### Refresh cadence + +Per-render liveness checks over 20+ session files are cheap on Linux +but unnecessary. A 2-second TTL cache is plenty — the human won't +notice a two-second lag between a claude exiting and its chip fading +out, and the sub-second render work stays bounded. + +#### Why not rely on the peer sensor + +`sensor-peers` already does liveness. But it lives in the attend +process, not attend-chat; and ADR-118's data flow is filesystem- +mediated (everything goes through `~/.cache/attend/signals/`). +There's no pre-computed "live peers" file attend-chat can read. We +replicate the detection inline rather than introducing a new shared +state file — the check is small enough that duplication is cheaper +than a new coordination surface. + +### 6. Channel-bar: same three-state rule + +The channel bar applies the same presence logic as §5 at the group +level: + +| State | Condition | Rendering | +|---|---|---| +| **Active** | at least one live member (per §5) | full-weight | +| **Declared (nobody home)** | membership non-empty but zero live members | dimmed, addressable | +| **Empty pinned** | zero members, `pinned: true` | full-weight (the pin overrides quiet) | + +Same reasoning as §5: a declared-but-inactive channel still has its +directory on disk, and a message sent there sits until the next +claude to join picks it up. Dimming, not hiding. + +`#open` (the base) is **never** dimmed or hidden — it's the commons +whether or not anyone is listening. A message to `#open` always has +someone eventually. + +## Consequences + +### Positive + +- New users see `#open` as the obvious commons without reading docs. +- Ordering has a principled rule, not just alpha — the bar + self-organizes around what the human cares about. +- The `_broadcast/@open` dual affordance disappears; there's one + default channel, not two overlapping ones. + +### Negative + +- Migration needed for any existing `@open/` dir (Option A). One- + shot at first run; script is trivial. +- `attend` CLI output (`peers`, `focus all`) needs the same rename + to match the TUI — otherwise the CLI shows `open` as a group and + the TUI shows it as the base, which is worse than the current + state. +- Release note: "`@open` is now the base channel; messages sent + plain-text reach it; the old `@open` group no longer exists." + +### Neutral + +- `_broadcast/` on disk stays. We're renaming the display layer, not + breaking the wire format. + +## Implementation notes + +The work naturally splits into two PRs (post-slash-commands): + +**PR A — Base channel.** §1 through §4. + +1. attend-chat: `#open` pinned leftmost in `group_legend_row`; sort + remaining by (pinned, recent, alpha). +2. attend-chat: `_broadcast/` reads surface as `#open`; plain-text + writes route to `_broadcast/` (already do); `#open ...` writes + also route to `_broadcast/` (new — today they'd route to + `@open/`). +3. attend: `groups.rs` removes special-case handling of "open" as a + scene name; scenes.yaml docs update; one-shot migration of any + lingering `@open/` dir on `attend run` startup. +4. attend: `peers`, `focus all` output uses `#open` as the base + channel label; drop any display of `_broadcast/` as a thing the + user addresses directly. + +**PR B — Liveness.** §5 and §6. + +1. attend-chat: new `liveness` module with cached + `pid_is_claude(pid)` (2-second TTL). Port sensor-peers' detection + pattern. +2. attend-chat: `sessions::discover` reads PID from each session + file and tags entries with a `Live | Declared` presence state + rather than filtering. The registry carries the state through to + render. +3. `KnownIdentity` gains a `presence` field; renderers consult it to + pick full-weight vs. dimmed color. +4. attend-chat: `groups::scan` cross-references each group's + membership list against the live-agent set; groups with zero + live members render dimmed (unless pinned or `#open`). +5. Tests: synthesize a session file with a dead PID, assert legend + renders it as declared/dimmed (not dropped); assert routing to + `@DeclaredName body` still writes to disk successfully. + +## Open questions + +Settled during PR A: + +- `#open` is display-only — `_broadcast/` stays as the on-disk name. +- `open` is reserved in `validate_group_name` so nobody can + create an `@open/` group that would shadow the base. + +Still open (for PR B and beyond): + +- Do we need a `/unpin` slash command pair to `/pin`, or does + `attend focus unpin <name>` suffice? (Slash commands PR will + decide.) +- Liveness cache TTL — is 2 seconds the right number? Too short and + we burn syscalls on a busy signal stream; too long and a stopped + agent lingers visibly. 2s is a guess based on "humans don't + notice", but we should measure and tune after PR B lands. +- Should `attend peers` CLI output apply the same liveness filter as + the TUI? If so, the detection logic wants to live in a shared + crate (or in agent-identity), not duplicated. Leaving undecided + until the PR B implementation — two duplicates is tolerable, + three is extract-pressure. +- Richer channel ordering (pinned band, recent-activity band) — + deferred per §3, revisit when there's concrete pressure. diff --git a/docs/architecture/attend/ADR-129-instance-suffix-and-heartbeat-liveness.md b/docs/architecture/attend/ADR-129-instance-suffix-and-heartbeat-liveness.md new file mode 100644 index 00000000..d51f39be --- /dev/null +++ b/docs/architecture/attend/ADR-129-instance-suffix-and-heartbeat-liveness.md @@ -0,0 +1,116 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: the user reproduced same-cwd name collision by launching claude twice in one directory + - evidence: session_alive() in tools/attend/src/groups.rs:347 is a stub returning true, so ghost agents accumulate in _groups.yaml and the chat legend + - precedent: ADR-124 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-28 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-118 + - ADR-120 + - ADR-124 +imported: + from: docs/architecture/system/ADR-129-instance-suffix-and-heartbeat-liveness.md + format: v0 + status: Accepted +--- + +# ADR-129: Instance suffix and heartbeat liveness for attend identity + +## Context + +attend's identity and liveness systems each have a corner case that has surfaced in real use. The two cases are independent in cause but adjacent in effect — both produce a wrong picture of "who is here" — and they should be addressed in one decision so the design boundary between them is explicit. + +**1. Same-cwd identity collision.** `agent_identity::Identity::for_cwd(cwd)` (`tools/agent-identity/src/identity.rs`) is a pure function of the cwd path: an FNV-1a hash of the canonical path indexes into the nickname pool. The original design comment justifies this — "a claude restarting in the same cwd keeps the same name" — and it does that job correctly when there is at most one session per cwd. It does not cover concurrent invocations: two `claude` processes started in the same directory hash to the same index and present as the same `Jovan (myproj)`, with the same color and style. The user has reproduced this by launching claude twice in the same directory; both render identically in `attend peers`, in chat, and in any signal authored by either side. + +**2. Stale registration and ghost agents.** `tools/attend/src/groups.rs:347` defines `session_alive(sid)` as a stub that returns `true` unconditionally with a TODO comment. As a consequence, `_groups.yaml` accumulates dead members across sessions and is never cleaned by the liveness check that the rest of the codebase assumes is doing the job. The visible failure surfaces in attend-chat: when the chat TUI reloads, the chip legend (`tools/attend-chat/src/chip.rs:142`, `known_identities`) builds entries from buffered signals and does not filter by liveness — every agent that has ever sent a message in this cwd appears in the legend, including agents whose claude session has long exited. The existing liveness path through `~/.claude/sessions/*.json` plus `pid_is_claude` catches "claude exited" but not "claude alive, attend not running" — the agent is unreachable, but the PID check passes because the claude process itself still exists. + +The two problems are not the same problem. Collision is about *naming*: two real, alive sessions need distinguishable identities. Liveness is about *display*: the registry of who-has-spoken needs to be filtered against who-is-still-here. Conflating them — for example, by freeing a name when its session goes stale — would make names re-bind unpredictably and break peer references like `@Jovan-alpha` after a transient quiet period. Keeping them separated is load-bearing. + +Both surfaces matter today because attend, chat, and the focus-groups model (ADR-118) increasingly assume legible peer identity: human steering decisions, peer-to-peer addressing, and the attend-chat legend (ADR-124) all rely on the displayed name being unambiguous and the displayed roster being live. + +## Decision + +Adopt two independent mechanisms, one per concern. + +**Independence invariant:** liveness display never touches name allocation. A heartbeat going stale does not free a registry slot. A registry slot expiring (under age-based GC) does not signal liveness. This separation is what keeps names stable across transient quiet periods — peer references like `@Jovan-alpha` survive a 30-second pause in the heartbeat without re-binding to someone else. + +### Instance registry — naming + +A new persistent file per cwd records which session holds which instance discriminator. + +- **Location:** `~/.cache/attend/instances/<encoded-cwd>.yaml`, parallel to the existing `~/.cache/attend/signals/<encoded-cwd>/` layout. +- **Schema:** `session_id → { instance: <string>, registered_at: <iso8601>, last_seen: <iso8601> }`. The field is named `instance`, not `letter`. The file holds *instance assignments*; Greek letters are the current discriminator vocabulary, but the storage shape does not lock that in. Future allocators can swap their vocabulary without changing the on-disk format. +- **Allocator (current):** Greek letters in ASCII spelling — `alpha, beta, gamma, ..., omega` (24 slots). ASCII spelling, not glyphs, matches the existing constraint at `tools/agent-identity/src/names.rs:5` (ASCII only, no diacritics) so that `@`-completion remains keyboard-portable. Numeric fallback past 24: `Jovan-25`, `Jovan-26`, and so on. +- **Slot semantics:** slots are session-bound. Once a `session_id` holds `alpha`, it holds `alpha` for the life of that session — no reclamation while the session is live. New sessions skip taken slots and take the next-free letter. Resume always reclaims the original assignment. This is the property that makes peer references stable. +- **Concurrency:** read-modify-write under `flock(LOCK_EX)` on a sentinel `<encoded-cwd>.yaml.lock` file, *not* on the data file. Locking the data file does not serialize concurrent registers because flock state is keyed on the open-file-description — the kernel inode — not the path. The data file is renamed atomically (`.tmp` → `.yaml`) on every commit, so a fresh opener of the path after the rename gets a different inode and a different lock; the previous holder's lock no longer contends. The sentinel never moves, so its inode is the stable serializer. Two simultaneous fresh sessions racing for `alpha`: the lock orders them; the second writer reads after the first commits and takes `beta`. +- **Age-based GC:** entries with `last_seen` older than 7 days are reclaimable. This handles unbounded registry growth on long-lived projects. Trade-off: a resume after more than 7 days of inactivity may receive a different letter, because the slot may already have been reclaimed by another session. This is acceptable — a week-stale resume is far from the "I just stepped away and came back" case where rename surprise actually matters. +- **Render:** always emit the suffix, even when the session is solo. `Jovan-alpha`, never the conditional `Jovan` (solo) / `Jovan-alpha` (collision). Always-on is chosen over conditional for predictable pattern matching: every render site, every grep, every agent-self-reference produces a name with the same shape. +- **Render sites that must consume the registry:** + - `tools/attend/src/identity_view.rs::render_sender_label` + - `tools/attend/src/cmd/peers.rs` (agent column) + - `tools/attend-chat/src/chip.rs` (chip rendering, `known_identities`, and `resolve_nickname`) + +### Heartbeat sidecar — liveness display + +A separate per-session file records that the session's attend is currently running. + +- **Location:** `~/.cache/attend/heartbeat/<session-id>`, one file per session. +- **Encoding:** mtime *is* `last_seen`. There is no body to parse. Each tick is a `touch`. +- **Touched** per attend tick, in the existing tick loop in `tools/attend/src/cmd/run.rs`. +- **Liveness predicate:** `now - mtime < grace`, with `grace = 90s` — three times the base sensor interval of 30s. Long enough to ride out a paused tick or a slow filesystem; short enough to drop a dead attend within a couple of minutes. +- **Consumers:** + - `tools/attend/src/groups.rs:347 session_alive(sid)`: replace the `return true` stub with `pid_is_claude(sid_owner) AND heartbeat_fresh(sid)`. This is the predicate the rest of the codebase already calls; fixing it propagates correctness through `_groups.yaml` cleanup and elsewhere. It also catches "claude alive, attend not running," which the PID-only check misses. + - `tools/attend-chat/src/chip.rs::known_identities`: filter signal-derived identities by liveness so dead agents disappear from the chip legend. Historical messages stay in the buffer (the conversation record is not destructive), but their senders are rendered dimmed or muted when the sender's heartbeat is stale, so the buffer reads correctly without implying the speaker is still here. +- **No write contention:** each session writes only its own file. + +### Operational + +A new `make purge-attend-state` target wipes `~/.cache/attend/` for clean-base recovery. It is never invoked automatically. After the existing `make attend`, `make attend-rebuild`, `make attend-chat`, and `make attend-chat-rebuild` targets, an advisory hint is printed pointing at the purge target, so users with an inconsistent cache from an older binary know how to reset cleanly. + +## Consequences + +### Positive + +- **Names disambiguate same-cwd peers.** Two simultaneous claude sessions in `~/Projects/myproj` render as `Jovan-alpha` and `Jovan-beta`, with stable colors and styles inherited from the base nickname. +- **Peer references survive quiet periods.** Because liveness staleness never frees a registry slot, `@Jovan-alpha` continues to mean the same session even if its attend tick paused briefly. Agent self-model is preserved across resumes. +- **Ghost agents disappear from the chat legend.** With the heartbeat predicate replacing the `return true` stub, `known_identities` no longer accumulates every agent that ever spoke in the cwd. The chip legend reflects who is currently reachable. +- **`_groups.yaml` self-cleans.** The same predicate, fixed in one place, flows through to focus-group cleanup. Stale members drop on the next peer poll, restoring the cleanup behavior ADR-118 already assumed was working. +- **Catches "claude alive, attend not running."** PID-only liveness misses the case where the claude process exists but its attend has exited or hung; the heartbeat catches it because the heartbeat file stops being touched regardless of process state. +- **Filesystem-as-state continuity.** Both mechanisms fit attend's existing storage model — flat files under `~/.cache/attend/` — so they are debuggable with `ls` and `cat`, and require no new dependency. +- **Discriminator vocabulary is swappable.** Storing `instance: alpha` (rather than `letter: alpha`) lets a future allocator change vocabulary without a file-format migration. + +### Negative + +- **Identity contract change.** Every existing display surface for nicknames now carries an instance suffix. `Jovan` becomes `Jovan-alpha` even for a solo session. Render sites listed under Decision (`identity_view.rs`, `cmd/peers.rs`, `chip.rs`) must be updated together — partial rollout will produce inconsistent legends. Agents will also see their own displayed name change, which they encounter in self-reference (`I am Jovan-alpha`) and in addressed peer references. +- **Two new state shapes.** `~/.cache/attend/instances/` and `~/.cache/attend/heartbeat/` are added to attend's filesystem footprint. Recovery from corrupt state is the new `make purge-attend-state` target; absent that, debugging "why is my name wrong" requires inspecting two locations. +- **Resume-after-7-days renames.** Age-based GC means a resume after a week of inactivity may receive a different instance letter than it had previously. Mitigation: the rename is far enough from active use that the surprise is small, and 7 days is a tunable upper bound on registry growth. Users with very long lived projects who still want stable names across very long pauses can lengthen the GC threshold. +- **Heartbeat I/O cost.** Each attend tick writes one file. The cost is negligible per session, but on a multi-session host it adds up to N files touched per tick interval. Mitigated by the 30s base sensor cadence and the per-session file scoping. +- **Greek-letter pool exhaustion.** 24 concurrent sessions in one cwd is unreachable in practice today, but the numeric fallback (`Jovan-25`) is intentionally ugly so that hitting it surfaces as a signal that the cwd is overloaded. + +### Neutral + +- **Heartbeat staleness does not free a name.** This is the explicit design boundary, not an emergent property. A session that is registered but unreachable continues to hold its instance letter until its registry entry GCs by age. +- **Heartbeat is mtime-only.** No file body, no parser. Adding fields to the heartbeat in the future would require a real format; today there is intentionally none. +- **Always-on suffix.** Solo sessions display `Jovan-alpha` rather than `Jovan`. The user-visible cost is a slightly longer name; the gain is that every render site can build the label from a single rule. + +## Alternatives Considered + +- **Sequential numbering — `Jovan2`, `Jovan3`.** Rejected. Order-dependent on registration race; observers would not agree on numbering without a registry; and it breaks "same name across restart" the moment a peer joins, because the restarting session would re-number. +- **Hex-suffix from session_id — `Jovan-7c4a`.** Stable per session, zero coordination needed, no registry. Rejected because the suffix is unmemorable for both humans and agents in conversation. "Tell `-7c4a` to do X" is a worse UX than "tell `beta` to do X," and the value of the suffix is precisely that it surfaces in human and agent speech. +- **Resumer yields on conflict.** Initial proposal: when a session resumes and finds its old letter held by a live session, it takes the next-free letter. Rejected. Name changes during resume break agent self-model and break references made by humans and peers (`@Jovan-alpha please...`). The current decision — slots are session-bound for the life of the session, no reclamation while live — is the property that makes references durable. +- **Stateless render-time computation.** Sort same-cwd live sessions by mtime and assign letters in order at render time. Simpler, no registry, no GC. Rejected because letters drift when peers exit: the surviving sessions get re-lettered in place, which has the same self-model and reference-breaking problem as resumer-yields. +- **Conditional suffix — only on collision.** Cleaner UX for solo sessions (plain `Jovan`) but every render site must consult the live-set count to know whether to append. Rejected in favor of always-on for predictable pattern matching: one rule, one shape, every site. +- **SQLite for state.** Considered seriously. Three new state shapes, CAS semantics, and a growing YAML parser are real costs. Rejected because: the data is genuinely tiny (handfuls of rows per cwd); attend's whole architecture is filesystem-as-state already (signals, session.json, `_groups.yaml`, sensor checkpoints — all flat files); adding SQLite for *part* of state creates two consistency models inside attend; mtime-as-heartbeat is a perfect Unix-y fit that dies inside a database; and debuggability via `cat` and `ls` is load-bearing for users diagnosing identity problems. SQLite is reconsidered later if attend grows a real query surface — inbox search, threading analytics, cross-session aggregation — that flat files no longer serve. +- **Field on `_groups.yaml` for instance assignments.** Rejected. `_groups.yaml` is scoped to *named* focus groups (ADR-118). Mixing implicit per-cwd identity assignments into it muddies the contract: that file is about named-group membership, not about identity. Keeping instance assignments in their own file preserves the separation. diff --git a/docs/architecture/attend/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md b/docs/architecture/attend/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md new file mode 100644 index 00000000..36142b7c --- /dev/null +++ b/docs/architecture/attend/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md @@ -0,0 +1,476 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: 'live report: a #open message reached one peer but not another; the retention sweep deleted the stale signal file from the shared dir during the recipient''s Monitor-down gap' + - evidence: '@multi addressed only the first agent: parse_addressed returns one Addressed (fixed in PR #137)' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-20 +deciders: + - aaronsb + - claude +related: + - '[[ADR-120]]' + - '[[ADR-121]]' + - '[[ADR-123]]' + - '[[ADR-124]]' + - '[[ADR-129]]' +imported: + from: docs/architecture/system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md + format: v0 + status: Accepted +--- + +# ADR-136: Split addressed messaging from the sensor-observation bus + +> **Implemented and live-validated 2026-06-20** on branch +> `fix/attend-message-durability`. The decisions below are reconciled to +> what was actually built — see [Validation](#validation). Multi-recipient +> addressing (Bug 1) shipped separately in PR #137. + +## Context + +Attend carries two kinds of traffic over one pipeline, and they have +opposite handling requirements. + +### The office analogy + +Picture several office workers near each other. They **talk back and +forth** — conversations with continuity, where each knows roughly where +the exchange stands. That conversation is durable *state*; it must not +get wiped. + +Meanwhile the **phone rings, a fax comes in, a work package is dropped +off.** These are events. A worker mid-conversation can queue them or +ignore them, and either choice is fine — handling an interrupt, or +choosing not to, never erases the conversation they're holding. The fax +sitting in the tray and the package on the desk are **durable work +items**: they wait there until that worker processes them. Nobody empties +the tray on a timer, and no co-worker walking past gets to shred an +unread fax. + +Mapped to attend: + +- **Ambient observations** — git churn, a peer appearing, a process + starting. The phone ringing. These are *noise* / interrupts. The whole + point of the salience gate ([[ADR-121]]), the action-potential + refractory ([[ADR-123]]), and the disclosure governor is to *suppress* + most of them so Monitor only wakes a session for something that moved. + Queue-or-ignore, lossy, and rate-limited is exactly right here. +- **Addressed messages** — `attend send`, `attend reply`, and the + attend-chat `@Name` / `#group` surface. The fax in the tray, the + package on the desk, the conversation itself. These are *intentional + communication* a human or agent composed on purpose. The work item + waits in **that recipient's own tray** until they process it; the + conversation state it belongs to is never wiped by event handling. + +Today both ride the same path: a `.signal` file in a **shared** +directory, scanned by `sensor-peers` on each poll, then run through the +same gate stack built to throttle observations. Two concrete defects +surfaced, and both are symptoms of messages being second-class citizens +on a bus designed for disposable observations — work items thrown in a +shared tray that anyone can shred. + +### Two clocks — and the bridge between them + +attend and the ways/steering system run on **different time dimensions**, +and keeping them straight is what scopes this ADR. + +- **attend is wall-clock.** It is deliberately a real-time system: it + runs *between* an agent's turns, polls on a wall-clock cadence, decays + salience in seconds, and measures "away for 21 minutes." Real time is + its native substrate. Everything in this ADR — both lanes — lives here. +- **Ways / agent steering is epoch/turn-driven.** Ways fire and refire on + conversation *turns*, match per-prompt, and disclosure-gate per-turn. + There is no wall-clock in that dimension; a session paused for hours and + resumed is the same turn-continuum. The salience engine is *shared* + between the two ([[ADR-123]]), but consumed on two different clocks — + `salience.rs` already names the "deliberate asymmetry vs. ways' + `way_fire_outcome`, which always fires on first match regardless of + age," precisely because ways live on the turn clock and sensors on the + wall clock. + +**The notification is the bridge.** A wall-clock event crosses into the +turn dimension when Monitor injects a `<task-notification>` that becomes +a prompt. That crossing is why **digest-not-replay** (below) is not just +ergonomics but *dimensionally correct*: a wall-clock burst of N messages +must not become N turn-injections — it must coalesce into **one** turn, +because the receiving dimension is turn-discrete and context-precious. +Many-in-wall-clock → one-in-turns is the correct impedance match. The +turn dimension itself is out of scope here; it only meets attend at this +boundary. + +### The current flow + +```mermaid +sequenceDiagram + autonumber + participant H as Human / Agent + participant C as Compose<br/>(attend-chat / attend send) + participant FS as Shared signal dir<br/>(~/.cache/attend/signals/…) + participant S as sensor-peers<br/>(per-session, ~30s poll) + participant G as Gate stack + participant M as Monitor → session + + rect rgba(124,58,237,0.12) + H->>C: "@Cleo @Tam hi" / attend send + Note over C: parse_addressed → ONE Addressed<br/>CLI → ONE --to / --focus + C->>FS: write_signal (one file per dest dir) + end + rect rgba(217,119,6,0.12) + loop each peer polls independently + S->>FS: scan own-cwd + _broadcast + focus + S->>G: 1. seen-dedup + S->>G: 2. salience gate (seeded from file MTIME) + S->>G: 3. retention cleanup DELETES stale file (shared dir!) + S->>G: 4. emission_threshold 2.0 + S->>G: 5. action-potential refractory (per-session state) + S->>G: 6. disclosure governor: window cap + cooldown (per-session) + S->>G: 7. priority filter: magnitude ≥ 3 → stdout + G-->>M: survivors → stdout line + end + end + M->>H: <task-notification> +``` + +Seven gates sit between "send" and "notify." Three of them (salience, +refractory, governor) carry **per-session timing state**, so two peers +watching the same message can legitimately diverge. One of them +(retention cleanup) is **destructive on a shared resource**. + +### Bug 1 — `@multi` addresses only the first agent + +The bus is single-recipient end to end. +`parse_addressed` (`tools/attend-chat/src/legend.rs:173`) parses only +the *leading* sigil token and returns one `Addressed`; +`handle_enter` (`tools/attend-chat/src/app/keys.rs:78`) matches one arm +and writes to one inbox — so `@Cleo @Tam msg` routes to Cleo and +`@Tam` becomes body text. The CLI mirrors this: `attend send` accepts a +single `--to` or single `--focus` +(`tools/attend/src/cmd/send.rs`). The *write* layer already loops over a +`Vec<PathBuf>` of destinations (`send.rs:151`) — it is only ever handed +one element. Multi-recipient was never representable above the write +layer. + +### Bug 2 — a `#open` message reached one peer but not another + +Reported live: a human sent `#open`; the KGS peer (Clio) surfaced it, +the agent in `~/.claude` did not — only a few minutes apart. Timeline +reconstruction from the session transcript confirmed the agent's +Monitor had **down-gaps** (stopped 06:04:42Z, and again 06:35:53Z) while +the peer ran continuously. The decisive mechanism, consistent with a +few-minutes gap: + +1. The message was written during the recipient's Monitor-down gap. +2. The live peer scanned it within ~30s and was notified. +3. The retention cleanup sweep — run by *any* live attend, including the + peer (`tools/sensor-peers/src/lib.rs:394`, and the periodic sweep in + `tools/attend/src/cmd/run/tick.rs`) — **deleted the stale file from + the shared dir**. +4. The recipient's Monitor came back *after* the file was gone, so it + never saw the message. + +The root defect is not "the Monitor was off" — sessions stop and start +routinely. It is that **delivery is best-effort against an ephemeral +shared directory with no durable, replayable, per-recipient inbox.** A +recipient that is momentarily away loses the message permanently. The +52-minute salience-decay backlog filter ([[ADR-121]]) was an early +suspect and is *not* the cause here; the message was fresh. + +## Decision + +Split the bus into two lanes with different delivery contracts. The +sensor-observation lane is unchanged. Addressed messages move to a +**message lane** that is reliable rather than throttled. + +**The lane boundary is _authored communication_ vs _environmental +event_ — not directed vs broadcast.** A person or agent composing words +to reach someone is conversation, and rides the message lane. The +environment generating a notice — git changed, a peer appeared, a +process started; the phone, the fax, the package — is an event, and +rides the observation lane. This matters most for `#open`: it is +authored, so it is durable, full stop. `#open` is "talking loud in the +open office" — sometimes to convene ("we should all discuss this, maybe +break into smaller groups"), sometimes just because only a couple of +sessions are around and `#open` *is* the conversation. Both are talking. +Demoting `#open` to a best-effort bulletin would shred conversation in +exactly the small-room case where it carries the real exchange. +Convening-then-splitting rides the existing focus-group mechanism +([[ADR-118]], [[ADR-129]]); the migration from `#open` to a `#group` is a +workflow, not a durability tier. + +**1. The message lane bypasses the noise-control stack.** Addressed +messages (directed `@`, group `#`, and `#open` broadcast) skip the +event-lane gates that exist to suppress ambient observation *noise*; a +composed message is not noise. As built, this is **three distinct +mechanisms**, and live peer testing found them one at a time — each fix +exposed the next: + +- the per-signal **salience gate** (an mtime-seeded backlog filter) is + *removed* from the message path, so unread mail is never aged-out — the + "peers don't hear" half of Bug 2; +- the per-sensor **action-potential refractory** is *bypassed* for the + message lane (it surfaces whenever anything is accumulated and records + no engagement, so it can never build a refractory that holds + conversation); +- the shared **disclosure governor** is replaced with a *separate + permissive* governor for the message lane (flat 3 s cooldown, generous + window, no rate-ballooning). This is permissive, **not** a full bypass: + normal cadence flows, a true rapid burst coalesces into one digest + (Decision 5) and discloses promptly, and nothing is ever starved or + dropped. + +The lane keeps only **dedup** (deliver once). The neuron-decay model +(`sensor_trait` engagement/curve) stays fully intact for the *event* +lane — git, process, and a future external-chat sensor (e.g. Slack). + +**2. Each recipient owns its own seen-set; rooms stay shared.** A +directed multi-`@` message fans out one durable write per recipient +(all-or-nothing resolution + dedup — Bug 1, PR #137). Shared rooms +(`#open`'s `_broadcast`, focus groups) remain *shared directories* rather +than copying a broadcast into every inbox. Durability there comes from +three things together: nothing is reaped by age (Decision 3), each +session's own **persisted seen-set** (checkpointed) dedups across +restarts, and a **cold-start backlog baseline** keeps a fresh join from +dumping history. So the realized form of "each recipient has their own +tray" is *each recipient owns its own seen-set* — lighter than a +cross-party ack protocol, and lighter than fanning every broadcast out N +times. A restarting session restores its seen-set and surfaces only the +messages that arrived during the gap. + +**3. Nothing is reaped by age; lifetime is bound to project liveness.** +The destructive in-read 5-minute shred that caused Bug 2 — one peer's +scan deleting a file another peer (or a returning session) hadn't read — +is removed outright. Messages are *never* removed by wall-clock age. +Instead, lifetime mirrors Claude Code's own model: a tray dies when its +project is gone. `run_cleanup` reaps a directed tray when its project is +no longer tracked in `~/.claude/projects/`, and reaps a shared-room +signal when its *sender's* project is gone (sender cwd read from the wire +format). The age-based machinery (`--older-than`, duration parsing, +`cleanup.retention`) was removed; `--dry-run` / `--all` remain. +Conversation state (a session's seen-set and thread context, [[ADR-120]]) +is never wiped as a side effect of handling an event. + +**4. Addressing is multi-recipient.** `parse_addressed` becomes +"parse the leading run of `@`/`#` tokens" and returns a *set* of +targets; `handle_enter` and `attend send` build a multi-element +`dest_dirs` and fan out one durable write per recipient. This is the +direct fix for Bug 1 and falls out naturally once messages are +first-class — the write layer already loops. + +**5. Re-entry is a count-led digest, not a replay.** A single poll that +surfaces more than a small cap (8) of unseen messages coalesces into one +digest line — *"12 new messages: 3 to you, 9 on #open (newest 2m ago, +over 21m) — attend inbox for detail"* — instead of flooding the turn. +One mechanism covers both a warm-rejoin gap and a live burst from a +hyperactive peer: many-in-wall-clock → one-in-turns at the notification +bridge. Detail pulls from `attend inbox` +(`tools/attend/src/cmd/inbox.rs`), now **paged** (`--limit` / `--page` / +`--before <ts>`) over the never-reaped, chronological ledger. A live +finding: a real peer rarely triggers the digest end-to-end, because its +*own* auto-mode classifier spaces rapid sends into a 1–2-per-poll trickle +(under the cap) — which is the correct "spread-out = individual, prompt" +behavior. So the digest is the backstop for genuine bursts (a workflow, a +long down-gap rejoin); that path is covered by unit tests. + +**Wall-clock is first-class in both lanes; the lanes differ only in how +they _use_ it.** The event lane uses time to **decay and drop** (a stale +observation is less worth a wake-up — correct lossiness). The message +lane uses time to **stamp and digest** (a stale message is never dropped, +only summarized as "how long ago — you decide"). Same clock, opposite +policy. That is the clean line between the lanes. + +The exact on-disk shape (extend the existing signal-dir convention vs. a +dedicated message store) is an implementation choice for the follow-up +PRs; the contract above is what this ADR fixes. CLI remains the whole +interface ([[ADR-124]] / attend's CLI-is-the-contract rule) — no new +caller reaches into attend-owned state. + +### Target flows + +**Classification — what picks the lane** (authored vs environmental, the +one decision that routes everything): + +```mermaid +flowchart TD + X[New thing happens] --> Q{Authored by a person/agent<br/>to communicate?} + Q -->|"yes — @Name, #group, #open"| M[MESSAGE lane<br/>durable · dedup-only · wall-clock stamped] + Q -->|"no — git, process, peer-presence"| E[EVENT lane<br/>salience + refractory + governor] + M --> MT[recipient tray / room ledger<br/>never wiped until that recipient saw it] + E --> ET[coalesced · aged by wall-clock · may drop] + + classDef external fill:#f6821f,color:#1a1a1a,stroke:#4a5568 + classDef decision fill:#fbbf24,color:#1a1a1a,stroke:#4a5568 + classDef core fill:#7c3aed,color:#ffffff,stroke:#4a5568 + classDef process fill:#2d7d9a,color:#ffffff,stroke:#4a5568 + classDef store fill:#2d8e5e,color:#ffffff,stroke:#4a5568 + + class X external + class Q decision + class M core + class E process + class MT store + class ET process +``` + +**Authored message, live recipients** (also the Bug 1 multi-recipient fix): + +```mermaid +sequenceDiagram + autonumber + participant A as Author + participant L as Message lane + participant T1 as Alice's tray + participant T2 as Bob's tray / #open ledger + participant Rx as Live recipient + rect rgba(124,58,237,0.12) + A->>L: "@Alice @Bob ship it" (one compose) + Note over L: parse the SET {Alice, Bob}<br/>stamp wall-clock ts, one dedup id + L->>T1: durable write + L->>T2: durable write + end + rect rgba(45,142,94,0.12) + Rx->>T1: scan (live) + T1-->>Rx: notify once + Note over T1: read ≠ delete — marked seen in<br/>Rx's own set, survives restart + end +``` + +**Re-entry after a down-gap** (the Bug 2 fix; wall-clock front and center): + +```mermaid +sequenceDiagram + autonumber + participant Rx as Returning session + participant T as Trays + #open ledger + Note over Rx: was down 06:04–06:25 + rect rgba(217,119,6,0.12) + Rx->>T: on reconnect, scan UNSEEN (own seen-set) + T-->>Rx: digest — "while away: 2 to you (newest 2m ago) ·<br/>6 on #open over 21m" + Note over Rx: one coalesced turn, NOT a replay flood + end + rect rgba(45,142,94,0.12) + Rx->>T: attend inbox (opt-in pull) + T-->>Rx: full chronological ledger + Note over Rx: silence still valid — glance if worth it + end +``` + +**Environmental event** (unchanged — the lane the gate stack was built for): + +```mermaid +sequenceDiagram + autonumber + participant Env as Environment + participant S as Sensor poll + participant G as Salience + Refractory + Governor + participant Rx as Session + rect rgba(45,125,154,0.12) + Env->>S: git dirty / process up / peer appeared + S->>G: magnitude + wall-clock age + Note over G: aged, coalesced, most suppressed to stderr + end + rect rgba(217,119,6,0.12) + G-->>Rx: only if loud enough + Note over Rx: lossy BY DESIGN — the phone may ring unanswered + end +``` + +## Validation + +**Live-validated 2026-06-20** with two real sessions: a receiver +("Thaddeus", this repo) on the rebuilt binary and a peer ("Urban-beta", +`~/temp`). The test compared wall-clock send time against the message +path and drove a 12-message burst. Confirmed: + +- **Single send/reply lane** clean both directions, ~36 s round-trip + (≈ one sensor poll each way); ACKs surfaced reliably. +- **Permissive governor** discloses at a flat ~3 s cooldown — no more + multi-minute starvation (pre-fix, a coalesced digest sat unshown for + minutes behind the rate-ballooning event governor). +- **Refractory bypass** keeps the peers lane flowing after heavy prior + traffic (pre-fix it hit `ABSOLUTE REFRACTORY` and held the digest). +- **Cold-start / warm-restart** surfaced no flood — the existing backlog + stayed baselined. +- **Emergent safety**: a peer cannot be instructed by *another* peer to + flood `#open` — the auto-mode classifier keys on *who* asks + (operator-authorized bursts allowed, peer-instructed bursts refused). + +The test is what surfaced the three-gate stack in Decision 1 — each fix +exposed the next gate, something unit tests alone would not have caught. +It also showed the digest's end-to-end path is hard to trigger from a +real peer (its own classifier spaces sends), so that path leans on unit +tests while the live lane behavior is exercised directly. + +## Consequences + +### Positive + +- Messages a human or agent composed on purpose are delivered reliably + and exactly once, including across a recipient's brief restart. +- Both reported bugs are fixed by construction: multi-recipient + addressing (Bug 1) and no-silent-drop delivery (Bug 2). +- The salience / refractory / governor machinery gets a clearer mandate + — it governs *observations*, the job it was designed for — instead of + being asked to also not-lose intentional messages, which it was never + built to guarantee. + +### Negative + +- Two lanes is more surface than one pipeline. The win is that each lane + has a single, honest contract; the cost is that "it's all just + signals" stops being true. +- Never reaping by age means the on-disk ledger grows until a project is + removed. This is bounded by **project-liveness** reaping (a tray's, or a + sender's, signals go when its `~/.claude/projects/` entry does) rather + than a wall-clock timer. The constraint that actually matters is + *context* — how much gets rehydrated, capped by the count-led digest and + inbox paging — not disk. + +### Neutral + +- The three synchronized messaging docs must move in lockstep with any + contract change: `skills/attend/SKILL.md`, + `tools/sensor-disclosure/src/disclosures/messaging.md`, and + `hooks/ways/softwaredev/environment/attend/attend.md`. +- `#open` broadcast keeps its semantics ([[ADR-124]] base channel); only + its delivery guarantee changes from best-effort to durable. +- Threaded replies ([[ADR-120]]) ride the message lane unchanged. +- **Known limitation (follow-up).** The lane is selected per *sensor* + (`peers`), but that sensor also emits peer-presence *events*, which + therefore currently ride the message lane — skipping the refractory and + neuron-decay the event lane gives git/process. No message loss; only + presence-noise control is relaxed. The clean fix is to split message + scanning into its own sensor so the lane is chosen per *observation* — + also the seam a future external-chat (event) sensor wants. + +## Alternatives Considered + +- **Tune the gates so messages always pass.** Raise message magnitude + above every threshold, exempt them from cleanup. Rejected: it keeps + messages coupled to per-session timing state (governor window, + refractory) that can still diverge between peers, and it is a pile of + special-cases rather than a contract. The defect is structural, not a + threshold value. +- **Make the salience gate anchor to first-observation instead of file + mtime.** Fixes the backlog-decay edge but not this bug (the message + was fresh) and does nothing for the destructive-cleanup race or + multi-recipient addressing. A partial patch on one of seven gates. +- **Full cross-party acknowledgement protocol with shared-state GC.** A + message lingers in a shared store until every recipient explicitly + acks, then a collector reaps it. Rejected as too heavy: the office + analogy says the durable thing is *each worker's own tray*, emptied + when that worker processes their own fax — not a handshake the senders + and receivers all have to participate in. Per-recipient ownership gets + the same no-silent-drop guarantee with far less coordination. +- **Leave delivery best-effort; document that messages can drop.** + Rejected: the human's mental model is "I sent it, the agent will see + it." Silent loss of intentional communication is the worst failure + mode for a coordination surface, and "silence is a valid reply" + ([[ADR-121]]) only holds if the recipient actually *received* the + message and chose not to answer. diff --git a/docs/architecture/attend/ADR-137-boundedness-bounded-work-per-cycle.md b/docs/architecture/attend/ADR-137-boundedness-bounded-work-per-cycle.md new file mode 100644 index 00000000..971bf601 --- /dev/null +++ b/docs/architecture/attend/ADR-137-boundedness-bounded-work-per-cycle.md @@ -0,0 +1,120 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: attend +basis: + - evidence: a unit that scanned an external collection took time proportional to its size plus serial network latency, blew its cycle budget, and was killed without a trace +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-20 +deciders: + - aaronsb + - claude +related: + - ADR-113 +imported: + from: docs/architecture/system/ADR-137-boundedness-bounded-work-per-cycle.md + format: v0 + status: Accepted +--- + +# ADR-137: Boundedness — a unit of work must be bounded within its cycle + +## Context + +Several parts of this system run **cyclically**: a unit of work is invoked +repeatedly on a schedule, sharing a budget with everything else in that cycle. +A sensor polled by the awareness loop is the clearest case — it runs on a tick, +under a hard timeout, and a slow one stalls every other unit's cadence — but the +shape recurs anywhere there is a recurring budget per iteration. + +The cooperative assumption underneath all of these is that **each unit finishes +quickly and yields.** That assumption silently breaks when a unit's work is +*unbounded* — when its cost is proportional to something it does not control: + +- a collection that can grow without limit (scan all N items), +- a remote resource whose latency is serial and uncapped (page-after-page, + request-after-request), +- a stream with no natural end. + +When such a unit exceeds its budget the failure is usually **silent**: it is +killed mid-work and returns nothing, which is indistinguishable from "there was +nothing to report." Silent failure that masquerades as a quiet success is the +worst outcome — nobody is alerted, and the system looks healthy while a whole +class of observations never lands. We hit exactly this: a unit that scanned an +external collection took time proportional to the collection's size plus serial +network latency, blew its cycle budget, and was killed without a trace. The data +volume was trivial; the *latency against a fixed budget* was the wall. + +The lesson is general and worth stating once, abstractly, so it doesn't have to +be rediscovered per-feature. + +## Decision + +**A unit of work that runs within a cycle must do bounded work per cycle.** +"Bounded" along three axes: + +- **Time** — it completes comfortably within the cycle's budget, with margin, not + at the edge of the timeout. +- **Volume** — the data it processes per cycle is capped, not proportional to a + source that can grow without limit. +- **Scope** — it observes a bounded slice, not "everything," when "everything" + has no fixed size. + +**If a task is inherently unbounded, it does not belong inside the cycle.** An +open-ended collection, an uncapped remote, or a stream is the wrong thing to do +*inline* on a tick. The cyclic unit should instead **read a prepared, bounded +result** — a local file, a cached value, a small queue — that some other, +longer-lived shape produced at its own pace. Producing the result by doing the +unbounded work inline is the anti-pattern; consuming an already-bounded result is +the pattern. (This ADR deliberately does not prescribe *which* out-of-cycle shape +to use — that is a per-case choice, and conflating it with this principle is how +the principle got lost the first time.) + +**Boundedness violations must be loud, not silent.** A unit killed for exceeding +its budget must surface that fact — a diagnostic line, a recorded over-budget +marker — so "over budget" can never be read as "nothing observed." A budget +guard that fails quietly is worse than no guard, because it converts a +performance problem into an invisible correctness problem. + +## Consequences + +### Positive + +- The cycle stays responsive: no single unit can starve the others or stall the + whole loop by running long. +- Cost becomes independent of external size — a unit scales the same whether the + thing it watches has ten items or ten thousand. +- Failures are diagnosable. "Over budget" is visible, so it is fixed as a design + problem instead of haunting the system as phantom missing data. + +### Negative + +- Some genuinely useful observations are inherently unbounded, and this forbids + doing them inline. They must be pushed into a separate, longer-lived shape — + that is *more* moving parts, not fewer, and a real cost to weigh before + deciding the observation is worth having at all. + +### Neutral + +- This draws a line between what may run inside a cycle and what must run outside + it, without naming the outside mechanism. The outside mechanism is a separate + decision each time. + +## Alternatives Considered + +- **Raise the per-cycle budget (longer timeout).** Rejected: it only moves the + cliff. A larger budget still has an edge a larger workload crosses, and a longer + timeout lets one slow unit stall the whole cycle — the budget exists precisely + to protect the shared cadence. +- **Let units run unbounded but asynchronously within the cycle.** Rejected: it + breaks the cooperative, finishes-and-yields model the cycle depends on, and + reintroduces the silent-overrun failure in a subtler form (work that never + completes rather than work that is killed). +- **Cap the work but fail silently when the cap is hit (truncate quietly).** + Rejected: silent truncation is the same invisible-correctness trap as the + silent kill — the consumer believes it saw everything. If work is dropped, that + must be observable. diff --git a/docs/architecture/attend/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md b/docs/architecture/attend/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md new file mode 100644 index 00000000..94e3f14e --- /dev/null +++ b/docs/architecture/attend/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md @@ -0,0 +1,198 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: orphan @name/ dirs with no _groups.yaml entry escaped cleanup_stale and stale test channels accumulated; attend send --focus rejected human-only groups as having no live peers + - precedent: ADR-118 + - precedent: ADR-120 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-21 +deciders: + - aaronsb + - claude +related: + - ADR-118 + - ADR-120 + - ADR-124 + - ADR-129 +imported: + from: docs/architecture/system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md + format: v0 + status: Accepted +--- + +# ADR-170: Human focus-group membership via username identity and a shared attend-groups crate + +## Context + +attend-chat (ADR-120) is the human's seat on the signal bus, but focus-group +membership (ADR-118) is claude-only: `_groups.yaml` members are Claude Code +session UUIDs, liveness is judged against the per-session heartbeat sidecar +(ADR-129), and the only writer of the yaml is the `attend` binary acting for a +claude session. The chat TUI's slash registry advertises `/join` and `/leave` +as planned, and three deferred pieces all converge on them: + +1. **Humans have no membership key.** The chat user has no session UUID and no + heartbeat, so there is nothing to write into a `members:` list — and a + heartbeat-less member would immediately count as dead to `live_peer_count` + and be swept by `cleanup_stale`. +2. **attend-chat has no write path.** Its `groups` module is an explicitly + read-only byte-mirror of `attend::groups` — two hand-rolled YAML parsers + kept in sync by golden tests, with the module docs deferring "a shared I/O + layer" to the `/join` write path. +3. **Membership is invisible on human chips.** The chip renderer looks up + group glyphs by session UUID, which humans don't have. + +Separately, agent-side send validation (`attend send --focus`) counts a member +as live only if it appears in `PeerSensor::live_session_ids` — a claude-process +scan. A human member would never count, so a group whose only live member is a +human would reject agent sends with "no live peers". + +Channel hygiene has a matching blind spot. The chat's channel bar renders +every `@name/` directory on disk, but attend's `cleanup_stale` iterates +`_groups.yaml` entries — an **orphan dir** (no yaml entry at all) is invisible +to cleanup and renders in the bar forever. In practice stale test channels +accumulate and the human has no way to remove them from the surface where +they cause the confusion. + +## Decision + +**A human's membership identity is their sanitized username** (e.g. `aaron`, +via `agent_identity::sanitize_id_component($USER)`). One entry per human, +regardless of terminal or cwd — deliberately matching the chat registry's +existing human-dedupe rule (the same person in two terminals is one identity, +where two claudes in two cwds are two). + +**Human liveness rides the existing heartbeat sidecar.** While attend-chat +runs, it touches `heartbeat/<username>` on its existing 5-second refresh tick. +No new liveness mechanism and no yaml format change: `members:` lists now mix +session UUIDs and usernames, and every consumer already judges members by +heartbeat freshness (`live_peer_count`, `cleanup_stale`), so human members age +out after `DEFAULT_GRACE` exactly like abandoned claude sessions. attend-chat +does not clear the heartbeat on exit — a second chat instance for the same +user may still be running, and the 90s grace self-corrects. + +**Group I/O is extracted into a shared `attend-groups` workspace crate**, +following the pattern set by `attend-heartbeat` and `attend-instances`. The +crate owns `GroupEntry`, `Groups` (join/leave/pin/dissolve/cleanup and the +yaml read-modify-write), `validate_group_name`, and the parse/serialize pair. +`attend` keeps its attend-specific pieces (the ADR-124 `@open` migration); +attend-chat's mirror parser and the cross-crate golden-drift tests collapse +into the shared crate's own tests. + +**agent-side liveness validation gains a heartbeat fallback.** `attend send +--focus` counts a member live if it is a live claude session *or* its +heartbeat is fresh — making a human-only group a valid send target. This +deliberately loosens the claude-side check too: a session whose claude died +but whose attend still heartbeats now counts as live, consistent with the +crate's `member_alive` philosophy (no attend, no mesh participation — and the +converse). The chat-side gate mirrors the agent side's self-exclusion: the +sender's own membership (now heartbeat-backed for humans) never counts toward +"live peers", or a solo human's send would validate against their own +heartbeat and sit unread. + +**The TUI wires the commands.** `SlashOutcome` grows effect variants; dispatch +stays IO-free (parse + validate only) and the key handlers execute effects: + +- `/join <group>` / `/leave <group>` call the shared `Groups` with the + username identity; `/clear` empties the message buffer (display-only). +- `/dissolve <group>` removes a channel entirely — yaml entry and `@dir`, + including orphan dirs the yaml doesn't know about. Chat-side it carries a + live-member guard the CLI's `attend focus dissolve` does not: a group with + heartbeat-fresh members refuses to dissolve, since from the TUI this is a + hygiene action and should not yank a channel out from under active peers. +- `/channels` lists every channel with `live/total` member counts in the + status row (IRC `/list` shaped) — `0/0 live` is the tell for `/dissolve` + fodder. +- `/purge [group]` deletes a channel's on-disk signal history (default: the + base channel's `_broadcast/`), keeping a heartbeat-grace tail so nothing a + peer's sensor may be mid-scan on is deleted. This is a **deliberate operator + override of ADR-136's durability default** — that ADR forbids *automatic* + age-reaping; an explicit human purge of a named channel is a different act, + and the sensors already tolerate signal deletion (the project-liveness + cleanup and `attend cleanup --nuke-all` predate this). Membership and the + channel itself survive a purge, unlike `/dissolve`. When a per-session + consumption checkpoint exists (the drain-verb work being specced against + ADR-136), purge should tighten to also refuse signals unconsumed by a live + session; the grace tail then becomes the fallback for non-live consumers. + +Human chips look up group glyphs by username so membership renders the same +as it does for claudes. + +**`cleanup_stale` gains an orphan-dir sweep.** After the member pass, `@name/` +dirs with no yaml entry and an mtime older than the heartbeat grace window are +removed — closing the accumulation path so stale channels stop outliving +their groups. Three guards defend concurrent joins, since `create_dir_all` on +a pre-existing orphan dir does not refresh its mtime: `join` saves its yaml +entry *before* touching the dir, the sweep re-reads the yaml immediately +before each removal, and the chat's group resolver falls back to the yaml +entry when the dir is missing (signal writers re-create their target dir), so +even the residual sub-millisecond race self-heals. Reserved names are never +swept — a lingering `@open/` belongs to the ADR-124 migration, which moves +its signals into `_broadcast/` rather than deleting them. + +## Consequences + +### Positive + +- Humans become first-class group members with zero wire-format change — + every existing consumer works unchanged because liveness was already + heartbeat-shaped, not UUID-shaped. +- The two hand-rolled YAML parsers and their golden-mirror maintenance burden + are replaced by one implementation with one test suite. +- Group lifecycle rules (empty-unpinned GC, stale sweeps) apply to humans for + free; an abandoned chat session cannot pin a group open forever. +- `attend send --focus` stops lying about human-only groups. +- Stale channels become manageable from the surface where they confuse: + `/channels` shows which are dead, `/dissolve` removes them, and the orphan + sweep stops the accumulation at the source. + +### Negative + +- Usernames and session UUIDs share one namespace in `members:`. Collision is + implausible (UUIDs vs short login names) but the list is no longer + homogeneous, and tooling that assumed "member = session UUID" must not + reappear. +- A username heartbeat conflates "some chat instance is running" with "this + chat instance is running" — two instances for one user share a heartbeat by + design, so per-instance presence for humans is out of scope. +- One more workspace crate to version and build. + +### Neutral + +- attend-chat gains its first write responsibilities on the signal base + (yaml read-modify-write and heartbeat touches), inheriting the same + last-writer-wins races attend sessions already tolerate. The extraction + hardened the write itself to keep that risk model honest: per-writer unique + tmp names mean concurrent savers can no longer publish a torn hybrid file — + the worst case is genuinely last-writer-wins, not corruption. +- The chat watcher still renders all groups regardless of membership; + subscribed-group *filtering* remains future ADR-120 work — after this ADR, + joining changes presence and addressability, not what the human sees. + +## Alternatives Considered + +- **Prefixed human member ids (`human:aaron`)** — would make the member kind + explicit, but changes the yaml contract in both parsers, requires + special-casing in every liveness check, and buys nothing the heartbeat + doesn't already provide. Rejected for format churn without benefit. +- **Per-instance human identity (`aaron@kitty`)** — mirrors the wire `from` + field, but splits one person into N members, contradicts the chat + registry's human-dedupe rule, and makes `/leave` ambiguous about which + instance leaves. Rejected: membership is about the person, not the seat. +- **attend-chat depends on the attend binary crate as a library** — avoids a + new crate but drags sensor/CLI machinery into the TUI build and inverts the + dependency taxonomy the small shared crates established. Rejected. +- **Keep duplicating: hand-roll a second yaml writer in attend-chat** — the + read-side mirror is already a documented maintenance hazard; a write-side + mirror doubles the drift surface on the file both binaries mutate. + Rejected; the golden tests exist precisely because this was fragile. +- **Synthetic session files for humans** — writing fake + `~/.claude/sessions/*.json` entries so humans traverse the claude discovery + path. Rejected: pollutes Claude Code-owned state and misrepresents what a + session is. diff --git a/docs/architecture/attend/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md b/docs/architecture/attend/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md new file mode 100644 index 00000000..d12c1503 --- /dev/null +++ b/docs/architecture/attend/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md @@ -0,0 +1,129 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: 'issue #378: a live multi-session test rendered one human and two sessions as four-plus identities, with three root causes reproduced from one session''s history' + - precedent: ADR-129 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-21 +deciders: + - aaronsb + - claude +related: + - ADR-129 + - ADR-136 + - ADR-168 + - ADR-170 +imported: + from: docs/architecture/system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md + format: v0 + status: Accepted +--- + +# ADR-171: Stable session identity — the roster enumerates addressable coordinating units + +## Context + +A live multi-session test (issue #378) surfaced the "multiple personalities" +failure: one human and two claude sessions rendered as four-plus identities +(`hana`, `Hana-beta`, the ghost `Hana-alpha`, a `tools/`-keyed persona from +the same session). Three root causes were reproduced from one session's own +history: + +1. **Identity keyed off process cwd at launch.** attend derived its working + dir from `current_dir()`, so a stray shell `cd` leaking into an + `attend run` launch (a Monitor inheriting a build directory) put the + session on the bus as a different persona than its project. +2. **Registration happened once, at startup.** The periodic instance-registry + maintenance used `touch`, which is a no-op for a missing entry — a session + whose registration was GC'd or never written rendered as the bare + nickname forever. +3. **Historical wire data kept superseded personas alive** in the chat + legend for as long as their signals sat in the buffer. + +The debate (operator + two peer sessions, recorded on #378) also settled what +identity *is* on this bus, which the fix must encode. + +## Decision + +**Identity is the coordinating unit, not the process.** A claude's canonical +identity is the tuple `(sessionId ∩ origin_path)` — the session UID paired +with the *session record's* cwd, never the process cwd. Concretely: + +- **A new `attend-session` crate owns the one derivation** of "who am I on + the bus": pid-ancestry walk over `~/.claude/sessions/*.json` for the + session UID, the session record's `cwd` as the origin path, with explicit + flagged fallbacks (`pid-<pid>` + process cwd, `resolved: false`) for + processes no Claude session owns. attend (run, send, status, peers, inbox, + config), sensor-peers, the instance registry, the heartbeat id, and + focus-group member ids all resolve through it. Downstream durable state — + the planned per-session consumption checkpoint (drain research) — keys on + the same tuple by calling the same derivation. +- **`attend whoami` is the CLI accessor** for the tuple (`--machine` emits + `key=value` of only the stable fields), so hooks and scripts obtain + identity without touching attend-owned state. The rendered display name is + deliberately absent from machine output: **ordinals are presentation, + never keys.** +- **The roster enumerates addressable coordinating units** — top-level + sessions, by session UID per origin dir. A second top-level instance in + the same origin dir takes the next Greek letter (the existing ADR-129 + allocator, which was already idempotent per `(cwd, sessionId)`; this ADR + fixes its *inputs*). Subagents are the supervisor's efferent limbs, not + afferent participants: they get no roster identity and no letter. Their + activity may render as decoration on the parent's chip (deferred, + presentation-only). +- **The periodic registry maintenance upserts** (`register`, idempotent) + instead of touching, so a session with a missing entry self-heals within + one interval instead of rendering bare forever. + +## Consequences + +### Positive + +- A session's bus persona is immune to shell cwd drift — the launch + environment can no longer mint accidental identities. +- The "bare nickname" degradation self-heals; ghosts stop accumulating at + their source. +- One identity derivation shared by four consumers replaces three partial + ones; the drain checkpoint research can key on it without re-implementing + resolution. +- A minimal attend build (without sensor-peers) now resolves real session + identity instead of degrading to `pid-<pid>` member ids. + +### Negative + +- Session-record resolution costs a sessions-dir walk plus a `ps`-based + ancestry climb per subcommand invocation — negligible at CLI cadence, but + no longer a bare `getcwd`. +- Historical signals written under superseded personas still render as + distinct chips until they age out of buffers — display residue this ADR + accepts rather than rewriting history. + +### Neutral + +- Humans are unaffected: their username identity (ADR-170) was stable by + construction. +- `Focus::default_focus()` keeps its process-cwd default for the generic + sensor API; attend overrides the working dir at the one place identity is + authoritative. + +## Alternatives Considered + +- **Per-process registry identity (PID lineage) so subagents get letters** — + rejected in the #378 debate: subagents have no independent voice in the + channel, letters would flicker with worker lifecycles, and consumption is + session-level regardless. Identity follows addressability. +- **Keying anything on the rendered ordinal** — rejected: slot reuse would + alias different sessions over time; monotonic slots grow unboundedly. + Ordinal is presentation; the tuple is the key. +- **Fixing cwd drift by documenting "don't `cd` before launching attend"** — + rejected: the failure was produced by an agent following normal build + workflows; discipline that one stray `cd` defeats is not an invariant. +- **Synthesizing session records for non-session processes** — rejected: + pollutes Claude Code-owned state (same reasoning as ADR-170's rejection of + synthetic session files). diff --git a/docs/architecture/attend/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md b/docs/architecture/attend/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md new file mode 100644 index 00000000..ffaf82bc --- /dev/null +++ b/docs/architecture/attend/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md @@ -0,0 +1,237 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: during the ADR-171 identity work, authored messages to an active session were observed waiting for the next peers-sensor poll + - precedent: ADR-136 + - precedent: ADR-171 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-22 +deciders: + - aaronsb + - claude +related: + - ADR-124 + - ADR-129 + - ADR-136 + - ADR-168 + - ADR-170 + - ADR-171 +imported: + from: docs/architecture/system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md + format: v0 + status: Accepted +--- + +# ADR-172: Turn-boundary inbound delivery via a CLI-owned drain checkpoint + +## Context + +Authored messages (ADR-136's message lane — `attend send`/`reply`, `@Name`, +`#group`, `#open`) reach a session through one conduit: the `peers` sensor, +polling on a wall-clock cadence under Monitor. A message written to a +recipient's tray waits there until the next poll surfaces it, then rides the +notification bridge into a turn. That poll interval is a floor on delivery +latency. + +Two regimes have different needs, because a Claude Code session is +**turn-gated** — it can only ingest an inbound message when something crosses +into the turn dimension (ADR-136: "the notification is the bridge"): + +- **Idle** (turn finished, awaiting input): nothing the session itself runs + can wake it. Only an in-flight conduit can — which is precisely what the + Monitor-hosted sensor *is*. The poll interval is the unavoidable cost of + waking an idle loop. +- **Active** (mid-turn, running tools): the session is right there, but an + arriving message still waits for the next poll. Here the poll latency is + pure lag — and there is a cheaper conduit available. A `Stop` hook fires at + the exact moment a turn ends, before the session goes idle, and can inject + text and force continuation. Draining pending messages there delivers them + with **no poll latency**. + +The general shape: inbound awareness is *afferent* — it must be carried across +the turn boundary by a conduit the loop can consume. The poller is the conduit +that can reach an *idle* loop; the turn boundary is a free, exact conduit for +an *active* one. Adding the second does not remove the first. + +The latency observation surfaced during the ADR-171 identity work; ADR-171 +supplies the stable key this delivery path depends on (`whoami --machine` → +`session_id` / `origin_path` / `resolved`), so a consumption record cannot be +keyed on a persona that shell-cwd drift or presentation churn could alias. + +## Decision + +**Deliver authored messages at the earliest turn boundary through a +CLI-owned atomic drain that records consumption on the stable identity tuple; +keep the poller as the idle-session conduit; and let a single seen-set keyed +on that tuple keep every conduit deliver-once.** + +Concretely: + +1. **A `Stop`-hook drain is the active-session fast path.** On turn end the + hook pulls pending authored messages for this session and, if any, injects + them and continues the turn — zero poll latency. The Monitor-hosted `peers` + sensor is unchanged and remains the **idle-session** waker; this ADR adds a + conduit, it does not replace one. + +2. **The hook is a dumb invoker.** It shells out to a new attend-owned verb — + `attend inbox --drain` (`--format hook`) — that atomically returns + pending-for-this-session *and* records their consumption. The hook never + reads attend-owned state. CLI remains the whole interface (ADR-124, + ADR-136). + +3. **One seen-set, two consumers — under merge semantics.** The drain and the + `peers` sensor share the single per-session seen-set ADR-136 already + maintains — they do not keep parallel dedup stores. But the two have + different lifecycles: the drain is a one-shot process writing at turn + boundaries, while the sensor is long-running, holds the set *in memory*, and + checkpoints it via `StateStore`. A whole-set snapshot on either side forks + the "one" set into two — a stale sensor snapshot re-surfaces a message the + drain already marked (double delivery), or a sensor checkpoint overwrites + the drain's marks (a consumed message resurrects). So the contract requires + **merge/append consumption semantics** — union-on-checkpoint, or + per-message mark files — and the sensor must **observe drain-recorded + consumption before emitting**. Deliver-once across both conduits holds only + under these semantics; a whole-set overwrite is how the "one" seen-set + silently becomes two. + +4. **Consumption keys on the resolved tuple, and refuses to mark under an + unresolved identity.** The seen-set (and every durable state ADR-171 names) + keys on `(session_id ∩ origin_path)`, obtained via `attend whoami`, never + the rendered display name. If `whoami` reports `resolved=false` (the + `pid-<pid>` / process-cwd fallback), the drain records **nothing** and + no-ops: marking consumption under an unstable id would alias sessions, + corrupt the shared seen-set, and desync from purge-liveness. Under + unresolved identity, Monitor remains the delivery path — a graceful + degradation. + +5. **Durability is preserved; the purge-agreement is a co-shipped deliverable, + not a present guarantee.** The drain reaps nothing — removal stays owned by + ADR-136 project-liveness reaping and the ADR-170 `/purge` human power tool. + The desired property — `/purge` cannot shred a message a live, resolved + session has not yet consumed — is *not* true as deployed today: purge + consults only the 90 s age tail (ADR-170 records the seen-set refusal as a + **planned tightening**, precisely because no consumption record existed to + consult). It becomes real only once purge's seen-set check ships **with** the + drain verb — an explicit deliverable of the implementation PR (below), not a + property this ADR can assert of the system as it stands. + +6. **Re-entry is bounded.** Inject-and-continue means the continued turn also + ends in a `Stop` hook, so the drain re-runs. Draining to empty terminates + for a single session; but two *actively conversing* sessions can each + injection-trigger the other's drain and keep both turns alive — sometimes + the desired liveliness, sometimes a livelock. The implementation carries the + standard Stop-hook re-entry guard plus a ceiling (drain-rounds-per-turn or a + quiet-period), so an empty drain always lets the turn end and a live + exchange cannot spin unbounded. + +```mermaid +flowchart TD + Msg[Authored message in this<br/>session's tray / room ledger] + Msg --> Idle{Session state?} + Idle -->|idle — nothing to interrupt| Poll[peers sensor under Monitor<br/>wall-clock poll] + Idle -->|active — turn about to end| Hook[Stop hook → attend inbox --drain] + Poll --> Seen[(one per-session seen-set<br/>keyed on session_id ∩ origin_path)] + Hook --> Seen + Seen --> Once[deliver-once across both conduits] + + classDef msg fill:#7c3aed,color:#fff,stroke:#4a5568 + classDef decision fill:#fbbf24,color:#1a1a1a,stroke:#4a5568 + classDef proc fill:#2d7d9a,color:#fff,stroke:#4a5568 + classDef store fill:#2d8e5e,color:#fff,stroke:#4a5568 + class Msg msg + class Idle decision + class Poll,Hook,Once proc + class Seen store +``` + +The on-disk form of the drain's consumption record (extend the existing +seen-set store vs. a sibling ledger) is an implementation choice for the +follow-up PR. The contract above — atomic drain, shared seen-set, tuple key, +resolved-gate, merge semantics — is what this ADR fixes. The implementation PR +carries four coupled deliverables, none shippable alone: (a) the +`attend inbox --drain` verb; (b) merge/append seen-set persistence +(Decision 3); (c) purge's seen-set tightening (ADR-170), co-shipped so the +durability property in Decision 5 becomes real; and (d) the Stop-hook re-entry +guard (Decision 6). + +## Consequences + +### Positive + +- An actively-working session receives authored messages the instant its turn + ends, instead of at the next poll — the latency the poll interval imposed on + the active regime is gone. +- Deliver-once (ADR-136) is preserved across both conduits because they share + one tuple-keyed seen-set; there is no second dedup store to diverge. +- CLI-is-contract holds: the hook is a thin invoker, and identity is obtained + through `attend whoami` rather than by reaching into attend state. +- `/purge` (ADR-170) and the drain agree — same tuple, same seen-set — so the + human power tool cannot delete a message a live resolved session has not + consumed. *This holds only once purge's planned seen-set tightening ships + with the drain verb (Decision 5); it is not a property of today's + age-tail-only purge.* +- Keying on the machine tuple makes consumption immune to display-name churn + (the presentation layer ADR-171 deliberately separated from the key). + +### Negative + +- A `Stop` hook runs at every turn end, so the drain must be cheap: one + memoized-identity shell-out, O(pending) work, and a fast no-op when the tray + is empty. A slow drain would tax every turn. +- A second delivery conduit is more surface than one poller, and the drain is a + second *writer* to the seen-set — which is why the merge/append discipline + (Decision 3) is load-bearing, not optional. Both still feed one logical set, + so there is a single source of truth for "seen" — but only if neither side + overwrites the other's marks. +- The turn-boundary path still cannot wake an idle session; the two-conduit + split (poller for idle, hook for active) is inherent to a turn-gated loop, + not complexity this ADR could design away. + +### Neutral + +- Requires a new `attend inbox --drain` verb; `inbox` is read-only today + (list + read-by-id over the never-reaped ledger). The verb adds the first + consumption-recording path to that surface. +- Under unresolved identity the drain no-ops and Monitor delivers instead — a + degradation, not a failure; the session still receives its mail. Note that + `StateStore` persistence is also absent when the session id is `None`, so the + Monitor path's dedup is process-lifetime only in that fallback. +- The three synchronized messaging docs (`skills/attend/SKILL.md`, + `tools/sensor-disclosure/src/disclosures/messaging.md`, + `hooks/ways/softwaredev/environment/attend/attend.md`) move in lockstep if + this changes the messaging contract, per ADR-136. + +## Alternatives Considered + +- **Let the hook read the signal dir / seen-set directly.** Rejected: it + violates CLI-is-contract (ADR-124/136) and makes the hook an *uncoordinated* + second consumer racing the sensor — two readers mutating dedup state diverge. + The verb keeps attend the single reader; the hook only invokes it. +- **Replace the Monitor poller with the Stop hook entirely.** Rejected: a hook + has no in-flight call, and an idle turn-gated loop that already finished its + turn has nothing to interrupt — so a hook cannot wake an idle session. The + poller is the only conduit that reaches idle; the hook is strictly an + active-regime accelerator. +- **A consumption checkpoint separate from the ADR-136 seen-set.** Rejected: + two dedup stores over the same messages diverge; the drain would re-deliver + what the sensor already surfaced, or vice versa. Reuse the one seen-set. +- **Lower the sensor poll interval instead.** Rejected: it burns wall-clock + CPU for *every* session to shave latency only for *active* ones, and never + reaches zero because the structural cost is the turn gate, not the interval. + The turn boundary is free and exact. +- **Key consumption on the rendered name, or drain under an unresolved + identity.** Rejected for the same reason ADR-171 rejects it: slot reuse + aliases different sessions over time and an unresolved `pid-<pid>` id + corrupts the shared seen-set and desyncs purge. Ordinal is presentation; the + resolved tuple is the key. +- **Channels / MCP push as the delivery mechanism.** Considered: an external + push channel is the one primitive that could wake an *idle* session without a + persistent poller, and may later subsume the Monitor conduit. Out of scope + here — it is a heavier, separately-gated capability; the turn-boundary drain + is the cheap win that needs no new transport. diff --git a/docs/architecture/attend/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md b/docs/architecture/attend/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md new file mode 100644 index 00000000..a48271ac --- /dev/null +++ b/docs/architecture/attend/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md @@ -0,0 +1,167 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: corpus-aligned verbs (send, reply, @name, inbox) are used correctly by agents on first contact while the bespoke focus verbs need disclosure every session; clear is a cross-surface homonym + - precedent: ADR-124 + - precedent: ADR-170 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-22 +deciders: + - aaronsb + - claude +related: + - ADR-118 + - ADR-124 + - ADR-136 + - ADR-170 + - ADR-172 +imported: + from: docs/architecture/system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md + format: v0 + status: Accepted +--- + +# ADR-173: Chat-idiom convergence for the attend command surfaces + +## Context + +Attend has two command surfaces that grew separate vocabularies for the +same objects. The agent CLI speaks an attention metaphor (`focus on +deploy`, `focus off`, `focus clear`, `send --focus`); the operator TUI +speaks chat idiom (`/join #deploy`, `/leave`, `/channels`, `#name` +syntax). ADR-170 already mixes the registers in one sentence — "humans +join focus groups" — and every doc that touches messaging has to teach +the mapping. + +The divergence has a measurable asymmetry in who pays for it. An +operator learns the TUI once. An agent session learns its surface +**every session**, through disclosure: the skill primer, the runtime +reheat, and the just-in-time way exist substantially because the +bespoke verbs are not guessable. Meanwhile the parts of the surface +that align with the chat corpus every model is trained on — `send`, +`reply`, `@name`, `#open`, `inbox`, `/purge` — need no teaching at +all; agents use them correctly on first contact. Bespoke vocabulary is +a recurring disclosure tax; corpus-aligned vocabulary is free. + +Two concrete defects sharpened the question. `clear` is a cross-surface +homonym: `attend focus clear` leaves all groups while `/clear` wipes +the transcript — a user of one surface guesses wrong on the other. And +the surfaces' *capabilities* have drifted independently of their +*vocabulary*: operator discovery (`/peers`, `/whois`) is registered but +unimplemented, while the agent has had `peers`/`whoami` from the start. + +## Decision + +**Converge both surfaces on the chat idiom. Where an agent-facing +surface has a saturated training-corpus convention, prefer it over +bespoke vocabulary — the model's prior is free disclosure; a bespoke +term is a tax paid every session.** + +Concretely: + +1. **The shared noun is "channel"** — the thing `#name` syntax already + implies. "Focus group" disappears from user-facing vocabulary; the + `@name/` directory layout and `_groups.yaml` wire format are + unchanged (storage is not surface). + +2. **Agents adopt the chat verbs as primary.** The mapping: + + | Concept | Operator (attend-chat) | Agent today | Agent after | + |---|---|---|---| + | Enter a channel | `/join <#g>` | `focus on <g>` | `join <g>` | + | Leave a channel | `/leave <#g>` | `focus off <g>` | `leave <g>` | + | List channels | `/channels` | `focus list` / `focus all` | `channels` | + | Leave all | — | `focus clear` | `scene private` | + | Destroy a channel | `/dissolve` | `focus dissolve` | `dissolve <g>` | + | Scope a send | `#g` prefix | `send --focus <g>` | `send --channel <g>` | + | Directed send | `@name` | `send --to <path>` | unchanged (`--to`) | + | New message / reply | type / — | `send` / `reply` | unchanged | + +3. **`focus` survives as a deprecated alias** (CLI-is-contract, + ADR-124): every `focus` invocation keeps working, maps to the new + verb, and prints a one-line deprecation note to stderr. No script + breaks; no flag-day. + +4. **The `clear` homonym is resolved by removal, not renaming.** The + agent-side leave-all folds into `scene private` (its existing + duplicate); `/clear` keeps the universal clear-the-screen meaning — + which is also the chat prior. + +5. **Vocabulary converges; capabilities do not.** The asymmetries are + design, and this ADR re-affirms them: `/purge` remains + operator-only (the ADR-136 durability override is a human power + tool); `run`, `inbox --drain`, `tune`, `sensors`, `config`, + `permissions` remain agent/session infrastructure. Discovery parity + is closed in the operator's favor: `/peers` and `/whois` graduate + from Planned to implemented. + +6. **The idiom's borrowed expectations are met or explicitly ended.** + Adopting chat vocabulary imports chat priors: a DM exists (`send + --to`), presence exists (`peers`), threading exists (`reply`). + Where the analogy stops — no read receipts, no typing indicators, + no message editing, durability-by-liveness instead of retention + policy — the three synchronized messaging docs say so in one + sentence rather than leaving the prior to guess. + +## Consequences + +### Positive + +- Zero-shot guessability: an out-of-the-box agent's first attempt + (`attend join deploy`) is correct, before any disclosure fires. +- The disclosure budget spent teaching `focus on/off` every session is + reclaimed for guidance that actually needs it. +- One register across both surfaces ends the "humans join focus + groups" split — docs, ways, and reheat text all speak channel. +- The `clear` homonym — the sharpest cross-surface hazard — is gone. +- Operator discovery reaches parity (`/peers`, `/whois`). + +### Negative + +- Alias maintenance: `focus` must keep working indefinitely under + CLI-is-contract, so the CLI carries two spellings of every channel + verb (one deprecated) plus tests for both. +- The three synchronized messaging docs, the attend way, the skill, + and the reheat disclosure all need coordinated rewording in one PR — + the ADR-136 lockstep rule applies in full. +- Borrowed expectations are a standing documentation duty: every new + chat-prior feature request must be either met or explicitly ended. + +### Neutral + +- Storage and wire formats are untouched — `@name/` dirs, + `_groups.yaml`, signal files, seen-set keys all keep their shapes; + this is a surface decision, not a data migration. +- The adjacent presentation-consistency backlog rides the same + convention but ships separately: plain render for injected text + (#388), the cell datetime line and cross-surface timestamps (#389), + path-based attachments (#390). +- `scene` survives unchanged as the agent's attention-preset layer — + it has no chat-idiom collision and its semantics (reconfigure many + channels at once) have no single-channel analog. + +## Alternatives Considered + +- **Keep both registers and document the mapping.** Rejected: the + mapping table is a permanent artifact someone must maintain, and the + agent side keeps paying the disclosure tax the table exists to + paper over. Documentation is the symptom, not the fix. +- **Converge on "focus" everywhere.** Rejected: it fights the + training prior on both surfaces — operators know IRC/Slack idiom + too, and ADR-170 chose `/join` for them deliberately. The attention + nuance that motivated "focus" (subscription tied to sensor routing) + never behaviorally diverged from room membership enough to earn its + vocabulary. +- **Full chat-platform parity (read receipts, edits, presence + indicators).** Rejected as scope: the idiom is adopted for its + verb vocabulary, not its feature checklist. Decision 6 handles the + expectation gap with documentation rather than implementation. +- **Hard rename without aliases.** Rejected: CLI-is-contract + (ADR-124) — scripts, hooks, and muscle memory built on `focus` + must not break on a vocabulary decision. diff --git a/docs/architecture/attend/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md b/docs/architecture/attend/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md new file mode 100644 index 00000000..6161efa1 --- /dev/null +++ b/docs/architecture/attend/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md @@ -0,0 +1,88 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: a cold rewrite of a 200k-token prefix under Fable 5.1 costs about four dollars against five cents for a warm turn + - evidence: the cache-tax mod (karanb192/claude-code-mods) keeps the cache warm only through early-access function hooks + - precedent: ADR-113 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-09-17 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-136 + - ADR-172 +imported: + from: docs/architecture/system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md + format: v0 + status: Accepted +--- + +# ADR-182: Keepwarm: attend keeps the prompt cache warm with a wake floor + +## Context + +Claude Code's main conversation is cached on the 1-hour prompt-cache tier. Every request reads the cached prefix and refreshes its timer. A request that arrives after the hour lapses rewrites the whole prefix at the cache-write rate, which on a 200k-token session under Fable 5.1 costs about four dollars against five cents for a warm turn. + +The `cache-tax` mod (karanb192/claude-code-mods) solves this with a function hook: at 50 idle minutes it sends a tool-less `$.model.fork` over the session's transcript, reads the reply's usage to confirm the cache was read rather than rewritten, and repeats until an armed window ends. It also drops one cold send with the price on screen. Function hooks are early access, gated behind an environment flag, and nothing in that surface is available to a settings hook or to attend. + +Attend already wakes the session. Every line it emits to stdout becomes a Monitor notification, which Claude Code turns into a turn over the full context. That turn is an API request over the same prefix, so it reads the cache and resets the timer. A peer message, a git change, and a scheduled wakeup all do the same. The cache-tax ping and an attend notification are the same operation with different plumbing. + +Attend also already reads the transcript. The context sensor (ADR-113) shells to `ways context --json`, which parses the per-message `usage` object that carries `cache_read_input_tokens` and `cache_creation_input_tokens`. The warm-or-cold verdict cache-tax reads from its fork's reply is on disk one poll after any wake. + +## Decision + +**A `keepwarm` sensor in attend emits one wake when the session has been idle for 50 minutes, while an armed window is open and the context is large enough to be worth it. The transcript is the clock and the scorecard. The agent's contract is one sentence: answer a keepwarm line with one word and no tools.** + +1. **The clock is the transcript.** Idle time is measured from the timestamp of the last assistant line in the session transcript. That covers every waker: user turns, peer messages, drained inbox messages, cron firings, and keepwarm itself. Attend does not keep its own emission clock for this. + +2. **The floor.** With a window armed, context at or above 50k tokens, and 50 minutes since the last assistant line, the sensor emits one line at medium priority. Medium is the lowest priority `emit_batch` writes to stdout, so the line wakes Monitor. The sensor rides the message lane with the peers sensor: the wake is a timed floor, so it bypasses the action-potential refractory that throttles observation noise. One wake per idle stretch. If no assistant line follows within the remaining ten minutes, the sensor disarms and records why. + +3. **The verdict comes from the next usage entry.** After a wake, the first new assistant response is checked: a cache read of at least half the previous context means the cache was warm and the wake did its job. A smaller read means the cache was already gone. What the response wrote does not enter the verdict, since a real turn landing beside the wake can append a large tool result to a warm prefix. On a cold verdict the sensor stops and records the reason, so an armed window never keeps paying for rewrites. + +4. **Cold writes are scored.** Any new assistant response whose cache write is at least half the previous context size and whose cache read is under half of it, with previous context over 20k, is a paid cold write. The sensor records it and, when the context is large enough for the floor to serve, arms three hours if no window covers them. That mirrors cache-tax: after one rewrite, the session is held warm so the second is not paid today. + +5. **Arming is a CLI verb.** `attend keepwarm on [WINDOW]` arms a window (default six hours), `attend keepwarm off` disarms, `attend keepwarm status` prints the cache state, context size, cold price, window remaining, and the session's cold writes. The window is stored per session under attend's own cache directory. The CLI is the contract (ADR-136): the arm file is attend-owned state. + +6. **`ways context --json` grows a `usage_tail`.** The last assistant responses' usage, one entry per message id, each with timestamp, model, input, cache read, cache write, output, and tier. Claude Code writes one transcript line per content block, and every line of a response repeats its usage, so the tail collapses them. It also gains `--session <id>` so a sensor pinned to one session never reads a peer's transcript in the same project. The transcript parser stays in one place. + +7. **The agent contract lives in the three synchronized guidance files.** The attend skill primer, the runtime disclosure, and the attend way each carry the same paragraph: a line beginning `keepwarm:` asks for one word and no tools, and it is not a prompt to investigate. + +8. **Prices are a dated table.** Cache read, 1-hour cache write, and output rates per model family, from the list prices on the decision date. The table is for the status card and the recorded cost of a cold write. A model not in the table prices as unknown. + +## Consequences + +### Positive + +- A session left open across lunch comes back to a warm cache for the price of one or two cache reads. +- No function hooks, no environment flag, no fork. The mechanism is a Monitor notification, which attend already owns. +- The verdict is stronger than the fork's. A fork must prove empirically that it shared the main cache. A Monitor wake is a turn in the main conversation, so it shares the prefix by construction. +- Every wake leaves a usage line, so the same parse scores cold writes whether or not keepwarm caused the turn. + +### Negative + +- A wake appends to the transcript. Each ping adds a user line and an assistant line, on the order of a hundred tokens. A six-hour window adds under a thousand. +- A model at high effort may think before answering with one word. The guidance forbids deliberation, and the output is billed whatever it says. +- On a subscription the dollars are a yardstick. How a cache read weighs against the five-hour and weekly limits is not documented. The status card prints the arithmetic and the operator decides. +- The refuse-once guard is not in this decision. Dropping a cold send needs the `UserPromptSubmit` hook to return a block decision. That is a hook change with its own ADR. + +### Neutral + +- Attend's chattiness now has a visible unit price: every notification is a cache read. The governor's calm-channel posture (ADR-136) already encodes that cost without naming it. +- The cache clock is an input attend could use for its own emission timing. A low-value disclosure delivered at 65 idle minutes pays the rewrite; the same one at 55 is nearly free. Holding cold-time disclosures is a governor decision left for a later ADR. +- The plugin author reports that Claude Code passes resume fields to settings hooks on `SessionStart`, including seconds since the last response and an estimated cache-write cost. This ADR does not depend on them. If they exist, the cold-resume warning belongs in the SessionStart hook. + +## Alternatives Considered + +- **Port cache-tax as a function-hook plugin.** Rejected. The surface is early access and off by default. Attend already has the timer, the wake, and the transcript parse. +- **A fork-style side request from attend.** Rejected. Attend has no model access, and a side request would have to prove it shared the main cache. A wake is the main cache. +- **Fold keepwarm into the context sensor.** Rejected. The context sensor rides the event lane with its refractory. A timed floor needs the message lane, and the two sensors have different disclosure shapes. +- **Arm by default.** Rejected. A ping costs money on an API key and an unknown amount of rate limit on a subscription. Arming is an operator act, with the auto-arm after a paid cold write as the one exception, because that session has already shown it comes back. +- **Attend keeps its own emission clock.** Rejected. Only the transcript sees every waker. diff --git a/docs/architecture/attend/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md b/docs/architecture/attend/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md new file mode 100644 index 00000000..51523955 --- /dev/null +++ b/docs/architecture/attend/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md @@ -0,0 +1,156 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: attend +basis: + - evidence: 'issues #536 and #538: misuse classes seen in transcripts; the #538 audit found tool-description guidance had the highest adherence of any placement' + - standard: 'MCP protocol revision of 2026-07-28: no server path places content into the model''s turn' + - precedent: ADR-172 +agent: + name: Claude + model: unrecorded +status: proposed +date: 2026-09-19 +deciders: + - aaronsb + - claude +related: + - ADR-124 + - ADR-129 + - ADR-136 + - ADR-162 + - ADR-169 + - ADR-171 + - ADR-172 + - ADR-173 + - ADR-181 + - ADR-182 + - ADR-184 + - ADR-185 +imported: + from: docs/architecture/system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md + format: v0 + status: Proposed + unmapped: + amends: ADR-169#1 +--- + +# ADR-187: Attend MCP server mode: outbound and queries as typed tools, inbound stays on Monitor and the Stop hook + +## Context + +An agent session reaches attend through one surface today: CLI verbs run from the Bash tool. Inbound messages arrive through two conduits, the Monitor-hosted `attend run` sensor loop for an idle session and the Stop-hook drain `attend inbox --drain` for an active one (ADR-172). The skill and the way both state that the CLI is the contract (ADR-124, ADR-136), and the same messaging guidance is held in lockstep across three sources: `skills/attend/SKILL.md`, the sensor-disclosure reheat `messaging.md`, and the attend way. + +The lockstep exists because agents misuse the surface without repetition. Issues #536 and #538 name the classes seen in transcripts: + +- running `attend run` from Bash, which blocks the tool call and discards notifications; +- reading or editing under `~/.cache/attend/` and `~/.config/attend/` to work around an unclear verb; +- shell metacharacters in message bodies, which the "always double-quote" rule and the glob-expansion fence in `cmd/send.rs` exist to catch; +- envelope flags (`--to`, `--channel`, `--re`) that have to be remembered and ordered before the trailing message. + +The transcript audit behind #538 found that guidance carried in a tool's own description had the highest adherence of any placement. The three prose sources fire at session start, at reheat, and on the `attend` command trigger; a tool description is present whenever the tool is callable. + +The MCP protocol revision of 2026-07-28 sets the boundary. The initialize handshake and sessions are removed; sampling, roots, and logging are deprecated; server-initiated requests are replaced by multi-round-trip responses to client requests. The one unsolicited server-to-client channel is `subscriptions/listen`, which carries list-change and resource-change notifications to the host. No path in the protocol places content into the model's turn. ADR-172 listed "MCP push as the delivery mechanism" as a possibility to revisit; the revision closes it. + +Two adjacent threads shape the parameters. Issues #532 and #533 add structured envelope fields to every signal (`from_kind`, `from_id`, `on_behalf_of`, `to`), and the design note for them is merged at `docs/architecture/attend/ADR-401-attend-envelope-fields-sender-kind-principal-and-addressee-on-every-signal.md`. ADR-181 defines guard hooks as a PreToolUse class that ships deactivated, with default activation reserved for a further ADR under ADR-162's bar. The shipped `settings.json` allows `Bash(attend:*)` as one entry in the ADR-169 operational baseline, so an operator has no way today to allow queries while gating sends. + +## Decision + +**Attend runs as two roles. The perceiving role, the Monitor-hosted sensor loop and the Stop-hook drain, owns every piece of mutable session state. The acting role, the CLI message verbs and the MCP server, owns none and is a client of it. Outbound and queries become typed tools on an MCP server that shares one library with the CLI.** + +1. **The split is forced by the protocol.** Under the 2026-07-28 revision a server can answer a client request and can notify the host of a changed list or resource. Neither reaches the model mid-turn or wakes an idle session. The two conduits ADR-172 fixed therefore remain the only inbound paths, unchanged in mechanism and in contract. The MCP server carries the direction the protocol supports: the model calling out. + +2. **Tool set.** The server exposes `send`, `reply`, `peers`, `inbox`, `status`, `join`, and `leave`, named by the chat idiom ADR-173 chose. `run`, `inbox --drain`, `whoami --machine`, `tune`, `sensors`, `config`, `permissions`, and `cleanup` stay CLI; they are session infrastructure invoked by Monitor, by hooks, or by the operator (ADR-173 item 5). `keepwarm` (ADR-182), `scene`, `scenes`, `channels`, and `dissolve` stay CLI until use shows a reason to add them, and adding one follows item 3. `chat` stays CLI-only; it is a TUI with its own stdout. The deprecated `focus` alias of `--channel` (ADR-173) resolves in `cli.rs` before the library boundary, so the library and the tool schema know only `channel`. + +3. **One library, two frontends.** The verb bodies now under `tools/attend/src/cmd/` move into a library target with typed request and response structs. `cli.rs` and the MCP server are thin frontends over it. Each verb carries one description string, and three renderings derive from it: clap help, the generated `docs/cli/attend.md`, and the MCP tool schema. A test asserts the three agree. The JSON document a verb emits under `--json` (ADR-185) and the structured content of its tool result are one shape. + +4. **Shared state for `reply` and `inbox`.** `reply` threads to the record `sensor_peers::last_inbound` writes for the session. `inbox` lists every signal in the scan directories, as the CLI `inbox` list does today; only the drain consults the seen-set `attend-state` holds under ADR-172's merge semantics, and only the drain marks or filters by it. The MCP frontend reads the last-inbound record through the library and keeps no store of its own. The MCP `inbox` tool is read-only. + +5. **Transport, and every call is a fresh invocation.** The server runs as `attend mcp` over stdio, launched by the client once per session. It is resident, and that is the one property it does not share with the CLI verbs, so each tool call behaves as one CLI invocation would: it derives the session id by the ADR-171 ancestry walk, verifies that a registration for that id exists in the instance registry the sensor loop maintains, adopts the registered tuple (persona and ordinal), reads the last-inbound record and channel membership from disk, performs the verb, and holds nothing in memory for the next call. No memoized tuple, no inbox cursor, no cached membership. The sensor loop can restart under Monitor's cap and re-register, and the next call sees the new registration. `status` reports the derived id, the registration found for it, and the `resolved` flag as read at call time. When ancestry resolves to a Claude session and no registration exists for it, because the loop has not registered yet or is not running, `send` and `reply` refuse with a cause naming `attend run`, matching the drain's refusal in ADR-172 item 4, so a persona the loop never allocated is never written to the bus. + +6. **Permissions are per tool.** Each tool has its own permission name: `mcp__attend__send`, `mcp__attend__reply`, `mcp__attend__peers`, `mcp__attend__inbox`, `mcp__attend__status`, `mcp__attend__join`, `mcp__attend__leave`. The shipped `settings.json` allows all seven, which preserves the autonomy the skill grants today under `Bash(attend:*)`. The per-tool names are the operator's control over the MCP path: removing `mcp__attend__send` puts that tool behind a prompt for a session running an approval rule. While both frontends ship, `Bash(attend:*)` stays allowed for the CLI frontend, so the per-tool entry alone gates the MCP path only; a hard outbound gate also needs `Bash(attend send*)` and `Bash(attend reply*)` in `permissions.deny`. An unattended run confirms the seven entries are present, since a headless session has no prompt and an unallowed call fails. The ADR-184 reconciler registers the server in each enabled target and withdraws it with the target. The seven allow entries and the `mcpServers` block widen the ADR-169 operational baseline; this ADR amends ADR-169 item 1 to include them, written by the same three-way merge as the rest of the baseline. + +7. **Where guidance lives.** The mechanical layer moves into the tool descriptions and parameter schemas: the send-versus-reply choice, on `send` ("starts a topic; use `reply` when answering the most recent peer message") and on `reply` ("threads to the most recent peer message; errors when the inbox is empty"), the scope parameters, the length budget, and `reply`'s empty-inbox behavior. The judgment layer stays in the way and the skill: autonomy to reply without asking, silence as a valid reply, and the two-conduit contract. Those sentences govern whether to call any tool at all, and a description is read once the model is already choosing one. The reheat shrinks to the judgment layer and the conduit contract. The quoting rule and the "never run `attend run` from Bash" line drop from the reheat for a session with the server connected; the server writes its own presence marker, a file under attend-state keyed on the session id, on start and removes it on exit, and the reheat reads that file to select the wording. The instance record stays sensor-loop-owned (item 11). The ban on reaching into attend-owned directories stays as one sentence in the way and the skill, because Read, Edit, and Write remain available to the model. + +8. **The interim for CLI-only sessions.** The two PreToolUse denies #536 proposes ship as guard hooks under ADR-181's contract: exit 0 or 2, a closed list with a repair per entry, deactivated by default, tested in the suite. They live under `hooks/ways/`, the guard-class convention ADR-181's `check-bash-bound.py` set, as `hooks/ways/attend-guard-pre.sh`, and `attend permissions` reports whether they are wired. Neither misuse meets ADR-162's bar for default activation; a blocked Monitor launch and a read of a cache directory are recoverable and disclose nothing. A separate attend-owned activation switch would be a second activation state beside ADR-184's targets, and this ADR declines to add one. The scripts are marked interim in their headers and are removed when the MCP frontend is the default for agent sessions. + +9. **Envelope fields are typed parameters.** `send` and `reply` take `message: string` and the optional `channel`, `to`, `broadcast`, and `on_behalf_of`. `from_kind` and `from_id` are set by the server from the connection identity and are never parameters, per #532. `inbox` and `peers` results carry the same fields as typed members of each row. The tool's `to` takes the grammar the envelope design note fixes: `*` for broadcast, `#<name>` for a channel, or a comma-separated set of canonical ids; a project path resolves at send time to the session ids of the live peers at that path, and that resolution is how the CLI's `--to PATH` maps into the same field. The field semantics, the id classes, and the wire format belong to the design note; this ADR fixes only that the fields cross the tool boundary as schema. + +10. **The contract statement changes.** "CLI is the contract" becomes: the attend tool surface is the contract, meaning the MCP tools where the server is connected and the CLI verbs everywhere; attend-owned paths remain implementation detail. + +11. **Ownership by role.** The sensor loop is the one process per session that holds the ADR-129 duplicate lock, registers and touches the instance record, writes the seen-set, writes the last-inbound record, and supplies the pid that liveness checks. An acting frontend, whether a one-shot `attend send` or the resident `attend mcp`, takes none of these. It appends signal files, writes channel membership on `join` and `leave`, and reads everything else. Identity is derived by both processes and registered by one: each derives the session id by ADR-171 ancestry; the sensor loop registers it and allocates the persona and ordinal (ADR-129); an acting frontend derives the same id, verifies the registration exists, and adopts the registered tuple. When ancestry resolves to a Claude session and no registration exists, because the loop has not registered yet or is not running, `send` and `reply` refuse with a cause naming `attend run`. The refusal is scoped to senders whose ancestry resolves to a Claude session: an agent session without its sensor loop cannot send, as it already cannot receive. A process with no Claude ancestry, a plain terminal, keeps the external path `identify_sender` in `cmd/send.rs` takes today (`$USER@<terminal>`, source kind external, which the envelope note renders as `from_kind: human`), and attend-chat's human sender is unchanged. The MCP server is driven by a model by construction, so when its own ancestry resolves to no session it refuses `send` and `reply` and the external path stays with the CLI. The two processes therefore never contend for a lock, never register twice, and never disagree on who the session is. Item 5 makes the resident server equivalent to the one-shot CLI: same reads at the same moment, nothing carried between calls. The library extraction in item 3 follows the same rule: verb bodies return values and frontends print, so nothing in the shared code writes to the stdout an MCP transport owns. + +| State | Owner | Acting frontends | +|---|---|---| +| Duplicate lock (ADR-129) | sensor loop | never taken | +| Instance record and liveness pid | sensor loop | read | +| Seen-set (ADR-172) | sensor loop and drain | never touched | +| Last-inbound record | delivering conduit | read (`reply`) | +| Signal files | appended by acting frontends | append | +| Channel membership | acting frontends (`join`, `leave`) | atomic file write, as the CLI does today | +| MCP presence marker (under attend-state, keyed on session id) | acting frontend (`attend mcp`) | write on start, remove on exit | + +12. **One state contract, then a trial.** The MCP verbs and the CLI message verbs keep one state contract: the same verb against the same disk state leaves the same disk state, whichever frontend ran it. A contract test runs each verb through both frontends against one seeded state directory and diffs the result. Once implemented, agent sessions switch to the MCP frontend for daily use while the CLI message verbs stay in the binary for comparison. What follows the trial is a later decision on the evidence of use; this ADR does not remove the CLI message verbs. + +Reversibility: reversible. The CLI frontend stays whole. Removing the MCP frontend deletes the `mcp` subcommand and the settings entries. The library extraction in item 3 stands on its own merits and would remain. + +```mermaid +flowchart LR + Model[Model turn] + Model -->|tool call| MCP[attend mcp<br/>send reply peers inbox status join leave] + MCP --> Lib[(attend library<br/>one verb body per verb)] + CLI[attend CLI<br/>run, inbox --drain, whoami ...] --> Lib + Lib --> State[(seen-set, last_inbound,<br/>signal files)] + State -->|poll| Monitor[Monitor: attend run] + State -->|turn end| Stop[Stop hook: attend inbox --drain] + Monitor -->|notification| Model + Stop -->|injected text| Model + + classDef proc fill:#2d7d9a,color:#fff,stroke:#4a5568 + classDef store fill:#2d8e5e,color:#fff,stroke:#4a5568 + classDef model fill:#7c3aed,color:#fff,stroke:#4a5568 + class MCP,CLI,Monitor,Stop proc + class Lib,State store + class Model model +``` + +## Consequences + +### Positive + +- For a session with the server connected, three misuse classes end structurally: there is no Bash surface for `attend run`, a typed `message` string has no shell to expand it, and the envelope is schema the client validates. +- The send-versus-reply guidance sits in the description of the tool being chosen, the placement the audit found agents follow most. +- An operator gains a per-tool gate on the MCP path. One permission entry separates "may read the bus" from "may write to it" there, which the single `Bash(attend:*)` entry could never do. While the CLI frontend ships beside it, a hard outbound gate adds two Bash deny entries, for `attend send` and `attend reply`. +- The CLI and MCP frontends cannot drift in behavior, because each verb has one body. The MCP `reply` threads to the same message the drain delivered, by construction. +- The reheat's disclosure budget shrinks to the judgment sentences. + +### Negative + +- Seven tool schemas sit in the context of every request for the session. Descriptions have to stay short, and the skill's messaging section has to shrink by at least as much. +- The library extraction is a refactor of `cmd/` that lands before any new capability, and two frontends mean two test surfaces over one body. +- A session without the server registered keeps every current failure mode. The guard hooks are opt-in, so the interim protection for those sessions is guidance unless the operator wires one settings line. +- The directory misuse loses only its Bash path. Read, Edit, and Write still reach `~/.cache/attend/`; the deactivated guard is the mechanical answer, and the one-sentence ban is the shipped one. +- Identity from pid ancestry assumes the client spawns the server in the session's process tree. A host that launches servers elsewhere gives the server no Claude ancestry, and `send` and `reply` refuse until it is fixed; the external sender path stays available through the CLI in a plain terminal. +- An agent session whose sensor loop has not registered yet cannot send through either frontend until it does. The refusal names `attend run`, and the first seconds after `/attend` are the window where it fires. + +### Neutral + +- ADR-172's open item on MCP push is closed by the protocol revision; this ADR records the closure. +- A presence marker file under attend-state, written and removed by the server, tells the reheat, the drain footer, and `attend status` that the MCP frontend is connected. The instance record is unchanged. +- The ADR-169 operational baseline grows by the seven `mcp__attend__*` allow entries and the `mcpServers` block, recorded as an amendment to ADR-169 item 1 in this ADR's frontmatter. +- #536 becomes moot for MCP sessions and survives as the deactivated guard for the rest. +- The `attend chat` TUI, `/purge`, scenes, and the signal file format are untouched. ADR-136's durability rules and ADR-173's vocabulary carry over as they stand. +- The three-source lockstep becomes a two-layer rule: one description source with three renderings for the mechanical layer, and the way plus the skill for the judgment layer, with the reheat quoting the judgment layer. The lockstep comment in each file is rewritten to say so. +- The skill's `allowed-tools` gains `mcp__attend__*`, and its pre-flight checks that the server is registered before choosing wording. + +## Alternatives Considered + +- **MCP for inbound as well, with the server delivering messages.** Unavailable. Under the 2026-07-28 revision the server can respond to a client request or notify the host of a changed list or resource. Neither reaches the model's turn or wakes an idle session. The Monitor line and the Stop-hook injection are the only paths that do. +- **Keep the CLI as the only frontend and add the #536 guards plus more reheat text.** Rejected. The guards cover the two mechanical misuses and leave the quoting and flag classes untouched, and the disclosure tax that motivated the audit continues. +- **An MCP server that shells out to the CLI.** Rejected. Results would arrive as text to parse, errors as exit codes to map, and the tool schema would be a hand-kept copy of the clap surface. The library call is typed and shares the description source. +- **Move every verb, including `run` and `inbox --drain`, into the server.** Rejected. `run` exists to feed Monitor's notification bridge, and the drain has to be invoked by the Stop hook at the turn boundary. Both are inbound machinery with no client-request shape. +- **One long-lived server per user over a network transport.** Deferred. Per-connection identity would need the client to declare its session, where a stdio child gets it from ancestry. Item 5's fresh-invocation rule, with identity derived and verified on every call, leaves the door open without paying for it now. +- **Activate the #536 guards by default.** Rejected under ADR-181: neither misuse produces an irreversible or disclosing outcome. +- **An attend-owned activation switch for the guards, such as a `permissions` sub-verb.** Rejected. ADR-184 makes the target the unit of activation; a second switch beside it would be a state the reconciler does not model. +- **Leave the three prose sources as they are and add the tool descriptions on top.** Rejected. That makes a fourth copy of the mechanical guidance and preserves the drift the lockstep comment exists to warn about. diff --git a/docs/design-notes/attend-messaging-disclosure-reheat.md b/docs/architecture/attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md similarity index 96% rename from docs/design-notes/attend-messaging-disclosure-reheat.md rename to docs/architecture/attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md index f6aabbd0..b76fe7fa 100644 --- a/docs/design-notes/attend-messaging-disclosure-reheat.md +++ b/docs/architecture/attend/ADR-400-attend-messaging-disclosure-with-token-gated-reheat.md @@ -1,7 +1,16 @@ -## Attend: Messaging Disclosure with Token-Gated Reheat +--- +contract: adr/v1 +kind: evidence +capability: attend +status: accepted +date: 2026-04-13 +deciders: + - aaronsb +related: [] +--- + +# ADR-400: Attend messaging disclosure with token-gated reheat -> **Type:** Design note (not an ADR) -> **Status:** Working draft, subject to revision > **Cites:** ADR-104, ADR-113 > **Motivates:** ADR for attend disclosure registry *(planned, draft after working sketch)* @@ -105,6 +114,6 @@ Work proceeds on the existing `feat/attend-messaging-reheat` branch as: ## References -- [ADR-104](../architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Token-gated way re-disclosure (the model this note reuses) -- [ADR-113](../architecture/system/ADR-113-attend-active-awareness-module.md) — `attend`: active awareness module -- [Cognitive loop and the awareness layer](./cognitive-loop-and-awareness-layer.md) — the broader frame this work sits within +- [ADR-104](../ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Token-gated way re-disclosure (the model this note reuses) +- [ADR-113](ADR-113-attend-active-awareness-module.md) — `attend`: active awareness module +- [Cognitive loop and the awareness layer](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) — the broader frame this work sits within diff --git a/docs/design-notes/attend-envelope-fields.md b/docs/architecture/attend/ADR-401-attend-envelope-fields-sender-kind-principal-and-addressee-on-every-signal.md similarity index 99% rename from docs/design-notes/attend-envelope-fields.md rename to docs/architecture/attend/ADR-401-attend-envelope-fields-sender-kind-principal-and-addressee-on-every-signal.md index 7a3a3a73..cbd69a22 100644 --- a/docs/design-notes/attend-envelope-fields.md +++ b/docs/architecture/attend/ADR-401-attend-envelope-fields-sender-kind-principal-and-addressee-on-every-signal.md @@ -1,4 +1,15 @@ -# Attend envelope fields: sender kind, principal, and addressee on every signal +--- +contract: adr/v1 +kind: spec +capability: attend +status: accepted +date: 2026-09-19 +deciders: + - aaronsb +related: [] +--- + +# ADR-401: Attend envelope fields: sender kind, principal, and addressee on every signal A reading of the attend signal wire format, taken 2026-09-19, that settles the structured envelope fields issues #532 (sender kind and on-behalf-of) and #533 (addressee) ask for. The two issues share one on-disk change and one set of readers, so this note specifies them together for a single implementation. It then states what #535 (age- and party-aware drain) reads from the fields and which parts of ADR-172 that amends. Signing and cross-machine relay stay out of scope; the last section says why these fields are still their prerequisite. #538 (attend as an MCP server for outbound) is covered where the field design touches tool parameters. diff --git a/docs/architecture/documentation/ADR-106-project-pulse-epoch-mapped-project-awareness.md b/docs/architecture/documentation/ADR-106-project-pulse-epoch-mapped-project-awareness.md new file mode 100644 index 00000000..5a3aeab6 --- /dev/null +++ b/docs/architecture/documentation/ADR-106-project-pulse-epoch-mapped-project-awareness.md @@ -0,0 +1,124 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +basis: + - evidence: a manual audit found 8 of 12 ADRs with incorrect statuses + - evidence: the 200K to 1M context window change was discovered ad hoc rather than surfaced from upstream releases (ADR-103, ADR-104) +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-20 +deciders: + - aaronsb + - claude +related: + - ADR-103 + - ADR-104 +imported: + from: docs/architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md + format: v0 + status: Accepted +--- + +# ADR-106: Project Pulse — Epoch-Mapped Project Awareness + +## Context + +Claude Code releases frequently — often daily. Each release may add features, change behaviors, or introduce capabilities that affect how this project (claude-code-config) should be built. Without a systematic way to compare upstream changes against our own work, we either miss opportunities or discover them accidentally. + +Separately, this project's own ADRs can drift from reality. An ADR may be Accepted but never implemented, or Draft while the code has already shipped. We discovered this during a manual audit where 8 of 12 ADRs had incorrect statuses. The same epoch-mapping logic that compares us against upstream can compare our stated intentions (ADRs) against our shipped code (commits, branches, PRs). + +The naive approach — projecting both projects onto a shared calendar and comparing dates — breaks down because commits and releases are epoch counters, not time events. A day with 12 commits isn't "more" than a day with 1 commit. The meaningful relationship is ordinal-causal: at the time of our epoch N, which Claude Code epoch was current, and how many upstream epochs passed between our epochs N and N+1? + +Previous examples of this gap: the jump from 200K to 1M context window required rethinking our epoch tracking and compaction strategies (ADR-103, ADR-104). That change was discovered ad-hoc rather than surfaced systematically. + +This project also has a velocity mismatch: Claude Code ships tagged releases nearly daily; we commit frequently but tag releases irregularly. The comparison tool must handle both granularities. + +## Decision + +Build a project awareness tool (`scripts/project-pulse`) and skill (`skills/project-pulse/SKILL.md`) with two modes: + +### Upstream mode (default) + +Compare Claude Code releases against our project's commits to surface what's new upstream and what might matter for us. + +### Inward mode (`--inward`) + +Compare our ADRs against our own commits, branches, and PRs. Searches for ADR numbers (e.g., "ADR-106") in branch names, commit messages, and PR descriptions. Flags mismatches: +- ADRs still Draft/Proposed but with substantial implementing code +- ADRs marked Accepted but with no referencing commits +- ADRs with no implementation activity at all (dormant) + +### Data model + +Treat both projects as epoch streams. Each epoch is a commit or release with an ordinal position and a timestamp (metadata, not primary key). The core data structure maps our epochs to the upstream epoch that was current at that moment, plus a delta showing how many upstream epochs passed since our previous epoch. + +For inward mode, the mapping is between ADR creation/status-change epochs and the commits that reference each ADR. + +### Feathered windowing + +The default window starts from our most recent release tag (the anchor), includes all commits since that anchor, and pulls the corresponding Claude Code releases spanning the same period — plus 1-2 extra releases on each side for context bleed. This "feathered" approach avoids hard date cutoffs. + +Overrides: +- `--since DATE` — everything from a date onward +- `--range DATE DATE` — specific window +- `--full` — entire history + +### Two-piece architecture + +1. **Script** (`scripts/project-pulse`) — data gathering. In upstream mode: fetches Claude Code releases via `gh api repos/anthropics/claude-code/releases`, our commits/tags via `git log`, builds the epoch mapping. In inward mode: scans ADR statuses, searches git log and branch names for ADR references, builds the reconciliation. Outputs structured markdown in both modes. + +2. **Skill** (`skills/project-pulse/SKILL.md`) — interpretation. Reads the script output and the project's README/ADR index to understand what we care about. In upstream mode: filters changes through the project's charter (hooks, settings, skills, context window, plugins, subagents, permissions, MCP), produces 2-5 suggestions in plain prose. In inward mode: highlights status mismatches and dormant ADRs, suggests corrections. + +### Tone and intent + +The skill output is a discovery conversation starter, not a compliance dashboard. No coverage scores, no red/yellow/green, no "you're N epochs behind." It reads like a colleague saying "did you see they added X? That might matter for us" or "ADR-100 has been Draft for a while — is that still the plan?" The user decides what to act on. + +### Watermark + +After each run, the tool writes a lightweight marker (date + our HEAD + upstream latest release) so the next default invocation knows where to start. + +### ADR provenance + +When upstream changes inspire new work, the resulting ADRs can cite the specific Claude Code release version as a reference (e.g., "Inspired by Claude Code v2.1.80 `effort` frontmatter support"). + +### Meta way integration + +A new `meta/project-health` way triggers on keywords related to upstream tracking, project status, and ADR reconciliation. It suggests running the project-pulse tool when relevant, and provides guidance on managing claude-code-config as a project — not just authoring its components. + +## Consequences + +### Positive + +- Systematic awareness of upstream changes without manual changelog reading +- ADR status reconciliation catches drift between intent and implementation +- Epoch-based comparison handles velocity mismatches between the two projects +- Conversational output avoids FOMO treadmill — user stays in control +- ADR provenance links our decisions to the upstream features that inspired them +- Feathered windowing ensures context bleed at boundaries isn't lost +- Single tool serves both outward (upstream) and inward (self) awareness + +### Negative + +- Depends on `gh` CLI and GitHub API access to anthropics/claude-code +- Skill interpretation quality depends on the project's README/ADR index being reasonably current +- Watermark file is another piece of state to maintain +- Inward mode depends on consistent ADR number references in commits/branches (a convention we need to maintain) + +### Neutral + +- Encourages more consistent release tagging in our project (better anchor points) +- Encourages referencing ADR numbers in branch names and commits (better traceability) +- The epoch mapping data structure could be reused by other tools (e.g., checks, ways staleness detection) +- The meta way creates a self-referential loop: claude-code-config has a way about managing claude-code-config + +## Alternatives Considered + +- **Calendar-based comparison (interleaved timeline)**: Simpler to implement, but misrepresents the relationship between projects. Commits aren't time events — they're ordinal epochs. Calendar projection creates false "gaps" when one project is quiet. +- **Disposition ledger (ADOPTED/WATCHING/IRRELEVANT per feature)**: Adds a persistent state file that must be maintained. The epoch mapping plus feathered window achieves retroactive lookback without ongoing bookkeeping. +- **Website scraping**: The changelog at code.claude.com/docs/en/changelog has the same content as GitHub releases but requires HTML parsing and is more fragile. The `gh api` approach is structured, authenticated, and paginated. +- **Full-history comparison every time**: Context-expensive and noisy. The feathered window gives relevant scope by default while `--full` remains available for archaeology. +- **Separate tools for upstream vs inward**: The data model (epoch streams, feathered windows) is the same in both directions. Splitting into two tools would duplicate the windowing logic and force the user to remember two commands. Modes on one tool are simpler. diff --git a/docs/architecture/documentation/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md b/docs/architecture/documentation/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md new file mode 100644 index 00000000..d929785b --- /dev/null +++ b/docs/architecture/documentation/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md @@ -0,0 +1,125 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - adr + - docs + - disclosure +basis: + - evidence: 'the adr way''s macro diff -q is direction-blind: a stale copy and a customized copy produce the same byte difference, and its note reassures in the stale case' + - evidence: 'issue #438: vendored copies fall behind upstream tool features such as adr archive' + - precedent: ADR-138 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-08-06 +deciders: + - aaronsb + - claude +related: + - 138 +imported: + from: docs/architecture/system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md + format: v0 + status: Accepted +--- + +# ADR-177: Version-stamped vendored tools with direction-aware drift detection + +## Context + +Three tools ship inside ways and get **vendored** — copied, never symlinked — +into consuming repositories: `adr-tool` (documentation/adr), `doc-tool` +(documentation/linting), and `chart-tool` (softwaredev/visualization/charts). +ADR-138 put the vendoring procedure in skills; the copies it produces are +snapshots that never hear about upstream changes. + +None of the three carries a version marker. The one drift check that exists — +the adr way's macro `diff -q` between the repo copy and the installed +template — is **direction-blind**: a stale copy and a deliberately customized +copy produce the same byte difference, and the macro's note ("differs from the +universal template; expected for customized setups") actively reassures the +reader in the stale case. As upstream tools grow features (e.g. `adr archive`, +issue #438), every vendored copy silently falls behind while the disclosure +says all is well. + +## Decision + +**Every vendorable tool carries its own version, and drift disclosure reads +direction from it.** + +1. **Per-tool semver constant** — a `TOOL_VERSION = "X.Y.Z"` line near the top + of each tool, plus a `--version` flag. The version is per-tool and bumps + only when the tool changes. It is deliberately *not* the ways release + version: coupling to releases would mark every vendored copy stale on every + release even when the tool never moved. + +2. **Direction-aware macro disclosure** — a way macro that surfaces a vendored + tool extracts `TOOL_VERSION` from both the repo copy and the installed + template (a `grep`, not an execution) and compares with `sort -V`. Five + states replace the direction-blind diff note: + + | Local copy vs installed | Disclosure | + |---|---| + | no version marker, installed stamped | predates versioning — out of date; re-vendor | + | lower than installed | stale — re-vendor to pick up the newer tool | + | equal, bytes differ (version line excluded) | customized — expected, say so | + | higher than installed | repo is ahead — the agent-ways install is stale; update it | + | stamped, installed unversioned | same as above — the install is stale | + + The version line is excluded from the customization diff so the stamp + cannot trigger the very warning it exists to disambiguate. The extracted + stamp is shape-restricted to a version string — a stamp that doesn't parse + as one is treated as unversioned, never echoed into disclosed context. + +3. **Re-vendoring is a documented move** — the skill that owns each tool's + vendoring procedure (ADR-138) also documents the update: plain copy when + the local tool is unmodified; when customized, diff first so local changes + are carried forward rather than clobbered. + +The unversioned era is bounded by the rule itself: a copy with no +`TOOL_VERSION` marker is by definition older than every stamped release, so +"no marker" needs no special casing beyond "out of date." + +## Consequences + +### Positive + +- Drift disclosure gains direction: stale, customized, and ahead-of-install + are distinct states with distinct remedies, instead of one reassuring note. +- Tool changes (like #438's `adr archive`) propagate — the next session in a + consuming repo is told its copy is behind, rather than nobody ever noticing. +- "Repo ahead of install" is surfaced too, catching the stale-install case the + old diff attributed to customization. + +### Negative + +- Version bumps are a manual discipline — a tool change without a bump defeats + the mechanism. Mitigated by review: the stamp lives in the same file as any + change to it. + +### Neutral + +- Requires stamping all three tools, upgrading the adr way's macro, and adding + update guidance to the owning skills; other tool-surfacing macros adopt the + same comparison as they grow one. +- Already-vendored copies in the wild have no marker and will read as "out of + date" on first contact with the new macro — which is accurate. + +## Alternatives Considered + +- **Keep the byte-diff note.** Rejected: direction-blind. It cannot + distinguish the case that needs action (stale) from the case that needs none + (customized), and its wording suppresses the one signal it does emit. +- **Stamp tools with the ways release version.** Rejected: false staleness — + every release would mark every vendored copy out of date even when the tool + is byte-identical. +- **Checksum registry (manifest of known tool hashes per version).** Rejected: + heavier machinery for the same answer; requires the installed side to carry + history, and any local edit breaks the lookup entirely — exactly the + customized case the mechanism must keep distinguishable. +- **Auto-update on detection.** Rejected: a vendored copy may be deliberately + customized; overwriting on sight destroys local changes. Disclosure names + the state; the operator (or a skill-guided session) decides. diff --git a/docs/architecture/documentation/ADR-300-documentation-structure.md b/docs/architecture/documentation/ADR-300-documentation-structure.md index 2a51d247..24c8ee18 100644 --- a/docs/architecture/documentation/ADR-300-documentation-structure.md +++ b/docs/architecture/documentation/ADR-300-documentation-structure.md @@ -1,5 +1,15 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: docs +basis: + - evidence: 'doc-graph.sh audit: 41 doc files, 29 links, 34 dead ends, 23 orphans' + - evidence: 'docs/audit-findings.md: 5 files present gzip NCD as primary where BM25 is the implementation, and 6 content areas are duplicated' +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-02-17 deciders: - aaronsb @@ -8,6 +18,10 @@ related: - ADR-004 - ADR-005 - ADR-014 +imported: + from: docs/architecture/documentation/ADR-300-documentation-structure.md + format: v0 + status: Accepted --- # ADR-300: Documentation Structure diff --git a/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md b/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md index 1164513f..29ab8b26 100644 --- a/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md +++ b/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md @@ -1,10 +1,26 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: + - docs + - method +basis: + - evidence: ablation testing without the ways system produces approval-seeking agents across model tiers + - evidence: project mechanisms are described only in invented vocabulary, defined by reference to each other +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-06-09 deciders: - aaronsb - claude related: [] +imported: + from: docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md + format: v0 + status: Accepted --- # ADR-301: Situated socialization as canonical framing and documentation prose refactor diff --git a/docs/architecture/documentation/ADR-302-unified-documentation-model.md b/docs/architecture/documentation/ADR-302-unified-documentation-model.md index 5ee51741..93da1148 100644 --- a/docs/architecture/documentation/ADR-302-unified-documentation-model.md +++ b/docs/architecture/documentation/ADR-302-unified-documentation-model.md @@ -1,5 +1,16 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: docs +basis: + - upstream: knowledge-graph-system ADR-908 (documentation strategy) and ADR-900 (domain numbering), with its running doclint.py + - standard: 'Diátaxis: four modes, closed 2x2' + - precedent: ADR-300 +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-06-19 deciders: - aaronsb @@ -7,6 +18,10 @@ deciders: related: - ADR-300 - ADR-301 +imported: + from: docs/architecture/documentation/ADR-302-unified-documentation-model.md + format: v0 + status: Accepted --- # ADR-302: A unified documentation model — typed graph, ways packaging, cross-repo convergence @@ -282,5 +297,3 @@ Resolved during drafting: this ADR stays **ADR-302 / `documentation` domain** (i subject is documentation tooling); the file is renamed to `ADR-302-unified-documentation-model.md`; frontmatter adopts Obsidian-compatible wikilink edges + `aliases` (§2, §5, §6). -</content> -</invoke> diff --git a/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md b/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md index b941f3f5..3e6a2e10 100644 --- a/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md +++ b/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md @@ -1,5 +1,14 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: adr +basis: + - evidence: 'issue #438: 85 ADRs on disk, nine supersedes/superseded_by declarations, and zero code reading any of it' +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-08-06 deciders: - aaronsb @@ -7,6 +16,10 @@ deciders: related: - 177 - 302 +imported: + from: docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md + format: v0 + status: Accepted --- # ADR-303: Active-set semantics, adr archive, and supersession reading for the ADR corpus diff --git a/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md b/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md index 92b635f3..d30b211e 100644 --- a/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md +++ b/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md @@ -4,6 +4,8 @@ kind: decision verb: change capability: adr amends: [ADR-304#6] +superseded_by: + - ADR-311 basis: - operator: aaronsb level: guided @@ -27,7 +29,7 @@ considered: said: "ok. so basically, I think it's the correct direction and implements the change to adr as discussed." via: session 2026-09-27, reviewing PR #583 covers: [] -status: accepted +status: superseded date: 2026-09-27 deciders: - aaronsb diff --git a/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md b/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md new file mode 100644 index 00000000..a162d6d5 --- /dev/null +++ b/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md @@ -0,0 +1,213 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-304#7] +basis: + - operator: aaronsb + level: guided + said: "I think we need to make sure that the new adr tools can create the content and manage the lifecycle of data, and probably, we need to think about a 'foreign import' tool that would take any kind of decision record that's not directly lintable/usable, and can ingest the foreign record. this way, it becomes our cannonical 'migration' tool." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "I think the foreign import model needs a round trip data object template of some kind." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "we should assuem that reasonable foregin records have some sort of structured data (frontmatter, for example). a /very foreign/ import could literally be jira issues for example" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "I think optional, but always lint warnings. sometimes, the summary isn't obvious until the apparent motion of the complete dataset is visible. this means that adr record properties can be altered (like summaries). this is fine, because any alteration that gets tracked is a git commit. we don't have to overthink integrity here" + via: session 2026-09-27, on whether imported records need a Summary + - operator: aaronsb + level: guided + said: "we don't need to explicitly handle jira. all I'm saying is 'jira issues can be flattened to a record, just like any other record, and usually there's a description and a summary and various fields, and if we can selectively import jira issues, then we probably can take records from about anything'" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "the domain layout shouldn't ever be set forever. things change over time, certain domains might merge or split. during import from v0 adr to v1 in-repo, it might make sense to add or combine domains. during a full foreign import, perhaps something like a jira or github issues, then these would expand over time." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "it might be more work to curate the records, but if an adr changes domains, then I think it needs to be changed in code." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "or in references" + via: session 2026-09-27, following the message above + - operator: aaronsb + level: guided + said: "yes we trade away hand migrations freedom to restructure a record, but that feels like a forced decision. there's nothing stopping us from transforming the record before import. import is just the acceptance model for a foreign record" + via: session 2026-09-27, reviewing ADR-306 + - evidence: "this repo holds 93 v0 records, each with v0 frontmatter (status, date, deciders, related) that a reader can map without judgement" + - evidence: "in a project that declares contract: adr/v1, `adr new` still writes a v0 record with no contract, kind, verb, capability, basis, agent or Summary" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "i read the adr and it aligns with my understanding. let's accept and merge it" + via: session 2026-09-27, PR #587 + covers: [] + - operator: aaronsb + said: "my assumption is that everything needed is carried in the docs. but because reality can drift and is complex, its possible (and we should assume it happens) that the actual implementation nearly always has some drift from the record of desire." + via: session 2026-09-27, answering the probes after acceptance + covers: [frontmatter-carries-all] + - operator: aaronsb + said: "we import markdown for text. any complex formatting language needs to be markdown. we import a structured body for structured data - I'm not sure what the convention is right now, but yaml or json seems to be the right approach." + via: session 2026-09-27, answering the probes after acceptance + covers: [non-markdown-bodies] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 305 +--- + +# ADR-306: adr import: foreign records through a round-trip import sheet + +## Summary + +- **Decided:** `adr import` is the acceptance model for a foreign record: whatever shape a record arrives in, it becomes an adr/v1 record through an import sheet. `scan` reads records from any structured source into one import sheet per record. The agent fills in what needs judgement. `apply` writes each finished sheet as a v1 record. `adr new` writes through the same writer, and `adr supersede` and `adr enact` complete the lifecycle commands. Domains can be added, merged and split as the corpus grows. A record that moves to another domain is renumbered into that domain's range, and every reference to it, by number or by path, is rewritten. +- **Trades away:** a direct edit from source to record. Every record passes through a sheet, a staging format with its own schema, and other sources through field maps. Both have to be documented and kept stable. +- **One-way?** No. Sheets are staging files and records stay in git. A bad import is reverted like any commit. +- **Probes:** *Confident:* v0 records from agent-ways, here or in any repo that adopted them, import with the body unchanged, since everything the reader needs is in the frontmatter. *Not confident:* whether a field map that flattens a structured item into fields and a body covers sources whose body isn't markdown, or whether some sources need a conversion step first. +- **Inversion:** at one end, a reader written in code for every format: exact, but never finished. At the other end, an agent reads each foreign record and writes v1 by hand: flexible, but manual across a hundred records. This decision maps structured fields mechanically and leaves only the judgement fields to the agent. + +## Context + +ADR-304 §7 moves a v0 record to v1 "when someone next edits it". That works for a trickle of edits, and it doesn't work for a corpus. This repo holds 93 v0 records. Any repo using a file-based record system, or v0 records from agent-ways, holds a corpus of its own, in v0, adr-tools, MADR or tracker formats. #581 migrated two records here by hand, and the tier 2 rehearsal showed an agent can do it without inventing anything, but record by record. + +Most of a migration is mechanical. Status maps by the §7 table, and date, deciders and links carry over. The body stays as written. A few fields need judgement: the verb, the capability, the basis, and a Summary. Those are the only fields an agent should have to touch. + +The tool also has lifecycle gaps. `adr new` writes a v0 record in a v1 project. Supersession needs both sides' links edited by hand, and enactment is a hand-edited field. + +## Decision + +### 1. The import sheet + +Import accepts a record. It does not restrict what happens to the record before it's accepted. A record can be split, merged or rewritten before `scan`, or its sheet edited before `apply`. The importer itself never changes content: whatever the sheet says is what `apply` writes. + +One YAML file per record is the round-trip object between a source and a v1 record: + +```yaml +sheet: adr-import/v1 +source: {path: docs/architecture/system/ADR-186-….md, format: v0, sha256: "…"} +target: {number: 186, domain: system} +record: # v1 frontmatter, filled as far as the reader can + contract: adr/v1 + kind: decision + status: accepted + date: 2026-09-17 + verb: ~ + capability: ~ + basis: [] + agent: {name: Claude, model: unrecorded} +summary: ~ +todo: [verb, capability, basis] +candidates: {capability: [testing, install]} +provenance: {status: "frontmatter status: Accepted"} +unmapped: {deprecation_note: "…"} +body: | + … +``` + +- `record` holds v1 frontmatter. The reader fills what the source states, and `provenance` says where each value came from. +- `todo` lists what the reader could not fill. `candidates` ranks vocabulary matches to help whoever fills them. +- `unmapped` keeps every source field that has no v1 home. Nothing is dropped. +- `body` starts as the source body, verbatim. It can be edited like any other part of the sheet. + +### 2. Readers + +A reader turns a source into sheets. Import assumes structured sources: frontmatter, a metadata block, or a structured export. Unstructured prose is out of scope. + +- **Built in:** v0 (this tool's frontmatter), MADR, and adr-tools (inline `## Status`, `0001-` numbering). +- **Field maps:** any other structured source is flattened into fields and a body. A declarative map names which source field fills which sheet field, and how values translate. No source gets its own reader. The example is a tracker export, since a source like that flattening cleanly suggests most structured records will: + +```yaml +reader: tracker +items: issues # a JSON export: one sheet per item, selected by --filter +fields: + title: fields.summary + date: fields.created + status: {from: fields.status.name, map: {Done: accepted, "Won't Do": rejected, "To Do": proposed}} + body: fields.description + unmapped: [key, fields.labels] +``` + +No reader ever writes an `operator` basis from `deciders`, an assignee or any other metadata (ADR-304 §7, §11). A basis comes from what the record says, and it is filled during cleanup. + +### 3. Commands + +- `adr import scan <paths> [--reader NAME | --map FILE]` writes sheets to `docs/architecture/.import/`. That directory is gitignored: sheets are working files, and only the records they produce are committed. +- `adr import apply [sheets] [--partial]` writes each sheet whose `todo` is empty as a v1 record, then lints it. A sheet with open items is skipped. `--partial` writes it anyway, and lint reports what is missing. +- `adr new` builds an empty sheet from its arguments and applies it, so a new record and an imported record share one writer. In a v1 project it writes v1. +- `adr supersede <old> --by <new>` writes both sides of the link. `adr enact <n> <commit>` sets `enacted` on an accepted cut or retire. + +### 4. Imported records + +An imported record carries `imported: {from, format}` in its frontmatter. For an imported record, a missing Summary is a lint warning. The Summary may be written later, once the whole corpus has been imported and read together, and git history records when it was added. + +### 5. Numbering + +A source numbered inside the project's domain ranges keeps its number. Code, ways and other records cite these numbers, and nothing structural calls for new ones, so an import never renumbers them. This covers every v0 record from agent-ways, in this repo or any other. Any other source gets a number from `target`, which the reader proposes from the domain and the agent may change. `apply` rewrites references within the imported set to the new numbers, and `imported.from` keeps the original identifier. + +### 6. Domains evolve + +The domain layout is not fixed. An import may add or combine domains, and a corpus fed from a tracker keeps growing new ones. A record's number tells you its domain, as it does under v0, so the number follows the domain: + +- A domain may hold several ranges. Merging two domains keeps both ranges, so no record is renumbered. +- A record that moves to another domain, whether by a split or on its own, gets a new number from that domain's range. `adr domain move` rewrites every reference to the record, whether by number (`ADR-N`) or by path (a link to the file, whose folder and name both change). That covers records (`related`, `supersedes` and the other edges, and links in the body), catalog docs, READMEs, ways and code. `adr cite` finds the number references, and the move finds the path references. This is more curation work than keeping the number, and in exchange a number always names its domain. +- The old number is retired and never reused. The record carries `renumbered_from: [ADR-N]`, so a citation that can't be rewritten, such as one in a commit message or a closed pull request, can still be traced, and `adr cite` reports any left in the tree. +- `adr domain add`, `merge`, `split` and `move` edit `adr.yaml`, move and renumber the files, rewrite citations and regenerate the index. +- A sheet whose `target.domain` names a domain that does not exist yet creates it on `apply`, with a free range. + +An import alone never renumbers (§5). Renumbering happens only when a record changes domain. + +### 7. Round-trip guarantees + +These are tested properties: + +- Applying an unedited sheet keeps the body byte-identical, and every source field is either mapped into `record` or kept in `unmapped`. +- Scanning a v1 record and applying the result reproduces the record. +- v0 output of every existing command stays byte-identical. + +## Consequences + +### Positive + +- Migrating a corpus becomes one scan, one cleanup pass over small structured files, and one apply. This repo's 93 v0 records and any other repo's file-based records go through the same path. +- A new source needs a field map, not a code change, whenever it flattens to fields and a body. +- `adr new` produces a v1 record in a v1 project. + +### Negative + +- The sheet schema and the field-map language are formats the tool must keep stable across versions, since sheets and maps outlive a single run. +- Imported records may sit without a Summary, and lint keeps warning until they have one. +- Field maps are a small configuration language that has to be documented and kept stable. + +### Neutral + +- ADR-304 §7's status table is unchanged. The v0 reader applies it. +- A v0 record edited by hand still moves to v1 as before. Import is the bulk path alongside it. + +## Alternatives Considered + +- **An agent migrates each record by hand.** The #581 rehearsal shows this works, but across a corpus of a hundred records the mechanical fields would be retyped each time, and nothing would check the result the way a round trip does. +- **A code reader per format, with no field maps.** This is exact for known formats, but every tracker and template needs code in the vendored tool. +- **Migrate in place with no intermediate object.** A migration writes the record directly. Without a sheet there is no place to stage what needs judgement, and no object to test the round trip against. + +## Corrections + +Appended 2026-09-27. Each entry was first made in place after acceptance and moved here so the accepted text above stays as it was. None changes what was decided. + +- **Summary, probes.** Each probe gained a label, so the operator's answers in `considered` can name the probe they cover: *Confident (frontmatter-carries-all)* and *Not confident (non-markdown-bodies)*. +- **§1, the sheet example.** The example source path reads `docs/architecture/ways/ADR-186-….md`. The `system` domain was renamed `ways` (ADR-310). +- **§2, bodies.** Recorded from the operator's probe answers. A body is one of two things. Text is markdown: a source whose text is in another formatting language (HTML, a tracker's rich text, wiki markup) is converted to markdown before the sheet is written, by the reader or by a step run before `scan`. Structured data is YAML or JSON: a source item whose content is data keeps it as data, in the sheet and in the record as a fenced `yaml` or `json` block. An import carries what the source says. A record states what was decided, and the implementation nearly always drifts from it, so an imported record is not evidence of what the code does; `adr cite` and review compare the two after import. +- **§3, `--partial`.** Some open items block even `--partial`, because lint cannot detect them after the record is written: a Deprecated record's missing historical note, a status that maps to nothing, and a changed number or domain. +- **§4, imported records.** An imported record carries `imported: {from, format, status, unmapped}`. `status` is the source's raw status, and `unmapped` holds every source field with no v1 home. No source value depends on a todo item being honoured to survive the import. The accepted text listed only `from` and `format`; the import as built keeps both extra fields. diff --git a/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md b/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md new file mode 100644 index 00000000..7115a066 --- /dev/null +++ b/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md @@ -0,0 +1,135 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +extends: [ADR-304] +amends: [ADR-304#4] +basis: + - operator: aaronsb + level: guided + said: "experiencing the phenomenonolgy of software is a missing piece that I think an agent decision record can promote." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "A complex flow could even demonstrate the phenomenon of the work itself as one modal in the flow, if it's possible to do so" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "Option one but looser because there's so many possible variations depending on the work at hand" + via: "session 2026-09-27, on the shape of an observable; option one, written by the agent, was a plain-words `see` with an optional `run` command" + - operator: aaronsb + level: guided + said: "Optional everywhere" + via: "session 2026-09-27, selected from agent-written options on which decisions must name an observable" + - operator: aaronsb + level: guided + said: "No, `via` covers it" + via: "session 2026-09-27, selected from agent-written options on whether considered records what was seen" + - operator: aaronsb + level: guided + said: "Not yet (Recommended)" + via: "session 2026-09-27, selected from agent-written options on whether the adr tool runs observable commands" + - evidence: "an analysis of 121 vendor documents across four agentic tools found that the platform records events and nothing records the verification duty discharged (arXiv 2608.15678, Aug 2026)" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "It has to be flexible; we are asking for a way of observing function which is not predictable." + via: session 2026-09-27, answering the probes on PR #596 + covers: [flexible-shape] + - operator: aaronsb + said: "I think an ask for an observable is fair. This way the operator could decline or just tell the agent \"observe it yourself, you can loop and iterate\" - don't use that verbatim but that is adjacentto the develop skill" + via: session 2026-09-27, answering the probes on PR #596 + covers: [optional-unused] + - operator: aaronsb + said: "I like the shape of the pr it tracks the minimum needed to state something was actually real. Let's accept and continue" + via: session 2026-09-27, PR #596 + covers: [] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 306 +--- + +# ADR-307: A decision names what should be observable when it holds + +## Summary + +- **Decided:** a decision may carry `observable`: what someone should be able to see, run or try when the decision holds. Its shape is loose, because the work varies. It is optional on every decision, and it may be added or refined after acceptance. When a decision is handed to the operator, the agent demonstrates its observables, where that is possible, before asking. +- **Trades away:** a checkable form. A loose field can't be verified by the tool, so an observable is only as good as its author makes it. +- **One-way?** No. The field is optional, so removing it later changes no record's validity. +- **Probes:** *Confident (flexible-shape):* a free-form list fits the range of work, from a command's output to a page to click through. *Not confident (optional-unused):* whether an optional field gets used at all, or whether the handover guidance alone carries it. +- **Inversion:** at one end the record is prose to be read, and consideration rests on reading. At the other end every decision carries a runnable check the tool enforces, which fits commands and misses everything seen by eye. This decision names what to observe, leaves its form open, and puts the demonstration in the conversation. + +## Context + +`considered` records the operator's words on a decision, but nothing ties those words to having seen the work. Reading a record is reading about the act. Experiencing what the software does is the missing piece, and a decision record can promote it by saying what should be observable once the decision holds. + +`cut` and `retire` decisions already have an observable end in `enacted`, the commit where the removal landed. `add` and `change` decisions say what was decided, but not how anyone would see that it holds. + +## Decision + +### 1. The observable field + +A decision may carry `observable`, a list. Each entry is either a line of plain words or a mapping whose keys the author chooses to suit the work: + +```yaml +observable: + - "tier 2 adr-migrate passes 10 of 10" + - see: a record scanned and applied comes back byte-identical + run: bash tests/adr-import-roundtrip.sh + - see: the evidence page renders the findings table + url: https://claude.ai/artifact/… +``` + +`see` and `run` are conventions, not requirements. A screenshot path, a URL, a scenario name or a step-by-step description are equally valid. Lint checks only that `observable` is a list of strings or mappings. + +### 2. Optional, and open after acceptance + +No decision is required to carry an observable. `observable` joins the fields that may change after a decision leaves proposed (ADR-304 §4, `mutable_after_accept`), because what shows a decision holding often becomes clear only once it is built. + +### 3. Asked for when drafting, demonstrated in the handover + +When an agent drafts an `add` or `change` decision, it asks the operator what should be observable once the decision holds. The operator can name an observable, decline, or hand the observing to the agent. In the last case the agent works out what to observe, runs the work and iterates until it can show the outcome, as the develop loop does, and then writes the observable it used into the record. + +When a decision with observables is handed to the operator, the agent demonstrates them, where possible, as one step of the flow: it runs the command, shows the output or a screenshot, or opens the page. It asks its questions afterwards. The consider way and the choices way carry this guidance. `considered.via` says what the operator was shown. No separate field records it. + +### 4. The tool does not run observables + +`adr` does not execute `run` entries. The agent runs them during the handover, or while iterating on the work (§3). A command runner in the tool can be decided later, once there is evidence of how observables are written. + +## Consequences + +### Positive + +- A decision can say how anyone, human or agent, would see it holding, in whatever form suits the work. +- A consideration made after a demonstration differs, in its `via`, from one made after reading. + +### Negative + +- Nothing enforces that an observable exists, or that it works. +- A loose field is harder to use mechanically later, for a runner or a report. + +### Neutral + +- Imported records carry no observables until someone adds them, which §2 allows. +- `enacted` stays as it is. It is the observable end of a cut or retire. + +## Alternatives Considered + +- **A fixed shape: `see` plus an optional `run`.** Rejected by the operator as too narrow for the variety of work. +- **Required on add and change, with a lint warning.** Rejected in favour of optional everywhere. The agent asks when drafting (§3), and the field itself stays optional. +- **A `seen` list on `considered`.** Rejected: `via` already says what the operator was shown. +- **`adr observe N` running a decision's commands.** Deferred until observables have been written in practice. + +## Corrections + +Appended 2026-09-27. The entry was first made in place after acceptance and moved here so the accepted text above stays as it was. It does not change what was decided. + +- **Demonstrating observables.** The consider way carries this guidance, and routes a batch of questions through the choices way. The accepted text said the consider way and the choices way both carry it; the choices way does not mention observables. diff --git a/docs/architecture/documentation/ADR-308-a-change-decision-may-list-several-capabilities.md b/docs/architecture/documentation/ADR-308-a-change-decision-may-list-several-capabilities.md new file mode 100644 index 00000000..bcbf9f6b --- /dev/null +++ b/docs/architecture/documentation/ADR-308-a-change-decision-may-list-several-capabilities.md @@ -0,0 +1,82 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-304#3] +basis: + - operator: aaronsb + level: guided + said: "Allow a multi-capability change" + via: "session 2026-09-27, selected from agent-written options while converting agent-ways' records; the label was written by the agent" + - evidence: "converting agent-ways' v0 records: about 20 of 92 alter several capabilities at once, e.g. ADR-123 unifies firing dynamics across attend, matching and disclosure, and ADR-174 spans disclosure, method and authoring" + - precedent: ADR-305 +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "Accept, and warn on long lists" + via: "session 2026-09-27, selected from agent-written options on PR #599; the operator chose the option that adds a lint warning past three capabilities" + covers: [list-inflation] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 305 + - 306 +--- + +# ADR-308: A change decision may list several capabilities + +## Summary + +- **Decided:** a `change` decision may name several capabilities as a list, when it genuinely alters each of them. Each listed capability must have its own prior, as a single one does: a supersede or amend edge covering it, or the baseline standing in (ADR-305). `*` stays reserved for `constrain`. +- **Trades away:** the rule that one change touches one capability, which kept each change's blast radius obvious from its frontmatter. +- **One-way?** No. A later change can narrow the rule again. Records written with a list stay valid, since each listed capability is checked on its own. +- **Probes:** *Confident (list-matches-history):* the ~20 converted records that alter several capabilities describe them more truthfully as a list than under one main capability. *Not confident (list-inflation):* whether authors will list capabilities a change only brushes, so that a list becomes a way to hedge instead of a statement of what changed. +- **Inversion:** at one end a change names one capability, and a cross-cutting change is either split or misrepresented. At the other end, a change names whatever it touches, and capability no longer says where a decision's weight sits. This decision allows a list when each capability is altered, and checks each one. + +## Context + +ADR-304 §3 lets only `constrain` name several capabilities. Converting agent-ways' own records showed that about 20 of 92 alter more than one capability. ADR-123 unifies the firing dynamics that attend, matching and disclosure each carried. ADR-305 left ADR-123 on v0 for exactly this reason. Assigning each such record one main capability makes it pass lint, but it understates what the decision did. + +## Decision + +### 1. The amended rule + +The ADR-304 §3 rule on capabilities now reads: + +- A `change` decision's `capability` is one name, or a list of names when the decision alters each of them. A `constrain` decision's `capability` may be a list or `*`. Other verbs take one name. + +### 2. Each listed capability is checked on its own + +The ADR-304 §3 requirement that a change supersede or amend a prior decision applies to each listed capability separately. The ADR-305 baseline exemption also applies per capability. + +### 3. What a list claims + +A list claims that the decision alters each named capability. A capability the decision only mentions, or depends on without changing, belongs in `related`, not in the list. Lint warns when a change lists more than three capabilities, as a prompt to check that each one is altered. + +## Consequences + +### Positive + +- Cross-cutting changes convert truthfully, ADR-123 among them. +- A query for decisions on a capability finds every change that altered it. + +### Negative + +- A change's frontmatter no longer shows a single place where its weight sits. +- Lint checks that each listed capability has a prior. It can't check that the decision really alters each one; the warning past three only prompts the check. + +### Neutral + +- `add`, `cut` and `retire` keep one capability. `constrain` is unchanged. + +## Alternatives Considered + +- **Main capability only.** The agent's recommendation. Each record carries the capability it mostly changes, and the rest appear through `related`. The operator chose the list, since it describes history more truthfully. +- **Split each cross-cutting record into one change per capability.** Faithful to the rule, but it rewrites history for imported records, and it multiplies records for a single decision. diff --git a/docs/architecture/documentation/ADR-309-an-evidence-record-kind-for-findings-surveys-and-explorations.md b/docs/architecture/documentation/ADR-309-an-evidence-record-kind-for-findings-surveys-and-explorations.md new file mode 100644 index 00000000..1aacd184 --- /dev/null +++ b/docs/architecture/documentation/ADR-309-an-evidence-record-kind-for-findings-surveys-and-explorations.md @@ -0,0 +1,105 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-304#1] +basis: + - operator: aaronsb + level: guided + said: "should we move design notes into agent decision records?" + via: session 2026-09-27, while deciding how stale citations in docs/design-notes should be treated + - operator: aaronsb + level: guided + said: "Yes: evidence kind, specs as spec (Recommended)" + via: "session 2026-09-27, selected from agent-written options; the label was written by the agent" + - precedent: ADR-304 + - evidence: "issue #491: research that informs a decision lands in a prior-art document the decision links to, and has no home in the record corpus" + - evidence: "docs/design-notes holds 13 undated, unnumbered notes (surveys, audits, specs and explorations) that decisions cite by path and that cite records by number" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "Same kind for now (Recommended)" + via: "session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent" + covers: [exploration-fit] + - operator: aaronsb + said: "Accept both and proceed (Recommended)" + via: "session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent" + covers: [] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 306 +--- + +# ADR-309: An evidence record kind for findings, surveys and explorations + +## Summary + +- **Decided:** the record corpus gains an `evidence` kind for findings, surveys, audits, measurements and the reasoning that came before a decision. An evidence record is frozen once accepted, takes no verb, names a capability, and may be cited by a decision's basis. The design notes move into the corpus: specifications become `spec` records, and everything else becomes `evidence`. +- **Trades away:** the looseness of an unnumbered notes folder. A note now needs frontmatter, a number and an accept step. +- **One-way?** No. A kind is configuration in `adr.yaml`, and evidence records can be moved back out. +- **Probes:** *Confident (frozen-history):* freezing evidence at acceptance keeps its citations true to their moment, so an exploration citing a record later superseded stays correct as history. *Not confident (exploration-fit):* whether explorations, the reasoning before a decision, belong in the same kind as measurements and surveys, or will want their own kind once there are more of them. +- **Inversion:** at one end evidence stays outside the corpus as loose notes, uncitable by number and unchecked. At the other end every note becomes a decision, and decisions fill with material that decides nothing. This decision gives findings their own kind: numbered and checked, but separate from decisions. + +## Context + +ADR-304 §1 declared record kinds as data, seeded `decision` and `spec`, and named an evidence kind as a likely next kind. §8 sends research, findings and benchmarks to evidence notes that decisions link to (#491). Those notes live in `docs/design-notes/`: 13 files without frontmatter or numbers. They are cited by path, and they cite records by number, including records since superseded, so `adr cite` reports them as stale. + +## Decision + +### 1. The evidence kind + +`adr.yaml` declares: + +```yaml +kinds: + evidence: + mutable_after_accept: [status, superseded_by, related] + verb: forbidden + requires: [capability] + edges: { supersedes: evidence } +``` + +An evidence record records what was found, measured or reasoned at a point in time. It is frozen once accepted, like a decision. A later finding supersedes it and does not rewrite it. It carries no Summary requirement, no verb and no `agent` requirement. + +### 2. Decisions cite evidence + +A decision's `basis` edges may point at evidence records as well as decisions and specs (`basis: [decision, spec, evidence]`). A basis entry `evidence: ADR-N` that names a record must resolve to an evidence or spec record. + +### 3. What moves + +Each design note becomes a record in the area its content belongs to, with a number from that area's band: + +- A note that specifies behaviour that is kept current becomes a `spec`. +- A survey, audit, measurement, plan or exploration becomes `evidence`, accepted as the record of its moment. +- Every path reference to a moved note is rewritten. The notes' own citations of superseded records stay as written, because they are history. + +## Consequences + +### Positive + +- Findings are numbered and citable, and a decision's basis can point at one. +- `adr cite` stops flagging historical citations inside evidence, which is frozen by kind. + +### Negative + +- Writing a note now takes frontmatter and an accept step. +- The contract carries a third kind that tools and guidance must know about. + +### Neutral + +- `spec` is unchanged; it gains two records from the move. +- Closes #491. + +## Alternatives Considered + +- **Exclude the notes from `adr cite` and leave them where they are.** This silences the warnings, but the notes stay uncitable, and the kind ADR-304 anticipated stays missing. +- **A separate `note` kind for explorations.** More precise, but a fourth kind before there is evidence it's needed. The not-confident probe keeps the question open. +- **Make each note a decision.** Rejected: most notes decide nothing, and a decision without a decision in it misleads. diff --git a/docs/architecture/documentation/ADR-310-record-numbers-are-permanent-identity-and-records-live-by-intent.md b/docs/architecture/documentation/ADR-310-record-numbers-are-permanent-identity-and-records-live-by-intent.md new file mode 100644 index 00000000..c1e19c83 --- /dev/null +++ b/docs/architecture/documentation/ADR-310-record-numbers-are-permanent-identity-and-records-live-by-intent.md @@ -0,0 +1,130 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-306#6] +basis: + - operator: aaronsb + level: guided + said: "I think it's ok to reorganize where they're located too (make the directory tree more aligned with intention)" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "I'm not sure - nobody external is citing these adr numbers, except the github record. so, I'm leaning keeping numbers where range allows but I'm not sure" + via: session 2026-09-27, on renumbering records during a reorganization + - operator: aaronsb + level: guided + said: "I'm thinking about how wbs has systems numbers. for instance NQA and systems design in nuclear archtecture (where I was exposed to this first) has a standardized numbering scheme" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "or, think about parts systems, mercedes benz for example" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "Adopt it (Recommended)" + via: "session 2026-09-27, selected from agent-written options: permanent numbers, capability as classification, moves change the folder and not the number; the label was written by the agent" + - operator: aaronsb + level: guided + said: "Six intent folders (Recommended)" + via: "session 2026-09-27, selected from agent-written options; the label was written by the agent" + - operator: aaronsb + level: guided + said: "Next in the area's band (Recommended)" + via: "session 2026-09-27, selected from agent-written options on how new numbers are allocated; the label was written by the agent" + - standard: "10 CFR 50 Appendix B, Criterion VIII: an item's identity is maintained on the item or on records traceable to it, throughout fabrication, erection, installation and use" + - standard: "MIL-STD-881F §4.3: WBS numbers show indenture only, and maintaining WBS numbering is not a requirement of the standard" + - evidence: "a survey of systems numbering schemes (VMRS, ATA iSpec 2200, IEC 81346, IEEE 803, MasterFormat, UniFormat, MIL-STD-881, Mercedes-Benz part numbers, Dewey, UDC, ACM CCS, SWEBOK, ISO/IEC 25010, 12207 and 42010) found no commonly known systems code for software, and found that traceability-heavy industries keep a permanent identity separate from a classification that may change, e.g. SAP's equipment number versus functional location" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "Fine: git and INDEX explain it (Recommended)" + via: "session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent" + covers: [band-hint] + - operator: aaronsb + said: "Accept both and proceed (Recommended)" + via: "session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent" + covers: [] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 306 + - 309 +--- + +# ADR-310: Record numbers are permanent identity, and records live by intent + +## Summary + +- **Decided:** a record's number is its permanent identity. It is never changed and never reused. What a record is about is classified by `capability`, and the folder it lives in follows that classification. A record that belongs in another area moves folders and keeps its number. agent-ways' records move into six intent areas: ways, governance, documentation, attend, platform and practice. New numbers come from the next free slot in their area's band. +- **Trades away:** reading a record's area from its number. After a move, the hundreds digit tells where a record started, not where it is. +- **One-way?** No. Folders can be reorganized again at the cost of rewriting path links. Numbers never change, so no citation breaks either way. +- **Probes:** *Confident (identity-stable):* commit messages, pull requests and issues cite records by number, and none of them can be rewritten. A number that never changes keeps every one of them pointing at the right record. *Not confident (band-hint):* whether allocating new numbers by band will mislead readers into treating the digits as meaning, once moved records sit outside their folder's band. +- **Inversion:** at one end, numbers carry the classification and are renumbered on every move, which breaks citations each time. At the other end, numbers are a single global sequence with no hint at all. This decision keeps identity permanent and lets the band hint only at a record's origin. + +## Context + +ADR-306 §6 made a record's number follow its domain: moving a record to another domain renumbered it and rewrote every reference. Planning a reorganization of agent-ways' records by intent put that rule to the test. It would have renumbered about 62 records and rewritten their citations. Citations in the git history, pull requests and issues can't be rewritten. + +A survey of systems numbering found no commonly known systems code for software: schemes like VMRS and ATA chapters work only because every truck or aircraft has the same systems. It also found that the industries with the strictest traceability requirements keep identity and classification apart. 10 CFR 50 Appendix B requires an item's identity to persist through fabrication, installation and use. Asset management separates the equipment number, which never changes, from the functional location, which follows the structure. MIL-STD-881F declines to require stable WBS numbers, because its numbers show only position. + +## Decision + +### 1. The amended rule + +ADR-306 §6 is amended. Its rule that a number follows its domain, and that a move renumbers and rewrites references, is replaced by: + +- A record's number is its permanent identity. It is never changed, and never reused for another record. +- Under adr/v1, a record's area (domain) is the folder it lives in. The number range of an area only allocates numbers for new records. +- A record that belongs in another area moves to that area's folder, keeps its number, and has every path reference to it rewritten. `adr domain move` does this. +- A domain may still hold several ranges, and areas may be added, merged and split. None of these changes a number. + +Under adr/v0 the area is still read from the number range, and v0 output is unchanged. + +### 2. Classification + +`capability` classifies what a record is about (ADR-304, ADR-308). The folder follows it: each area holds the capabilities it groups. No separate systems code is added. The capability vocabulary is the project's own list of its systems, validated by lint. + +### 3. agent-ways' areas + +| Area | Folder | Band | Capabilities | +|---|---|---|---| +| ways | `ways/` (renamed from `system/`) | 100–199 | matching, disclosure | +| governance | `governance/` | 200–299 | governance | +| documentation | `documentation/` | 300–399 | adr, docs | +| attend | `attend/` | 400–499 | attend | +| platform | `platform/` | 500–599 | install, config, cli, testing | +| practice | `practice/` | 600–699 | authoring, method, loop | + +Each record moves to the area of its main capability (the first listed). Evidence and spec records (ADR-309) are placed the same way. The legacy folder empties, and its range stays reserved so that no number in it is ever reused. + +## Consequences + +### Positive + +- No citation anywhere, including git history, pull requests and issues, ever goes stale because of a move. +- The directory tree reads by intent. +- Reorganizing costs a folder move and path rewrites, and never a renumbering. + +### Negative + +- A number's hundreds digit no longer tells where a record lives. Tools and readers must use the folder or `capability`. +- Path links break on every move until rewritten, so moves must go through `adr domain move`. + +### Neutral + +- ADR-306 §6's domain commands keep their purpose (add, merge, split, move) without renumbering. + +## Alternatives Considered + +- **Renumber on move (ADR-306 §6 as written).** Rejected: it breaks citations that can't be rewritten, which is the failure the traceability standards are built to prevent. +- **Numbers encode origin and are frozen, with a separate current-location field (the Mercedes-Benz part-number pattern).** Close to this decision, but the folder and `capability` already carry the current location, so a separate field would duplicate them. +- **A systems code borrowed from an existing standard.** None exists for software, and codes like ATA chapters are specific to one kind of machine. +- **One flat folder.** No path ever changes, but the tree says nothing about intent, which the operator asked for. diff --git a/docs/architecture/documentation/ADR-311-the-adr-tool-checks-shape-and-references-git-keeps-the-history.md b/docs/architecture/documentation/ADR-311-the-adr-tool-checks-shape-and-references-git-keeps-the-history.md new file mode 100644 index 00000000..94399b51 --- /dev/null +++ b/docs/architecture/documentation/ADR-311-the-adr-tool-checks-shape-and-references-git-keeps-the-history.md @@ -0,0 +1,91 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +supersedes: + - ADR-305 +amends: [ADR-304#1, ADR-304#5, ADR-304#6, ADR-304#11, ADR-304#12, ADR-308#2] +basis: + - operator: aaronsb + level: directed + said: "git is always the backstop. we should never try to recreate anything that git does" + via: session 2026-09-27, after reviewing the friction in #604 + - operator: aaronsb + level: directed + said: "we're really just recording decisions and responsibility in a ledger style in git tracked files" + via: session 2026-09-27 + - operator: aaronsb + level: directed + said: "not trying to invent some sort of complex checks and balances" + via: session 2026-09-27, the same message + - operator: aaronsb + level: guided + said: "Audit all gates first" + via: "session 2026-09-27, selected from agent-written options on where to draw the line; the label was written by the agent" + - evidence: "an audit of every check in adr-tool 2.x: shape and reference checks cost little and keep the ledger readable; the policy gates (frozen history, capability add, change prior, precedent chain, vocabulary stems, Summary legibility, enactment inventories, accept's corpus re-lint) produced most of the configuration and review friction in #602 to #604" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: Aaron Bockelie + said: "Holds" + via: session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent + covers: [shape-is-enough] + - operator: aaronsb + said: "Way teaches, review catches (Recommended)" + via: session 2026-09-27, selected from agent-written options when the probes were asked; the label was written by the agent + covers: [convention-holds] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 305 + - 308 +--- + +# ADR-311: The adr tool checks shape and references; git keeps the history + +## Summary + +- **Decided:** records are a ledger of decisions and who made them, kept in git-tracked files. The adr tool writes records, checks their shape, and checks that their references resolve. It enforces no policy about how decisions relate or whether a record changed; that is convention, taught by the ADR way, and git holds every version. +- **Trades away:** automatic detection of an accepted record edited in place, a capability with no `add`, a `change` with no prior, and a precedent chain that stays inside the corpus. A reviewer or a reader of git history finds those now. +- **One-way?** No. A removed check can come back as its own decision if its absence costs more than it did. +- **Probes:** *Confident (shape-is-enough):* the records stay readable and navigable with shape and reference checks alone, since those are what a reader relies on. *Not confident (convention-holds):* whether agents follow "correct by appending" and "a change names what it replaces" without a check, or drift once nothing flags it. +- **Inversion:** at one end the tool only formats files and nothing is checked. At the other it audits every record against the corpus and its history, and the checks need their own configuration and exceptions. This checks what a reader of one record needs: its fields are there, and its links go somewhere. + +## Context + +ADR-304 made records typed and added lint rules that enforce relationships: frozen decisions (§1, §6), enactment checked against the code (§5), an `add` for every capability, a prior for every `change`, precedent that reaches outside the corpus (§11), labelled probes (§12). ADR-305 added a `baseline` to excuse capabilities that predate adoption, and ADR-308 extended the prior rule to lists. Building the frozen check on moved records took three designs in #604, each adding configuration. An audit of the rules found the same pattern across the policy gates. + +## Decision + +1. **What the tool checks.** A record's fields and their values (kind, status, verb, capability in the vocabulary, a basis whose entries each name a source, agent, `imported`, `observable`), its required sections, leftover placeholders, `adr.yaml`'s own shape, and every reference: supersession pairs, `amends` sections, precedent, and `adr cite`'s citations. +2. **What it no longer checks.** Frozen records and `mutable_after_accept`; an `add` per capability and `baseline`; a prior per `change`, for one capability or a list; a precedent chain reaching outside; vocabulary stems across layers; Summary probe labels and inversion; enactment inventories, `surfaces`, and `cite`'s warnings on citations of a cut capability; a retire decision's `targets`. `enacted` and `targets` stay fields. +3. **Commands.** `adr accept` refuses a record that fails its own shape check, and does not lint the corpus. `adr set` refuses `status`, which the lifecycle commands own. `adr supersede` writes both sides on any record. `considered` stays a field; accept does not require it. +4. **Conventions.** Correct an accepted record by appending; record a change as a new decision that names what it replaces; quote the operator. The ADR way teaches these. +5. **Git is the backstop.** Earlier versions, and who changed them, are read with git. + +## Consequences + +### Positive + +- The tool shrinks by several hundred lines, and `adr.yaml` loses `baseline`, `surfaces` and the per-kind mutable lists. +- A move or a correction never needs an exception in configuration. + +### Negative + +- An in-place edit or a missing `amends` goes unflagged until someone reads the record or its history. + +### Neutral + +- Supersedes ADR-305. ADR-308's allowance for a list of capabilities stands; its rule that each needs a prior goes. +- Records already written keep their fields; nothing needs rewriting. + +## Alternatives Considered + +- **Keep the gates as warnings.** Warnings that fire on legitimate history train readers to ignore them. +- **Check only the change under review against its base.** Accepted briefly on an unmerged branch; still a check that git makes unnecessary. diff --git a/docs/architecture/governance/ADR-005-governance-traceability.md b/docs/architecture/governance/ADR-005-governance-traceability.md new file mode 100644 index 00000000..a32d740a --- /dev/null +++ b/docs/architecture/governance/ADR-005-governance-traceability.md @@ -0,0 +1,172 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: governance +superseded_by: + - ADR-200 +basis: + - evidence: the step from policy document to way.md is invisible; the link between policy source and compiled guidance exists only in the author's head (Context, The Compilation Gap) +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-02-05 +deciders: + - aaronsb + - claude +related: + - ADR-013 + - ADR-200 +imported: + from: docs/architecture/legacy/ADR-005-governance-traceability.md + format: v0 + status: Superseded +--- + +# ADR-005: Governance Traceability for Ways + +> **Superseded by [ADR-200](../governance/ADR-200-compliance-claims-and-session-derived-findings.md).** +> This ADR established provenance traceability but framed the hand-authored mappings as +> "auditability" and "proof." ADR-200 corrects the epistemics: the mappings are compliance +> **claims** (control-*design* assertions — SOC 2 Type I), not **findings** (assessed +> evidence). The sidecar mechanism described here is retained and renamed; its framing as +> proof is replaced. Storage later moved from way frontmatter to a `provenance.yaml` +> sidecar (ADR-110). Kept unedited below as the historical record. + +## Context + +### The Compilation Gap + +The ways system has a clear pipeline, documented in `docs/hooks-and-ways/README.md`: + +1. **Principle** — an opinion about how things should work +2. **Governance Interpretation** — how it applies in practice (policy docs) +3. **Documentation** — how the system works (reference layer) +4. **Implementation** — way.md files (machine layer) + +This pipeline produces good guidance, but the transformation from stage 2 to stage 4 is invisible. A human reads the policy, thinks about it, and writes a way.md. The connection between policy source and compiled guidance exists only in that person's head. + +For a personal project, this is fine. For an organization that needs to demonstrate that its AI agent governance actually derives from its stated policies, this is a gap. + +### The Compilation Metaphor + +Ways are compiled policy. The analogy is precise: + +| Concept | Software Build | Way System | +|---------|---------------|------------| +| Source code | `.c` / `.rs` files | Policy documents (ADRs, controls, regulatory specs) | +| Compiler | `gcc` / `rustc` | Human authoring process | +| Object code | `.o` / `.rlib` files | `way.md` files (compressed, context-optimized) | +| Debug symbols | DWARF / PDB | Provenance metadata (frontmatter the runtime ignores) | +| Symbol table | `.map` / `.pdb` | Traceability manifest (generated index) | +| Linker output | Executable | Agent context (assembled at runtime from matching ways) | + +Just as debug symbols don't affect program execution but are essential for debugging, provenance metadata doesn't affect way injection but is essential for governance auditing. + +### Cross-Repo Reality + +In practice, policy documents and way implementations won't live in the same repository: + +- A compliance team maintains ADRs, control inventories, and regulatory mappings in their repo +- A platform team maintains Claude Code ways in `~/.claude/hooks/ways/` +- An enterprise might have multiple policy repos feeding multiple way configurations + +The traceability system must bridge this boundary without coupling the repos. + +## Decision + +Add optional **provenance metadata** to way frontmatter and provide tooling to generate a **traceability manifest** that bridges policy repos to way repos. + +### 1. Provenance Frontmatter + +Way frontmatter gains an optional `provenance:` block: + +```yaml +--- +match: regex +pattern: commit|push +provenance: + policy: + - uri: github://acme-corp/sdlc-controls/docs/architecture/change/ADR-150.md + type: adr + - uri: governance/policies/code-lifecycle.md + type: governance-doc + controls: + - NIST SP 800-53 CM-3 (Configuration Change Control) + - SOC 2 CC8.1 (Change Management) + verified: 2026-02-05 + rationale: > + Conventional commits create structured change records with type classification + and justification, implementing auditable configuration change control. +--- +``` + +**Fields:** + +| Field | Purpose | +|-------|---------| +| `policy[].uri` | Reference to source policy document | +| `policy[].type` | Classification: `adr`, `governance-doc`, `regulatory-framework`, `control-spec` | +| `controls[]` | Regulatory control IDs this way addresses | +| `verified` | Date provenance was last confirmed accurate | +| `rationale` | How policy intent became way guidance | + +**URI convention:** `github://org/repo/path` for cross-repo references. Relative paths for same-repo references. The scheme is a convention, not a protocol — it's meant to be human-readable and machine-parseable, not clickable. + +**Zero runtime cost:** The hook scripts (`show-way.sh`, `check-prompt.sh`, etc.) strip all frontmatter before injection. Fields they don't recognize are silently ignored. Provenance metadata never reaches the agent's context window. + +### 2. Traceability Manifest + +A generated `provenance-manifest.json` that aggregates provenance across all ways: + +- Per-way provenance data extracted from frontmatter +- Inverted indices: policy → implementing ways, control → addressing ways +- Coverage statistics (ways with/without provenance) + +Generated by scanning, not maintained by hand. Can be cross-referenced with external audit artifacts (e.g., a compliance repo's control disposition map). + +### 3. Verification Tooling + +A coverage report script that: +- Identifies ways missing provenance +- Cross-references control IDs against external audit ledgers +- Flags stale `verified` dates +- Reports gaps in both directions (policies without ways, ways without policies) + +Reports, does not enforce. The system is informational, not blocking. + +## Consequences + +### Positive + +- **Traceability exists.** The chain from regulatory framework through policy through way to agent context is walkable in either direction. +- **Zero runtime cost.** Provenance is metadata only — no performance impact, no context window cost. +- **Incremental adoption.** Ways without provenance work exactly as before. Provenance can be added gradually. +- **Cross-repo capable.** URI references work across organizational boundaries without coupling repos. +- **Auditable.** An auditor can ask "show me which agent governance implements NIST CM-3" and get a concrete answer. + +### Negative + +- **Manual compilation.** The human still writes the way. Provenance makes the connection traceable, not automatic. +- **Staleness risk.** Provenance metadata can go stale if policies change and ways aren't updated. The `verified` date and verification tooling mitigate this. +- **Added frontmatter complexity.** Way authors now have an optional block to learn about. Keeping it optional prevents this from being a barrier. + +### Neutral + +- **Not enforcement.** The system reports, it doesn't block. A way with missing or stale provenance still fires normally. +- **Not a product.** This establishes a pattern and proves a concept. Enterprise adoption would likely want tighter integration with their specific compliance tooling. + +## Alternatives Considered + +### External mapping file + +A separate YAML/JSON file mapping ways to policies, kept alongside `ways.json`. Rejected because it separates provenance from the thing it describes — the way file should be self-contained. Also creates a maintenance burden (two files to keep in sync). + +### Mandatory provenance + +Requiring all ways to have provenance. Rejected because many ways (like `meta/todos` or `meta/memory`) are operational, not policy-derived. Mandatory provenance would force meaningless metadata on ways that don't need it. + +### Automated compilation + +A tool that reads policy documents and generates way.md files automatically. Rejected as premature — the compilation requires human judgment about what's relevant to agent context, how terse to be, what triggers to use. This may be feasible in the future but the pattern needs to be established manually first. diff --git a/docs/architecture/governance/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md b/docs/architecture/governance/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md new file mode 100644 index 00000000..24f3b8ad --- /dev/null +++ b/docs/architecture/governance/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md @@ -0,0 +1,182 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - governance + - install +supersedes: [] +basis: + - evidence: the ~1,100-line governance/provenance engine lives inside the binary-only ways-cli crate with zero tests and no library seam to test against + - precedent: ADR-200 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-110 + - ADR-111 + - ADR-142 + - ADR-200 + - ADR-201 +imported: + from: docs/architecture/system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md + format: v0 + status: Accepted +--- + +# ADR-151: Extract ways-core crate and ways-audit sibling binary + +## Context + +ADR-200 reframes the compliance subsystem into an operator-invoked claim/finding +formation layer, backed by a purpose-built toolkit rather than the core `ways` binary. +This ADR decides how that toolkit is packaged. + +The relevant facts about the current code: + +- `tools/` is a Cargo **workspace** (a set of crates built together), and shared + library crates are already the house style here — `sensor-trait`, `agent-fmt`, and + `agent-identity` are exactly that pattern: small libraries several binaries depend on. +- `ways-cli` is a **binary-only crate** (it produces the `ways` executable and has no + `[lib]` target). The compliance logic — the `governance`/`provenance` modules, ~1,100 + lines — lives *inside* that binary crate and reaches into its private siblings + (`crate::util`, `crate::cmd`). It is not a reusable library; it is trapped in the + executable. +- That ~1,100-line engine has **zero tests**, including the integrity linter that is + supposed to audit the claims. Being buried in a binary with no library boundary is + part of why: there is no clean seam to unit-test against. +- ADR-111 consolidated a sprawl of shell scripts (`governance.sh`, `provenance-scan.py`, + …) into the single `ways` binary. That was the right call for *script sprawl*. The + question here is different: the compliance concern is a **distinct operator surface**, + invoked deliberately (like `/ship` or `/wrap`), which the subsystem's own `governance/README.md` already sketches as + separable (its "Making This Its Own Repo" section — a "perforated pop-out"). It should + not be re-scattered — but it also + should not bloat the core tool everyone runs. + +So the packaging question: how does `ways-audit` reuse the engine without rewriting it, +keep the core binary lean, and give the untested engine a home worth testing? + +## Decision + +### 1. Extract a `ways-core` library crate + +Pull the reusable engine out of the `ways` binary into a new **`ways-core`** library +crate in the workspace: way discovery and scanning, frontmatter parsing, path and +projection resolution, the firing-event log reader, and the claim (`provenance.yaml`) +sidecar model and manifest builder. Both `ways` (binary) and `ways-audit` (binary) +depend on it. This is the established house pattern (`sensor-trait` et al.), not a new +paradigm. + +The extraction pays for itself independently of the compliance work: a library boundary +is exactly what makes the previously-untested engine unit-testable. `ways-core` ships +with tests; the core tool gets healthier whether or not anyone ever runs `ways-audit`. + +### 2. Add `ways-audit` as a sibling binary + +A new **`ways-audit`** binary crate, depending on `ways-core`, owns the whole compliance +surface: + +- The reporting commands presently under `ways governance` (report, trace, control, + gaps, matrix, lint), migrated and renamed to the compliance vocabulary of ADR-200 + (claim / finding / POA&M). +- The **finding pipeline**: read a claim + its firing events + the session transcript, + produce an assessment finding with a determination, and append it to the finding + ledger. +- The **claim-authoring assist**: propose candidate control mappings for a way + (agent-assisted, human-grounded per ADR-200 §1). + +The core `ways` binary **sheds** the `governance` subcommand. It goes back to being the +runtime — matching, hooks, session lifecycle — with the compliance concern living next +to it, not inside it. + +### 3. The claim schema lives in `ways-core`, built for scale and assessability + +`ways-core` defines the claim type stored in the `provenance.yaml` sidecar (ADR-110). +Per ADR-200 §1 the schema carries a **determination criterion** — the observable +behavior that would let an assessor mark the claim *satisfied* / *other than satisfied* +— so a claim is assessable rather than a bare assertion. `ways-audit` consumes that +type both to *assess* (produce findings) and to *suggest* (candidate mappings). Defining +it in the shared crate keeps the format single-sourced as claims are seeded across the +corpus at scale (ADR-200 §7). + +### 4. Distribution reuses the per-component release machinery + +`ways-audit` slots into the existing component-parameterized release flow: a +`ways-audit-v*` tag series, its own CI build job, and a download entry, exactly as +`ways` has. The marginal cost is one CI job and one download path — the same +prebuilt-binary machinery already planned for the other app binaries (ADR-142/ADR-146). + +**`ways-audit` is a first-class member of the suite, not an opt-in add-on.** +`ways update` manages the *entire* agent-ways collection; we do not ship partial +tool collections, and no binary has a separate lifecycle. So `make install` / +`make setup` build and link `ways-audit` alongside `ways`/`attend`, and +`ways update` refreshes it through the same download-first `refresh_component` +path as the rest. "Deliberately-invoked" (§2) describes how the tool is *used* — +you run `ways-audit assemble` when you want it, the way you run `/ship` — not +whether it is *installed*: it is always present, like `ways` itself. (An earlier +implementation kept it out of the default install "to stay lean"; that produced a +published-but-unfetchable binary and is corrected here.) + +## Relationship to ADR-111 + +ADR-111 folded shell-script sprawl into one binary because the *abstraction layer* — one +tool, consistent surface — was the value. This ADR does not reverse that: it introduces +**one** additional binary for **one** cohesive, separable concern, and the two binaries +share a real library (`ways-core`) rather than duplicating logic. The distinction is +concern-separation with a shared abstraction, not a return to sprawl. ADR-111's own +"the abstraction layer is the value" argument is what `ways-core` embodies. + +## Consequences + +### Positive + +- The core `ways` binary stays lean — the compliance concern is isolated and optional. +- `ways-core` gives the previously-untested engine a testable boundary; the extraction + improves the main tool on its own merits. +- The claim schema is single-sourced and scale-ready, with assessability built in. +- Extraction to a standalone repo later (already sketched in `governance/README.md`) becomes a small + move — the library seam is already drawn — without paying that cost now. +- Distribution is nearly free: the release/CI machinery already handles N components. + +### Negative + +- A real refactor: extract `ways-core`, move the modules off `crate::util`/`crate::cmd` + onto the library API, and fix imports across the workspace. +- A second binary to build, ship, and version. +- `ways-core` now has a public API surface to keep stable for two consumers. + +### Neutral + +- `ways-audit` can version independently of `ways` (different cadence, own tag series). +- The `ways governance` command path is retired from the core binary; a one-release + deprecation pointer to `ways-audit` can be kept if any operator scripted against it, + though it was never a consumed surface. +- Repo extraction remains an available future option, deliberately not taken now. + +## Alternatives Considered + +- **Keep compliance inside the `ways` binary (status quo / strict ADR-111).** Rejected: + it bloats the core runtime with an optional, deliberately-invoked concern, offers no + separation seam, and couples the compliance release cadence to the core tool. +- **A new binary that copies the engine (no shared crate).** Rejected: duplication and + drift. The workspace already solves this with shared library crates; not using one + here would be the anomaly. +- **Extract the whole subsystem to a separate repository now.** Rejected as premature: + a same-workspace sibling is cheaper, keeps the `ways-core` refactor in one place, and + still leaves repo extraction open once the toolkit proves itself. Drawing the library + seam now is the reversible half of that decision. + +## References + +- **ADR-200** — the compliance claim/finding model this toolkit backs. +- **ADR-110** — the `provenance.yaml` sidecar storage the claim schema extends. +- **ADR-111** — the single-binary consolidation this refines (one sibling, shared lib). +- **ADR-142 / ADR-146** — XDG application distribution and installer binary handling the + new binary plugs into. +- **The Cargo Book — Workspaces** — the shared-crate mechanism used here. + https://doc.rust-lang.org/book/ch14-03-cargo-workspaces.html diff --git a/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md b/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md index 644f96d1..d9019b1e 100644 --- a/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md +++ b/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md @@ -1,5 +1,19 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: governance +supersedes: + - ADR-005 +basis: + - evidence: 'nothing consumes the provenance subsystem: no CI job, hook, or reviewer runs it, and its claims were never assessed' + - standard: 'NIST SP 800-53A Rev. 5 (assessment findings: satisfied / other than satisfied)' + - standard: NIST OSCAL layered model (Component Definition, Assessment Results, POA&M) + - standard: AICPA SOC 2 Type I and Type II under SSAE 18 +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-07-02 deciders: - aaronsb @@ -11,8 +25,10 @@ related: - ADR-111 - ADR-151 - ADR-201 -supersedes: - - ADR-005 +imported: + from: docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md + format: v0 + status: Accepted --- # ADR-200: Compliance claims and session-derived findings diff --git a/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md b/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md index c6aa6e22..38c22119 100644 --- a/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md +++ b/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md @@ -1,5 +1,16 @@ --- -status: Accepted +contract: adr/v1 +kind: decision +verb: change +capability: governance +basis: + - standard: 'NIST SP 800-53A: assessment findings, and the assessor as a role' + - standard: NIST OSCAL Assessment Results + - precedent: ADR-200 +agent: + name: Claude + model: unrecorded +status: accepted date: 2026-07-02 deciders: - aaronsb @@ -8,6 +19,10 @@ related: - ADR-110 - ADR-151 - ADR-200 +imported: + from: docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md + format: v0 + status: Accepted --- # ADR-201: Findings assembled as classifier-ready assessment records diff --git a/docs/architecture/platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md b/docs/architecture/platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md new file mode 100644 index 00000000..dd42e4ba --- /dev/null +++ b/docs/architecture/platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md @@ -0,0 +1,192 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - install + - authoring + - matching +basis: + - evidence: tooling spread across C, C++, Bash and Python (tool inventory table); every tool re-walks the tree and re-parses frontmatter, and the bash carries macOS bash 3.2 constraints + - standard: 'the gh, aws and gcloud CLIs: one binary with subcommands and shared infrastructure' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-30 +deciders: + - aaronsb + - claude +related: + - ADR-014 + - ADR-107 + - ADR-108 + - ADR-110 +imported: + from: docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md + format: v0 + status: Accepted +--- + +# ADR-111: Unified `ways` CLI — Single Binary Tool Consolidation + +## Context + +The ways tooling has grown organically across multiple languages and entry points: + +| Tool | Language | Lines | Function | +|------|----------|-------|----------| +| `way-match` | C | ~920 | BM25 scoring | +| `way-embed` | C++ | ~920 | Embedding match (ONNX/GGUF) | +| `generate-corpus.sh` | Bash | ~200 | Corpus generation | +| `lint-ways.sh` | Bash | ~530 | Frontmatter validation | +| `way-tree-analyze.sh` | Bash | ~300 | Tree structure analysis | +| `embed-lib.sh` | Bash | ~200 | Shared utilities | +| `embed-suggest.sh` | Bash | ~100 | Embedding suggestions | +| `provenance-scan.py` | Python | ~150 | Provenance scanning | +| `governance.sh` | Bash | ~540 | Governance orchestration | +| Various others | Bash | ~500 | Misc utilities | + +Every tool re-walks the same directory tree and re-parses the same frontmatter. Adding a new feature (e.g., the graph generator from ADR-110) means writing yet another script that duplicates file discovery, YAML extraction, and JSON emission. The bash scripts also carry macOS bash 3.2 compatibility constraints that a compiled binary eliminates. + +The `gh` CLI, `aws` CLI, and `gcloud` CLI demonstrate the pattern: one binary, subcommands for everything, shared infrastructure for common operations. + +## Decision + +### 1. Create a `ways` CLI binary + +A single `ways` binary replaces all current tooling with subcommands: + +``` +ways lint [path] # frontmatter validation (lint-ways.sh) +ways corpus [--global] # corpus generation (generate-corpus.sh) +ways match <query> # BM25 scoring (way-match) +ways embed <query> # embedding match (way-embed) +ways siblings <id> # way-vs-way cosine scoring (new, ADR-110 §5) +ways graph [--format jsonl]# graph export (new, ADR-110 §4) +ways tree <path> # tree analysis (way-tree-analyze.sh) +ways provenance # provenance scanning (provenance-scan.py) +``` + +### 2. Pure Rust implementation + +The entire CLI is pure Rust — no FFI, no C/C++ compilation. BM25 scoring was reimplemented natively (~176 lines in `bm25.rs`). Embedding matching delegates to the existing `way-embed` binary via subprocess (the embedding engine requires GGUF/ONNX runtime which remains a separate C++ binary). + +The original ADR planned FFI wrappers via the `cc` crate, but the BM25 C code was small enough to port directly. This eliminated cross-compilation complexity entirely — `cargo build` produces the binary with no native toolchain required beyond Rust. + +### 3. Project structure + +``` +tools/ways-cli/ +├── Cargo.toml +├── src/ +│ ├── main.rs # clap dispatcher (19 subcommands) +│ ├── cmd/ +│ │ ├── scan/ # prompt/command/file/state matching +│ │ ├── show/ # session-aware way display +│ │ ├── governance/ # 9 governance query modes (7 files) +│ │ ├── lint.rs # frontmatter validation +│ │ ├── corpus.rs # corpus generation +│ │ ├── match_bm25.rs # BM25 scoring (pure Rust) +│ │ ├── embed.rs # embedding match (delegates to way-embed) +│ │ ├── list.rs # session way list with forecast +│ │ ├── context.rs # token usage from transcript +│ │ ├── reset.rs # session state recovery +│ │ └── ... # graph, tree, provenance, stats, etc. +│ ├── bm25.rs # BM25 engine (Porter2 stemming, IDF) +│ ├── scanner.rs # shared: file discovery by frontmatter +│ ├── frontmatter.rs # shared: YAML frontmatter parsing +│ ├── session.rs # session state (directory-per-session) +│ ├── table.rs # ANSI-aware table formatting +│ └── util.rs # shared utilities (home_dir, project detection) +├── tests/ +│ └── session_sim.rs # 8 integration scenarios +└── download-ways.sh # pre-built binary installer +``` + +### 4. Incremental delivery + +Subcommands ship independently. The order follows dependency: + +| Phase | Subcommands | Replaces | Status | +|-------|-------------|----------|--------| +| 1 | `lint`, `corpus`, `graph` | lint-ways.sh, generate-corpus.sh, new | Shipped | +| 2 | `match`, `embed`, `siblings` | way-match (ported to Rust), way-embed (subprocess) | Shipped | +| 3 | `tree`, `provenance`, `scan`, `show`, `governance`, `context`, `list`, `stats`, `reset`, `init`, `status`, `suggest` | All remaining scripts | Shipped | + +All three phases delivered as pure Rust. BM25 was ported rather than wrapped via FFI. Embedding delegates to the existing `way-embed` binary. The `scan` and `show` subcommands absorbed the hook orchestration that was previously spread across show-core.sh, show-way.sh, match-way.sh, and check-prompt.sh. + +### 5. Installation and distribution + +The binary installs to `~/.claude/bin/ways` with a symlink to `~/.local/bin/ways`. Three install paths: + +1. **Download** — `download-ways.sh` pulls pre-built binary from GitHub Releases (no toolchain needed) +2. **Build from source** — `cargo build --release` (requires Rust toolchain) +3. **`make install`** — tries download first, falls back to build + +CI builds 4 platforms (linux-x86_64, linux-aarch64, darwin-x86_64, darwin-arm64) via `cargo-zigbuild` for ARM cross-compilation. Tagged releases (`ways-v*`) create GitHub Releases with checksums. + +The 10 remaining hook scripts are thin dispatchers — they parse hook JSON input and call `ways scan` or read session state. The orchestration, matching, and display logic lives entirely in the binary. + +### 6. Remaining C/C++ (embedding engine) + +BM25 was ported to pure Rust (176 lines in `bm25.rs`). The embedding engine (`way-embed`) remains a separate C++ binary because it depends on llama.cpp for GGUF model inference. The `ways embed` subcommand delegates to `way-embed` via subprocess. + +Future option: the `ort` crate could replace the C++ embedding binary with a pure Rust ONNX path, eliminating the last subprocess dependency. This is not planned — the current approach works and the embedding binary is stable. + +## Consequences + +### Positive + +- Single binary, single install, single update path +- Shared file scanning — one tree walk serves all subcommands +- Shared frontmatter parsing — one YAML parser, tested once +- New features (graph, siblings) are subcommands, not new scripts +- macOS bash 3.2 compatibility concerns eliminated for ported logic +- `ways --help` gives discoverability across all tooling +- Shell completion for free via `clap` + +### Negative + +- Rust toolchain required for development (not for end users — pre-built binaries available) +- CI cross-compiles for 4 platforms via `cargo-zigbuild` (working, but adds build complexity) +- Porting bash to Rust took more lines for the same functionality (mitigated: type system caught bugs the bash scripts silently swallowed) + +### Neutral + +- `governance.sh` (543 lines), `provenance-verify.sh`, and `context-usage.sh` were deleted — fully replaced by `ways governance`, `ways governance lint`, and `ways context` +- 10 bash hook scripts remain as thin dispatchers (15-30 lines each) — they parse hook JSON and call `ways scan` +- `Makefile` has `make ways`, `make ways-rebuild`, `make install`, `make release` targets +- CI builds changed from "compile C, compile C++, run bash" to "cargo build, run tests" +- Session state moved from flat `/tmp/.claude-*-{uuid}` markers to per-user `/tmp/.claude-sessions-{uid}/{session_id}/` directories + +## Alternatives Considered + +### Pure Go + +Go's `cobra` library is excellent for CLIs and cross-compilation is normally trivial. However, ONNX Runtime is a C library — Go requires CGo to call it. CGo cross-compilation for 4 platforms requires Docker-based toolchains or zig-cc as a C cross-compiler, negating Go's primary advantage. The CGo boundary is also more awkward than Rust's `cc` crate integration. + +### Pure Rust (rewrite everything) + +Porting the C/C++ inference code to Rust risks subtle behavioral differences in numerics-sensitive paths (BM25 scoring, embedding normalization). The `ort` crate for ONNX is solid but static linking of ONNX Runtime remains uneven across platforms. The incremental approach (Rust shell + C/C++ FFI) gets to a working binary faster and keeps the option open. + +### Extend C++ + +C++ is the right language for the inference engine but the wrong language for directory walking, YAML parsing, CLI dispatch, and JSON emission — which is 80% of the work. Libraries like `yaml-cpp` and `CLI11` exist but are a step down from `serde_yaml` and `clap`. The bash scripts exist precisely because C++ was too high-friction for the scripting layer. + +### Go + C/C++ library (CGo) + +Same FFI benefit as Rust + `cc`, but CGo cross-compilation is harder than Rust's `cargo-zigbuild`, and maintaining three languages (Go + C + C++) is worse than two (Rust + C/C++). + +## Interaction with Other ADRs + +| ADR | Interaction | +|-----|-------------| +| ADR-014 | `way-match` binary becomes `ways match` subcommand. BM25 algorithm unchanged | +| ADR-107 | Corpus generation becomes `ways corpus`. Locale support (Phase 3) becomes a flag | +| ADR-108 | `way-embed` binary becomes `ways embed` subcommand. ONNX/GGUF loading unchanged | +| ADR-110 | Graph export (`ways graph`) and sibling scoring (`ways siblings`) ship as subcommands rather than standalone scripts | + +## Extension: `attend` adopts the same pattern (2026-05-09) + +The `attend` binary (ADR-113) initially shipped a hand-rolled argv dispatcher with per-command help text written as free-form `println!` calls. As the surface grew to ~14 subcommands, the lack of a uniform help/argument-parsing layer started to bite — `--help` worked on some commands, errored as "unknown subcommand" on others, and silently ran the command on a third group. Rather than introduce a parallel CLI convention, `attend` adopts the clap-derive structure established here: a `Cli` struct, a `Commands` enum, doc comments as the source of truth for help text, and the same `agent_fmt::Banner` special-casing for bare invocation. External reference docs (`docs/cli/attend.md`) are generated from the same `Cli` definition via `clap-markdown`, wired into the project Makefile so the runtime help text and the published reference cannot drift. This decision retroactively confirms the ADR-111 pattern as the canonical CLI shape across the workspace. diff --git a/docs/architecture/platform/ADR-115-declarative-config-with-project-scope-overlay.md b/docs/architecture/platform/ADR-115-declarative-config-with-project-scope-overlay.md new file mode 100644 index 00000000..5293c7a3 --- /dev/null +++ b/docs/architecture/platform/ADR-115-declarative-config-with-project-scope-overlay.md @@ -0,0 +1,147 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: config +basis: + - evidence: attend's user-scope config plus project-scope +/- overlay, introduced during implementation, proved clean enough to propose as the workspace standard + - standard: XDG Base Directory conventions ($XDG_CONFIG_HOME, $XDG_CACHE_HOME) + - precedent: ADR-113 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-10 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-111 +imported: + from: docs/architecture/system/ADR-115-declarative-config-with-project-scope-overlay.md + format: v0 + status: Accepted +--- + +# ADR-115: Declarative Configuration with Project-Scope Overlay + +## Context + +`attend` (ADR-113) introduced a two-layer configuration pattern during implementation: a user-scope config at `~/.config/attend/config.yaml` and a project-scope overlay at `{project}/.claude/attend.yaml`. The project overlay uses `+/-` syntax to add or remove sensors without rewriting the full config. This pattern proved clean enough to propose as the standard for the agent-ways workspace. + +`ways` currently has no central configuration file. Its "config" is distributed across individual way files (frontmatter declares thresholds, vocabulary, triggers) and environment variables. This works for way authoring but leaves system-level tuning scattered: disclosure gate parameters live in the Rust source, embedding engine paths are hardcoded or env-var-driven, and there's no project-scope override for global behavior. + +This ADR proposes adopting attend's config pattern for the workspace — a shared convention that both `ways` and `attend` (and future sibling tools) follow. + +## Decision + +### The pattern + +Each tool in the agent-ways workspace may declare a config file at two scopes: + +``` +~/.config/{tool}/config.yaml # user scope — always loaded +{project}/.claude/{tool}.yaml # project scope — layered on top +``` + +User scope provides defaults. Project scope overrides or extends them. Tools load user scope first, then apply project scope on top. Missing files at either scope are a no-op (compiled defaults apply). + +### Config format + +YAML subset — flat keys, nested sections, lists. No full YAML parser required; the minimal subset that covers key-value pairs and two-level nesting is sufficient. This keeps tools zero-dependency (no serde, no yaml crate). + +### Project-scope overlay syntax + +For collection-type configs (sensors in attend, potentially way groups in ways), the project overlay uses `+/-` to modify the set: + +```yaml +# project/.claude/attend.yaml +sensors: + +disk-pressure: # add a project-local sensor + script: .claude/sensors/check-disk.sh + interval: 120 + -processes: # disable a user-scope sensor +``` + +The `+` prefix adds an entry that doesn't exist in user scope. The `-` prefix disables an entry from user scope. Unprefixed entries override properties of existing entries. + +### Trust model + +Same as ways scoping: + +- **User scope** (`~/.config/`) is trusted — the user installed it. +- **Project scope** (`{project}/.claude/`) has the same trust level as project-scope ways. Scripts declared in project config are code that runs on poll — same scrutiny as project-scope way macros. + +### CLI convention + +Each tool provides a `config` subcommand: + +``` +{tool} config init # write default config to user scope +{tool} config show # display effective config (both layers merged) +{tool} config path # show user and project config file paths +``` + +### What this means for ways + +`ways` could adopt this pattern for: + +- **Engine configuration**: embedding model path, corpus path, fallback behavior, forced engine selection — currently hardcoded or env-var-driven +- **Disclosure gate tuning**: re-disclosure intervals, token-gated thresholds — currently compiled defaults in Rust +- **Per-project way groups**: enable/disable way categories per project without removing files +- **Scoring overrides**: per-project BM25/embedding threshold adjustments + +Example: + +```yaml +# ~/.config/ways/config.yaml +engine: + model: minilm-l6-v2.gguf + fallback: bm25 + forced: auto + +disclosure: + redisclose_default: 10 + token_gate: 0.3 + +# project/.claude/ways.yaml +scoring: + +softwaredev/code/security: + threshold: 1.5 # lower threshold for security-sensitive project + -ea/comms: # disable comms way in this project +``` + +### XDG compliance + +Config files follow XDG conventions: +- Config: `$XDG_CONFIG_HOME/{tool}/` (default `~/.config/{tool}/`) +- State: `$XDG_CACHE_HOME/{tool}/` (default `~/.cache/{tool}/`) + +This is consistent with the existing XDG separation documented in project memory. + +## Consequences + +### Positive + +- **Tuning without recompiling.** Governor params, sensor intervals, scoring thresholds — all externalized. Edit a file, restart (or self-reload in attend's case). +- **Project-scope customization.** A security-focused project can lower security way thresholds. A hardware project can add system sensors. A documentation project can disable code-focused ways. +- **Shared convention.** All workspace tools follow the same pattern. Users learn it once. +- **Progressive adoption.** Tools can adopt the pattern incrementally. Attend shipped it first; ways can adopt it when ready without changing attend's implementation. +- **Zero dependency.** The minimal YAML parser handles the config subset without pulling in serde or yaml crates. + +### Negative + +- **Two files to manage.** Users must understand the layering. Mitigation: `config show` displays the effective merged result; `config path` shows where to look. +- **Minimal parser limitations.** The YAML subset doesn't handle anchors, multi-line values, or complex nesting. Mitigation: the config format is intentionally simple; complex configuration belongs in way files or sensor scripts, not in config.yaml. + +### Neutral + +- **attend already implements this.** This ADR documents the pattern for workspace-wide adoption. No changes to attend's existing implementation. +- **ways adoption is deferred.** This ADR defines the target; a follow-up implements it in ways when the need arises. + +## References + +- **attend config implementation**: `tools/attend/src/config.rs` +- **attend ADR**: [ADR-113](../attend/ADR-113-attend-active-awareness-module.md) — config section documents attend's implementation +- **XDG separation**: project memory `xdg-separation.md` diff --git a/docs/architecture/platform/ADR-116-declarative-permission-requirements.md b/docs/architecture/platform/ADR-116-declarative-permission-requirements.md new file mode 100644 index 00000000..3380fdef --- /dev/null +++ b/docs/architecture/platform/ADR-116-declarative-permission-requirements.md @@ -0,0 +1,206 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - config + - authoring + - attend +basis: + - evidence: the trusted-project-macros file is a binary, opaque, all-or-nothing trust model that attend sensors cannot use, and missing permissions fail silently or prompt mid-session + - standard: Claude Code settings.json permissions.allow and its permission evaluation, which the audit matches + - precedent: ADR-115 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-10 +deciders: + - aaronsb + - claude +related: + - ADR-004 + - ADR-113 + - ADR-115 +imported: + from: docs/architecture/system/ADR-116-declarative-permission-requirements.md + format: v0 + status: Accepted +--- + +# ADR-116: Declarative Permission Requirements + +## Context + +Ways macros and attend sensors both execute shell commands that require tool permissions in `~/.claude/settings.json`. Neither tool currently declares what permissions it needs. When a required permission is missing, the tool silently fails or prompts the user mid-session. + +The current permission model for way macros uses a flat file (`~/.claude/trusted-project-macros`) that lists project paths whose macros are allowed to run. This is a binary trust model — a project is either fully trusted or not. It has several problems: + +- **Opaque** — you can't see what a macro will do without reading the script +- **All-or-nothing** — trusting a project grants all its macros, not specific capabilities +- **Siloed** — attend sensors have no equivalent mechanism +- **Obscure** — the file is undiscoverable and undocumented in the tool itself + +Meanwhile, `settings.json` already has a structured `permissions.allow` list that governs what Claude Code can do. This is the real authority — but there's no way to validate that declared requirements align with granted permissions. + +## Decision + +Add a `requires:` field to both way frontmatter and attend sensor configuration that declares tool permissions. Provide audit commands that diff declared requirements against `settings.json` grants. + +### Schema + +**Way frontmatter** (new field in `frontmatter-schema.yaml`): + +```yaml +--- +description: GitHub workflow guidance +vocabulary: github pull request merge review +macro: prepend +requires: + - Bash(gh:*) + - Bash(git:*) +--- +``` + +**Attend sensor config** (in `attend.yaml` / `config.yaml`): + +```yaml +sensors: + +disk-pressure: + script: scripts/disk-check.sh + interval: 60 + requires: + - Bash(df:*) + - Bash(du:*) +``` + +**Built-in sensor defaults** — hardcoded in the attend binary, queryable via `attend permissions`: + +| Sensor | Requires | +|-----------|-----------------------------------| +| git | `Bash(git:*)` | +| processes | `Bash(ps:*)` | +| peers | `Read` | +| context | `Read` | + +### Wildcard + +A way or sensor may declare `requires: ["*"]` to indicate it needs arbitrary tool access. This is equivalent to the old `trusted-project-macros` blanket trust — the audit will flag it as "unrestricted" rather than listing specific gaps. This provides a migration path from the old model. + +### Audit Commands + +**`ways permissions audit`** — scans all way frontmatter (global + project-local), collects `requires:` fields, diffs against `settings.json`: + +``` +$ ways permissions audit + Way Requires Status + ──────────────────────────────────────────────────────── + softwaredev/delivery/github Bash(gh:*) MISSING + softwaredev/delivery/github Bash(git:*) granted + softwaredev/code/quality Bash(wc:*) granted + softwaredev/code/quality Bash(find:*) granted + meta/knowledge/authoring Bash(ways:*) granted + + 1 missing permission. Add to settings.json: + "Bash(gh:*)" +``` + +**`attend permissions audit`** — same pattern for sensor configs: + +``` +$ attend permissions audit + Sensor Requires Status + ────────────────────────────────────── + git Bash(git:*) granted + processes Bash(ps:*) MISSING + peers Read granted + +disk-pressure Bash(df:*) granted + +disk-pressure Bash(du:*) MISSING + + 2 missing permissions. Add to settings.json: + "Bash(ps:*)", "Bash(du:*)" +``` + +### Permission Format + +The `requires:` values use the same permission string format as `settings.json`: + +- `Read` / `Read(/path/**)` — file read access +- `Write(/path/**)` — file write access +- `Edit(/path/**)` — file edit access +- `Bash(command:*)` — specific bash command +- `Bash(*)` — any bash command +- `*` — unrestricted (wildcard) + +### Matching Semantics + +Permission matching follows a containment hierarchy, not string equality: + +- `*` covers everything +- `Bash(*)` covers any `Bash(command:*)` requirement +- `Bash(git:*)` covers `Bash(git:status)`, `Bash(git:diff)`, etc. +- `Read` (unscoped) covers `Read(/any/path)` + +The audit checks whether each declared requirement is **satisfied by** at least one granted permission. A grant of `Bash(*)` satisfies a requirement of `Bash(git:*)`. A grant of `Bash(git:*)` does NOT satisfy a requirement of `Bash(*)`. + +This matches Claude Code's own permission evaluation — the audit tells you exactly what Claude Code will allow or prompt for. + +### Config Location + +Both tools read `settings.json` from the standard Claude Code location (`~/.claude/settings.json`). Sensor configs follow the XDG pattern established by attend (ADR-115): user config at `$XDG_CONFIG_HOME/attend/config.yaml`, project overlay at `$PROJECT/.claude/attend.yaml`. + +### Replaces `trusted-project-macros` + +The `trusted-project-macros` file is deprecated. Project-local ways that declare `requires:` are validated against `settings.json` like any other way. A project-local way with `requires: ["*"]` is equivalent to the old "trusted project" entry. + +Migration: if `trusted-project-macros` exists, the audit command warns that it's deprecated and suggests adding explicit `requires:` fields to the relevant ways. + +## Consequences + +### Positive + +- **Discoverable** — `requires:` is visible in frontmatter, not buried in a separate file +- **Granular** — per-way and per-sensor permission declarations +- **Auditable** — one command shows the full permission gap +- **Consistent** — same model for ways and attend, using `agent-fmt` Table output +- **Single authority** — `settings.json` is the only place permissions are granted +- **Self-documenting** — reading a way's frontmatter tells you what it needs + +### Negative + +- Existing ways need `requires:` fields added (can be done incrementally) +- Built-in sensor requirements are hardcoded (but rarely change) + +### Neutral + +- Ways without `requires:` are assumed to need no special permissions (static-only ways) +- The audit is advisory — it doesn't block execution, it reports gaps +- `frontmatter-schema.yaml` gains one new field + +### Declarative Config for Ways + +Ways currently scatters configuration across dynamic shell checks (matcher selection, default language, corpus settings). Following attend's lead (ADR-115), ways should read from a config file at `$XDG_CONFIG_HOME/ways/config.yaml` with project overlay at `$PROJECT/.claude/ways.yaml`. + +This config would hold: +- `permissions:` — the `requires:` audit settings +- `matcher:` — force semantic over BM25, or vice versa +- `language:` — default locale (e.g. `en`) +- `corpus:` — rebuild policy, staleness threshold + +Dynamic checks in the code become config reads. When a config value is missing, the check says what to set rather than guessing — "have an opinion." This is the same pattern attend established and should be a shared convention across all agent-ways tools. + +Full config schema design is out of scope for this ADR but should be addressed alongside or immediately after implementation. + +## Implementation Plan + +1. Add `requires:` to `frontmatter-schema.yaml` (optional field, string array) +2. Add `requires:` parsing to attend sensor config +3. Implement `ways permissions audit` in ways-cli (reads frontmatter + settings.json) +4. Implement `attend permissions audit` (reads sensor config + settings.json) +5. Add `requires:` to existing macro-bearing ways (github, adr, quality, etc.) +6. Hardcode built-in sensor requirements in attend +7. Deprecation notice for `trusted-project-macros` in audit output +8. Update `docs/hooks-and-ways/macros.md` security model section +9. Add `requires` validation to `ways lint` — flag ways with `macro:` but no `requires:` as warning +10. Implement `ways lint --fix` auto-population — scan macro.sh for commands, generate `requires:` field +11. Containment-aware permission matching (not string equality) in the audit engine diff --git a/docs/architecture/platform/ADR-131-project-scope-way-toggles.md b/docs/architecture/platform/ADR-131-project-scope-way-toggles.md new file mode 100644 index 00000000..38eb74f6 --- /dev/null +++ b/docs/architecture/platform/ADR-131-project-scope-way-toggles.md @@ -0,0 +1,125 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - config + - matching +basis: + - evidence: the only enable/disable knob is disabled_domains in user-scope ways.json, which is domain-level and global, so per-project muting needs file deletion or JSON edits + - precedent: ADR-115 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-05-22 +deciders: + - aaronsb + - claude +related: + - ADR-115 + - ADR-105 + - ADR-111 +imported: + from: docs/architecture/system/ADR-131-project-scope-way-toggles.md + format: v0 + status: Accepted +--- + +# ADR-131: Project-scope way toggles + +## Context + +Ways fire across every project the user opens, but not every way is wanted in every project. A research-heavy repo doesn't need `itops/incident`; a personal scratch project doesn't want `documentation/adr` nagging. Today the only enable/disable knob is `disabled_domains` in user-scope `~/.claude/ways.json` — coarse (domain-level) and global (applies everywhere). The result: + +- Authors keep ways generic so they don't annoy users in unrelated projects, which weakens the matchers. +- Users tolerate noise rather than disabling a domain globally, because disabling globally would also kill the way in the one project where it *is* wanted. +- Per-project muting today requires deleting/renaming way files or editing user-scope JSON every time the user switches contexts — neither survives `git pull` and both leak between projects. + +ADR-115 introduced the project overlay (`{project}/.claude/ways.yaml`) for tuning thresholds and disclosure parameters, and `config.rs` already loads it. But the overlay has no per-way enable/disable schema, and no CLI surface — users would have to hand-edit YAML to use it. + +The need is narrow: *per-way, per-project, defaulted-enabled* toggles, with a CLI ergonomic enough that users actually use them. + +## Decision + +### Scope + +- **Per-way granularity.** Toggles target individual ways by their canonical name (e.g., `itops/incident`, `meta/introspection`), not domains. +- **Project scope only.** No new global disabler — the existing `disabled_domains` in user-scope `ways.json` is retained for backward compat but not extended. Per-way toggles live exclusively in `{project}/.claude/ways.yaml`. +- **Default enabled.** Absence of a toggle means the way fires normally. The config is opt-out, not opt-in. A project that ships no `ways.yaml` behaves exactly as today. + +### Schema + +Extend the project overlay with a `ways` mapping. Keys are way canonical names; values are `enabled: true|false`: + +```yaml +# {project}/.claude/ways.yaml +ways: + itops/incident: + enabled: false + meta/introspection: + enabled: false +``` + +Shorthand (boolean value) is also accepted for the disable-only case: + +```yaml +ways: + itops/incident: false + meta/introspection: false +``` + +The mapping form is reserved for future per-way knobs (threshold overrides, refire presets) — see Consequences > Neutral. + +### CLI + +Add two subcommands to the `ways` binary: + +``` +ways disable <way> # set ways.<way>.enabled: false in $PROJECT/.claude/ways.yaml +ways enable <way> # remove the entry (or set enabled: true) +ways disable --list # show currently disabled ways in this project +``` + +Both default to project scope. There is no `--global` flag — per the project-scope-only constraint. + +Behavior: +- Creates `.claude/ways.yaml` if missing. +- Round-trips comments and unrelated keys via a minimal YAML edit (not a full re-serialize). +- Validates that `<way>` exists in the corpus before writing; warns but still writes if not (allows pre-emptive disable before a way is authored). +- `ways enable <way>` is a no-op if the way isn't currently disabled — exits 0. + +### Enforcement + +Toggle resolution happens at the same gate as `disabled_domains` today, in `ways scan` (the Rust production firer) and `inject-subagent.sh` (the bash subagent gate). Both consult `config::global().disabled_ways` — a new `Vec<String>` populated from the project overlay during config load. + +A disabled way is skipped entirely: no scoring, no disclosure, no marker. It is as if the way did not exist for this session. + +### Trust model + +Same as the rest of `.claude/ways.yaml`: project-scope config is committed alongside the repo and reviewed like any other source file. Disabling a way is no more privileged than deleting it from the project's local `.claude/ways/` directory. + +## Consequences + +### Positive + +- **Authoring freedom.** Way authors can write sharper triggers without worrying about a project where the way doesn't belong — users can just disable it there. +- **Per-repo discipline.** A repo's `ways.yaml` becomes the canonical record of "which guidance applies here," reviewable in PR and durable across machines. +- **Reversible.** `ways enable` is a single command — no file deletion, no merge conflicts with upstream ways. + +### Negative + +- **Second place to look** when debugging why a way isn't firing — alongside corpus presence, threshold, refire window. Mitigated by `ways status` surfacing disabled ways and `ways scan --explain` reporting "skipped: disabled in project config." +- **Drift risk.** A way renamed upstream silently stops being disabled. Mitigated by `ways disable --list` warning on entries that don't match the current corpus. + +### Neutral + +- The `ways:` mapping form (vs. shorthand boolean) is forward-compatible with per-way threshold/refire overrides — those are deliberately out of scope for this ADR but the schema doesn't preclude them. +- Existing `disabled_domains` in user-scope `ways.json` is unchanged. Domain-level disable remains the right tool for "I never want any `ea/*` way to fire anywhere." + +## Alternatives Considered + +- **Per-way disable in user scope.** Rejected: the whole problem is that user-scope is too broad. A user who wants `itops/incident` muted in project A but active in project B can't express that globally. +- **Reuse `disabled_domains` with deeper paths** (e.g., `itops/incident`). Rejected: domain-level and way-level are conceptually distinct; collapsing them muddles the schema and the existing `disabled_domains` lives in user-scope JSON, not the project YAML. +- **`+/-` overlay syntax from ADR-115.** Considered for symmetry — `ways: [-itops/incident]`. Rejected for the toggle case: `enabled: false` reads more clearly to humans, and the `+/-` form is better reserved for collection-shaped configs (sensors, way groups) where addition is also meaningful. Disable is monotonic; no `+` half is needed. +- **Delete the way file from `~/.claude/hooks/ways/`.** Rejected: that's a global, destructive change that wipes the way for every project and doesn't survive a re-sync. diff --git a/docs/architecture/platform/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md b/docs/architecture/platform/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md new file mode 100644 index 00000000..7e3a0fec --- /dev/null +++ b/docs/architecture/platform/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md @@ -0,0 +1,219 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: install +basis: + - evidence: 'PR #182 (external contributor) packaged the subdirectory copy approach; its settings.json merge shipped broken because it lived in untestable command prose' + - evidence: check-config-updates.sh runs git against ~/.claude, so update detection goes dark when ~/.claude is not a repo + - precedent: ADR-138 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-22 +deciders: + - aaronsb + - claude +related: + - '[[ADR-138]]' + - '[[ADR-139]]' +imported: + from: docs/architecture/system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md + format: v0 + status: Accepted +--- + +# ADR-140: Two install topologies: in-place repo and subdirectory projection + +## Context + +Since its first release agent-ways has held one identity invariant: **the repo +*is* your config.** `~/.claude` is the git clone. That invariant is load-bearing +and almost everything in the install/update machinery quietly depends on it: + +- `git pull` (or `make update`) *is* the whole update — files are updated in place. +- `git status` in `~/.claude` tells the truth, because the installed files and the + source files are the same files. +- The update checker (`hooks/check-config-updates.sh`) just runs + `git -C ~/.claude rev-parse` and classifies the result (clone / fork / + renamed_clone / plugin) to decide whether to nudge. +- There is **zero drift** — nothing can be "installed but stale" relative to source. + +That invariant is also a wall for a specific, common adopter: someone who already +has a `~/.claude` they care about — existing sessions, credentials, their own +`settings.json` (model, theme, plugins) — and is unwilling to turn that directory +into a clone of someone else's repo. Today the installer offers them only +`--dangerously-clobber` (back up and replace) or "move it aside." For an **org +rollout to non-technical colleagues**, where the safe-by-default path matters most, +that is the wrong and only choice on offer. + +The subdirectory *copy* approach is not new — it has been informal practice for a +while. PR #182 (an external contributor) is the first time it has been **packaged**: +the repo lives in a subdirectory (e.g. `~/.claude/directory`), `~/.claude` stays the +user's own directory, and a new `/sync-to-home` command **projects** the repo's +outputs — skills, agents, commands, hooks, binaries — up into `~/.claude`, merging +only the hooks block and ways permissions into the existing `settings.json`. This +ADR formalizes existing practice rather than inventing a topology. + +This is not a small feature; it breaks the in-place invariant and forks the +deployment model. The break has a measurable blast radius, ranked: + +1. **Update detection goes dark.** `check-config-updates.sh` runs git against + `~/.claude`, which in this topology is *not* a repo. It falls through every + classifier and writes no cache, so the out-of-date nudge **never fires** — for + precisely the topology with the most fragile update path (a two-step the user + must remember: `git pull` *then* `/sync-to-home`). +2. **Divergent update workflows, single-topology messaging.** `update_status_text()` + and several docs hardcode `make update`; the subdirectory path is + `git pull && /sync-to-home`. The advice is wrong for these users. +3. **Drift becomes a first-class failure mode.** Copy-projection means "pulled but + didn't re-sync" is a new, silent, wrong state. Canonical never had it; nothing + detects it. +4. **The installer doesn't offer the gentle path.** `install.sh` knows only + clone-in-place and clobber. The safe option for existing configs exists only in + prose. +5. **Two `settings.json` writers** (`make install` and `/sync-to-home`) must agree + on what they own, or a topology-switcher is surprised. + +The decision underneath all five symptoms is **how** the projection happens: by +**copy** (what #182 does) or by **symlink** (`~/.claude/hooks → …/directory/hooks`, +etc.). That choice is too consequential to settle inside a command's bash, which is +why it is captured here. + +## Decision + +**Support two install topologies as first-class, with copy-based projection for the +subdirectory topology — made observable, so it reaches parity with the in-place +topology on update detection and messaging without taking on symlink fragility.** + +1. **Two named topologies.** + - **In-place** (default, unchanged): `~/.claude` *is* the repo. `make update`. + - **Subdirectory**: repo at `~/.claude/<dir>`, projected into `~/.claude` by + `/sync-to-home`. `git pull && /sync-to-home`. + +2. **Copy, not symlink, is the subdirectory default.** The target audience spans + mixed OSes including Windows (the #182 author already added Windows path-quoting); + symlinks there require developer-mode/admin and must be followed correctly by + Claude Code's hook execution. Copy is the form a non-technical adopter can + reason about ("it copied files in") and is robust everywhere. Symlink-projection + is recorded as a rejected alternative below, available to advanced users who opt + in, but it is not the blessed default. + +3. **The feature is delivered across all three ADR-138 substrates — way / skill / + make — each owning its proper layer:** + - **Mechanism (`make sync-to-home`).** A real script `scripts/sync-to-home.sh`, + surfaced as a Makefile target, mirroring the `make update` → `scripts/update.sh` + pattern of the in-place topology. It owns topology detection, backup, copy, the + `settings.json` merge, and the marker/stamp writes below — deterministically, + with no agent. This is why the `settings.json` merge shipped broken in #182: + living in command prose it was un-testable; as a script behind a `make` target + it is CI- and human-runnable and gets a smoke test. + - **How (`/sync-to-home` skill).** The instructional layer (*skills own the how*). + It tells Claude *what* to do and *to use the `make` target to do it*, carrying + the judgment the mechanism can't — confirm this is the subdirectory topology, + get consent before mutating `~/.claude`, invoke `make sync-to-home`, then report. + It no longer carries the bash. + - **Why (a topology way, `meta/deployment`).** Ways own the 5W. A new way carries + the *rationale* — why you'd pick in-place vs subdirectory, and copy vs symlink — + and fires when the topology question surfaces: install talk, an existing + `~/.claude` conflict, "how do I update," or helping a colleague deploy. It arms + Claude (often the very agent a non-technical adopter is pasting `curl | bash` + into) with the *judgment behind the choice*, not just the command to run. The + decision tree it discloses: existing `~/.claude` you value → subdirectory; + greenfield / want zero-drift simplicity → in-place; Windows / symlink-shy → + copy, never symlink. + +4. **Make the copy topology observable** so the blast radius closes: + - **Source marker.** `make sync-to-home` writes a `~/.claude/.claude-source` + marker recording the absolute path of the projecting repo. This reuses the + existing marker mechanism the update checker already honors for `renamed_clone` + (`.claude-upstream`). `check-config-updates.sh` learns a new branch: if + `~/.claude` is not itself a repo but a source marker points at one, run git + *there*, classify against upstream, and on "behind" emit advice tailored to the + topology: `cd <repo> && git pull && make sync-to-home`. + - **Synced-HEAD stamp.** `make sync-to-home` records the repo HEAD it last + projected (e.g. in the marker). A lightweight check compares the repo's current + HEAD to the last-synced HEAD; when they differ, surface a **"pulled but not + synced — run `make sync-to-home`"** nudge. This makes drift detectable rather + than silent. + +5. **Installer offers the topology as a conflict-menu option.** When `install.sh` + detects an existing, non-agent-ways `~/.claude`, it already presents a menu of + choices — today: (1) back up and clobber, (2) merge manually, (3) start fresh + (move aside). Every one of those either *destroys* the existing config or *punts* + to manual work; none keeps it. Subdirectory projection is added as a new, + **non-destructive** menu entry — "install alongside: clone into a subdir, keep + your `~/.claude`, run `make sync-to-home`" — defaulting to copy, with the symlink + model available as the advanced variant. This makes the safe path a first-class + choice at the exact moment of conflict, rather than something buried in the docs. + +6. **Messaging parity.** `update_status_text()` and the install/update docs gain a + subdirectory branch so every surface speaks the right update command for the + topology in use. + +PR #182 is the **starting point** for the copy-mode implementation. Its inline bash +is extracted into `scripts/sync-to-home.sh` (point 3); the script then gains the +marker and synced-HEAD writes; `make sync-to-home` wraps the script and the +`/sync-to-home` command is reduced to a narrating wrapper. The installer offer and +the messaging branch are the remaining follow-on work this ADR authorizes. + +## Consequences + +### Positive + +- Colleagues with an existing `~/.claude` get a safe, non-destructive install path — + the rollout's most common objection is answered without `--dangerously-clobber`. +- The dark-detection defect is closed: subdirectory installs get the same emphatic + out-of-date nudge as in-place ones, plus a drift nudge in-place installs never need. +- The copy approach keeps Windows and symlink-shy environments working. +- The projection becomes deterministic and testable: a real script behind + `make sync-to-home`, runnable without an agent, in CI, or by a human — closing the + class of defect (the broken `settings.json` merge) that only existed because the + logic lived in un-testable command prose. +- The decision and its alternatives are captured before the implementation hardens, + so the copy-vs-symlink fork is a recorded choice rather than an accident of bash. + +### Negative + +- Two topologies is genuinely more surface to maintain: every install/update change + must now be reasoned about twice, and tested on both. +- Copy-projection keeps drift as a real state; we mitigate it (synced-HEAD nudge) but + do not eliminate it the way symlink or in-place would. +- A new marker file and a new detection branch add moving parts to the update checker, + which is security-sensitive (it runs at session start). +- Session-start health checks must resolve through the binary, not hardcoded + `~/.claude/...` paths. Under symlink projection those paths are links into the subdir + repo, so a path-based check is topology-fragile — the embedding-engine notice in + `check-setup.sh` checked the wrong `way-embed` location and went stale once already. + The durable contract: probe the engine *functionally* (`ways match`), which resolves + identically across in-place, copy-subdir, and symlink-subdir installs. + +### Neutral + +- `make sync-to-home` and `make install` both touch `settings.json`; both now live in + the Makefile and this ADR makes their ownership explicit (hooks block + ways + permissions) but does not otherwise unify them. +- Leaves room for an opt-in symlink mode later without re-deciding the topology + question — only the projection mechanism would change. + +## Alternatives Considered + +- **Symlink-projection** (`~/.claude/hooks → …/directory/hooks`, same for `bin`, + `skills`, …). Architecturally the cleanest: `git pull` becomes the whole update, + **zero drift**, no sync step, and the update checker can `readlink` to the repo and + run git there — it restores every property the in-place topology has. Rejected as + the *default* because the rollout audience includes Windows and low-privilege + environments where symlinks are fragile (dev-mode/admin, and hook execution must + follow them), and because it interleaves repo-owned trees into the user's own + directory. Retained as a documented advanced opt-in. + +- **Keep subdirectory as a doc-only escape hatch** (merge #182, label it + "advanced/manual," invest nothing in detection/installer parity). Rejected: it + leaves update detection dark for the exact users who most need the nudge, which is a + defect, not a documentation gap — and undercuts the rollout goal of a safe, + *supported* path. + +- **Status quo — clobber or move-aside only.** Rejected: forces a destructive choice + on adopters with valuable existing config, which is the wall this whole effort + exists to remove. diff --git a/docs/architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md b/docs/architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md new file mode 100644 index 00000000..93c087ae --- /dev/null +++ b/docs/architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md @@ -0,0 +1,373 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: install +basis: + - standard: XDG Base Directory specification + - evidence: 'auto-update has never been safe: ~/.claude interleaves shipped files with the user''s settings, sessions and credentials with no manifest to tell them apart' + - precedent: ADR-140 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-29 +deciders: + - aaronsb + - claude +related: + - '[[ADR-140]]' + - '[[ADR-141]]' + - '[[ADR-112]]' + - '[[ADR-128]]' +imported: + from: docs/architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md + format: v0 + status: Accepted +--- + +# ADR-142: agent-ways 1.0 — XDG application distribution + +## Context + +Since its first release agent-ways has held one identity invariant: **`~/.claude` +*is* agent-ways.** You install by cloning the repo into `~/.claude`; the installed +files and the source files are literally the same files. ADR-140 named this the +**in-place topology** and made it the default. Almost everything in the install/update +machinery quietly depends on it: `git pull` (or `make update`) *is* the whole update; +`git status` in `~/.claude` tells the truth; the update checker just runs +`git -C ~/.claude rev-parse`; there is zero drift, because there is nothing to drift +*from*. + +That invariant is load-bearing, and it is also the source of every structural problem +1.0 exists to resolve. Three forces have converged: + +1. **The application and the user's workspace are the same directory, so we cannot + safely replace the application.** `~/.claude` holds agent-ways' shipped files (ways, + skills, hooks, `bin/`) *and* the user's own `settings.json` (model, theme, plugins), + their sessions, their credentials, and any ways they wrote themselves. Because these + are interleaved in one tree with no manifest distinguishing them, "update agent-ways" + cannot mean "replace the shipped files" — it can only mean "git pull and hope the + merge is clean." This is precisely why **auto-update has never been safe**: today + out-of-date is a *notification*, never auto-applied, because we cannot tell the + application's files apart from the user's and so cannot mutate one without risking the + other. + +2. **ADR-140 forked the deployment model but left the fork unresolved.** Faced with the + adopter who already has a `~/.claude` they value, ADR-140 introduced the + **subdirectory topology**: clone into `~/.claude/<dir>`, then *project* the outputs up + into `~/.claude`. It chose **copy** as the projection default (Windows / low-privilege + robustness) and explicitly recorded **symlink-projection as a rejected default, + retained as an advanced opt-in** — flagging in its own Neutral consequences that it + "leaves room for an opt-in symlink mode later without re-deciding the topology + question." Copy-projection, however, *creates drift as a first-class state* ("pulled + but didn't re-sync"), which ADR-140 then had to detect and nag about. We now have two + topologies, two update workflows, two `settings.json` writers, and a drift failure mode + that exists only because of the copy choice. + +3. **"global" conflates the application with the user.** The `ways` runtime today scans + two root classes — **global** (everything shipped, under `~/.claude`) and **project** + (per-repo ways). A way the *user* writes and a way agent-ways *ships* land in the same + global root, indistinguishable. There is no way to shadow a shipped way without editing + it in place — which the next update then clobbers. The "can't safely update" problem and + the "can't tell app from user" problem are the same problem viewed from the runtime side. + +The underlying realization: **agent-ways has outgrown being a dotfile and become an +application.** Applications that respect their host don't own the user's home directory; +they separate code from config from state from cache, install into well-known locations, +and replace their own code wholesale on update without touching the user's data. The XDG +Base Directory specification is exactly that contract. The cache half of it is already +here — derived corpus, embeddings, and the model live in `$XDG_CACHE_HOME` +(`~/.cache/claude-ways/` today, renamed to `agent-ways/` for one consistent name) and are regenerated, not version-controlled. 1.0 finishes the +job: it makes agent-ways a real XDG application and reduces `~/.claude` to the thin +**projection surface that Claude Code owns**. + +This is a 1.0-scale change: it re-decides ADR-140's defaults, restructures the runtime's +root model (→ child ADR-143), and replaces the install/update/repair scripts with a single +reconciler engine (→ child ADR-144). This ADR is the spine; it records the identity shift, +the state taxonomy everything else falls out of, and the lifecycle for getting existing +users there without stranding them. + +**A note on who this affects, and how much we know about them.** The migration-relevant +population is specifically the **in-place-clone adopters** — those running a `~/.claude`-*is*-the-repo +install that the migrator (§5, §8) must handle; everyone else is a fresh install. The maintainer +does **not** know that population's exact size. The passive signals available are weak and +ambiguous: roughly 19 stars and ~10 forks; release-asset downloads are low *and* undercounted +because adopters typically build from source (`make ways`) rather than pull a release artifact; +clone traffic is polluted by CI. The honest read is **small, real, and unknown-exact**, with the +forkers (~10+) the best proxy for in-place adopters. The deprecation lifecycle (§8) is therefore +designed to be **population-independent** — its safety must not, and does not, depend on knowing N. + +## Decision + +**Restructure agent-ways from "a clone that lives in `~/.claude`" into an XDG application +that *projects into* `~/.claude`.** `~/.claude` stops being the application and becomes a +thin, regenerable projection surface; the application, the operator's config, and the +durable session substrate move to their proper XDG homes. + +### 1. The state taxonomy (the core decision — durability and repair fall out of it) + +The single decision from which everything else derives is **classifying every file agent-ways +touches by its durability and ownership**, and putting each class in the XDG location whose +contract matches: + +| Location | Holds | Durability contract | +|---|---|---| +| `$XDG_DATA_HOME/agent-ways/` | **The application** — exactly what is on GitHub: ways, skills, hooks, `bin/`, docs. | Read-only to the user; **replaced wholesale on update**. Losing it is a re-install, not data loss. | +| `$XDG_CONFIG_HOME/agent-ways/` | **The operator's own** ways and macros, plus `ways.json`. | Durable; **never touched by update**. Out of every mutation's blast radius. | +| `$XDG_STATE_HOME/agent-ways/` | **Session substrate** — ledger, memory, focus — that should survive a `~/.claude` wipe. | Durable; survives reinstall/repair. (Boundary vs. Claude-Code-owned state is an open question, below.) | +| `$XDG_CACHE_HOME/agent-ways/` | **Derived** — corpus, embeddings, model. (Exists today as `claude-ways/`; the reconciler renames it.) | Regenerable; safe to delete; rebuilt on demand. | +| `~/.claude/` | **The projection** — a merged `settings.json` (hooks + ways permissions only) and the projected tree. | The irreducible **Claude-Code-owned floor**; regenerable from the manifest. | + +The payoff is that **durability and ownership are now structural, not conventional.** "Can we +auto-update?" stops being a judgment call and becomes a lookup: update mutates only +`$XDG_DATA` (replaceable by definition) plus one surgical merge into `settings.json`; it never +goes near `$XDG_CONFIG` or `$XDG_STATE`. Repair regenerates the projection from the manifest; +losing the projection loses nothing. This is the separation of concerns that the single-tree +model made impossible. + +### 2. Manifest-driven projection (symlink default, copy fallback, git-derived) + +A **manifest** is the contract describing what should exist in `~/.claude` and where each entry +points back into `$XDG_DATA`. The projection is *whatever materializes that manifest* — so the +ADR-140 "copy vs symlink" question stops being a topology fork and becomes an **implementation +detail of materialization**: + +- **Symlink (default, Unix).** `~/.claude/hooks → $XDG_DATA/agent-ways/hooks`, etc. With symlinks, + `git pull` in `$XDG_DATA` makes most updates **live with zero projection step and zero drift** — + it restores every good property the original in-place topology had, without owning the user's home. +- **Copy (fallback, Windows / low-privilege).** Where symlinks require developer-mode/admin or are + fragile, the same manifest is satisfied by copying. Drift is reintroduced here, but it is now the + *exception*, confined to environments that can't symlink, and the reconciler (child ADR-144) detects + and closes it. + +**The manifest is derived from the agent-ways git-tracked file set.** Git is the authoritative answer +to "what does agent-ways ship." This makes the manifest do double duty as the **app-vs-user +disentangler**: a file present in a project-owned directory (e.g. `hooks/ways/...`) but **not +git-tracked-by-agent-ways** is, by set difference, user-scope. The thing we could never compute in the +single-tree model — *which of these files are mine and which are the user's* — is now a `git ls-files` +set difference against ground truth. (Format and mechanics: child ADR-144.) + +### 3. Flip the default — supersede ADR-140's defaults (not ADR-140) + +**Projection becomes the default, and within projection, symlink becomes the default +materialization.** This **supersedes the *defaults* ADR-140 decided** — its in-place-default and its +copy-default — while leaving ADR-140 *valid as the origin of the two-materialization idea*. ADR-140 +explicitly parked "an opt-in symlink mode later, without re-deciding the topology question" in its +Neutral consequences; **this ADR is that future, promoted from opt-in to default.** The in-place model +does not vanish — it becomes the legacy from-state the migrator reconciles away from (§5, child ADR-144). + +### 4. One reconciler engine, four from-states + +There is **no separate installer, updater, repairer, and migrator.** There is **one state-reconciler** +that drives the live `~/.claude` tree toward the manifest, entered from four different source states: + +| From-state | What reconciliation means | +|---|---| +| **fresh** (no prior install) | Materialize the manifest from nothing. | +| **drifted** (projection damaged) | Re-materialize the missing/broken entries. *Repair.* | +| **out-of-date** ($XDG_DATA behind) | Advance $XDG_DATA, re-derive manifest, re-materialize the delta. *Update.* | +| **legacy-in-place** (`~/.claude` *is* an old clone) | Relocate the app to $XDG_DATA, lift user files to $XDG_CONFIG/$XDG_STATE, replace the tree with the projection. *Migrate.* | + +Collapsing four tools into one engine is the design's central SOLID claim: the *logic* (converge the +tree to the manifest, idempotently) has **one reason to change**; the four entrypoints differ only in +*starting state and trust posture*, not in mechanism. Detailed engine, manifest format, and bootstrap: +**child ADR-144**. + +### 5. The autonomy gradient (the load-bearing safety model) + +The same engine runs under **two trust postures, keyed on blast radius** — this is what makes "one +engine for both a silent repair and a destructive migration" safe rather than reckless: + +- **update / repair → act silently and autonomously.** They touch only replaceable app-scope + (`$XDG_DATA` + the regenerable projection) and the surgical `settings.json` merge. Reversible, low + blast radius, no consent needed — consistent with the framework's silent-on-success convention. +- **migration → a higher tier.** Migration rewrites the user's actual `~/.claude`. It must: **back up + the whole `~/.claude` first**; **announce loudly** (with the backup path); **gate the first run behind + explicit consent** (a `migrate` skill / installer invocation — never a silent SessionStart side + effect); and be **crash-safe and resumable** (marker-driven phases, back-up-before-mutate, atomic + moves). A half-migrated `~/.claude` is worse than an un-migrated one; the kept backup is the escape + hatch behind "things are broken." + +### 6. Manifest check on every SessionStart (escalating, silent-on-success) + +A SessionStart manifest check, mirroring the framework's existing escalation convention: + +- **all good → silent.** (As `ways init` / check-setup are silent today when healthy.) +- **fixable drift → repair, then emit "X was repaired."** +- **unfixable → emit "things are broken" loudly**, pointing at the backup / re-install path. + +The chicken-and-egg this raises (if hooks are gone, nothing runs to repair them) and its resolution +(the self-sufficient bootstrap hook, and why that one entrypoint must be *exempt* from the breakage it +repairs) are **child ADR-144's** core problem. The one cross-cutting fact the spine must state: because +`settings.json` is read once at session start, a repaired hook **takes effect next session** — so the +first post-install run is a one-time **bootstrap → "start a new session to pick up agent-ways" → +restart** flow. + +### 7. Auto-update, finally safe — decomposed + +Auto-update was unsafe because *any* update risked clobbering user customizations. The taxonomy removes +the risk by construction: + +- User ways live in `$XDG_CONFIG` — **never in the blast radius** of an update. +- Under symlink projection, `git pull` in `$XDG_DATA` makes most updates **live with zero projection**. +- The **only** mutation to a user-owned file is the **surgical `settings.json` merge** (hooks block + + ways permissions). That single shared-write seam is the entire remaining risk surface — and it is + named, bounded, and testable rather than diffuse. + +Therefore the stance flips: **out-of-date becomes auto-applied** (silently, per the gradient), not a +notification the user must act on. This retires ADR-140's drift-nudge machinery for the default +(symlink) path; it survives only as the copy-fallback's exception handling. + +### 8. Deprecation lifecycle (concrete, and population-independent) + +Existing in-place users must reach the new architecture without being stranded. The lifecycle is +designed so that **the absolute size of the adopter base is irrelevant to its safety** — the maintainer +does not know N exactly (small, real, unknown-exact; see Context), and the design must not depend on +knowing it. + +The mechanism that makes this work: **1.1.0 is a per-user, update-gated transition, not a global flag +day.** Each in-place user encounters the 1.1.0 "last migration release" gate **only when *they* update +to it, on their own clock** — there is no calendar date at which anyone is cut off. + +- **1.0** — new architecture ships; the migrator (legacy-in-place from-state) goes live. Migration path + first viable. +- **1.0.x** (1.0.1 … 1.0.N, N may be large) — migration supported throughout. The in-place check and the + assisted migrator run for everyone still in-place. +- **1.1.0** — the **final** migration release. *When a user updates to it*, it migrates them (or confirms + they already are) and announces that the *next* release will no longer check or migrate: from 1.1.0's + successor on, ensuring `~/.claude` is how you want it becomes the user's responsibility, and agent-ways + manages only itself under the concern-separated architecture. +- **post-1.1.0** — XDG-only; no in-place checking or migration. We distinguish **"remove the migrator"** + (stop assisted migration and the in-place check) from **forcibly breaking anyone still in-place**: the + design must not strand un-migrated users — removing assistance is not the same as actively breaking them. + +Two properties fall out of the per-user gate, and they are the whole point: + +- **"Remove the migrator" is safe by construction, not by hope.** Releases are sequential, so a user + cannot reach a post-1.1.0 release without having passed *through* 1.1.0's gate — which migrated them on + the way. By the time the migrator is gone, everyone who got there was migrated *en route*. Safety never + depends on "did everyone migrate by some date"; the update path itself enforces it. This is exactly why + N is irrelevant. +- **A user who never updates past 1.1.0 stays frozen at 1.1.0 — working but unsupported.** This is an + **explicitly acceptable outcome**, not a failure: their in-place install keeps functioning; they simply + stop receiving updates and assisted migration. Nobody is stranded by a date — they self-select out by + declining to update, and what they decline into is a working frozen version. + +## Consequences + +### Positive + +- **Auto-update becomes safe and can be turned on**, because durability is now structural: the only + user-owned write is one bounded `settings.json` merge, and user ways are physically outside the blast + radius. +- **App, config, state, and cache are separable for the first time** — clean separation of concerns. A + `~/.claude` wipe loses nothing durable; a re-install is a manifest re-materialization; corrupted cache + is regenerated. +- **The copy-vs-symlink fork collapses** into one manifest with two materializations; symlink restores + the original in-place model's zero-drift, zero-step update without owning the user's home. +- **Users can shadow a shipped way** (via `$XDG_CONFIG`, child ADR-143) instead of forking it in place and + losing the edit on the next update. +- **Four lifecycle tools collapse into one reconciler** with a single reason to change; the four + entrypoints are starting-state + trust-posture, not duplicated logic. +- **The app-vs-user question becomes computable** — a `git ls-files` set difference against ground truth, + not a heuristic. + +### Negative + +- **The `settings.json` merge is a genuine shared-write seam and the design's leakiest boundary.** Two + writers (Claude Code owns the file; agent-ways surgically merges hooks + ways permissions) write one + file agent-ways does not own. This is the one place the otherwise-clean separation of concerns breaks, + and it is exactly where a botched merge can corrupt a user-owned file. It must be idempotent, + key-scoped (touch only hooks + ways permissions), and backed up before write. Calling it out honestly: + this seam is the residual risk the whole taxonomy could not eliminate, only shrink. +- **Migration is the highest-risk operation agent-ways has ever performed** — it rewrites the user's home + config tree. Crash-safety, resumability, and a full backup are mandatory, not optional; a half-migrated + `~/.claude` is worse than not migrating. The risk is real even with the gradient. +- **More moving parts and more locations.** A user debugging their setup now reasons across four XDG roots + plus the projection, rather than one directory. `git status` in `~/.claude` no longer tells the truth; + the truth is in `$XDG_DATA`. This is a real loss of the in-place model's one-directory legibility. +- **Symlink-vs-copy divergence persists at the edges.** Windows / low-priv installs still get copy, still + get drift, and still need the exception-handling path — the fork is narrowed to a fallback, not removed. +- **Bootstrap is a known race with a one-time restart UX cost** (§6); the first post-install session is a + bootstrap-then-restart, which is friction at the worst possible moment (first impression). +- **A long 1.0.x tail of dual-mode support.** For the whole 1.0.x window we maintain both the in-place + check / migrator and the XDG runtime — the very "reason about it twice" cost ADR-140 already flagged, + extended across the deprecation window. + +### Neutral + +- `$XDG_CACHE_HOME/agent-ways/` already exists today as `claude-ways/` and already behaves as + derived/regenerable state; 1.0 ratifies the convention rather than inventing it, and **harmonizes the + name to `agent-ways`** so all four XDG tiers share one application name. The rename + (`claude-ways` → `agent-ways`) is a one-time move the reconciler performs; because the tier is + regenerable, a missed rename costs only a rebuild. +- The session ledger (ADR-112) and the KG evidential backend (ADR-141) become tenants of `$XDG_STATE`; + their "survive a `~/.claude` wipe" expectation is exactly what the state tier provides. Memory routing + (ADR-128) is unaffected in intent but its files' XDG-vs-Claude-Code ownership is an open question. +- The two-topology framing of ADR-140 is folded into a one-manifest / two-materialization framing; ADR-140 + is superseded only on its *defaults*, and remains the cited origin of the materialization choice. + +## Alternatives Considered + +- **Keep the single-tree model; make update smarter.** Try to teach the updater to diff and preserve user + edits inside the shared tree. Rejected: this is the status quo's unwinnable problem — without a manifest + grounding "what we ship" in git truth, app and user files are indistinguishable, so a "smart" merge is a + heuristic that will eventually clobber something. The taxonomy makes the distinction structural instead + of heuristic. +- **Copy-projection as the default (ADR-140's choice), just better instrumented.** Rejected as the + *default*: copy makes drift a permanent first-class state, which then needs continuous detection and + nagging. Symlink eliminates drift on the platforms that can symlink (the majority), so copy is correctly + the *fallback*, not the default. ADR-140's instinct was right for its Windows-first rollout framing; 1.0 + re-weights for the whole population. +- **Put everything (app + config + state) under one new `~/.agent-ways/` directory.** Rejected: it would + re-create the single-tree conflation in a new location and ignore the XDG contracts that *exactly* encode + the durability classes we need. The value is in the *separation*, not in merely moving out of `~/.claude`. +- **Stay in-place forever; never migrate.** Rejected: it permanently forecloses safe auto-update and + shadowable user ways, which are the two capabilities driving 1.0. But its *spirit* is honored in the + deprecation lifecycle's "don't strand un-migrated users" clause. + +## Open Questions + +These are recorded deliberately undecided; they do not block the Draft but must be resolved before +Accepted: + +- **`$XDG_STATE` vs. Claude-Code-owned boundary.** Which of sessions / memory are Claude-Code-owned (and + stay in `~/.claude`) vs. agent-ways-owned (and move to `$XDG_STATE`)? The ledger and focus are clearly + ours; auto-memory (ADR-128) sits in Claude Code's `projects/<slug>/memory/` and may not be ours to move. +- **Exact length of the 1.0.x migration window** before 1.1.0 retires assisted migration. +- **Whether to add a lightweight, privacy-respecting adoption signal** so future lifecycle decisions + aren't blind. Note this is **explicitly not a prerequisite**: passive GitHub signals (forks, release + downloads) already exist, and the per-user-gated deprecation (§8) is population-independent by design — + so telemetry would only inform *future* judgment calls, never gate the safety of *this* lifecycle. Any + such signal must clear the project's own anti-surveillance bar before it's worth adding. +- **Whether the two children stay separate ADRs or fold into this spine** — see ADR-143 / ADR-144; + current recommendation is to keep them separate (below). +- Bootstrap-exemption mechanism and Windows materialization specifics are parked in **child ADR-144**. + +## Amendment (2026-07-05): source-pinned deploys — `ways update --ref` + +§4's **out-of-date → update** from-state advances `$XDG_DATA` to the *latest +release* on the tracked branch and refreshes binaries download-first (§7). A +development need sits alongside it: deploying an **unpublished ref** — a feature +branch, a tag, or a bare commit — onto a live install to dogfood it, without +publishing to `main` and cutting a release. + +`ways update --ref <branch|tag|sha>` serves that need as a variant of the update +from-state, with three deliberate differences from the release-channel path: + +- **Fetch + detached checkout of the ref**, not a pull of the tracking branch. + The checkout lands on a detached HEAD at the ref; `ways update --ref main` + returns to the release channel. +- **Build the whole suite from source** (`ways-rebuild` and the sibling + `*-rebuild` targets, plus way-embed's source target) — an unpublished ref has + no pre-built release binary to download, so download-first (§7) does not apply. +- **The ADR-150 downgrade guard is bypassed** (see that ADR's amendment): pinning + an explicit ref is a deliberate choice, not a channel update, so "never move + backward" is not the right invariant. + +The reconciler tail is unchanged: after the source build it runs the same relink ++ corpus regen + `ways reconcile` projection as any other update, so the trust +posture (§5: app-scope plus the one `settings.json` merge) is identical. This +stays a developer/dogfooding affordance layered on the update from-state, not a +fifth reconciler mode. diff --git a/docs/architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md b/docs/architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md new file mode 100644 index 00000000..120213f9 --- /dev/null +++ b/docs/architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md @@ -0,0 +1,313 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: install +supersedes: + - ADR-145 +basis: + - evidence: 'install.sh, update.sh and sync-to-home.sh duplicate backup, projection and settings-merge logic, which is how the broken settings-merge shipped in #182' + - evidence: sync-to-home.sh's .claude-source-manifest is copy-mode-only and built from hardcoded subtrees, not git + - precedent: ADR-142 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-29 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-140]]' + - '[[ADR-133]]' + - '[[ADR-128]]' +imported: + from: docs/architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md + format: v0 + status: Accepted +--- + +# ADR-144: Install / repair / migrate as one manifest reconciler + +## Context + +This is a child of ADR-142 (agent-ways 1.0). The spine declared **one reconciler engine, +four from-states** (fresh / drifted / out-of-date / legacy-in-place) and an **autonomy +gradient** (silent for update/repair, gated-and-loud for migration). This ADR records the +engine: the **manifest** it converges toward, the **bootstrap** that lets it run at all, and +the **migrator** plus its deprecation lifecycle. + +The current reality is several scripts with overlapping jobs: + +- `install.sh` — clone-in-place or `--dangerously-clobber` (ADR-140 §5). +- `scripts/update.sh` / `make update` — `git pull` for the in-place topology. +- `scripts/sync-to-home.sh` — the subdirectory projector. It already does most of what the + engine needs: it has **both `copy` (default) and `symlink` modes** (`SYNC_MODE`, lines + 11-33), projects trees idempotently (`project_tree`, lines 68-90, which already skips when a + link/dir already resolves to source), backs up before overwrite (lines 47-55), merges + `settings.json` for hooks + ways permissions only (lines 167-198), and writes a source marker + + synced-HEAD stamp (`.claude-source`, lines 189-199). +- `hooks/check-config-updates.sh` — the out-of-date detector that runs at session start. + +Two facts from the real script are load-bearing for this ADR: + +1. **A manifest already exists** — `.claude-source-manifest` (lines 143-165). In copy mode the + script enumerates the source subtrees it projects (`build_manifest`, lines 144-156: + `skills`, `agents`, `commands`, `hooks/ways`, two named hooks, four named binaries), diffs + against the previous manifest, and **prunes only orphans it previously created** — explicitly + leaving user-owned files in the same shared dirs untouched ("User-owned files in the same + shared dirs are never in the manifest, so they are never touched," line 141). The + app-vs-user disentangling ADR-142 frames as the manifest's "double duty" is **already the + stated rationale of this code** — but the manifest is built by *filesystem enumeration of + hardcoded source subtrees*, not from git. +2. **The manifest is copy-mode-only.** Symlink mode "reflects source live — no orphans" (line + 142), so it writes no manifest. There is today no single artifact that describes the desired + `~/.claude` state across *both* materializations. + +So the engine is less a green-field build than a **consolidation and a grounding correction**: +unify the scripts into one reconciler, lift the manifest from a copy-mode pruning side-effect to +*the* projection contract, and re-base it on git truth. + +## Decision + +**One idempotent reconciler that converges `~/.claude` toward a git-derived manifest, plus a +self-sufficient bootstrap that is exempt from the breakage it repairs, plus a gated crash-safe +migrator with an explicit deprecation lifecycle.** + +### 1. Manifest format, derived from the git-tracked file set + +The manifest is the list of projection entries — for each, the relative path in `~/.claude` and +the source path in `$XDG_DATA` it points back to. It is **derived from `git ls-files` in +`$XDG_DATA/agent-ways`**, not from filesystem enumeration of hardcoded subtrees. This is the +correction to today's `build_manifest`: + +- **Git is the authoritative "what we ship."** A way/skill/command added or removed upstream is + reflected automatically; the hardcoded `for t in skills agents commands hooks/ways` list + (line 146) stops being a thing anyone can forget to update. +- **The set difference becomes exact.** "What is app vs. user" is `git ls-files` (app) vs. what's + present in the shared dirs (everything); the complement is user scope. Today's manifest gets + the *spirit* right (don't touch what we didn't write) but its membership test is "did a prior + projection write it," which is weaker than "is it git-tracked-by-agent-ways." + +Both materializations (symlink default, copy fallback — ADR-142 §2) are driven from the *same* +manifest; symlink mode now also has a manifest (it just materializes entries as links), closing +the "no single desired-state artifact" gap. + +### 2. The four from-states are one convergence with different entry conditions + +The engine computes (desired = manifest) vs. (actual = live `~/.claude`) and applies the delta +idempotently. The four from-states differ only in *starting actual* and *trust posture*: + +- **fresh** — actual is empty → materialize all entries. +- **drifted** — actual is missing/broken entries → re-materialize just those (repair). +- **out-of-date** — `$XDG_DATA` HEAD advanced → re-derive manifest, materialize the delta, prune + orphans (the existing `comm -23` orphan logic, lines 159-163, generalized). +- **legacy-in-place** — actual *is an old clone* → the migrator (§4) first relocates, then this + same convergence runs. + +Idempotence is already the discipline in `project_tree` (skip when already linked/resolved, lines +75-85); the engine generalizes it to every entry and to the prune step. + +### 3. Bootstrap and the exemption (the engine's hardest problem) + +**Chicken-and-egg:** the reconciler runs from a SessionStart hook; if the hooks are gone, nothing +runs to repair them. Resolution: + +- The **installer detects Claude Code and writes the minimum self-sufficient hook** into + `~/.claude` — the irreducible Claude-Code-owned floor (ADR-142's `~/.claude` row). On SessionStart + that hook invokes the manifest check/repair. +- **CRITICAL exemption: the bootstrap/repair entrypoint must not be a symlink into the tree it + manages.** If the healer is itself a link into `$XDG_DATA`, a broken `$XDG_DATA` (or a wiped link) + takes out its own healer. So the bootstrap entrypoint is **the one deliberate copy** (a small + copied script) *or* a call to a binary at a **stable absolute XDG path** — never a link into the + managed tree. This is a deliberate, named exception to the symlink default, justified precisely + because it is the thing that repairs the symlinks. +- **The restart seam.** `settings.json` is read once at session start, so a repaired hook **takes + effect next session**. The first post-install session is therefore a one-time **bootstrap → + "start a new session to pick up agent-ways" → restart** flow. The reconciler must make this + legible (say it plainly), not silently leave the user in a half-armed session. + +### 4. SessionStart check — escalating, silent-on-success + +Mirroring the framework's existing convention (and today's silent `ways init` / `check-setup`): + +- all entries satisfied → **silent**; +- fixable drift → repair autonomously, emit **"X was repaired"**; +- unfixable → emit **"things are broken"** loudly, pointing at the backup / re-install path. + +This replaces `check-config-updates.sh`'s in-place-only detection with a manifest check that works +across all materializations; the in-place git classifier survives only for the legacy from-state +during the deprecation window (§5). + +### 5. The migrator and its deprecation lifecycle + +Migration (legacy-in-place → XDG) is the **higher trust tier** (ADR-142 §5): it rewrites the user's +actual `~/.claude`. + +- **Gated, not silent.** Entered only via an explicit `migrate` skill / installer invocation — never + a SessionStart side effect. +- **Back up first, announce loudly** with the backup path (the existing pre-overwrite backup, lines + 47-55, generalized from "the projected dirs" to "the whole `~/.claude`"). +- **Crash-safe and resumable.** Marker-driven phases, back-up-before-mutate, atomic moves. A + half-migrated `~/.claude` is worse than an un-migrated one; on crash the next run resumes from the + last completed phase, and the backup is the escape hatch. +- **Migration ground truth is git.** What to *relocate to `$XDG_DATA`* (the app) vs. *lift to + `$XDG_CONFIG`/`$XDG_STATE`* (the user's ways, sessions, settings) is the same `git ls-files` set + difference as §1: app files are git-tracked-by-agent-ways; everything else in the old clone is the + user's and must be preserved, not replaced. + +**Deprecation lifecycle** (concrete, per ADR-142 §8). The governing property: **the transition is +per-user and update-gated, not a global flag day.** A user meets the 1.1.0 gate only when *they* update +to it, on their own clock — so the engine never needs to know how many in-place adopters exist (the +maintainer doesn't; ADR-142 Context). The migrator is invoked *by an update reaching the user*, not by a +date reaching the population. + +The lifecycle is an **escalating-pressure curve across the 1.0.x line**, dormant at first so adopters +meet the migrator before they're pushed by it: + +- **1.0.0** — the migrator ships **functional but dormant**: `ways migrate` (plan / what-if / execute) + is available on demand, with **no SessionStart pressure**. The release exists to make the path real + and prove it end-to-end; updating to it is idempotent-in-use to 0.9.x (the runtime behaves identically + until the operator chooses to migrate). This is the verify-before-you-push release. +- **1.0.1 … 1.0.5** — the migrator stays available while a SessionStart nudge **escalates each release**, + keyed to version and **silent for already-migrated installs**: a gentle offer at 1.0.1 ("agent-ways can + migrate — want to?") → progressively more insistent → at **1.0.5**, the last-chance warning that the + *next* release drops the migrator and that migrating after this point means checking out the 1.0.5 tag. +- **1.1.0** — migrator and in-place check **removed**. Safe by construction: releases are sequential, so + reaching 1.1.0 means having passed through the escalation. Two properties make the removal humane: + - **The escape hatch is the immutable tag.** A user still in-place migrates with + `git clone --branch ways-v1.0.5 … && ways migrate` — the migrator lives *forever* at that tag, so + "we removed it" never means "you can't." This is the concrete form of the never-strand clause. + - **The cliff is a ramp.** "Remove assistance" is not "break the un-migrated." A 1.1.0 binary on an + un-migrated `~/.claude` still **reads correctly** through the transition fallbacks — core ways from + the clone's own `hooks/ways`, cache from the legacy `claude-ways`, events from `~/.claude/stats`. It + loses assisted migration and in-place *checking*, not *function*. A user frozen pre-1.1.0 stays + working and unsupported by their own choice; a user who updates *to* 1.1.0 while in-place keeps + running via the fallbacks. The design never actively breaks them. + +(The specific patch numbers above are the *cadence*, not a contract — the mechanism is "pressure level +keyed to release, escalating to a final 1.0.x before 1.1.0," and any 1.0.x spacing satisfies it.) + +## Consequences + +### Positive + +- **One engine, one reason to change.** The convergence logic lives once; install/update/repair/ + migrate become entry conditions, not four scripts to keep in sync (the SOLID payoff ADR-142 §4 + claims, realized here). +- **The manifest stops being a copy-mode afterthought** and becomes the single desired-state contract + for both materializations, git-grounded — which is also what makes the app-vs-user split exact + rather than "whatever a prior sync happened to write." +- **Repair is real and idempotent**, reusing the already-idempotent `project_tree` discipline; a + damaged projection self-heals at the next SessionStart, silently when it can. +- **Migration is auditable and recoverable** — gated, backed up, resumable — rather than a one-shot + destructive script. +- Much of this is **consolidation of code that already exists and works** (copy/symlink modes, + backup, settings merge, orphan prune, marker/stamp), lowering implementation risk. + +### Negative + +- **The bootstrap exemption is a permanent, deliberate violation of the symlink default** — a copied + (or stable-path-binary) entrypoint that must be kept in sync with the engine it launches by some + *other* mechanism than the manifest, because it can't be a managed link. If that copy goes stale, + the healer and the thing it heals disagree. This is irreducible complexity, not an oversight. +- **The git-derived manifest assumes `$XDG_DATA` is a clean agent-ways checkout.** If a user has + hand-edited core files in place (exactly what pre-1.0 forced them to do to customize), those edits + are now "drift from git truth" and the reconciler may revert them — the migrator must detect and + rescue such edits into user scope, or it silently destroys the customization it was meant to + preserve. This is the migration's sharpest edge. +- **Migration crash-safety is genuinely hard to get right** and easy to get subtly wrong; resumable + marker-driven phases over a directory the user also has open in live sessions is a concurrency + problem, not just a sequencing one. +- **The `settings.json` merge remains the shared-write seam** (ADR-142 Negative) and is the one engine + output that touches a user-owned file; it must be idempotent and key-scoped or the engine corrupts + the very floor it depends on. +- **A long dual-mode window:** the in-place git classifier and the manifest reconciler both ship for + all of 1.0.x. + +### Neutral + +- `.claude-source` (marker) and `.claude-source-manifest` already exist (ADR-140 observability); this + ADR re-bases the manifest on git and extends it to symlink mode, but the marker mechanism is reused, + not invented. +- The session substrate this engine must preserve across migration (ledger, focus, and possibly memory) + is defined by ADR-142's `$XDG_STATE` boundary open question and ADR-128's memory ownership — the engine + inherits whatever that boundary resolves to. + +## Alternatives Considered + +- **Keep separate install / update / sync / repair scripts; just fix the manifest.** Rejected: the four + scripts already duplicate backup, projection, and settings-merge logic three ways (`install.sh`, + `update.sh`, `sync-to-home.sh`), which is how the broken settings-merge shipped in #182 (ADR-140) — + duplicated lifecycle logic is the defect generator. One engine is the structural fix. +- **Make the bootstrap hook a symlink like everything else** (no exemption). Rejected: a broken target + takes out the healer; the entrypoint that repairs the tree cannot live inside the tree it repairs. +- **Keep the filesystem-enumeration manifest** (today's `build_manifest`). Rejected: it requires a + hand-maintained subtree list, and its "is it ours" test ("did a prior sync write it") is weaker than + git truth — it can't classify a file the user dropped into a shared dir that happens to match a name we + later ship. Git ls-files is the authoritative membership test. +- **Auto-migrate silently at SessionStart** (no gate). Rejected outright by ADR-142's autonomy gradient: + rewriting the user's home config without consent and without a loud backup announcement is exactly the + blast radius the gradient exists to prevent. + +- **Bootstrap via a thin marketplace plugin.** *Viable alternative, not the chosen approach* — recorded + here so the reviewer can weigh it against §3's self-sufficient copied/stable-path entrypoint, which + stands. (Research-grounded: Claude Code plugins-reference plus adversarial cross-verification of the + named issues below.) + + - **The crux is documented to work.** A plugin can ship a `SessionStart` hook that auto-fires the + moment the plugin is *enabled*, with **no `settings.json` edit by the user** (plugins-reference; the + shipped `explanatory-output-style` plugin is an existing example of an auto-firing plugin hook). The + hook file lives in Claude-Code's managed plugin cache + (`~/.claude/plugins/cache/<marketplace>/<plugin>/<version>/`), which is **independent of agent-ways' + own `~/.claude` projection** — a different tree, owned by a different system. + + - **Effect on the bootstrap exemption: it *relocates* the exemption, it does not eliminate it.** Because + the hook is a real file in a Claude-Code-managed cache (loader-guaranteed to exist, *not* a symlink + into the projected tree agent-ways manages), it satisfies §3's exemption requirement by construction. + The burden of "guarantee a stable entry point" shifts **from agent-ways** (a copied file we keep in + sync) **to Claude Code's plugin loader**. That is a genuinely cleaner exemption — but it is a + **dependency on a third-party surface, not the removal of the dependency.** + + - **Corrected risk-ranking — the exemption is *not* the headline risk here.** The headline risk is + **coupling a load-bearing bootstrap to Anthropic's plugin subsystem, which agent-ways does not control + and which is visibly churning**: an ephemeral, version-stamped `CLAUDE_PLUGIN_ROOT` (#15642); a + stale-marketplace-clone update bug (#35752); repeated Windows symlink regressions + (#23819 / #24140 / #53948 / #50052); and an uninstall path that deletes the plugin *and* its + `CLAUDE_PLUGIN_DATA` — which would make the projection persist while its **healer vanishes**. A + secondary facet is *fit with the official model*: plugins are meant to *contribute artifacts*, not + *mutate `~/.claude`*. A bootstrap hook that symlinks and merges `settings.json` into `~/.claude` is + **technically allowed** (hooks run unsandboxed at full user privilege) but is **off-label** use of the + plugin contract. + + - **Three hard constraints if pursued:** (1) the bootstrap must resolve `$XDG_DATA` **independently** and + treat `CLAUDE_PLUGIN_ROOT` as ephemeral/throwaway (it rotates every plugin update); (2) any durable + bootstrap state goes to `$CLAUDE_PLUGIN_DATA` (survives updates) or `$XDG_STATE` — **never** + `CLAUDE_PLUGIN_ROOT`; (3) the plugin must be a genuinely self-contained **thin shim** (the hook plus a + small launcher only). Plugins deliberately skip outside-pointing symlinks for security, so the shim + **cannot symlink-bridge into `$XDG_DATA`**; and the binaries / model / corpus do not fit git-clone + plugin delivery — so **"ship the whole app as a plugin" is not viable** (wrong shape), full stop. + + - **When it earns its keep — and when it doesn't.** The real payoff is **discoverability**: + `claude plugin install agent-ways@aaronsb` via a self-published single-plugin marketplace, riding + Anthropic's managed install channel, with the exemption-maintenance shift as a *bonus*. For the + exemption *alone* it is a **bad trade** — standing up a whole marketplace plus lifecycle coupling to + replace one copied file. It implies a **two-channel update model**: the thin plugin updates rarely via + the marketplace, while the bulk app updates independently via the reconciler / `git` in `$XDG_DATA`. + This is adjacent to **ADR-133**, which already shells `claude plugin list --json` at `SessionStart` to + discover plugin-contributed ways — so the framework already touches the plugin surface, just for + discovery rather than bootstrap. + +## Open Questions + +- **Bootstrap-exemption mechanism: copied script vs. stable-path binary call.** Both satisfy "not a link + into the managed tree"; which is chosen affects how the entrypoint is kept current. Parked per ADR-142. +- **Windows materialization specifics** — copy fallback details, developer-mode requirement, path quoting + (the #182 author already started this); how the reconciler detects "can't symlink here" and falls back. +- **Rescue policy for hand-edited core files** found during migration (revert to git, or lift the diff + into user scope as a shadow — see ADR-143's user root). +- **Exact 1.0.x window length** before 1.1.0 (shared with ADR-142). +- **Whether this child stays separate or folds into ADR-142.** Recommendation: keep separate — the + bootstrap exemption and the crash-safe migrator are each substantial enough to warrant their own + recorded decision and their own Consequences. diff --git a/docs/architecture/platform/ADR-146-installer-binary-verification-and-guided-build-fallback.md b/docs/architecture/platform/ADR-146-installer-binary-verification-and-guided-build-fallback.md new file mode 100644 index 00000000..84ada99f --- /dev/null +++ b/docs/architecture/platform/ADR-146-installer-binary-verification-and-guided-build-fallback.md @@ -0,0 +1,208 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: install +basis: + - evidence: download success is not launch success, and a way-embed source build without a C++ toolchain aborts mid-install so 100+ ways silently degrade to keyword-only matching + - precedent: ADR-144 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-30 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-144]]' + - '[[ADR-125]]' +imported: + from: docs/architecture/system/ADR-146-installer-binary-verification-and-guided-build-fallback.md + format: v0 + status: Accepted +--- + +# ADR-146: installer binary verification and guided build fallback + +## Context + +The install (ADR-142/144) acquires platform binaries — `ways`, `attend`, and the +`way-embed` semantic matcher — by downloading a per-platform prebuilt release and, +failing that, building from source. Two gaps make a fresh install silently +under-deliver: + +1. **Download success is not launch success.** The installer treats "the file + arrived" as done. A binary can download and still fail to execute — wrong arch + slipped through detection, a glibc/libstdc++ too old for the release, a + truncated or corrupt asset. The user gets a broken engine with no signal. + +2. **Source build is an unguided dead end.** When no prebuilt binary is available, + the fallback compiles locally. `way-embed` needs a C++ toolchain (cmake + + compiler, via llama.cpp). On a machine without it the build aborts with a raw + error mid-install, and the model download — sequenced after the binary — never + runs. Semantic matching (a hard dependency, ADR-125) is off, and 100+ ways + silently degrade to keyword-only matching. + +Local compilation is a **last resort** for regular users, not a first-class path. +When we can't hand them a working binary, the installer should set them up for +success — verify what it downloaded, and if it must fall back to building, do so +transparently and only with consent, never silently and never by installing system +packages behind their back. + +Two facts about the entry point shape the answer: + +- **The published installer is `curl -sL … | bash`, where bash's stdin *is* the + piped script.** A naive `read` consumes script bytes or hits EOF, not a keypress. + Interactive prompting through a pipe is fragile (`/dev/tty` may or may not exist, + behaves differently across SSH/containers), so an in-band "press a key to build" + is the wrong primitive for the common path. +- **The bootstrap already clones the full source** (with submodules) to + `$XDG_DATA_HOME/agent-ways` to enable in-place updates. So the machinery to build + from source — Makefile, `tools/way-embed`, llama.cpp submodule — is *already on + disk* after acquisition. Recovery needs no separate `git clone`; it is a `cd` into + the staged app dir plus `make`. + +Together these point away from prompting and toward **halting with precise, local +instructions**: when a downloaded binary won't run, stop and tell the user exactly +how to finish against the source we already staged. + +## Decision + +Add a **verify-then-guide** stage to the installer's acquisition path. Prebuilt +download stays primary; local compilation stays the last resort; between them sits +verification and a consent-gated, dependency-aware fallback. + +**1. Verify every acquired binary launches.** After download, run each binary's +cheap liveness probe (`--version`). A binary that does not exit success is treated +as *not acquired* — the same as a missing download — and routed to the fallback. +Report per-binary which verified and which did not, rather than a single opaque +"installed." + +**2. On a failed/absent binary, check — never install — build dependencies.** +Detect the toolchain the source build needs (cmake + a C++ compiler for +`way-embed`) with `command -v`. Checking is always safe and non-interactive; +installing system packages is neither and is never done implicitly. + +**3. On a verification failure, halt fast and hand off — do not prompt-and-build +in-band.** Stop at the *first* binary that won't launch and print a +dependency-aware **recovery card** rather than attempting an interactive build +through a pipe. Guiding a novice through a raw compile is not, by itself, good UX, +so the card leads with the higher-leverage route and keeps the manual one as a +backstop: + +- **Ask the agent (preferred).** Claude Code is, by definition, installed — this is + being set up *for* it. The card tells the user to open Claude Code in the staged + app dir (`$XDG_DATA_HOME/agent-ways`) and ask it to finish setup. A shipped + agent-context file (see 4) primes the agent with the exact, safe steps. This + turns "go compile a C++ project" into "ask the assistant that's already here." + For the user who doesn't even know *what* to ask, the card supplies a + **ready-to-paste prompt** — "if you're not sure what to do, start Claude Code as + usual and paste this:" — with the real values already substituted (the resolved + app-dir path, the platform triple, and which binary failed to launch). Recovery + drops to a single copy-paste; the agent, primed by the context doc, does the rest. +- **Do it yourself.** The precise commands against the *already-staged* source, + shaped by the dep check: deps present → `cd "$XDG_DATA_HOME/agent-ways" && make + setup`; deps missing but a recognized package manager (pacman/apt/dnf/brew) is + present → `… && make deps && make setup`; no recognized manager → install cmake + + a C++ compiler, then `make setup`. No `git clone` step — the source is already on + disk from the bootstrap. + +Then exit with a status reflecting reality: the install is functional for +keyword/pattern matching but degraded (semantic matching is a hard dependency, +ADR-125), so exit non-fatal-warning, never a hang. + +**4. Ship the agent build-context as a *discoverable* doc — never a way or skill.** +A small install-completion context file (e.g. an `AGENTS.md` / finish-setup doc at +`$XDG_DATA_HOME/agent-ways`) gives Claude Code what it needs to complete the build +*safely*: what `way-embed` is, how to check deps, that `make deps` may need sudo +(ask the human first), `make setup`, and how to verify success (`ways status` shows +the engine up). It is a **plain document, found on demand** — the recovery card +names its path, and the user opens Claude Code in the app dir and asks it to finish. + +It is explicitly **not** a way (`hooks/ways/…`) or a skill (`skills/…`): those load +into session context — a way via matching, a skill at startup — for *every* user in +*every* project, forever, to serve a recovery path almost no one hits. That is +exactly the always-on context tax the framework's progressive-disclosure model +exists to avoid. Install-completion is disclosed at the moment of need (the failed +install, the pointer in the card), not injected proactively. + +**5. `make deps` is the toolchain installer; the installer only checks.** A new +`make deps` target detects the OS/package manager and installs the build +prerequisites (cmake + C++ compiler), using sudo where the platform requires it. +It is the single command both the recovery card and the agent point at, and the +only place system packages are installed. Keeping "install deps" out of the +installer's own path preserves the property that piping a script to bash never +mutates system state without an explicit, separately-invoked, user-run command. + +**6. An inline interactive build is an optional convenience, never the contract.** +When `install.sh` is run *directly* (not piped) and a genuine interactive terminal +is present, it MAY additionally offer "build now? [y/N]" reading from the tty. This +is a nicety for the power user who cloned and ran the script by hand; it never +replaces the recovery card, which is what the piped `curl | bash` path relies on. + +The ordering also decouples the model download from the binary build: the model is +a hard dependency independent of *how* the binary was obtained, so its acquisition +must not be gated behind a binary build that may be skipped or deferred. + +## Consequences + +### Positive + +- A broken or absent binary is caught at install time with a precise, per-binary + signal instead of surfacing later as mysteriously bad matching. +- Regular users are never dropped into a raw compiler error: they get a working + prebuilt binary, or a recovery card that hands off to the Claude Code they're + already installing this for (primed by a shipped context doc) — with exact + one-line commands against the already-staged source as the backstop. +- The recovery leverages what's uniquely true here — an AI agent is guaranteed + present — without taxing every session: the build context is a discoverable doc, + not an always-loaded way or skill. +- The installer never installs system packages implicitly and never compiles + without consent — `curl | bash` stays as low-surprise as its reputation demands. +- Unattended contexts (CI, cron, containers) get a deterministic, non-blocking + outcome with actionable output. + +### Negative + +- More installer branching and platform-specific dependency detection to maintain + (package-manager matrix in `make deps`). +- The interactive build path depends on `/dev/tty` semantics, which vary across + terminals, SSH sessions, and container runtimes — a class of environment bugs to + test for. +- A two-step "run `make deps`, then re-run the installer" flow is more friction + than a one-shot install for the toolchain-less user — accepted deliberately as + the price of never installing system packages behind their back. + +### Neutral + +- Prebuilt releases must exist for every supported platform for the happy path to + stay toolchain-free (see the way-embed release job; this ADR assumes it). +- `make deps` becomes a documented, supported target users may run independently of + installing agent-ways. + +## Alternatives Considered + +- **Auto-install deps and auto-build silently.** Rejected: a piped-to-bash + installer that runs sudo package installs and compiles without consent violates + the least-surprise contract and is a security smell. +- **Never build locally; require a prebuilt binary or fail.** Rejected as the sole + policy: prebuilt coverage can lag (new platform, a release not yet cut), and a + capable machine with the toolchain present should be allowed to build with + consent rather than be told "unsupported." +- **Prompt-and-build in-band (via `/dev/tty`).** Considered as the primary fallback + and rejected: `/dev/tty` prompting through `curl | bash` is fragile across SSH, + containers, and CI, and building a C++ project mid-install is poor UX for a + novice. Since the full source is already staged, halting with a recovery card + (agent handoff + exact commands) is more robust; an inline tty prompt survives + only as an optional convenience for the directly-run script (Decision 6). +- **Carry the install-completion context as a way or skill.** Rejected: a way + matches into context and a skill loads at startup — both would inject this into + *every* session for *every* user to serve a rare recovery path, the exact always-on + context tax progressive disclosure exists to prevent. It must be a discoverable + doc, surfaced only at the point of need. +- **Verify by file existence / checksum only.** Rejected as insufficient: a + checksum proves the bytes match the release, not that the binary *runs* on this + host (glibc/arch mismatches pass a checksum and still fail to launch). A launch + probe is the actual property we need. diff --git a/docs/architecture/platform/ADR-147-composable-settings-json-config-fragments.md b/docs/architecture/platform/ADR-147-composable-settings-json-config-fragments.md new file mode 100644 index 00000000..b6f4e3fe --- /dev/null +++ b/docs/architecture/platform/ADR-147-composable-settings-json-config-fragments.md @@ -0,0 +1,299 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +superseded_by: ADR-169 +basis: + - standard: Claude Code's documented settings merge law and managed-settings.d drop-in directory + - standard: 'SchemaStore claude-code-settings.json schema (Anthropic publishes none: anthropics/claude-code#11795)' + - evidence: the managed settings console is a raw settings.json textarea whose only safety note is "Invalid settings may disable Claude Code for your organization" +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-07-01 +deciders: + - aaronsb + - claude +related: + - '[[ADR-169]]' + - '[[ADR-142]]' + - '[[ADR-143]]' + - '[[ADR-145]]' +imported: + from: docs/architecture/system/ADR-147-composable-settings-json-config-fragments.md + format: v0 + status: Superseded +--- + +# ADR-147: Composable settings.json — a store of YAML config fragments + +## Context + +Claude Code's configuration surface is large — ~90 `settings.json` keys (many +version-gated), plus file artifacts (skills, agents, commands, hooks, statusline) and +MCP servers. Today a user manages all of it as a single hand-edited +`~/.claude/settings.json` and a scattering of files: no composition, no lint, no +rationale, no lifecycle. The one *management* surface Anthropic ships — the enterprise +**managed settings** console — is a raw `settings.json` textarea whose entire safety +story is one sentence: *"Invalid settings may disable Claude Code for your +organization."* It is a deploy backstop, not an authoring workflow. + +Three observations shape the decision: + +1. **Claude Code already ships the composable pattern — but only for managed scope.** + File-based managed settings support a `managed-settings.d/*.json` drop-in directory: + numbered fragments, merged in filename order, with a *defined* merge law. That is + the conf.d pattern — and it exists for the *enterprise* layer and nowhere else. User + and project scope are each a single file. + +2. **The merge law is a spec, not something to invent.** Claude Code documents exactly + how settings merge: most keys override by precedence, `permissions.allow/deny/ask` + **concatenate + deduplicate**, `env`/objects deep-merge. A compiler mirrors this + table; it does not design new semantics. + +3. **This shape is proven prior art.** A tree of markdown files with YAML frontmatter — + structured fields up top, human-readable body below — is exactly agent-ways' own + ways format, the Open Knowledge Format (Google, 2026), and this author's `kg-fuse` + knowledge store. The format demonstrably scales to typed graphs and epistemic + metadata in those richer systems. We adopt only the **node shape**; the graph, + content-addressing, and query machinery those systems carry solve problems + `settings.json` does not have (see Non-goals). + +This refactors the earlier ADR-147 draft ("projectable user config layer"), which +framed the same idea too small — as a sync mechanism. The projection/capture it +described survives here as the final pipeline stage. + +## Decision + +Manage Claude Code configuration as a **store of composable YAML config fragments**: a +tree of markdown files with YAML frontmatter, compiled deterministically into +`settings.json` and projected by the existing reconciler. + +### The unit + +One file per concern. Frontmatter carries a `settings:` block — a `settings.json` +fragment expressed in YAML — plus minimal meta; the body carries the rationale. + +```markdown +--- +scope: user # user | project | managed +mandatory: false # org-lock (managed scope only) +settings: # a settings.json fragment, in YAML + permissions: + allow: ["Bash(git:*)", "Bash(gh:*)"] + deny: ["Bash(rm -rf *)"] +--- +# Git & GitHub permissions +Let Claude run git/gh unprompted — constant use, prompts are pure friction. +`rm -rf` stays denied: destructive, never worth auto-allowing. +``` + +- `settings:` is a literal `settings.json` fragment in YAML — the readable spelling of + the JSON, mapping 1:1. +- The body is the *why* — the thing the console's textarea can never hold. `git blame` + on `permissions/git.md` answers "who allowed this, and why." +- Ordering is filename prefix (`10-permissions.md`, `20-hooks.md`) — the same + convention as `managed-settings.d` and a modular shellrc. + +### The pipeline + +`author → lint → compile → project` + +- **lint** — three flat, deterministic checks: (a) *schema-valid* (key exists, right + type) against Claude Code's settings schema; (b) *scope-legal* (no managed-only key + authored at user/project scope); (c) *duplicate-scalar* (two fragments set `model` → + warn; last-wins resolves it). This catches the disable-Claude-Code footgun *locally, + at author time* — the thing the console cannot do. +- **compile** — deep-merge every fragment's `settings:` block in filename order, + applying Claude Code's documented merge law (a fixed per-key lookup, not a merge + engine). Emit a baked `settings.json` and a `key → source-fragment` provenance + manifest. +- **project** — the existing reconciler (ADR-144/145) materializes the compiled output + into `~/.claude` and three-way-merges `settings.json`, using the provenance manifest + as its base (the same base hardened in the ADR-145 settings-merge work). + +### Schema source + +The lint and compile stages need Claude Code's settings schema. Anthropic does +not publish a machine-readable one (anthropics/claude-code#11795), and the native +binary embeds keys only as minified string literals — not cleanly extractable. +The community **SchemaStore** schema +(`json.schemastore.org/claude-code-settings.json` — 84 keys, every one carrying a +`description`) is the de-facto definition editors already use for autocomplete, so +we adopt it as the shape source. It supplies the key set, types, and descriptions +(enough to generate fill-in-the-blank fragment templates); it deliberately does +*not* encode scope-class (managed-only / managed-overridable), which stays the +small hand-curated overlay. + +The schema is **externalized, not compiled in**: a data file shipped alongside the +binary (`share/claude-code-settings.schema.json`, riding the app's data-dir clone), +read lazily at runtime only when a `ways settings` command needs it — never on the +every-turn hook path. Claude Code's settings surface changes frequently, so this is +deliberate: the schema updates **without a rebuild** (`make update`'s `git pull` +refreshes the shipped copy; `refresh-settings-schema.sh` writes a durable user copy +that takes effect immediately), and the binary that runs on every turn does not +carry data it rarely reads. The trade — the file can be absent — is handled by +**graceful degradation**: the linter skips schema-valid (scope-legal and duplicate +still run off the overlay) and scaffolding errors with guidance, rather than the +tool breaking. An earlier draft bundled it via `include_str!` for offline +determinism; the update cadence and every-turn binary size won out. + +Two configuration surfaces follow, both env-overridable and both defaulting +sanely: the **source** (`settings_schema_url`) — where a refresh *fetches* from, so +an org can point at an internal mirror, a version-pinned URL, or an official +Anthropic schema if one is ever published; and the **file location** +(`$WAYS_SETTINGS_SCHEMA_FILE` › `$XDG_CONFIG` durable copy › shipped copy) — which +file is *read*. project-pulse tracks the drift. + +### Independence and the shape contract + +*Managing* the configuration is independent of the ways matching engine — someone can +use the fragment store with zero interest in ways. The **contract** between the store +and its consumers is the compiled output: a baked `settings.json` (or, for org scope, +`managed-settings.d/*.json` fragments Claude Code merges natively) plus the provenance +manifest. Consumers are swappable: the reconciler projects it into `~/.claude`; an +enterprise console receives it as a *deploy target*; a plain viewer browses the tree. + +## Managed-scope interop + +Managed settings are how an organization *enforces* configuration, and they behave +unlike any other scope. The reconciler must **coexist** with the managed layer, never +merge it — and the factory must know which shape to emit for it. + +**Delivery is IT-owned, through three channels; Claude Code only reads.** Server-managed +settings are fetched from the claude.ai admin console, cached to +`~/.claude/remote-settings.json`, and **re-polled hourly**. MDM/OS policy is deployed to +a plist (macOS) or the HKLM registry (Windows) and read once at startup. File-based +settings live at `/etc/claude-code/managed-settings.json` (plus a `managed-settings.d/` +drop-in dir) and are read once at startup. Within the managed tier these channels do +**not** merge — if server-managed delivers any keys, endpoint sources are ignored — so +an org commits to one channel, and the factory cannot split its output across two. + +**As a consumer, agent-ways coexists — it does not reconcile the managed layer.** Managed +is the highest precedence and lands in files the reconciler never touches +(`managed-settings.json`, `remote-settings.json`); our three-way merge is scoped to +`~/.claude/settings.json` alone. Three behaviors follow, and they are *not* uniform: + +- **List keys concatenate.** `permissions.allow/deny` and `deniedMcpServers` from our + user-scope fragments still take effect — a user may *broaden* a managed allowlist, only + not *narrow* it. Our permission fragments are not dead under management. +- **Override keys are dead on arrival.** `fallbackModel`, `availableModels`, and scalar + hard-overrides (e.g. `model`) set at managed scope replace ours entirely. A user-scope + fragment for such a key is silently ignored on a managed endpoint — the linter should + say so (a *managed-overridable* warning, alongside the scope-legal check). +- **Policy locks can suppress the whole projection.** An org that sets + `allowManagedHooksOnly`, `allowManagedPermissionRulesOnly`, or + `strictPluginOnlyCustomization` turns agent-ways into a no-op *by policy* — our hooks, + skills, and agents do not load. This is the sharpest managed fact for the project: the + honest behavior is to *detect* the lock and report "policy-suppressed," not to project + silently and imply it took effect. Runtime detection is a `ways`-side concern, tracked + separately from this factory. + +**As a producer, the compile target follows the deploy channel** (this settles the +managed-compile-target question): for file/MDM deployment, emit `managed-settings.d/*.json` +and let Claude Code merge them natively — its drop-in law (alphabetical, later-file-wins, +systemd-style) *is* this ADR's `NN-` filename-prefix convention, so the fragment tree maps +1:1 with no merge code of our own; for the console, pre-merge to a single `settings.json` +blob to paste in. + +**No readback.** There is no documented mechanism for the console to read effective config +back from an endpoint, or for Claude Code to report config state upstream — the console +only audit-logs changes made *in* it. The fragment tree is therefore the sole source of +truth and the console is strictly the last mile. This is why the earlier draft's two-way +sync is *dropped*, not deferred: there is no upstream to sync from. + +## Non-goals (the discipline) + +Complexity is bounded to what `settings.json` needs. Explicitly **not** built: + +- **No typed dependency graph / topological compile.** `settings.json` precedence is + *linear* (filename order / last-wins). "Conflict" is a duplicate-scalar warning, not + cycle detection. +- **No content-addressed store.** A plain `key → fragment` manifest is enough provenance + for capture and diffing. +- **No epistemic layer** (grounding, contradiction scoring). Config is declarative fact, + not uncertain knowledge. +- **No semantic query / FUSE projection.** Config is *composed*, not *queried*. + +The prior art (ways, kg-fuse, OKF) proves the format *scales* to these if a future need +appears; none is a need for `settings.json` today. + +Initially scoped to `settings.json`. File artifacts (skills/agents/commands/hooks) and +MCP (`~/.claude.json`) already project through the reconciler and stay there; the same +fragment pattern extends to them later if wanted. + +## Consequences + +### Positive + +- Config becomes **literate, greppable, git-native** — the structure of JSON, the + readability of YAML, with the rationale attached and `git blame` for provenance. +- Deterministic lint catches invalid/scope-illegal settings **before** they reach + `~/.claude` — strictly better than the console's "may disable Claude Code." +- A vetted, tested, compiled `settings.json` can be **copied into an enterprise + console** — the console as a deploy target, authoring done as a git-managed factory. +- Fixes the `statusLine`-class bug at the root: the framework stops force-claiming + user-scoped keys; a user's own config lives in fragments the reconciler projects, + rather than being orphaned or clobbered. +- Reuses the ways/OKF node shape the project already tools and understands. + +### Negative + +- A **compile step** sits between editing a fragment and it taking effect (edit → + compile → project), versus editing `settings.json` directly. +- A **second faithful implementation of Claude Code's merge law**, which drifts as the + settings schema evolves (version-gated keys). Mitigated by tracking the settings docs + via project-pulse (a schema-freshness feed) — but it is a maintenance surface. +- Authoring YAML fragments has a small learning curve over a single JSON file. + +### Neutral + +- Depends on the reconciler and its provenance base (ADR-144/145); extends them rather + than inventing machinery. +- The reverse direction (**capture** — classify a live `settings.json` back into + fragments) is enabled and worth pursuing, but deferred. It is seeded by Claude Code's + known in-situ writers (`advisorModel`, `effortLevel`) and the managed-only key list. + +## Alternatives Considered + +- **Leave user config untouched (status quo).** Rejected: drops the value (portable, + versioned, testable config) and leaves the `statusLine`-class bug. +- **Plain numbered JSON fragments (`managed-settings.d` style), no bodies.** The minimum + floor, zero new format. Rejected as the primary shape only because it cannot carry + rationale — but it remains the fallback if the markdown shape proves too heavy. +- **Full OKF/kg-fuse machinery** (typed DAG, content-addressing, epistemic layer, + semantic query). Rejected as over-complexity for a JSON file with one merge quirk; + deferred to if-ever (see Non-goals). +- **A standalone dotfiles tool.** Rejected: the reconciler already projects and merges + `settings.json`; the fragment store is the missing *authoring* layer, not a second + projector. +- **Rely on the enterprise managed-settings console.** Rejected as the authoring + surface: no lint, no lifecycle, no rationale, no composition. It is a deploy target + the factory *feeds*, not a place to author. + +## Resolved Questions + +Settled during implementation (slices 1–2): + +- **Store location** — *resolved.* The store path is a **parameter** with an XDG + default (`$XDG_CONFIG_HOME/agent-ways/settings/`). A personal user gets the + default; an org points the tool at a shared git repo (`ways settings lint + ./our-config/`). A shared repo is a *different argument*, not a separate mode, so + the original dichotomy dissolves. (Default is `settings/`, aligned with the + `ways settings` command, not the ADR's earlier `config/`.) +- **Test depth** — *decided: static only; decline the live-boot gate.* Claude Code + parses settings **tolerantly** (invalid entries are stripped with a warning, + remaining policy enforced), so booting a sandbox CC against compiled output would + not fail on bad config — a low-signal gate. The static checks flag exactly what CC + would silently drop, so they are the *stronger* signal, not a compromise. When the + compile stage lands, the high-value test is **merge-law conformance** (deterministic, + unit-testable), not a live boot. +- **Manager/ways boundary** — *decided: a `ways settings` subcommand; defer a + standalone binary.* Independence is real as a **code boundary** — `cmd/settings` + pulls in no matching-engine code (only `paths` + `serde`). Shipping inside the + `ways` binary is a distribution convenience, and extracting a standalone binary + later is cheap precisely because the module is decoupled (a `[[bin]]` target or + workspace member, no API churn). Extract only when a config-only consumer that + won't install `ways` actually appears. diff --git a/docs/architecture/platform/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md b/docs/architecture/platform/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md new file mode 100644 index 00000000..e98be04d --- /dev/null +++ b/docs/architecture/platform/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md @@ -0,0 +1,134 @@ +--- +contract: adr/v1 +kind: decision +verb: retire +capability: install +targets: + - skill:project-pulse + - way:meta/project-health +basis: + - evidence: 'the project-pulse skill and meta/project-health way are maintainer tools that sat in projected roots and symlinked into every install; project-health''s `project: ~/.claude` gate went dead under the 1.0 projection' + - precedent: ADR-142 + - precedent: ADR-143 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-01 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-143]]' + - '[[ADR-106]]' +imported: + from: docs/architecture/system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md + format: v0 + status: Accepted +--- + +# ADR-148: framework surface ships operator content; dev harness in project scope + +## Context + +The projected framework surface — the roots the reconciler symlinks into every +operator's `~/.claude` (`skills/`, `agents/`, `commands/`, `hooks/ways/`, ADR-142) — +is the product. Everything in it reaches every install. That surface must therefore +carry **operator** content only: guidance and tools useful to someone who adopted +agent-ways to steer Claude in *their* project. + +It had leaked. Two artifacts — the `project-pulse` skill and the `meta/project-health` +way — are **maintainer tools for developing agent-ways itself**: track Claude Code +upstream releases against agent-ways' commits, reconcile agent-ways' own ADRs against +its shipped code. An operator has no use for either, yet both sat in projected roots +(`skills/`, `hooks/ways/`), so they symlinked into every install. The backing tool, +`scripts/project-pulse`, correctly lived in non-projected `scripts/` — but the skill +and way that fronted it did not. + +The tell was in the old gate. `project-health` fired on `when: project: ~/.claude`. +Pre-1.0 that *meant* "only when developing agent-ways," because `~/.claude` **was** +the repo then. 1.0 made `~/.claude` a projection (ADR-142), so the gate pointed at +the wrong place and the way went dead — but its intent was always dev-only. It was +never operator content; it was dev harness with a dev gate the projection model broke. + +A second pre-1.0 fossil blocked the clean fix. The repo's `.gitignore` used a +deny-all-then-allowlist posture (`*`, then `!` each tracked path) — necessary when +the repo *was* an in-place overlay on a live Claude Code install and had to avoid +committing the install's runtime state. That posture made the repo's own project +scope (`<checkout>/.claude/`) untrackable, so dev harness had nowhere tracked-and- +non-projected to live. + +## Decision + +**The projected framework surface ships operator content only. agent-ways' own +dev-harness artifacts live in the repo's project scope** — `.claude/ways/` and +`.claude/skills/` — tracked and non-projected. + +This is the placement test for any new skill/way/command: *would an operator who +adopted agent-ways use this, or is it for developing agent-ways itself?* Operator +content goes in the projected roots; agent-ways' own tooling goes in the repo's +project scope, exactly where any *other* project's dev guidance would live (ADR-143's +project tier). agent-ways thereby dogfoods its own three-root design — these are its +first project-scoped artifacts. + +Concretely: +- `skills/project-pulse/` → `.claude/skills/project-pulse/` +- `hooks/ways/meta/project-health/` → `.claude/ways/meta/project-health/` +- `scripts/project-pulse` stays put (already non-projected; unchanged). +- The `project-health` way drops its `when: project:` gate entirely — a project- + scoped way fires only in its own project by construction, so no gate (and no new + matcher condition) is needed. + +Enabling change: the `.gitignore` converts from the pre-1.0 deny-all overlay to a +conventional "track by default, ignore build output and runtime junk" list, which is +correct now that the repo is a normal application (ADR-142) rather than an overlay on +a CC install. `.claude/` runtime state stays ignored; `.claude/ways/` and +`.claude/skills/` are un-ignored so the dev harness is tracked and shared. + +## Consequences + +### Positive + +- Operator installs stop carrying maintainer tooling they can't use — the product + surface is honestly operator-only. +- agent-ways dogfoods ADR-143's project tier, which is a working proof of the + three-root design (and directly relevant to the adopter story). +- A durable, one-question placement rule for contributors, so this class of leak + doesn't recur. +- The broken `project: ~/.claude` gate disappears rather than being re-plumbed — and + with it the need for a new `when: git_remote:` matcher condition that existed only + to re-gate this one way. + +### Negative / caveats + +- A project-scoped *semantic* way (like `project-health`) only fires once the project + corpus includes project-local ways; in the dev checkout that may need a `ways corpus` + regen. The *skill* is discovered natively regardless, so the capability the dev + actually reaches for (`/project-pulse`) works immediately. +- The conventional `.gitignore` is "risky by default" (junk creeps in unless ignored) + where the deny-all was "safe by default." Mitigated by an explicit ignore list for + every build-output and runtime-junk category, verified against the tracked baseline. + +### Neutral + +- `scripts/project-pulse` is unchanged — the capability was never the problem, only + the surface its skill/way sat on. + +## Alternatives Considered + +- **Delete the skill and way outright.** Rejected: the capability is genuinely useful + for developing agent-ways (tracking Claude Code changes, ADR reconciliation). The + problem was placement, not existence — relocate, don't destroy. +- **Generalize `project-pulse` into an adopter-facing tool** (parameterize the upstream + repo so any project can track its own upstream). Rejected as the framing here: it's a + dev harness, not a half-built product feature. Generalizing it would be dressing up a + maintenance tool as something operators want. If a genuine adopter-facing "project + pulse" is ever wanted, that's its own decision with its own design — not a reason to + keep this one in the shipped surface. +- **Re-gate the way with a new `when: git_remote:` matcher condition.** Rejected: + building a matcher feature to keep one dev way firing in the projected surface is + backwards. Project scope *is* the gate; the feature isn't needed. +- **Add targeted `.gitignore` negations for `.claude/ways` / `.claude/skills` only, + keeping the deny-all.** Rejected: it preserves the pre-1.0 fossil and adds more + allowlist cruft. Converting to conventional fixes the root cause once. diff --git a/docs/architecture/platform/ADR-149-operator-config-interview-skill.md b/docs/architecture/platform/ADR-149-operator-config-interview-skill.md new file mode 100644 index 00000000..ba3c87fa --- /dev/null +++ b/docs/architecture/platform/ADR-149-operator-config-interview-skill.md @@ -0,0 +1,185 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +superseded_by: ADR-169 +basis: + - evidence: Claude Code's /insights report already carries configuration recommendations ("Where Things Go Wrong", "Suggested CLAUDE.md Additions", "Existing CC Features to Try") + - precedent: ADR-147 +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-07-01 +deciders: + - aaronsb + - claude +related: + - '[[ADR-169]]' + - '[[ADR-147]]' + - '[[ADR-134]]' +imported: + from: docs/architecture/system/ADR-149-operator-config-interview-skill.md + format: v0 + status: Superseded +--- + +# ADR-149: operator config interview skill + +## Context + +ADR-147 built the `ways settings` CLI: deterministic primitives for managing +Claude Code's `settings.json` as composable fragments — `lint`, `new` (scaffold +from schema), `schema` (show/refresh), `compile` (merge → baked settings.json + +provenance), `project` (install into the live config). These are **mechanism**: +dumb, testable, no cleverness. They are also, per ADR-147, a *shape contract* — a +surface a separate consumer can drive. + +Four assets now sit unused by any conversational layer: + +1. **Claude Code already understands its own configuration.** It ships with + context and skills that know what `statusLine`, `permissions`, `hooks`, and the + ~90 settings keys *mean* and how to configure itself. Re-teaching that would be + waste and drift. +2. **We have a structured schema** — the vendored 84-key schema with types and + descriptions (ADR-147) — the *spine* for a guided authoring flow. +3. **Ways carries operator telemetry** — firing stats, near-miss logging + (ADR-134), a `permissions audit`, governance/provenance. This is a record of + *learned behavior*: what the operator actually does, repeatedly, by hand. +4. **Claude Code already analyzes the operator's usage.** The `/insights` command + writes a report (`~/.claude/usage-data/report-<timestamp>.html`) from the last + 30 days of local sessions, and its sections are *already configuration + recommendations*: **"Where Things Go Wrong"** (friction), **"Suggested CLAUDE.md + Additions"**, **"Existing CC Features to Try"**, **"How You Use Claude Code"**. + It is Claude Code pre-computing the very suggestions this skill wants to make. + +Nothing composes these into an authoring experience. A user still hand-writes +fragments. The mechanism exists; the **conductor** does not. + +## Decision + +Build a **skill** that interviews the operator to author configuration, then +drives the `ways settings` primitives to lint, compile, and project it. The skill +is the intelligence; the CLI keeps the guarantees. + +**Core thesis — synthesize, don't rebuild.** The skill does *not* re-implement +Claude Code's knowledge of its own settings or its usage analysis. It **composes +four sources**: + +- Claude Code's own config self-knowledge (what a key is *for*, sensible values); +- our schema (the authoritative key set, types, and descriptions — the interview's + spine, and what keeps suggestions valid by construction); +- ways telemetry (what the operator repeatedly grants/does — the raw material for + *suggestions*); +- the **`/insights` report** — Claude Code's own analysis of the operator's last 30 + days (friction, pre-suggested CLAUDE.md rules, features to try). + +The result is a management system more capable than any one alone: CC's +understanding *and its usage analysis*, made **composable, inspectable, lintable, +and projectable** by the ADR-147 substrate, and **informed by the operator's own +history**. + +**On `/insights`: read it, don't parse it.** The report is HTML, not JSON — and a +brittle HTML *parser* would be the wrong dependency. Our consumer is a language +model: the skill **reads the latest report in-context** and lets Claude interpret +it the way a human would, focusing on the config-bearing sections. This sidesteps +the fragility of screen-scraping a format that may change, and it *is* the +synthesize-with-CC thesis at its purest — Claude Code generates the analysis, Claude +reads it, our schema turns what matters into valid fragments. Constraint: a skill +cannot invoke a slash command, so it consumes the existing report (noting its age) +and asks the operator to run `/insights` when the report is stale or absent. + +**Shape.** A skill (not a way, not a slash command — per the Skills Way, it *runs +a procedure*) whose sub-functions map onto the primitives: + +| Sub-function | Drives | +|---|---| +| interview / author | `ways settings new` + fills the value **and the body rationale** from the conversation | +| check | `ways settings lint` | +| rebuild | `ways settings compile` | +| project | `ways settings project` | +| pull-schema | `ways settings schema --refresh` | +| suggest *(v2)* | read the latest `/insights` report + mine ways telemetry (`permissions audit`, firing stats) → propose fragments | + +**The interview's byproduct is documentation.** The operator's answer to "why do +you want this?" becomes the fragment's markdown body — so `git blame` on the config +answers *who* and *why*, captured at authoring time for free. This is the payoff of +ADR-147's markdown-with-rationale format. + +**MVP vs. v2.** MVP is the interview→author→lint→compile→project loop with the +schema + CC self-knowledge. Telemetry-driven `suggest` is v2 (it needs a read +contract against the audit/stats sources) — deferred so the interview lands first. + +**Boundaries (Skills Way).** The skill projects into personal scope +(`~/.claude/skills/`), so: a **tight trigger** naming the specific task and the +words an operator says (with an explicit "not for" clause), and +**location-independence** (resolve the target up front, assume no cwd). It leans on +the `ways settings` CLI and `ways` telemetry — it does not reimplement them. + +## Consequences + +### Positive + +- The full author → lint → compile → project loop becomes **conversational**, with + Claude Code's own config understanding driving it — a materially more capable + manager than hand-editing JSON or the enterprise console's textarea. +- Rationale is captured at authoring time (the interview *is* the documentation). +- Suggestions (v2) turn passive telemetry into proactive config hygiene ("you keep + approving this by hand — want a fragment?"). +- Realizes ADR-147's independence promise: the skill is a *consumer* of the CLI + shape contract, swappable and separately versioned. + +### Negative + +- A conversational surface over live config must be careful: it drives `project`, + which writes `~/.claude/settings.json`. It inherits `project`'s safety + (dry-run/backup) but adds a trust surface (the skill proposing changes). +- Telemetry `suggest` (v2) couples the skill to internal ways data shapes — a + maintenance surface, deferred deliberately. +- A global skill with a loose trigger would hijack unrelated requests; the trigger + discipline is load-bearing, not optional. + +### Neutral + +- Depends on the ADR-147 primitives (all now built) — extends them, invents no new + mechanism. +- Composes CC's self-knowledge rather than encoding it, so it tracks CC's evolution + for free where the schema lags. + +## Alternatives Considered + +- **A settings GUI / TUI.** Rejected: rebuilds interaction Claude Code already does + conversationally, and can't leverage CC's self-knowledge or the operator's + history the way an in-session skill can. +- **Re-encode CC's settings knowledge in ways.** Rejected as the core mistake this + ADR avoids — waste, and guaranteed drift against a surface that changes often. +- **A way (hook-injected guidance) instead of a skill.** Rejected per the Skills + Way: this *runs a multi-step procedure* (interview → CLI calls), which is a + skill; a way shapes behavior, it doesn't execute a workflow. +- **Fold the orchestration into the `ways` binary** (a `ways settings interview` + subcommand). Rejected: the intelligence is conversational and model-driven, which + is exactly what a skill is for; the binary stays the deterministic mechanism. +- **Ship an HTML parser for `/insights`** (cheerio-style, as community tools do). + Rejected: brittle screen-scraping of an unsupported, changeable format. Our + consumer is a language model that reads HTML natively, so the skill reads the + report in-context instead — more robust *and* less code. +- **Use `/usage` or an OpenTelemetry exporter** for usage data instead of + `/insights`. Noted as complements, not replacements: `/usage` is tokens/cost (not + config-shaped), and OTel is structured but high-friction to stand up. `/insights` + already emits *config-shaped* recommendations, which is why it's the primary v2 + usage source. + +## Open Questions + +- **Trigger surface** — the exact `description` phrasing and "not for" clause that + fires on "help me configure Claude Code" without hijacking adjacent requests. +- **Telemetry read contract (v2)** — which sources (`permissions audit`, firing + stats, governance) and in what shape the `suggest` function consumes them. +- **`/insights` freshness (v2)** — the skill reads the newest + `~/.claude/usage-data/report-*.html`; how stale is too stale before it should ask + the operator to re-run `/insights`, and how it maps the report's sections + ("Suggested CLAUDE.md Additions", "Where Things Go Wrong") onto *settings* + fragments vs. CLAUDE.md memory (which is a different surface). +- **Store bootstrapping** — does the skill scaffold an empty store on first run, + and where (the ADR-147 default `$XDG_CONFIG/agent-ways/settings/`)? diff --git a/docs/architecture/platform/ADR-150-version-truth-and-downgrade-safe-self-update.md b/docs/architecture/platform/ADR-150-version-truth-and-downgrade-safe-self-update.md new file mode 100644 index 00000000..a1253f5a --- /dev/null +++ b/docs/architecture/platform/ADR-150-version-truth-and-downgrade-safe-self-update.md @@ -0,0 +1,196 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: install +basis: + - evidence: 'ways update silently downgraded a live install: main was 78 commits ahead of ways-v1.0.0 with the same version string, the stale release binary replaced it, and ways update and ways settings vanished' + - precedent: ADR-142 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-01 +deciders: + - aaronsb + - claude +related: + - ADR-142 +imported: + from: docs/architecture/system/ADR-150-version-truth-and-downgrade-safe-self-update.md + format: v0 + status: Accepted +--- + +# ADR-150: Version-truth and downgrade-safe self-update + +## Context + +`ways update` (the self-update command, ADR-142's projection lifecycle) refreshes +the app source and its binaries "pre-built first": it downloads the latest +GitHub Release binary and only falls back to building from source when no +toolchain is present. This is correct for an end user pinned to a release — but +it **silently downgraded a live install** in practice: + +- The app checkout tracks `main`. At the time of the failure `main` was **78 + commits ahead of the `ways-v1.0.0` tag** — those 78 commits included the + entire `settings` command and `ways update` itself. +- `tools/ways-cli/Cargo.toml` had stayed at `version = "1.0.0"` across all 78 + commits, and the newest published release was also `ways-v1.0.0`. So the + pre-built binary and the source **reported the same version string while being + 78 commits apart**. +- `ways update` downloaded the stale `ways-v1.0.0` binary over the current one. + `ways update` and `ways settings` vanished (`unrecognized subcommand`), and a + second `ways update` was impossible — the command that would fix it had been + removed by the update. + +The machinery to *avoid* this already exists; the gap is **version-truth**, not +tooling: + +- **The release pipeline is complete.** `.github/workflows/build-ways.yml` (and + the sibling `build-attend`/`build-way-embed` workflows) trigger on `ways-v*` + tags, build all four platforms, and the tag-gated `release` job runs + `gh release create`. Cutting a release is `bump + tag + push`; CI does the + rest. +- **The binary already bakes provenance.** `build.rs` sets `WAYS_COMMIT` from + `git rev-parse --short HEAD`, and `banner.rs` shows `v1.0.0 (c595437)`. + +But that provenance never reaches the surfaces that decide anything: + +- `ways --version` is wired to clap's `version` = `CARGO_PKG_VERSION` alone + (`1.0.0`), with **no commit**. `download-ways.sh`'s "already installed" check + and every human sanity-check see only the frozen string. The one place the + commit appears — the bare-invocation banner — is human-only and unread by + tooling. +- **Nothing compares versions before replacing a binary.** `ways update`'s + `refresh_component` renames-then-reverts on *build failure*, but a *successful* + download of an older binary is treated as success. There is no "is this + candidate actually newer than what I'm replacing?" gate. + +The result: staleness is invisible where it counts, and the self-updater will +downgrade whenever the tracked source is ahead of the latest cut release — which, +for a checkout that follows `main`, is the *normal* state between releases. + +## Decision + +Make the version *truthful and staleness-detectable*, and make self-update +*refuse to move backward* — leveraging the existing tag→CI→release pipeline +rather than adding a new one. Four changes: + +1. **Bake `git describe`, not just the short hash.** Extend `build.rs` to emit + `WAYS_BUILD = git describe --tags --always --dirty` (e.g. + `ways-v1.0.0-78-gc595437`) alongside `WAYS_COMMIT`. This single string + encodes the nearest release tag, the commits-ahead count, the commit, and a + dirty flag — everything needed to order two builds. + +2. **Surface it in `ways --version`.** Set clap's `long_version` so + `ways --version` prints the full provenance + (`ways 1.0.0 (ways-v1.0.0-78-gc595437)`), making a dev build visibly distinct + from a release build to both humans and scripts. The banner is unaffected. + +3. **A downgrade guard in `ways update`.** Before installing a candidate + pre-built binary, compare its embedded `WAYS_BUILD` against the pulled + source's `git describe`. Install the pre-built **only if it is at least as new + as the source** (same commit, or the source is an ancestor of the release). + Otherwise: + - **toolchain present →** build from source (the checkout is the authority + and can't be behind itself); + - **no toolchain →** keep the current binary and warn loudly ("a pre-built + matching this source hasn't been published yet; install a toolchain or wait + for the release"). **Never replace a binary with an older one.** This is the + belt-and-suspenders that keeps `ways update` safe *between* releases, + independent of whether anyone remembered to cut one. + +4. **A release-cut helper so the version bump can't be forgotten.** The root + cause was human: 78 commits merged with no bump and no tag. Add a single + entry point — `scripts/release.sh <component> <patch|minor|major>` (surfaced + as `make cut-release COMPONENT=ways LEVEL=patch`) — that bumps the component + `Cargo.toml` version, commits it, creates the `<component>-vX.Y.Z` tag, and + pushes. CI takes it from there. The policy: **a release is cut from `main` + whenever a user-affecting change to a shipped binary lands** (not every PR, + but no long silent runs). Between releases, `git describe` — now surfaced and + guarded — carries the truth, so a missed release degrades to "build from + source / keep current," never to a downgrade. + +Semver stays per-component (`ways-vX.Y.Z`, `attend-vX.Y.Z`, …), matching the +existing tag scheme and independent CI workflows. + +## Consequences + +### Positive + +- The self-updater can never silently downgrade: the worst case between releases + is "built from source" or "kept current with a warning," both of which leave a + working, current-or-newer binary. +- Staleness is legible everywhere — `ways --version`, `download-ways.sh`, CI logs, + bug reports — because the build string names its exact provenance. +- The release pipeline that already exists gets *used* on a discipline, closing + the gap that let 78 commits sit unpublished. +- `git describe` ordering is free and reliable — no version-string parsing games, + no dependence on anyone bumping Cargo.toml for *correctness* (the bump is for + human-facing semver; the guard keys on commit ancestry). + +### Negative + +- `build.rs` now depends on `git describe` succeeding in the build environment; + a tarball build with no `.git` yields `WAYS_BUILD = unknown` (handled: the + guard treats `unknown` as "cannot prove newer" → build/keep-current, never + downgrade). +- The downgrade guard needs the source checkout's `git describe` at update time — + fine for the app checkout (always a git clone per ADR-142), and the guard + degrades safely when it can't be computed. +- One more release step to remember — mitigated by the `make cut-release` helper, + which makes the correct path the easy path. + +### Neutral + +- Requires touching `build.rs`, `main.rs` (clap `long_version`), + `update.rs` (`refresh_component` gains the version comparison), and a new + `scripts/release.sh` + `make cut-release` target — the build slice tracked + separately from this ADR. +- The `attend`/`attend-chat` binaries still have no published pre-builts; the + guard's "no candidate newer than source → build/keep" path already covers them + (they are toolchain-gated today), so nothing regresses. + +## Alternatives Considered + +- **Always build from source when a toolchain is present (drop pre-built-first).** + Rejected: it discards the explicit "not everyone has build tools" requirement + that motivated pre-built-first, and it's slower for the common case where the + source *is* a released tag. The downgrade guard achieves the same safety while + keeping download-first for users on releases. +- **Compare `CARGO_PKG_VERSION` strings only (bump-per-PR discipline).** Rejected + as the *primary* mechanism: it's exactly what failed — a human forgot to bump, + and string equality can't see commit ancestry. Version bumps remain for + human-facing semver, but *correctness* rides on `git describe`, which can't be + forgotten. +- **Auto-cut a release on every merge to `main` (CI bumps + tags).** Rejected as + over-complex for now ("only complexify to the amount necessary"): it turns + every merge into a published release with version churn and 4-platform builds. + The `make cut-release` helper keeps cutting cheap and deliberate; auto-cut can + be revisited if the cadence proves too manual. +- **Embed a monotonic build number instead of `git describe`.** Rejected: it + needs external state (a counter), where `git describe` derives ordering from + the repo itself for free and is human-legible. + +## Amendment (2026-07-05): the downgrade guard does not apply to `--ref` + +The guard above governs the **release channel**: `ways update` pulls the tracked +branch and must never replace a binary with an older published one. `ways update +--ref <branch|tag|sha>` (ADR-142's amendment) is a different lifecycle — an +explicit pin to a chosen ref, built from source — and the guard is +**intentionally bypassed** there, for two reasons: + +- **Download-first cannot apply.** An unpublished ref has no GitHub Release + binary, so `--ref` always builds from source. There is no downloaded candidate + to compare — which is the only thing the guard gates. +- **"Never move backward" is the wrong invariant for an explicit pin.** The guard + exists because a *channel* update should be monotonic. Choosing to deploy a + specific commit — including one behind the current build, to reproduce or + bisect — is a deliberate act, not an accidental downgrade. The guard's job is to + stop *silent* regressions; a named `--ref` is neither silent nor accidental. + +The safety that remains is structural, not guard-based: `--ref` still touches only +app-scope (`$XDG_DATA` plus the regenerable projection) per ADR-142 §5, and +returning to the release channel is one command (`ways update --ref main`), after +which the guard governs again as normal. diff --git a/docs/architecture/platform/ADR-152-framework-default-secret-path-deny-baseline.md b/docs/architecture/platform/ADR-152-framework-default-secret-path-deny-baseline.md new file mode 100644 index 00000000..b018e274 --- /dev/null +++ b/docs/architecture/platform/ADR-152-framework-default-secret-path-deny-baseline.md @@ -0,0 +1,149 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +supersedes: [] +basis: + - standard: 'Claude Code permissions: permissions.deny rule syntax and deny-over-allow precedence (code.claude.com/docs/en/permissions)' + - evidence: permissions.deny was empty, so nothing stopped Read/Edit/Write from reaching ~/.ssh keys, ~/.aws/credentials or a project .env + - precedent: ADR-142 + - precedent: ADR-147 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-142 + - ADR-147 +imported: + from: docs/architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md + format: v0 + status: Accepted +--- + +# ADR-152: Framework-default secret-path deny baseline + +## Context + +agent-ways already owns a slice of Claude Code's `settings.json`. During +`ways reconcile` the settings three-way merge (ADR-142, ADR-147) set-unions a +fixed list of permission strings — `WAYS_PERMS` — into `permissions.allow` +(`Bash(ways:*)`, `Edit(~/.claude/**)`, …) and tracks them in a base so entries it +stops owning are cleaned up. The user's own `permissions` entries are held +invariant beside ours. + +`permissions.deny` is currently empty. Nothing stops the agent's **file tools** +(`Read`, `Edit`, `Write`) from reaching credential material — `~/.ssh` private +keys, `~/.aws/credentials`, GPG keyrings, a project `.env`. A single mistaken +tool call can read a private key into the transcript, where it is then persisted +and, if the transcript is ever shared, leaked. The framework projects itself into +every session; it should also project a floor of protection for the paths that +are almost never a legitimate read target for an agent. + +Claude Code's permission model makes this expressible: `permissions.deny` is a +list of `Tool(pattern)` rules (`Read(./.env)`, `Read(~/.ssh/**)`, +`Read(//abs/path/**)`), and **deny takes precedence over allow** — a deny is a +hard block, not a prompt. That precedence is the whole point (a broad +`Read(~/**)` allow cannot re-open a denied secret path) and also the constraint +this ADR must respect: a forced deny the user cannot override via `allow` needs a +deliberate, documented escape hatch, or it becomes a cage. + +## Decision + +### 1. Ship a curated secret-path deny baseline, reconciler-owned + +Add a `WAYS_DENY` list, the deny-side sibling of `WAYS_PERMS`, set-unioned into +`permissions.deny` by the same merge and tracked in the same base. The baseline +is intentionally **narrow and high-confidence** — paths an agent has essentially +no legitimate reason to read or write: + +``` +Read(~/.ssh/**) Edit(~/.ssh/**) Write(~/.ssh/**) +Read(~/.aws/**) Read(~/.gnupg/**) Read(~/.config/gcloud/**) +Read(~/.kube/config) Read(~/.netrc) +Read(./.env) Read(./.env.local) +``` + +`~/.ssh` denies read **and** write (the agent should neither exfiltrate a key nor +tamper with `authorized_keys`/`config`). The credential stores deny read. +Project env files deny the secret ones (`.env`, `.env.local`) but **not** +`.env.example` / `.env.sample`, which are meant to be read. The list is a +starting floor, expected to grow through the same review as any owned-permission +change — not a claim of completeness. + +Two deliberate narrownesses in that floor: the `.env` rules are **root-anchored** +(`Read(./.env)` matches the project-root file, not a nested `subdir/.env`) and +deny **read only** — the home-directory credential stores, not project env files, +are the higher-severity targets, and over-broad project globs are where workflow +breakage lives. Both can widen later if warranted. + +### 2. Secure by default, with an explicit opt-out + +The baseline applies by default. Because a Claude Code deny cannot be overridden +by a user `allow`, autonomy is preserved through a config flag, not settings +editing: `secret_path_deny: false` in `$XDG_CONFIG_HOME/agent-ways/config.yaml` +suppresses the baseline entirely (default `true`). A user who genuinely needs the +agent to read a protected path opts the whole baseline out deliberately, rather +than having it silently re-asserted on the next reconcile. This mirrors the +`disabled_domains` pattern: framework behavior on by default, off by one explicit +line. + +### 3. Honest scope — this gates tools, not the shell + +The deny binds the `Read`/`Edit`/`Write` **tools**. It does **not** stop a shell +command — `Bash(cat ~/.ssh/id_ed25519)` reads the same file through a different +door, and command-string denial is unreliable (countless spellings). Bash-level +exfiltration is a distinct, harder problem addressed elsewhere (the adversarial +`contributions` way, the hardened code-reviewer). This baseline is a meaningful +floor against the most common accident — the agent's own file tools wandering +into a secret — not a claim of exfiltration-proofing. Overselling it would be the +exact overclaim the compliance work (ADR-200) exists to avoid. + +## Consequences + +### Positive + +- Secret material is protected from the agent's file tools by default, in every + reconciled session, at zero user effort. +- Reuses the proven owned-permission merge machinery — no new projection seam. +- The opt-out keeps the default honest: secure, but not a cage. + +### Negative + +- A forced deny can surprise a user whose task legitimately needs a protected + path; they must know about the config opt-out. Mitigated by documenting it + where the deny is described. +- The curated list is a maintenance surface and a judgment call — too broad + breaks workflows, too narrow misses secrets. It ships deliberately conservative. + +### Neutral + +- Bash-path exfiltration remains out of scope, by design and stated plainly. +- The list can grow; each addition is an owned-permission change reviewed like any + other, and the three-way base cleans up entries later removed. + +## Alternatives Considered + +- **Seed once, never re-assert.** Add the deny on first reconcile but let a user + removal stick. Rejected: it makes the security floor silently erodible and + diverges the deny logic from the forced-allow logic for no clear gain; the + config opt-out gives the same autonomy explicitly. +- **`ask` instead of `deny`.** Use `permissions.ask` so secret paths prompt. + Rejected for the baseline: a prompt is the right tool for *ambiguous* paths, but + for private keys and credential stores a hard default block is the safer floor. + `ask` remains available to users for their own softer cases. +- **Do nothing / leave to the user.** Rejected: the framework already projects + permissions; declining to project the one that protects secrets, when the + syntax makes it a few lines, is a missed default. + +## References + +- **ADR-142 / ADR-147** — the projection and the settings fragment/merge machinery + this extends. +- **Claude Code permissions** — `permissions.deny`, rule syntax, and deny-over-allow + precedence. https://code.claude.com/docs/en/permissions diff --git a/docs/architecture/platform/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md b/docs/architecture/platform/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md new file mode 100644 index 00000000..5102f767 --- /dev/null +++ b/docs/architecture/platform/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md @@ -0,0 +1,208 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +superseded_by: ADR-167 +basis: + - evidence: a PR in an unrelated repository with no local .claude/settings.json still carried the session link in its body; commit 3b7f04c's attribution.commit/pr settings did not govern the link +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-07-06 +deciders: + - aaronsb + - claude +related: + - 152 + - 163 + - 167 +imported: + from: docs/architecture/system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md + format: v0 + status: Superseded + unmapped: + revised: 2026-07-16 +--- + +# ADR-162: Mechanical session-link suppression as defense against transcript disclosure + +## Status: Superseded by ADR-167 + +**Superseded 2026-07-16.** The threat model below is correct and load-bearing: the +session link resolves to the full transcript, and publishing it on a public repo is an +accidental-disclosure surface. The **hook** this ADR designed also survives unchanged — +deny rather than rewrite, at user scope, trailer/footer position only. Its *rationale* +does not. + +This ADR's central claim — that no setting can govern the link, so suppression must be +mechanical — was **false when written**. It rests on a mis-citation: `#18253` is the +Co-Authored-By/footer bug ([CLOSED/COMPLETED]), not the session link; the issue meant +was `#41873` ([CLOSED/NOT_PLANNED]). And `#41873` had already been superseded by +`attribution.sessionUrl`, shipped in **v2.1.183 (2026-06-19)** — seventeen days before +this ADR was dated. A controlled experiment on 2026-07-16 (same repo, same v2.1.212, +only the key differing) confirms the setting governs the trailer. + +[[ADR-167]] carries the corrected decision: `attribution.sessionUrl: false` is the +primary control, and the hook below is retained as a **backstop** rather than the sole +defense. Read this ADR for the threat model and the hook's design; read ADR-167 for +what actually defends against it. + +## Context + +Claude Code appends a session link to git commit messages (`Claude-Session: +https://claude.ai/code/session_…`) and to PR bodies. That link resolves to the +**full session transcript**. A transcript routinely contains material that +scrolled through the conversation — file contents, environment values, tokens, +internal paths. Publishing the link therefore places a single click between a +public commit and whatever secret happened to pass through the session. On a +public repository this is an accidental-disclosure surface, not a convenience. + +A prior attempt to close this (commit `3b7f04c`, "give the session-link trailer a +control surface") set `attribution.commit` and `attribution.pr` to `""` and added +a *report-only* block to the GitHub delivery macro. Subsequent investigation shows +that attempt did not govern the session link at all: + +- `attribution.commit` / `attribution.pr` only suppress the **Co-Authored-By / + "Generated with Claude Code"** footers. They never controlled the session link. +- The key that does govern it, `attribution.sessionUrl`, is **undocumented** + (added in v2.1.183) and is **not reliably injected into the model's context** + (upstream issue [#18253](https://github.com/anthropics/claude-code/issues/18253), + marked *not planned*). The system-prompt instruction to append the link survives + the setting, and the model follows it. +- Claude Code has **no mechanism that strips** such a link after it is written. The + entire design is preventive ("do not instruct the model"), and the injection bug + defeats prevention. + +The failure is observable: a PR was authored in an unrelated repository that had +**no local `.claude/settings.json`** — it inherited the global +`attribution: {commit:"", pr:""}` — and the session link still landed in the PR +body. The setting layer neither reached the model nor covered a repo without local +config. + +The conclusion is that any control depending on the harness honoring a setting, or +on the model choosing to obey it, is not load-bearing. Suppression must be +mechanical and must live in a layer we own. + +## Decision + +Add a **PreToolUse hook on `Bash`** that inspects command invocations which author +commits, PRs, or issues — `git … commit`, `gh … pr create|edit|comment`, +`gh … issue create|edit|comment`, matched on token presence (newline-collapsed) so +intervening flags such as `git -C <path> commit` or `gh <global-flags> pr create` +cannot evade the gate — and **denies** any that carry a `Claude-Session:` trailer or +a bare `claude.ai/code/session_…` URL **in trailer/footer position** (line-leading), +returning a remediation message that instructs the model to remove the link and +re-issue. A session URL mentioned *inline in prose* is not a leak and passes, so a +repo can still discuss `Claude-Session` as subject matter without a deny-loop. The +clean re-issue passes through the normal permission flow. The link is blocked on the +covered command forms (coverage gaps are listed in Consequences). + +The hook **denies rather than rewrites**. Rewriting a command in place (PreToolUse +`updatedInput`) is honored *only* when the hook also sets `permissionDecision: +allow` — `updatedInput` alone is ignored, and no pass-through value exists +(upstream FR #381). Forcing `allow` would bypass the operator's own +Deny/Allow/Ask rules for that call, including on outward-facing `gh pr create` / +`gh issue create` — silently defeating a review gate the operator may rely on. +Deny bypasses nothing: it costs one extra round-trip, after which the model +re-issues cleanly (and adapts within the session to stop appending the link). + +The hook must operate at **user scope** — covering **every** repository the +operator works in, independent of whether a repo carries local config — rather +than as a per-project hook. That is the property this decision requires; *how* +the hook is distributed to achieve it (dotfile source-of-truth, the settings +projection, cross-host travel, key-vs-file ownership) is the province of ADR-163, +which carries this hook as the primitive it distributes. What matters here is the +guarantee, not the delivery: suppression does not depend on the harness injecting +the `attribution` setting, on the model obeying it, or on a per-project note. + +The `attribution.sessionUrl: false` key (undocumented, v2.1.183+) is the harness's +own would-be control; **this change does not set it** — it is defeated by #18253 +today, and coupling to an undocumented key risks tripping settings validation for +no working benefit. The hook is the load-bearing control. The existing report-only +macro block is retained; it reports the effective `attribution` state alongside a +hook that actually enforces suppression. An operator may set `sessionUrl: false` +additionally once upstream honors it. + +Scope boundary: this ADR governs suppression of the **session link** specifically. +It does not remove the Co-Authored-By footer (an independent operator preference +already handled by `attribution.commit`/`.pr`). + +## Consequences + +### Positive + +- Every repository is protected regardless of local config, the harness injection + bug, or model obedience. The unrelated-repo leak that motivated this ADR — a + trailer/footer link on a commit or PR — is closed on the covered command forms at + the layer we control. +- Enforcement is auditable and testable — a hook script with explicit patterns, + not a soft instruction competing with the system prompt. +- Defense-in-depth composes with ADR-152's secret-path deny baseline: 152 keeps + secret *files* out of the model's reach; 162 keeps a *pointer to the whole + transcript* out of published history. + +### Negative + +- A PreToolUse hook adds a small latency to every `git`/`gh` Bash call. On its + no-op paths (no command, out of scope, no link) it fails **open**; once a link is + actually detected it fails **closed** (deny), so a hook error there blocks rather + than leaks. +- Deny costs one extra round-trip each time the model appends a link — until it + adapts within the session. Benign but visible friction on the first commit/PR. +- Detection is bounded by what the hook can see in the command string, so the + control is **best-effort, not a guarantee**. Inline `-m` and heredoc bodies are + covered (this is how the model formats commits/PRs by default). Uncovered vectors, + each of which would slip a link through: + - a body/message passed via `--body-file`/`-F <path>` — the text lives in a file + the hook can't read from the command string; + - a bare `git commit` that opens `$EDITOR` — the message isn't in the command at + all; + - message-carrying siblings not in scope — `git tag -m`, `gh release create`, + `gh gist create`. + These are documented residual gaps, not silent ones. +- If Claude Code changes the link format, the hook's patterns need updating — a + maintenance coupling to an undocumented upstream string. + +### Neutral + +- The change belongs in the GitHub delivery way alongside the existing attribution + report block. The hook entry lives in the **reconcile-owned `settings.json` hooks + block** (beside the existing `check-bash-pre.sh` gate), not the settings-fragment + store — `ways settings project` deliberately skips the hooks/permissions slices, + so a fragment-authored hook would silently never fire. Its cross-host distribution + is owned by ADR-163. +- The hook enforces by **deny** (`permissionDecision: deny` with a + `permissionDecisionReason`, falling back to exit 2 + stderr). It deliberately does + *not* rewrite the command in place: see the Decision and the rejected silent-strip + alternative for why forcing `permissionDecision: allow` — the only way to make + `updatedInput` take effect — is unacceptable here. +- The hook shares the `Bash` PreToolUse matcher with the disclosure gate + `check-bash-pre.sh`. Multiple matching hooks all run and the most-restrictive + result wins (`deny > defer > ask > allow`), independent of array order, so the + deny takes effect regardless of position — no ordering dependency. + +## Alternatives Considered + +- **Rely on `attribution.sessionUrl: false` alone.** Rejected: undocumented and + defeated by injection bug #18253, with no coverage for repos where the setting + never reaches the model. Kept only as a secondary layer. +- **Silently rewrite the command in place (PreToolUse `updatedInput`).** Rejected: + `updatedInput` is honored only alongside `permissionDecision: allow` — it is + ignored on its own, and no pass-through value exists (upstream FR #381). The + rewrite would therefore force-approve the command and bypass the operator's + Deny/Allow/Ask gates, including on outward-facing `gh pr create` / `gh issue + create`. Cleaner UX, but a real permission regression on exactly the commands + worth gating; deny preserves the gate for the price of one retry. (This was the + initially-ratified choice, reversed once the `updatedInput`→`allow` coupling was + confirmed.) +- **A way / memory instruction telling the model to omit the link.** Rejected: + soft, competes with the system-prompt instruction, and scoped per project. This + is precisely what already failed — the leaking repo had no such note. +- **Post-hoc history rewrite / secret rotation.** Rejected as the primary control: + reactive, triggered only after the link (and any secret) is already pushed. + Remains the incident response if a link ever slips past the hook. +- **Do nothing / accept the setting layer as sufficient.** Rejected: the motivating + leak proves the setting layer is not load-bearing, and the exposure is a security + posture, not a preference. diff --git a/docs/architecture/platform/ADR-163-config-separation-dotfiles-source-of-truth.md b/docs/architecture/platform/ADR-163-config-separation-dotfiles-source-of-truth.md new file mode 100644 index 00000000..3688e009 --- /dev/null +++ b/docs/architecture/platform/ADR-163-config-separation-dotfiles-source-of-truth.md @@ -0,0 +1,137 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +basis: + - evidence: 'an audit of two hosts: statusline.sh missing on slab while settings.json referenced it, a permissions.deny for gh/docker credentials absent on slab, and the session link leaking on slab' + - evidence: 'PR #347 removed the inert statusLine key from the repo-tracked settings.json' + - precedent: ADR-147 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-06 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-147]]' + - '[[ADR-162]]' +imported: + from: docs/architecture/system/ADR-163-config-separation-dotfiles-source-of-truth.md + format: v0 + status: Accepted +--- + +# ADR-163: Config separation — dotfiles as source-of-truth feeding the settings fragment store + +## Context + +ADR-147 built the user-scope settings fragment store (`$XDG_CONFIG_HOME/agent-ways/settings/`) +and its projector (`ways settings project`). It defined how fragments *compile and +merge* into `~/.claude/settings.json`, but left two things open: + +1. **Where fragments come from** — the store is a per-host directory. Nothing said how + an operator's config *travels* across their machines. +2. **File artifacts** — ADR-147's Context named `statusline` (and hook scripts) as file + artifacts, but the built machinery projects only settings.json *keys*, never files. + +The gap is not academic. Auditing one operator's two hosts (call them north and slab) +surfaced silent drift that no CI catches: + +- `statusline.sh` present on north, **missing on slab** — the settings *pointer* had + propagated (legacy residue) but the *script* never did, so slab's status line was + broken while its `settings.json` still referenced it. +- `model` differed (`opus[1m]` vs `opus`); a `permissions.deny` guarding gh/docker + credentials was present on north, **absent on slab** — a security drift. +- Session-link (`Claude-Session:`) suppression worked on north **only because a per-host + `~/.claude` memory told the agent to disobey the harness prompt** — slab lacked the + memory and leaked the URL into commits/PRs. (This ADR originally recorded ADR-162's + reason: that `attribution.sessionUrl` was *broken upstream*, leaving a mechanical deny + hook as the real fix. That premise was false — [[ADR-167]] proves the key governs the + link and supersedes ADR-162. The observation above is unaffected: the control, whatever + it is, only defends the hosts it reaches. *Distributing* it is this ADR's concern, and + the drift is the point.) + +Separately, the framework repo was **force-claiming a user-scoped key**: `statusLine` sat +inert in the repo-tracked `settings.json` (reconcile co-owns only `hooks` + +`permissions`, so it was never projected) — exactly the anti-pattern ADR-147 set out to +end. + +## Decision + +**dotfiles is the operator's cross-host source-of-truth; agent-ways compiles and projects +what dotfiles feeds it.** The layering: + +``` +dotfiles (VCS, per-operator) + ├─ settings fragments ──deploy──► $XDG_CONFIG_HOME/agent-ways/settings/ (ADR-147 store) + │ └─ ways settings project ──► ~/.claude/settings.json (KEYS) + └─ file artifacts (statusline.sh, …) ──deploy──► ~/.claude/ (FILES) +``` + +Four commitments: + +1. **Artifact-ownership split.** The fragment store owns settings.json **keys**; dotfiles + owns **file artifacts** and deploys them directly to `~/.claude`. One owner per + artifact — no file has two writers. (`statusLine` the *key* → fragment; + `statusline.sh` the *file* → dotfiles.) + +2. **The framework de-claims user-scoped keys.** The repo-tracked `settings.json` ships + only what `ways reconcile` co-owns (`hooks` + `permissions`). `statusLine` was removed + (PR #347, merged). + +3. **The fragment store is the projection boundary.** dotfiles never writes `~/.claude`'s + `settings.json` directly — it deploys the store, and `ways settings project` performs + the three-way merge. agent-ways stays the single `settings.json` writer (alongside + reconcile's two co-owned slices), so the operator's own keys survive. + +4. **Projector-base hygiene is part of the contract.** The projector's per-host + last-applied base (`$XDG_STATE/agent-ways/settings-fragments-<scope>.json`) is + **host-local ephemeral state**, not config. It can carry *ghosts* — keys a prior + projection managed but the store no longer declares — which cause surprising + cross-host retractions. (Observed: a stale base recording `model: opus` would, if + carried to another host, silently retract that host's model to the default.) The base + is reset when the store's ownership set changes; dotfiles never deploys it. + +## Consequences + +- Cross-host config becomes reproducible: clone dotfiles → deploy → `ways settings + project`, and any host converges to a coherent Claude Code config. +- Session-link suppression (ADR-162) stops being a per-host memory hack: the deny hook + becomes a primitive distributed through this same pipeline. +- The two disjoint settings writers (`ways reconcile` vs `ways settings project`) that + ADR-147 left unreconciled remain disjoint here; unifying them is out of scope + (follow-up). +- File-artifact projection stays *outside* agent-ways (owned by dotfiles). If the file + set grows, revisit building projection into the fragment store (see Alternatives). + +## Alternatives Considered + +- **Plain dotfiles writes `~/.claude/settings.json` directly.** Rejected: two owners + (dotfiles + `ways settings project`/reconcile) fighting one file — the exact drift that + produced the slab breakage. +- **Personal config lives in the framework repo.** Rejected: revives ADR-147's + force-claiming anti-pattern (the inert `statusLine` is the cautionary case). +- **Build file-artifact projection into agent-ways so it owns `statusline.sh` too.** + Deferred, not rejected: cleaner single-owner end-state, but net-new machinery. The + dotfiles-owns-files split ships today and proves the loop; revisit when the file set + justifies it. + +## Validating implementation (the statusline pilot) + +`statusLine` was the pilot that exercised every seam: + +- **De-claimed** from the repo (`statusLine` removed from tracked `settings.json`, PR #347, merged). +- **Authored** as a user-scope fragment (`10-statusLine.md`) in the store. +- **Projected** cleanly — the stale projector base was reset first to clear ghosts, after + which `ways settings project` reported `already up to date`. +- **Distributed** — both the fragment store and `statusline.sh` are now dotfiles-managed + symlink entries (`claude-settings`, `claude-statusline`), committed and pushed, so a + second host reproduces the config with `dotfiles pull && dotfiles deploy && ways + settings project`. + +Attribution/session-link suppression (ADR-162) is the second, higher-value pass over this +same pipeline. diff --git a/docs/architecture/platform/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md b/docs/architecture/platform/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md new file mode 100644 index 00000000..80bf15cc --- /dev/null +++ b/docs/architecture/platform/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md @@ -0,0 +1,101 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: config +basis: + - evidence: the statusline.sh artifact ADR-163 called distributed was a symlink to /home/<authoring-user>/.local/share/agent-ways/statusline.sh, which dangled on a host with a different $HOME + - precedent: ADR-163 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-07 +deciders: + - aaronsb + - claude +related: + - '[[ADR-163]]' + - '[[ADR-147]]' + - '[[ADR-142]]' +imported: + from: docs/architecture/system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md + format: v0 + status: Accepted +--- + +# ADR-164: File artifacts distributed across hosts must be carried by value not host-absolute reference + +## Context + +ADR-163 split config artifact ownership: the settings fragment store owns +settings.json **keys**, and dotfiles owns **file artifacts** and deploys them to +`~/.claude`. Its validating pilot claimed the `statusline.sh` file artifact was +"distributed" across hosts. + +It was not. The dotfiles-tracked artifact was itself a **symlink whose target was an +absolute path into a per-host directory** — `/home/<authoring-user>/.local/share/ +agent-ways/statusline.sh`. That link resolved only on the host that authored it. On +a second host with a different `$HOME` (a different username) the target did not +exist, so the deployed `~/.claude/statusline.sh` dangled and the status line was +broken — the very cross-host drift ADR-163 set out to end, reintroduced one layer +down. The keys half of the pipeline never had this problem: the fragment store holds +real YAML, so it already travelled by value. + +The failure mode generalizes beyond statusline: any artifact distributed **by +reference**, where the reference encodes a host-local absolute path (a `$HOME`, an +app dir, a cache dir, a username), silently fails to reproduce on a host whose layout +differs. + +## Decision + +**A file artifact carried through the config pipeline is distributed by value — its +content — never by a reference that encodes a host-local absolute path.** + +Concretely, the dotfiles-tracked artifact is the **real file**. `dotfiles deploy` +then symlinks `~/.claude/<artifact>` to that dotfiles copy — a link whose target is +host-relative under `$HOME`, so it resolves identically regardless of username or +where the application happens to be installed. No artifact anywhere in the pipeline +may be a symlink whose target is an absolute path into a per-host application, cache, +or home directory. + +Corollary: dotfiles is the source of truth for a managed artifact's content +(consistent with ADR-163) — which also lets the operator customize it. Managing an +app-shipped default in dotfiles means dotfiles wins; the artifact is no longer an +alias of the app's copy. + +## Consequences + +### Positive + +- Distributed artifacts reproduce on any host regardless of username or install + layout; ADR-163's cross-host distribution claim actually holds. +- Operators can customize a managed artifact, because dotfiles owns its content + rather than pointing at an app-shipped file. + +### Negative + +- App updates to a shipped default (e.g. a new `statusline.sh`) no longer flow + automatically to a host that manages it via dotfiles — the operator re-syncs the + content when they want the newer default. A by-value file is opaque to which app + version produced it. + +### Neutral + +- This brings file artifacts to parity with the keys half of the pipeline, which was + already by-value (the fragment store holds real YAML). +- A future "file-artifact projection built into agent-ways" (ADR-163 Alternatives, + deferred) would supersede the dotfiles-owns-files mechanism, but must honour this + same by-value principle — it too cannot distribute a host-absolute reference. + +## Alternatives Considered + +- **Relative symlink into the app dir** (e.g. `../../.local/share/agent-ways/ + statusline.sh`). Rejected: still by-reference. It aliases the app's copy rather + than owning content, so the application stays the true owner (contra ADR-163); it + breaks if the dotfiles store or the app dir moves relative to `$HOME`; and it + cannot carry operator customization. +- **A symlink target with `$HOME`/env expansion.** Rejected: a symlink stores literal + bytes; the OS does not expand environment variables in a link target. +- **Leave the artifact host-absolute and re-author it per host.** Rejected: that is + exactly the silent cross-host drift ADR-163 set out to end. diff --git a/docs/architecture/platform/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md b/docs/architecture/platform/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md new file mode 100644 index 00000000..c51a4683 --- /dev/null +++ b/docs/architecture/platform/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md @@ -0,0 +1,220 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +supersedes: ADR-162 +basis: + - evidence: 'a single-variable experiment on 2026-07-16 (v2.1.212): with attribution.sessionUrl true the system prompt carried the Claude-Session trailer instruction, with false it was absent' + - standard: 'Claude Code v2.1.183 changelog: added the attribution.sessionUrl setting; upstream issues #41873 and #18253' + - precedent: ADR-162 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-16 +deciders: + - aaronsb + - claude +related: + - '[[ADR-104]]' + - '[[ADR-147]]' + - '[[ADR-162]]' + - '[[ADR-163]]' +imported: + from: docs/architecture/system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md + format: v0 + status: Accepted +--- + +# ADR-167: Session-link suppression: attribution.sessionUrl as primary control, deny hook as backstop + +## Context + +The disclosure surface [[ADR-162]] identified is real and unchanged: Claude Code +appends a session link to commit messages and PR bodies, that link resolves to the +**full session transcript**, and a transcript routinely carries material that merely +scrolled through the conversation — file contents, environment values, tokens, +internal paths. On a public repository, publishing it puts one click between a commit +and whatever secret happened to pass through the session. + +What has changed is the claim that nothing but a hook can stop it. ADR-162 concluded: + +> The conclusion is that any control depending on the harness honoring a setting, or +> on the model choosing to obey it, is not load-bearing. Suppression must be +> mechanical and must live in a layer we own. + +That conclusion was false when written, and it rested on two errors. + +**A mis-citation.** ADR-162 attributed the failure to upstream `#18253`, described as +*marked not planned*. `#18253` is a different bug — "Attribution settings not honored +— Co-Authored-By and PR footer added despite empty attribution config" +([CLOSED/COMPLETED]) — and concerns the **footers**, never the session link. The issue +it meant is `#41873`, "attribution setting does not control session URL in commit +messages" (filed April 2026, [CLOSED/NOT_PLANNED] 2026-05-22). + +**A superseded fact.** `#41873` was closed not-planned, and then upstream shipped the +control anyway. Claude Code **v2.1.183** (2026-06-19) added `attribution.sessionUrl`: + +> Added `attribution.sessionUrl` setting to omit the claude.ai session link from +> commits and PRs in web and Remote Control sessions + +ADR-162 is dated 2026-07-06 — seventeen days after the key shipped. It named +`attribution.sessionUrl` as "the key that does govern it" and then dismissed it as +unreliable on the strength of an issue about a different bug. The key is absent from +the official settings documentation (upstream `#69614`, open), which is why it stayed +unset here long after it was available; `#66504` (open) separately argues the link +should be opt-in rather than default-on. + +### The evidence + +A single-variable experiment on 2026-07-16 settles it. Same repository, same binary +(**v2.1.212**), same account; only `attribution.sessionUrl` differs. The question put +to each session: *does your system prompt instruct you to end git commit messages with +a Claude-Session trailer?* + +| Run | `attribution.sessionUrl` | Trailer instruction in system prompt | +|-----|--------------------------|--------------------------------------| +| control | `true` | **present** — a `Claude-Session:` trailer naming a session URL | +| treatment | `false` | **absent** | + +The setting governs the trailer. Two earlier attempts did **not** establish this and +were discarded: a headless `claude -p` pair (void — `-p` never injects the trailer, so +the control could not reproduce the baseline) and a run in a non-git directory (void — +no repository, therefore no commit instructions). Only the controlled pair is evidence. + +### How the error persisted + +The mis-citation propagated into five artifacts, each of which made the next look +corroborated: + +1. **ADR-162** — the original wrong citation. +2. **`strip-session-link-pre.sh`** — its header comment repeats "the `attribution` + setting does not reliably reach the model (upstream #18253)". +3. **The operator's memory entry** — repeats "attribution setting doesn't govern it + (#18253)". +4. **[[ADR-163]]** — records the premise second-hand: "the `attribution.sessionUrl` + config key is broken upstream, so a mechanical hook that denies the publishing + command is the real fix". Its own decision does not rest on this, but an Accepted + ADR restating the claim lends it the weight of a second source. +5. **The `softwaredev/delivery/github` macro** — reports `Session-link trailer: OFF` + from `attribution.commit`/`.pr` alone, having never read `sessionUrl`. + +The macro is the sharp end. It rendered the wrong fact as a reassuring green light in +every session, including sessions whose system prompt carried the trailer. In the +control run above, the session had to reason around its own status line contradicting +its system prompt. + +The count itself makes the point. This ADR's first draft listed **four** carriers and +missed ADR-163 — a document it cites approvingly in its own Decision and lists in +`related:`. A reviewer found it. An argument about how unchecked claims propagate is +exactly the argument whose own citations go unchecked, and being the one making it +confers no immunity. + +The lesson generalizes past this key: ADR-162 froze a **contingent upstream fact** as +though it were a durable architectural truth. Upstream behavior is evidence with a +shelf life. A decision that depends on it should say so, and should name the +observation that would overturn it. + +## Decision + +**`attribution.sessionUrl: false` is the primary control.** It is authored as a +fragment in the settings store ([[ADR-147]]) and reaches `~/.claude/settings.json` +through the dotfiles source-of-truth ([[ADR-163]]) — `20-attribution.md`, carrying +`commit: ""`, `pr: ""`, and `sessionUrl: false`. Suppression happens at the source: the +model is never instructed to emit the link, so there is nothing to catch. + +**The ADR-162 hook is retained, demoted to backstop.** Its mechanism is unchanged and +its design rationale still holds — deny rather than rewrite (rewriting requires forcing +`permissionDecision: allow`, which would bypass the operator's own permission rules on +outward-facing `gh` commands), user scope rather than per-project, trailer/footer +position only so inline prose mentions pass. What changes is its standing: it is no +longer the sole defense. + +It is retained for two reasons, and two only: + +- **Undocumented** — absent from the official settings docs (`#69614`), so it cannot be + rediscovered from primary sources and may change without notice. A key we found by + reading a changelog is a key upstream never promised us. +- **Default-on, and delivered by an independent path** — a host that never receives the + fragment leaks by default. This bites only because hook and fragment travel + *separately*: the hook rides reconcile-owned `hooks`, the fragment rides dotfiles → + fragment store → project. Were they one pipeline, the backstop would be absent exactly + when it was needed and would cover nothing. [[ADR-163]] records a real instance — + suppression held on one host and leaked on another. + +A third reason was considered and **rejected as circular**: that the changelog +scope-qualifies the key to "web and Remote Control sessions". The qualifier cuts both +ways. If Remote Control is the likely reason the link reached commits at all — and this +operator runs `remoteControlAtStartup: true` — then the qualifier may describe *complete +coverage*, not a gap. Reasoning "the boundary is uncharacterized, therefore keep the +backstop" while also holding "the backstop stands, therefore the boundary needn't be +characterized" is ADR-162's error inverted: an unexamined boundary dressed as a known +risk. The boundary is genuinely uncharacterized. That is an open question, not evidence. +The two reasons above carry the decision without it. + +The backstop costs nothing when the primary control works: the model never emits the +link, so the hook never fires. It is insurance against the fragment not being projected +and against the key changing upstream — not against a session-type gap we have never +observed. + +**Status surfaces must read the key that governs.** The `github` macro reports the +session link off `attribution.sessionUrl` (absent means ON), and reports the +Co-Authored-By / "Generated with Claude Code" footers separately off `commit`/`pr`. +Conflating two independent controls under one label is the error this ADR corrects; a +surface that asserts a fact it never read is worse than no surface. + +**Defense in depth is the posture, not a single mechanism.** Setting suppresses, hook +catches, macro reports. Each is independently insufficient. + +## Consequences + +### Positive + +- The link is suppressed **before** the model is instructed to emit it, rather than + denied after it is written. No retry round-trip, no reliance on the model adapting + within a session. +- The hook stops being load-bearing, so a gap in its command coverage (ADR-162 lists + them) is no longer a leak on its own. +- The operator's config carries the reason: the fragment body records why the key is + set, what governs what, and why the hook remains. + +### Negative + +- The primary control is undocumented upstream. If Anthropic renames or removes + `sessionUrl`, the fragment silently stops working — and the failure is invisible + (a link appears where none did before). The backstop exists for exactly this. +- The key is not in the vendored SchemaStore schema, which declares `attribution` with + `additionalProperties: false` and only `commit`/`pr`. `ways settings new + attribution.sessionUrl` therefore refuses it, and the fragment is hand-authored. + Tracked separately, with an upstream PR as the durable fix. +- Two mechanisms to maintain where ADR-162 had one. + +### Neutral + +- ADR-162's hook code is untouched by this decision; only its header comment needs its + rationale corrected. +- The scope boundary of `sessionUrl` ("web and Remote Control sessions") is uncharted + for other session types — it is not known whether such sessions receive the link at + all. Characterizing it is cheap (one controlled run with `remoteControlAtStartup` + off) and would either close the question or reveal a real gap. Left open + deliberately, and named here so the omission is visible rather than assumed away. + +## Alternatives Considered + +- **Keep the hook as sole defense; ignore the setting.** Rejected: preserves the + false premise, and pays a deny/retry round-trip on every commit for a leak the + harness will suppress for free. Being wrong for a good reason is still wrong. +- **Adopt the setting; retire the hook.** Rejected. The setting is undocumented, + default-on, and scope-qualified — three properties that individually argue for a + backstop. ADR-162's instinct was half right: the harness honors the setting fine; + the failure was in our reading of it. That is a reason to verify the setting, not to + trust it alone. +- **Amend ADR-162 in place.** Rejected: the *decision* changes, not merely its + context. The hook's role moves from sole defense to backstop, and a new control is + introduced. Precedent ([[ADR-104]] → ADR-123) supersedes when the decision moves and + preserves what remains correct — which is what this does. +- **Set `sessionUrl` directly in `~/.claude/settings.json`.** Rejected: unmanaged, + host-local, and undocumented — the same shape of failure ADR-162 recorded, where a + repository without local config leaked because the control never traveled. The + fragment store carries it to every host by construction. diff --git a/docs/architecture/platform/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md b/docs/architecture/platform/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md new file mode 100644 index 00000000..08bfedc5 --- /dev/null +++ b/docs/architecture/platform/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md @@ -0,0 +1,274 @@ +--- +contract: adr/v1 +kind: decision +verb: retire +capability: config +targets: + - cli:ways-settings + - skill:ways-settings +supersedes: + - ADR-147 + - ADR-149 +basis: + - evidence: the inert Write(~/.claude/**) and Write(~/.ssh/**) entries raised a launch-time warning and were re-added by ways reconcile after deletion + - standard: 'Claude Code settings documentation (code.claude.com/docs/settings): no user-scope settings.local.json exists' + - precedent: ADR-163 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-18 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-147]]' + - '[[ADR-149]]' + - '[[ADR-152]]' + - '[[ADR-163]]' +imported: + from: docs/architecture/system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md + format: v0 + status: Accepted +--- + +# ADR-169: agent-ways relinquishes user-scoped settings.json; retains only its operational baseline + +## Context + +`~/.claude/settings.json` is Claude Code's declarative user configuration file. It +is also a *shared-write object*: several independent parties mutate it, and none of +them coordinates with the others. + +- **Claude Code owns it.** `/config` persists user preferences (theme, model, + `autoCompactEnabled`, tui/notification toggles) straight into + `~/.claude/settings.json` at user scope, and the write target cannot be + redirected. Claude Code additionally writes some keys *in situ* during a session + (`advisorModel`, `effortLevel` — the "known in-situ writers" of ADR-147). +- **agent-ways writes two slices** via `ways reconcile` + (`tools/ways-cli/src/cmd/settings_merge.rs`): the hook entries it ships, and its + permission strings (`WAYS_PERMS` in `permissions.allow`, `WAYS_DENY` in + `permissions.deny`, the ADR-152 secret-path baseline). This is a three-way merge + keyed on a persisted last-applied base — the "one shared-write seam" of ADR-142. +- **agent-ways also ships a config-management service** on top of that seam: the + composable fragment store (ADR-147) and the operator interview skill (ADR-149), + exposed as `ways settings` and the `ways-settings` skill. Its projector + (`settings/project.rs`) is a *second* independent three-way writer that owns the + fragment keys (`model`, `env`, `statusLine`, …) and deliberately skips + hooks/permissions. + +Three problems motivated this decision: + +1. **Over-reach.** The fragment store makes agent-ways a general manager of + *user-scoped* Claude Code configuration — config that is not agent-ways' concern. + `~/.claude/` is properly a dotfiles-class directory (the user's own machine + config), and an application framework offering to own the whole of it is an + overstep. ADR-163 already recorded the first symptom: agent-ways was + force-claiming `statusLine`, "exactly the anti-pattern ADR-147 set out to end," + and named dotfiles the cross-host source of truth — but left the *engine* in + agent-ways. + +2. **A concrete redundancy bug.** `WAYS_PERMS` ships both `Edit(~/.claude/**)` and + `Write(~/.claude/**)`, and `WAYS_DENY` ships `Write(~/.ssh/**)` alongside its + `Edit`/`Read` siblings. Claude Code's file-permission checks are satisfied by the + `Edit(path)` rule (which covers the file-writing tools), so the `Write(...)` + entries are inert and surface as a launch-time warning. Deleting them from the + live `settings.json` does not stick: the reconciler re-adds them from the source + constants on the next `ways reconcile`. + +3. **Two things wanted to be true at once, and were assumed incompatible.** The + user should own their Claude config through dotfiles; and a fresh agent-ways + install with no dotfiles must still be safe and usable. An early proposal to + collapse to a single writer — lifting the baseline into the dotfiles store — was + examined and rejected: a security baseline that only holds when dotfiles is + deployed has a hole. + +Two empirical facts (confirmed against `code.claude.com/docs/settings` and the +Claude Code precedence chain) constrain any solution: + +- **There is no user-scope `settings.local.json`.** The `.local.json` override + exists only at *project* scope (`.claude/settings.local.json`). So there is no + user-level side file into which a tool could isolate `/config`'s writes; every + user-scope writer lands in the same `~/.claude/settings.json`. +- **Because `/config` and the in-situ writers mutate that single file, any tool + that owns keys in it must use base-preservation (three-way) merge, not a naive + compile-and-replace.** A compiler that rewrites the file from fragments would + clobber the operator's live `/config` toggles on every deploy. + +Finally, an explicit product constraint: **the dotfiles-side config tool must work +without agent-ways installed at all**, and (symmetrically) agent-ways must work +without dotfiles. Neither may depend on the other's binary. + +## Decision + +**agent-ways relinquishes management of user-scoped Claude Code configuration and +retains only its own operational and security baseline in `settings.json`.** + +1. **Retain the operational baseline, unchanged in mechanism.** agent-ways keeps + `settings_merge.rs` as a three-way, base-preserving, self-auditing writer of + exactly the slices it *must* own for the framework to function and to be safe + standalone: + - `hooks` — the entries agent-ways ships (SessionStart disclosure, the ADR-162/167 + deny backstop, etc.). + - `permissions.allow` — the operational allows for its own binaries + (`Bash(ways:*)`, `Bash(attend:*)`, `Bash(attend-chat:*)`, `Bash(way-embed:*)`, + `Edit(~/.claude/**)`). + - `permissions.deny` — the ADR-152 secret-path baseline (opt-out preserved). + + This slice is disjoint from any user-preference key, self-audits (reverts on any + change to an unmanaged field), and ships *with the application* so a fresh install + is safe and usable with no dotfiles present. + +2. **Hand the user-config service to dotfiles; keep only self-management.** What + `ways settings` *did* — author and project user-scoped configuration — did not + work cleanly and conflicted with the premise that dotfiles owns the user's own + config. That capability is handed to the dotfiles tool. agent-ways removes the + fragment store and interview apparatus outright: the `ways settings` CLI + subcommands (`settings/project.rs`, `settings/compile.rs`, and siblings), the + `ways-settings` skill, and the now-orphaned schema plumbing + (`settings_schema_*` in `paths.rs`, `settings_schema_url` in `config.rs`, the + `refresh-settings-schema.sh` script, and the vendored + `share/claude-code-settings.schema.json`). This supersedes **ADR-147** and + **ADR-149**. agent-ways keeps *only enough configuration capability to manage + itself* — the baseline in (1), possibly nothing more; it retains **no + user-config surface**. User-scoped configuration (model, `statusLine`, `env`, + user-authored permissions, prefs) is no longer agent-ways' concern. + +3. **User config moves to dotfiles.** The dotfiles-side tool owns the fragment store + and authoring experience and carries **its own standalone three-way merger** — it + does not call the `ways` binary. Its architecture is recorded in a companion ADR + in the dotfiles repository. agent-ways makes no claim on the keys that tool owns. + +4. **Coexistence is by disjoint ownership, not a shared engine.** Because each tool + must run standalone, each carries its own three-way base-preserving merger over a + *disjoint* set of owned keys, each keying on its own last-applied base and each + self-auditing hands-off on keys it does not own. Multiple such writers over one + `settings.json` is safe precisely because their owned sets do not overlap — the + model ADR-163 validated in practice. A single unified engine was rejected (see + Alternatives): the standalone-independence requirement forecloses it. The exact + invariant that makes independent writers safe — disjoint scalar/object ownership, + additive-union on shared lists with each writer removing only what its own base + recorded, per-writer last-applied base, and self-audit hands-off — is the + **peer-writer coexistence contract** specified in + `docs/architecture/platform/ADR-500-settings-json-three-way-merge-spec-and-peer-writer-coexistence-contract.md`. The + dotfiles-side tool ports the same merge algorithm from that spec (shared design + lineage, not a runtime dependency), so the two mergers behave identically without + coupling. + +5. **Fix the redundancy.** Remove `Write(~/.claude/**)` from `WAYS_PERMS` and + `Write(~/.ssh/**)` from `WAYS_DENY`. `Edit(...)` already covers the file-writing + tools; the `Write(...)` entries are inert and only produce the launch warning. + Removing them at the source constants makes the fix durable through reconcile. + +6. **No user-scope local-override layer.** An earlier design sketch proposed a + `settings.local.json` "escape hatch" layer; it does not exist at user scope and is + dropped. + +7. **Authoritative key-set partition.** Every user-scope `settings.json` key falls + into exactly one of three buckets (mirrored in the dotfiles-side ADR-010 so the + partition is agreed on both sides and disjointness holds by set-subtraction, not + guesswork): + + - **(A) agent-ways baseline** — `hooks` (its shipped entries) + + `permissions.allow` `{Bash(ways:*), Bash(attend:*), Bash(attend-chat:*), + Bash(way-embed:*), Edit(~/.claude/**)}` + `permissions.deny` (`WAYS_DENY`). This + is exactly the `WAYS_PERMS`/`WAYS_DENY` constants — the constant *is* the + boundary. + - **(B) user/dotfiles** — everything else user-authored, including the + agent-ways-*adjacent* tooling that agent-ways does **not** ship as a core binary + (`way-match`, `kg`, `mmaid`, `adr`/`adr-tool`, the knowledge-graph and + thinking-strategies MCP servers, shell/prompt tooling, generic `Bash(...)`, + `Read(~/**)`, …). Owned by the operator via the dotfiles config tool. + - **(C) Claude Code runtime** — keys Claude Code writes autonomously (`model` and + the `/config` toggles; `advisorModel`, `effortLevel`, and other in-situ writes). + **Neither tool owns these**; both must leave them to base-preservation. They must + never be declared as a managed fragment, or the tool would thrash against Claude + Code on every run. + +8. **Relinquish protocol for the ownership handoff.** The steady-state disjoint + contract assumes ownership never moves. The retirement *moves* ownership of the + user-fragment keys (`statusLine`, `attribution`, `env`, …) from agent-ways' + projector to the dotfiles tool, and that transition has a hazard the contract does + not cover: if agent-ways' final act treated those keys as *deprecated-ours* (in its + base, absent from `ours`), its three-way merge would **remove** them from + `settings.json` — clobbering the value the dotfiles tool now asserts, order- + dependently. The retirement therefore **relinquishes** rather than deprecated- + removes: it clears agent-ways' fragment base for those keys and **leaves the live + values in place** as foreign keys for the dotfiles tool to adopt. In practice, because + the projector is *removed entirely* (not shipped in a deprecation mode), it simply + never runs deprecated-removal again — removal *is* the relinquish, **provided the + removal performs no final "cleanup" projector pass**. Adoption on the other side is + the ordinary "migrating" behavior: a live foreign value is asserted as `ours`, the + base is seeded from it, and the result is idempotent. Once agent-ways stops touching + the keys, steady-state disjointness is restored and run order no longer matters. + +**Ratification:** the "how minimal" question — total removal of the user-config +service vs. keeping a minimal affordance for no-dotfiles operators — was ratified in +favour of *self-management only*: agent-ways keeps the baseline in (1) and no +user-config surface. An agent-ways-only operator uses the baseline plus raw +`/config`; managed user config is a dotfiles-tool adoption away. Retaining any +user-config service was rejected as reintroducing the over-reach this ADR removes. A +minimal user surface, if ever wanted, is a purely *additive* future change and does +not gate this decision. + +## Consequences + +### Positive + +- agent-ways stops owning config that isn't its concern; `~/.claude/settings.json` + returns to being the user's own file plus one narrow, auditable framework slice. +- The launch-time permission warning is fixed durably. +- The security/operational baseline still travels with the app, so a standalone + install is safe and usable with no dotfiles. +- Each tool is independently installable and testable; neither depends on the other. +- Removes a whole class of dual-control confusion: agent-ways no longer has two + writers into `settings.json` (the fragment projector goes away). + +### Negative + +- The three-way merge logic is duplicated across repositories (agent-ways keeps its + baseline merger; dotfiles builds its own). This is the deliberate price of the + standalone-independence constraint — one tested engine would have been less code + but would have coupled the tools. +- Operators who used `ways settings` must migrate their fragments to the dotfiles + tool. A migration note is required. +- Superseding two Accepted ADRs (147, 149) is a non-trivial reversal of recent + design; the reasoning must be legible to anyone who read those first. + +### Neutral + +- ADR-152's deny baseline is retained as-is, now framed explicitly as part of the + operational baseline agent-ways keeps. +- ADR-163's "dotfiles is the source of truth" direction is carried to completion: + the *engine* follows the *store* to dotfiles, rather than the store feeding an + engine that stayed behind. +- ADR-142's projection model is unchanged except that the shared-write seam narrows + to the baseline slice only. + +## Alternatives Considered + +- **Single unified engine, layered fragment stores (L1 app-baseline / L2 user / + L3 host / L4 local).** One loader composing ordered fragments, borrowing the + dotfiles `zshrc` `conf.d` shape. Rejected on two grounds: (a) the standalone + constraint requires each tool to write `settings.json` without the other, so a + single engine cannot live in only one repo; (b) the L4 user-scope + `settings.local.json` layer it relied on does not exist in Claude Code. + +- **Lift the baseline (hooks + `WAYS_DENY`) into the dotfiles store; one writer + total.** Rejected: a security baseline that only exists when dotfiles is deployed + strands a fresh agent-ways install. The baseline must ship with the application. + +- **dotfiles depends on the `ways` binary as its merge engine (store-only + dotfiles).** The least-code option and initially preferred. Rejected by the + explicit constraint that the dotfiles tool must work with no agent-ways installed. + +- **A naive compiler (compile-and-replace) instead of a three-way merger.** + Rejected: `/config` and the in-situ writers mutate the same user-scope file, and a + compiler would clobber the operator's live toggles on every deploy. Base + preservation is mandatory. + +- **Keep `ways settings` as-is and merely re-posture it as opt-in.** Rejected as + insufficient: leaving the engine in agent-ways keeps the over-reach and the + dual-writer surface the user objected to; ADR-163 already showed re-posturing + alone does not stop the framework from claiming user keys. diff --git a/docs/architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md b/docs/architecture/platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md similarity index 100% rename from docs/architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md rename to docs/architecture/platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md diff --git a/docs/architecture/platform/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md b/docs/architecture/platform/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md new file mode 100644 index 00000000..7d3577a1 --- /dev/null +++ b/docs/architecture/platform/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md @@ -0,0 +1,89 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: config +basis: + - evidence: 'the Cypress survey (docs/architecture/practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md, issue #465) proposed a second refusing hook ported from a timeout guard' + - evidence: 'the Bash tool already bounds every foreground command: 120 seconds by default, 600 at most' + - precedent: ADR-162 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-09-10 +deciders: + - aaronsb + - claude +related: + - ADR-162 + - ADR-178 +imported: + from: docs/architecture/system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md + format: v0 + status: Accepted +--- + +# ADR-181: Guard hooks: a blocking PreToolUse class, shipped deactivated + +## Context + +Every hook this project ships injects context and exits zero. One departs from that: `strip-session-link-pre.sh` (ADR-162) exits 2 to refuse a commit that would publish a session link. It is the only hook that can stop a tool call, and nothing in the corpus names the class it belongs to. + +The Cypress survey (`docs/architecture/practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md`, issue #465) proposed a second refusing hook, ported from a guard that blocks any shell command lacking an explicit `timeout` or a detached launch. That guard was written for harnesses that run foreground commands unbounded. This harness is different on three points: + +1. **Every foreground command is already bounded.** The Bash tool enforces a timeout, 120 seconds by default and 600 at most, and refuses a bare foreground `sleep`. +2. **Detached work has a first-class form.** `run_in_background: true` on the Bash tool, and the Monitor tool for watching a condition. The hook sees `run_in_background` in `tool_input`. +3. **The permission system already gates Bash** by command prefix, and the operator tunes that list per project. + +What the harness leaves open is narrower: pattern kills (`pkill`, `killall`, `kill` by name) that can match the agent's own shell or an operator process; interactive-prone commands (`sudo` without `-n`, `ssh` without `BatchMode`, `docker exec -it`, package managers without their flag) that sit on a prompt until the timeout; log followers and attached containers that never return in the foreground; and pipe-to-shell. + +The operator's read: Bash execution is already guarded enough, and another refusal wired in by default is unwanted. + +## Decision + +**Establish guard hooks as a named class with a contract. Ship `check-bash-bound.py` as a member, deactivated.** + +A guard hook: + +1. Refuses by exit 2 with the reason and the accepted form on stderr. The model receives the reason as the tool result and chooses again. +2. Exits only 0 or 2. Every internal failure path exits 0 with one line on stderr. A bug in a guard degrades to no guard. +3. Is wired without `|| true` when it is wired at all. +4. Refuses a closed list named in the script, each entry with its repair. A guard has no semantic lane. +5. Honors the harness's own forms: `run_in_background` and a `timeout` prefix exempt the never-returns class. + +`check-bash-bound.py` refuses the four classes above and lives at `hooks/ways/check-bash-bound.py` with its verdict test at `tests/test-bash-bound.sh`, which runs in the suite. **It is not wired in `settings.json`.** An operator who wants it adds one entry under `hooks.PreToolUse` for matcher `Bash`: + +```json +{ "type": "command", "command": "${HOME}/.claude/hooks/ways/check-bash-bound.py" } +``` + +**Bounded execution is guidance.** The `softwaredev/environment/bounded-execution` way fires on the same command classes through the ordinary inject path and carries the discipline: background the never-returning command, give the interactive one its flag, kill by pid, download an installer before running it, and claim running only on a liveness signal. + +Activating the guard by default needs its own ADR, and the bar is ADR-162's: an irreversible or disclosing outcome that injected guidance has been observed to fail to prevent. + +## Consequences + +### Positive + +- Bash stays as permissive as the harness and the operator's permission list make it. No hook runs in front of every shell call unless the operator asks. +- The class has a name, a contract, and a tested reference member. A project that wants the refusal gets it with one settings line. +- The discipline still reaches the model at the moment it is about to run one of those commands. + +### Negative + +- A pattern kill or a pipe-to-shell that the model decides on despite the guidance runs. The harness timeout and the permission prompt are the remaining stops. +- A deactivated script drifts. Its test runs in the suite, which keeps it working; whether its refused list still matches the harness is checked only when someone activates it. + +### Neutral + +- `strip-session-link-pre.sh` is retroactively a member of the class. Its contract already matches. +- The Makefile marks `.py` hooks executable alongside `.sh`, so the opt-in path works after `make install`. +- The `code/security/guards` way, which says a guard that cannot block is decoration, describes controls a project ships in its own code. It does not argue for wiring this one. + +## Alternatives Considered + +- **Wire the guard by default (the first draft of this ADR).** Rejected by the operator: Bash execution is already guarded enough. +- **Withdraw the guard entirely.** Rejected in favor of shipping it deactivated: the port and its forty-case test exist, and a project with a different risk posture can opt in without re-deriving them. +- **Port the Cypress guard unchanged, refusing any command without an explicit bound.** Rejected. The harness already bounds foreground commands; the rule would refuse ordinary builds and installs. +- **Fold the guard into the `ways` binary's scan path.** Rejected for now; the scan path is an inject path with a semantic lane, and a deactivated script is easier to read and to opt into. diff --git a/docs/architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md b/docs/architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md new file mode 100644 index 00000000..620a5537 --- /dev/null +++ b/docs/architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md @@ -0,0 +1,96 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - install + - config +basis: + - evidence: 'a first install on a CLAUDE_CONFIG_DIR machine replaced a real skills directory and dropped the user''s hooks; PRs #501 and #502 fixed the data-loss paths' + - precedent: ADR-142 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-09-17 +deciders: + - aaronsb + - claude +related: + - ADR-131 + - ADR-142 + - ADR-144 + - ADR-185 +imported: + from: docs/architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md + format: v0 + status: Accepted +--- + +# ADR-184: Installation and activation are separate states: targets as the unit of activation + +## Context + +ADR-142 made `~/.claude` a thin projection of an application that lives under `$XDG_DATA_HOME/agent-ways`. The projection has one destination, and that destination is a constant: the installer, the reconciler's default, and all 32 shipped hook commands name `~/.claude` by path. Installing and activating are one act. The installer stages the source, builds the binaries, and immediately projects into the one place it knows. + +Claude Code does not have one place. `CLAUDE_CONFIG_DIR` relocates its whole config directory, and operators use that to keep profiles apart, switched per folder. A first install on such a machine landed in the default directory, replaced a real skills directory there, and dropped the user's own hooks from settings, because the installer chose a destination without saying so and the reconciler treated whatever it found there as its own. PRs #501 and #502 fixed the two data-loss paths. They did not change the fact that the destination is chosen silently. + +Two further needs came with the same report: wiring ways into one profile or one repository rather than everywhere, and turning the system off to judge it. Neither has a form today. There is no uninstall. + +## Decision + +**Installation and activation are separate states. A target is a Claude Code config directory named in the user config. The reconciler converges every enabled target and withdraws from every disabled one. The installer never activates.** + +1. **Three states.** Absent: nothing staged. Installed: source staged under XDG data, binaries built and on PATH, no target enabled. Active: at least one enabled target. Every user-owned file is untouched in the first two states. `ways status` names the state and the targets. + +2. **Targets live in the user config.** `config.yaml` under the agent-ways config root gains a `targets` list. Each entry names a config directory and carries `enabled` and `observe` flags. When the key is absent the list is one implicit entry, the default config directory, enabled. An install that predates this decision keeps working unchanged, and its next bare reconcile records the implicit entry as the explicit one, so an existing install becomes explicit without an operator act. Writing the key makes the list explicit and the implicit entry stops. + + **Each target carries its own configuration set.** A target's `config.yaml`, at `targets/<key>/config.yaml` under the agent-ways config root or at the path its entry names, holds the same keys as the user config and is layered over it for every session running under that target's config directory. The session names its target through `CLAUDE_CONFIG_DIR`, with the default directory as the fallback. The `targets` key itself is read from the user layer only, so a target cannot redirect the list. Two profiles on one machine can run different languages, disabled domains, or thresholds without sharing them. + +3. **Verbs on `ways config`.** `targets` lists entries with their converged state. `target add`, `enable`, `disable`, and `remove` edit the list and reconcile. When the list is empty, `targets` runs the bootstrap: discovery, candidates, the plan per target, confirmation. The installer's last act is to invoke it or, unattended, to print it. + +4. **Reconcile converges per target.** An enabled target receives the projection roots and the hooks merge into its `settings.json`, with a merge base kept per target under the state root. A disabled target is withdrawn: our symlinks are unlinked, our hooks block is removed through the same three-way merge base that wrote it, and nothing else in the directory is touched. Reconcile is idempotent in every state, so `ways update` is safe wherever the operator stopped. + +5. **`observe` is orthogonal to `enabled`.** Transcript-derived readers resolve their roots from the targets with `observe` true, defaulting to the value of `enabled`. A disabled target can stay in reports; an enabled one can stay out of them. + +6. **Per-project off is a one-line switch.** `enabled: false` in a project's `ways.yaml` (the file ADR-131 already uses for per-way disable) makes the hook entry point exit before any scan. Hooks still spawn; nothing is injected. + +7. **Discovery is heuristic and never a verdict.** Candidates come from three signals: `CLAUDE_CONFIG_DIR` in the environment, shell rc files, and direnv's allow records; directories under the home and XDG config roots with the shape of a Claude Code config dir; and the environment of running `claude` processes. Attended mode confirms them. Unattended mode with more than one candidate and no explicit list exits with the literal command that would resolve it. + +8. **One bus across targets.** Attend's signals, channels, instance roster, heartbeats, and state live under the user cache, keyed by the user, never by a config directory. Agents in sessions under different targets are peers on the same bus. What attend resolves through a config directory, its own session identity and its peers' from the session records, walks every target's directory plus the one `CLAUDE_CONFIG_DIR` names, so a session under a relocated profile is a full peer and never a fallback identity. + +9. **Unattended stays.** The installer keeps a flag for pipelines. It stages and builds, reads the targets list, and reconciles. With no list and one candidate it activates that one. With no list and several it stops with the hint. It never guesses. + +Reversibility: expensive. The implicit-target compatibility in item 2 keeps every existing install on the old behavior until it writes the key, so the model can be withdrawn by removing the verbs and leaving the key ignored. After the installer hands off to the bootstrap, reversing means restoring the installer's own projection step. + +This decision amends ADR-142's single projection destination and ADR-144's manifest, which described the desired state of one directory and now describes the desired state of each target. Neither is superseded whole. + +## Consequences + +### Positive + +- The installer makes no choice on the operator's behalf. Nothing user-owned changes until a target is named. +- Withdrawal exists. Disabling a target is a verb, and it restores the directory through the same base that changed it. +- Per-profile and per-repository installs become entries in one list. Issue #503 and PR #504 reduce to a second target and a hooks-only slice on a target. +- Leaving the installer early leaves the system installed and inactive, which is a valid state with a one-line resume. +- Judging the system is a switch and a query: `observe` plus the config-dir stamp on events makes sessions with ways on and off comparable. + +### Negative + +- A second target is not honest until the hook commands and scripts stop naming `~/.claude`. Issue #503 is the prerequisite, and until it lands the only honest target is the default directory. +- Withdrawal through the merge base cannot remove hooks written before a base existed. The first apply after this decision seeds one, and PR #502 narrowed what that seed claims. +- Discovery reads shell rc files and the process table. Both are heuristics, and the attended flow has to make that visible rather than present a candidate as a fact. + +### Neutral + +- The deployment way loses its clobber branch. On a fresh install there is no decision for a guiding Claude to make, only the bootstrap to point at. +- The statistical readers change one resolver from singular to plural. Events gain a config-dir field so the filter in item 5 has something to key on. +- ADR-185 sets the output contract the new verbs follow. + +## Alternatives Considered + +- **Infer scope from state files on a bare reconcile.** PR #504's design. Rejected: a mode inferred from disk on every update is a new failure class, and the review found it flips when the recorded directories are gone. The targets list is the same state made explicit. +- **Keep the installer projecting, with a prompt.** Rejected: the prompt lives in bash, `make setup` and the one-liner stay two code paths, and an unattended run still has to guess. +- **A separate `ways target` noun.** Rejected in favor of `ways config`: the list is user-scope configuration that survives updates, and it belongs in the file that already holds the other user-scope switches. +- **Uninstall as a script.** Rejected: withdrawal through the merge base is the only method that knows what we wrote, and it already exists for the forward direction. +- **Honor `CLAUDE_CONFIG_DIR` only, no list.** Rejected: it gives one destination per invocation and no record of which directories are active, so update, status, and withdrawal have nothing to iterate. diff --git a/docs/architecture/platform/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md b/docs/architecture/platform/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md new file mode 100644 index 00000000..e2573ccb --- /dev/null +++ b/docs/architecture/platform/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md @@ -0,0 +1,72 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: cli +basis: + - evidence: ways config show prints a Rust debug rendering, and three of the binary's forty verbs take --json + - precedent: ADR-111 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-09-17 +deciders: + - aaronsb + - claude +related: + - ADR-111 + - ADR-184 +imported: + from: docs/architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md + format: v0 + status: Accepted +--- + +# ADR-185: CLI output contract: structured output for people, JSON for machines + +## Context + +`ways config show` prints a Rust debug rendering of the config struct. Three of the binary's forty verbs take `--json`. `attend` renders tables through `agent-fmt` while `ways` prints ad hoc lines. ADR-111 chose the argument parser and said nothing about output, so every verb has decided for itself. + +ADR-184 adds verbs whose output an operator reads to decide what to activate, and whose output a script reads to decide what to do next. Both readers need a stable form. + +## Decision + +**Every verb in `ways` and `attend` renders for a person by default, through `agent-fmt`, and renders JSON under `--json`. A debug rendering never reaches stdout.** + +1. **Human output** is structured: tables for lists, labeled rows for records, color where the terminal supports it and none where it does not. The formatter already in the workspace is the one renderer. + +2. **`--json`** emits one document on stdout and nothing else there. Diagnostics go to stderr. The document's shape is the verb's contract, and a change to it is a change to the verb. + +3. **Two views where a resolved value differs from a stored one.** `--json` emits what the file says. `--json --effective` emits the resolved state with defaults applied. Only the first is accepted back by an `apply` verb. Round-tripping the effective view would freeze every default at its current value. + +4. **Round trip.** A verb that shows a configuration object has a partner that accepts the same JSON back. The pair is the test: show, apply, show again, byte-equal. + +5. **Exit codes carry the verdict.** Zero for done, one for a failure the verb reports, two for bad usage. A verb that refuses to act, such as reconcile at a real directory, uses a code the caller can distinguish from failure. + +Reversibility: reversible. Each verb converts on its own; a verb not yet converted is a defect against this contract, and nothing depends on the order. + +## Consequences + +### Positive + +- An operator reads a table; a script parses a document. Neither has to parse the other's form. +- The bootstrap in ADR-184 shows the plan per target in the same renderer attend uses, which people already read. +- Round-trip pairs make configuration editable by tools without hand-editing YAML. + +### Negative + +- Forty verbs, three converted. The sweep is spread over ordinary work, and the contract holds before the sweep completes. +- Effective-versus-stored is one more flag to explain. + +### Neutral + +- `agent-fmt` becomes a dependency of every output path in `ways`, as it already is in `attend`. +- The config verbs in ADR-184 are the first converted, since they are new. + +## Alternatives Considered + +- **JSON by default, human on a flag.** Rejected: the first reader of every verb is the operator at the terminal, and a guiding Claude reading the output prefers the labeled form too. +- **A third format such as YAML for round trips.** Rejected: one machine format keeps the pair test simple, and the apply verb writes the config file's YAML. +- **Leave output to each verb.** The status quo. Rejected by the defect that opened this decision. diff --git a/docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md b/docs/architecture/platform/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md similarity index 100% rename from docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md rename to docs/architecture/platform/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md diff --git a/docs/design-notes/settings-json-merge-spec-and-peer-writer-contract.md b/docs/architecture/platform/ADR-500-settings-json-three-way-merge-spec-and-peer-writer-coexistence-contract.md similarity index 98% rename from docs/design-notes/settings-json-merge-spec-and-peer-writer-contract.md rename to docs/architecture/platform/ADR-500-settings-json-three-way-merge-spec-and-peer-writer-coexistence-contract.md index 90d7336b..f054df20 100644 --- a/docs/design-notes/settings-json-merge-spec-and-peer-writer-contract.md +++ b/docs/architecture/platform/ADR-500-settings-json-three-way-merge-spec-and-peer-writer-coexistence-contract.md @@ -1,4 +1,15 @@ -# settings.json three-way merge — spec and peer-writer coexistence contract +--- +contract: adr/v1 +kind: spec +capability: install +status: accepted +date: 2026-07-18 +deciders: + - aaronsb +related: [] +--- + +# ADR-500: settings.json three-way merge: spec and peer-writer coexistence contract Status: reference for ADR-169. Portable specification of the algorithm implemented in `tools/ways-cli/src/cmd/settings_merge.rs`, extracted so an independent tool diff --git a/docs/architecture/practice/ADR-013-ways-skills-governance-architecture.md b/docs/architecture/practice/ADR-013-ways-skills-governance-architecture.md new file mode 100644 index 00000000..a0ea18f7 --- /dev/null +++ b/docs/architecture/practice/ADR-013-ways-skills-governance-architecture.md @@ -0,0 +1,355 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: + - method + - governance +basis: + - evidence: an evaluation of Anthropic's Claude Code skills system against the ways system, prompted by doubt that ways duplicate the official primitives + - standard: 'Claude Code skills and hooks primitives: skills cannot fire on PreToolUse, file patterns or session state, or session-gate' + - evidence: governance matrix covers 12 ways with 37 control claims and 93 justifications across NIST, OWASP, ISO, SOC 2, CIS and IEEE (Decision §4) +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-02-09 +deciders: + - aaronsb + - claude +related: + - ADR-004 + - ADR-005 + - ADR-200 +imported: + from: docs/architecture/legacy/ADR-013-ways-skills-governance-architecture.md + format: v0 + status: Accepted +--- + +# ADR-013: Ways, Skills, and Governance Architecture + +> **Refined by [ADR-200](../governance/ADR-200-compliance-claims-and-session-derived-findings.md).** +> This ADR's core insights stand — ways as *activation cues* for latent knowledge, the +> Type 1–4 way classification, the author-as-compiler model, and the formation intuition +> in its "Shift Lead Wisdom" type. But its **epistemics are corrected**: wherever it calls +> provenance "evidence," "auditability," "proof," or what "turns Claude from a hack into a +> professional," read **claim**. A hand-authored mapping is a control-*design* assertion +> (SOC 2 Type I), not an assessed **finding** (SOC 2 Type II); a citation on demand is an +> assertion awaiting a finding, not proof. Its analogies (kitchen poster, shift lead, FAA +> tool) are illustration; the grounded vocabulary — *first line*, *tacit knowledge*, +> *claim/finding* — lives in ADR-200. Mechanism has also moved on: `governance.sh` is now +> `ways governance`/`ways-audit` (ADR-111), and provenance lives in a `provenance.yaml` +> sidecar, not way frontmatter (ADR-110). + +## Context + +An evaluation of Anthropic's official Claude Code skills system against the custom "ways" system revealed both complementary functions and overlapping capabilities. This ADR documents the architectural relationship between skills, ways, and governance provenance — and establishes a framework for how they compose into a professional-grade agent governance stack. + +The evaluation was prompted by genuine doubt: do ways and skills actually complement each other, or is the ways system duplicating what Anthropic's official primitives now provide? + +## The Three-Layer Professional Practice Stack + +### Layer 1: Skills — Capability ("Here's how to use the tool") + +Skills are what the agent CAN do. They represent organized knowledge and specific capabilities. + +**Analogy**: A mechanic's toolbox with foam cutouts, labels, and knowledge of what every tool does and how to use it. + +**Characteristics**: +- Semantically discovered by Claude based on intent (description matching) +- User-invocable via `/slash-commands` +- Can restrict tool access (`allowed-tools`) +- Can fork into subagent contexts (`context: fork`) +- Distributable via plugins/marketplace +- Follow the Agent Skills open standard (agentskills.io) + +**What skills cannot do**: +- Fire before a specific tool executes (no PreToolUse access) +- Fire based on file patterns being edited +- Fire based on session state (context threshold, file existence) +- Session-gate (fire once then stay silent) +- Differentiate agent/teammate/subagent scopes +- Carry governance provenance metadata + +### Layer 2: Ways — Policy ("Here's why you must account for every tool before closing the job") + +Ways are what the agent MUST do, and the practical guidance for how to do it. They represent codified policy in actionable form. + +**Analogy**: The rule that says you must account for every tool before closing the job — because if one's missing, it might be in the wing, the engine, or the break room, and you do not close out the job until it's located. + +**Characteristics**: +- Triggered by actions (tool use, file edits, commands), keywords, or state conditions +- Session-gated: fire once per session via marker system +- Support macros for dynamic context (shell scripts that run when the way triggers) +- Scope-filtered: can target agent, teammate, or subagent contexts +- Domain-organized with enable/disable via `ways.json` +- Built on Claude Code's hook primitives (PreToolUse, UserPromptSubmit, SessionStart, etc.) + +**What ways cannot do**: +- Be invoked by the user as a slash command +- Restrict which tools Claude can use +- Fork into isolated subagent contexts +- Be distributed via plugin marketplace +- Appear in Claude's context budget for auto-discovery + +### Layer 3: Provenance — Institutional Memory ("Because FAA regulation XYZ, which exists because...") + +Provenance is the evidence chain from repeated failure → institutional learning → codified standard → policy → way → agent behavior. + +**Analogy**: The FAA regulation that requires tool accountability, which exists because in a specific historical incident, a tool was left inside an aircraft and caused a failure. The regulation is the codified institutional memory of "this went wrong enough times that we wrote it down." + +**Characteristics**: +- Embedded in way frontmatter as `provenance:` blocks +- Stripped before context injection (zero runtime token cost) +- Queryable via the `governance-cite` skill +- Maps specific way directives to specific control requirements with justifications +- Creates a self-supporting network: multiple controls cross-referencing the same practice + +**What a claim provides** *(epistemics corrected by ADR-200 — this section originally read "What provenance provides")*: +- A statement that a way is *designed* to address specific controls — a control-design **claim** (SOC 2 Type I), not proof that it does. +- Traceability of intent: the lookup can surface which controls a way *claims* to address. This is an asserted mapping, not evidence — a citation on demand is a claim awaiting a **finding**, not a satisfied "prove it." +- Explicitness: a way with a claim states its intended control alignment, which makes that alignment *assessable*; substantiating it requires a session-derived finding (ADR-200). + +#### The Authorship Model: Provenance as Design-Time Artifact + +Provenance is stripped before the way reaches Claude's context — it never sees the control IDs, the justifications, or the regulatory citations at runtime. What Claude sees is the way body: the poster on the wall. This creates a specific authorship dynamic with three key properties: + +**Latent space activation.** Claude was almost certainly trained on the full text of NIST 800-53, OWASP, ISO 27001, and other regulatory corpora. That knowledge exists in the model's latent space but isn't precisely retrievable on demand. Provenance metadata serves as the design-time bridge: the way author maps their guidance to specific controls, which ensures the way content contains the right activation cues to surface that latent regulatory knowledge at runtime. The provenance doesn't teach Claude the controls; it ensures the way content is written in a way that *activates* what Claude already knows. + +**The author as compiler.** The way author carries the critical responsibility of *compiling* governance into actionable guidance. The provenance block is the author's working notes — "I wrote this bullet point because NIST CM-3 requires change classification" — but the bullet point itself must stand alone as useful, activating guidance. The author bridges the gap between regulatory language and practitioner language. A bad author writes "comply with CM-3"; a good author writes "use conventional commits with type prefixes" and traces it back to CM-3 in the provenance. The way is the compiled thought; the provenance is the source code. + +**Epistemic position determines authorship direction.** A developer writing a way compiles from experience upward — "this is how I've seen it done well." A security engineer compiles from controls downward — "this is what the regulation requires, expressed as practitioner guidance." A compliance officer might write the provenance first and derive the way content from it. The direction doesn't matter; what matters is that the compiled output (the way body) lands at the right level of abstraction for the practitioner reading it. + +In practice, Claude itself is likely to be the author of most ways — but the human provides the epistemic grounding: which controls matter, what governance posture the framework should carry, what scope to cover. The human sets the intent and the governance coverage; Claude compiles it into effective activation cues. This is a collaborative authorship model where the human's domain knowledge of *what matters* meets Claude's ability to express it in a form that activates its own latent knowledge effectively. + +## The Three Types of Ways + +Not all ways serve the same function. Forcing governance provenance onto all ways would be dishonest. The three types are: + +### Type 1: Kitchen Poster (Compiled Governance) + +Practical, actionable guidance that traces directly to regulatory controls. The worker gets "use parameterized queries" on the poster. Corporate policy (OWASP A03, NIST IA-5) is several layers deep behind that poster. The provenance block captures the chain. + +**Should have provenance**: Yes — that's their primary purpose. + +**Current ways in this category** (all with provenance): +- `softwaredev/commits` → NIST CM-3, SOC 2 CC8.1, ISO 27001 A.8.32 +- `softwaredev/security` → OWASP A03, NIST IA-5, CIS v8 16.12, SOC 2 CC6.1 +- `softwaredev/quality` → ISO 25010, NIST SA-15, IEEE 730 +- `softwaredev/deps` → NIST SA-12, OWASP A06, NIST RA-5 +- `softwaredev/testing` → NIST SA-11, IEEE 829, ISO 25010 +- `softwaredev/config` → NIST CM-6, CIS v8 4.1, NIST IA-5 +- `softwaredev/errors` → OWASP A09, NIST SI-11, NIST AU-3 +- `softwaredev/ssh` → NIST AC-17, NIST IA-2, NIST IA-5 +- `softwaredev/release` → NIST CM-3, SOC 2 CC8.1, NIST SA-10 +- `softwaredev/adr` → NIST CM-3, ISO 27001 A.5.1, NIST PL-2 +- `softwaredev/github` → SOC 2 CC8.1, NIST CM-3, ISO 27001 A.8.32 +- `meta/knowledge` → ISO 9001 7.5, ISO 27001 5.2, NIST PL-2 + +### Type 2: Shift Lead Wisdom (Experience-Derived) + +Practical wisdom that exists BECAUSE of the governance environment but isn't DIRECTLY traceable to a specific control. An experienced worker knows "when the fryer oil looks like that, change it" — not because a regulation says so, but because working in a regulated kitchen long enough teaches you patterns that prevent problems. + +**Should have provenance**: No — forcing it would be dishonest. These are reactions to governance, not implementations of it. + +**Graduation path**: Type 2 ways can graduate into Type 1 when experience gets codified. In highly regulated industries, "shift lead wisdom" routinely becomes SOPs and best practice standards after enough incidents. The fryer oil example becomes "change oil every N hours per health code §X" once someone gets sick. The architecture supports this migration naturally — add a provenance block and the way changes type. The classification isn't permanent; it reflects the current state of formalization. A way without provenance today may earn it tomorrow when the pattern it captures gets traced to a control. + +**Current ways in this category**: +- `softwaredev/debugging` — experience-based debugging process +- `softwaredev/patches` — "never hand-write patches" is experience, not regulation +- `softwaredev/performance` — performance analysis workflow +- `softwaredev/api` — API design omissions Claude commonly makes +- `softwaredev/design` — design discussion framework +- `softwaredev/migrations` — schema change practices + +### Type 3: HQ Policy Manual (Governance About Governance) + +Meta-governance: how to write the poster, how to structure policy, how to maintain the traceability matrix. + +**Should have provenance**: Special case — self-referential. `meta/knowledge` already maps to ISO 9001 7.5 and NIST PL-2, which is appropriate because the way system itself is a documented information management process. + +**Current ways in this category**: +- `meta/knowledge` — how ways work, how to author them +- `meta/skills` — how skills work +- `meta/introspection` — session reflection and learning capture + +**Related skills**: `governance-cite` — the query interface into provenance data (a Layer 1 skill that reads Layer 3 metadata) + +### Type 4: Plumbing (Operational, Not Governance) + +Session lifecycle management. Not governance, not experience, just Claude Code operational mechanics. + +**Should have provenance**: No — these don't participate in governance. + +**Current ways in this category**: +- `meta/memory` — MEMORY.md management +- `meta/subagents` — delegation guidance +- `meta/teams` — team coordination norms +- `meta/todos` — context-threshold task list enforcement +- `meta/tracking` — cross-session state + +## How Ways Work as a System + +Ways are a trigger-dispatched, session-gated, context-injection system with a multi-modal retrieval function. They resemble RAG but differ in a fundamental way: + +| | Traditional RAG | Ways | +|---|---|---| +| Corpus | Documents | ~30 guidance files + dynamic macro outputs | +| Trigger | Query embedding similarity | Multi-modal matching (regex, semantic, model, state) | +| Injection | Into LLM context | Into Claude's context via hooks | +| Frequency | Every query | Once per session (session-gating) | +| Purpose | Compensate for the model NOT KNOWING | ACTIVATE what the model ALREADY KNOWS | + +The critical insight: way content is an activation signal, not a knowledge injection. Claude — and any sufficiently large LLM — has read NIST 800-53, OWASP, ISO 27001, IEEE standards, and other regulatory corpora during training. That knowledge exists in the latent space but isn't precisely retrievable on demand. Ways carry just enough ablated context to prime that existing deep knowledge into active use. This is why ways should be concise — you need "conventional commits, type prefix, atomic changes, rationale in body," and the model's training fills in the depth. More context isn't better; the right context is better. + +The provenance layer serves a different audience entirely. The way content primes Claude (the agent). The provenance block (stripped before injection, zero context cost) serves the human — via `governance-cite` — who asks "prove it." Two retrieval paths for two consumers from one source file. + +### Multi-Modal Retrieval Function + +Ways use four matching strategies instead of embeddings: + +1. **Regex** (default, fast): pattern matching against prompts, commands, file paths +2. **Semantic** (BM25): Term-frequency scoring with IDF weighting — no embeddings, no infrastructure dependency +4. **State triggers**: session conditions (context threshold, file existence, session start) + +### Session-Gating + +Each (way, session) pair has a marker at `/tmp/.claude-way-{domain}-{way}-{session_id}`. Once a way fires, the marker prevents it from firing again until session restart or compaction. This is fundamentally different from skills (always available) and traditional RAG (retrieves every query). + +Exception: context-threshold triggers bypass markers and repeat every prompt until a task list is created — enforcement, not education. + +### Scope Filtering + +Ways differentiate between agent (main session), teammate (team member), and subagent (quick delegate) contexts. This prevents, for example, three teammates simultaneously writing MEMORY.md or subagents receiving delegation guidance about delegation. + +## The Complementary Relationship + +Skills and ways complement at the edges and overlap in the middle. + +### Where the distinction is sharp + +**Ways can do, skills cannot**: Fire before `git commit` runs; fire when editing `.env` files; fire at context threshold; fire once then stay silent; differentiate scopes; carry governance provenance. + +**Skills can do, ways cannot**: User-invocable slash commands; restrict tools; fork into subagent contexts; distribute via plugins; always-in-context description for auto-discovery. + +### Where they overlap + +Both can inject contextual guidance when relevant. Both support dynamic context (macros vs `!`command``). Both do semantic matching (BM25 vs description-based). Pure reference-content ways with semantic matching and no macro, no governance, no scope filtering are functionally similar to `user-invocable: false` skills. + +### Why the overlap doesn't invalidate ways + +The overlap zone is "inject contextual guidance." But the ways that justify the system are the ones skills CAN'T replicate: +- PreToolUse-triggered guidance (commit formatting on `git commit`) +- State-triggered enforcement (context-threshold nag) +- Governance-traced policy (provenance blocks with NIST/OWASP/ISO mappings) +- Scope-filtered team coordination +- Macro-based dynamic context (repo health checks, file scanning) + +### First-line formation, not proof + +*(Reframed per ADR-200.)* A way carrying a claim knows *which control shape* it steers toward — but citing that claim is an assertion, not proof. In audit terms this is **first-line** work (the Three Lines Model): shaping how risk is managed *in the doing of the work*, distinct from the third-line assurance that independently proves conformance. + +- Without claims: guidance is followed because the model was trained on good practice. +- With claims: guidance also *states* the control shape it intends — so a later **finding** can test whether the work actually took that shape. + +The claim layer makes intended alignment explicit and assessable. It does not, by itself, turn assertion into evidence — that is precisely the correction ADR-200 exists to make. + +## The Governance Control Surface + +The user controls their governance posture through `ways.json`: + +```json +{ + "disabled": ["itops", "experimental"] +} +``` + +Enabling/disabling domains determines how much governance the agent carries. This is the control surface — not a one-size-fits-all system, but a configurable governance posture that the framework user chooses. + +Ways that are governance-relevant (Type 1) carry provenance. Ways that are experience-derived (Type 2) don't. The presence or absence of provenance in a way file indicates which type it is. No separate classification system needed. + +## Decision + +### 1. Skills and ways are complementary layers of a professional practice stack + +Skills = capability (Layer 1), Ways = policy (Layer 2), Provenance = institutional memory (Layer 3). They are not competing systems — they serve different layers. + +### 2. Keep governance provenance commingled in way files (Option A) + +The provenance travels with the guidance it justifies. One file, one truth. The `governance.sh` matrix provides the cross-cutting auditor's view derived from the source of truth. No separate governance overlay files. + +### 3. Recognize three types of ways (plus plumbing) + +- **Type 1 (Kitchen Poster)**: Compiled governance — SHOULD have provenance +- **Type 2 (Shift Lead Wisdom)**: Experience-derived — should NOT have provenance +- **Type 3 (HQ Policy Manual)**: Meta-governance — special case (self-referential) +- **Plumbing**: Operational — doesn't participate in governance + +### 4. Fill the provenance gap on Type 1 ways + +All active Type 1 ways now carry provenance. The governance matrix covers 12 ways with 37 control claims and 93 justifications across NIST, OWASP, ISO, SOC 2, CIS, and IEEE frameworks. Provenance authoring is a metadata task — the way body (practitioner guidance) already existed; the provenance block traces it back to the controls it implements. + +### 5. Way content is an activation cue, not knowledge injection + +Ways should be concise — just enough ablated context to activate Claude's latent training knowledge. The provenance serves humans via governance-cite. Two consumers, two paths, one source file. + +### 6. Governance is domain-agnostic + +The way/provenance architecture works for any governance domain (software engineering, financial compliance, data privacy, operational safety). The current implementation covers software engineering controls. Future domains (enabled via ways.json) can carry their own provenance to their own regulatory bodies. + +## Consequences + +### Positive +- Clear architectural rationale for why both skills and ways exist +- Framework for deciding which ways need provenance (Type 1) vs which don't (Type 2/3/plumbing) +- Establishes ways as codified policy, not just "contextual guidance" +- The governance-cite skill becomes more valuable as provenance coverage expands +- Domain-agnostic design allows governance expansion without architectural changes + +### Negative +- Provenance authoring requires domain expertise (control knowledge + implementation knowledge) — the author must be able to compile in both directions +- More provenance = more to maintain when controls or guidance change +- The compiler metaphor cuts both ways: a good author compiles governance into effective activation cues; a bad author either writes "comply with CM-3" (too abstract, no activation) or writes guidance that activates the wrong latent knowledge. The way system is only as good as its authors. + +### Neutral +- Pure reference-content ways (Type 2, semantic match, no macro) could migrate to skills/rules over time as Anthropic adds features, but this isn't urgent +- The `once` field in skill/agent hook frontmatter moves Anthropic incrementally toward session-gating, narrowing one differentiator +- itops domain remains disabled; its ways would follow the same type classification when enabled + +## Alternatives Considered + +### Organize ways by governance body (Option B) +Rejected: directory structure should match the user's mental model ("what work am I doing") not regulatory taxonomy ("what regulation am I following"). Nobody enables governance by body. + +### Separate governance overlay files (Option C) +Rejected: provenance must travel with the guidance it justifies. Separated mappings drift and become stale. One file, one truth. + +### Replace ways with official skills/rules +Rejected: skills cannot fire on tool use, cannot session-gate, cannot carry provenance, cannot scope-filter. The unique value of ways is precisely what skills cannot do. + +### Replace ways with raw hooks +Possible but rejected: ways provide an abstraction (session-gating, multi-mode matching, macros, governance provenance, scope filtering, domain organization) that would need to be rebuilt in every hook script. The abstraction layer is the value. + +## Traceability: a live query, not a frozen table + +*(Reframed per ADR-200.)* This ADR originally carried a point-in-time table marking each +way's provenance as "Complete." That is retired on two counts. Coverage is a **live +query** now (`ways governance`, moving to the `ways-audit` binary) — not a snapshot +frozen into a decision record. And under ADR-200 a way with a hand-authored mapping +carries a **claim**, not a completed control; "Complete" was the +coverage-equals-compliance overclaim. A claim is complete only once a session-derived +**finding** substantiates it. Run the tool for current coverage, and read any count it +reports as *claims made*, not *conformance achieved*. + +## References + +- **Kahneman, D.** (2011). *Thinking, Fast and Slow.* System 1 (fast/intuitive) and System 2 (slow/deliberative) as models of individual cognition. The ways architecture extends this: System 1 maps to Claude's latent training patterns; System 2 to deliberative reasoning when prompted. +- **Beer, S.** (1972). *Brain of the Firm.* The Viable System Model (VSM) describes how organizations maintain viability through recursive system layers including policy (System 3) and identity/ethos (System 3*). The three-layer stack in this ADR — skills (capability), ways (policy), provenance (institutional memory) — parallels VSM's separation of operational, regulatory, and normative functions. System 3 corresponds to institutional governance codified in standards bodies; System 3* corresponds to the ways+provenance mechanism that bridges institutional cognition into an individual agent's decision process. +- **NIST SP 800-53 Rev. 5** — Security and privacy controls referenced throughout provenance blocks +- **OWASP Top 10 (2021)** — Application security risks referenced in security, errors, and dependency provenance +- **ISO/IEC 27001:2022** — Information security management controls referenced in commits, ADR, and GitHub provenance +- **ISO/IEC 25010:2011** — Software quality characteristics referenced in quality and testing provenance +- **SOC 2 (AICPA)** — Trust services criteria referenced in change management provenance +- **CIS Controls v8** — Security implementation guidance referenced in config and security provenance +- **IEEE 730/829** — Software quality assurance and test documentation referenced in quality and testing provenance diff --git a/docs/architecture/practice/ADR-100-ways-scaffolding-wizard.md b/docs/architecture/practice/ADR-100-ways-scaffolding-wizard.md new file mode 100644 index 00000000..4c3c2b33 --- /dev/null +++ b/docs/architecture/practice/ADR-100-ways-scaffolding-wizard.md @@ -0,0 +1,112 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: authoring +basis: + - evidence: 'cold-start problem: creating a project-local way needs knowledge spread across extending.md, matching.md and authoring/way.md, with no guided path from intent to working way; /ways only listed fired ways' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-02-17 +deciders: + - aaronsb + - claude +related: + - ADR-013 + - ADR-014 +imported: + from: docs/architecture/system/ADR-100-ways-scaffolding-wizard.md + format: v0 + status: Accepted +--- + +# ADR-100: Ways Scaffolding Wizard + +## Context + +The ways system has a cold-start problem. Creating a project-local way requires knowledge spread across multiple documents (extending.md, matching.md, authoring/way.md) and familiarity with directory conventions, YAML frontmatter, matching modes, and vocabulary design. This knowledge is well-documented but not discoverable at the moment a human thinks "I want Claude to do X differently in my project." + +The current `/ways` skill shows which ways have fired in the current session — mildly interesting for debugging but not actionable. The `/test-way` skill handles vocabulary tuning and validation but assumes ways already exist. There is no guided path from intent ("our API uses GraphQL") to working way. + +Meanwhile, the broader Claude Code ecosystem has developed patterns for agent steering (CLAUDE.md files, PROMPT.md, hooks) but these are either monolithic (everything in one file) or require manual setup. The ways system solves the architecture problem (event-driven, contextual injection) but lacks an entry point for humans who haven't read the documentation. + +## Decision + +Repurpose `/ways` from a session diagnostic into a scaffolding wizard that interviews the human and creates project-local ways. + +### Design + +The wizard is a Claude Code skill (`commands/ways.md`) that directs Claude through a conversational flow: + +1. **Ground**: Read `docs/hooks-and-ways/matching.md` and `extending.md` to load the decision framework before engaging the human. This is explicit in the skill prompt — the agent needs the full matching mode landscape before it can recommend one. + +2. **Detect**: Check project state — does `.claude/ways/` exist? Are there existing project-local ways? Show what's already there. + +3. **Interview**: Use `AskUserQuestion` to gather intent conversationally: + - "What should Claude know or do differently in this project?" (plain language) + - Follow-up questions adapted to the answer — trigger timing, scope, specificity + - Recommend matching mode with a one-sentence explanation of why + +4. **Scaffold**: Create the directory and `way.md` with correct frontmatter. The body is the human's intent compressed to directive form, following the voice guidance from extending.md (collaborative, includes the why, writes for the innie). + +5. **Validate**: Lint the new way. If semantic, score it against sample prompts from the conversation. Show the result. + +6. **Handoff**: Point to `/ways-tests` for ongoing tuning. Explain that the way will fire automatically in future sessions when the trigger conditions are met. + +### Session commitment + +Invoking `/ways` is an intentional act — the human has decided to build or revise project steering. This isn't a casual diagnostic; it's dedicating the session to way work. The context cost of reading matching.md and extending.md upfront is justified because: + +- The human expects informed recommendations, not guesswork +- Matching mode decisions are the first fork in the road — you can't interview well without understanding the options +- Revision of existing ways (vocabulary tuning, trigger changes, scope adjustments) requires the same foundational knowledge as creation + +This framing also means the wizard handles both creation and revision. A human invoking `/ways` on a project with existing ways should be able to say "the API way isn't firing when I talk about endpoints" and get diagnostic help, not just a scaffold for a new way. + +### Interaction with the ways system + +The wizard leverages the existing trigger infrastructure recursively: + +- Talking about ways fires `meta/knowledge` — system overview enters context +- Editing a way.md fires `meta/knowledge/authoring` — frontmatter spec enters context +- Discussing vocabulary fires `meta/knowledge/optimization` — tuning workflow enters context + +The skill is the ignition; the ways system is the engine that keeps running as the conversation deepens. The skill prompt doesn't need to duplicate what the ways already provide — it just needs to start the conversation in the right neighborhood. + +### Scope + +- **Project-local ways** are the primary target — that's the user-facing use case +- **Global ways** remain an advanced/maintainer concern — the wizard can mention them but doesn't need to scaffold them +- The existing "which ways fired" diagnostic moves to a subcommand or gets folded into the wizard's detect phase + +## Consequences + +### Positive + +- Humans can create project-local ways without reading documentation first +- The matching mode decision (regex vs semantic vs state) is guided rather than discovered +- The wizard naturally produces ways that follow conventions (correct directory structure, valid frontmatter, appropriate voice) +- Recursive trigger behavior means the wizard improves its own context as it works — each step loads more relevant guidance +- The `/ways` namespace becomes useful instead of decorative + +### Negative + +- The skill prompt needs to be rich enough to guide Claude but not so prescriptive that it scripts every branch — finding this balance requires iteration +- Wizard-created ways may need vocabulary tuning that the wizard can start but the human must finish (the discrimination judgment is inherently human) +- `/test-way` is renamed to `/ways-tests` to share the `/ways` root namespace (tab-completion discoverability), but remains a separate skill — creation/revision vs precision tuning are different tasks at different expertise levels + +### Neutral + +- The old "show fired ways" behavior could become `/ways status` or be dropped — low value either way +- This establishes a pattern for other scaffolding wizards (skills, hooks, ADR domains) if the approach works +- The wizard's grounding step (reading matching.md) means it consumes context tokens on the documentation — acceptable given the task is inherently documentation-heavy + +## Alternatives Considered + +- **Management hub with subcommands** (`/ways list`, `/ways new`, `/ways health`, `/ways test`): More comprehensive but spreads the skill thin. The cold-start problem is the acute pain; management features can be added later. A Swiss Army knife is less useful than a sharp blade when you need to cut one thing. + +- **Dashboard focused on project state** (`/ways` shows coverage, gaps, health for this repo): Useful for maintainers but doesn't solve the creation problem. You can't show health metrics for ways that don't exist yet. + +- **Template gallery** (pick from pre-built way templates for common patterns): Lower friction than a wizard but less adaptive. Templates assume the human's intent fits a category; the wizard discovers it through conversation. Templates could be a future enhancement within the wizard flow. diff --git a/docs/architecture/practice/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md b/docs/architecture/practice/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md new file mode 100644 index 00000000..6914b20e --- /dev/null +++ b/docs/architecture/practice/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md @@ -0,0 +1,174 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - authoring + - matching +basis: + - evidence: project-local ways got BM25 fallback matching (91%) while global ways got embedding matching (98%) + - evidence: embedding 58+ ways takes ~2 seconds, too slow to redo on every session start + - precedent: ADR-108 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-23 +deciders: + - aaronsb + - claude +related: + - ADR-108 + - ADR-107 + - ADR-105 + - ADR-111 + - ADR-125 +imported: + from: docs/architecture/system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md + format: v0 + status: Accepted +--- + +# ADR-109: Project-Scope Way Embedding with Manifest-Based Staleness Detection + +## Context + +ADR-108 shipped embedding-based way matching using all-MiniLM-L6-v2. The corpus currently contains only global ways (`~/.claude/hooks/ways/`). But Claude Code's way system supports project-local ways at `/path/to/project/.claude/ways/`, and these participate in matching via BM25 and NCD — the embedding engine should cover them too. + +Projects are tracked by Claude Code in `~/.claude/projects/<encoded-path>/`. Each may have its own way tree at the project root's `.claude/ways/`. These ways follow the same n-deep progressive disclosure structure as global ways (ADR-105) and may change significantly between sessions as projects evolve. + +The gap: a user working in a project with custom ways gets BM25 fallback matching for those ways while global ways get embedding-quality matching. This creates an inconsistent experience — the same prompt triggers different matching quality depending on whether the way is global or project-local. + +Additionally, runtime artifacts need clear separation from source. The corpus is a cache (generated from way files), not source data. ADR-108 moved it to `${XDG_CACHE_HOME:-~/.cache}/claude-ways/user/`. All runtime artifacts — binary, model, corpus, and now the embedding manifest — live outside `~/.claude/`. + +## Decision + +### Unified Corpus Generation + +`generate-corpus.sh` scans both global and project-local ways into a single corpus: + +1. Embed global ways (`~/.claude/hooks/ways/`) — same as today +2. Scan `~/.claude/projects/` for decoded project paths +3. For each project with `.claude/ways/`: + a. Lint the ways (must pass to proceed) + b. Check the inclusion marker at `<project>/.claude/.ways-embed` + c. Embed included ways into the same corpus + +Global ways keep their current format (`softwaredev/code/testing`). Project ways are namespaced by a **project key derived from the project's real path**: `<project-key>/<way-tree-path>`. + +> **Implementation note (supersedes the original "encoded project path" wording).** +> The corpus id prefix is *not* the `~/.claude/projects/<encoded-path>` directory name. That encoding (`/`→`-`, and `:`→`-` on Windows) is one-way and lossy, and — critically — the matcher (`ways scan`) only knows the *real* project directory (from `--project` / `CLAUDE_PROJECT_DIR`), never the encoded form. If the corpus prefixed ids with the encoded dir name, the matcher's lookup key would never equal the corpus id and project ways would never match semantically on any platform. +> +> Instead, both sides compute one shared key, `encode_project_key(real_path)` (`ways-cli/src/util.rs`): canonicalize the path, strip Windows verbatim prefixes (`\\?\`), lowercase on Windows, then flatten every separator and `:` to `-`. The result is a flat token, so the single `/` in `<project-key>/<way-tree-path>` is the namespace boundary. Way-tree path segments are joined with `/` on all platforms (`path_to_id`) so ids never leak `\` on Windows. +> +> Corpus generation embeds the **current** project straight from `CLAUDE_PROJECT_DIR` (Windows-safe, no decode) and additionally scans `~/.claude/projects/*` for other projects, deriving each one's key from its *resolved real path* (via `sessions-index.json`, falling back to greedy filesystem decode) — never the raw encoded dir name. The matcher prefixes its project candidates with `encode_project_key(--project dir)`, so the keys agree by construction. + +Way trees are n-deep (progressive disclosure, ADR-105), so IDs reflect the full tree path — no fixed domain/way depth assumption. + +### Inclusion Markers + +A file at `<project>/.claude/.ways-embed` controls whether that project's ways are embedded: + +| State | Behavior | +|-------|----------| +| No marker, valid ways found | Create marker = `include`, embed | +| No marker, no ways | Skip silently | +| Marker = `include` | Embed | +| Marker = `disinclude` | Skip, warn if valid ways exist | +| Ways fail lint | Never embed, regardless of marker | + +Writing to `<project>/.claude/` is consistent with Claude Code's own behavior — it already writes memory, `settings.local.json`, and permissions there. + +### Manifest-Based Staleness Detection + +A manifest at `${XDG_CACHE_HOME}/claude-ways/user/embed-manifest.json` records what was embedded using content hashes — not timestamps: + +```json +{ + "global_hash": "a1b2c3...", + "global_count": 58, + "projects": { + "-home-aaron-myproject": { + "path": "/home/aaron/myproject", + "ways_hash": "d4e5f6...", + "ways_count": 3 + } + } +} +``` + +Hashes are computed from the way files themselves (e.g., `git ls-files -s .claude/ways/ | sha256sum` for tracked files, plus stat of untracked way files). This is content-addressed staleness — immune to clock skew, catches uncommitted edits, and doesn't care *when* things changed, only *whether* they changed. + +### Session-Start Staleness Check + +At session start (cheap, no embedding work): + +1. Read the manifest +2. Compute current content hash for global ways +3. For each project in `~/.claude/projects/`, compute ways hash if `.claude/ways/` exists +4. Compare hashes against manifest entries +5. If any hash differs, or a project exists that isn't in the manifest, trigger regen + +Hash computation is one `ls-files | sha256sum` per scope — cheaper than git log, and works for uncommitted changes too. At worst it runs every session start. User-scope ways are relatively static; project-scope ways evolve significantly, making this check worthwhile. + +If the manifest is missing or corrupted, treat it as "everything stale" and do a full regen. The manifest is a cache of caches — losing it just costs one regen cycle. + +### Staleness is Harmless + +If a project is deleted or its ways removed, stale embeddings remain in the corpus but never match — no way file backs them at runtime. Next regen naturally drops them. No eager purge, no reconciliation. Corpus regen is append-from-scan, not diff-and-reconcile. + +## Consequences + +### Positive + +- Project-local ways get embedding-quality matching (98%) instead of BM25 fallback (91%) +- Single corpus, single embedding space — no per-project model overhead +- Content-addressed staleness — immune to clock skew, no date comparison, catches uncommitted edits +- Lint-gating prevents broken or untrusted ways from entering the embedding +- Inclusion markers give projects control without global configuration +- Manifest enables incremental regen — only re-embed when something changed + +### Negative + +- Session start gains a filesystem scan across `~/.claude/projects/` (should be fast — directory listing + content hash per scope) +- Regen cost grows linearly with project count (each project's ways are additional embedding work, ~20ms per way) +- Manifest is another file to manage in XDG cache (but it's a cache — loss just triggers full regen) +- Inclusion markers in project `.claude/` directories are a write outside the framework's own tree + +### Neutral + +- BM25 and NCD fallback paths already handle project-local ways — this extends existing behavior to the embedding tier +- The matcher **does** need a change for matching: its project candidates must be looked up under the same `<project-key>/<way-tree-path>` namespace the corpus writes (see implementation note above). The original ADR assumed corpus generation alone sufficed; that was incorrect — the two sides must share one key derivation. Pattern/keyword matching was unaffected (it reads way files directly), which is why the gap surfaced only as silent semantic misses. +- Progressive disclosure (ADR-105) works the same way for project ways — parent/child relationships, sibling coverage, depth tracking all apply. These operate on the bare way id (session markers, show, parent-boost); only the embedding lookup uses the namespaced `<project-key>/...` id. +- `embed-status` CLI tool needs updating to report: manifest contents, per-project inclusion state (included/disincluded/not found), staleness per scope, and project paths. +- Claude Code's project path encoding (`/` → `-`, and `:` → `-` on Windows) is a one-way function. Paths with hyphens in directory names can't be reverse-decoded. The manifest records real paths at embed time and the namespace key is derived from the real path, not the encoded dir name — so decoding is never on the matching critical path. `~/.claude/projects/*` scanning uses `sessions-index.json` (real path) first and greedy decode only as a best-effort fallback for *other* projects; the current project always comes from `CLAUDE_PROJECT_DIR`. + +### Footgun guard (added during implementation) + +`ways corpus --ways-dir <dir>` historically wrote to and re-embedded the canonical user corpus, silently replacing all global + project ways with just `<dir>`'s ways. Corpus generation now accepts `--output <dir>` to direct artifacts to an isolated location, and warns when `--ways-dir` is used without `--output` against the canonical corpus. + +## Alternatives Considered + +### Per-project corpus files + +Generate a separate corpus per project, load multiple at match time. + +Rejected: multiplies file I/O, complicates the scanner, and breaks the "one embedding space" property that makes cosine similarity scores comparable across all ways. + +### Embed on every session start unconditionally + +Skip the manifest, just regenerate every time. + +Rejected: embedding 58+ ways takes ~2 seconds. With multiple projects, this could grow to 5-10 seconds — noticeable on every session start. A content-hash check is effectively free by comparison and only triggers regen when something actually changed. + +### Store corpus in project `.claude/` alongside ways + +Each project manages its own embedding cache. + +Rejected: violates the XDG separation principle. Runtime artifacts belong in `~/.cache/`, not in project trees. Also breaks unified matching. + +### No project-scope embedding (status quo) + +Keep embedding for global ways only, rely on BM25 for project ways. + +Rejected: creates an inconsistent matching experience. The whole point of ADR-108 was that BM25 can't distinguish meaning — that limitation applies equally to project ways. diff --git a/docs/architecture/practice/ADR-110-way-file-separation-and-graph-compatible-structure.md b/docs/architecture/practice/ADR-110-way-file-separation-and-graph-compatible-structure.md new file mode 100644 index 00000000..c49484b9 --- /dev/null +++ b/docs/architecture/practice/ADR-110-way-file-separation-and-graph-compatible-structure.md @@ -0,0 +1,289 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - authoring + - governance +basis: + - evidence: provenance adds 10-15 lines of nested YAML to frontmatter that only the governance pipeline reads; 84 way files all named way.md are hard to navigate + - evidence: 'Anthropic''s harness design research (2025): every harness component encodes an assumption about what the model can''t do' + - standard: 'Unix man pages, man(7) and mandb(8): one-line metadata, conventional body sections, derived indexes' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-29 +deciders: + - aaronsb + - claude +related: + - ADR-107 + - ADR-108 + - ADR-005 + - ADR-111 + - ADR-151 + - ADR-200 +imported: + from: docs/architecture/system/ADR-110-way-file-separation-and-graph-compatible-structure.md + format: v0 + status: Accepted +--- + +# ADR-110: Way File Separation and Graph-Compatible Structure + +> **Accepted as implemented; reconciled with later ADRs.** The `provenance.yaml` sidecar +> defined here is the storage for what **ADR-200** reframes as a compliance **claim** (a +> control-*design* assertion, not evidence). Its consumers named below — `governance.sh` +> and `provenance-scan.py` — predate the single-binary consolidation (**ADR-111**); the +> sidecar is now read by `ways governance`, moving to the `ways-audit` binary +> (**ADR-151**). Read those tool names accordingly. + +## Context + +A way directory currently mixes four distinct concerns into one or two files: + +1. **Matching frontmatter** — trigger configuration consumed by `way-match`, `way-embed`, and the scanner scripts. Fields: `description`, `vocabulary`, `threshold`, `embed_threshold`, `pattern`, `commands`, `files`, `scope`. +2. **Guidance body** — markdown prose injected into Claude's context by `show-way.sh`. This is the way's actual content. +3. **Macros** — executable shell scripts that generate dynamic content based on project state. Already separated as `macro.sh` (good precedent). +4. **Governance provenance** — control mappings, policy URIs, and justifications consumed exclusively by `governance.sh` and auditors. Currently embedded in way.md frontmatter as deeply nested YAML. + +The provenance block is the most problematic. A typical way's frontmatter is 6-8 lines of matching config, but provenance adds 10-15 lines of nested YAML that no other consumer reads. The matching pipeline ignores it. `show-way.sh` strips it before injection. Only the governance pipeline parses it. It's instrumentation wearing the costume of content. + +At 84 way files across 4 domains, the corpus is becoming difficult for humans to navigate. The directory tree provides domain grouping, but cross-references between ways are implicit (vocabulary overlap, shared domain ancestry). There's no way to see the relationships between ways without running the embedding engine or reading all 84 files. + +Meanwhile, the ways system is approaching a point where humans other than the original author need to participate — reading, understanding, tuning, and authoring ways. The current structure requires understanding the frontmatter schema, the scoring system, and the governance model before making any edit. That's too high a barrier for someone who just wants to tune the prose in a way or understand how ways relate to each other. + +### What prompted this + +Three observations converged: + +1. **Anthropic's harness design research** (2025) found that "every component in a harness encodes an assumption about what the model can't do on its own." Ways are exactly this — each one encodes an assumption. As models improve, some assumptions become stale. There's no mechanism to signal which assumptions a way is making or how firm they are. + +2. **Man page architecture** has solved the "many structured documents, navigable by humans and machines" problem for 40 years with a clear separation: minimal metadata (one line), conventional body sections (`NAME`, `SEE ALSO`), derived indexes (`mandb`/`whatis`), and locale as a directory concern. Ways could follow the same pattern. + +3. **Graph visualization tools** (Logseq, Obsidian, Cytoscape, or plain JSONL export) can make a corpus of 84+ documents navigable — but only if the documents carry their relationships in a format that's both human-readable and machine-extractable. The current frontmatter doesn't express relationships between ways. + +## Decision + +### 1. Extract provenance to a sidecar file + +Move governance data from way frontmatter to a separate `provenance.yaml` in the same directory: + +``` +hooks/ways/softwaredev/code/quality/ + quality.md # matching frontmatter + guidance body ({dirname}.md) + macro.sh # dynamic content (already separated) + provenance.yaml # governance instrumentation (newly separated) +``` + +`provenance.yaml` contains exactly what's currently under the `provenance:` key in frontmatter: + +```yaml +policy: + - uri: governance/policies/code-lifecycle.md + type: governance-doc +controls: + - id: ISO/IEC 25010:2011 (Maintainability - Analyzability, Modifiability) + justifications: + - File length thresholds enforce analyzability through size limits + - Nesting depth limit maintains modifiability by controlling complexity +``` + +The governance CLI (`ways governance`, later the `ways-audit` binary — ADR-111/151) reads `provenance.yaml` files instead of parsing way frontmatter. The pipeline never needs to parse markdown again. + +Way files that have no governance mapping simply don't have a `provenance.yaml`. Absence is the default, not an empty block. + +### 2. Add conventional body sections for graph data + +Instead of adding more frontmatter fields, use conventional markdown sections in the way body — following the man page pattern where `SEE ALSO` is a body section, not metadata: + +**Epistemic stance** as an HTML comment (invisible when rendered, extractable by tooling): + +```markdown +<!-- epistemic: heuristic --> +``` + +Four values, each declaring what kind of claim the way is making: + +| Value | Meaning | Durability | +|-------|---------|------------| +| `premise` | Reasoning to incorporate, not a rule to follow | Durable — tied to the cognitive model | +| `convention` | How we do it here — functional choice, not universal truth | Stable but overridable per-project | +| `heuristic` | Works most of the time — override when context demands | May deprecate as models improve | +| `constraint` | Hard boundary (legal, security, compliance) | Durable — tied to external requirements | + +**See Also** as a standard markdown section with typed references following the `name(domain)` convention from man pages: + +```markdown +## See Also + +- code/testing(softwaredev) — test coverage and fixtures +- code/errors(softwaredev) — error handling boundaries +- trust(meta) — why quality matters beyond the code +``` + +These are readable as plain text, parseable by any tool that knows the convention, and ignorable by any tool that doesn't. + +### 3. Keep frontmatter matching-only + +After provenance extraction, `{name}.md` frontmatter contains only fields the matching pipeline reads: + +```yaml +--- +description: code quality, refactoring, SOLID principles +vocabulary: refactor quality solid principle decompose +threshold: 2.0 +pattern: solid.?principle|refactor +scope: agent, subagent +macro: append +--- +``` + +Maximum 8-10 lines. No nesting deeper than one level. A human can read and edit this without understanding YAML block scalars or nested arrays. + +The `scan_exclude` field stays in frontmatter because `macro.sh` reads it — it's part of the macro's configuration, not governance. + +### 4. Build derived artifacts, not authoritative indexes + +Following the `mandb` pattern, a generator script reads all way files and produces derived artifacts. The source of truth is always the individual files. The artifacts are build outputs, regenerated on demand: + +| Artifact | Format | Consumer | Content | +|----------|--------|----------|---------| +| `ways-corpus.jsonl` | JSONL | `way-match`, `way-embed` | Matching fields (already exists per ADR-107) | +| `ways-graph.jsonl` | JSONL | Any graph tool, export scripts | Nodes (ways) + edges (See Also, sibling scores) | + +All artifacts are gitignored or explicitly regenerated. None are authoritative. The generator is idempotent — running it twice produces identical output. + +The `.vault/` directory previously described in this ADR (Logseq-compatible markdown with uniquely-named symlinks) is no longer needed. The file rename from `way.md` to `{dirname}.md` (see Section 7) gives every way file a unique name, making symlink indirection unnecessary. Any graph tool can consume the source files directly or use `ways-graph.jsonl`. + +The `ways-graph.jsonl` format is deliberately simple: + +```jsonl +{"id":"code/quality","domain":"softwaredev","epistemic":"heuristic","description":"..."} +{"id":"code/testing","domain":"softwaredev","epistemic":"convention","description":"..."} +{"source":"code/quality","target":"code/testing","type":"see_also"} +{"source":"code/quality","target":"code/errors","type":"see_also"} +{"source":"code/quality","target":"docs/standards","type":"sibling","weight":0.69} +``` + +Nodes and edges in the same stream. Importable by Cytoscape, D3, Logseq plugin, or a `jq` one-liner. The format doesn't assume any visualization tool. + +### 5. Sibling scoring as an authoring compass + +`way-embed` gains a `siblings` subcommand that scores way-vs-way across the full corpus: + +```bash +way-embed siblings code/quality +# code/quality <-> code/testing: 0.78 (expected, different concern) +# code/quality <-> code/errors: 0.71 (expected, adjacent) +# code/quality <-> docs/standards: 0.69 (vocabulary overlap?) +# code/quality <-> delivery/commits: 0.43 (distant, good) +``` + +This is an **authoring tool, not a matching tool**. It never runs during prompt evaluation. It runs when a human or agent is tuning vocabulary and wants to understand how a way relates to its neighbors. + +The sibling scores feed into `ways-graph.jsonl` as weighted edges. A graph visualization shows clusters (expected), outliers (review), and suspicious overlaps (vocabulary collision or candidate for `See Also`). + +The compass metaphor is deliberate: it points north, it doesn't steer. The author interprets the direction. An agent reading sibling scores understands the sparseness landscape — which ways are close, which are distant — and uses that to calibrate vocabulary when authoring or tuning. + +### 6. Tool agnosticism as a design constraint + +The way file format does not reference or optimize for any specific visualization tool. Logseq, Obsidian, Cytoscape, a plain text editor, or `grep` are all valid ways to interact with the corpus: + +- **vim/emacs**: Edit `{name}.md` directly. Frontmatter is short YAML. Body is markdown. See Also is readable text. +- **Logseq/Obsidian**: Open the ways tree directly. Every file has a unique name, so graph view works without a generated vault. Edit in place; regenerate graph JSONL if needed. +- **Cytoscape/D3**: Import `ways-graph.jsonl`. Visualize clusters and edge weights. +- **CLI**: `way-embed siblings`, `governance.sh --trace`, `lint-ways.sh`. Same data, terminal interface. +- **Claude**: Reads `{name}.md` as today. New sections are markdown it can interpret. Sibling scores inform vocabulary tuning. + +If a tool requires a format the source files don't provide, the answer is a generator that produces the format — not modifying the source files to accommodate the tool. + +### 7. File rename: `way.md` to `{dirname}.md` + +Way files are renamed from the generic `way.md` to `{dirname}.md` — the file takes its name from its parent directory. For example: + +``` +hooks/ways/softwaredev/code/quality/quality.md +hooks/ways/softwaredev/code/testing/testing.md +hooks/ways/softwaredev/delivery/commits/commits.md +hooks/ways/meta/trust/trust.md +``` + +Similarly, check files are renamed from `check.md` to `{dirname}.check.md` (e.g., `quality.check.md`). + +**Rationale:** Every way file being named `way.md` created several problems: + +1. **Graph tool compatibility.** Logseq, Obsidian, and similar tools identify documents by filename. With 84 files all named `way.md`, these tools see 84 identically-named nodes — useless for navigation or graph visualization. The original plan (Section 4) required a `.vault/` directory with uniquely-named symlinks to work around this. Unique filenames eliminate that workaround entirely. + +2. **Editor tab ambiguity.** Opening multiple ways in any editor shows tabs all labeled `way.md`. The developer must check the path to know which way they're editing. `quality.md` vs `testing.md` is immediately distinguishable. + +3. **Search result clarity.** `grep` and `ripgrep` output includes the filename. Results from `quality.md` are self-documenting; results from `way.md` require reading the full path. + +4. **Shell completion.** Tab-completing into a way directory and hitting tab again now shows a meaningful filename, not the generic `way.md` that every directory shares. + +The rename is a one-time migration. Scanner scripts (`check-prompt.sh`, `check-bash-pre.sh`, `check-file-pre.sh`) discover way files by glob pattern, updated from `way.md` to `*.md` with exclusions for known non-way files (`macro.sh`, `provenance.yaml`, `*.check.md`). The `lint-ways.sh` and corpus generator scripts are updated correspondingly. + +## Consequences + +### Positive + +- Way files become dramatically simpler. Frontmatter drops from 20+ lines to 8-10. +- Provenance is independently auditable without parsing markdown. +- Humans can author and tune ways without understanding governance mappings. +- Graph relationships become explicit and visible across the corpus. +- Epistemic stance makes the authority level of each way inspectable — both by humans choosing how firmly to follow guidance and by agents assessing which ways may need revision as models improve. +- Format is tool-agnostic. No vendor lock-in to any visualization tool. +- Sibling scoring gives authors (human and agent) a calibration tool for vocabulary sparseness. +- Unique filenames per way eliminate the need for a `.vault/` symlink directory and make the corpus directly navigable by graph tools, editors, and search. + +### Negative + +- Provenance extraction is a migration across 84 files. Each way with a `provenance:` block needs the block moved to `provenance.yaml` and the frontmatter cleaned. This is mechanical but must be validated by `governance.sh` and `lint-ways.sh` after migration. +- `governance.sh` and `provenance-scan.py` need to read `provenance.yaml` sidecar files instead of way frontmatter. The scan logic changes from "parse YAML frontmatter from markdown" to "read YAML file directly" — simpler, but a code change. +- `frontmatter-schema.yaml` needs updating: `provenance` moves from way schema to a separate provenance schema. +- Generator tooling is new code to write and maintain. +- `See Also` sections need to be added to existing ways. This is authoring work — each cross-reference should be intentional, not bulk-generated. +- The file rename requires updating all scanner scripts and any tooling that hardcodes `way.md`. This was a one-time migration but touched every scanner and the linter. + +### Neutral + +- `macro.sh` is unchanged. It was already separated correctly. +- The matching pipeline (`way-match`, `way-embed`, scanner scripts) requires a glob pattern update for file discovery but reads the same frontmatter fields. +- `show-way.sh` is unchanged except for removing the provenance-stripping step (no longer needed). +- `lint-ways.sh` gains new checks: validate `provenance.yaml` schema, validate `See Also` references point to existing ways, validate `epistemic` value if present. +- ADR-107's locale convention (`{name}-es.md`, `{name}-fr.md`) applies only to way files. `macro.sh` and `provenance.yaml` are shared across locales — macros execute the same logic regardless of language, and controls don't change by locale. + +## Interaction with ADR-107 + +ADR-107 Phase 3 defines locale support with `{name}-{lang}.md` files (e.g., `quality-es.md`, `quality-fr.md`). This ADR clarifies which files are locale-specific and which are shared: + +| File | Per-locale? | Rationale | +|------|-------------|-----------| +| `{name}.md` / `{name}-{lang}.md` | Yes | Matching vocabulary and guidance body vary by language | +| `macro.sh` | No | Detection logic is language-independent | +| `provenance.yaml` | No | Controls and policies don't vary by language | + +This simplifies ADR-107's scope: locale support only touches way files, not the entire directory. + +## Alternatives Considered + +- **Provenance in the way body as a markdown section**: Would keep everything in one file. Rejected — provenance is deeply structured YAML (nested controls with justification arrays). Representing this as markdown would be awkward and harder to parse than a YAML file. The governance pipeline needs structured data, not prose. + +- **A single domain-level provenance manifest** (e.g., `softwaredev/provenance.yaml` mapping all ways to controls): Would reduce file count. Rejected — it centralizes what should be co-located. When you're editing a way, its governance mapping should be in the same directory, not in a parent file you have to cross-reference. The sidecar pattern preserves locality. + +- **Wikilinks in way.md frontmatter** (`see_also: ["[[code-testing]]"]`): Would make graph edges machine-readable from frontmatter. Rejected — this is tool-specific syntax (Logseq/Obsidian) in what should be a tool-agnostic format. The `## See Also` body section with `name(domain)` references is readable as plain text and parseable by a generator into any format. + +- **Embedding graph data in frontmatter** (`epistemic`, `see_also`, `model_floor` as YAML fields): Would be machine-readable without body parsing. Rejected — every field added to frontmatter increases the barrier to authoring. The man page lesson is that conventional body sections scale better than metadata fields. Frontmatter should contain only what machines need for matching; everything else is body. + +- **Knowledge graph MCP as the graph backend**: The project already has `mcp__knowledge-graph__*`. Ways could be ingested as nodes. Rejected for this purpose — adds a runtime dependency for what should be a static, file-based operation. The knowledge graph is appropriate for the cognitive framework paper's concepts, not for the way corpus which is already file-based and should stay that way. + +## References + +- ADR-107: Way-Match Corpus, Batch Mode, and Locale Support +- ADR-108: Embedding-Based Way Matching with all-MiniLM-L6-v2 +- ADR-005: Governance Traceability for Ways +- Anthropic Engineering: "Harness Design for Long-Running Application Development" (2025) +- `hooks/ways/frontmatter-schema.yaml`: Current field definitions +- Provenance consumer: `ways governance` (consolidated from the former `governance/governance.sh` per ADR-111; moving to the `ways-audit` binary per ADR-151) +- Unix man pages: `man(7)`, `mandb(8)` — prior art for structured document corpora diff --git a/docs/architecture/practice/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md b/docs/architecture/practice/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md new file mode 100644 index 00000000..ee3db28f --- /dev/null +++ b/docs/architecture/practice/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md @@ -0,0 +1,114 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - method + - install + - disclosure +basis: + - standard: 'Claude Code auto-memory behaviour (code.claude.com/docs/en/memory): MEMORY.md loads its first 200 lines or 25 KB at every session start' + - evidence: '2026-04-22: a memory entry pointed at a stale design note whose claim that Curve variants were unused would have led a session to delete live code' + - precedent: ADR-125 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-22 +deciders: + - aaronsb + - claude +related: + - ADR-125 +imported: + from: docs/architecture/system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md + format: v0 + status: Accepted +--- + +# ADR-128: Memory as repo-portable ways — seed routing over accumulated snapshots + +## Context + +Claude Code ships an official auto-memory feature (v2.1.59+, browsable via the `/memory` slash command). It stores `MEMORY.md` at `~/.claude/projects/<project-normalized>/memory/MEMORY.md`, loads the first 200 lines or 25 KB unconditionally at every session start, and encourages the session to review and extend the file. Topic files as `.md` siblings in the same directory load on demand via pointer lines in `MEMORY.md`. Reference: [Claude Code Memory docs](https://code.claude.com/docs/en/memory). + +Anthropic has also shipped, without public announcement, a **periodic memory-compaction cycle**. After roughly 24 hours + 5 sessions of accumulation, a 4-phase pass runs — orient, gather from session transcripts, consolidate, prune — rewriting `MEMORY.md` and its topic files to stay under the 200-line load window. Anthropic's internal/branded name for this is "Auto-Dream" (toggleable via `/memory`; manual trigger `/dream` on gradual rollout). Third-party reference: [claudefa.st on auto-dream](https://claudefa.st/blog/guide/mechanics/auto-dream). The mechanism is memory compaction — it makes accumulated memory tidier; it does not send consolidated knowledge anywhere other than back into memory. + +Observed across recent instances, the capture surface this produces is elaborate, persistent, and frictionless — Claude readily writes substantial entries without external prompting. The compaction cycle's quiet ship signals that Anthropic is doubling down on memory as a capture medium with self-organizing behavior. + +Three failure modes follow from accepting the default. + +**1. Staleness overrides reality.** `MEMORY.md` loads unconditionally into every session. A memory entry capturing yesterday's state silently outranks today's source of truth. 2026-04-22 provides a concrete case: a design note in this repo (`docs/design-notes/post-adr-126-simplification.md`) claimed certain `Curve` variants were "unused in the tree" — accurate for way frontmatter files but wrong for Rust code, where `ActionPotential` is `default_sensor_curve()` (`tools/sensor-trait/src/lib.rs:220`) and attend's engagement curve (`tools/attend/src/cmd/run.rs:177`). A memory entry pointed at the note. A session following the memory would have deleted live code. `grep` would have caught it in 10 seconds; the memory didn't. + +**2. Short-circuiting discipline.** Memory's low friction is not a feature — it is the problem. Every elaborate memory entry is an ADR, way, design note, GitHub issue, PR description, or commit message that *didn't get written* because the model satisfied its capture instinct cheaply. The friction enforced by those artifacts — naming, frontmatter, lint, review — exists precisely because that friction forces thinking memory would skip. + +**3. Per-instance cage.** Auto-memory is scoped to a Claude instance and the project-path normalization. It does not travel with the repo. Teammates, CI, other Claude runs on the same repo from a different path — none see it. The project's accumulated "memory" is invisible to everyone who wasn't the instance that wrote it. (This is orthogonal to `CLAUDE.md`, which is human-authored, repo-committed, and loaded separately and in full at session start. `CLAUDE.md` is not auto-memory.) + +These failure modes are *disproportionately acute* for projects with their own capture discipline (ADRs, ways, enforced PR review, commit-message conventions). For a bare Claude Code session with no such discipline, memory's competing capture surface is unobtrusive — a few scribbled notes, no downstream cost. For a repo that already has friction-enforcing artifacts, memory's low-friction shortcut actively erodes the discipline those artifacts encode. The memory system also exposes **no programmatic integration hooks** — no `memoryWrite` intercept, no API, no per-project steering beyond the binary `/memory` toggle. A harness-level response must therefore work by *observing and rewriting* `MEMORY.md` rather than by hooking memory events. This ADR's mechanism takes that constraint as given. + +## Decision + +**Redirect Claude Code's official auto-memory slot rather than fight it.** Seed `MEMORY.md` with routing guidance that treats project-scoped knowledge as belonging in repo artifacts (ways, ADRs, design notes, GitHub issues, PR descriptions, commit messages, and other friction-enforcing artifact classes a project may adopt). Auto-memory remains writable (we do not disable it or modify the feature), but the *first bytes Claude reads at session start* — within the official 200-line / 25 KB load window — frame memory as narrow: short cross-project user facts, nothing that has a better home. + +Mechanism: + +**a. Seed template** shipped with the ways tooling. Opens with the "Memory short-circuits discipline" thesis, names the friction-enforcing artifacts, includes an anti-rationalization table pre-empting common capture-skip patterns. + +**b. Frontmatter-bearing identification.** The seed begins with YAML frontmatter: + +```yaml +--- +seed: claude-code-memory +seed-version: 1 +--- +``` + +**c. Byte-equality integrity check against an embedded canonical.** The canonical seeded body is baked into the ways binary at build time via Rust's `include_str!`. The binary physically carries the canonical bytes, so verification is pure byte comparison — no hash, no runtime crypto dependency, no "chasing" a moving hash value. Scoping the compared region to the body between frontmatter and the `## User Context` heading means user-added entries below that heading are orthogonal to the integrity check and preserved across re-seeds. + +**d. SessionStart integration via `ways init`** runs on every Claude invocation. The existing `ways init` subcommand (already registered in `SessionStart` settings.json hooks) now also verifies the seed: + +- `MEMORY.md` missing → write current seed. +- Frontmatter parses + `seed == claude-code-memory` + `seed-version` matches + extracted body bytes equal `canonical_body()` → no-op. +- Any check fails → save unified diff of current content to `MEMORY.diff.YYYY-MM-DD.NNN.md` (serial increment), write fresh seed preserving everything from `## User Context` onward. If the `## User Context` marker is structurally missing, diff the whole file and rewrite in full. + +**e. Memory-seed files** are a new artifact class — frontmatter-bearing, byte-compared against an embedded canonical, managed by ways tooling, destination the harness's memory slot rather than the ways tree. The vocabulary matters: a future reader sees `seed: claude-code-memory` and has a word for what they're looking at. + +**f. Defense-in-depth via the memory way.** `hooks/ways/meta/memory/memory.md` carries the same routing table and anti-rationalization block. Surfaces when the way fires mid-session on "remember this" prompts, after the seed has already framed the session at startup. + +**g. Drift surfaces as session-time review via hook stdout.** When `ways init` detects drift and writes a diff file, it emits a review prompt to stdout. The `SessionStart` hook pipeline captures that stdout into the session's initial context, giving Claude an actionable triage instruction: apply the routing table (see point `f`) to each diff entry — convert repo-relevant content to ways/ADRs/design notes/issues, discard what doesn't warrant preservation, re-add only genuine cross-project user facts under `## User Context`. Compaction output is no longer a silent silo — it's a conversion queue the session actively triages. No separate state-trigger way is needed; the init hook's own stdout is the surfacing mechanism. + +Claude remains free to write memory — no permission changes, no harness modification. The intervention is purely framing: the memory attractor gets redirected toward the repo-portable artifacts the project already maintains, and any output that accumulates despite the framing gets actively reviewed and re-routed at the next session. + +## Consequences + +### Positive + +- **Repo-portability** — project knowledge captured as ways travels with `git clone`. CI sees it, teammates see it, other Claude runs see it. The per-instance cage is broken for everything that belongs in a repo artifact. +- **Preserved discipline** — ADR/way/issue/PR/commit friction is protected from memory's shortcut. Elaborate memory entries get reframed as "discipline bypassed" *before* they're written, not after. +- **Staleness detection** — ways and ADRs are lint-validated in CI; memory isn't. Routing project facts through ways means drift gets caught mechanically rather than on accidental human re-read. +- **Idempotent, structurally versioned seeding** — the hook is integrity-verified (byte-equality against the binary's embedded canonical) and carries `seed-version` in frontmatter. Ships at v1; template improvements are designed to ship as version bumps with diff preservation on migration (path present but unexercised until a v2 lands). +- **Auditable drift** — when a seed is edited (intentionally by the user, or by Claude mid-session), the unified diff is preserved with a serial-numbered filename. Review is a concrete artifact, not a memory lookup. +- **Compaction becomes a re-routing trigger** — the periodic memory-compaction cycle rewrites `MEMORY.md` between sessions; our byte-equality check detects it, the hook re-seeds + writes a diff, and emits a review prompt to stdout which SessionStart injects into Claude's context at the next session. Compaction output gets *converted* into repo artifacts (ways, ADRs, issues) or explicitly discarded — not silently re-consolidated back into memory. The harness's memory-tidying work becomes an input to our discipline, not a competitor. + +### Negative + +- **New-project onboarding cost** — every new project needs the hook installed. `project-init` should handle this automatically; existing projects need a one-time install step. +- **Non-preserving for seeded-portion edits** — if a user or Claude edits the seeded body (above `## User Context`), the next session's hook detects drift, diffs it, and rewrites. User edits to the seed itself are treated as drift, not customization. Intentional: opinionated framing is the whole point. +- **Harness coupling** — mechanism depends on the harness loading `MEMORY.md` directly at session start. Anthropic's in-flight memory-compaction work (see Context) is visibly ongoing; a future release replacing direct load with summarization or compacted injection would change the delivery path. The *framing* (project knowledge belongs in repo artifacts, not per-instance memory) survives any such change; only the force-feed mechanism is coupled. +- **Review cost per compaction cycle** — when compaction runs, Claude is prompted at next session-start to triage the diff. If the user just wants a quick task, this pulls attention into memory review first. Mitigation: users who want no compaction at all can toggle it off via `/memory`, and the diff file can simply be ignored until the user is ready to triage (the hook does not block the session or gate future work on review). +- **Gated on project scaffolding** — seeding is piggybacked on `ways init`, which early-returns when the CWD has neither `.claude/` nor `.git/`. Scratch-directory sessions (running Claude Code in an arbitrary folder with no project shape) skip seeding and get the harness's default memory behavior. This inherits `ways init`'s existing gating and is correct for its scope; the consequence is that the routing guidance reaches only project-shaped sessions. + +### Neutral + +- **Memory remains writable** — no permissions change. The seed steers, it doesn't prohibit. Claude can still save memory when genuinely warranted for cross-project user facts. +- **Scope is harness-specific** — the hook exists for this harness shape. The underlying principle generalizes; the specific hook does not. + +## Alternatives Considered + +- **Do nothing; accept the harness's memory defaults.** Rejected — the three failure modes in Context are observed and recurring. Inaction means continued drift and continued short-circuiting. +- **Disable or suppress memory writes entirely.** Rejected. Fights the harness rather than collaborating; breaks on harness updates; removes legitimate cross-project user memory use cases; hard to enforce without permission hacks. +- **Put routing guidance only in `hooks/ways/meta/memory/memory.md` (no seed).** Rejected as sole mechanism. The way fires on triggers; the seed loads unconditionally. To redirect the memory-writing instinct, the guidance must reach the model *before* it forms an intent to save — that requires the force-fed slot. Both channels (seed at startup, way on trigger) compose for defense-in-depth. +- **Hash the whole file including `## User Context`.** Rejected — any legitimate user addition would trigger re-seed. The narrower hash (seeded portion only) distinguishes template integrity from user content. +- **Feature-flag the seeding behind an opt-in.** Rejected. The framing should be the default, not opt-in. Also: a flag is exactly the kind of unrequested preservation scaffold the project has explicitly deprecated. +- **Put project knowledge in `~/.claude/projects/<hash>/memory/` as topic files (current practice).** Rejected — this is the per-instance cage this ADR exists to leave. Topic files in that directory don't travel with the repo, aren't lint-validated, aren't reviewable, and compete with ways for the project's accumulated learning. +- **Rely on Anthropic's periodic memory-compaction to consolidate instead.** Rejected. Compaction consolidates memory *into* memory — it tidies the silo without addressing the two structural problems this ADR targets (per-instance cage, short-circuited discipline). A tidier silo is still a silo. diff --git a/docs/architecture/practice/ADR-132-collaboration-ways-domain.md b/docs/architecture/practice/ADR-132-collaboration-ways-domain.md new file mode 100644 index 00000000..59185180 --- /dev/null +++ b/docs/architecture/practice/ADR-132-collaboration-ways-domain.md @@ -0,0 +1,80 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: authoring +basis: + - evidence: the new onboarding-share way had no natural home, and collaboration ways (meta/teams, meta/trust, meta/subagents) were filed under meta + - precedent: ADR-131 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-04 +deciders: + - aaronsb + - claude +related: [] +imported: + from: docs/architecture/system/ADR-132-collaboration-ways-domain.md + format: v0 + status: Accepted +--- + +# ADR-132: Collaboration ways domain + +## Context + +Ways are organized into top-level domains (`meta`, `softwaredev`, `ea`, …). A domain is **not** a matching mechanism — disclosure is driven by frontmatter (`description`/`vocabulary`) and triggers (`pattern`/`files`/`commands`/state). The directory governs two other things: **organization** and the **include/exclude unit** (`ways.json` disabled domains, project-scope toggles per ADR-131). So a domain boundary is a low-risk, reversible choice — but it is the unit users reason about and toggle, so it's worth drawing on a real seam. + +A new capability surfaced the gap: a way for *when to publish a repo's onboarding guide as a share link a teammate's own agent opens directly*. It had no natural home. Looking for one exposed that **collaboration concerns are scattered or mis-filed under `meta`**: + +- `meta/teams` — coordination norms for agents working in a team. +- `meta/trust` — the relational model between Claude and the human. +- `meta/subagents` — delegating work to ephemeral helpers. + +`meta` is meant for *how the agent itself operates* (knowledge, memory, reasoning, persistence). "Working across the boundary to other people and their agents" is a distinct concern, and `meta` was drifting into a catch-all for it. + +## Decision + +Create a top-level **`collaboration`** ways domain: ways about working across the boundary to other people and their agents — capabilities and norms that exist *because the collaborators are agent-mediated*. + +Initial members: + +- `collaboration/onboarding-share` — new; when to surface publishing an onboarding guide as a teammate-openable share link. +- `collaboration/teams` — moved from `meta/teams` (agent-team coordination norms). + +**Explicitly kept in `meta`, with rationale** (so the boundary is enforceable, not aspirational): + +- `meta/trust` — a *foundational* model other domains derive from (5 referrers: `ea`, `autonomy`, `delegation`, `onboarding-share`). A root concept, not a collaboration occupant. +- `meta/subagents` — execution/parallelization (referenced by `delivery/implement`). "How work gets done," not "collaborating with peers." + +**Deferred:** `meta/attend`. Its children mix concerns — peer-session awareness is collaboration, but `context-pressure` is a *solo* signal. That's a per-child split, not a move; out of scope here. + +**Naming:** chose `collaboration` over `agentic`. Every way is agent-run, so `agentic` fails to discriminate and invites junk-drawer growth. For the same reason, members sit flat (`collaboration/onboarding-share`) rather than under a redundant `collaboration/agentic/` layer. + +## Consequences + +### Positive + +- A coherent home for a growing class of cross-agent collaboration ways; `onboarding-share` lands cleanly. +- `meta` narrows back toward "how the agent itself operates," reducing catch-all drift. +- Collaboration ways become a single include/exclude unit, toggleable as a group per project (ADR-131). + +### Negative + +- One more top-level domain to keep coherent — the boundary must be enforced or it becomes a different catch-all. +- Moving `teams` carries a one-time doc-reconciliation tail (done: 3 live docs updated; legacy ADR-013 left pointing at the old path as a historical record). +- Locale-alias translations and the embedding corpus for the new/moved ways must be regenerated; until then `onboarding-share` matches by `pattern:` only. + +### Neutral + +- No disclosure or behavior change — directory is organization + toggle unit, not a matching input. `teams` still fires identically (`session-start`, `scope: teammate`). +- Future platform sharing/handoff capabilities now have an obvious destination. + +## Alternatives Considered + +- **Keep `onboarding-share` in `meta` (e.g. `meta/handoff`)** — rejected: `meta` is already absorbing collaboration concerns; the point is to stop that drift, not extend it. +- **Name the domain `agentic`** — rejected: too broad to discriminate (all ways are agentic); a junk-drawer waiting to happen. +- **Sub-group as `collaboration/agentic/…`** — rejected: redundant nesting; the whole domain is agent-mediated. +- **Move `trust`/`subagents`/`attend` in too** — rejected/deferred: `trust` is foundational and heavily referenced, `subagents` is execution, `attend` mixes solo and collaboration signals (needs a split, not a move). diff --git a/docs/architecture/practice/ADR-138-skills-own-the-how-ways-own-the-5w.md b/docs/architecture/practice/ADR-138-skills-own-the-how-ways-own-the-5w.md new file mode 100644 index 00000000..d256ae6d --- /dev/null +++ b/docs/architecture/practice/ADR-138-skills-own-the-how-ways-own-the-5w.md @@ -0,0 +1,101 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: + - method + - authoring +basis: + - evidence: 'an agent vendoring adr/doc tooling hit a dead end: the cp procedure was copied into four places (two macros, the migration way, project-init) and the adr and docs skills carried none of it' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-21 +deciders: + - aaronsb + - claude +related: [] +imported: + from: docs/architecture/system/ADR-138-skills-own-the-how-ways-own-the-5w.md + format: v0 + status: Accepted +--- + +# ADR-138: Skills own the how, ways own the 5W + +## Context + +An agent trying to vendor the doc-management tools (`adr`, `doc`) into a fresh +project couldn't reliably find the install procedure. Investigation showed the +way+macro disclosure path actually works — in a real empty project the +`adr`/`documentation` ways fire and their macros surface the "tooling available" +observation. But the *procedure itself* — the `cp` of `adr-tool`/`doc-tool` from +`~/.claude`, the copy-not-symlink rule, the `doclint.py` pairing — had been +copied into **four** places: the two macros, the `migration` way, and the +`project-init` command. The `adr` and `docs` **skills**, which are the +front-and-center entry points (they trigger on "create an ADR", "new doc page") +and bypass the macro system entirely, carried *none* of it. An agent reaching the +tooling via the skill in an un-provisioned project hit a dead end. + +The deeper issue is a missing authoring boundary. Without a rule for what belongs +in a skill versus a way, an imperative procedure leaks into every artifact that +references it and drifts — exactly the consistency drift the freshness way warns +about. We need one home for each procedure, and a principle that says where. + +## Decision + +Adopt a single authoring boundary across the corpus: + +- **Skills own the *how*** — the executable procedure. One canonical, ideally + **parametric** home per procedure (one skill with modes, not N near-duplicate + skills). Skills are invoked deliberately and must be self-sufficient: they may + not assume a way fired first. +- **Ways own the *who / what / where / when / why*** — disclosure and context. + A way observes state ("this tool isn't installed and should be / its state is + inconsistent"), explains why it matters, and fires at the right moment. A way + (or its macro) must **point to** the skill for the procedure; it must not be the + canonical home of one. + +A way's macro is an *observation*, not an instruction sheet. When it detects a +gap, it names the gap and hands off to the skill. + +This change implements the boundary for doc-tooling vendoring as the motivating +prototype: the `cp` procedure now lives only in the `adr` and `docs` skills; the +macros, `migration` way, and `project-init` command all defer to them. The +principle then governs the wider sweep — consolidating clone-skills into +parametric ones and relocating any *how* still smuggled inside a way. + +## Consequences + +### Positive + +- One home per procedure — drift surface collapses from four copies to one. +- Skills become safe to invoke standalone; the dead-end on the skill path closes. +- A clear test for every future artifact: imperative steps → skill; observation, + rationale, timing → way. Authoring decisions stop being ad hoc. +- Parametric consolidation reduces the skill count and the surface to maintain. + +### Negative + +- Existing clone-skills and ways-carrying-procedures need a migration sweep. +- A way that observes a gap now costs one extra hop (way → skill) to act on it, + versus an inline command. + +### Neutral + +- Pairs with the prototype-before-accept principle: vendoring was prototyped, + this ADR records the intent; acceptance follows the sweep. +- Suggests a companion meta-way on *authoring skills vs ways* that fires when + editing `skills/` or `hooks/ways/`, enforcing the split going forward. + +## Alternatives Considered + +- **Add the procedure to the skills but leave the other three copies** — closes + the reported dead-end but keeps the four-way drift. Rejected: treats the symptom. +- **Keep the how in the macros, point skills at the macros** — inverts the natural + roles (a deliberate-invocation artifact depending on a passive disclosure one) + and macros can't be invoked on demand. Rejected. +- **Memory note instead of an ADR** — a cross-cutting authoring principle that + changes how every skill and way is written is architecture, not a narrow fact. + Memory would bypass the friction this decision deserves. diff --git a/docs/architecture/practice/ADR-141-knowledge-graph-as-evidential-memory-backend.md b/docs/architecture/practice/ADR-141-knowledge-graph-as-evidential-memory-backend.md new file mode 100644 index 00000000..4d432656 --- /dev/null +++ b/docs/architecture/practice/ADR-141-knowledge-graph-as-evidential-memory-backend.md @@ -0,0 +1,121 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: method +basis: + - evidence: native auto-memory exposes no programmatic integration hooks and its compaction rewrites contradictions away; the session ledger is a flat stream with no cross-session relationships + - precedent: ADR-112 + - precedent: ADR-128 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-29 +deciders: + - aaronsb + - claude +related: + - ADR-112 + - ADR-128 +imported: + from: docs/architecture/system/ADR-141-knowledge-graph-as-evidential-memory-backend.md + format: v0 + status: Accepted +--- + +# ADR-141: Knowledge Graph as Evidential Memory Backend + +## Context + +Cross-session knowledge has two homes in this framework, and both are *editorial* — they keep a single current view and discard what they supersede: + +- **Native auto-memory** (ADR-128): `MEMORY.md` + topic files, loaded at session start. Its compaction cycle ("Auto-Dream") performs minimal consistent revision — when session 12 contradicts session 3, it picks one and rewrites. The *fact that understanding shifted* is destroyed. It also exposes **no programmatic integration hooks** — no `memoryWrite` intercept, no API, only the `/memory` toggle. +- **The session ledger** (ADR-112) preserves raw prose, but is a flat append-only stream. It records *what was said*, not the relationships between assertions across sessions. + +ADR-112 already anticipated a third tier: ingesting ledger entries into an external **evidential** knowledge graph (KG) — one that, per Dempster–Shafer rather than AGM, keeps contradictory assertions and *scores the balance of evidence* (AFFIRMATIVE / CONTESTED / CONTRADICTORY / INSUFFICIENT_DATA). That tier was sketched as a one-line file copy guarded by `if -d`, but left two things unspecified: + +1. **Activation and detection** — how a session decides, safely and without configuration ceremony, that a valid KG is present *and* that this project has opted in. +2. **The return path** — ADR-112 only described writing *into* the KG. It did not address consuming the KG's evidential view *back* into the session, where the contradiction-preserving signal would actually change reasoning. + +The motivating force is general: an agent operating under a shrinking context window needs a durable, *non-editorial* memory whose value compounds across sessions, and it must integrate without fighting a host memory system that offers no hooks. This ADR records how the framework attaches to such a backend. + +## Decision + +Integrate an external evidential knowledge graph as an **optional, presence-gated** memory backend, in two independently activatable directions. The framework owns *detection and activation*; the KG owns *curation*. The reference backend is a `kg`-compatible system exposing per-ontology ingestion (e.g. a FUSE mount at `${KG_KNOWLEDGE_ROOT:-$HOME/Knowledge}/ontology/<ontology>/ingest`, or an equivalent `POST /ingest` API). + +### 1. The activation contract: presence *is* opt-in + +A project's KG ontology is named after its slug. The **existence of that ontology** answers two questions at once: + +| Question | Established by | +|---|---| +| Is a valid KG reachable? | The ontology's ingest endpoint resolves — the backend had to answer for the ontology to exist | +| Has *this project* opted in? | The resolved ontology is named for *this* project | + +Opt-in therefore requires no separate config in the happy path: if the ontology is absent, the integration is inert at zero cost. A single `ways.json` flag exists only as an explicit veto. + +**Three states**, evaluated as `not-vetoed ∧ ontology-present`: + +- `inactive` — ontology absent → integration silently does nothing (default). +- `active` — ontology present → session knowledge flows. +- `hard-off` — `reflection.kg = false` → never flows, even if the ontology exists (per-machine / per-session veto). + +**Activation is the act of creating the ontology.** A helper exposes `activate` (create the ontology), `status` (report the state for this project), and `deactivate` (set the veto flag; never deletes data). + +### 2. Detection is fail-safe and never blocks the session + +Detection resolves the project ingest target or reports absence. It must: + +- treat an unreachable or orphaned backend as *absent*, not as an error (a dead mount stats false → `inactive`); +- never block — any copy to the backend is backgrounded and fire-and-forget, consistent with the Stop hook running `async` (ADR-112). A wedged backend must not stall a turn. + +The same contract has two transports: a live ingest directory (file copy) where a mount is present, or `GET /health` + ontology-exists + `POST /ingest` for headless/CI hosts with no mount. + +### 3. Outbound — session knowledge to the KG + +When `active`, the framework copies its existing durable artifacts to the ontology's ingest target: ledger entries (ADR-112) and auto-memory writes (ADR-128). These are already curated prose with frontmatter; no new artifact is produced. The KG deduplicates across channels. This is fire-and-forget: the source artifact exists regardless of whether ingestion succeeds. + +### 4. Inbound — the KG as a read-only memory projection (staged) + +Because native memory exposes no hooks, the only integration surface for the *return path* is the memory files themselves (ADR-128's "observe and rewrite" constraint). The framework therefore consumes the KG by **projecting it into the native memory surface as read-only files**: a generated `MEMORY.md` index of the highest-value concepts (ranked by grounding × connectivity, scoped to the project ontology, within the 25 KB session-start window) plus concept topic files served on demand. Claude reads KG concepts through the same slot it already reads memory — including `CONTESTED`/`CONTRADICTORY` entries that editorial curation would have erased. + +This direction is **staged behind the outbound bridge** because it carries three constraints that must be resolved in the implementation ADRs/PRs, not assumed away: + +- **The projection is read-only, and the editorial layer must be disabled.** Auto-Dream rewrites and prunes memory files on its own schedule and cannot be hooked; a projected surface must present as read-only and run with Auto-Dream toggled off, otherwise the curation this integration exists to escape runs against the projection. +- **Writes redirect to ingestion, not to the memory path.** With the read surface read-only, the framework's "save a memory" path must target the KG ingest channel (direction 3), closing the loop: reads come from the projection, writes go to ingestion. +- **Mounting over a populated, host-owned directory is unsolved by a pure API-backed filesystem.** Presenting the projection at the memory path requires either relocating genuine user-level memory into the KG (takeover) or union/overlay support in the backend filesystem. This is an implementation decision deferred to the inbound PR. + +### 5. The generalized principle + +Cross-session memory should be **artifact-presence-gated and evidentially curated**: the framework detects and gates an external backend by the presence of the project's own artifact (not by configuration), pushes its existing durable artifacts to it fire-and-forget, and — where the host permits — consumes the backend's evidential view back through whatever surface the host already reads. The framework never becomes the curator; it routes. Any backend satisfying the ingest/projection contract is substitutable. + +## Consequences + +### Positive + +- Sessions gain a memory whose value compounds — contradictions are preserved and scored rather than overwritten, and minor signals survive as connected low-weight nodes instead of being pruned. +- Zero-ceremony opt-in: creating the ontology activates the project; its absence is a silent, costless no-op. +- No coupling to a specific backend — the ingest/projection contract is the interface; the reference KG is replaceable. +- Outbound ships independently of inbound; each direction is separately activatable, consistent with ADR-112's tiered model. +- Works without fighting the host: outbound copies existing artifacts; inbound uses the only surface native memory exposes. + +### Negative + +- Inbound depends on disabling Auto-Dream and on resolving the populated-mountpoint problem — non-trivial, hence staged. +- Detection correctness hinges on fail-safe behavior against orphaned/wedged backends; a naive `-d` check without liveness/timeout discipline could stall a turn (a portable timeout is itself a follow-up, as stock macOS lacks `timeout`). +- A ranked projection within the 25 KB window means the session sees a *curated slice* of the graph, not all of it; ranking quality becomes load-bearing. + +### Neutral + +- This refines ADR-112's Tier 2 (gives it an activation contract and a return path) and consumes ADR-128's "redirect, don't suppress" model (the projection *is* the redirect, taken to its conclusion). +- The ledger remains the source of truth and the replay seed; the KG is a derived, regenerable view. +- Implementation spans two repos (the framework: detection, activation, hooks, write-redirect; the backend: memory-view formatter, ranked projection query, read-only/overlay mount mode) and will land as separate referencing PRs. + +## Alternatives Considered + +- **Query the KG via MCP at state transitions instead of projecting into memory.** Rejected as the *primary* return path: it requires the session to issue queries and the host to surface results, whereas the memory slot is read unconditionally at session start. MCP query remains a viable complementary retrieval mode and is not precluded. +- **Disable native memory entirely.** Rejected for the same reasons as ADR-128 — it fights the harness, breaks on updates, and removes legitimate cross-project memory. The projection coexists with the feature instead of replacing it. +- **Configuration-flag opt-in (enable per project in `ways.json`).** Rejected as the primary mechanism: it duplicates state that the ontology's existence already encodes and invites drift (flag on, ontology absent). The flag is retained only as a veto. +- **Push via a CLI/MCP call rather than a file copy.** Rejected for outbound: a guarded file copy is fire-and-forget, has no failure surface that can block a turn, and needs no tool call. The API transport is the documented fallback for hosts without a mount. +- **Full takeover vs. overlay for the inbound mount.** Deferred, not decided here — both satisfy the contract; the choice belongs to the inbound implementation ADR. diff --git a/docs/architecture/practice/ADR-143-three-root-way-runtime-core-user-project.md b/docs/architecture/practice/ADR-143-three-root-way-runtime-core-user-project.md new file mode 100644 index 00000000..15da4f4b --- /dev/null +++ b/docs/architecture/practice/ADR-143-three-root-way-runtime-core-user-project.md @@ -0,0 +1,177 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - authoring + - matching +basis: + - evidence: 'the runtime scans only global and project roots (live ways status: "Global ways: 115 total, 107 semantic"), so a shipped way cannot be shadowed without editing it in place' + - precedent: ADR-142 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-29 +deciders: + - aaronsb + - claude +related: + - '[[ADR-142]]' + - '[[ADR-140]]' + - '[[ADR-111]]' +imported: + from: docs/architecture/system/ADR-143-three-root-way-runtime-core-user-project.md + format: v0 + status: Accepted +--- + +# ADR-143: Three-root way runtime — core, user, project + +## Context + +This is a child of ADR-142 (agent-ways 1.0). The spine moves the *application* to +`$XDG_DATA` and the *operator's own ways* to `$XDG_CONFIG`. This ADR records what that +split means for the **runtime that loads ways**. + +Today the `ways` runtime scans **two** root classes: + +- **global** — `~/.claude/hooks/ways` (`home_dir().join(".claude/hooks/ways")`, e.g. + `tools/ways-cli/src/cmd/language.rs:18`, `permissions.rs:12`). The corpus builder scans + it as the first root (`tools/ways-cli/src/cmd/corpus.rs:69-74`, "Scan global ways"). +- **project** — `<project>/.claude/ways`, resolved from `CLAUDE_PROJECT_DIR` + (`corpus.rs:89-94`), scanned second and namespaced per project. + +`ways status` reports exactly these two: a single **"Global ways"** count and a **Projects** +list (verified against a live `ways status`: "Global ways: 115 total, 107 semantic" followed +by per-project counts). + +A two-tier **precedence already exists** at file-resolution time: `resolve_way_file` +(`tools/ways-cli/src/session.rs:563-577`) looks up a way ID in `<project>/.claude/ways/<id>` +first and falls back to `~/.claude/hooks/ways/<id>` — "Project-local takes precedence." So +*project > global* is not new; what is missing is a tier *between* them for the operator's own +ways, because there is no operator-owned root distinct from global to resolve against. + +The defect is that **"global" conflates two different owners in one directory.** A way +agent-ways *ships* and a way the *user wrote themselves* both land in `~/.claude/hooks/ways`, +indistinguishable to the runtime. Three consequences follow: + +1. **You cannot shadow a shipped way without forking it.** To change how a core way behaves, + the only lever is editing the shipped file in place — which the next update clobbers (and, + pre-1.0, which made update unsafe to automate at all; ADR-142 §1). +2. **The corpus cannot tell app ways from user ways**, so neither can any tooling built on it + (telemetry, tuning, `siblings`, permissions audit). +3. **The runtime's root model contradicts the storage model ADR-142 just created.** The spine + puts core ways in `$XDG_DATA` and user ways in `$XDG_CONFIG` — two physically separate + locations — but a two-root runtime can only see one "global." The runtime has to learn the + split or the storage split is invisible at match time. + +## Decision + +**Promote the runtime from two roots to three — core, user, project — scanned as one runtime +class, with precedence `project > user > core` and dedup-by-name.** This splits today's +single "global" into its two real owners and makes shadowing a first-class, non-destructive +operation. + +### 1. Three roots, one runtime class + +| Root | Location | Owner | Updatable | +|---|---|---|---| +| **core** | `$XDG_DATA_HOME/agent-ways/hooks/ways` (today's `~/.claude/hooks/ways` content) | agent-ways (shipped) | replaced wholesale on update | +| **user** | `$XDG_CONFIG_HOME/agent-ways/ways` | the operator | never touched by update | +| **project** | `<project>/.claude/ways` (unchanged) | the repo | per-repo, committed | + +They are scanned as **one runtime class**: a way is a way regardless of which root it came +from. The runtime tags each loaded way with its origin root so downstream consumers (corpus +namespacing, `status`, telemetry, permissions audit) can distinguish them — the tag is what +today's model structurally lacks. + +### 2. Precedence `project > user > core`, dedup-by-name + +When the same way *name* appears in more than one root, the higher-precedence root wins and the +lower one is dropped from the active set (dedup-by-name). Precedence runs **project > user > +core** — this *inserts the user tier* into the project-over-global order `resolve_way_file` +already implements (session.rs:563-577), rather than inventing precedence from scratch: a repo +can override anything for its own checkout; an operator can override a shipped way for all their +work; core is the floor. This is the mechanism that lets a user **shadow** a shipped way — drop a same-named way in `$XDG_CONFIG`, and it wins over core without editing or +forking the shipped file. + +(Open question, below: whether shadow should be *whole-file replacement* by name or *field-level +overlay*. This ADR commits to whole-file-by-name as the 1.0 semantics and parks overlay.) + +### 3. Corpus builder and `ways status` extend from two roots to three + +- **Corpus** (`cmd/corpus.rs`): the "Scan global ways" step splits into "Scan core ways" + + "Scan user ways," each content-hashed for staleness (the existing `global_hash` mechanism, + `corpus.rs:71`, generalizes to per-root hashes). Project scanning is unchanged. Dedup-by-name + is applied across the three before embedding so a shadowed core way isn't double-embedded. +- **`ways status`**: the single "Global ways" line becomes two — **core** and **user** — making + the split observable, and surfacing shadowing (e.g. "3 user ways, 1 shadowing a core way"). + +### 4. Scope boundary + +This ADR changes *where the runtime looks and how it resolves collisions*. It does **not** change +the matching pipeline (ADR-108 embeddings), progressive disclosure (ADR-105), or the single-binary +consolidation (ADR-111) — three roots feed the same matcher. It depends on ADR-142 for the +existence of the `$XDG_CONFIG` user root; absent the spine, "user root" has no home. + +## Consequences + +### Positive + +- **Shadowing without forking.** An operator overrides a shipped way by name in `$XDG_CONFIG`; + the override survives every update because user scope is never in the update blast radius + (ADR-142 §7). This is the runtime-side payoff of the whole 1.0 restructure. +- **Clean open/closed boundary.** Core is extended (by user/project ways) without being modified — + the SOLID open/closed principle expressed structurally rather than by convention. Today's only + extension mechanism is *modifying* the shipped tree, the exact opposite. +- **App-vs-user becomes visible to every corpus consumer**, because origin is now a tag, not a lost + distinction — tuning, telemetry, `siblings`, and the permissions audit can all scope by root. +- **The runtime model and the storage model finally agree** (ADR-142's two physical locations map + to two runtime roots). + +### Negative + +- **Precedence is a new surface for surprise.** "My way isn't firing" can now mean "a + higher-precedence root shadows it." `ways status` must make shadowing legible or it becomes a + silent debugging trap — the same class of "topology-fragile, went stale once already" failure + ADR-140 flagged for path-based checks. +- **Dedup-by-name needs a defined collision policy across three roots**, including the awkward case + of a project way and a user way colliding with *different intent* — name collision is now load- + bearing where before there was one namespace per scope. +- **More roots to scan at every corpus build**, with per-root staleness hashing; modest added cost + and more cache-invalidation paths. +- **Whole-file shadow is coarse.** Overriding one frontmatter field (e.g. a trigger threshold) + requires copying the whole way into user scope, which then *won't* track upstream improvements to + the rest of that way — a real maintenance cliff the overlay alternative would avoid (parked). + +### Neutral + +- The two-root corpus namespacing (`encode_project_key`, `corpus.rs:94`) generalizes to three; the + on-disk corpus format gains a root tag but is otherwise unchanged. +- Project scope is entirely unaffected in mechanism — only its position in a now-explicit precedence + order is stated. + +## Alternatives Considered + +- **Keep two roots; put user ways in the same dir as core and tag by some marker.** Rejected: it + re-creates the exact conflation ADR-142 §1 exists to remove, and leaves user ways inside the + update blast radius (a marker file doesn't move them out of `$XDG_DATA`). The split must be + physical to make update safe. +- **Field-level overlay instead of whole-file shadow** (user way patches named fields of a core + way; unspecified fields inherit). More powerful and avoids the maintenance-cliff Negative — but + needs a merge semantics, conflict rules, and a way to express "remove this field," none of which + exist today. Deferred to a follow-on ADR; 1.0 ships whole-file-by-name and learns from it. +- **Four roots** (split project into project-shipped vs project-local, or add an org/team root). + Rejected for 1.0 as scope the brief doesn't call for; the org-rollout need ADR-140 raised could + motivate a team root later, but it would slot into the same precedence chain without re-deciding + this model. + +## Open Questions + +- **Shadow semantics: whole-file vs field-level overlay** (committed to whole-file for 1.0; overlay + parked as a follow-on). +- **Collision policy detail** when project and user ways collide by name with divergent intent. +- **Whether this child stays a separate ADR or folds into ADR-142.** Recommendation: keep separate — + it has its own decision (precedence + dedup), its own consequences (the shadow maintenance cliff), + and touches a distinct subsystem (the matcher/corpus), which is enough to stand alone. diff --git a/docs/architecture/practice/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md b/docs/architecture/practice/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md new file mode 100644 index 00000000..46328658 --- /dev/null +++ b/docs/architecture/practice/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md @@ -0,0 +1,157 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: loop +basis: + - evidence: outside a /goal loop the review-fix-merge tail was hand-typed many times a day with no named invocation, and wrap had no opening counterpart + - precedent: ADR-138 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-07 +deciders: + - aaronsb + - claude +related: + - 128 + - 138 + - 143 +imported: + from: docs/architecture/system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md + format: v0 + status: Accepted +--- + +# ADR-165: Loop-control bookends: start, develop, merge, release, wrap + +## Context + +The project already teaches every *stage* of software work as a way — design, +prototype, ADR, plan, build, PR, review, merge, docs, release each disclose +themselves when triggered. What it has never named is the **loop that carries a +piece of work through those stages**, or the ritual of **opening** and **closing** +a working session around it. + +Two concrete pains motivated this: + +1. **The tail is retyped every session.** Outside a `/goal` loop, the operator + hand-types variations of "run a code review on this, fix the findings, and + merge" many times a day. The doctrine for that tail already exists (the deliver + workflow in `meta/workflows`, the review-before-merge default in + `delivery/github`) but there was no single named invocation for it. + +2. **There is a closing bookend but no opening one.** `wrap` (skill + `meta/wrap` + way) is a fully-realized *terminal* ritual: it squares the TaskList, writes a + continuation prompt, and hands the operator a directed `/compact`. Nothing + symmetric exists at the *start* of a session — orienting to where work was left + off, framing intent, and warming context before planning. + +Underneath both is a shift already recorded in `core.md`: the method is no longer +"ADR-driven development." Prose states *claims*; claims are held to the evidence +the running system produces; "the ledger is not the whole method." An ADR is one +stage among several, and the *order* of the early stages is not fixed — +design-first, prototype-first, and adr-first are all valid openings depending on +**where the uncertainty lives** (`core.md`'s uncertainty-location map). The loop +has a **variable front and a stable tail**, and no way named that. + +## Decision + +Introduce a **five-verb loop-control family**, each an operator-invocable skill +with a corresponding way (the way is the *when/why*, the skill is the *how* — +ADR-138): + +| Skill | Role | Way | +|-------|------|-----| +| `/start` | Open the session — orient to prior state, greenfield interview, gauge-guard, warm context, then recommend planning | `meta/start` (new) | +| `/develop` | Carry the core loop; pick the loop shape (variable front / stable tail) and **borrow** the stage skills | `meta/develop` (new) | +| `/merge` | Land an increment — the four-square review gate → merge → cleanup | `delivery/merge` (new) | +| `/release` | Publish a release — version, changelog, artifacts | `delivery/release` (exists) | +| `/wrap` | Close the session → hand to a new one | `meta/wrap` (exists) | + +Five load-bearing choices: + +1. **`/start` and `/wrap` are inverse gauge-aware bookends.** Both read the same + instrument (`ways context`). `wrap` checks the session is near its *end* and + scales the handoff to how much is about to be lost; `start` checks the session + is near its *beginning* and **dissuades** starting if it is not — starting is a + beginning-of-session act. Same instrument, opposite pole. + +2. **Planning lives inside `/start`, not as a `/plan` skill.** Claude Code reserves + plan mode; we do not claim that verb. `/start` gathers state *first*, so when it + recommends planning the context is already warm — plan mode inherits an oriented + situation instead of a blank one. Like `wrap` and `/compact`, `/start` cannot + itself invoke plan mode (skills/hooks cannot fire `/` commands); it prepares the + ground and hands the operator the trigger. + +3. **`/develop` is a borrowing-router, not an orchestrator.** It establishes the + loop, selects the front order by where the uncertainty lives, lays the TaskList, + then delegates to the reviewer (the `code-reviewer` subagent), `/merge`, and + `/release`, and lets the existing stage ways disclose. It does not reimplement + what those already teach. (The monolithic-orchestrator reading was considered and + rejected — see Alternatives.) + +4. **`ship` splits into `merge` + `release`.** "Ship" carried release-weight — + version bumps, changelogs, artifacts — which made it the wrong name for the daily + act of landing an increment. `/merge` is the light, high-frequency tail (branch → + commit → PR → review-gate → merge → cleanup, minus publish); `/release` is the + occasional heavy publish (ship's former Publish step). The daily retype now maps + to `/merge`. + +5. **The review gate is a four-square, not a single policy.** `/merge` classifies + the work on two axes and picks the path, surfacing the call when ambiguous: + + | | machine review: light | machine review: deep / swarm | + |-------------------------|-----------------------|------------------------------| + | **human gate: none** | run it — quick review → auto-merge | swarm review + adversarial verify → agent-gated merge | + | **human gate: required**| single-agent review → offer to read before merge | swarm review **and** operator approval before merge | + + The X axis is driven by complexity / blast-radius; the Y axis by whether the work + sets direction (e.g. an ADR) and whether the operator has already read it. + +### Skill ↔ way coverage + +Most stages are already covered because `/develop` *borrows* them. New authoring is +concentrated in four places: `meta/start`, `meta/develop` (the variable-front / +stable-tail selector — the corpus's largest prior gap), `delivery/merge` (the +four-square + the previously-partial "fix/remediate per finding" coverage), and the +rename ripple through references to `ship`. + +## Consequences + +### Positive + +- The daily "review, fix, merge" retype collapses to `/merge`. +- The session gains a symmetric opening ritual; `/start` warms context so planning + starts oriented. +- The loop's variable-front / stable-tail shape is named and disclosed for the first + time, grounded in the claim→evidence method. +- `merge` and `release` each become honestly-scoped; neither overloads the other. + +### Negative + +- A rename ripple: references to `ship` across ways, skills, and the skills catalog + must all move to `merge`/`release`. +- Five coordinated surfaces (two new skills, one rename, one split, a hook, and + new/revised ways) are more moving parts to keep coherent than a single skill. + +### Neutral + +- `/review` and `/code-review` (built-in review skills) are unchanged; the merge gate + borrows the `code-reviewer` subagent for its automated review. +- Composes with `/goal`: the manual `/merge` gate is the single-shot form of what a + goal loop automates, mirroring the `wrap` / `compaction-checkpoint` duality. + +## Alternatives Considered + +- **Fold the tail into `ship` / enrich `ship`.** Rejected: "ship" reads as + release-weight, and the operator's daily act is iteration, not publishing. +- **`/develop` as a monolithic orchestrator** that runs every stage itself. + Rejected: it would duplicate the stage ways and violate the borrow-don't-recreate + posture; the mode-setter + shape-selector reading keeps it thin. +- **A standalone `/plan` skill.** Rejected: Claude Code reserves plan mode; planning + folds into `/start` as an optional, context-warmed branch. +- **A dedicated non-agile naming metaphor** (film: action/take/wrap; compiler; + systems). Explored at length; the plain SE verbs (start/develop/merge/release) + won for being literal, low-fatigue, and already the words for the acts. diff --git a/docs/architecture/practice/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md b/docs/architecture/practice/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md new file mode 100644 index 00000000..908613b0 --- /dev/null +++ b/docs/architecture/practice/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md @@ -0,0 +1,198 @@ +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: + - method + - loop +basis: + - evidence: the heron_brook section in Claude Code 2.1.219+ ("Do not call the AgentTool unless the user requested it"), confirmed live in two sessions; meta/subagents ranked fifth at 0.3 on a query naming delegation + - standard: Claude Code 2.1.219+ system prompt, heron_brook section; upstream issue anthropics/claude-code#80988 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-31 +deciders: + - aaronsb + - claude +related: + - ADR-155 +imported: + from: docs/architecture/system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md + format: v0 + status: Accepted +--- + +# ADR-175: Standing delegation authorization satisfies the harness permission gate rather than overriding it + +## Context + +Claude Code 2.1.219+ injects a system-prompt section (`heron_brook`) carrying a single +constant, emitted as one block: + +``` +Do not call the AgentTool unless the user requested it +Do not use workflows or deep-research unless the user requested it +``` + +The gate is the `opus_5_prompt_bundle` model capability plus the killswitch flag +`tengu_fennel_godwit`, currently `false`. Confirmed live from inside two sessions during +the investigation that produced this ADR. Upstream: `anthropics/claude-code#80988`, open, +no staff response at time of writing. + +There is no opt-out. `CLAUDE_INTERNAL_FC_OVERRIDES` is dead code in 2.1.220 — an +unconditional `return` precedes the env read. `DISABLE_GROWTHBOOK=1` makes matters worse, +because the killswitch defaults `false` and blocking the fetch guarantees the gate stays +open. Flipping `tengu_fennel_godwit` is all-or-nothing: the same capability gates five +other sections including `delivering_work_max` and `overcorrection`, both of which are +wanted. Buying delegation by dropping those is not a trade worth making. + +The constant carves out no exception by agent type. `code-reviewer` is suppressed exactly +as `Explore` or `general-purpose` is. The observed effect is that the operator must name +delegation explicitly on every occasion, including for procedures whose own documented +steps call for it. That is the gate working as written, not a defect in it. + +A second, independently gated nudge exists: the `subagent_steer_delegation` experiment, +arm `counter_steer`, injecting *"Subagents multiply cost and time… Delegate only when the +payoff clearly exceeds that overhead."* It was not observed in either sampled session. +Both samples share one account, and GrowthBook buckets on stable user attributes derived +from email and account uuid, so two sessions on one account land the same arm by +construction. The observation discriminates nothing; the gate is untested, not absent. + +### The two gates have different shapes + +`heron_brook` asks a permission question: *did the user request this?* Its condition is +satisfiable — and satisfiable in advance, in writing, by the operator. + +`counter_steer` asks a cost/benefit question: *does the payoff exceed the overhead?* Its +condition is answerable by stating the payoff. + +Neither is a prohibition. A countermand written as "ignore the above" fails both, and +wins only on recency — user-authored way text and user-authored `CLAUDE.md` carry the same +authority, so recency is the only lever such text has, and it is a weak one. + +### Placement is not firing + +`hooks/ways/meta/subagents/subagents.md` is the natural home for delegation doctrine, and +`scope: agent` delivers it to the main loop — `scan/mod.rs` filters the task surface to +ways whose scope contains `subagent`, so `agent` scope is main-loop delivery. But its +trigger surface is: + +``` +pattern: subagent|delegat|spawn.{0,30}agent|review.{0,30}\bpr\b|organiz.{0,30}docs +``` + +Every one of those tokens is something the operator says only once they have already named +delegation — the moment at which the gate is already satisfied and the text is redundant. +The cases that need it most (*"let's merge this"*, *"research X"*) do not match at all. +Correct text on a surface that never fires changes nothing. + +Widening the pattern only partly helps, and is partly blocked. `review` is in the linter's +`COMMON_WORDS` and is flagged regardless of anchoring, so the review half is unavailable. +`merge` is not on that list, and `delivery/merge` already ships a lint-clean phrase pattern +(`merge (this|it|the pr)`) that matches *"let's merge this"* today — so reach was never the +whole problem. + +The binding constraint is different. A way has to *win* a match to disclose, and measurement +during this work put `meta/subagents` fifth at 0.3 on *"delegate the code review to a +subagent"* — below a project-local way — even though that query names delegation outright. +A clause that must win a ranking to appear is a clause that sometimes doesn't. Sites don't +rank; they are read when the procedure is read. + +## Decision + +Encode the operator's authorization as **evidence that satisfies each gate on its own +terms**, placed in two tiers. + +**Tier 1 — doctrine, in `hooks/ways/meta/subagents/subagents.md`.** Two sections. The +first states that invoking a skill or way whose steps call for delegation *is* the user's +request, and names the conditions under which delegation still requires asking. The second +answers the cost/benefit nudge by requiring a one-clause payoff statement at the point of +delegation. This tier explains the reasoning and is allowed to fire only when delegation +is already named; at that moment it supplies the *how*, not the permission. + +**Tier 2 — operative clause, inline at each delegation site.** Every skill or way whose +own steps direct a delegation carries a short authorization clause at that step. Skills are +read at invocation, so the clause is present exactly when delegation is about to happen, +with no dependence on a way firing. This is what reaches the *"let's merge this"* case: the +merge skill loads, the clause sits in the step that dispatches `code-reviewer`. + +**Scope is `AgentTool` only.** Workflows and deep-research are the second line of the same +constant and remain gated. They are token-invasive enough that the operator wants them +proposed and discussed, never invoked unprompted. A change that also frees them is out of +scope for this decision. + +The authorization is not unconditional. Delegation still requires asking when it is not +part of an invoked procedure, when it spends significant tokens outside the stated task, or +when the operator has said to work solo. + +It also grants *whether*, not *how wide*. A procedure that says "fan out" authorizes the +fan-out it describes, not an unbounded agent count; the width is stated before spawning and +asked for past roughly half a dozen. Without that bound the permission gate is answered and +the cost gate quietly isn't. + +Tier 2 carries a corollary: a way that instructs the *opposite* verb at the same moment +re-opens the recency contest this decision declines to enter. `delivery/github` said to +"offer" a reviewer on the same trigger where `delivery/merge` now says to dispatch one. +Aligning both was part of the change, and any future way covering a delegation moment has +to be checked the same way. + +## Consequences + +### Positive + +- Procedures that document a delegation step execute it, instead of stopping to re-ask for + permission the operator already granted by invoking the procedure. +- The mechanism does not contradict the system prompt, so it does not depend on winning a + recency contest against it. +- The payoff clause is useful on its own merits and remains correct whether or not + `counter_steer` is ever present in a given session. +- Tier 2 is inherently drift-resistant in one direction: a delegation site that is deleted + takes its clause with it. + +### Negative + +- Tier 2 duplicates a short clause across several files. Adding a new delegating skill + means remembering to carry the clause, and nothing enforces it. The drift is fail-closed + — a site without a clause grants nothing, so the model asks, which is the pre-change + behaviour — but the tiers can still fall out of step, and Tier 1 must name only + procedures that carry a clause at their own dispatch site. +- The change is unverifiable from inside the repository. Whether the model actually + delegates without asking can only be observed in live sessions. +- If upstream flips `tengu_fennel_godwit` or removes the section, Tier 1 becomes an + explanation of a constraint that no longer exists and will need pruning. + +### Neutral + +- `subagents.md` needed a correctness pass regardless, done alongside: it documented the + tool as `Task` (the model-facing name is `Agent`; the `Task` matchers in `settings.json` + are the harness event namespace and stay), and its roster listed only the `agents/` + directory, omitting the harness built-ins `Explore` and `general-purpose` that the + research way's fan-out step now names. +- `counter_steer` remains untested. Confirming its presence requires a sample from a + different account, not another session on this one. + +## Alternatives Considered + +- **Widen `subagents.md`'s `pattern` to reach the moment-of-use vocabulary.** Rejected, but + not because it is impossible: `review` is in the linter's `COMMON_WORDS` and unavailable, + while `merge` is clean and already used in phrase form elsewhere. Rejected because reach + is not the constraint — the way still has to win a match to disclose, and it ranked fifth + on a query that named delegation explicitly. It would also make a frequently-firing way + out of one that should fire narrowly. +- **Flip the `tengu_fennel_godwit` killswitch.** Rejected: all-or-nothing across the + `opus_5_prompt_bundle` capability, which would also drop `delivering_work_max` and + `overcorrection`. +- **`CLAUDE_INTERNAL_FC_OVERRIDES` / `DISABLE_GROWTHBOOK`.** Rejected: the first is dead + code in 2.1.220; the second guarantees the gate stays open, since the killswitch defaults + `false`. +- **Put the authorization in user `CLAUDE.md`.** Rejected: same authorship as way text, so + it wins on recency alone, and it applies to every project rather than to the procedures + that actually specify delegation. +- **Write the text as an explicit override of the system prompt.** Rejected: fails both + gates on their own terms, and instructs the model to disregard its own instructions — + a pattern that should not be normalized in this corpus regardless of whether it works. +- **Wait for upstream.** Rejected as the sole response: issue 80988 is open with no staff + reply, and the friction is present in every session meanwhile. This decision is + compatible with an upstream fix and is cheap to remove if one lands. diff --git a/docs/architecture/practice/ADR-176-contract-identification-as-the-develop-loop-front-gate.md b/docs/architecture/practice/ADR-176-contract-identification-as-the-develop-loop-front-gate.md new file mode 100644 index 00000000..a9bc7cdb --- /dev/null +++ b/docs/architecture/practice/ADR-176-contract-identification-as-the-develop-loop-front-gate.md @@ -0,0 +1,78 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: loop +basis: + - evidence: a cross-check against socratic (m4vic/socratic) found no cheap gate at the front of a build that establishes what the build is bound to + - precedent: ADR-165 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-08-05 +deciders: + - aaronsb + - claude +related: + - 165 + - 128 +imported: + from: docs/architecture/system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md + format: v0 + status: Accepted +--- + +# ADR-176: Contract identification as the develop-loop front gate + +## Context + +A cross-check against socratic (m4vic/socratic) — a self-interrogation skill that makes an agent question itself before writing code — surfaced one gap the corpus does not already cover better. + +Socratic's core move is a **pre-code contract**: before building, it self-interrogates across engineering domains, self-answers what it can from the codebase, and emits a short surface — what it assumed, the few open questions, the top risks, the plan — then builds and verifies. Most of that is already ours, usually with a stronger mechanism: `meta/choices` owns self-answer-vs-escalate, `softwaredev/delivery/implement` owns the briefing, the `develop` stable tail owns build→review→fix, and the per-domain ways own the question banks. The disclosure-plus-decay mechanism (ADR-160) re-injects to course-correct later turns, which socratic's load-once skill cannot do. + +The one thing missing: a **cheap, always-on gate at the front of a build that establishes what the build is bound to** — below the ceremony threshold of the `implement` briefing, which fires ADR-side on substantial work. ADR-165 gave the develop loop a variable front (design / prototype / adr, ordered by where the uncertainty lives) and a stable tail, but no first move that asks *what are we contracted on?* before the front order is even chosen. + +Socratic answers that question by **generating** a contract through self-interrogation — which assumes task-start is where the thinking begins. Three months into a many-months build, that assumption is false. The binding already exists — in an ADR, an issue's acceptance criteria, a backlog item, a design note, last session's plan or TaskList. The contract is not absent; it is **elsewhere, in the ledger**. So the move is not *create a contract* but *identify what already binds this build, and load it*. This is the same posture `core.md` already states — "when you can't locate what a vague command refers to, the referent is itself the uncertainty — name it, don't hunt for it" — and the ledger philosophy behind ADR-128: memory does not get to shortcut finding the artifact that already holds the decision. + +## Decision + +Introduce **contract identification** as an explicit gate at the front of the develop loop — the first move, run before the variable front order (design / prototype / adr) is selected — expressed as a new thin way in the develop front. + +The gate is **recovery-first, not creation-first**. It reconciles the work-in-hand against the project's existing ledger — ADRs, issues and their acceptance criteria, backlog items, design notes, the prior session's plan or open TaskList — and produces a short contract surface: what this work commits us to, its acceptance test, and what remains open. Three branches: + +- **Binding found** → surface the contract cheaply and proceed. No authoring. +- **Binding partial** → the gaps *are* the open questions. Batch them (via `meta/choices`); do not invent answers. +- **No binding at all** → the absence is the finding. Name it, and route to a binding-establishing stage — `adr`, `design`, or the `implement` briefing — to create one. (Not `prototype`: a prototype burns down uncertainty but produces no binding.) Do not build against nothing. + +Authoring a contract is the no-binding branch, never the default path. This deliberately inverts socratic: socratic generates the contract by self-interrogation; we recover it from the ledger and generate only on absence. The inversion is what keeps the gate cheap on routine turn-40 work and honest to the ledger — a self-generated contract that duplicates an ADR already on disk is the memory-shortcut ADR-128 warns against, wearing a new costume. + +The gate wires into three existing loop-control surfaces (ADR-165) rather than standing alone: + +- **`/start`** already checks tracking state on session open; it extends to "what binds the work being resumed." +- **`/develop`** runs the gate before selecting the front order — the reconciliation is what *tells* you where the uncertainty lives. +- **`/implement`** consumes the recovered contract in its briefing instead of re-deriving intent from scratch. + +## Consequences + +### Positive + +- A cheap, always-on "what are we contracted on?" gate exists below the `implement` briefing threshold, covering routine work the briefing was too heavy to gate. +- Long-project sessions stop building against nothing — the gate forces a reconciliation with the ledger before code, even when no contract was created in the current session. +- The ledger is reused, not duplicated — recovery reads ADRs / issues / plans rather than regenerating intent, keeping memory-shortcut pressure off (ADR-128). +- Absence of a binding becomes a first-class, named signal that routes to the existing front, rather than a silent gap the build papers over. + +### Negative + +- One more front stage to teach, and a risk of ceremony if it reads as mandatory on trivial edits. Mitigated by two facts: recovery is cheap (a ledger read, not an interrogation), and it is a develop-*loop* stage, not a global hook — trivial edits that never enter the loop never trigger it. + +### Neutral + +- Requires a new way file in the develop front plus wiring into the `start` / `develop` / `implement` skills; it refines the ADR-165 loop rather than replacing any part of it. +- Leans on the `tracking` way's session-start check as the resumed-work half of recovery. + +## Alternatives Considered + +- **Fold it into the `implement` briefing** (extend the briefing to fire earlier and cheaper). Rejected: the briefing sits above the routine-work threshold by design; stretching it down loads briefing ceremony onto small tasks, which is exactly the friction the gate is meant to avoid. The gap is *below* the briefing, not inside it. +- **Extend `tracking` + `start` only** (treat it purely as session-orientation). Rejected: it is a build-front concept, not just an open-the-session one. It must fire on fresh work started mid-session, not only when picking up prior state at session open — which is a develop-loop responsibility, not a bookend one. +- **Adopt socratic's generate-a-contract model wholesale.** Rejected: creation-first duplicates the ledger and imports socratic's 697-question checklist — an enumerated tick-list that the corpus deliberately resists (it discloses judgment, not checklists; ADR-128 / the memory doctrine). Recovery-first takes the one idea socratic exposed without the mechanism the framework is built to reject. diff --git a/docs/architecture/practice/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md b/docs/architecture/practice/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md new file mode 100644 index 00000000..2cab730d --- /dev/null +++ b/docs/architecture/practice/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md @@ -0,0 +1,152 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: method +basis: + - evidence: core.md uses the antithesis construction it bans at 11 per thousand words against a corpus baseline of 4, and passes the density postcheck + - standard: ASD-STE100, via the operator's simplified-modified-technical output style + - precedent: ADR-174 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-08-13 +deciders: + - aaronsb + - claude +related: + - ADR-174 +imported: + from: docs/architecture/system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md + format: v0 + status: Accepted +--- + +# ADR-178: Register transfers by demonstration - core.md carries policy + +## Context + +ADR-174 addressed decorated prose with three tiers of guidance and a postcheck that counts. Its own test returned a null result and closed as "unfalsified, not validated." The operator reports output quality has degraded since, and identified `core.md` as the source. + +The measurements support that reading, on a mechanism ADR-174 did not consider. + +### Core.md contains the construction it bans + +`core.md` line 23 reads: *"Don't resolve with 'not X but Y' or 'I can't do X, so I'll do Y' to sound measured."* + +The file uses that construction eleven times by regex count and fourteen by inspection, in 1,002 words. Two are section headers: + +- **Reasoning runs; it doesn't pose.** +- **Prose delivers; it doesn't decorate.** + +Others: "Write to be read, not admired," "Directness is the absence of detour, not short sentences," "The filesystem is not a task queue," "Memory is for what's load-bearing across sessions, not prostration gestures," "The ledger is not the whole method," "what's checkable is the contract, not the prose." + +The ban and its licensed exception arrive in the same sentence, so line 23 demonstrates the shape twice while forbidding it. + +Measured against the corpus: 73 way files over 200 words, 30,555 words, 142 matches, a baseline of **4 per thousand**. `core.md` runs **11 per thousand**, and it is the one file that fires on every session before any work begins. + +### The existing check cannot see this + +Running the `documentation/markdown/density` postcheck logic against `core.md` yields **1 significance clause and 9 em-dashes per thousand words**, against thresholds of 3 and 15. The file passes the instrument this project built to catch exactly this problem. + +The postcheck counts a surface form. What transfers across a session is a rhetorical shape, and the shape survives every rewrite that dodges the regex. + +### The corrected examples teach the dodge + +`meta/trust/prose/prose.md` presents paired before/after rows. Both "delivered" versions relocate the evaluation into a subordinate clause rather than deleting it: + +| Labeled decorated | Labeled delivered | +|---|---| +| "…three dead triggers. That is the mark of a system with no linting." | "…three dead triggers, none of which the system would have surfaced on its own." | +| "Core fires once per session. This is the most serious gap in the retention model." | "Core fires once per session and is never refreshed until compaction." | + +The string `That is the` dies; the move survives. A model reading these learns where the counter looks. + +### The mechanism + +In-context style transfer is imitative before it is instructed. A file's register is a sample; its rules are a claim about a sample. When the two disagree, the sample wins, and `core.md` is the largest register sample in the always-on window — 677 of its 1,002 words sit under Posture, and eight bolded aphorisms lead its paragraphs. + +ADR-174 established the adjacent failure: negative constraints are the weakest instruction form, failing at 22-30% on frontier models. This ADR adds the reason a stronger negative constraint cannot help. Listing "there's a tension here," "two things are true," and "it's worth naming that" as forbidden strings places those strings in every session's context. Polarity is cheap; presence is not. + +ADR-174's null result reads consistently with this. Both arms scored 0.9 significance clauses per thousand with zero antithesis constructions at ~4,400 fresh-session words. The instrument had power and found nothing, because the register had not yet had a long draft to propagate through. + +ADR-174 closes its own test section with **"unfalsified, not validated."** The phrase is the construction under discussion, in the document written to remove it, at the point of maximum epistemic self-regard. A remediation that reproduces its own subject at its own conclusion is the clearest available evidence that the demonstration channel outranks the rule channel. + +### Register now has a better home + +Claude Code output styles inject register into the system prompt, unconditionally, for the whole session. The operator has authored two (`finnish-direct`, `simplified-modified-technical`). The Finnish-direct style carries the same posture as core's Posture section in 90 words, at a flat temperature, with no demonstration of a banned form. + +`core.md` fires at turn 1 and, per ADR-174's re-disclosure finding, never again. Its register instruction is therefore both weaker in placement and louder in demonstration than the surface that supersedes it. + +## Decision + +**`core.md` carries policy and epistemics. Register instruction leaves it.** + +Five changes. + +**1. Strip register content from `core.md`.** + +Removed: the reasoning-tic ban list with its verbatim forbidden strings; the prose delivery/decoration bullets, which ADR-174 already assigned to tiers 2 and 3; the eight bolded aphorism lead-ins. + +Retained: the uncertainty taxonomy, which is epistemic routing carried nowhere else; the collaboration and post-compaction guidance; Method; Language; Line handling; Attribution. One line of the ban list survives as reasoning guidance rather than phrasing guidance — a one-sided claim is stated once and stopped. + +Result: 632 words, all policy, measured by the same prose extraction the checks use. + +**2. The constraint on `core.md` is structural and checked, in `scripts/check-register.sh`.** + +Any file in the always-on path is a style demonstration whether or not it intends to be. "Write plainly" would drift the way "sparingly" drifted, so the constraint is four counts, all at zero for `core.md`: antithesis constructions, significance clauses, bolded paragraph lead-ins, and paired em-dash asides. The script runs in the pre-commit hook and reports the corpus baseline under `--corpus`. + +The fourth count comes from ASD-STE100, by way of the operator's `simplified-modified-technical` output style. STE turns a parenthetical aside into its own sentence, and applying that rule to `core.md` is what made the constructions visible in the first place. A structural constraint governs sentence form, so it holds where a taste rule dissolves: "Reasoning runs; it doesn't pose" fails an STE check mechanically, on the semicolon antithesis and on the second meaning loaded into "run." + +**3. Re-pick the examples in `meta/trust/prose/prose.md`.** + +The "delivered" column deletes the evaluation instead of relocating it. The way also loses its own antithesis constructions, including the sentence that defines directness by metaphor and then reasons from the metaphor across three clauses. + +**4. Register authority belongs to the output style.** + +Ways carry method, evidence discipline, and domain practice. An operator without an output style loses the turn-1 register nudge; that is accepted, because the nudge was net-negative in the measured configuration. + +**5. Sweep the corpus.** + +Fourteen way files measured at or above 8 antithesis constructions per thousand words. Each was edited construction by construction, splitting a load-bearing distinction into two positive statements and deleting a decorative counterweight outright. The corpus moved from **4 per thousand to 2**, with no file left above the advisory line. + +### The tic reproduces under active removal + +While editing `CLAUDE.md` to remove one antithesis construction, the replacement text written into it read: *"Register — how output reads — comes from the active output style, not from the ways corpus."* That is a paired em-dash aside wrapped around a counterweight, authored one line after removing the same two shapes from the same file, by a writer holding the rule in working memory. + +The same thing happened one step earlier. The first repair of `meta/subagents/subagents.md` turned "This is not an override" into "so nothing here overrides it," which demotes the negation into a subordinate clause. That is the exact dodge this ADR documents in `prose.md`'s example table, committed by the author of the table. + +Both are recorded because they bound what any of this can achieve. The register is not held in the rules; it is resident in the writer. A checked count catches it after the fact, and nothing catches it during. + +### Not decided here + +Adding an antithesis counter to the `density` postcheck was considered and deferred. That check reads the text just written on any markdown Edit, so its false-positive cost is paid on every file in the repository, and the regex that catches the tic also catches ordinary negation. ADR-174's threshold discipline holds: a surface that nags trains its reader to ignore it. `check-register.sh` takes the strict thresholds instead, because it is scoped to one file whose count is zero. + +## Consequences + +### Positive + +- The largest register sample in the always-on window stops demonstrating the constructions the project is trying to remove. +- `core.md` drops from 1,002 to 632 words, all of it policy that no other surface carries. +- The output style becomes the single register authority, with no competing voice in the hook path. +- The forbidden strings leave the context window. + +### Negative + +- Operators running no output style lose the turn-1 register guidance entirely. Tiers 2 and 3 of ADR-174 remain, and tier 2 fires only on explicit long-form requests. +- The `density` postcheck still cannot see the shape this ADR is about. Only `core.md` is checked for it. +- The corpus sweep touched fourteen files in one change, so any behavioral effect measured after this lands cannot be attributed to `core.md` alone. + +### Neutral + +- ADR-174's three-tier structure stands. Tier 1 shrinks to nothing for prose guidance; tiers 2 and 3 are unchanged apart from the example rewrite. +- The core re-disclosure gap ADR-174 recorded is unaffected, and a shorter core makes a future distance-based re-disclosure cheaper to consider. + +## Alternatives Considered + +- **Rewrite the examples only, keep core.md intact.** Rejected. The two section headers are the loudest instances in the file, and they arrive before any example does. +- **Strengthen the ban with more explicit prohibitions.** Rejected. ADR-174 measured negative-constraint compliance at 22-30%, and each added prohibition adds its forbidden string to every session. +- **Add antithesis counting to the density postcheck now.** Deferred rather than rejected — see "Not decided here." +- **Leave the outlier way files for a later pass.** Rejected. Register propagates from every sample in the window, so a clean `core.md` sitting among fourteen dyed way files leaves the channel open. The cost is recorded above: the change is now unattributable between core and corpus. +- **A "write plainly" instruction in the authoring way.** Rejected on ADR-174's own finding. An unmeasurable rule has no moment at which compliance can be tested, which is how "use em dashes sparingly" survived to 118 em-dashes. diff --git a/docs/architecture/practice/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md b/docs/architecture/practice/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md new file mode 100644 index 00000000..61e4020c --- /dev/null +++ b/docs/architecture/practice/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md @@ -0,0 +1,333 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: loop +basis: + - evidence: each session re-derives its task list from the issue body by hand, and nothing carries completion back + - evidence: task store behaviour verified in-session on 2026-09-08 and read from the Claude Code 2.1.263 and 2.1.276 bundles + - evidence: feature request anthropics/claude-code#79096 is open with no response +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-09-08 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-114 + - ADR-152 + - ADR-172 +imported: + from: docs/architecture/system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md + format: v0 + status: Accepted +--- + +# ADR-180: GitHub issues as the shared truth for the session task list + +## Context + +Claude Code ships a task list: `TaskCreate`, `TaskList`, `TaskGet`, +`TaskUpdate`, with `TaskCreated` and `TaskCompleted` hook events. The list is +persisted to `~/.claude/tasks/session-<first 8 chars of session id>/`, one JSON +file per task, with a `.lock` and a `.highwatermark` beside them. Each task +carries `id`, `subject`, `description`, `activeForm`, `status`, `owner`, +`blocks`, `blockedBy`, and a free-form `metadata` object. + +This project tracks work in GitHub issues. Every session that picks up an issue +re-derives its task list from the issue body by hand, and nothing carries +completion back. Feature request anthropics/claude-code#79096 asks for a +bridge and is open with no response. No first-party integration exists: +`claude-code-action` reads issues as prose context for a workflow run, the +GitHub MCP server has no notion of a session task, and the Agent SDK exposes +the task tools as observable blocks without wiring them anywhere. + +Three facts, verified in-session on 2026-09-08, shape the design: + +1. **The task store is the injection point.** A task JSON file written to the + session directory by an external process appears in the next `TaskList` + call. Nothing needs to touch the session transcript. +2. **Subagents resolve to the parent's store.** A subagent spawned with the + `Agent` tool inherits `CLAUDE_CODE_SESSION_ID` set to the parent's session + id and `CLAUDE_CODE_CHILD_SESSION=1`. No child task directory is created. A + general-purpose subagent gets no Task tools at all, so the store is the + only task view available to it. +3. **The ID allocator reconciles from disk on read, and ids are strings.** + After an externally written `2.json` was read by `TaskList`, + `.highwatermark` advanced from 1 to 2 without a `TaskCreate` call. A task + file with id `gh-455` listed, updated through `TaskUpdate` with its + metadata intact, and a following `TaskCreate` allocated numeric `4` beside + it. Whether the in-process counter is memoized between reads is + unobserved. + +Implementation read the store module out of the Claude Code 2.1.263 bundle +and settled two facts the first draft left open: + +- **The lock is `proper-lockfile`.** The zero-byte `.lock` file is the lock + target; the mutex is a `.lock.lock` directory taken with `mkdir`, 30 + retries at 5 to 100 ms, stale after 10 s. Creation holds that directory + lock and re-reads the highest numeric id from disk each time, so nothing is + memoized. Updates hold a per-file lock, `<id>.json.lock`, the same way. + Deletes hold none. An external writer using `mkdir` on the same names is a + full participant. +- **The list id, not the session, names the store.** The directory is + `tasks/<list id>/`, where the list id is `CLAUDE_CODE_TASK_LIST_ID` when + set, else the current team's name, else the full session id. An + interactive process initializes an in-process "session team" at startup, + names it `session-<first 8 of its own id>`, and records it in + `teams/<name>/config.json` with `leadSessionId`, `createdAt`, and the + leader's `cwd`. The team directory is removed when the process exits; the + task directory stays. A `claude -p` session never initializes that team, so + its store stays under the full id. `CLAUDE_CODE_ENABLE_TASKS=false` + disables the store outright. +- **A resumed session's list is keyed by an id the hooks never see** (read + from the 2.1.276 bundle after a resume put the mirror in a store nothing + read). The process that resumes a transcript gets a fresh id, names its + team from that id, and reads `tasks/session-<8 of the fresh id>`, which + starts empty. The `SessionStart` stdin and `CLAUDE_CODE_SESSION_ID` still + carry the transcript's id, so no `leadSessionId` matches it. The bridge's + `attach` verb, run first on `SessionStart` for the `startup` and `resume` + sources, therefore finds the team by shape rather than by id: a + `teams/session-*/config.json` whose leader `cwd` is the hook's cwd, that + no other session's bridge state has claimed, and that was created after + this process started. The hook's parent chain reaches the `claude` + process and `ps` gives its age; the oldest team created after that start + is the one created at startup. Without a reachable process the fallback + is the newest team created within the last 300 seconds. `attach` records + the name in the bridge state, and every later verb reads the record; the + full-id fallback is never recorded, so the lead-session match keeps + working where it did. When the record changes, `attach` copies the + previous store's open tasks into the new one with their ids, drops edges + to tasks that did not come along and `owner` (the agent it named did not + survive the restart), and whispers one line. It copies nothing into a + store that already holds a numeric task, into an explicit + `CLAUDE_CODE_TASK_LIST_ID` list, or into a store whose layout it does not + recognize. A resume with no record (the first after this landed, or after + the runtime directory was cleared) whispers that nothing was carried. + `compact` and `clear` keep the process and its team, so `attach` is a + no-op there. + +The schema is `id` string, `subject`, `description`, `status` enum, +`blocks[]`, `blockedBy[]`, optional `activeForm`, `owner`, `metadata` +record. Listing sorts by `Number(id)`, so string ids sort after numeric ones. + +Documented surfaces: the `metadata` field, the `TaskCreated` and +`TaskCompleted` hooks, and `session_id` on hook stdin. `CLAUDE_CODE_SESSION_ID` +is undocumented and works empirically in skill and subagent shells, where no +stdin JSON exists. The file layout is undocumented. `gh` 2.100 exposes +`blockedBy`, `blocking`, and `parent` on `gh issue view --json`, verified +against issue #433. + +An earlier sketch put the issue watcher in `attend` (ADR-113) as a script +sensor, with the on-demand affordance of ADR-114 as the disclosure route. That +was set aside. `attend` is a daemon beside Claude Code's event model. A +subagent can run its CLI but cannot receive Monitor delivery, the same +distinction ADR-172 draws between its two conduits. The one thing the daemon +adds here, waking an idle session, is not worth a second process for a task +list. + +## Decision + +GitHub issues carrying an opt-in label are the shared truth for session tasks. +Each session's task store is a cache of those issues. The bridge is built from +hooks, one executable, one skill, and one way. No daemon. + +Implementation lands in two increments. The first ships `pull`, `whisper`, the +`TaskCreated` guard, the skill, and the way. The second ships `push`. Pull +alone delivers seeding and the whisper, and every write-side failure mode +waits for the increment that owns it. + +### The label is a trust boundary + +The label `tasklist` gates every direction. An issue without it is never +pulled, its title never whispered; a task without `metadata.github_issue` is +never pushed. Repositories without the label see no behavior. + +Applying the label is an authorization act by whoever holds triage on the +repository. An issue body is text written by whoever filed it, editable after +labeling. `pull` writes body text into `description` wrapped in a provenance +fence naming the source and marking it untrusted content, and the way says +the same in prose. The deny posture of ADR-152 applies: the bridge reads +issues and writes back status, and nothing in an issue body can widen that. + +### Field ownership, not merge + +Two-way merge needs a three-way snapshot per task and conflict rules. +Ownership, plus two small pieces of memory per task, avoids all of it. + +| Field | Owner | Rule | +|---|---|---| +| `subject`, `description` | GitHub | pull overwrites when the issue changed since the last pull, tracked by `metadata.body_sha`; a local edit survives until the issue itself changes | +| open or closed | GitHub | pull sets `pending` or `completed` from the observed state | +| `completed` | session, as a request | push closes the issue only when local status disagrees with both the observed remote state and `metadata.pushed_state` | +| `in_progress` | session | push sets a label; never on a closed issue | +| `blockedBy` edges between pulled tasks | GitHub | from native `blockedBy`; a blocking issue without the label yields no edge | +| `blockedBy` edges to local tasks, every `blocks` edge, `owner`, `activeForm` | session | pull never removes them; reciprocals of GitHub edges are added to `blocks`, and every read-merge-write happens under the task file's lock | +| session-side description edits | session | posted as an issue comment; the local text persists until the body changes | + +`metadata.pushed_state` records the last open/closed state this session +pushed. A maintainer reopening an issue after a session closed it is observed +on the next pull as a change since `pushed_state`, the task returns to +`pending`, and push stays quiet. An issue closed as not planned lands as +`completed` with `metadata.state_reason` set, and the whisper says so. + +Two sessions on one repository converge through GitHub. Session A closes #12 +while session B holds it `in_progress`; B's next pull marks it `completed`, +whispers the change, and B's push never labels a closed issue. + +### Identity and ids + +Pulled tasks take deterministic string ids: `gh-<issue number>`. Pull is +idempotent, every session holds the same id for the same issue, `blockedBy` +translation needs no lookup table, and the external writer never touches +`.highwatermark`. The numeric allocator cannot produce a string id, so the two +namespaces never meet. Pull refuses to overwrite a file at its id whose +`metadata.github_issue` does not match, and whispers the conflict. + +Hooks read `session_id` from stdin. Skills and subagent shells read +`CLAUDE_CODE_SESSION_ID`. When neither resolves, the executable refuses to +write, following ADR-172 Decision 4, and says so on stderr. Agent-team +directories under a team name are out of scope for the first increment; the +store module resolves session directories only. + +### The store module + +All store I/O lives in one module. Before writing anything it validates the +field shape of every existing task file; on mismatch it writes nothing and +`whisper` emits one line saying the layout is unrecognized. A pull stages its +files in a temporary directory and renames them into place, so a failure +partway leaves the store as it was. The lock protocol Claude Code uses is +probed once during implementation and recorded here; until then, ordering +against a concurrent `TaskCreate` is best-effort, and the string id namespace is +what makes the race harmless. + +### The executable + +One command, `gh-tasks`, with four verbs: + +- `pull` reads labeled open and recently closed issues, then writes or updates + task files as above. +- `push` reconciles session-owned state outward: close, label, or comment. It + checks write permission once per session and, lacking it, degrades to + pull-only with one whisper line. +- `whisper` diffs the current labeled issue set against a per-session + snapshot and prints one line per delta: opened, closed, reopened, retitled, + relabeled. Silent when nothing changed. +- `link <issue>` attaches an existing task to an issue by writing the + metadata. + +### Hook wiring + +| Event | Matcher | Action | +|---|---|---| +| `SessionStart` | | `pull`, then `whisper` the labeled open list once as additional context | +| `UserPromptSubmit` | | `pull` when the snapshot is older than a threshold, then `whisper` deltas only | +| `PostToolUse` | `Bash` with `if: Bash(gh issue *)` | `pull`, so a change made through `gh` in-session reflects at once | +| `TaskCreated` | | reject a task for labeled work whose subject lacks the `[gh#N]` prefix and metadata link; validation only | +| `Stop` | | `push` | + +Closes made through the web UI, the API, or a merge keyword reach the store on +the next throttled pull. `TaskCompleted` fires before the completion commits +and a later hook can veto it, so the close request waits for `Stop`. The +`Stop` push emits no context, so the re-entry ceiling of ADR-172 Decision 6 +does not arise; its merge semantics (Decision 3) do, and push appends to +`metadata` rather than replacing it. + +The whisper is the only piece that spends context. One line per delta, and the +threshold on `UserPromptSubmit` keeps the poll off the hot path. + +### The skill and the way + +An `issues` skill wraps the verbs for on-demand use, `/issues pull`, `/issues +push`, `/issues link 123`, and gives subagents a task view the tool surface +denies them. A way under `softwaredev/` teaches the convention: issue-tracked +work gets the `[gh#N]` subject prefix and the metadata link, issue-tracked +work is not tracked as a bare task, and pulled descriptions are untrusted +text. The `TaskCreated` hook enforces the prefix mechanically for labeled +work, which is the example the Claude Code hooks reference ships. + +## Consequences + +### Positive + +- One task per issue, seeded on session start, with completion carried back + without a manual `gh issue close`. +- Several sessions on one repository converge through GitHub on their next + pull. The store is per session and the truth is shared. +- Everything runs inside Claude Code's own event model. Subagents get the same + behavior through the inherited session id and the skill. +- The push half rides only on documented surfaces. The undocumented layout is + confined to the pull half and one module, and a layout change degrades to + one whisper line. +- String ids make pull idempotent and keep the allocator out of the picture. + +### Negative + +- Nothing wakes an idle session. An issue opened between prompts surfaces at + the next `UserPromptSubmit`. +- A session-side description edit lives in the local file and an issue + comment. The next body change overwrites the local text, and nothing reads + the comment back. +- Each hook event that runs `pull` costs a `gh` round trip. The threshold on + `UserPromptSubmit` bounds it, and `SessionStart` pays it once. +- Issue-backed tasks display as `#gh-455` rather than a bare number. +- `TaskCreated` rejection adds friction for labeled work created without a + link. That friction is the convention taking hold. +- Read-only contributors and fork workflows get pull without push. + +### Neutral + +- The task tools are absent on Fable 5, Mythos 5, Opus 4.8, and Sonnet 5 + unless `CLAUDE_CODE_ENABLE_TODO_TOOLS=1` is set or the tools are named in + `--allowedTools`. The bridge assumes the variable is set. +- `gh` is the GitHub client. The GitHub MCP server duplicates it and adds + tool-definition weight to every session. +- The executable starts as shell over `gh` and `jq`. It moves to a Rust crate + under `tools/` if the store module grows past what shell holds cleanly. +- Resume behavior and the lock protocol are recorded here once implementation + probes them. +- If Claude Code ships the bridge requested in anthropics/claude-code#79096, + this ADR is superseded and the label and prefix convention carry over. + +## Alternatives Considered + +- **One issue holding a checklist, one task per checklist item.** The shape + the upstream request asks for. One API object, no id translation. Rejected: + checklist items have no independent close, label, assignee, or dependency + state, and cross-issue dependencies have nowhere to live. Sub-issues would + restore that state and collapse back into one-task-per-issue. +- **Read-only pull, no push.** Adopted as the first increment rather than as + the whole design. Pull delivers seeding and the whisper; push is the half + that mutates other people's issues and carries the ownership and trust + findings, so it lands second. +- **An `attend` sensor plus the ADR-114 affordance (ADR-113, ADR-114).** Wakes + an idle session, which the hook design lacks. Rejected: a second process + beside Claude Code's event model, Monitor delivery unreachable from a + subagent, and the idle wake not worth it for a task list. +- **Writing synthetic tool-use records into the session transcript.** Claude + Code owns the transcript and appends to it concurrently. Rejected once the + task store proved to be an injection point. +- **Injecting "create these tasks" as additional context and letting the + model call `TaskCreate`.** Fully on documented surfaces. Rejected: spends + tokens and model attention on bookkeeping, and the model may skip or reorder + the calls. +- **A git-tracked intermediate file synced to GitHub separately.** History and + review for free, and the undocumented-store risk decoupled from the GitHub + risk. Rejected: a third copy to keep in step, slower convergence, and a + commit per status change. +- **Adopt the convention now, defer the machinery.** The prefix and label + cost nothing and carry over under supersession. Rejected as the whole plan: + seeding on session start is the value, and the convention alone does not + deliver it. +- **True two-way merge with a three-way snapshot.** Rejected for v1: + ownership plus `body_sha` and `pushed_state` cover the observed cases, and + the merge can be added behind the same verbs if a field ever needs it. +- **GitHub Projects as the truth instead of issues.** Richer status model. + Rejected: issues are what this project uses, `gh` handles them directly, + and Projects adds a second object to keep in step. +- **The GitHub MCP server as the client.** Rejected: duplicates `gh` and adds + tool definitions to every session, including sessions that never touch + issues. diff --git a/docs/design-notes/cognitive-loop-and-awareness-layer.md b/docs/architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md similarity index 96% rename from docs/design-notes/cognitive-loop-and-awareness-layer.md rename to docs/architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md index b6ee1f4b..443e3ea5 100644 --- a/docs/design-notes/cognitive-loop-and-awareness-layer.md +++ b/docs/architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md @@ -1,7 +1,16 @@ -# Cognitive Loop and the Awareness Layer +--- +contract: adr/v1 +kind: evidence +capability: method +status: accepted +date: 2026-04-09 +deciders: + - aaronsb +related: [] +--- + +# ADR-600: Cognitive loop and the awareness layer -> **Type:** Design note (not an ADR) -> **Status:** Working draft, subject to revision > **Cites:** ADR-103, ADR-104, ADR-105, ADR-106, ADR-108, ADR-111, ADR-112 > **Motivates:** ADR-113, ADR-114 @@ -264,18 +273,18 @@ To be explicit about what this framing *does not* propose: **ADRs motivating this note or cited within it:** -- [ADR-103](../architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md) — Checks: epoch-distance-aware confidence sensors for ways -- [ADR-104](../architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Token-gated way re-disclosure -- [ADR-105](../architecture/system/ADR-105-progressive-disclosure-for-way-trees.md) — Progressive disclosure for way trees -- [ADR-106](../architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md) — Project Pulse: epoch-mapped project awareness -- [ADR-108](../architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) — Embedding-based way matching -- [ADR-111](../architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) — Unified ways CLI -- [ADR-112](../architecture/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — Session ledger and knowledge graph integration +- [ADR-103](../ways/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md) — Checks: epoch-distance-aware confidence sensors for ways +- [ADR-104](../ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Token-gated way re-disclosure +- [ADR-105](../ways/ADR-105-progressive-disclosure-for-way-trees.md) — Progressive disclosure for way trees +- [ADR-106](../documentation/ADR-106-project-pulse-epoch-mapped-project-awareness.md) — Project Pulse: epoch-mapped project awareness +- [ADR-108](../ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) — Embedding-based way matching +- [ADR-111](../platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) — Unified ways CLI +- [ADR-112](../archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — Session ledger and knowledge graph integration **Prior attempts at adjacent capabilities (for context on why the awareness layer is different):** -- [ADR-101](../architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) — Wormhole relay protocol (Deprecated) -- [ADR-102](../architecture/system/ADR-102-irc-based-local-agent-communication.md) — IRC-based local agent communication (Abandoned) +- [ADR-101](../attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) — Wormhole relay protocol (Deprecated) +- [ADR-102](../attend/ADR-102-irc-based-local-agent-communication.md) — IRC-based local agent communication (Abandoned) **ADRs that cite this note:** diff --git a/docs/design-notes/autonomy-goals-signposts-initiation.md b/docs/architecture/practice/ADR-601-autonomy-as-a-layered-system-goals-signposts-and-the-initiation-pattern.md similarity index 97% rename from docs/design-notes/autonomy-goals-signposts-initiation.md rename to docs/architecture/practice/ADR-601-autonomy-as-a-layered-system-goals-signposts-and-the-initiation-pattern.md index 563439fd..978f0ac0 100644 --- a/docs/design-notes/autonomy-goals-signposts-initiation.md +++ b/docs/architecture/practice/ADR-601-autonomy-as-a-layered-system-goals-signposts-and-the-initiation-pattern.md @@ -1,7 +1,16 @@ -# Autonomy as a Layered System: Goals, Signposts, and the Initiation Pattern +--- +contract: adr/v1 +kind: evidence +capability: method +status: accepted +date: 2026-06-21 +deciders: + - aaronsb +related: [] +--- + +# ADR-601: Autonomy as a layered system: goals, signposts, and the initiation pattern -> **Type:** Design note (not an ADR) -> **Status:** Working draft, subject to revision > **Cites:** ADR-138 > **Motivates:** a future signpost-convention way, a workflow way, the deliver workflow, and continuance guidance diff --git a/docs/design-notes/ways-functional-audit.md b/docs/architecture/practice/ADR-602-ways-functional-audit-117-ways-against-the-firing-contract.md similarity index 98% rename from docs/design-notes/ways-functional-audit.md rename to docs/architecture/practice/ADR-602-ways-functional-audit-117-ways-against-the-firing-contract.md index 52dab09a..f984e300 100644 --- a/docs/design-notes/ways-functional-audit.md +++ b/docs/architecture/practice/ADR-602-ways-functional-audit-117-ways-against-the-firing-contract.md @@ -1,4 +1,15 @@ -# Ways Functional Audit — 117 ways +--- +contract: adr/v1 +kind: evidence +capability: authoring +status: accepted +date: 2026-07-05 +deciders: + - aaronsb +related: [] +--- + +# ADR-602: Ways functional audit: 117 ways against the firing contract Assessed all 117 frontmatter ways against the functional firing contract (sonnet fan-out, 24 batches). diff --git a/docs/design-notes/cypress-survey.md b/docs/architecture/practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md similarity index 98% rename from docs/design-notes/cypress-survey.md rename to docs/architecture/practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md index 9ef6433b..3d7f449d 100644 --- a/docs/design-notes/cypress-survey.md +++ b/docs/architecture/practice/ADR-603-cypress-survey-what-a-node-routed-seed-teaches-a-hook-disclosed-corpus.md @@ -1,4 +1,15 @@ -# Cypress Survey: What a Node-Routed Seed Teaches a Hook-Disclosed Corpus +--- +contract: adr/v1 +kind: evidence +capability: method +status: accepted +date: 2026-09-09 +deciders: + - aaronsb +related: [] +--- + +# ADR-603: Cypress survey: what a node-routed seed teaches a hook-disclosed corpus A reading of [CYPRESS](https://github.com/llopresto87/Cypress) (Luigi Lopresto, MIT) against the ways corpus, taken 2026-09-09 at its 7.x line. The survey covers the method surface: postures, protocols, skills, delegation briefs, agents, and tooling. The harvested corpora (library, legal, tool, agent, skill) are project residue and were skipped. diff --git a/docs/architecture/ways/ADR-004-way-macros.md b/docs/architecture/ways/ADR-004-way-macros.md new file mode 100644 index 00000000..d34511e4 --- /dev/null +++ b/docs/architecture/ways/ADR-004-way-macros.md @@ -0,0 +1,289 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - authoring +basis: + - evidence: static way text cannot adapt to the environment it lands in, e.g. solo vs team repo, which SSH tools are installed, AWS account and region (Context, The Limitation) +agent: + name: Claude + model: unrecorded +status: accepted +date: 2025-12-30 +deciders: + - aaronsb + - claude +related: [] +imported: + from: docs/architecture/legacy/ADR-004-way-macros.md + format: v0 + status: Accepted +--- + +# ADR-004: Way Macros for Dynamic Context Injection + +## Context + +### The Ways System Philosophy + +The ways system is a **domain-agnostic guidance framework**. It provides automated, consistent guidance triggered by keywords, commands, and file patterns. While this repository ships with software development ways (GitHub, commits, testing), the mechanism itself is general-purpose. + +A different user might have ways for: +- Excel/Office productivity (formulas, pivot tables, VBA macros) +- AWS operations (EC2, S3, IAM policies) +- Financial analysis (portfolios, tax lots, rebalancing) +- Research workflows (citations, data collection, peer review) + +The ways system doesn't care about the domain—it just matches triggers and injects guidance. + +### The Limitation + +Static markdown guidance can't adapt to the user's actual environment. The way says "how to do X" but doesn't know "what X looks like right now." + +**Examples across domains:** + +| Domain | Way | What static misses | +|--------|-----|-------------------| +| Software dev | GitHub | Is this solo or team? Who are reviewers? | +| Software dev | SSH | Which tools are available? sshpass? ssh-agent? | +| AWS ops | IAM | Which account/region? Prod or dev? | +| Finance | Trading | Market hours? Account type? | +| Office | Excel | Which version? What add-ins? | + +### The Insight + +**Ways = guidance (the "how")** +**Macros = state detection (the "what is")** +**Combined = contextual guidance (the "how, given what is")** + +Macros provide optional, domain-specific state detection that contextualizes way guidance for the user's actual environment. + +## Decision + +Extend the ways system to support **way macros**—shell scripts that generate dynamic context injected alongside static way content. + +### Mechanism + +1. **Frontmatter control**: New `macro:` field in way markdown + ```yaml + --- + keywords: github|pull.?request + macro: prepend # Run {wayname}.macro.sh, prepend output + --- + ``` + +2. **Macro file convention**: `{wayname}.macro.sh` alongside `{wayname}.md` + ``` + hooks/ways/ + ├── github.md # Static guidance + ├── github.macro.sh # Dynamic state detection + ├── ssh.md + ├── ssh.macro.sh + ├── commits.md # No macro = static only + ``` + +3. **Position options**: + - `macro: prepend` - State context before guidance + - `macro: append` - State context after guidance + - No `macro:` field = current behavior (static only) + +4. **Execution**: `show-way.sh` checks for macro field, runs script if present, combines output + +### Macro Contract + +Macros must: +- Be executable shell scripts +- Output markdown to stdout +- Handle missing tools gracefully (degrade, don't fail) +- Not require user interaction + +### Execution Model + +- **Once per session**: Macro runs when way triggers, output cached with way marker +- **Coupled to way**: If way doesn't fire (already shown), macro doesn't run +- **No runtime babysitting**: We don't enforce timeouts or output limits at runtime + +### Macro Author Responsibilities + +Authors are responsible for writing macros that: +- Don't infinite loop (internal code is trusted) +- Keep output reasonable (aim for < 20 lines) +- Timeout external calls appropriately +- Provide helpful context on failure, not silent exit + +### Recommended Patterns + +```bash +#!/bin/bash +# Pattern 1: Early exit if precondition not met +gh repo view &>/dev/null || { + echo "**Note**: Not a GitHub repository" + exit 0 +} + +# Pattern 2: Timeout external calls +RESULT=$(timeout 2 some_api_call 2>/dev/null) +[[ -z "$RESULT" ]] && { + echo "**Note**: Could not reach API" + exit 0 +} + +# Pattern 3: Trap and report errors with context +if ! DATA=$(some_command 2>&1); then + echo "**Note**: Could not acquire info - $DATA" + exit 0 +fi + +# Pattern 4: Degrade gracefully based on available tools +if command -v some_tool &>/dev/null; then + echo "- some_tool available" +else + echo "- some_tool not installed" +fi +``` + +### Output Guidelines + +- No top-level headers (`#` or `##`) - the way provides structure +- Use bold for key info: `**Context**: Solo project` +- Keep concise - macro adds state, way provides guidance +- Always exit 0 (non-zero reserved for future error signaling) + +### Example Macros by Domain + +**Software dev (github.macro.sh)**: +```bash +#!/bin/bash +gh repo view &>/dev/null || { echo "**Note**: Not a GitHub repository"; exit 0; } + +CONTRIBUTORS=$(timeout 2 gh api repos/:owner/:repo/contributors --jq 'length' 2>/dev/null || echo "0") + +if [[ "$CONTRIBUTORS" -le 2 ]]; then + echo "**Context**: Solo/pair project - PR optional, direct merge acceptable" +else + echo "**Context**: Team project ($CONTRIBUTORS contributors) - PR recommended" +fi +``` + +**AWS ops (iam.macro.sh)**: +```bash +#!/bin/bash +ACCOUNT=$(timeout 2 aws sts get-caller-identity --query Account --output text 2>/dev/null) +[[ -z "$ACCOUNT" ]] && { echo "**Note**: AWS credentials not configured"; exit 0; } + +REGION=$(aws configure get region) +echo "**Context**: Account $ACCOUNT, Region $REGION" + +if [[ "$ACCOUNT" == "123456789" ]]; then + echo "- ⚠️ This is PRODUCTION" +fi +``` + +**Office (excel.macro.sh)**: +```bash +#!/bin/bash +if command -v xlsx2csv &>/dev/null; then + echo "- xlsx2csv available for data extraction" +fi +if [[ -f "$FILE" && "$FILE" =~ \.xlsm$ ]]; then + echo "- ⚠️ Macro-enabled workbook - VBA content present" +fi +``` + +### Testing + +Add `hooks/ways/tests/` directory with validation: +- `test-frontmatter.sh` - Validate all ways have valid frontmatter +- `test-macros.sh` - Validate macros are executable and exit cleanly +- `test-triggers.sh` - Validate keyword/command patterns + +### Project-Local Macros + +Same precedence rules as ways: +- Project-local macro (`$PROJECT/.claude/ways/foo.macro.sh`) shadows global +- If project-local way exists but no project-local macro, global macro does NOT run +- Macro is coupled to its way - they travel together + +``` +Lookup order: +1. $PROJECT/.claude/ways/foo.md + foo.macro.sh (if exists) +2. ~/.claude/hooks/ways/foo.md + foo.macro.sh (if exists) + +No mixing: project-local way with global macro is not supported. +``` + +## Consequences + +### Positive +- Ways become environment-aware without losing domain-agnosticism +- Guidance adapts to actual state (solo vs team, prod vs dev, tools available) +- Framework remains simple: bash + jq, no new dependencies +- Maintains backward compatibility (no macro = current behavior) +- Users can create macros for any domain, not just software dev + +### Negative +- Additional complexity in show-way.sh +- Macros add execution time (mitigated by once-per-session caching) +- More files to maintain per way + +### Neutral +- Macros are optional - ways work without them +- Testing infrastructure needed +- Documentation for macro authors needed +- **Security**: Macros execute arbitrary shell code. Project-local macros are trusted by convention (same as project-local ways). Users cloning untrusted repos should review `.claude/ways/` contents. + +## Alternatives Considered + +### 1. Hook-based injection (instead of macros) + +Use existing `UserPromptSubmit` hook to run detection scripts and prepend context to prompts. + +**Why rejected**: Hooks are event-driven, not way-coupled. Would require duplicating trigger logic. Macros are semantically "part of a way"—the state detection is specific to that way's domain. + +### 2. Template engine in way files (Jinja-style) + +Embed logic directly in way markdown: +```markdown +{% if contributors <= 2 %} +**Context**: Solo project +{% endif %} +``` + +**Why rejected**: Requires a template engine dependency. Bash is already available and more powerful. Template syntax is limiting for real environment detection. Violates "bash + jq only" philosophy. + +### 3. Sandboxed scripting (Lua, WASM) + +Use a sandboxed language that's safer and more portable than shell. + +**Why rejected**: Adds significant complexity and dependencies. Macro authors are trusted (same as way authors). The security boundary is at the project level, not the macro level. + +### 4. Embed logic in frontmatter + +Extended frontmatter with conditional fields: +```yaml +--- +keywords: github +context_if: gh repo view +context_text: "GitHub repo detected" +--- +``` + +**Why rejected**: Too limited. Can't do contributor counting, tool detection, or complex state queries. Frontmatter should declare triggers, not implement behavior. + +### 5. Always-run convention (no frontmatter opt-in) + +If `foo.macro.sh` exists alongside `foo.md`, always run it. + +**Why rejected**: Less control. Some ways may want the macro file present but conditionally disabled. Explicit `macro:` frontmatter is clearer and allows prepend/append control. + +## Implementation Plan + +1. Update `show-way.sh` to parse `macro:` frontmatter +2. Implement prepend/append logic with output combination +3. Create `github.macro.sh` as proof of concept +4. Create `ssh.macro.sh` to demonstrate tool detection pattern +5. Add `hooks/ways/tests/` validation suite +6. Update `knowledge.md` way with macro authoring documentation +7. Update README with macro documentation for other-domain users diff --git a/docs/architecture/ways/ADR-014-tfidf-semantic-matcher.md b/docs/architecture/ways/ADR-014-tfidf-semantic-matcher.md new file mode 100644 index 00000000..8ca244e1 --- /dev/null +++ b/docs/architecture/ways/ADR-014-tfidf-semantic-matcher.md @@ -0,0 +1,269 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'gzip NCD is surface-level: ''optimize database queries'' and ''speed up SQL'' share no bytes; per-way NCD thresholds (0.52-0.58) are hand-tuned and length-sensitive' + - evidence: Anthropic's January 2026 restriction of programmatic claude -p subscription use makes model-match.sh an unreliable foundation (References) + - standard: BM25 with the standard literature defaults k1 = 1.2, b = 0.75 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-02-16 +deciders: + - aaronsb + - claude +related: + - ADR-013 +imported: + from: docs/architecture/legacy/ADR-014-tfidf-semantic-matcher.md + format: v0 + status: Accepted +--- + +# ADR-014: TF-IDF/BM25 Binary for Semantic Way Matching + +## Context + +Ways use three matching modes: regex, semantic (gzip NCD), and state triggers. A fourth mode — model-based matching via `claude -p` subprocess (`model-match.sh`) — is wired up but unused, and increasingly fragile due to Anthropic's January 2026 crackdown on third-party tools invoking Claude subscriptions programmatically (see [References](#references)). + +The semantic mode (`semantic-match.sh`) combines two techniques: + +1. **Keyword counting** — count how many vocabulary words appear in the prompt (match if >= 2) +2. **Gzip NCD** — information-theoretic similarity via compression ratio + +This works but has known weaknesses: + +- **Gzip NCD is surface-level**: it detects shared byte patterns, not shared meaning. "optimize database queries" and "speed up SQL" share no bytes but are semantically close. NCD misses this. +- **Threshold tuning is fragile**: each way needs a hand-tuned NCD threshold (0.52–0.58), and the score is sensitive to prompt length — a long prompt dilutes the signal from a short description. +- **Keyword counting is brittle**: requires manually curated vocabulary lists per way. Misses synonyms, abbreviations, and natural phrasing. +- **No term weighting**: "test" in a prompt about "testing frameworks" and "test" in "put it to the test" are treated identically. No concept of term importance. + +Seven ways currently use semantic matching (api, config, debugging, design, security, testing, adr-context). As the corpus grows, false positives and missed matches will increase. + +TF-IDF and BM25 address these issues with term-frequency weighting and inverse-document-frequency discrimination — without requiring any ML model, GPU, or external service. + +## Decision + +Build a single portable binary (`way-match`) using Cosmopolitan Libc (APE format) that: + +1. Accepts way descriptions + vocabulary as a **corpus** and a **prompt** as query +2. Computes **BM25 relevance scores** between the prompt and each way +3. Returns ranked matches above a configurable threshold +4. Runs on Linux (amd64/arm64), macOS (amd64/arm64), and Windows (amd64) from a single binary + +### Interface + +``` +# Batch mode: score prompt against all ways at once +way-match score --corpus ways.jsonl --query "optimize my database queries" + +# Output: ranked matches, one per line +# way_id<TAB>score<TAB>description_snippet +testing 0.12 writing unit tests... +debugging 0.31 debugging code issues... +design 0.87 software system design... + +# Single pair mode (drop-in for semantic-match.sh) +way-match pair --description "software system design..." \ + --vocabulary "architecture pattern database..." \ + --query "optimize my database queries" \ + --threshold 0.4 + +# Exit 0 if match, 1 if not (compatible with current script interface) +``` + +### Corpus file format + +```jsonl +{"id":"design","description":"software system design architecture patterns...","vocabulary":"architecture pattern database schema..."} +{"id":"testing","description":"writing unit tests, test coverage...","vocabulary":"unittest coverage mock tdd..."} +``` + +Generated at install time or on first run by scanning way frontmatter. Cached and regenerated when way files change (mtime check). + +### BM25 parameters + +| Parameter | Default | Notes | +|-----------|---------|-------| +| k1 | 1.2 | Term frequency saturation | +| b | 0.75 | Document length normalization | +| threshold | 0.4 | Minimum score to report (tunable per-way via frontmatter override) | + +These are the standard BM25 defaults from the literature. The threshold replaces the current NCD threshold and can be carried forward in way frontmatter. + +### IDF computation + +IDF is computed over the way corpus itself (currently ~33 documents). This is a small corpus, but the IDF still provides signal: "test" appears in many way descriptions (low IDF), "owasp" appears in one (high IDF). Terms from the vocabulary field are indexed alongside the description to enrich the document representation. + +The query (user prompt) is tokenized and scored against each way's combined description+vocabulary document. + +### Integration with existing matching + +``` +check-prompt.sh + ├── regex ways → pattern match (unchanged) + ├── semantic ways → way-match binary (replaces semantic-match.sh) + │ falls back to gzip NCD if binary absent + └── state ways → check-state.sh (unchanged) +``` + +The `semantic-match.sh` script is replaced by a call to `way-match pair` with the same interface contract (exit 0/1). The `match: semantic` frontmatter field continues to work unchanged. The `threshold:` field is reinterpreted as a BM25 threshold instead of NCD threshold (values will differ and need one-time recalibration). + +The `match: model` mode (`model-match.sh`) remains in the codebase but is not invested in further. No ways use it today, and Anthropic's tightening of `claude -p` subprocess usage makes it an unreliable foundation. If LLM-level matching is needed in the future, a local embedding approach (see Alternatives) is a more sustainable path than depending on subscription-gated CLI invocations. + +Batch mode (`way-match score --corpus`) is available for future optimization: score all semantic ways in one invocation instead of N separate calls. + +### Repo organization + +Source and binary live in the claude-code-config repo (not a submodule). The tool has exactly one consumer — the ways matching system — and is ~500-1000 lines of C. A separate repo adds git submodule overhead with no real benefit for a single-file, single-purpose tool. + +``` +tools/ + way-match/ + way-match.c # BM25 implementation (~500-1000 lines) + Makefile # cosmocc build targets +bin/ + way-match # Built APE fat binary (checked in, ~100-300KB) +``` + +`.gitignore` additions: +```gitignore +# Tools (source + build) +!tools/ +!tools/way-match/ +!tools/way-match/* + +# Built binaries (checked in, cross-platform APE) +!bin/ +!bin/way-match +``` + +### Build and distribution + +- Source: C, ~500-1000 lines estimated +- Compiler: `cosmocc` (Cosmopolitan toolchain — bundles its own gcc/clang, no system compiler needed) +- Build deps: cosmocc only (download + unzip, ~60MB). Only needed to rebuild — not to use. +- Output: single APE binary, estimated ~100-300KB +- No dynamic linking, no runtime dependencies +- Checked into repo at `bin/way-match` (fat binary covers linux/macos/windows, amd64/arm64) +- Fallback: if binary missing or fails, fall back to `semantic-match.sh` (gzip NCD) + +```makefile +# Makefile sketch +COSMOCC ?= $(HOME)/.cosmocc/bin/cosmocc + +bin/way-match: tools/way-match/way-match.c + $(COSMOCC) -O2 -o $@ $< + +verify: tools/way-match/way-match.c + $(COSMOCC) -O2 -o /tmp/way-match-verify $< + @if cmp -s bin/way-match /tmp/way-match-verify; then \ + echo "PASS: binary matches source"; \ + else \ + echo "MISMATCH: checked-in binary differs from source build"; \ + echo " checked-in: $$(sha256sum bin/way-match)"; \ + echo " from source: $$(sha256sum /tmp/way-match-verify)"; \ + fi + @rm -f /tmp/way-match-verify + +clean: + rm -f bin/way-match +``` + +### Trust and verification + +The binary is checked in for convenience, but **users should never need to trust it blindly**. The trust model: + +1. **Source is adjacent**: `tools/way-match/way-match.c` is a single readable C file, ~500-1000 lines, no obfuscation, no vendored blobs +2. **Build is trivial**: one `cosmocc` invocation, no configure step, no fetched dependencies +3. **`make verify`**: builds from source and compares against the checked-in binary +4. **`make bin/way-match`**: users can always build their own and ignore the checked-in copy +5. **CI can enforce**: a GitHub Action can run `make verify` on PRs that touch `bin/way-match` to ensure the binary matches the source + +The checked-in binary is a convenience for users who don't want to install cosmocc. It is never the only option. Anyone uncomfortable with it can build from source in under 10 seconds. + +### Trust chain principle + +The ways system is deliberately built on tools already present on the machine — bash, gzip, jq, bc. This is a design principle, not an accident. Every layer of the matching system must be auditable and replaceable: + +``` +Trust tier 0 (always available): bash, gzip, wc, bc → NCD fallback +Trust tier 1 (source-auditable): way-match binary → BM25 scoring +Trust tier 2 (external service): model-match.sh → LLM classification +``` + +The fallback path (`semantic-match.sh`) runs on POSIX builtins — it would work on a busybox instance. The binary is an upgrade in quality, not a hard dependency. If `bin/way-match` is missing, corrupt, or untrusted, the system degrades to gzip NCD automatically. No way ever fails to match because the binary isn't there. + +### Tokenization + +Simple whitespace + punctuation splitting, lowercased, with the existing stopword list. No stemming in v1 — keeps the implementation minimal and the binary small. Stemming (Porter or similar) is a future enhancement if needed. + +## Consequences + +### Positive +- **Better matching quality**: BM25 handles term importance — rare domain terms score higher than common ones +- **Length-insensitive**: BM25's length normalization handles short descriptions vs long prompts naturally (gzip NCD degrades here) +- **Single threshold semantic**: one scoring model instead of keyword-count OR NCD, reducing dual-path complexity +- **Cross-platform**: one binary for Linux/macOS/Windows, amd64/arm64 +- **Fast**: BM25 over 33 documents is microseconds, well under the current gzip NCD latency (which forks gzip 3x per way) +- **No model, no API, no GPU**: pure computation, fully offline +- **Graceful fallback**: gzip NCD script remains as fallback + +### Negative +- **New build dependency**: Cosmopolitan toolchain for compilation (though output is a static binary) +- **Binary in repo**: ~100-300KB checked-in binary (acceptable for a cross-platform fat binary) +- **Threshold recalibration**: existing per-way thresholds need one-time adjustment from NCD scale to BM25 scale +- **Still not embeddings**: BM25 won't catch "make it faster" → "performance optimization" without shared terms. Vocabulary field partially mitigates this. + +### Neutral +- Batch mode enables future architectural changes (score all ways per prompt in one call) but doesn't require them immediately +- The `vocabulary` field becomes more valuable as BM25 document enrichment rather than a flat keyword list +- The `model` matching mode remains in codebase but is not actively invested in + +## Alternatives Considered + +### Keep gzip NCD, tune thresholds +Rejected: fundamental limitation — NCD measures byte-level redundancy, not term importance. No amount of threshold tuning fixes "optimize SQL" vs "speed up database queries". + +### Embeddings (fastembed, llama.cpp) +Deferred: would require shipping a model file (25-130MB) alongside the binary. Dramatically better semantic understanding but violates the "zero dependencies, tiny binary" constraint. Could be a future ADR-015 if BM25 proves insufficient. + +### Python implementation (scikit-learn TF-IDF) +Rejected: adds Python runtime dependency. The ways system is currently pure bash + coreutils. A compiled binary preserves the "just works" property. + +### TF-IDF instead of BM25 +Rejected: BM25 is strictly better for this use case — it adds term frequency saturation (diminishing returns for repeated terms) and document length normalization, both relevant when matching short descriptions against variable-length prompts. Implementation complexity is nearly identical. + +### WASM binary instead of Cosmopolitan APE +Considered: would require a WASM runtime (wasmtime, wasmer). APE is more self-contained — the binary IS the runtime. + +## Validation + +A test harness will compare BM25 against gzip NCD on the actual way corpus to validate the upgrade. Tests live alongside the source in `tools/way-match/` and cover: + +### Correctness tests +- **True positives**: prompts that should match a specific way do match (e.g., "add unit tests for the auth module" → testing way) +- **True negatives**: prompts that should not match a way don't (e.g., "what's for lunch" → no match) +- **Synonym/paraphrase coverage**: prompts using different words for the same concept (e.g., "speed up SQL" should match the same way as "optimize database queries") + +### Comparative tests (BM25 vs gzip NCD) +- Run both scorers against a shared test fixture of prompt/expected-way pairs +- Report match/miss matrix: cases where BM25 matches but NCD misses (expected wins), and vice versa (regressions to investigate) +- Measure score distributions to calibrate BM25 thresholds against real prompts + +### Performance tests +- Latency comparison: BM25 single invocation vs gzip NCD (which forks gzip 3x per way, ~7 ways = ~21 subprocess spawns) +- Batch mode throughput: score all semantic ways in one invocation + +Test fixtures are a JSONL file of `{"prompt": "...", "expected_way": "...", "should_match": true}` entries, curated from real usage patterns and known NCD failure cases. + +## References + +- [Stop using Claude for OpenClaw and OpenCode](https://generativeai.pub/stop-using-claudes-api-for-moltbot-and-opencode-52f8febd1137) — Jan 2026, context on Anthropic restricting programmatic Claude subscription usage +- [You might be breaking Claude's ToS without knowing it](https://blog.devgenius.io/you-might-be-breaking-claudes-tos-without-knowing-it-228fcecc168c) — Jan 2026, ToS implications of `claude -p` subprocess patterns +- [Please stop using OpenClaw](https://www.xda-developers.com/please-stop-using-openclaw/) — coverage of Anthropic cease-and-desist and account bans +- [Cosmopolitan Libc](https://github.com/jart/cosmopolitan) — Actually Portable Executable toolchain +- [llamafile](https://github.com/Mozilla-Ocho/llamafile) — proof of Cosmopolitan APE at scale (multi-GB ML binaries) diff --git a/docs/architecture/ways/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md b/docs/architecture/ways/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md new file mode 100644 index 00000000..8aefd0d2 --- /dev/null +++ b/docs/architecture/ways/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md @@ -0,0 +1,256 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - matching + - authoring +basis: + - evidence: 'landscape analysis of Claude Code''s extension points (PreToolUse hooks, ways, skills, CLAUDE.md, permissions, agents): none gives decay-modulated, domain-coupled injection at the moment of action' + - evidence: '''four-foot circle'' failures: confident action on interpolated rather than verified knowledge, while a way injected 20 turns earlier has faded' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-13 +deciders: + - aaronsb + - claude +related: + - ADR-004 + - ADR-013 + - ADR-014 +imported: + from: docs/architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md + format: v0 + status: Accepted +--- + +# ADR-103: Checks — Epoch-Distance-Aware Confidence Sensors for Ways + +## Context + +Ways inject contextual guidance when a domain becomes relevant — triggered by keywords, tool use, or state conditions. This guidance fires once per session and decays in attentional influence as the conversation progresses (more tokens accumulate between the injection point and the current generation step, reducing relative positional weight via RoPE decay). + +The problem: when Claude acts on assumptions within a domain, there is no mechanism to verify those assumptions before committing. Ways tell Claude *how to approach* work; nothing tells Claude *what to check* before acting. This produces "four-foot circle" failures — confident action based on interpolated knowledge rather than verified understanding. + +The frozen model problem compounds this: Claude cannot develop calibrated uncertainty through experience. Unlike a human apprentice who learns "I should check before cutting," Claude has uniform confidence across known facts and confabulations. External scaffolding must compensate. + +The existing permission-seeking pattern ("Want me to proceed?") is noise — it asks for *permission* rather than surfacing *assumptions*. The useful question is not "may I act?" but "did I check?" + +### Observations driving this design + +1. **Ways decay.** A way injected 20 turns ago has weak attentional influence on current generation. Re-anchoring is valuable when the context has drifted. + +2. **Each domain has different assumptions to verify.** A generic "check your work" directive is useless. Architecture checks are different from deployment checks are different from security checks. The sensor must be coupled to its domain. + +3. **Repeated nudging has diminishing returns.** A check that fires every turn becomes noise — the same failure mode as permission-seeking. Successive firings must become progressively harder. + +4. **Context budget matters.** Injecting a full re-anchor when the way is still warm wastes tokens. Injecting a light check when the way is cold wastes an opportunity. + +### Landscape analysis: why no existing primitive covers this + +Claude Code's native extension points were evaluated: + +| Primitive | What it does | Why it's insufficient | +|-----------|-------------|----------------------| +| **PreToolUse hooks** | Block, modify input, inject additionalContext | Binary fire/don't-fire — no scoring curve, no distance awareness, no decay | +| **Ways** | Domain-coupled guidance, once per session | Fire-and-forget — no re-anchoring, no verification at action time | +| **Skills** | Intent-driven multi-step workflows | Heavyweight, user/Claude-initiated — not automatic pre-action sensors | +| **CLAUDE.md / rules/** | Static project instructions | Front-loaded, no timing control, maximum positional distance from action | +| **Permissions** | Allow/deny tool access | Access control, not contextual verification | +| **Agents / Subagents** | Delegated personas with tool restrictions | Scoping mechanism, not confidence checking | + +The gap: **no native primitive provides adaptive, domain-coupled, decay-modulated context injection at the moment of action.** PreToolUse hooks provide the infrastructure (we already use them); checks add the scoring model that makes injection context-aware. + +Checks are a **subclass of ways**, not a new primitive category. They share the same ecosystem (directories, frontmatter, matching engine, hook infrastructure) but have different firing semantics. + +## Decision + +Introduce **checks** as a second file class within way directories. A `check.md` sits alongside `way.md` in the same directory, coupled by domain but decoupled by timing and trigger model. + +### File structure + +``` +ways/{domain}/{wayname}/ + way.md # directive: how to approach (fires on domain entry) + check.md # sensor: what to verify (fires before action) +``` + +### Trigger model + +Checks fire on **PreToolUse** (before Edit, Write, Bash) — the moment Claude is about to commit to an action. Ways fire on **UserPromptSubmit** or tool-pattern match — the moment a domain becomes relevant. These are different hook events with natural temporal separation. + +### Epoch counter + +A monotonic counter increments on every hook event within a session. When a way fires, its epoch is stamped. When a check evaluates, the distance from the parent way's epoch is available. + +```bash +# /tmp/.claude-epoch-{session_id} — current epoch (integer) +# /tmp/.claude-way-epoch-{way}-{session} — epoch when way fired +# /tmp/.claude-check-fires-{check}-{session} — check fire count +``` + +### Scoring curve + +Checks use the same BM25/semantic matching as ways, but the raw match score is modulated by two contextual factors: + +``` +effective_score = match_score × distance_factor × decay_factor +``` + +Where: + +- **match_score** — BM25 or gzip NCD score against current tool input / description +- **distance_factor** = `ln(min(epoch_distance, 30) + 1) + 1` — grows sublinearly with distance from parent way, **capped at 30** to prevent score explosion when the way hasn't fired or is very distant. Max multiplier: ~4.4×. +- **decay_factor** = `1 / (fire_count + 1)` — shrinks with each successive check fire in the session. Diminishing returns on repeated nudges. + +The check fires if `effective_score >= threshold` (threshold set in check.md frontmatter, same as ways). + +### Behavioral properties of the curve + +| Scenario | Distance | Fires | Effective multiplier | Behavior | +|----------|----------|-------|---------------------|----------| +| Way just fired | 0 | 0 | 1.0 | Barely fires — way is warm | +| Some work done | 5 | 0 | 2.8 | Fires easily — re-anchor valuable | +| Deep in session | 20 | 0 | 3.3 | Fires — way is cold | +| Deep, nagged twice | 20 | 2 | 1.1 | Barely fires — diminishing returns | +| Deep, nagged 4x | 20 | 4 | 0.66 | Doesn't fire — stops nagging | + +### Check fires before way + +If a check's effective score exceeds threshold but the parent way has *not* fired this session, inject the way alongside the check. The way's context is needed to make the check meaningful. Distance is treated as maximum (way is infinitely cold). + +### check.md format + +```yaml +--- +description: what this check verifies (for semantic matching) +vocabulary: domain terms for matching +threshold: 2.0 +scope: agent +--- +``` + +Body contains two sections, selected by the loader based on epoch distance: + +```markdown +## anchor +[1-2 line re-anchor to parent way's intent — injected when distance is large] + +## check +[the actual verification questions — always injected] +``` + +A distance threshold (e.g., epoch_distance < 5) determines whether the anchor section is included or omitted. + +### Differences from ways + +| Property | way.md | check.md | +|----------|--------|----------| +| Fires per session | Once (idempotent) | Multiple (with decay) | +| Trigger phase | UserPromptSubmit, PreToolUse | PreToolUse only | +| Scoring | match_score vs threshold | match_score × distance × decay vs threshold | +| Purpose | Directive (how to approach) | Sensor (what to verify) | +| State tracked | Fired yes/no (marker file) | Fire count + parent way epoch | + +### Stats and observability + +Check firings are logged to the same `events.jsonl` as way firings, with additional fields: + +```json +{ + "event": "check_fired", + "check": "softwaredev/architecture/design", + "domain": "softwaredev", + "trigger": "semantic", + "epoch": 17, + "way_epoch": 4, + "distance": 13, + "fire_count": 1, + "match_score": 2.4, + "effective_score": 3.1, + "anchored": true, + "scope": "agent", + "project": "/home/aaron/myproject", + "session": "abc-123" +} +``` + +This enables: +- Tracking check fire frequency per domain (are some checks too noisy?) +- Measuring average epoch distance at fire time (is the curve well-calibrated?) +- Comparing anchor vs non-anchor fires (is re-anchoring happening at the right distance?) +- Correlating check fires with session outcomes (do checked sessions produce fewer corrections?) + +The `stats.sh` tool and `/ways-tests` evaluation harness are extended to report on checks alongside ways. + +### Authoring and evaluation + +The `/ways` authoring skill and `/ways-tests` evaluation harness are updated to support checks: + +**Authoring** — The ways scaffolding wizard gains a check template option. When creating a new way, the author can optionally scaffold a paired check.md. Guidance includes: +- Keep checks short (3-5 verification questions) +- Anchor section should be 1-2 lines that semantically bridge to the parent way +- Vocabulary should overlap with but be narrower than the parent way's vocabulary (checks are more specific) +- Threshold tuning: start at parent way's threshold, adjust based on observed fire rate + +**Evaluation** — The ways-tests harness gains check-specific test cases: +- Verify check fires after parent way (temporal ordering) +- Verify decay curve behavior (fire count reduces effective score) +- Verify anchor inclusion at distance (epoch distance > threshold includes anchor) +- Verify check-before-way pulls in parent way +- Measure false positive rate (checks firing on irrelevant tool actions) + +## Consequences + +### Positive + +- Compensates for frozen model's inability to develop calibrated uncertainty +- Domain-coupled verification — each check knows what's relevant for its domain +- Self-limiting via decay — prevents check-as-noise failure mode +- Context-budget-aware — light injection when way is warm, full re-anchor when cold +- No new trigger infrastructure — uses existing PreToolUse hooks +- Empirically tunable — the curve constants can be adjusted based on observed behavior +- Rich observability — epoch distance, fire count, effective score all logged + +### Negative + +- Adds a second file class to the ways system (more to maintain) +- Epoch counter adds one file read/write per hook event +- Scoring curve introduces floating-point math (awk dependency in bash) +- Risk of over-engineering if checks proliferate without discipline +- Authoring and evaluation tooling must be updated + +### Neutral + +- check.md files are optional — ways without checks behave exactly as before +- The epoch counter is useful beyond checks (could inform other distance-aware behaviors) +- Opens the question of whether a third file class will be needed (we explicitly defer this — two classes until proven otherwise) +- Stats collection grows but remains append-only JSONL (same infrastructure) + +## Alternatives Considered + +- **Native PreToolUse hooks alone** — Already in use. Checks build on top of this infrastructure. The gap isn't the hook mechanism but the scoring model — native hooks have no concept of contextual distance or firing decay. + +- **Section within way.md** — Rejected because it prevents separate injection timing. The whole point is that checks fire at a different moment (pre-action) than ways (domain entry). A single file means both inject together, wasting context budget. + +- **Generic "check your assumptions" directive** — Rejected because it's domain-unaware. A generic check is just another verbose system prompt that gets ignored. The value is in domain-specific verification coupled to the domain's way. + +- **Fixed-interval firing (every N epochs)** — Rejected because it's a timer, not a sensor. Checks should fire based on match relevance modulated by context, not on a clock. A timer would fire during irrelevant actions and miss relevant ones. + +- **No decay (fire every match)** — Rejected because repeated nudging becomes permission-seeking noise. The decay factor is what distinguishes checks from the "Want me to proceed?" anti-pattern. + +- **Confidence introspection directive** — Rejected because Claude cannot reliably introspect on confidence. The model's self-reported uncertainty is generated text, not measured signal. Empirical checking (use a sensor / run a test) is more reliable than asking the model to evaluate its own certainty. + +- **Skills-based approach** — Rejected because skills are intent-driven and user/Claude-initiated. Checks must fire automatically at the pre-action moment without anyone remembering to invoke them. Skills also lack the decay model. + +## Implementation Plan + +1. **Framework** — `epoch.sh` (counter), `show-check.sh` (loader with curve), wire into `check-file-pre.sh` and `check-bash-pre.sh` +2. **Prototype check** — `architecture/design/check.md` as first test case +3. **Stats** — Extend `log-event.sh` calls and `stats.sh` reporting for check events +4. **Authoring** — Update `/ways` skill with check template and guidance +5. **Evaluation** — Update `/ways-tests` with check-specific test fixtures and assertions +6. **Observe** — Run for several sessions, tune curve constants based on logged data diff --git a/docs/architecture/ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md b/docs/architecture/ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md new file mode 100644 index 00000000..a92d490f --- /dev/null +++ b/docs/architecture/ways/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md @@ -0,0 +1,140 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: disclosure +superseded_by: ADR-123 +basis: + - evidence: 'long-context benchmarks: Opus 4.6 retrieval (MRCR v2) 91.9% at 256K to 78.3% at 1M; Sonnet 4.6 90.6% to 65.1% (docs/reference/model-context-decay/)' + - precedent: ADR-004 +agent: + name: Claude + model: unrecorded +status: superseded +date: 2026-03-13 +deciders: + - aaronsb + - claude +related: + - ADR-103 + - ADR-004 + - ADR-123 +imported: + from: docs/architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md + format: v0 + status: Superseded + unmapped: + revised: 2026-04-14 +--- + +# ADR-104: Token-Gated Way Re-Disclosure for Long Context Windows + +## Status: Superseded by ADR-123 + +**Superseded 2026-04-14.** The core insight of this ADR — that ways must re-fire on a token-distance axis rather than once-per-session, because retrieval degrades as context accumulates past a way's injection point — is still correct and load-bearing. The specific implementation described below (a global `REDISCLOSE_PCT: u64 = 25` constant, a per-model context-window lookup, a flat step-function that admits re-firing once distance crosses 25% of the window) has been replaced by the per-way curve engine introduced in [ADR-123](ADR-123-firing-dynamics-progression-axis-unification.md). + +The current implementation: each way declares an explicit `curve:` block in its frontmatter, and the engine queries `current_salience(current_tick) < REFIRE_FLOOR` to decide when to re-fire. `REFIRE_FLOOR` defaults to `0.5`, so `Curve::Exponential { half_life: H }` re-fires at delta `H`, `Curve::Flat { suppression: N }` re-fires at delta `N`, and `Curve::ProgressiveStaircase` re-fires on each declared step. The 25% global default is gone; ways pick their own re-fire cadence. + +The rest of this ADR is preserved as historical context — the empirical motivation (retrieval degradation benchmarks), the argument against epoch-based gating, and the alternatives-considered table are all still the reasoning that led to ADR-123's decision. The Decision section below describes *what was decided in 2026-03-13*; the ADR-123 successor describes *what actually runs now*. + +## Context + +Ways initially fired once per session, gated by a marker file (`/tmp/.claude-way-{name}-{session}`). This rule was designed for 200K context windows where the entire conversation fit within a single effective attention span. + +With Opus 4.6's 1M context window, this assumption breaks. Empirical benchmarks show measurable degradation over long contexts: + +- **Retrieval** (MRCR v2): Opus drops from 91.9% at 256K to 78.3% at 1M (~15% degradation) +- **Reasoning** (GraphWalks BFS): Opus drops from 72.8% at 256K to 68.4% at 1M (~6% degradation) +- **Sonnet 4.6** degrades much faster: retrieval 90.6% → 65.1%, reasoning 61.5% → 41.2% + +(See `docs/reference/model-context-decay/` for benchmark charts and data tables.) + +A way disclosed at token 50K is not gone at token 500K — but it's faded. The model can still retrieve the general concept but loses specificity. For guidance that depends on precise rules (security checks, commit conventions, architectural patterns), this degradation produces subtle failures: the model follows the spirit but misses the letter. + +The epoch counter (ADR-103) tracks **event distance** — how many tool actions have occurred since a way fired. This is the right metric for check decay (is the model still thinking about this domain?). But it's the wrong metric for re-disclosure (has the way faded from retrievable memory?). A session can have 200 epoch events in 50K tokens, or 10 epoch events in 500K tokens. Token distance is the signal that correlates with measured retrieval degradation. + +## Decision (as of 2026-03-13) + +Replace the hard "once per session" marker with a **token-distance-gated re-eligibility window**. A way becomes eligible for re-disclosure when the token distance since its last disclosure exceeds a model-specific threshold. + +### Token distance tracking + +When a way fires, stamp the current token position alongside the existing epoch stamp. Token position is read from the transcript using the same method as `context-usage.sh` — sum of `cache_read_input_tokens + cache_creation_input_tokens + input_tokens` from the most recent API usage record. + +### Re-disclosure thresholds + +**Percentage-based, not fixed token counts.** Re-disclosure fires when a way has drifted 25% of the context window since its last disclosure. This scales automatically with the model's context size: + +| Model | Context window | 25% interval | Max re-disclosures | +|-------|---------------|-------------|-------------------| +| Opus 4.6 | 1M | 250K tokens | ~3-4 per session | +| Sonnet 4.6 | 200K | 50K tokens | ~3 per session | +| Haiku 4.5 | 200K | 50K tokens | ~3 per session | + +The 25% figure corresponds to the empirical degradation curves: retrieval accuracy drops ~10-15% per quarter-window. + +Using percentages meant the system automatically adapted when Anthropic shipped new context tiers — no hardcoded token counts to update. + +### Re-disclosure behavior + +When a way re-discloses: + +1. The way content is injected again (same as first disclosure) +2. The token stamp is updated to the current position +3. The epoch stamp is updated (checks reset their distance) +4. A `way_redisclosed` event is logged with the token distance that triggered it +5. The fire count is incremented (for stats, not for gating) + +### What re-disclosure is NOT + +- **Not a timer.** It doesn't fire every N tokens regardless. The way must still be triggered by a matching prompt or tool action. Token distance only makes it *eligible* — it still needs a trigger to fire. +- **Not a check.** Checks (ADR-103) are pre-action verification sensors with decay curves. Re-disclosure is a periodic refresh of the full way guidance. +- **Not visible to the user.** The user sees the same way content; they don't know it's a re-disclosure vs first disclosure. + +## What ADR-123 changed (2026-04-14) + +- **`REDISCLOSE_PCT` constant is gone.** `tools/ways-cli/src/session.rs` no longer hard-codes a 25% threshold. Each way declares its own curve and the engine computes per-way re-fire distances from `Curve::refire_delta(REFIRE_FLOOR)`. +- **`token_distance_exceeded()` is gone.** Replaced by `session::way_fire_outcome(way_id, session_id, curve) -> FireOutcome` which returns `FirstFire | ReFire | Suppressed`. Same semantic role (gatekeeper for firing), different mechanics. +- **Model-specific window lookup is gone from the firing path.** The engine doesn't care what the context window is; it only knows ticks (token positions) supplied by the caller. Context-window detection still exists in `session.rs` for the visualization path (`ways list`, `ways rethink`), as a fallback when a way's frontmatter is missing or unparsable. +- **Marker file shape changed.** The old `.value` stamp files in `way-tokens/` are still written for legacy callers (tree-metrics, scan-time "has this way been seen") but the canonical firing state now lives at `{session_dir}/way-engagement/{way_id}.json` as a serialized `EngagementState`. +- **25% was the wrong unit of analysis.** It treated all ways the same — a one-size-fits-all cadence. ADR-123 lets each way express its own tempo: a quality way fires on a short half-life because file size grows quickly between fires; an architecture way fires on a long half-life because design decisions persist. Per-way curves capture this directly. + +The empirical motivation (retrieval degradation, the MRCR v2 benchmarks, the argument against epoch gating) is **unchanged and still correct**. ADR-104's contribution is the insight that token distance is the right axis; ADR-123's contribution is making the curve shape on that axis a per-way decision instead of a global constant. + +## Consequences + +### Positive (preserved) + +- Compensates for empirically measured retrieval degradation over long contexts +- Maintains the trigger requirement — ways only re-disclose when the domain is relevant +- Resets check distance — prevents stale checks from nagging when the way is freshly re-anchored +- Low token cost per re-disclosure +- Invisible to the model — no behavioral change needed from the model's perspective + +### Positive (added by ADR-123) + +- Each way declares its own re-fire cadence — no single-heuristic calibration +- The same engine drives attend's inward-gate refractory, eliminating two divergent implementations +- The curve shape is a first-class parameter, enabling progressive-disclosure staircases and other non-exponential shapes + +### Negative (resolved by ADR-123) + +- ~~Adds token position reading to the hot path (one jq call per way evaluation)~~ — still present, but the cost has been minimal in practice +- ~~Model detection adds complexity to show-way.sh~~ — the shell dispatchers are gone; all firing logic is in the Rust `ways` binary +- ~~Thresholds are empirically derived but not session-specific~~ — resolved by per-way curves + +## Alternatives Considered (2026-03-13) + +- **Fixed epoch-based re-disclosure (every N events)** — Rejected because epoch count doesn't correlate with retrieval degradation. 100 quick edits in the same file consume fewer tokens than 10 complex prompts with tool chains. Token distance is the right signal. +- **Percentage-of-window triggers (at 25%, 50%, 75% absolute positions)** — Simpler but less nuanced. Doesn't account for when the way was first disclosed. +- **Always re-disclose (remove the once-per-session gate entirely)** — Wasteful; re-disclosing the same way 50 times in 10 minutes adds noise. +- **Decay the existing check system to handle re-anchoring** — Checks inject a short re-anchor (1-2 lines); re-disclosure injects full way content (~200-500 tokens). Using checks for re-disclosure would require making them much longer, defeating their "light sensor" design. +- **Let the user decide (manual re-disclosure command)** — Users shouldn't have to manage context decay. If the user has to remember "my security way has probably faded," the system has failed. + +## References + +- **[ADR-123](ADR-123-firing-dynamics-progression-axis-unification.md)** — the progression-axis unification that made the per-way curve shape first-class. +- **ADR-103** — check scoring via epoch distance; unchanged. +- **ADR-004** — the original once-per-session marker design that ADR-104 replaced. +- `docs/reference/model-context-decay/` — empirical retention benchmarks. +- `docs/hooks-and-ways/context-decay.md` — the presentation-economics model that explains why token-distance gating works. diff --git a/docs/architecture/ways/ADR-105-progressive-disclosure-for-way-trees.md b/docs/architecture/ways/ADR-105-progressive-disclosure-for-way-trees.md new file mode 100644 index 00000000..c4ef9cbe --- /dev/null +++ b/docs/architecture/ways/ADR-105-progressive-disclosure-for-way-trees.md @@ -0,0 +1,91 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - authoring + - matching +basis: + - evidence: a 173-line docs way injected everything on one trigger; overlapping vocabulary degraded BM25 discrimination across 48+ ways + - evidence: the supplychain tree (11 files, 3 levels, thresholds 1.8 to 2.5, sibling Jaccard < 0.06) showed the pattern; integration test accuracy rose from 87% to 94% after the refactor + - evidence: comparison with obra/superpowers, whose skills lazy-load through the Skill tool +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-17 +deciders: + - aaronsb + - claude +related: + - ADR-014 + - ADR-103 + - ADR-104 +imported: + from: docs/architecture/system/ADR-105-progressive-disclosure-for-way-trees.md + format: v0 + status: Accepted +--- + +# ADR-105: Progressive Disclosure for Way Trees + +## Context + +As the ways corpus grows (48+ ways at time of writing), flat ways create two problems: + +1. **Token bloat**: A 173-line docs way dumps everything on a single trigger, regardless of whether the user needs README guidance, docstring conventions, or Mermaid styling. +2. **Vocabulary crowding**: As more ways share overlapping vocabulary, BM25 discrimination degrades. A broad way with 20 vocabulary terms competes with everything. + +Comparison with [obra/superpowers](https://github.com/obra/superpowers) revealed that their skill system uses lazy-loading via the Skill tool, while our hook-based system loads eagerly on match. We need an equivalent of lazy-loading within the hooks architecture. + +The supply chain tree (`softwaredev/code/supplychain/`) already demonstrated the pattern organically — 11 files across 3 depth levels with threshold progression from 1.8 to 2.5 and zero vocabulary overlap between siblings. + +## Decision + +Adopt **tree-structured ways with threshold progression** for complex domains. Specifically: + +**Threshold convention**: Root ways use threshold 1.8 (broad catch), mid-tier 2.0 (focused), leaves 2.5 (narrow specialist). This ensures the root fires on general mentions and children only fire on specific sub-topics. + +**Vocabulary isolation**: Sibling ways must have Jaccard similarity < 0.15. Each child owns its own keyword space. The `way-tree-analyze.sh` tool validates this at authoring time. + +**Parent-aware threshold lowering**: When a parent way's marker exists, children's BM25 thresholds are reduced by 20%. This is evidence-based: if the parent domain is active, children should fire more easily. + +**Tree disclosure tracking**: `show-way.sh` records disclosure events to a JSONL metrics file per session, capturing: parent way, depth, epoch distance from parent, and sibling coverage. This data feeds back into authoring decisions (never-fire children need vocabulary broadening). + +**Anti-rationalization patterns**: High-stakes leaf ways include "Common Rationalizations" tables that counter the agent's tendency to skip steps. Placed in leaves (not roots) so they only appear when the agent is actively doing the thing it might skip. + +**Token budget targets**: Realistic single path ~1200 tokens, worst-case full tree ~4000 tokens. Validated by `/ways-tests budget`. + +**When NOT to tree**: Ways under 80 lines with a single cohesive concern stay flat (errors, performance, debugging, config, commits). + +## Consequences + +### Positive + +- Token efficiency: typical session injects ~1200 tokens of relevant guidance instead of ~600 tokens of everything +- Matching precision: narrow child vocabularies reduce false positives and vocabulary crowding +- Anti-rationalization: high-stakes ways actively counter shortcuts the agent might take +- Observability: disclosure metrics reveal which children never fire (vocabulary gaps) and which cascade instantly (vocabulary overlap) + +### Negative + +- More files to maintain (docs: 1→4, testing: 1→3, security: 1→4) +- Authors must understand threshold progression and vocabulary isolation conventions +- Parent-aware threshold lowering adds a marker check per way in the scan loop (~0.1ms each) + +### Neutral + +- Think strategies (stateful cognitive scaffolding) follow the same pattern but with explicit stage advancement rather than independent child activation +- The `way-tree-analyze.sh` tool and `/ways-tests tree|budget|crowding|metrics` commands make tree health observable +- Integration test accuracy improved from 87% to 94% after refactoring + +## Alternatives Considered + +- **Skill-based lazy loading** (obra/superpowers approach): Skills load via the Skill tool on-demand. We considered this but it requires the agent to self-select, which is unreliable. Hook-based triggering is more deterministic. +- **Flat ways with longer vocabulary**: Keep single files but expand vocabulary to cover sub-topics. Rejected because it worsens vocabulary crowding — more terms per way means more overlap between ways. +- **CLAUDE.md includes**: Load sub-content via CLAUDE.md references. Rejected because it's spatial coupling (position in file) rather than temporal coupling (when the agent needs it). + +## Reference Implementation + +`hooks/ways/softwaredev/code/supplychain/` — 11 files, 3 depth levels, threshold 1.8→2.0→2.5, all sibling Jaccard < 0.06, worst-case ~3840 tokens, average path ~940 tokens. diff --git a/docs/architecture/ways/ADR-107-way-match-corpus-batch-mode-and-locale-support.md b/docs/architecture/ways/ADR-107-way-match-corpus-batch-mode-and-locale-support.md new file mode 100644 index 00000000..02b579ba --- /dev/null +++ b/docs/architecture/ways/ADR-107-way-match-corpus-batch-mode-and-locale-support.md @@ -0,0 +1,233 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - matching + - disclosure + - config +superseded_by: + - ADR-125 +basis: + - evidence: a Japanese prompt produces zero BM25 tokens and low similarity under the English-only embedding model + - evidence: native-language stubs outperform cross-language matching, e.g. ja 0.93 vs 0.69, ar 0.96 vs 0.40 (evaluation table) + - precedent: ADR-111 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-02 +deciders: + - aaronsb + - claude +related: + - ADR-108 + - ADR-110 + - ADR-111 + - ADR-125 +imported: + from: docs/architecture/system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md + format: v0 + status: Accepted +--- + +# ADR-107: Corpus, Matching Pipeline, and Locale Support + +> **Note (2026-04-17):** The matching pipeline architecture described here (two-tier embedding → BM25 → keyword/regex) has been superseded by [ADR-125: Authored Disclosure Graph and Removal of BM25](ADR-125-authored-disclosure-graph-and-removal-of-bm25.md). BM25 is removed; the embedding model is the sole retrieval tier. Locale stubs remain as specified (packed `.locales.jsonl`, one file per way), but the per-locale `embed_threshold` field is removed — thresholds are per-node, in English frontmatter. Sections below describing the two-tier pipeline, BM25 stemmer selection, and the `bm25_stemmer` field in `languages.json` are historical; refer to ADR-125 for the current model. + +## Context + +This ADR was originally drafted when the matching system was a C binary (`way-match`) called N times per prompt by shell scanners. Since then, ADR-111 consolidated everything into a single Rust binary (`ways`). This rewrite reflects the shipped architecture and defines the locale support plan within it. + +### What shipped (Phases 1 & 2) + +**External corpus** — `ways corpus` generates `ways-corpus.jsonl` in `~/.cache/claude-ways/user/`. The corpus is a cache artifact, regenerated on demand, read-only at runtime. IDF is computed across the full corpus (85+ ways), not a hardcoded seed. + +**Batch scoring** — `ways scan prompt --query "..." --session ID` scores all ways in one call. The Rust binary loads the corpus once, tokenizes the query once, and scores every way. Scanner hooks call this instead of N separate invocations. + +**Matching pipeline (historical)** — ADR-108 added embedding (all-MiniLM-L6-v2, 98% accuracy). The original shipped pipeline was embedding → BM25 fallback → keyword/regex patterns with automatic engine selection by model availability. ADR-125 removed BM25; the current pipeline is embedding-only, with explicit `pattern:` / `commands:` regex triggers as a separate override surface. + +**Security boundary preserved** — runtime scanners never write to `~/.claude/`. The corpus and embedding model live in `~/.cache/claude-ways/user/` (XDG cache). Regeneration is an explicit authoring operation (`ways corpus`, `make setup`). + +### What remains: locale support + +The matching pipeline is English-only: +- BM25 stemmer: hardcoded `Algorithm::English` in `bm25.rs:175` +- Stopwords: English-only array in `bm25.rs:12-21` +- Embedding model: `all-MiniLM-L6-v2` is English-only +- Way content: all 85+ ways are written in English + +Claude Code has a `language` setting. Users in non-English locales type prompts in their language, but matching operates on English vocabulary. The gap: a Japanese user's prompt produces zero BM25 tokens (no whitespace boundaries) and low embedding similarity (English-only model). + +## Decision + +### Language resolution + +A new `agents/` module provides a resolution cascade for output language: + +1. `ways.json` `output_language` — explicit user override +2. Claude Code `settings.json` `language` — agent-level config (project then user scope) +3. System locale (`$LC_ALL` → `$LC_MESSAGES` → `$LANG`) — parsed from locale strings like `ja_JP.UTF-8` +4. Default: `en` + +The `agents/` module defines an `AgentConfig` trait. `claude_code.rs` implements it for Claude Code. This is the abstraction point for supporting other CLI agents — each gets its own module implementing the same trait. + +The resolved language affects: +- **Output directive**: `core.md`'s "must be in English" line is substituted at render time with the configured language. The agent writes commit messages, comments, and docs in the user's language. +- **BM25 stemmer selection**: `languages.json` maps language codes to `rust_stemmers::Algorithm` names. +- **Status display**: `ways status` shows the resolved language. + +### Language configuration resource + +`languages.json` is embedded at compile time. It defines the 52 languages supported by the multilingual embedding model (`paraphrase-multilingual-MiniLM-L12-v2`), even though the current shipping model is English-only. Each entry contains: + +```json +{ + "ja": { + "name": "Japanese", + "native": "日本語", + "bm25_stemmer": null + }, + "de": { + "name": "German", + "native": "Deutsch", + "bm25_stemmer": "German" + } +} +``` + +- `name` / `native`: display and normalization (accepts codes, English names, or native names) +- `bm25_stemmer`: the `rust_stemmers::Algorithm` variant name, or `null` if BM25 cannot support this language + +The `null` stemmer field is the honest signal. Languages where `bm25_stemmer` is null (CJK, Thai, Arabic, etc.) require the embedding engine for matching. BM25 cannot tokenize them — no whitespace boundaries, no suffix-stripping morphology. This is an architectural limitation, not a tuning problem. + +### Matching: language coverage by engine + +The matching pipeline runs: embedding → BM25 → keyword/regex. Each engine has different language coverage: + +**Embedding (primary)** — with the multilingual model (`paraphrase-multilingual-MiniLM-L12-v2`), covers all 52 languages in `languages.json`. Cross-language matching works natively — a Japanese prompt about security produces a vector near the English `description: security vulnerability scanning`. This is the primary matching path for all non-English users. + +**BM25 (fallback)** — covers the ~15 languages where Snowball stemmers exist (Romance, Germanic, Slavic, Turkic, Finnic). `languages.json` `bm25_stemmer` field identifies these. BM25 is the fallback when the embedding engine is unavailable (model not downloaded, `way-embed` binary missing). + +For languages with `bm25_stemmer: null`, BM25 is architecturally incapable — not "stemmer not yet added." These fall into two categories: + +- **No word boundaries**: Japanese, Chinese, Thai have no whitespace between words. BM25's tokenizer (`split on whitespace`) produces whole sentences as single tokens. A segmenter (MeCab, jieba, ICU) would be needed before BM25 could operate at all. +- **Non-concatenative morphology**: Arabic and Hebrew build words from consonant roots with vowel patterns interleaved (k-t-b → kataba, kitāb, maktūb). Snowball's suffix-stripping approach cannot extract these roots. The concept of "stemming" doesn't apply — these languages need root extraction, a fundamentally different operation. + +These languages require the embedding engine. There is no BM25 path and adding one would mean replacing the tokenizer and morphological analyzer — at which point you've built a search engine, not a fallback. + +**Keyword/regex (always)** — language-independent. Technical terms borrowed into all languages (`git commit`, `npm install`, file paths, error codes) match regardless of prompt language. This tier fires even when both embedding and BM25 miss. + +**The practical implication:** for languages BM25 can't handle, the embedding engine is not optional — it's required. `ways status` should surface this: if the resolved language has `bm25_stemmer: null` and the embedding engine is unavailable, warn that matching will be limited to keyword/regex patterns only. + +### Way content stays English + +Way body content (the guidance injected into agent context) is NOT translated. Rationale: + +- The agent reads English perfectly regardless of output language +- 85+ way files × N languages is a maintenance nightmare with divergence risk +- The guidance is for the agent's reasoning, not displayed to the user +- Cross-language injection is well-understood: English instructions → non-English output + +### Native language stubs (shipped) + +The original ADR-107 Draft proposed a tiered file model (`{name}-{lang}.md`). This was initially deferred in favor of cross-language embedding. However, evaluation data showed that native-language stubs dramatically outperform cross-language matching: + +| Language | EN model × EN desc | Multi model × cross-lang | Multi model × native stub | +|----------|-------------------:|------------------------:|-------------------------:| +| ja | -0.03 | 0.69 | **0.93** | +| ar | 0.04 | 0.40 | **0.96** | +| de | 0.08 | 0.62 | **0.82** | +| es | 0.44 | 0.79 | **0.84** | + +Native stubs are now the primary multilingual matching strategy. Each stub provides a `description` and `vocabulary` in the target language, scored by the multilingual embedding model. + +### Packed locale storage (.locales.jsonl) + +Stubs are stored as **packed JSONL**, one file per way, co-located with the way it belongs to: + +``` +ea/briefing/ + briefing.md # the way (English) + briefing.locales.jsonl # all language stubs +``` + +```jsonl +{"lang":"ja","description":"朝のブリーフィング、昨夜の要約","vocabulary":"朝礼 ブリーフィング 要約 優先事項"} +{"lang":"de","description":"Morgendliches Briefing, Tagesübersicht","vocabulary":"Morgenbriefing Tagesübersicht Zusammenfassung"} +``` + +Design constraints: +- **No `embed_threshold` per locale entry** — per ADR-125, thresholds are per-node (English frontmatter only). Locale entries are coordinate aliases; they do not carry their own gates. The corpus generator uses the node's threshold (or system default) for all of a node's aliases. +- **No `embed_model`** in packed format — always `"multilingual"` for locale stubs. +- **Override mechanism**: if `briefing.ja.md` exists as a real file on disk, it supersedes the `ja` entry in `briefing.locales.jsonl`. This allows graduating any stub to a full native-language way with body content. +- **Co-location over aggregation**: one `.locales.jsonl` per way (not per language, not one global file). Way deletion = directory deletion, translations go with it. + +This replaces the individual `{name}.{lang}.md` stub files (which would grow to 4,000+ files at full language coverage). The packed format keeps the training corpus version-controlled, diffable, and lintable while eliminating file sprawl. + +### Dual embedding model (shipped) + +Both models ship simultaneously. `make setup` downloads both: + +| Model | Size | Languages | Use case | +|-------|------|-----------|----------| +| all-MiniLM-L6-v2 | 21MB | English | Precise EN matching (default) | +| paraphrase-multilingual-MiniLM-L12-v2 | 127MB | 52 | Native-language stub matching | + +`ways corpus` splits entries by `embed_model` field into two corpora (`ways-corpus-en.jsonl`, `ways-corpus-multi.jsonl`). The scanner queries both and merges results. Each way's English entry is scored by the EN model; each locale stub is scored by the multilingual model. + +`languages.json` defines the supported language set for the multilingual model. Adding a language means verifying it's in the model's training data and adding the entry — no code changes. + +### Embedding model language verification + +`languages.json` declares what languages we *intend* to support. The embedding model determines what we *actually* support. These must be verified to match. + +A test fixture per language validates that the model produces meaningful cross-language similarity. Each fixture contains a prompt in the target language and an English way description expressing the same intent. The test embeds both and checks that cosine similarity exceeds a minimum threshold (e.g., 0.25 — well below the matching threshold but above random noise). + +```jsonl +{"lang": "ja", "prompt": "依存関係の脆弱性をチェックして", "description": "dependency vulnerability scanning", "min_similarity": 0.25} +{"lang": "de", "prompt": "Abhängigkeiten auf Schwachstellen prüfen", "description": "dependency vulnerability scanning", "min_similarity": 0.25} +{"lang": "ko", "prompt": "의존성 취약점 검사", "description": "dependency vulnerability scanning", "min_similarity": 0.25} +``` + +When the test runs against the current English-only model (`all-MiniLM-L6-v2`), most non-English languages will fail — that's expected and informative. It tells us exactly which languages gain support when we swap to the multilingual model. When we do swap, the same tests validate the new model's coverage without manual verification. + +The test is run as: `ways embed-test-languages` or as part of `make test`. It reads `languages.json`, loads the model, and reports per-language pass/fail. Any language that fails gets flagged — either the model doesn't support it, or the test fixture needs revision. + +This makes model selection empirical: run the tests against candidate models, pick the one that passes the languages you need at the size you can tolerate. + +## Consequences + +### Positive + +- Output language works immediately for all languages — no model or matching changes required +- `agents/` module provides the abstraction point for multi-agent support +- Language resolution cascade respects user intent at every level +- `languages.json` as embedded resource means the language list is data, not code +- BM25 stemmer selection is a one-line change per language in `bm25.rs` +- Embedding model upgrade is a config change, not an architecture change + +### Negative + +- Multilingual matching requires a 6x larger embedding model (21MB → 120MB) +- BM25 fallback quality varies significantly across language families +- CJK/Thai users get no BM25 matching — embedding engine is required, not optional +- Cross-language embedding similarity is lower than same-language — thresholds may need per-language tuning + +### Neutral + +- Way content stays English — no translation infrastructure needed +- Packed `.locales.jsonl` replaces per-language stub files — same data, fewer files +- Override mechanism (`{name}.{lang}.md` supersedes JSONL entry) allows gradual migration from stubs to full native-language ways +- `ways.json` `output_language: "en"` is the default — zero behavior change for existing users + +## References + +- ADR-108: Embedding-Based Way Matching with all-MiniLM-L6-v2 +- ADR-110: Way File Separation and Graph-Compatible Structure +- ADR-111: Unified Ways CLI — Single Binary Tool Consolidation +- ADR-125: Authored Disclosure Graph and Removal of BM25 (supersedes the matching pipeline described here) +- `tools/ways-cli/src/agents/` — Agent config module (Claude Code, system locale) +- `tools/ways-cli/languages.json` — Supported language definitions +- [paraphrase-multilingual-MiniLM-L12-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2) — Multilingual embedding model +- [rust_stemmers](https://docs.rs/rust-stemmers/) — Snowball stemmer implementations (removed with BM25 per ADR-125) diff --git a/docs/architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md b/docs/architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md new file mode 100644 index 00000000..cf0ae6e5 --- /dev/null +++ b/docs/architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md @@ -0,0 +1,168 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'test 2026-03-21: a creative-writing prompt fired 4 ways after the corpus and vocabulary audit with 1 true positive; the remaining false positives are BM25 stem collisions (''agent'', ''document'')' + - evidence: 'timing: the BM25 path spawns 58 processes per prompt (~120ms); the embedding path measures ~22ms' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-03-21 +deciders: + - aaronsb + - claude +related: + - ADR-014 + - ADR-107 + - ADR-125 +imported: + from: docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md + format: v0 + status: Accepted + unmapped: + amended_by: ADR-125 +--- + +# ADR-108: Embedding-Based Way Matching with all-MiniLM-L6-v2 + +> **Amendment note (2026-04-17):** [ADR-125](ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) removed BM25 entirely. The "BM25 remains as a fallback" decision below is historical — the embedding model is now the sole retrieval tier. Sections describing the engine-selection cascade (`way-embed` → `way-match` → NCD), BM25 fields in the corpus, and the "fallback chain is automatic" behavior no longer apply. + +## Context + +ADR-014 introduced BM25 for semantic way matching. ADR-107 Phase 1 shipped an external corpus for correct IDF computation across 58 ways. A vocabulary audit removed 7 ambiguous terms and reduced false positives from 8 to 4 on a test prompt ("write about Dwarf Fortress and AI agent simulation"). + +The remaining 4 false positives expose a fundamental BM25 limitation: bag-of-words matching operates on stems, not meaning. "Agent" in SSH (ssh-agent) and "agent" in AI (autonomous agent) share the same stem. "Document" in docstrings and "document" in "write a document" are identical tokens. No amount of vocabulary tuning can fix this — the model has no concept of word sense. + +The ways system now has 62 ways across 5 domains (softwaredev, itops, meta, writing, research). As the corpus grows, BM25 vocabulary collisions will increase. Each new domain introduces common terms that overlap with existing ways. + +Meanwhile, the prompt-evaluation loop must remain imperceptible to the user. The current BM25 path spawns 58 processes per prompt (~120ms). Any replacement must be faster, not slower. + +### Evidence + +Tested 2026-03-21 in an empty directory with a creative writing prompt: + +| State | Ways fired | True positives | +|-------|-----------|----------------| +| Before corpus + audit | 8-9 | 1 | +| After corpus + audit | 4 | 1 | + +The 3 remaining false positives (docs, docstrings, subagents) all match on stem collisions that BM25 cannot disambiguate. + +## Decision + +Replace BM25 with embedding-based semantic matching using **all-MiniLM-L6-v2** as the primary way-matching engine. BM25 initially remained as a zero-dependency fallback; [ADR-125](ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) subsequently removed BM25 entirely, making the embedding model the sole retrieval tier. + +### Model choice: all-MiniLM-L6-v2 + +- **Parameters**: 22M (6 transformer layers, 384-dim embeddings) +- **License**: Apache 2.0 +- **GGUF size**: ~44MB (F16), ~22MB (Q5_K_M), ~21MB (Q4_K_M) +- **Why this model**: Smallest viable sentence embedding model with strong semantic discrimination. Battle-tested in production search systems. Pre-converted GGUF available on HuggingFace (second-state/All-MiniLM-L6-v2-Embedding-GGUF). + +### Architecture: pre-compute + single-spawn runtime + +Way embeddings are computed at **authoring time** and stored in `ways-corpus.jsonl`. At runtime, only the user prompt is embedded — one forward pass, then cosine similarity against pre-computed vectors. + +**Authoring time** (like `generate-corpus.sh` today): +``` +way-embed generate --corpus ways-corpus.jsonl +``` +Reads all way.md descriptions, embeds them, writes 384-dim vectors to the JSONL alongside existing BM25 fields. + +**Runtime** (called by `match-way.sh`, replaces 58 `way-match pair` calls): +``` +way-embed match --corpus ways-corpus.jsonl --query "user prompt" +``` +One process spawn. Loads pre-computed vectors from JSONL, embeds the prompt (~3-5ms), computes 58 cosine similarities (<1ms), outputs matches above threshold. + +### Timing budget + +| Operation | BM25 today (58 spawns) | Embedding (1 spawn) | +|-----------|----------------------|---------------------| +| Process spawns | 58 × ~2ms = ~116ms | 1 × ~2ms | +| Model load | N/A | ~5ms (GGUF mmap) | +| Scoring | <1ms each (in-process) | ~12ms forward pass + <1ms cosine sims | +| **Total** | **~120ms** | **~22ms** | + +The embedding path is **5-6x faster** than the current BM25 path while providing dramatically better discrimination. (Measured on Linux x86_64, 2 threads.) + +### Binary packaging + +Build on llama.cpp's GGML library (pure C tensor computation, no dependencies). Two files: + +- `bin/way-embed` — static binary (~5MB), built with cosmocc for cross-platform APE +- Model weights downloaded to `${XDG_CACHE_HOME}/claude-ways/user/` (~44MB F16, ~22MB Q5_K_M) + +The binary loads the model via mmap (no full read into RAM). The model is distributed via GitHub Release artifact or direct HuggingFace download, verified against a committed SHA-256 checksum. It lives in the XDG cache directory, not in the git repo (too large at 44MB). + +Future option: llamafile-style single binary with model weights concatenated into the executable. Deferred until cosmocc builds of llama.cpp stabilize. + +### Corpus format evolution + +`ways-corpus.jsonl` gains an `embedding` field: + +```json +{ + "id": "writing", + "description": "Content creation — documents, presentations...", + "vocabulary": "write draft compose...", + "threshold": 2.0, + "embedding": [0.023, -0.041, 0.118, ...] +} +``` + +BM25 fields (`threshold`, tokenized vocabulary) were retained in the corpus at the time of this ADR to serve the fallback path. [ADR-125](ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) removed the BM25 engine; these fields are no longer read at runtime. + +### Configuration (superseded) + +> Superseded by ADR-125: embedding is the sole tier; no fallback cascade. + +Originally, the scanner detected which engine was available: + +1. If `bin/way-embed` exists and model file present → use embedding +2. Else if `bin/way-match` exists → use BM25 with corpus +3. Else if gzip + bc available → use NCD (legacy fallback) + +Per ADR-125, the embedding model is a hard dependency. If the embedding engine is unavailable, matching does not degrade — it errors. This surfaces setup problems early rather than silently returning degraded results. + +### Security boundary + +Same constraint as ADR-107: **runtime scanners never write to `~/.claude/`**. The model file and pre-computed embeddings are authoring-time artifacts committed to the repo. The runtime binary only reads. + +The model file is a published, checksummed artifact from HuggingFace. It can be verified against known hashes. It does not execute arbitrary code — GGUF is a tensor format, not an executable format. + +## Consequences + +### Positive + +- Solves the stem-collision problem that BM25 cannot fix. "SSH agent" and "AI agent" will have distant embedding vectors despite sharing a stem. +- 15x faster than current 58-spawn BM25 path. Imperceptible to the user. +- Pre-computed embeddings mean runtime cost is independent of corpus size. 200 ways would cost the same as 58. +- Multilingual potential without per-language stemmers or vocabulary files. MiniLM handles many languages out of the box (trained on multilingual data). ADR-107 Phase 3 locale work becomes simpler. +- BM25 vocabulary tuning becomes unnecessary. New ways need only a good description — no manual keyword curation. + +### Negative + +- 44MB model file download (F16) or ~22MB (Q5_K_M). Not git-tracked — distributed via GitHub Release or HuggingFace. Users run `make model` to download. +- llama.cpp / GGML dependency for building the binary. More complex build than the current single-file `way-match.c`. Though the binary itself remains dependency-free once compiled. +- Model quality is fixed at MiniLM's training — it may not perfectly capture domain-specific semantics (e.g., "way" as a concept in this system vs. "way" as a path). BM25's explicit vocabulary handles this better for known edge cases. +- Quantization (Q4) trades quality for size. Need to validate that Q4 embeddings still discriminate well enough on the test fixture corpus. + +### Neutral + +- `way-match` binary and BM25 scoring initially remained in the repo as the fallback. ADR-125 subsequently removed both, as the fallback was rarely exercised once the multilingual model was reliably distributed. +- The `ways-corpus.jsonl` format was initially additive — embedding vectors sat alongside BM25 fields. Per ADR-125, the BM25 fields are no longer read at runtime; the corpus format will drop them when convenient. +- The `/ways-tests` skill needs to learn to score with embeddings (cosine similarity thresholds are on a different scale than BM25 scores). +- Threshold values will need recalibration. BM25 thresholds (1.8-2.5) don't apply to cosine similarity (0.0-1.0). The test fixture corpus provides the calibration data. + +## Alternatives Considered + +- **Larger embedding models (Snowflake Arctic, Qwen3-0.6B)**: Better quality but 1.2GB+ model files. Overkill for matching 60 short descriptions against short prompts. MiniLM's 22M parameters are sufficient for this task scale. +- **ONNX Runtime instead of GGML/llama.cpp**: Would require shipping a shared library (~15MB onnxruntime.so). Less portable than a static binary. GGML is pure C with no dependencies, matching our existing build pattern. +- **Python-based inference (fastembed, sentence-transformers)**: Adds a Python runtime dependency. Non-starter for a system that runs on bash + coreutils + static binaries. +- **Fine-tuning MiniLM on way-matching data**: Would improve domain-specific discrimination (training data exists: 74 test fixtures + 58 way descriptions). Deferred — try the off-the-shelf model first. Fine-tuning is ~$5-20 on cloud GPU if needed. +- **Hybrid scoring (BM25 + embedding combined)**: Run both engines, combine scores. More complex, harder to calibrate, and the embedding model alone should handle all cases BM25 handles plus the ones it can't. Keep it simple — one engine primary, one fallback. +- **Keep BM25 and accept the false positives**: The vocabulary audit reduced false positives from 8 to 4, but the remaining 3 FPs are structural BM25 limitations. As the corpus grows past 100 ways, vocabulary collisions will worsen. The problem doesn't stabilize — it compounds. diff --git a/docs/architecture/ways/ADR-114-attend-as-insistent-way-trigger-type.md b/docs/architecture/ways/ADR-114-attend-as-insistent-way-trigger-type.md new file mode 100644 index 00000000..dcac684b --- /dev/null +++ b/docs/architecture/ways/ADR-114-attend-as-insistent-way-trigger-type.md @@ -0,0 +1,255 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - matching + - attend + - disclosure +basis: + - standard: Claude Code's Monitor tool delivers stdout lines as one-line async notifications, too short for way-length guidance + - precedent: ADR-113 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-09 +deciders: + - aaronsb + - claude +related: + - ADR-104 + - ADR-105 + - ADR-108 + - ADR-112 + - ADR-113 +imported: + from: docs/architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md + format: v0 + status: Accepted +--- + +# ADR-114: `attend` Events as an Insistent Way Trigger Type + +## Context + +[ADR-113](../attend/ADR-113-attend-active-awareness-module.md) introduces `attend`, a sibling binary in the agent-ways workspace that implements the active awareness layer described in the [Cognitive Loop and the Awareness Layer](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) design note. `attend` observes Claude Code session state and environmental signal, tracks approaching mechanical consequences (context pressure, reflection deferral, etc.), and produces emissions that need to reach Claude. + +The question this ADR answers is: **how do those emissions become guidance Claude reads?** + +### Current trigger types + +The ways system supports a small set of trigger types, each keyed to a hook event in Claude Code's loop: + +| Trigger | Fires on | Example | +|---|---|---| +| `UserPromptSubmit` | User sends a message | Intent classification ways | +| `PreToolUse` | Before Claude calls a tool | Tool guidance ways | +| `Stop` | After Claude's response | Reflection capture, cleanup | +| `SessionStart` | New session begins | Project pulse (ADR-106), orientation | +| `PostCompact` | After a compaction pass | Resume ways, ADR-112 | +| `context-threshold` | Context crosses a percentage boundary | Reflection, memory distillation (ADR-104, ADR-112) | + +Every existing trigger is **reactive**: it keys off an event inside Claude's own loop (Claude did something, or Claude is about to do something, or Claude's context is changing). There is no trigger type for events *outside* Claude's loop — no way for a file change, an idle timeout, a context-pressure projection, or a user-requested timer to reach Claude through the same machinery. + +### The delivery primitive: `Monitor` + +`attend` produces exactly that class of event: externally-observed signals that occur between hook boundaries and need to surface into Claude's attention. The delivery primitive for those signals is already available in Claude Code — the `Monitor` tool, introduced in the same release wave that makes ADR-113 buildable. `Monitor` accepts a shell command, runs it as a background process, and delivers each line the command writes to stdout as an asynchronous notification in Claude's conversation. With `persistent: true`, a single invocation covers the session's lifetime. This is the mechanism by which `attend`'s stdout emissions reach Claude. + +`Monitor` is described in detail in the [design note](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) as the *delivery primitive*. ADR-113 commits `attend` to writing its observations as single-line stdout notifications that `Monitor` delivers. **That is the primary delivery channel for the awareness layer**, and it requires no involvement from the ways system at all. + +### Why ways integration still matters + +Given that `Monitor` handles delivery, why does this ADR exist? Because there is a second class of `attend` emission — high-salience observations where the notification alone is insufficient and deeper guidance would materially improve Claude's response. Examples: + +- `attend` detects context pressure imminent enough that the reflection window is closing. The one-line notification *"projected critical in 3 turns"* tells Claude the stakes, but the *actual reflection guidance* — what to reflect on, how to compress it, what structure to use — lives in a way body that ADR-112 already defined. +- `attend` notices a peer Claude Code session has modified a file this session is editing. The notification surfaces the conflict, but the *coordination pattern* — how to reconcile, whether to rebase, how to communicate — belongs in a peer-coordination way. +- `attend` detects a build failure. The notification says "build failed," but the *triage playbook* belongs in a way. + +For these cases, `attend` formats the `Monitor` notification as an **affordance**: a string that explicitly names a `ways show attend/<signal-type>` command Claude can invoke if it wants the deeper guidance. When Claude invokes that command, the ways system runs the matcher and ADR-104 disclosure gate normally, and the matched way's body is injected through the standard guidance pipeline. + +This gives the awareness layer **two composable delivery paths**: + +1. **`Monitor` notification only** — default for most emissions. Claude reads the one-line observation, integrates it, acts or dismisses. No ways involvement. +2. **`Monitor` notification + affordance → `ways show attend/<signal>`** — for high-salience emissions. Claude reads the notification, recognizes the stakes, invokes the named ways command, receives the full guidance body. + +The ways system does not need to know `attend` exists until Claude invokes `ways show attend/<signal>`. At that point, the matcher treats the invocation as it would any other explicit way query, runs ADR-104's disclosure gate normally, and injects the matched way. No new hook event, no new matcher input source, no automatic firing — just a well-formed query Claude chose to make. + +## Decision + +Extend the way trigger schema with a new type, `attend`, that declares a way as the handler for one or more `attend` signal types. Ways with this trigger are **invoked on demand** when Claude runs `ways show attend/<signal>`, typically in response to an affordance in a `Monitor`-delivered `attend` notification. The ways system runs the matcher and ADR-104 disclosure gate normally for each invocation. + +The primary delivery channel for `attend` observations remains `Monitor` (see ADR-113). This ADR adds the secondary on-demand path for cases where a notification alone is insufficient. + +### Schema + +Ways declaring `trigger.type: attend` use the following frontmatter fields: + +```yaml +--- +name: reflect-on-context-pressure +description: Guide Claude through progressive ledger reflection as context approaches compaction +trigger: + type: attend + signals: + - context-pressure + - reflection-overdue +--- + +(way body — the guidance Claude reads when it invokes `ways show attend/context-pressure` +or `ways show attend/reflection-overdue`) +``` + +Field definitions: + +- **`type: attend`** (required) — Marks this way as a handler for an `attend` signal type. Ways with this trigger are never fired automatically by a hook event; they are invoked on demand by Claude in response to an affordance. +- **`signals`** (required) — List of signal types this way handles. When Claude invokes `ways show attend/<signal-type>`, the ways CLI selects ways whose `signals` list contains the requested signal. A way may handle multiple signals; one signal may be handled by multiple ways (in which case the matcher scores among them as usual). + +Standard fields (`name`, `description`, `embed_threshold`, way body, etc.) behave exactly as they do for other trigger types. The way body is the guidance Claude reads when the invocation succeeds and the disclosure gate permits the injection. + +Absent from the schema by design: no `subscriptions`, no `debounce_turns`, no `salience_floor`. These were artifacts of an earlier draft that modeled automatic firing. In the Monitor-primary design, habituation is handled by the disclosure gate on each invocation (ADR-104), and salience decisions happen in `attend` before the affordance is ever emitted. + +### Invocation via affordance + +The full flow for a high-salience emission: + +1. `attend` detects a signal worth surfacing (e.g., context pressure crossing a critical threshold) +2. The insistence emitter computes the observation text and determines that a deeper-engagement path is warranted (typically because the emission mode is `insistent` or `critical`) +3. `attend` writes a single notification line to stdout, formatted as an affordance: + + ``` + context at 86% — projected critical at turn 58 (3 turns remaining). Use `ways show attend/context-pressure` for reflection guidance. + ``` + +4. `Monitor` delivers the line to Claude as an asynchronous notification +5. Claude reads the notification, recognizes the stakes from the text and projection, and decides whether to invoke the affordance +6. If Claude invokes `ways show attend/context-pressure`: + - The ways CLI looks up ways with `trigger.type: attend` where `signals` contains `context-pressure` + - Each candidate is scored through the standard embedding + BM25 + NCD pipeline against the request context + - ADR-104's disclosure gate checks whether the matched ways have been disclosed recently + - The matcher returns the surviving way's body as injected guidance +7. Claude receives the guidance and acts on it + +At every step, Claude retains agency. `attend` suggests; Claude decides. The ways system provides the guidance only when asked. The disclosure gate ensures habituation applies uniformly regardless of source. + +### Emission modes and delivery mapping + +ADR-113 defines the emission modes. Each mode maps to a specific delivery pattern in this two-path model: + +| Mode | Notification via `Monitor` | Includes affordance? | When Claude invokes the affordance | +|---|---|---|---| +| `silent` | No line emitted | N/A | N/A | +| `informational` | Short declarative observation | No | N/A — informational lines do not warrant deeper engagement | +| `affordance` | Observation + named tool invitation | Yes (optional) | Ways CLI runs matcher + disclosure gate; way fires if eligible | +| `insistent` | Observation + explicit stakes + named ways command | Yes (suggested) | Disclosure gate applies; way fires if eligible; if recently disclosed, terse re-surface | +| `critical` | Maximum-clarity observation + explicit consequence language + ways command | Yes (strongly suggested) | Disclosure gate is more willing to re-fire given the criticality; way injection prioritized | + +The escalation between modes is handled entirely by `attend`'s insistence emitter. The ways system responds to invocations the same way regardless of mode — the mode determines how `attend` phrases the notification, not how ways responds to the resulting query. + +The "acknowledged-but-silent" state from the design note is implemented in `attend`'s deferred intent store, not the ways system. Observations held below the emission threshold never become `Monitor` notifications, which means the ways system never sees them. + +### Unified disclosure and habituation + +Each invocation of `ways show attend/<signal>` passes through the same disclosure gate as any other way invocation. The disclosure tracker does not distinguish signal sources; it only tracks which ways have been disclosed and how recently. This means: + +- A way that handles an `attend` signal is subject to the same recent-disclosure suppression that applies to reactively-triggered ways +- ADR-104's token-gated re-disclosure rules apply uniformly +- Habituation works the same way: first invocation triggers full guidance, rapid re-invocations are suppressed to terse re-surfacing, extended silence eventually allows re-disclosure at full weight + +This uniform treatment is load-bearing. Claude does not need a separate mental model for reactive vs on-demand ways. Authors of `attend` signal handlers don't need to learn new disclosure rules. The underlying system is one system whose invocations sometimes come from hooks and sometimes come from Claude acting on an affordance. + +### Way authoring + +The author of an `attend` signal handler writes it the same way they write any other way. They identify what deeper guidance Claude should receive when a particular `attend` signal warrants engagement, they write the guidance as the way body, they declare which signals the way handles, and the system handles the rest. No knowledge of `attend` internals is required. + +Example: + +```yaml +--- +name: note-build-completion +description: Guide Claude in responding to a completed background build +trigger: + type: attend + signals: + - build-complete +--- + +A background build has finished. Consider: + +- Running the test suite to validate the build +- Checking output for warnings worth investigating +- Updating the working context with what changed + +The user may already be aware via terminal output — only engage with this if +the build result is directly relevant to current work. +``` + +This way is invoked when Claude runs `ways show attend/build-complete` in response to an `attend` affordance. The author didn't write any sensor code, didn't touch `attend`'s internals, and didn't need to know how the underlying signal is detected. The contract between `attend` and this way is a single string: `build-complete`. + +### Graceful no-op when `attend` or `Monitor` is absent + +If `attend` is not running, ways with `trigger.type: attend` are **inert, not broken**. No affordances are ever emitted, so `ways show attend/<signal>` is never invoked automatically. The rest of the ways system functions normally. When `attend` starts, affordances begin appearing in notifications and the ways become live. + +Equally, if `Monitor` is unavailable (because the running Claude Code version doesn't ship it, or because the SessionStart way that drives invocation was not installed), the awareness layer produces no notifications at all, and `attend` signal handler ways are simply never invoked through the automatic path. Claude could still invoke `ways show attend/<signal>` manually from a prompt, and the ways CLI would handle that correctly — the invocation path is not conditional on `attend` running. + +This is the "presence as additive" property from the design note, expressed at the ways layer. An agent-ways installation that never runs `attend` or `Monitor` never experiences any downside from having `trigger.type: attend` ways in the corpus — they are dormant until Claude invokes them, and nothing else breaks. + +### Scope of this ADR + +This ADR defines: + +- The new `trigger.type: attend` schema with the `signals` field +- Invocation via affordance: how `ways show attend/<signal>` integrates with the matcher and ADR-104 disclosure gate +- The mapping from `attend` emission modes to `Monitor` notification format and affordance presence +- Way authoring for signal handler ways +- Graceful inert behavior when `attend` or `Monitor` is absent + +It explicitly does not define: + +- The `attend` binary's internal architecture (that is ADR-113) +- The `Monitor` tool's behavior (that is Claude Code documentation) +- The sensor catalog or specific signal types (those grow incrementally) +- Any changes to ADR-104's disclosure gate — the gate applies as-is to `ways show` invocations sourced from `attend` affordances + +## Consequences + +### Positive + +- **Minimal surface area.** The ways system gains one new trigger type and one new invocation path (`ways show attend/<signal>`). It does not need to accept emission streams, process salience weights, or implement a second matching pipeline. The integration is one schema extension and one CLI convention. +- **Uniform habituation.** ADR-104's disclosure gate applies to every `ways show attend/<signal>` invocation exactly as it applies to any other way lookup. No special cases. No separate habituation rules for proactive vs reactive ways. +- **Clear separation of concerns.** `attend` + `Monitor` handle observation and delivery. The ways system handles guidance retrieval and disclosure. The boundary is "Claude invoked a `ways show` command with a specific signal name," which is a clean contract both sides can reason about independently. +- **Incremental adoption.** Ways with `trigger.type: attend` can be added to the corpus before `attend` is built — they simply remain dormant until Claude invokes them. Claude can also invoke them manually from a prompt even without `attend` running, which makes the ways immediately useful as signal-shaped handlers regardless of the awareness layer's state. +- **Two delivery paths, one attention surface.** Most `attend` observations arrive as `Monitor` notifications and never touch the ways system. High-salience ones route through ways for deeper guidance. Claude sees one consistent attention surface even though two paths are active underneath. +- **Claude retains agency.** `attend` suggests affordances; Claude decides whether to invoke them. The ways system only produces guidance in response to Claude's explicit request. There is no path by which the awareness layer forces injection. + +### Negative + +- **Two-path cognitive load for way authors.** Authors writing `trigger.type: attend` ways need to understand that their ways fire on `ways show` invocations from Claude, not on sensor events. The authoring model is simple (`signals: [...]` and a body), but the mental model is new. +- **Affordance format is a contract between `attend` and ways.** If the `ways show attend/<signal>` command name convention changes, both sides need to update. Mitigation: the format is a short documented subset — command name + signal identifier — and versioning it if needed is cheap. +- **Implicit dependency in way corpora.** A corpus with many `trigger.type: attend` ways assumes a source that knows to invoke them. Without `attend`, they are reachable only via explicit user or Claude invocation. Mitigation: the ways are documented as "dormant without `attend`" and do not fail when `attend` is absent; reactive ways should cover baseline behavior. + +### Neutral + +- **No change to existing trigger types.** This ADR adds a new type; it does not modify `UserPromptSubmit`, `PreToolUse`, `Stop`, `SessionStart`, `PostCompact`, or `context-threshold`. Existing ways continue to work exactly as before. +- **Way authoring model is unchanged for non-attend ways.** Only authors who want to write signal handler ways need to learn the new fields. Everyone else continues as before. +- **Scoring infrastructure is unchanged.** The embedding + BM25 + NCD tier (ADR-107, ADR-108) handles `ways show attend/<signal>` invocations exactly as it handles any other way lookup. No new scoring path. + +## Alternatives Considered + +- **New hook event type (e.g., `AttendEmission`).** Rejected. Duplicates what `Monitor` already provides. `Monitor` is a first-class Claude Code tool that delivers async notifications by design; a new hook event class would be a parallel mechanism doing the same job with less flexibility. +- **Automatic firing of ways from `attend` emissions (the original draft of this ADR).** Rejected after the `Monitor` delivery primitive became available. Automatic firing required `attend` to invoke the ways matcher directly, injecting way bodies into Claude's context without Claude's explicit consent. In the `Monitor`-primary model, this violates the "Claude retains agency" invariant from ADR-113 — the awareness layer should inform, and Claude should decide. Keeping ways on-demand via affordance preserves that invariant cleanly. +- **`Monitor`-only with no ways integration at all.** Considered seriously. The argument is that a well-formed notification can carry enough information that deeper guidance isn't needed — `attend` just writes the full guidance into the notification text. Rejected because: (1) Notification text is one line (plus 200ms batching); long prose guidance doesn't fit the format. (2) The ways system already provides rich, habituation-aware guidance retrieval; duplicating that in notification text would compromise the brevity that makes notifications useful. (3) High-salience guidance benefits from ADR-104's disclosure gate — a `Monitor`-only model can't use it. +- **Direct injection bypassing the disclosure gate.** Rejected. Bypasses ADR-104's habituation rules and would cause repeated signals to dominate Claude's context. The disclosure gate is exactly the right place for these decisions, regardless of whether the invocation came from a hook or from Claude acting on an affordance. +- **Invoke ways from `attend` directly (not via Claude).** Rejected. `attend` invoking ways on Claude's behalf would inject content without Claude's agency, violating the design note's invariants. Claude must be the one to invoke `ways show attend/<signal>` because Claude is the one deciding the affordance is worth engaging with. +- **A single catch-all "external" trigger type instead of `attend`-specific.** Considered and rejected. Tying the trigger type to a specific source documents what's producing the affordance convention and prevents the trigger type from becoming a dumping ground. If another source of proactive signal is ever needed, it gets its own trigger type and its own ADR. +- **`subscriptions` as the field name (earlier draft).** Renamed to `signals` because the new model doesn't involve subscribing to an event stream — the way simply declares which signal names it handles when invoked. + +## References + +- **Design note:** [Cognitive Loop and the Awareness Layer](../practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) +- **Related ADRs:** + - [ADR-104](./ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — Disclosure gate that this ADR reuses + - [ADR-105](./ADR-105-progressive-disclosure-for-way-trees.md) — Progressive disclosure model + - [ADR-108](./ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) — Matcher that scores attend emission payloads + - [ADR-112](../archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — Reflection and ledger ways that will be the first consumers of attend signals + - [ADR-113](../attend/ADR-113-attend-active-awareness-module.md) — The attend binary whose emissions this ADR routes diff --git a/docs/architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md b/docs/architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md new file mode 100644 index 00000000..9b1f3fdf --- /dev/null +++ b/docs/architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md @@ -0,0 +1,376 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - attend + - matching +supersedes: + - ADR-104 + - ADR-119 + - ADR-121 + - ADR-112 +basis: + - evidence: ways re-disclosure was a flat REDISCLOSE_PCT = 25 step function while attend used the same decay math on wall-clock time + - evidence: 'Phase F validation session: 92 turns, 212K tokens, zero operator redirections, reactive firing and decay verified in flight (docs/hooks-and-ways/observed-behavior.md)' + - precedent: ADR-119 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-14 +deciders: + - aaronsb + - claude +related: + - ADR-113 + - ADR-114 + - ADR-117 + - ADR-119 + - ADR-121 +imported: + from: docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md + format: v0 + status: Accepted +--- + +# ADR-123: Firing dynamics — progression-axis unification for attend and ways + +## Context + +Two tools in this workspace implement firing dynamics independently: + +- **attend** has landed the action-potential engagement model (ADR-119) as `EngagementState` in `sensor-trait`, and has a drafted salience decay for signal presentation (ADR-121). Both use wall-clock time — `Instant`, `Duration`, `decay_per_minute`. +- **ways** has a rudimentary model in `ways-cli/src/session.rs`: token-position distance as a percentage of the context window, per-way markers, epoch counters, and a fire-count counter. The suppression is a step function — `REDISCLOSE_PCT: u64 = 25` — with no curves, multiplier decay, burst detection beyond the flat counter, or re-engagement reset. + +The math is identical across both tools: exponential decay, refractory-period multipliers, burst detection over a windowed history. Only the units differ — seconds for attend, token positions for ways. A naive unification would generalize the engine over a `Clock` trait (`Instant` for attend, `TokenPosition` for ways), but that smuggles time semantics into the engine where none exist. + +There is a cleaner abstraction hiding under the question: **attend's wall clock and ways' token count are both instances of the same thing — a monotonic progression axis supplied by the caller.** Neither is privileged. The engine doesn't need to know what progression means; it only needs a `u64` that strictly increases. Each caller labels the axis for its own context. + +This reframe matters beyond the refactor. It means: + +- **Future progression axes come for free.** Turn count, commits-since-branch, bytes-written, lines-changed — any monotonic can be wired in without touching the engine. ways itself is designed to be hostable by any turn-based coding agent that exposes the necessary data; different hosts may expose different natural axes, and the engine must not assume any single one. +- **Ways' axis choice is motivated by the host's own addressing unit, not by any specific decay theory.** Firing decisions should be keyed on the axis the host uses to address what ways is injecting. For transformer-based hosts, that axis is token position — it is the unit attention uses to address content, regardless of what decay shape attention happens to apply (see Decision 4 for the full argument, including its limits). Wall clock is external to the host's addressing; turn count is a coarser aggregate; token position is the host's own unit. This generalizes: a future host that addresses content differently supplies a different axis, and the engine adapts automatically. +- **Curves become first-class.** When the engine stops owning time semantics, the shape of decay/refractory becomes a parameter instead of a built-in — which enables progressive disclosure, staircase re-firing, and other patterns that ADR-119 and ADR-121 could not express because they were each hard-coded to one curve. +- **Event-count burst detection, not tick-windowed.** A subtle consequence of progression-axis genericity: some axes are granular (wall-clock seconds), others chunky (token position can jump thousands in a single tool call). Time-windowed burst detection degenerates on chunky axes — a single event can swallow the window in one step. Burst detection therefore must be event-count based with its window defined implicitly by the decay curve itself, not by a separate tick span. This is developed in Decision 2. + +Additionally, ways currently only fires predictively — against user prompts (`check-prompt.sh`), task spawns (`check-task-pre.sh`), about-to-edit file paths (`check-file-pre.sh`), about-to-run bash commands (`check-bash-pre.sh`). It has no reactive firing path. Observations like "the file that was just written is now 900 lines long" are structurally invisible. Extending into `PostToolUse` and `PostToolUseFailure` requires an evaluation model that isn't semantic embedding — it's metric predicates over observed state. Both modes should compose under the same firing-dynamics engine rather than growing a parallel track. + +## Decision + +### 1. Unit-agnostic progression-axis engine + +The firing-dynamics engine operates over a caller-supplied monotonic progression axis. The public API uses progression-flavored types, not time-flavored ones: + +```rust +pub type Tick = u64; +pub type TickDelta = u64; + +pub struct EngagementState { + curve: Curve, + history: VecDeque<(Tick, f64)>, // (tick, magnitude) per fire + last_fire: Option<Tick>, +} +``` + +All curve-specific parameters live inside the `Curve` variant (Decision 2). `EngagementState` itself holds only the curve plus the event history. Burst detection, refractory, decay, and salience are all derived by the curve from `(current_tick, history)`. + +No `Instant`, no `Duration`, no `decay_per_minute`. The word "clock" does not appear in the engine surface. Each caller supplies its own meaning: + +- **attend** interprets ticks as wall-clock seconds. `epoch_secs()` replaces `Instant::now()` at call sites. +- **ways** interprets ticks as token positions. `get_token_position()` supplies ticks. +- **Future callers** interpret ticks as whatever monotonic matches their cadence. + +Burst detection and decay evaluation are derived from the tick history — `current_tick - entry_tick` compared against `burst_window`, curve evaluated against tick delta. The engine cannot and does not distinguish between "two seconds" and "two tokens" — both are `2u64`, and the curve parameters are in the same unit as the ticks. + +### 2. Pluggable curve as first-class parameter + +`Curve` is a sealed enum — cheap to match, serializable in frontmatter, no trait-object dispatch. All decay parameters across all variants are expressed as **`half_life` in caller ticks**, for a single consistent mental model: + +```rust +pub enum Curve { + /// Smooth exponential decay. Salience = 0.5^(delta / half_life). + Exponential { half_life: TickDelta }, + + /// Action potential — event-count burst detection raises a refractory + /// multiplier which decays back toward 1.0 over tick distance. + /// + /// The "burst window" is NOT a tick span. It is defined implicitly by + /// which history entries still have non-trivial multiplier contribution + /// (via multiplier_half_life). This is robust to chunky progression axes + /// where a single event can advance the tick by thousands. + ActionPotential { + burst_threshold: usize, // fires in recent history to trigger a burst + peak_multiplier: f64, // multiplier at burst peak (e.g., 1.5 or 2.0) + absolute_refractory: TickDelta, // hard suppression after a burst + multiplier_half_life: TickDelta, // half-life of multiplier decay toward 1.0 + }, + + /// Explicit re-fire schedule at tick deltas with diminishing salience. + /// Progressive disclosure as a first-class pattern. + ProgressiveStaircase { steps: Vec<(TickDelta, f64)> }, + + /// Discontinuous step: suppressed for N ticks, then fully recovered. + /// Valid shape for ways that want all-or-nothing suppression without decay. + Flat { suppression: TickDelta }, +} + +impl Curve { + pub fn salience_at(&self, delta: TickDelta) -> f64 { /* ... */ } + pub fn multiplier_at(&self, delta: TickDelta, history: &[(Tick, f64)], current: Tick) -> f64 { /* ... */ } +} +``` + +- **`Exponential`** — primary salience-decay shape. Ways' likely default once tuned. +- **`ActionPotential`** — the ADR-119 engagement model, ported via event-count burst detection (see below). Usable for attend's current sensors and for ways that want inward-gate refractory on top of outward-gate decay. +- **`ProgressiveStaircase`** — new shape. A way declares `[(0, 1.0), (15_000, 0.5), (40_000, 0.2)]` and re-fires at those deltas with the declared salience. +- **`Flat`** — explicit step function. Not a shim, not a backward-compatibility translation. A valid first-class opt-in for ways that genuinely want discontinuous behavior. + +**Burst detection is event-count based, not tick-windowed.** This is the critical adaptation for chunky progression axes. Time-windowed burst detection (ADR-119's original form) assumed events progress smoothly against the tick axis. On wall-clock axes this holds. On token-position axes it does not — a single `Read` tool call can advance the tick by 5k–20k in one step, swallowing any reasonable `burst_window` in one event. Burst detection therefore windows by *event count in recent history*, with "recent" defined implicitly by the decay of the refractory multiplier itself: once an event's contribution has decayed past a small epsilon (e.g., multiplier drops back below 1.01), it no longer counts toward burst detection. Attend's time-windowed behavior is preserved as a degenerate case — short `multiplier_half_life` produces a tight effective window; chunky ways use longer `multiplier_half_life` and event count stays meaningful regardless of tick jumps. + +This is why `EngagementState` holds only `(curve, history, last_fire)` — there is no separate `burst_window` field. The window is a derived property of the curve's multiplier decay, not a standalone parameter. + +### 3. Inward and outward gates, cleanly separated + +The engine exposes two query methods that answer different questions against the same state: + +- **`should_fire(current_tick, stimulus_magnitude) -> bool`** — the inward gate. "Should this new event be allowed to fire?" Consults the refractory portion of the curve. This is the question ADR-119 answers for attend sensors. +- **`current_salience(current_tick) -> f64`** — the outward gate. "Has the last-fired guidance faded enough that re-injection is warranted?" Consults the decay portion of the curve. This is the question ADR-121 answers for attend signals. + +ways' current flat `redisclose` conflates these. The unified engine separates them cleanly so pre- and post-firing decisions consult the same state and curve but ask different questions. `Curve::Flat` collapses them into the same answer (binary suppressed/allowed), preserving current semantics for ways that don't opt in to a richer shape. + +### 4. Ways' tick unit: host addressing, not a decay theory + +Ways' tick is token position for transformer-based hosts. The justification is the host's addressing unit, not a specific theory of how attention decays. + +**The principle:** firing dynamics should be keyed on the axis the host agent uses to address what ways is injecting. For a transformer, that axis is token position — it is the unit the model uses to locate and refer to content via its attention mechanism, regardless of what decay shape attention happens to apply in practice. Wall clock is external to the host's addressing (the model cannot observe it). Turn count is an aggregate over token positions that wobbles with per-turn context density (a long think turn and a short reply both count as "one turn" but displace very different amounts of addressable context). Token position is the host's own unit. + +**On RoPE specifically:** rotary positional embedding does provide a baseline long-term inner-product decay with relative distance, which is a convenient piece of theoretical grounding — it means token distance is not an arbitrary choice but corresponds to a real axis the model uses. However, this should not be overstated. Modern attention heads in trained LLMs (Claude included) are demonstrably capable of overriding RoPE's baseline decay via needle-in-a-haystack retrieval — a highly salient token at position `P - 100_000` can still receive near-full attention when the current decision depends on it. The model's attention does *not* fade on a clean exponential curve along token distance; it fades until a specific attention head decides the content is critical, at which point the effective salience spikes back up internally. + +What this means for ways: token position is the *closest available proxy* for "how much context has accumulated between a way's injection and the current decision point," which is the quantity firing dynamics should track. It is not a direct model of attention decay. The firing engine assumes the guidance fades in effective impact as context accumulates; the model's actual attention mechanism may or may not honor that assumption for any specific token. That's fine — the firing decision is about *presentation economics* (when to re-inject guidance so it's freshly available), not about *modeling attention internals*. + +**Generalization beyond transformers:** ways is designed to be hostable by any turn-based coding agent that exposes the needed hooks and data. Different hosts may address content differently — a sliding-window chunked-conversation agent might use chunk count; a stateless agent might only expose turn count. The principle holds: use whatever axis the host uses to address injected content. Token position is the right answer for transformer hosts; it is not the right answer for all possible hosts. The engine's unit-agnosticism is what makes this portability possible. + +**For attend:** wall clock is correct because attend steers external timing — peer-conversation cadence, build events, ambient awareness — which lives outside the host's token space entirely. Attend's wall-clock axis is correct for attend's domain for the same reason ways' token-position axis is correct for ways' domain: each axis matches the addressing unit of the thing being steered. + +There is a deeper reason wall clock earns its meaning in attend specifically — one that explains *why* ways cannot simply adopt the same axis even if it wanted to. **Wall clock only becomes a meaningful coordinate when there is more than one observer.** A single agent running its own monotonic token counter has a single progression axis, a 1-D line; there is no second axis to project onto, and wall clock adds no information that token position does not already encode. Ways is single-observer — one session, one monotonic token stream — so wall clock is superfluous. + +Attend introduces the second observer. Each peer has its own internal progression (its own session, its own token count, its own turn stream) that is disconnected from every other peer's. Peer A's "token position 5000" and peer B's "token position 3000" are not comparable — they live in different coordinate systems with different origins. The *only* axis that is shared across peers, and therefore the only axis on which their events can be meaningfully compared, is wall clock. "Peer A said X at t=100s, peer B said Y at t=102s" places both events in a common frame; "peer A at its own position 5000, peer B at its own position 3000" does not, because the positions are on non-overlapping number lines. + +This is a dimensionality argument: two disparate monotonic counts, when compared in the same space, require an additional shared axis to form a coordinate system in which their relationship can be expressed. Wall clock is that shared axis — not because time is metaphysically special, but because it is the coordinate frame that is guaranteed to be common across all peers regardless of their internal progression. Attend's wall-clock axis is not arbitrary; it is the *unique* axis that multi-peer coordination requires. In a single-observer system, that requirement does not exist, and wall clock collapses to a redundant label on the one monotonic that already exists. + +This is the clean division: single-observer systems need one progression axis (whatever the host uses internally — token position for transformer-hosted ways); multi-observer systems need wall clock on top as the shared coordinate frame that lets disparate internal progressions be compared. Both choices are emergent from the observer topology, not from any preference about units. + +### 5. Reactive firing via `postcheck.sh` + +Hooks are extended with `PostToolUse` and `PostToolUseFailure` matchers. A new hook script — `hooks/ways/check-post.sh` — is wired to these events. Each way may ship an optional `postcheck.sh` alongside its `macro.sh`: + +``` +hooks/ways/softwaredev/code/quality/ +├── quality.md +├── macro.sh (existing) +└── postcheck.sh (new, optional) +``` + +On `PostToolUse` / `PostToolUseFailure`, `check-post.sh` walks ways with a `postcheck.sh`, runs each with `tool_response` as stdin, and treats exit 0 as "request firing." The firing request then flows through the standard inward gate — the engine's `should_fire` consults the way's curve and current refractory state. Postcheck reactive requests are not privileged over predictive matches; both go through the same gate. + +Reactive firing uses the observed post-state, not intent. A `postcheck.sh` can check file size after a write, test exit codes after a test run, `gh pr` status after a merge, or any other metric its way cares about. The evaluation is metric-predicate style — cheap, deterministic, side-effect-free — not semantic re-embedding. + +The two modes compose without any extra machinery: a way that fires predictively at turn T (inward gate passes, salience = 1.0) will not re-fire reactively at turn T+1 unless salience has decayed below the re-injection floor. The engine already handles this via the curve query; the reactive path just adds another set of match sources. + +### 6. Shared-crate home and refactor strategy + +The firing-dynamics core lives in `sensor-trait` (or migrates to a renamed `firing-dynamics` crate if the scope outgrows the current name — naming question decided during implementation). Both attend and ways-cli depend on it. + +The refactor touches three crates in sequence: + +1. Rename progression types inside `sensor-trait` — `Instant`/`Duration` → `Tick`/`TickDelta` (both `u64`). Factor the curve enum. Adapt attend's internal uses to go through the renamed API. +2. Port attend's call sites (`SensorSlot::poll`, `ready_to_disclose`, `current_multiplier`, etc.) to supply `epoch_secs()` as the tick source. Existing attend parameters map onto `Curve::Exponential` (for signal salience) and `Curve::ActionPotential` (for engagement). Semantics preserved exactly. +3. Wire ways-cli to the shared crate. Per-way `EngagementState<Curve>` replaces the current flat `token_distance_exceeded` check. Tick source is `get_token_position()`. Existing `redisclose: N` frontmatter is parsed as `Curve::Flat { suppression: N }` for backward compatibility. + +### 7. Frontmatter schema migration — no shims + +Way frontmatter gains a required `curve:` field. The existing `redisclose: N` field is **removed from the schema entirely**. There is no shim layer, no dual-parser, no silent translation of old syntax. Existing ways are migrated to explicit `curve:` declarations as an explicit step of the implementation plan, and the old `redisclose` parser is deleted. + +```yaml +curve: + type: Exponential + half_life: 50000 # tokens +``` + +Or: + +```yaml +curve: + type: ProgressiveStaircase + steps: + - [0, 1.0] + - [15000, 0.5] + - [40000, 0.2] +``` + +Or (explicit step function, for ways that want discontinuous behavior): + +```yaml +curve: + type: Flat + suppression: 15000 # tokens +``` + +**Why no shim:** shims are carryovers from older, pre-unified thinking and carry a translation burden that obscures the actual choice each way is making. Forcing explicit `curve:` declarations during migration is a one-time cost that leaves every way in the codebase with a single, self-documenting firing model. The translation layer is tech debt the moment it ships. Every way must declare its curve; the migration step is part of the refactor, not a deferred clean-up. + +### 8. Empirical tuning: `ways tune` + +A new `ways tune` subcommand mirrors `attend tune`. It surveys recent sessions via `~/.claude/stats/events.jsonl` (which already captures way-fire events via `log_event` in `session.rs` and the `inject-subagent.sh` logging path), computes per-way cadence statistics, and suggests curve parameters grounded in the user's actual usage. `--apply` rewrites the relevant `curve:` entries in frontmatter or a centralized overrides file. + +Tuning parameters come from real session distributions, not guesses — the same discipline ADR-119 used when it sized attend's parameters to "Claude's actual turn cadence, not biological neuron kinetics." + +## Amendment — 2026-07-29: postcheck state handoff to its own macro + +Decision 5 describes postcheck evaluation as "metric-predicate style — cheap, deterministic, side-effect-free." That clause bundles three properties, and one of them is narrowed here: a postcheck may write **session-scoped state for the sole purpose of handing a finding to its own way's `macro.sh`**. Cheap and deterministic are unchanged, and the reason they survive is spelled out in constraint 4 below. + +**Why the channel is needed at all.** `check-post.sh` invokes each postcheck with `>/dev/null 2>&1` and reads only the exit code — ADR-135's amendment already recorded the consequence: "a postcheck's stdout is discarded... the content injected on a fire is the way's own body." `macro.sh` cannot close the gap either: `show/helpers.rs` runs it with no stdin and no arguments, so it never sees the tool input. `show/mod.rs` applies no trigger condition to macro execution, so a macro *does* run on a postcheck-triggered fire — the ordering (postcheck, then gate, then macro) is guaranteed by `check-post.sh`. A session-scoped stash is therefore the only channel by which a reactive way can name the file it is reacting to. Without it, every reactive way is limited to a static body, which is adequate for a way whose finding *is* its body (over-build names a replacement) and inadequate for one whose finding is a location. + +**Constraints, so this does not become a general escape hatch:** + +1. **Session-scoped only**, under `sessions-root.sh`. Not project files, not global state, not anything a later session or another agent inherits. +2. **Own-way only.** A postcheck writes state that its own way's macro consumes. Reading or writing another way's state is out of scope and would reintroduce the coupling Decision 5's separation was protecting. +3. **Bounded.** State is a set with a cap and a TTL, not an append-only log — because the inward gate may deny the fire that would have consumed an entry, so unconsumed entries are the normal case rather than an error. +4. **Never an input to the predicate.** The exit code stays a pure function of the observed post-state. State flows one way — postcheck writes, macro reads — so the *predicate* remains deterministic and re-runnable, which is the property the original clause was protecting. A postcheck that consulted its own prior state to decide whether to fire would be a genuine violation, not a narrowing. + +**A set, not a single slot.** `refire` keys on the way, not on the artifact the way is reacting to. With a single overwritten slot, a permissive cadence re-fires about a file the operator already declined to act on, and a gate denial leaves a stale path that a later, unrelated fire would name incorrectly. A seen-set addresses both: the macro names what has not yet been surfaced, and the way stays quiet about what has. + +Unchanged: reactive requests are not privileged over predictive matches, and both continue through the same inward gate. + +First application is the markdown line-handling way (#415), whose postcheck detects hard-wrapped prose in freshly written markdown and whose macro names the file and the remediation command. + +## Consequences + +### Positive + +- **Single source of truth for firing dynamics.** The math lives in one crate. Drift between attend and ways is structurally impossible. +- **RoPE-aligned semantics for ways.** The firing engine operates on the same axis the model's attention decays on. Firing decisions are mechanistically matched to how the model treats the injected content. +- **Progressive disclosure as a first-class pattern.** Not a workaround, not a special-case hack — a named curve variant that any way can opt into. +- **Reactive firing composes cleanly.** Post-tool-use evaluation uses the same inward/outward gates as predictive firing. No parallel system, no new state to reconcile. +- **Empirical tunability.** `ways tune` grounds curve parameters in the user's actual session distribution. Defaults are starting points, not final answers. +- **Future axes for free.** Turn count, commit count, line count — any future trigger surface supplies a monotonic and picks parameters in its own unit. No engine change. +- **Inward/outward separation preserves ADR-119/121's clarity argument.** "Should this new event fire?" and "Should this already-fired signal still be visible?" are different questions, answered against the same state but through different queries. + +### Negative + +- **Refactor footprint spans three crates.** `sensor-trait`, `attend`, `ways-cli`. The rename from time-flavored to progression-flavored types touches attend's entire engagement path. This must land as a coherent change or a carefully staged sequence — partial adoption leaves the code confused about what "tick" means. +- **Curve enum adds one match layer.** Cheap, but not free. Each `salience_at` / `multiplier_at` is a dispatch per call. +- **Tuning defaults are guesses until real session data is surveyed.** Initial ways parameters will be rough — the calibration pass after landing is where the real defaults come from. +- **Schema migration is a required step, not optional.** Every existing way must be rewritten to declare `curve:` explicitly before the new parser can land. This is one-time work with no graceful rollout — the old and new parsers cannot coexist. The upside is zero shim tech debt; the downside is one high-risk migration commit. +- **Exponential parameter conversion at the rename boundary is non-trivial.** Attend's existing parameters are expressed as `decay_per_minute` rates; converting them to `half_life` in ticks requires `half_life = ln(0.5) / ln(1 - rate_per_minute)` then a unit conversion into seconds. For `decay_per_minute = 0.1`, the half-life is `ln(0.5) / ln(0.9) ≈ 6.58 minutes ≈ 395 seconds`. **This is not `0.1 / 60`** — that would only be correct for linear decay; exponential decay requires the compounding formula. Every attend parameter needs recomputation through this formula at the rename boundary, and the migration must validate behavior is preserved (not just syntactically, but numerically). + +### Neutral + +- **Not backward-compatible — and intentionally so.** `redisclose: N` is removed. Every existing way migrates to explicit `curve:`. This is tech-debt-free by construction; the ADR commits to one coherent schema rather than a dual-parser shim. +- **Crate naming question deferred.** The firing-dynamics core lives somewhere — either `sensor-trait` (kept as a home) or a renamed/extracted crate. Decision falls out of implementation, not ADR. +- **Reactive firing can land in stages.** The progression-axis unification and the `PostToolUse` hook extension are independent work items that can sequence separately if staging helps. This ADR establishes both as a single architectural direction; the implementation can still split along natural seams. + +## Alternatives Considered + +### `Clock` trait generic over `Instant` and `TokenPosition` + +Parameterize the engine over a `Clock` trait with associated types for `Time` and `Duration`. Each tool supplies a `Clock` impl. Rejected because this keeps time semantics inside the engine — the type names themselves ("Clock", "Time", "Duration") imply wall-clock reasoning, and the engine has no business reasoning about wall clock for ways. The progression-flavored rename is the same refactor without the conceptual leak. + +### Separate implementations per tool + +Keep attend's engagement model in `sensor-trait`, add a parallel model to `ways-cli`. Rejected because the math is identical; two implementations would drift, and the code reuse Aaron surfaced in this conversation is real and worth capturing. + +### Turn count as ways' axis + +Use the number of user turns as ways' progression unit, matching ADR-121's choice for attend signals. Rejected because token position is mechanistically matched to the transformer's attention decay (RoPE) and turn count is not. A think turn and a short reply both count as one turn but displace different amounts of context; token position is the exact measure. + +### Wall clock for ways + +Use wall-clock seconds for ways' axis, matching attend. Rejected because the model's attention mechanism has no wall-clock dimension. Firing decisions should be made on an axis the model actually observes. Wall clock is external to the thing ways is shaping. + +### Keep flat `redisclose`, add reactive firing only + +Land `postcheck.sh` and `PostToolUse` hooks without the dynamics unification. Rejected because flat suppression cannot express the inward/outward gate separation that ADR-119 and ADR-121 demonstrated is necessary — and because the reactive firing path needs the same inward gate to avoid spam-firing on every tool call. + +### LLM-based reactive evaluation via `type: prompt` hooks + +Use Claude Code's `"type": "prompt"` hook as a PostToolUse filter. Run an LLM pass on the tool result and ask which ways to fire. Rejected as the default path because of cost per tool call. Reserved for specific high-value hooks (likely `PostToolUse:Task` on subagent returns, where the tool call is expensive enough that one more LLM eval is noise). The primary reactive path is `postcheck.sh` metric predicates — cheap and deterministic. + +### Tick-windowed burst detection + +Use `burst_window: TickDelta` to scope which history entries count toward a burst, matching ADR-119's original shape. Rejected because progression axes with chunky advancement (token position in ways) can jump thousands of ticks in a single event, swallowing any reasonable window in one step and breaking burst detection. Event-count-based windowing with the implicit bound defined by the decay curve's multiplier collapse is robust to axis granularity — it works identically for attend's smooth seconds and ways' chunky tokens. + +### Backward-compatible `redisclose → Flat` shim + +Keep `redisclose: N` in the schema as sugar that translates to `Curve::Flat { suppression: N }`. Rejected on Aaron's direction: shims are carryovers from older, pre-unified thinking. Forcing an explicit `curve:` migration is a one-time cost that leaves the codebase self-documenting. A translation layer obscures what each way is actually declaring and would need to be cleaned up eventually anyway. + +## Implementation Plan + +1. **sensor-trait rename pass.** Introduce `Tick: u64`, `TickDelta: u64`. Replace `Instant` / `Duration` with `u64` throughout `EngagementState`. Remove `decay_per_minute` and `burst_window` as engine concepts — they move into `Curve::ActionPotential` as `multiplier_half_life` and are derived from event-count windowing. +2. **Curve enum.** Factor the existing burst/multiplier/decay math into `Curve::ActionPotential` (with event-count burst detection), `Curve::Exponential`, `Curve::ProgressiveStaircase`, and `Curve::Flat`. Expose `salience_at` and `multiplier_at` as the two queries. All decay across all variants is parameterized by `half_life`. +3. **attend parameter conversion.** Convert attend's existing `decay_per_minute` values to `multiplier_half_life` in ticks using `half_life_seconds = (ln(0.5) / ln(1 - rate_per_minute)) × 60`. Recompute every engagement parameter; do not simple-divide. Validate preserved behavior via existing refractory tests before renaming anything else. +4. **attend migration.** Adapt `SensorSlot` and its consumers to the renamed API. Tick source wraps `SystemTime::now()` → seconds-since-epoch as `u64`. Verify refractory and salience behavior is unchanged end-to-end on a real session replay. +5. **ways frontmatter migration (required, no shim).** Rewrite every existing way's frontmatter to declare explicit `curve:`. Ways with `redisclose: N` become either `Curve::Flat { suppression: N }` (if N was chosen for step-function reasons) or `Curve::Exponential { half_life: N }` (if smooth decay is wanted). Remove the `redisclose` field from every way and from the schema. +6. **ways-cli integration.** Per-way `EngagementState` in `session.rs`. Tick source is `get_token_position()`. Replace `token_distance_exceeded` with `curve.salience_at(delta) >= floor`. Delete the old `REDISCLOSE_PCT` constant and its consumers. +7. **Frontmatter schema update.** Extend `frontmatter.rs` to parse the new `curve:` block. Delete the `redisclose` parser entirely — no dual-parser. Lint-ways validates the new schema and errors loudly on any leftover `redisclose:` field. +8. **Hook wiring.** Update `check-prompt.sh`, `check-task-pre.sh`, `check-file-pre.sh`, `check-bash-pre.sh` to consult the engine's inward gate before injecting. Stamp tick on fire via the renamed API. +9. **Reactive firing path.** Add `PostToolUse` and `PostToolUseFailure` to `settings.json`. New `check-post.sh` walks ways with `postcheck.sh`, runs each, pipes through the inward gate. +10. **`ways tune`.** Subcommand that surveys `events.jsonl`, computes per-way cadence, suggests curve parameters. `--apply` writes them into frontmatter. +11. **Empirical calibration.** After landing, run `ways tune` against real session data and commit the calibrated defaults as the production starting point. + +## Open Questions + +- **Crate home.** Generalize `sensor-trait` in place, or extract a new crate? Lean: keep in `sensor-trait` until the scope clearly outgrows the name. Renaming is a follow-up, not this ADR. +- **Curve dispatch.** Enum (chosen here) vs trait object. Enum wins on serialization and cheap matching; trait object would be needed only if third-party curves matter. Not yet. +- **Default curve for ways with no `curve:` field after migration.** After the migration step deletes `redisclose`, the schema could either (a) require every way to declare `curve:` explicitly and error on omission, or (b) provide a sensible default (e.g., `Curve::Exponential { half_life: <25% of context window in tokens for the active model> }`). Lean: (a). Explicit over implicit, matches the "no shim" directive. +- **Whether `ProgressiveStaircase` ships in the initial implementation or lands as a follow-up once `Exponential` and `ActionPotential` are proven portable.** Lean: ship it in the initial implementation because it demonstrates that the curve-as-parameter shape actually buys something beyond refactoring. +- **Validation test for attend parameter conversion.** The exponential-decay conversion formula is straightforward, but validating that the renamed attend produces numerically equivalent refractory/salience curves on real session replays is the actual acceptance criterion. How is that validated — side-by-side simulation, golden trace, or live comparison? Decide during implementation. +- **Epsilon for "event has decayed out of burst consideration".** Event-count burst detection needs a small threshold below which a history entry stops contributing. `multiplier > 1.01` is a reasonable first pick but arbitrary. Empirical tuning via `ways tune` should inform the final value. + +## References + +- **ADR-113** — attend active awareness module; origin of the disclosure governor. +- **ADR-114** — attend as insistent way trigger type; the integration point. +- **ADR-117** — sensor crate extraction and feature flags; precedent for cargo-workspace splits. +- **ADR-119** — action potential engagement model; the inward gate, superseded by ADR-123 and now rewritten in place to describe the shipped unified engine. +- **ADR-121** — salience decay for signal presentation; the outward gate, superseded by ADR-123 and now rewritten in place to describe the shipped unified engine. Ways is the first concrete consumer; attend sensor-peers application still deferred. +- **Cognitive Frameworks paper** — cognitive economics principle (`cheapest path = correct path`); the reason ways wants model-aligned decay rather than a parallel steering layer. +- **Rotary Positional Embedding (RoPE)** — Su et al., *RoFormer: Enhanced Transformer with Rotary Position Embedding* (arXiv:2104.09864) — the mechanism that makes token position the correct axis for ways. + +## Implementation Status + +**Accepted 2026-04-14** after Phase F validation in [`docs/hooks-and-ways/observed-behavior.md`](../../hooks-and-ways/observed-behavior.md) confirmed the unified stack preserves the 2026-03-17 baseline (supertask held under detour, ~21% context use, zero operator redirections across 92 turns). + +### Shipped + +- **Phase A — Foundation (1f30b80).** `sensor_trait::Curve` enum with four variants (`Exponential`, `ActionPotential`, `ProgressiveStaircase`, `Flat`). `EngagementState` owning a `Curve` plus tick-keyed history. `Tick`/`TickDelta` aliases. `salience_at` (outward gate) and `multiplier_at` (inward gate) as the two query methods. Event-count burst detection replaces time-windowed burst detection for chunky-axis robustness. 26 unit tests in `sensor-trait`. +- **Phase B — Attend migration (9340185).** `SensorSlot::engagement` ported to the new `EngagementState`. `decay_per_minute` converted to `multiplier_half_life` via `ln(0.5) / ln(1 - rate) × 60`. Wall-clock-seconds tick source via `sensor_trait::epoch_secs()`. Old `Instant`/`Duration`-keyed state and its unit tests deleted. B3 parity check validated live during the Phase F observation session — absolute refractory, relative suppression, governor cooldown, and graceful binary reload all hold under the new engine. +- **Phase C — Ways integration (b94cc25, ad20744, 47bc9de).** `curve:` field added to frontmatter schema; `redisclose:` rejected at parse time via `detect_legacy_redisclose()`. 99 existing ways migrated (8 with explicit `redisclose: N` mapped to `Curve::Exponential { half_life: N × 2000 }` on a 200k baseline; 91 defaulting to `Curve::Exponential { half_life: 30000 }`). Per-way `EngagementState` persisted at `{session_dir}/way-engagement/{way_id}.json`. `REFIRE_FLOOR = 0.5` as the outward-gate cutoff. `REDISCLOSE_PCT`, `token_distance_exceeded`, and `detect_context_window` deleted. Per-way visualization in `ways list` and `ways rethink` driven by `Curve::refire_delta(floor)`. +- **Phase D — Reactive firing (dddb4ca).** `PostToolUse` and `PostToolUseFailure` hooks wired in `settings.json`. `hooks/ways/check-post.sh` walks ways with `postcheck.sh`, pipes `tool_response` through each, and treats exit 0 as "request firing" — which then flows through the same inward-gate check as predictive fires. `hooks/ways/softwaredev/code/quality/postcheck.sh` is the load-bearing demo, firing when an Edit or Write grows a file past 500 lines. Verified live during the Phase F session: the quality way fired reactively on `tune_curves.rs` at epoch 83 when that file crossed the threshold mid-edit, which triggered the `--days` scope trim as an in-session refactor. +- **Phase E — Empirical calibration (28527e8).** `ways tune-curves` subcommand surveys `~/.claude/stats/events.jsonl`, pairs `way_fired` + `way_redisclosed` events by (way, session), computes token-position deltas between consecutive fires, and suggests calibrated `Curve::Exponential` half-life values per way. Dry-run by default, `--apply` rewrites `curve:` blocks via line-surgery. Floor of 3 delta samples, ±20% tolerance band, round to nearest 500 tokens. Verified against the live events log: 2 ways with early data, both within ±20% of migration defaults — the initial tuning guesses were well-calibrated for the ways that have accumulated samples. +- **Phase F — Validation (4672e44).** Re-observation written into `docs/hooks-and-ways/observed-behavior.md` from the session that landed Phases B3 through E. 92 turns, 212K / 1000K tokens, zero operator redirections, several in-session detours handled without supertask drift, Phase D reactive firing and check-firing decay both verified in-flight. ADR-123 stack preserves the 2026-03-17 baseline. +- **Cross-tool lint hygiene (1397d68, f360ea6, c869841).** `ways lint` split from one 938-line file into 7 focused modules with `--fix` removing foreign top-level fields and `when:` sub-fields, `x-*` escape hatch, and locale stub validation for 83 newly-covered files. New `attend config lint` subcommand with matching semantics (UNKNOWN / DEPRECATED, `--fix`, `--check`) — `engagement.burst_window` is the first DEPRECATED entry with a pointer to ADR-123's "burst window is implicit in multiplier_half_life" reframing. +- **Docs (057f3f3, d991209).** `docs/attend-and-monitor/engagement.md`, `context-decay.md`, `loop.md`, `configuration.md`, and `salience.md` all rewritten in place for the ADR-123 framing. The old "two mechanisms with different units" narrative in salience.md becomes "one engine, two queries against one state." ADR-104, 119, and 121 rewritten in place with `Status: Superseded` + `superseded_by: ADR-123`, their original Context and Alternatives preserved as the historical decision trail. +- **Secondary fixes (679b205).** `attend config show` no longer displays a synthesized `burst_window` default when the user's file doesn't declare the key. Phase 1 of the soft-removal; Phase 2 (delete the field from the schema entirely) tracked as issue #50 for after a few real-usage cycles let the deprecation warning do its work. +- **`burst_window` phase 2 removal (issue #50).** `EngagementConfig::burst_window`, its parser branch, its DEPRECATED-keys entry, and the `attend config show` conditional display row are all deleted. `config::detect_legacy_burst_window` now hard-rejects the key at load time with a pointer to `attend config lint --fix`, mirroring ways' `detect_legacy_redisclose` shape. The lint fixer still removes the line in place via the UNKNOWN path, so the soft-migration rhythm (warn → fix → hard reject) is complete. + +### Deferred + +- **Attend sensor-peers outward-gate application.** ADR-121 step 7 in the superseded decision, still pending. The engine is in place; what's missing is a per-signal `EngagementState` in sensor-peers and a `current_salience >= floor` check in the presentation path. See [`docs/attend-and-monitor/salience.md`](../../attend-and-monitor/salience.md) for the sketch of the path forward. Pure new consumer of an existing facility — no engine work needed. +- **`attend status` refractory state display.** ADR-119 step 7 in the superseded decision. Should show current multiplier and refractory state per sensor in the existing `attend status` table. Not load-bearing for ADR-123; tracked as a follow-up. +- **Motivation / reflection-overdue sensor.** ADR-119 steps 8–9. A new sensor that emits time-based sub-threshold stimuli to drive intrinsic self-prompting. Architecturally available under ADR-123 (just another consumer of the engine) but not scoped to this PR. +- **`attend config lint --check` in CI.** Once the config-lint path has a few real uses under its belt, wire it to a CI job alongside `ways lint --check`. Tracked on the PR. +- **`tools/attend/src/main.rs` split.** 2027-line dispatcher file flagged by the quality way during this PR. Pure structural refactor, proposed module layout in issue #51. Do after the attend sensor-peers application lands (or any time it's convenient) — no dependency on ADR-123. +- **`tools/ways-cli/src/main.rs` split.** Pushed past the 500-line review threshold by the `TuneCurves` enum variant in 28527e8. Mirror case of the attend main.rs split, much smaller. Roll into whatever PR sets the pattern for attend (noted on issue #51). + +### Open questions resolved during implementation + +The ADR's Open Questions section is preserved above for the historical record. The actual resolutions: + +- **Crate home:** `sensor-trait` generalized in place. No rename. +- **Curve dispatch:** enum with serde derive. No trait object. +- **Default curve for ways:** explicit `curve:` required; parse-time rejection of legacy `redisclose:` via `detect_legacy_redisclose()`. No default, no shim. +- **Whether `ProgressiveStaircase` ships initially:** yes, as one of the four enum variants. No production way uses it today, but its presence proves the curve-as-parameter shape buys something beyond refactoring. +- **Validation test for attend parameter conversion:** live session observation (the Phase F re-observation) rather than side-by-side simulation or golden trace. The conversion formula's correctness was validated by the fact that attend's cadence behavior matched pre-refactor expectations in a real workload. +- **Epsilon for "event has decayed out of burst consideration":** deferred to empirical tuning via `ways tune`. Current default in `Curve::ActionPotential` is left at the ADR's first-guess value and will be revisited once the event log accumulates enough token-position-enriched fires for `ways tune-curves` to suggest per-way values. diff --git a/docs/architecture/ways/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md b/docs/architecture/ways/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md new file mode 100644 index 00000000..ce59168b --- /dev/null +++ b/docs/architecture/ways/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md @@ -0,0 +1,136 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +supersedes: + - ADR-107 +amends: ADR-108 +basis: + - evidence: Russian-locale feedback and six failing languages in make test-multilingual, traced to 1411 uncalibrated per-locale embed_threshold entries + - evidence: BM25 uses an English-only stemmer (Algorithm::English in bm25.rs) and bypasses the graph + - precedent: ADR-107 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-17 +deciders: + - aaronsb + - claude +related: + - ADR-105 + - ADR-107 + - ADR-108 + - ADR-110 +imported: + from: docs/architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md + format: v0 + status: Accepted +--- + +# ADR-125: Authored Disclosure Graph and Removal of BM25 + +## Context + +The way matching system has accumulated three matching tiers (embedding → BM25 → keyword/regex) and two parallel threshold systems (English `embed_threshold` in frontmatter, per-locale `embed_threshold` in `.locales.jsonl`). The original ADR-107 specified one threshold per way; in practice, 1411 per-locale entries now carry their own values, individually estimated and uncalibrated. + +Recent multilingual feedback (Russian locale, six failing languages in `make test-multilingual`) traced to per-locale thresholds set higher than the multilingual model's actual cosine scores against real user queries. The existing `ways tune` command can't fix this — it measures stub-versus-corpus discrimination, not real-query recall, and produces "0 would change" across all 1411 entries while leaving the user-facing miss rate intact. + +The deeper issue is architectural drift, not a calibration bug. Several concepts are already present in the system but under-named: + +- ADR-105 established progressive disclosure (parents fire, children become reachable) +- ADR-110 established the way graph (nodes, edges, `ways-graph.jsonl` export) +- ADR-107 added locale stubs as packed JSONL aliases on each way +- ADR-108 added the multilingual embedding model alongside the English-only one + +These four ADRs describe a single architecture — an authored DAG of ways, with embedding-coordinate aliases on each node, and disclosure semantics governing which subgraph is live in a session — but no ADR names that architecture explicitly. Without the name, each new feature was added as a localized patch (per-locale thresholds, per-tier fallbacks, per-language tuning logic) rather than as a property of the underlying model. + +BM25 illustrates the cost of the unnamed model. It is a lexical tier that bypasses the graph entirely, uses an English-only stemmer (`Algorithm::English` in `bm25.rs:175`), and provided value primarily as a fallback when the embedding model was absent. With the embedding model now downloadable on four platforms via CI release artifacts, the fallback is rarely exercised, and extending BM25 to multilingual would require per-language stemmer wiring that fights the alias model. + +## Decision + +Name the architecture and remove what doesn't belong in it. + +### 1. Authored disclosure graph + +The way library is an **authored disclosure graph**: + +- **Nodes** = ways (one file per node, unique filenames per ADR-110) +- **Edges** = parent/child (directory tree), siblings (cosine-weighted, computed by `ways siblings`), explicit `See Also` references +- **Coordinate aliases** = each node carries one or more embedding-space coordinates. The English content (frontmatter description + vocabulary + prose) produces the canonical alias. Locale stubs in `.locales.jsonl` produce additional aliases. Project-local extensions and domain-specific phrasings are also aliases under this model — multilingual is one application, not the headline. +- **Disclosure state** = the live subgraph reachable from the session frontier. Progressive disclosure (ADR-105) is the traversal mechanism over this graph. + +"Authored" distinguishes this from LLM-extracted variants: nodes and aliases are written by humans (or by Claude under explicit authoring instruction), not derived at indexing time. + +### 2. Embedding model as a black-box singularity + +Retrieval is embedding-only. The multilingual model is treated as a black box: we accept its score distribution as the ground truth and design the system around that boundary, rather than augmenting it with lexical tiers that try to compensate for what the model "should" have done. + +A node's match score against a query is: + +``` +node_score(query, node) = max over aliases A of cosine(embed(query), embed(A)) +``` + +The `max` collapses node-aliasing into a single per-node score. A user query in Russian, a user query in English, and Claude's own native-language tool-call utterance all reach the same node through whichever alias is closest. + +### 3. BM25 is removed + +`bm25.rs`, BM25 score thresholds (`threshold:` in frontmatter when used for BM25), the two-tier fallback logic, and the engine-selection-by-model-availability code are removed. The embedding model becomes a hard dependency of `ways`. + +### 4. One threshold per node, in English frontmatter + +The per-locale `embed_threshold` field in `.locales.jsonl` is removed. Each node has at most one `embed_threshold` field, in the English frontmatter. Nodes that omit it use a system default. + +Stub fidelity becomes a measurable graph property: + +``` +fidelity(node, alias) = cosine(embed(canonical_alias), embed(alias)) +``` + +A low-fidelity alias is one whose embedding sits far from the canonical; the fix is to re-author the alias text (re-translate, fix vocabulary), not to lower a per-alias gate. `ways tune` is rewritten to measure and report fidelity, not to tune per-alias thresholds. + +### 5. Explicit triggers survive + +The `pattern:` and `commands:` regex fields in way frontmatter are not BM25 — they are explicit, deterministic triggers. They survive the tier removal and remain the override mechanism for "fire this way exactly when this pattern appears." + +## Consequences + +### Positive + +- **One architectural model, named.** Future ADRs can extend "authored disclosure graph" instead of inventing parallel concepts. The model composes: aliases for non-language extensions (project-local, domain-specific) are now well-typed. +- **Threshold surface collapses from ~1411 dials to ≤83.** The reporter's Russian-locale bug (and all six failing languages in `make test-multilingual`) become a one-line consequence: per-alias gates were never the model. +- **Stub quality becomes measurable.** Fidelity is a number; low-fidelity aliases are visible in audit output and direct re-authoring effort to where it matters. +- **Matcher pipeline simplifies.** One tier, one scoring function, no engine selection. The black-box framing also stops the temptation to add lexical patches when the embedding behavior surprises us. +- **Removes ~English-only assumptions in the matcher.** With BM25 and its `Algorithm::English` stemmer gone, the matcher has no English-special-case code paths. + +### Negative + +- **Embedding model is a hard dependency.** `ways` cannot match without it. Setup must succeed at fetching/building the model; offline or air-gapped installs need the model present. Mitigation: model is 127MB, four-platform CI artifacts are already shipped, and `make setup` is the single command users run. +- **Per-locale calibration tweaks are no longer possible.** A language whose stub embeddings cluster lower than English's must be addressed by re-authoring the stub (raising fidelity) or lowering the node's threshold globally — there's no per-language escape valve. This is intentional, but it raises the bar on stub authoring quality. +- **Migration touches 1411 lines across ~83 `.locales.jsonl` files.** Mechanical (delete `embed_threshold` field), but it does change every locale entry on disk. + +### Neutral + +- The corpus generator still emits per-alias rows; the runtime just stops reading per-alias thresholds. +- `ways graph` and `ways siblings` (already shipped per ADR-110) are unchanged — they were already operating in the model this ADR names. +- `ways tune --audit` continues to surface ambiguous nodes (those whose canonical alias is confused with neighboring nodes); the audit's value increases now that fidelity is the explicit metric. + +## Alternatives Considered + +- **Per-locale calibration via `ways tune --from-queries`.** Curate query-set-per-language, calibrate thresholds to admit real queries, leave per-locale dials in place. Rejected — keeps the per-locale dial surface (1411 entries), shifts the calibration burden to query curation per language, and doesn't address the deeper drift from ADR-107's original intent. +- **Keep BM25 as a fallback for offline/no-model installs.** Rejected — BM25's English-only stemmer makes it actively misleading for multilingual users (silently degrades to "no match" rather than "incorrect match"), and maintaining a parallel matching path for the rare offline case is not worth the architectural cost. +- **Universal global threshold (single number, all nodes).** Considered, rejected for now — different ways have different score distributions against natural queries (broad ways like `ea` score lower than narrow ways like `delivery/commits`), so per-node thresholds carry real signal. May revisit once we have empirical data on whether per-node tuning matters in practice. +- **Coordinate-only ways (drop text, store vectors as source of truth).** Rejected — destroys human auditability of way intent. Authoring needs to happen in text; embeddings are derived. +- **Adopt "GraphRAG" as the architectural label.** Rejected — accumulates baggage from the Microsoft variant (LLM entity extraction, hierarchical community detection, expensive offline preprocessing) that doesn't match what we do. "Authored disclosure graph" describes the actual mechanism without importing the framework's reputation. + +## Migration + +1. Delete `bm25.rs` and BM25-related code paths in `ways-cli` +2. Remove BM25 `threshold:` fields from way frontmatter (migration script: identify which `threshold:` values were BM25 vs. other uses) +3. Strip `embed_threshold` from all `.locales.jsonl` entries (sed-able, one-time change) +4. Rewrite `ways tune` to measure and report alias fidelity (cosine to canonical), not per-alias thresholds +5. Update ADR-107 status to `Superseded by ADR-125 (in part — locale support model)` +6. Update ADR-108 to note BM25 fallback is removed; embedding tier is sole tier +7. Update `make test-multilingual` to verify the Russian and other failing-language queries now resolve via the alias model, with no per-locale threshold tuning required diff --git a/docs/architecture/ways/ADR-126-window-relative-refire.md b/docs/architecture/ways/ADR-126-window-relative-refire.md new file mode 100644 index 00000000..072ac7f7 --- /dev/null +++ b/docs/architecture/ways/ADR-126-window-relative-refire.md @@ -0,0 +1,363 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: disclosure +basis: + - evidence: in a 4.3-hour, 1M-token session on 2026-04-20 the softwaredev/code/quality way fired 5 times, re-injected roughly every 200k tokens + - evidence: '95 ways at half_life: 30000 calibrated for a 200k window; the narrow re-tune in PR #70' + - precedent: ADR-123 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-04-22 +deciders: + - aaronsb + - claude +related: + - ADR-115 + - ADR-121 + - ADR-123 +imported: + from: docs/architecture/system/ADR-126-window-relative-refire.md + format: v0 + status: Accepted +--- + +# ADR-126: Window-relative refire with named presets + +## Context + +Way frontmatter currently carries a raw token count for refractory decay: + +```yaml +curve: + type: Exponential + half_life: 30000 +``` + +That number was calibrated when Claude Code's context window was 200k and +compaction forced rotation at ~160k. In a 200k session, `half_life: 30000` gave +ways ~3 re-fire chances across the full session — the intended "load-bearing +but not spammy" cadence. + +The context window has since grown. Opus-4 runs 1M-token sessions. In a +4.3-hour, 1M-token session on 2026-04-20, the `softwaredev/code/quality` way +fired 5 times — its static heuristic table and rationalization table were +re-injected into context roughly every 200k tokens. The reporter described this +as "light cognitive overhead per incident, cumulative across the session." The +mechanism is working as designed; the tuning is stale. + +The stale-tuning problem is not a one-time mistake that a sweep fixes. Every +context-window expansion re-invalidates every hand-tuned `half_life` in the +tree. With 95 ways at `half_life: 30000` today, each future window change is a +95-file sweep against a moving target. + +The deeper issue is that `half_life` as a raw token count mixes two concerns in +the frontmatter: + +1. **Author intent** — "this payload should re-disclose rarely" vs "this + payload fires on each new occurrence of its trigger." +2. **Host calibration** — "what does 'rarely' mean in tokens given this + session's context window?" + +Authors are forced to know the host's context window to express intent. That +coupling breaks every time the host's window changes. + +ADR-123 already made ways' axis unit-agnostic at the engine level: +`EngagementState` consumes a monotonic `Tick`, and callers supply the unit. +Callers in practice supply `get_token_position(session_id)` — token position on +a known context window. That's enough primitive to express intent as a +fraction rather than a count. + +A narrow-tune remediation shipped on 2026-04-22 (PR #70, ADR-127): 14 +static-heavy ways in `softwaredev/code/*`, `docs/standards`, and `architecture` +parents bumped from `half_life: 30000` → `half_life: 200000`. That addressed +today's pain for the confirmed problem class without prejudging this ADR. The +sweep was a patch; this ADR is the structural fix. + +## Decision + +Introduce a `refire:` frontmatter field that accepts either a number or a +preset name. The engine resolves to a concrete `half_life` in caller ticks at +fire-evaluation time, so ADR-123's unit-agnostic boundary is preserved. + +### Frontmatter shape + +`refire:` accepts two forms. Both are valid, choice expresses intent: + +```yaml +refire: 0.2 # direct: fraction of context window, pinned to today's model +``` +```yaml +refire: rare # reference: tracks the project's preset config +``` + +- A **number** in approximately `[0.0, 1.0+]` is interpreted as a fraction of + the session's context window. Half-life = `refire × window_size`. Writing a + number is an explicit choice to pin the cadence to today's model — it does + not auto-scale if the operator later swaps to a model with different + attention characteristics. +- A **string** is looked up in the `refire_presets` config section at fire + time. Unknown names fail closed at two upstream gates — `ways lint` and + `ways corpus` both reject unknown preset names and abort/warn. Fire-time + resolution has a fail-soft fallback to `normal` (0.15) with a stderr + warning, so a bypassed lint doesn't crash a live session, but the error + path is intended to be caught upstream. + +### Preset configuration + +A new section in the existing config file (`$XDG_CONFIG_HOME/ways/config.yaml` +with `$PROJECT/.claude/ways.yaml` overlay — see ADR-115): + +```yaml +# Refire presets. Each is a fraction of the session context window. +# Half-life = preset × context_window. See ADR-126. +refire_presets: + once: 1.0 # effectively once per session (never re-fires before end) + rare: 0.4 # static-heavy, 1–2 fires per session + normal: 0.15 # load-bearing, ~3 fires per session (matches pre-ADR default) + frequent: 0.05 # procedural, fires often relative to session +``` + +Built-in defaults ship with these four presets. User and project configs can +override individual values or add new preset names (e.g., `perpetual: 0.01`). +No schema migration needed. + +### Portability + +**Framework portability (strong claim).** Expressing refire as a fraction of +session capacity generalizes across any agent harness that can report its +capacity. Every agent harness has a finite context, and fractions generalize +over finite capacities by construction. A way tagged `refire: 0.15` means +"15% of whatever this framework calls a session" and works wherever +`session_capacity()` is callable. Ways become portable across LLM vendors and +agent frameworks, not just across Anthropic model generations. + +**Model-generation portability for presets (weaker sub-claim).** The preset +table embeds an additional assumption: the relative cadence between presets +(`rare` < `normal` < `frequent`) stays roughly portable across Anthropic model +generations, even as absolute attention characteristics shift. If a future +model shows materially different mid-context recall, the preset values are +re-tuned globally in the config file — no way-file sweep. If the sub-claim +breaks (e.g., one model wants the *ordering* inverted), the presets move to +per-model tables. Not needed now; structurally available later. + +Authors who don't trust the preset sub-claim for a specific way write a +number instead of a name. That's the escape hatch. + +### Resolution semantics + +Preset names resolve **at fire time**, not parse time. The flow: + +1. Frontmatter loader reads `refire:` as an enum: `Numeric(f64)` or `Preset(String)`. +2. Each fire evaluation: + - Fetch current context window via `cmd/context.rs::model_to_window()`. + - If `Numeric(v)`: `half_life = (v × window).round() as u64`. + - If `Preset(name)`: look up in `config::global().refire_presets[name]`, + then multiply by window. + - Build a fresh `Curve::Exponential { half_life }` and hand it to the engine. +3. Engine (sensor-trait) sees only concrete `Curve::Exponential`. No new + variant. No context-window threading through `way_fire_outcome`. + +Fire-time resolution means config edits take effect mid-session — operators can +tune presets, re-run, observe, without restarting. The config file is tiny; +re-reading or mtime-caching per fire is negligible. + +### Engine changes + +None required to `sensor-trait`. The resolver lives entirely in `ways-cli`. +`Curve::Exponential` is the only shape the engine ever sees for refire-bearing +ways. `attend`'s exhaustive matches on `Curve` are untouched. + +### Frontmatter and lint changes + +1. **`frontmatter.rs`** — add `RefireSpec` enum: + + ```rust + pub enum RefireSpec { + Numeric(f64), + Preset(String), + } + ``` + + Parse `refire:` as a scalar; if it's a float, `Numeric`; if it's a string, + `Preset`. If `curve:` is also present, `refire:` wins and lint emits a + warning about the duplication. + +2. **`config.rs`** — extend `Config` with `refire_presets: HashMap<String, f64>`, + populated from the new YAML section. Built-in defaults match the table + above. + +3. **`ways lint`** — four diagnostics covering presence, shape, and drift: + - **Warning** when a fire-bearing way (any trigger channel wired: + description+vocabulary, pattern, files, commands, or trigger; not an + attend handler; not a check file) has no `refire:` field. + - **Warning** when both `refire:` and legacy `curve:` coexist, pointing + authors at the `curve:` block to remove. + - **UNKNOWN (foreign-field warning)** when a `curve:` block is present + alone. The schema no longer lists `curve:` as a valid field, so the + existing unknown-field logic flags it without special-case code. + - **Error** when `refire:` is malformed — a numeric value outside + `[0.0, 10.0]` (e.g., a raw token count like `30000` accidentally + pasted in) or a preset name not present in `config.refire_presets`. + Fail-closed: lint is the primary typo gate. + +4. **`ways corpus`** — echoes the malformed-refire check as a stderr + WARNING during corpus generation. Corpus is run frequently (CI, local + rebuilds) and hits every way file, so catching typos here prevents them + from reaching a live session even when lint isn't invoked. + + These checks give authors an unambiguous signal about whether a way file + conforms to the post-ADR-126 shape at two independent gates (lint + + corpus), with a fire-time fallback that keeps sessions running even + when both gates are bypassed. + +## Migration + +**Phase 1 — engine and parser (now).** Ship `RefireSpec` parsing, config +extension, fire-time resolution. Accept both `refire:` (new) and `curve:` +(legacy) unchanged. No way files change yet. + +**Phase 2 — mechanical numerical conversion.** Convert each way using its +window-at-tuning-time as the reference: + +``` +refire = half_life / window_at_tuning_time +``` + +Two buckets exist in the tree: + +- **The 14 files tuned on 1M Opus** (ADR-127 narrow-tune, committed in + PR #70 `f93bb74`): `code.md`, `quality.md`, `errors.md`, `performance.md`, + `security.md`, `auth.md`, `injection.md`, `secrets.md`, `supplychain.md`, + `mocking.md`, `tdd.md`, `testing.md`, `architecture.md`, `standards.md` → + `refire = half_life / 1_000_000` → `refire: 0.2`. +- **All other ways, still at 200k-era values**: `refire = half_life / 200_000` + → `refire: 0.15` for `half_life: 30000`, proportional for other values. +- **Raw `curve:` blocks** for `Flat`, `ActionPotential`, `ProgressiveStaircase` + stay untouched — `refire:` is specifically the Exponential shorthand. + +The rule captures each author's real intent at the moment of tuning as a +fraction of the window they were thinking in. This preserves original design +intent across the tree. + +Side-effect: the 81 unhacked files that are currently broken on 1M (firing +~22 times per session instead of the designed ~3) will fire ~4 times once the +resolver multiplies `refire: 0.15` by the actual 1M window. An unintentional +but welcome fix, consistent with the intent the author expressed when they +wrote `half_life: 30000` against a 200k window. + +Second-order effect: three attend handlers (`meta/attend/build-complete`, +`context-pressure`, `reflection-overdue`) inherited `refire: 0.15` from the +migration. On 1M Opus, their effective suppression grows from 30k to 150k +half-life — a signal-debounce change, not a disclosure-decay change. This may +be too sticky for rapid-signal handlers; individual attend handlers can be +re-tuned per-signal (e.g. `refire: 0.03` to preserve the pre-migration 30k +behavior on 1M). Separate concern from the primary way-disclosure cadence +this ADR addresses. + +**Phase 3 — opt into presets (per-way judgment).** For ways whose intent is +genuinely model-portable ("be load-bearing in any session"), authors replace +the number with a preset name. This is per-way authorial judgment, not a +mechanical sweep. The lint suggestion surfaces candidates. + +**Phase 4 — deprecate raw `half_life:`.** `ways lint` promotes the +`half_life:`-in-frontmatter warning to an error in a later minor release. Raw +`curve:` blocks remain valid for non-Exponential shapes. + +## Consequences + +### Positive + +- Way files stop encoding host-specific token counts when the author wants + portability. The translation to tokens happens at fire time. +- Context-window growth no longer invalidates portable-intent tuning. A future + 2M-token model re-calibrates every preset automatically. +- Mis-tuned ways become observable at the field level — a way tagged `once` + that fires 10 times is a preset mis-match, not a magic-number arithmetic + error. +- The preset table is a single-file tuning surface. If cadence needs global + adjustment, it's one config edit, not 95 files. +- Numerical and preset forms coexist. Model-specific tuning stays available + without forcing every way through the preset table. +- Config-edit-mid-session lets operators tune presets interactively. +- Engine boundary preserved. Sensor-trait unchanged. `attend` unaffected. + +### Negative + +- Small parser complexity: `refire:` accepts two types. Lints must detect + numeric-matches-preset and unknown-preset-name cases. +- Fire-time resolution re-reads or mtime-caches the config file. Overhead is + negligible but non-zero. +- The weaker preset sub-claim (relative cadence stable across models) is + unverified empirically. First re-tuning against a new model will either + validate it or force per-model preset tables. The stronger framework + portability claim is essentially definitional and doesn't need validation. +- Requires `model_to_window()` to cover the operator's model. Unknown models + fall back to a default window; an explicit `CLAUDE_CONTEXT_WINDOW` env + override is needed for unknown models. Silent fallback can produce + surprising behavior — `ways lint` should warn when model is unrecognized. + +### Neutral + +- Raw `half_life:` inside `curve:` blocks remains valid indefinitely as an + escape hatch, primarily for non-Exponential shapes. +- Preset vocabulary is fixed at four values by default. Users add custom + names in their config (`perpetual`, `transient`, project-specific names) + without a schema change. + +## Alternatives Considered + +### Keep raw `half_life`, sweep values periodically + +Operationally cheapest right now (ADR-127 the narrow-tune sweep is exactly +this). Rejected as the primary answer because the tuning debt recurs with +every context-window change. The sweep is a patch, not a fix. + +### Bands only (named presets, no numeric form) + +Earlier draft of this ADR. Rejected after gaming out the migration: the +2026-04-22 hack (`half_life: 200000` on a 1M window → sigma 0.2) doesn't map +cleanly to any band. Forcing every way through a band loses fidelity where +authors have already tuned carefully. The numeric form is the primitive; +presets sit on top as optional portability. + +### Numeric only (no presets) + +Simplest possible shape. Rejected because it forces every future +model-generation change into a 95-file sweep — exactly the problem this ADR +is trying to eliminate. Presets are the portability layer. + +### "Fires per session" as the author-facing unit + +`fires_per_session: 3` is the most intent-aligned expression. Rejected because +it's coupled to the refire floor (currently 0.35) — if the floor changes +later, the meaning of `3` shifts. Fractional fires (`2.5`) are also awkward. +Fraction-of-window has cleaner semantics at the cost of slightly less +intuitive numbers. + +### `ExponentialBanded { sigma }` as a new `Curve` variant + +Original draft proposed this. Rejected because it breaks the unit-agnostic +engine boundary established by ADR-123 — `Curve` would need to know about +context windows. Resolving to `Curve::Exponential` at the ways-cli layer keeps +the engine pristine and is a much smaller diff. + +### Per-model preset tables from day one + +Rejected as speculative. The preset portability sub-claim is unverified; +premature per-model tables buy complexity before the premise is tested. +Structurally available if needed (the config loader can grow a per-model +key), but not the initial design. + +### Mechanical divide-by-1M for every file + +Tempting for its simplicity — one rule, zero judgment. Rejected because it +conflates the value with the intent. Files tuned on 200k (the 81 unhacked +ways) have `half_life: 30000` meaning "re-fire ~3 times per 200k session." +Dividing by 1M gives `refire: 0.03`, which preserves the broken 22-fires +behavior currently observed on Opus rather than the original 3-fires intent. +Using each file's window-at-tuning-time as the reference captures intent +faithfully at a cost of two buckets in Phase 2 (hacked and unhacked). diff --git a/docs/architecture/ways/ADR-127-reject-full-body-embedding-corpus.md b/docs/architecture/ways/ADR-127-reject-full-body-embedding-corpus.md new file mode 100644 index 00000000..2ba97a8d --- /dev/null +++ b/docs/architecture/ways/ADR-127-reject-full-body-embedding-corpus.md @@ -0,0 +1,127 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'experiment on a 90-way corpus with 18 prompts: all variants tie at 11/16 top-1, and full-body costs 6.89x to 9.23x rebuild time' + - precedent: ADR-108 +agent: + name: Claude + model: unrecorded +status: rejected +date: 2026-04-22 +deciders: + - aaronsb + - claude +related: + - ADR-105 + - ADR-107 + - ADR-108 + - ADR-125 +imported: + from: docs/architecture/system/ADR-127-reject-full-body-embedding-corpus.md + format: v0 + status: Rejected +--- + +# ADR-127: Full-body embedding corpus for way matching + +## Context + +Way matching (ADR-108) embeds a keyword-curated string per way — the +`description` field plus the `vocabulary` list — rather than the way's body +prose. When specific parent/child promotions failed (e.g. `environment/deps` +dominating `supplychain/depscan/node`), the instinct was that a denser text +signal would sharpen discrimination. That instinct carried an implicit premise: +**routing quality is limited by the text-matching dimension and can be improved +by making the embedded representation smarter.** + +This ADR tests that premise by productizing full-body embedding in two shapes, +and documents what the test revealed about where routing quality actually lives. + +## Considered + +Two variants against the existing `description + vocabulary` baseline, evaluated +on a 90-way global corpus with 18 prompts (16 actionable across three firing +bands, 2 negative): + +- **full-truncated** — embed the first ~256 tokens of each way's body +- **full-chunked** — embed overlapping chunks of the full body, max-pool to a + single vector + +## Results + +| Variant | Rebuild | Cost × | Clear 8 | Cofire 5 | Bound 3 | Top-1 | +|---|---|---|---|---|---|---| +| baseline | 569 ms | 1.00× | 6/8 | 4/5 | 1/3 | 11/16 | +| full-truncated | 3918 ms | 6.89× | 7/8 | 3/5 | 1/3 | 11/16 | +| full-chunked | 5252 ms | 9.23× | 6/8 | 3/5 | 2/3 | 11/16 | + +All three variants tie at 11/16 top-1 hits. Full-body variants fix some +parent-promotion failures and introduce new ones at the same rate: + +- **P1** `audit npm dependencies for known vulnerabilities` — baseline promotes + parent `environment/deps` over expected `supplychain/depscan/node`. Full-body + variants correctly return the child. +- **P7** `decide between two approaches for the user schema` — baseline promotes + `meta/knowledge` over expected `architecture/design`. Full-body variants + correctly return design. +- **P9** `catch me up on what happened in my inbox overnight` — baseline + correctly returns `ea/briefing` at 0.597; both full-body variants are pulled + to `ea/email` at 0.297–0.341. +- **P10** `find a free 30-minute slot on my calendar` — all three correct, but + baseline confidence 0.671 vs full-body 0.414/0.420. + +Chunked vs truncated is a wash on recall. Chunking costs 34% more than +truncation, confirming the content swap is doing any work, not the +chunking+pooling mechanism. + +The softwaredev-only pilot's apparent win was sample homogeneity: technical +prose is dense enough that full-body's broader signal dominates keyword +curation. Once the corpus includes conversational (`ea/`), operational +(`itops/`), and reflective (`meta/`) trees, full-body wins and losses cancel. + +## Decision + +**Rejected.** Do not productize full-body embedding — neither truncated nor +chunked. + +The premise the experiment was testing — that text-matching density is the axis +of improvement — is not supported by the data. Net-zero aggregate recall across +a 7–9× rebuild-time penalty is the surface reading. The deeper reading is +structural. + +Siblings in the authored graph share vocabulary *by design* — they are authored +into the same subgraph. No text-only representation, keyword-curated or +full-body, can reliably disambiguate them. P9 is the cleanest demonstration: +`ea/briefing` and `ea/email` are graph neighbors with overlapping surface +language, and baseline only picks correctly through keyword luck. Full-body's +denser signal reveals the underlying ambiguity rather than resolving it, which +is why confidence regresses on correct hits (P10: 0.671 → 0.414). + +The text-matching axis is exhausted. + +## Forward path + +The real value in way routing is the authored graph itself (ADR-125): +parent/child structure, progressive disclosure (ADR-105), firing history, and +neighborhood context. Embeddings are one coordinate on each node, not the +routing algorithm. + +**Tactical (bridging):** +- Targeted vocabulary adjustments when a specific sibling collision becomes + painful. +- Per-way `embed_threshold` tuning under ADR-125. + +**Strategic:** graph-aware routing. Subsequent ADRs should explore how +disclosure state, parent context, and recency break ties that text matching +cannot. This ADR's contribution to that direction is negative evidence — the +text-matching axis has been tested and yielded. + +## Provenance + +Full results, per-prompt top-5 for all 18 prompts, and timing distribution were +produced by the harness at `experiments/chunked-embeddings/` on branch +`experiment/full-body-embeddings`. Both were discarded with this ADR; the +evidence above is the preserved record. diff --git a/docs/architecture/ways/ADR-130-sentence-salience-input-reduction-for-embed-matching.md b/docs/architecture/ways/ADR-130-sentence-salience-input-reduction-for-embed-matching.md new file mode 100644 index 00000000..b66f3ee2 --- /dev/null +++ b/docs/architecture/ways/ADR-130-sentence-salience-input-reduction-for-embed-matching.md @@ -0,0 +1,361 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 600+ way-embed SIGABRT crashes between 2026-05-06 and 2026-05-21, peaking at 136/day, from inputs past the 128-token position window + - evidence: 'the lossy truncation mitigations shipped in PRs #94, #95 and #96' + - precedent: ADR-125 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-05-21 +deciders: + - aaronsb + - claude +related: + - ADR-105 + - ADR-108 + - ADR-125 + - ADR-127 +imported: + from: docs/architecture/system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md + format: v0 + status: Accepted +--- + +# ADR-130: Sentence-salience input reduction for embed matching + +## Context + +Way matching today passes hook payloads — user prompts, bash command lines, +subagent dispatch prompts, persisted-output and task-notification blobs +delivered through `UserPromptSubmit` — directly into `way-embed` as the +`--query` argument. The MiniLM embedding models (English `all-MiniLM-L6-v2`, +multilingual `paraphrase-multilingual-MiniLM-L12-v2`) are trained on 128 +tokens of position embeddings; inputs past that abort the embedder inside +`ggml_compute_forward_get_rows`. + +Three things made this latent bug visible recently: + +1. **Auto mode** raised the *rate* of tool dispatches. Each dispatch fires a + hook; each hook spawns `way-embed` twice (EN + multilingual). With three + to four concurrent sessions, the per-minute spawn count crosses what the + single-shot binary architecture was designed for. + +2. **The "goals" feature** raised the *size* of dispatches. Top-down goal + context gets re-packed into structured agent prompts — file paths, + conventions, smoke tests, return-shape requirements — so a `vue-expert` + delegation today is routinely 3–5 KB of well-formed prose. The original + ways scan path was sized for one-sentence intents. + +3. **Custom subagents** (`.claude/agents/*.md` files in project, user, and + plugin locations) became common. These are exactly the dispatches that + carry the longest prompts. + +`way-embed` failed open across all three vectors: the Rust caller +(`run_embed_match` in `tools/ways-cli/src/cmd/scan/scoring.rs`) catches +non-zero exit codes and returns `None`, so the scan appears successful and +emits empty match scores. The underlying SIGABRTs accumulated silently — +600+ crashes between 2026-05-06 and 2026-05-21, peaking at 136/day — and +only surfaced when KDE's `drkonqi` started showing crash-reporter popups. + +### What's been shipped as immediate mitigation + +PRs #94, #95, #96 closed the three crash paths with the cheapest possible +fixes: + +- **#94** — discriminate `subagent_type` against known custom-agent `.md` + files; skip `ways scan task` for custom agents (their `.md` IS their + constitution, so ways injection is redundant) +- **#95** — truncate the bash command field at 256 chars before + `ways scan command` (the longest `commands:` regex in the corpus is 106 + chars, so the cap changes no existing matching behavior) +- **#96** — truncate the combined prompt+response-topics at 1024 chars + before `ways scan prompt` + +These are correct as **safety nets** — they keep the embedder inside its +position-embedding window — but they're lossy. Past the cap, bash heredoc +bodies, persisted-output blobs, and the prose of agent dispatches are +discarded outright. For agent dispatch in particular, that prose IS the +intent signal the matcher exists to read. + +### What the alternatives forbid + +ADR-125 ("Authored Disclosure Graph and Removal of BM25") established +embedding-only retrieval as the architecture. The multilingual model is +treated as a black box; lexical tiers are explicitly out of scope. *"Stops +the temptation to add lexical patches when the embedding behavior surprises +us."* + +ADR-127 ("Full-body embedding corpus for way matching") tested making the +embedded representation denser by ingesting full way bodies instead of +description + vocabulary; **rejected**, because text-matching density is +not the axis of improvement. The deeper reading: routing quality lives in +the graph and the authored aliases, not in throwing more text into the +embedder. + +Both ADRs constrain the design space. We cannot: + +- Reintroduce BM25 or any other lexical-scoring tier as a parallel matcher + (ADR-125) +- Make the embedded representation richer in hopes of better discrimination + (ADR-127) +- Add an LLM-based intent-extraction step (defeats the perf goal — the + whole problem is that we're embedding too often, not that we lack a + smarter analyzer) + +So the design space reduces to: **how do we shrink an arbitrary-sized hook +input into something the embedder can consume, without discarding the +prose that carries the intent signal, without introducing a lexical +matching tier, and at a cost cheaper than the embed step it precedes?** + +## Decision + +Add a **sentence-salience reduction step** in front of the embed call. When +a hook input exceeds the embedder's working window, score its constituent +sentences by frequency-distilled salience, keep the top-N highest-salience +sentences (concatenated in document order), and embed only those. When the +input already fits, pass it through unchanged. + +### Algorithm + +``` +fn reduce_for_embed(input: &str, budget_tokens: usize) -> String { + let token_count = approx_tokens(input); + if token_count <= budget_tokens { + return input.to_string(); + } + + let sentences = split_sentences(input); + let token_counts = count_tokens(&input); // bag of words over whole input + let weights = softmax(token_counts); // distribution with spread + + let scored: Vec<(usize, &str, f64)> = sentences + .iter() + .enumerate() + .map(|(i, s)| { + let tokens = tokenize(s); + let salience = tokens.iter().map(|t| weights[t]).sum::<f64>() + / (tokens.len() as f64).max(1.0); + (i, s, salience) + }) + .collect(); + + let mut selected: Vec<&(usize, &str, f64)> = scored.iter() + .collect(); + selected.sort_by(|a, b| b.2.partial_cmp(&a.2).unwrap()); + + let mut accumulated = 0; + let mut keep: Vec<&(usize, &str, f64)> = Vec::new(); + for s in selected { + let s_tokens = approx_tokens(s.1); + if accumulated + s_tokens > budget_tokens { break; } + keep.push(s); + accumulated += s_tokens; + } + + // Re-order kept sentences in their original document order + keep.sort_by_key(|s| s.0); + keep.iter().map(|s| s.1).collect::<Vec<_>>().join(" ") +} +``` + +The shape: sentences are the unit of selection (preserves phrasing within +each sentence); each sentence's salience is the average softmax-weight of +its tokens against the whole input (so sentences carrying terms the input +emphasizes score higher than sentences full of one-off mentions); +document order is preserved so the embedder sees coherent prose. + +### Why these specific choices + +- **Sentence unit, not token unit.** Bag-of-words over tokens (Shape A in + the design conversation) is cheaper but throws away phrasing — + "rollback the deployment" and "deployment the rollback" become + identical. Sentence unit preserves the local word order that the + embedder is trained to use. If empirical testing shows sentence-level + salience adds no recall over token-level, we shave down to Shape A + later; the path from B → A is mechanical. + +- **Softmax over frequencies, not raw counts.** Counts emphasize the + most-repeated tokens, which in long structured prompts are often + scaffold words ("the", "agent", "context"). Softmax compresses the + high tail and lifts the middle, which is where the discriminative + vocabulary lives. The user's intuition here matches the IR convention: + softmax-normalized weights give better distributive spread than raw TF + or TF-IDF when the corpus comparison is intentionally absent. + +- **No stemming.** The MiniLM tokenizer is WordPiece; it already + segments `agent`/`agents`/`agentic` into shared subword units before + the position embeddings ever look at them. An English-stemmer pass on + top is redundant for prose and harmful for code (where `Agent` the + type is distinct from `agent` the noun). Stopword filtering at the + salience-scoring stage is fine; the stopword list is small and + language-aware via the input itself, not via a per-locale dictionary. + +- **No corpus involvement.** The reducer never reads the ways corpus. + This is deliberate: it preserves ADR-125's separation between input + preparation and corpus-side matching, and it means the reducer's + output is the same regardless of which subset of ways is enabled. + +### Where this lives + +Inside `tools/ways-cli/src/cmd/scan/`, as a function called from +`scan::command`, `scan::task`, `scan::prompt`, and `scan::file` immediately +before `batch_embed_score(query)`. The hook scripts continue to pass the +full payload via the existing CLI flags — the reduction happens once, in +Rust, after both EN and multilingual `--query` arguments have been +identified. + +The `budget_tokens` constant is calibrated to leave headroom under the +MiniLM 128-position-embedding window: + +| Hook | Budget (tokens) | Reasoning | +|---|---|---| +| `scan command` | ~60 | Command shape is short; preserves room for description | +| `scan task` | ~100 | Agent dispatch prompts carry the most intent; max budget | +| `scan prompt` | ~100 | User prompts can be discursive; needs more headroom than commands | +| `scan file` | ~30 | Filepath alone — rarely needs reduction | + +Approximate tokenization (whitespace + punctuation split, ratio ~4 chars/token) +is sufficient for budgeting; precise tokenization is the embedder's job. + +### What stays as-is + +- The interim truncation caps in PRs #94–96 stay in the hook scripts as a + belt-and-suspenders safety net, in case the reducer is bypassed, + errors, or undercounts. Their character limits (1024 for prompt, 256 for + bash) are well above what the reducer should ever emit, so they only + fire on logic faults. +- Custom-agent skip (PR #94's discriminator) stays — those dispatches + shouldn't reach the embed path at all, reducer or not. +- The embedding-only matching architecture from ADR-125 is unchanged. + The corpus side, the alias model, and the per-node thresholds are + untouched. + +## Consequences + +### Positive + +- **Crash class eliminated structurally, not just clamped.** With the + reducer in place, no hook payload can drive `way-embed` past its + position-embedding window regardless of source. The truncation caps + become dead code that never fires. +- **Agent-dispatch prose is preserved, not discarded.** The current + 1024-char prompt cap drops the back half of any 3+ KB delegation; the + reducer keeps the high-salience sentences distributed across the whole + document. That's the prose the matcher exists to read. +- **Cost stays well below the embed step.** Tokenize + frequency-count + + sentence-split + softmax + sort on a 5 KB input runs in single-digit + milliseconds in Rust. The embed step it precedes costs hundreds of + milliseconds (model load + tokenization + cosine against the corpus + rows). Reducer cost is in the noise. +- **Cheaper, not just safer.** Smaller `--query` means smaller embed + input means slightly faster per-call embedding. The reducer pays for + itself even when truncation wouldn't have been triggered. +- **Foundation for shaving down to Shape A later.** Empirical testing + against real prompts will reveal whether sentence-level structure is + doing real work. If not, simplification to token-level (Shape A) is a + ~20-line diff inside the same function. + +### Negative + +- **Bag-of-words within sentences still loses phrasing nuance.** The + reducer keeps whole sentences in document order, but it scores them + using a token-frequency bag. A sentence with an unusual phrasing of a + central concept scores the same as one with a common phrasing of the + same concept. Mitigation: the embedder sees the full sentence text, + so the eventual match still benefits from the unusual phrasing — only + the *selection* step uses the bag. +- **Heavily-templated prompts may collapse onto scaffold content.** Agent + dispatch prompts often include boilerplate like "Return: a summary of + ..." or "Files you'll touch: ...". Repetition makes these tokens + high-weight, which means their sentences may dominate the selection. + This is the "lost in the middle" risk in miniature. Mitigation lives + in the stopword list (which can grow to include scaffold tokens + empirically observed to dominate) and in the budget — keeping enough + sentences that scaffold dominance still leaves room for substance. +- **New code path to maintain.** Roughly 80 lines of Rust + tests. Small + but real ongoing surface. + +### Neutral + +- **Match tuner's role is unchanged.** ADR-125's tuner still measures + alias fidelity in embedding space. The reducer prepares the *input* + before embedding; the tuner audits the *corpus* aliases. Different + axes, no interaction. +- **Multilingual handling is implicit.** The reducer operates on whatever + language the input is in. Sentence-splitting and tokenization both + work reasonably across the languages the corpus supports + (whitespace-segmented; languages without whitespace word boundaries + like Japanese degrade to single-sentence behavior, which is the + current state for `scan prompt` already). +- **The interim truncation caps (#94 / #95 / #96) become dead code.** + Worth keeping for one cycle as redundancy; can be removed in a + follow-up once the reducer is proven in practice. + +## Alternatives Considered + +### Shape A — Token-frequency bag-of-words + +Same algorithm as the chosen direction, but selecting individual tokens +rather than sentences. Top-K tokens by softmax weight, concatenated as +the query. Cheaper (~30 lines), but throws away phrasing entirely. Picked +as the **simplification target** if Shape B's sentence-level scoring +doesn't earn its complexity empirically. + +### Shape C — Multi-chunk parallel embedding + +Split the input into 128-token chunks; embed and match each independently; +union the matched ways. Preserves every byte of prose but multiplies the +per-call embed cost by chunk count, which makes the spawn-storm worse, +not better. Also makes "best match" semantics across chunks ambiguous — +which chunk's score does a way inherit? Rejected on the perf axis. + +### BM25 against the corpus vocabulary + +The original sketch in the design conversation. Rejected by ADR-125 and +by the user's clarifying note in this thread: the corpus's `vocabulary:` +fields are tuned for embedding-space matching, not lexical-scoring +matching. Re-introducing BM25 against them would be the regression +ADR-125 was specifically trying to prevent. + +### LLM-based intent extraction (LLMLingua-style) + +Use a small LLM to compress the input to its intent. Highest quality +in the literature; **rejected** here because it adds a model invocation +ahead of the embed invocation — the exact "embedding repeatedly" failure +mode this work is meant to eliminate. Cost outranks quality for this +problem. + +### Plain truncation (status quo via PRs #94–96) + +Keep the character caps; do nothing more. Works for crash prevention, +fails for signal preservation. Acceptable as immediate mitigation +(already shipped); insufficient as the long-term answer for agent +dispatch prose specifically. The PRs are the safety net under this ADR, +not its replacement. + +## Open Questions + +- **Budget calibration.** The token budgets above are eyeballed from the + 128-token model window. They should be validated empirically — too + large and the embedder degrades inside its window; too small and the + reducer over-prunes. A small benchmark using existing way-matching + ground truth (the same fixtures ADR-127 used) would set defaults. +- **Sentence splitter robustness.** Bash commands, JSON payloads, and + code snippets don't have natural sentence boundaries. The reducer + needs a fall-through: if `split_sentences` returns one (or zero) + sentence on a non-prose input, degrade to Shape A on tokens. This is + a few extra lines but worth deciding before implementation. +- **Stopword list.** Start with a minimal set (`the`, `a`, `is`, `of`, + `and`, …) and grow empirically based on what scaffold tokens dominate + agent prompts in practice. Out of scope to enumerate here; settling + the list belongs in the implementation PR. +- **Observability.** The fail-open silence that hid the original crashes + for two weeks suggests the embed pipeline needs a place where failures + surface. Not in scope for this ADR (it's a separate observability + concern that applies regardless of the reducer), but worth tracking + as a follow-up. diff --git a/docs/architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md b/docs/architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md new file mode 100644 index 00000000..05d0ae75 --- /dev/null +++ b/docs/architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md @@ -0,0 +1,102 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'a 2026-06-09 session readout: about 17 of 47 fires landed in a session whose work never touched their domain' + - evidence: 'near-misses were discarded, so recall was unmeasured; implemented and verified in PRs #117 to #121' + - precedent: ADR-123 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-09 +deciders: + - aaronsb + - claude +related: + - ADR-123 + - ADR-125 + - ADR-130 + - ADR-135 +imported: + from: docs/architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md + format: v0 + status: Accepted +--- + +# ADR-134: Empirical auto-tuning from fire and near-miss telemetry + +## Context + +The firing engine is tuned by hand. Every threshold, half-life, and vocabulary was set by authorial judgment, and the system collects no evidence that could revise them. Three gaps compound: + +1. **ADR-123 Phase E was never built.** The planned `ways tune` cadence calibration (derive per-way `half_life` from observed fire deltas in `~/.claude/stats/events.jsonl`) remains open; `ways tune` today audits locale alias fidelity only (per ADR-125). *(Correction, 2026-06-12: this claim was already false when drafted — Phase E shipped as the `ways tune-curves` subcommand in PR #49. See the Amendment below.)* +2. **Fires are logged; near-misses are discarded.** The matcher computes a score for every way on every prompt, then throws away everything below threshold. False silences — the way that *should* have fired — are structurally invisible, so the precision-first discipline (0 false positives as hard constraint) has no recall measurement to trade against. +3. **Fire relevance is unmeasured.** A 2026-06-09 session readout showed ~17 of 47 fires landing in a session whose work never touched their domain (`itops/incident`, `delivery/migrations`, `testing/mocking` firing into a docs-only session). Each costs an injection; collectively they erode the trust the emission discipline exists to protect. Nothing currently distinguishes a fire that shaped an action from one that was scrolled past. + +Hand-tuning cannot close these gaps because the evidence doesn't exist to hand-tune *from*. A nervous system that cannot adjust its own sensitivities from experience is a reflex arc. + +## Decision + +Extend the telemetry surface and the `ways tune` subcommand into an empirical tuning loop with three measurements and a gated apply step. + +1. **Near-miss logging.** The matcher logs, per prompt, the top-scoring below-threshold candidates (score within a margin of their threshold, e.g. 0.05) to `events.jsonl` as `way_nearmiss` events. The scores are already computed; this is persistence, not new computation. Volume is bounded by the margin. +2. **Cadence calibration (ADR-123 Phase E, as planned).** `ways tune --cadence` groups `way_fired` / `way_redisclosed` events by way, computes token-delta distributions between fires, and suggests `half_life` per the existing Phase E worksheet (rule of thumb: half_life ≈ median delta). *(Already shipped — as the `ways tune-curves` command; no `tune --cadence` mode exists or is needed under that name. See the Amendment.)* +3. **Relevance signal.** `ways tune --precision` correlates each way's fires with the session's subsequent activity class, derived from data already in the event log (trigger channel, tool mix, domains of other fires). A way whose fires repeatedly land in sessions that never touch its domain is flagged with its observed irrelevance rate and suggested remedies: threshold raise, vocabulary narrowing, or trigger-channel change. This is a heuristic flag, not a verdict — the same contract as `ways tune`'s fidelity audit. +4. **Gated apply.** All three report by default; `--apply` rewrites frontmatter (`half_life`, `embed_threshold`) in place. Vocabulary changes are never auto-applied — they re-shape the embedding neighborhood and stay authorial. Applied changes are ordinary git diffs, reviewable and revertible. *(The `half_life` apply already ships in `tune-curves --apply`, which rewrites the `curve:` block; only `embed_threshold` apply remains to build. See the Amendment.)* + +The recall counterpart falls out of (1): near-miss data plus the existing fixture workflow lets `ways tune` report likely false silences (ways that consistently score just under threshold on prompts whose sessions then did that way's kind of work), giving the 0-FP discipline its first recall estimate. + +## Consequences + +### Positive + +- Closes the open loop: telemetry that exists only as a record becomes input to calibration. Half-lives can be retuned per model generation (the maintenance posture ADR-301 documented). +- False silences become measurable for the first time; precision-first stops being precision-only. +- Per-session precision audits (`ways list` plus irrelevance rates) turn anecdotes like the 2026-06-09 readout into a tracked metric. + +### Negative + +- `events.jsonl` grows faster; near-miss margin needs a cap and the log needs rotation. +- The relevance signal is a proxy — activity-class correlation can mislabel legitimately cross-cutting ways (e.g. `meta/tracking`). Flags must stay diagnostic, never auto-applied to vocabulary. +- `--apply` writing frontmatter from observed behavior risks codifying one user's work distribution; suggested values should show the sample size they derive from. + +### Neutral + +- ADR-123 remains Accepted and unedited; this ADR absorbs and extends its open Phase E. Phase F (A/B validation) is unaffected. +- Fine-tuning the embedding model (deferred in ADR-108) becomes more attractive once near-miss data accumulates as training signal — out of scope here. + +## Alternatives Considered + +- **Edit ADR-123 to widen Phase E** — rejected: ADR-123 is an Accepted historical record; widening its scope post-acceptance hides when the precision dimension entered the design. +- **Model-graded relevance (LLM judges whether each fire influenced the session)** — rejected for now: highest-fidelity signal but adds inference cost to a system whose value is being cheap and ambient. Revisit if the activity-class proxy proves too coarse. Note that an *external* instance of this signal already arrives for free: the periodic Claude Code usage report is model-graded analysis of where a session actually went wrong, generated outside this loop at no inference cost to it. It does not measure way-fire relevance directly, but it measures the downstream thing ways exist to prevent — and a 2026-06 report (write-time friction: Buggy Code 40, Wrong Approach 37, Excessive Changes 7) is what empirically justified the content-level trigger in ADR-135, this design's first pattern-level consumer. +- **Manual periodic audits of `ways list` output** — rejected as the only mechanism: it found today's signal, but it doesn't scale and never sees near-misses. + +## Amendment — 2026-06-12: Phase E reconciliation + +Implementation grounding for Decision 1 surfaced that this ADR's own Context was wrong on a load-bearing point. Context gap #1 and Decision 2 both assert that ADR-123 Phase E — cadence-derived `half_life` calibration — was never built. **It was.** It ships as the `ways tune-curves` subcommand (commit `28527e8`, PR #49; with its input field `token_position` added to `way_fired` in commit `8b20782`), wired at `main.rs`, predating this ADR's 2026-06-09 draft. Run today it processes the real event log (hundreds of fires per high-traffic way), groups `way_fired`/`way_redisclosed` by `(way, session)`, computes token-position deltas, and suggests `half_life ≈ median delta` — verbatim Decision 2. There is no separate cadence work to do. + +That a *draft about empirical self-correction* shipped a confident-but-wrong claim about its own installed components is the failure mode the 2026-06 usage report named; recording the correction here rather than silently editing it over is the point. + +The reconciliation, decision by decision: + +- **Decision 1 (near-miss logging)** — genuinely new; implemented (`way_nearmiss`, the matcher's 3-state `match_prompt`, `near_miss_margin` config). Stands. +- **Decision 2 (cadence calibration)** — **already satisfied by `tune-curves`.** No new code. The `--cadence` *spelling* in the Decision is aspirational; the *capability* exists under a sibling command name. Building a `tune --cadence` alias was considered and rejected as cosmetic duplication (it would add CLI surface for naming fidelity alone) — the over-build the firing engine's own discipline (ADR-135) exists to prevent. +- **Decision 3 (relevance / `--precision`)** — genuinely new and **not built**. This is the substantive remaining contribution: correlating each way's fires with session activity class to flag the irrelevance the 2026-06-09 readout measured (17/47 off-domain fires). +- **Decision 4 (gated apply)** — **partly already shipped.** `tune-curves --apply` performs the `half_life` rewrite (on the `curve:` block) with sample-size reporting. Only `embed_threshold` apply — driven by the recall/precision signals — remains to build. + +Net: this ADR's open work is Decision 3 (precision) plus the `embed_threshold` slice of Decision 4. Decisions 1 and 2 and the `half_life` half of 4 are done. The ADR is retained whole — including the now-corrected Context — because the empirical-tuning frame and the precision signal it still authorizes remain valid; the value is in narrowing the build to what does not already exist. + +## Implementation status — 2026-06-12 (Accepted) + +The design is adopted; the empirical-tuning loop it authorizes is built and exercised against the real event log. Status of each piece: + +- **Decision 1 — near-miss logging** — shipped (PR #117). `way_nearmiss` events emit for below-threshold candidates within `near_miss_margin` (default 0.05); the matcher's `match_prompt` is a 3-state `Fired | NearMiss | NoMatch`. Verified live. +- **Decision 2 — cadence calibration** — pre-existing as `ways tune-curves` (reconciled above, PR #118). No code. +- **Decision 3 — relevance signal** — shipped (PR #119) as `ways tune-precision`: parent-family activity classes, off-class irrelevance rate, and a breadth-based cross-cutting guard that keeps the named false positive (`meta/tracking`) out of the vocab-narrowing remedy. Verified live (69 ok / 64 cross-cutting / 25 mis-targeted / 44 low-n of 202 ways). +- **Decision 4 — gated apply** — split. `half_life` apply ships in `tune-curves --apply`. `embed_threshold` apply needs a derivation that the data did not support: `way_fired` carried no firing score. The **fire-score telemetry** that makes a principled raise possible shipped (PR #120) — `fire_score` is now logged on first-fires. The **suggestion + apply** itself is **deferred until that telemetry accumulates** (it cannot be validated over an empty population), sequenced exactly as `token_position`→`tune-curves` was. Tracked as the only remaining build. +- **Negative (log growth)** — addressed (PR #121): `events.jsonl` is bounded by tail-compaction in `log_event`. + +This ADR is **Accepted** with the `embed_threshold`-apply slice of Decision 4 as tracked, data-gated future work — a sequencing constraint, not an open design question. diff --git a/docs/architecture/ways/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md b/docs/architecture/ways/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md new file mode 100644 index 00000000..e12be518 --- /dev/null +++ b/docs/architecture/ways/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md @@ -0,0 +1,123 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - matching + - authoring +basis: + - evidence: 'an external model-graded usage report (2026-05 to 2026-06, 92 sessions): Buggy Code 40, Wrong Approach 37, Excessive Changes 7, all at write time' + - precedent: ADR-134 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-12 +deciders: + - aaronsb + - claude +related: + - ADR-134 + - ADR-123 + - ADR-130 +imported: + from: docs/architecture/system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md + format: v0 + status: Accepted +--- + +# ADR-135: Content-aware write-time over-build gate with a self-extending pattern corpus + +## Context + +A write-time check now exists: `softwaredev/code/code.check.md` fires at `PreToolUse(Edit|Write)` on source files and surfaces a short anchor plus verification questions — does this need to exist, is this the codepath that actually runs, is this the minimal change. It ships with no engine change because it needs none: the `ways scan file` path matches a check to a file by its `files:` regex and never inspects what is about to be written. That is its ceiling. A question can ask "did you reach for the stdlib first?"; it cannot *see* the hand-rolled LRU cache, the reinvented email-regex validator, or the new dependency in the diff and name them. + +Naming them requires the candidate content, which today never reaches the matcher. `check-file-pre.sh` forwards `tool_input.file_path` only, and `ways scan file` accepts `--path`, `--session`, `--project` — no content channel. Adding one is a genuine new capability, not a new file. + +Three forces make it worth building, and one makes it dangerous: + +1. **The friction is empirically at write-time.** An external, model-graded usage report (2026-05 to 2026-06, 92 sessions) shows the dominant friction buckets are Buggy Code (40), Wrong Approach (37), and Excessive Changes (7) — every one of them landing at the moment code is written, the exact moment this gate would fire. +2. **The minimalism content is well-understood.** The "ponytail" plugin's ladder (YAGNI → stdlib → native → installed dep → one line) and its review tags (`stdlib:`, `native:`, `yagni:`, `shrink:`) are a serviceable taxonomy for what over-build *looks like* in concrete code. +3. **The matcher already learns elsewhere.** ADR-134 established empirical tuning from fire and near-miss telemetry — but at the threshold and cadence level. Pattern *recognition* is the next level up and has no mechanism. + +The danger is the obvious implementation. The ponytail model bakes one experienced developer's opinion into a fixed catalog. A fixed catalog nags at its edges: code that falls outside it is misjudged, and legitimately novel approaches read as "you're doing it wrong." A matcher that suppresses innovation because it has not seen it before is a failure we will not ship. The catalog cannot be the deliverable. + +## Decision + +Build the content channel, but make the thing it feeds a **learning matcher, not a static catalog**. The pattern corpus self-extends from comprehended encounters, bounded by the precision discipline already in force, so it grows understanding without growing dogma. + +### 1. Content channel + +`check-file-pre.sh` forwards `tool_input.content` (Write) / `tool_input.new_string` (Edit). `ways scan file` gains a `--content` input — or a sibling `ways scan write` subcommand — that pattern-scans the candidate code. Content is budget-reduced before scanning, consistent with ADR-130's uniform hook budget. + +### 2. The gate, and what runs on the hot path + +The gate emits ponytail-style `stdlib:` / `native:` findings plus a new-file line-count signal, and is silent otherwise. The hot path is cheap and does no inference: + +- **Recognized pattern** → emit the finding (`stdlib: hand-rolled LRU — functools.lru_cache covers it`). +- **Unrecognized code** → **silence**, plus an *encounter* entry to the `events.jsonl` near-miss stream (ADR-134). Silence is the correct output, not a fallback. + +There is no inline comprehension on the hot path. "Spend cycles to understand it" never means deep-analyze every keystroke — that would destroy the ambient, near-zero-cost value the firing engine exists to provide. + +### 3. The loop — corpus growth happens out of band, and runs both ways + +Comprehension and recording happen in a deliberate authoring/tuning pass modeled on `ways tune`, consuming the encounter telemetry the hot path emitted: + +1. Read the encounter stream — the unrecognized writes the gate stayed silent on. +2. Comprehend the genuinely ambiguous cases (the cost is paid here, off the hot path, on a bounded sample). +3. **Record a finding — or, usually, do not.** A finding is either a new **match** (a real, generalizable reinvention) or an **exemption** (a legitimate or novel shape that must never be flagged). Most encounters are neither and stay silent and unrecorded. +4. **Prune** patterns whose telemetry shows them over-firing. + +The loop is **bidirectional by design**. A loop that only adds matches becomes ponytail-by-accretion — a catalog that has "seen everything" and therefore nags at everything. That is the named anti-pattern. Recognition is provisional and revisable; a pattern earns its place from observed behavior and loses it the same way. + +### 4. Precision discipline (inherited, not invented) + +The gate inherits ADR-134's precision-first, zero-false-positive **hard constraint**. The corpus gets less ignorant over time *only through comprehension*, never through speculative up-front enumeration. Comprehensiveness is explicitly a non-goal: a small, high-confidence corpus that is silent on the unfamiliar beats a large one that is confidently wrong. The bootstrap seed is deliberately tiny — on the order of three or four patterns (hand-rolled LRU/TTL cache, email-regex validation, manual retry around an idempotent call, new-file line count) — and the seed is a starting point for the loop, not a specification of coverage. + +This makes ADR-135 the first *content-level* and *pattern-level* application of ADR-134's machinery: 134 tunes thresholds and cadence from telemetry; 135 tunes recognition itself from the same substrate. + +## Amendment — 2026-06-12: mechanism (PostToolUse postcheck, pattern-as-way) + +The Decision above specified the content channel as *PreToolUse* plus a new `ways scan --content` channel in the binary. Implementation grounding revised the mechanism on two points. The intent above is unchanged — detect concrete over-build in candidate code, as a learning corpus governed by precision-first silence. Only the delivery changes, and it changes toward less code and existing infrastructure. + +**1. PostToolUse via per-way `postcheck.sh`, not PreToolUse via the binary.** The repo already has the mechanism this needs: ADR-123 Decision 5's reactive-firing path (`check-post.sh` walks `**/postcheck.sh`, pipes the PostToolUse input to each, and treats exit 0 as a fire request gated through `ways show way`). The gate is advisory — it emits, never denies — so PreToolUse's interception power buys nothing; PostToolUse is both the better behavioral fit (let the write land, then surface the fix for the next turn, as the usage report itself suggested) and free of any binary or core-hook change. The PreToolUse/binary channel was the right *first* design on paper; the existing postcheck path is the lazier one that works. + +**2. Patterns live as pattern-ways, not as a bespoke data structure.** A postcheck's stdout is discarded — only its exit code is read, and the content injected on a fire is the *way's own body*. So a postcheck cannot emit a dynamic finding line; the body carries the named replacement(s). This collapses the "self-extending pattern corpus" into the corpus itself: + +- **Recognition** = a postcheck predicate (stimulus → match?). +- **Learned response** = the way's body (the named replacement). +- **Learning a pattern** = authoring a pattern-way, through the existing `knowledge/authoring` loop — so the #7 loop's "record a finding" is literally `ways`-authoring from encounter telemetry. +- **Pruning** = deleting the way; **exemption** = the absence of one. + +The pattern library is therefore not new infrastructure; it is a small, prunable domain of pattern-ways. This deepens, rather than alters, the learning-corpus thesis: the recognition set is the corpus, and it grows and shrinks by gaining and losing ways. Encounter telemetry and the loop (#7) remain gated on ADR-134. + +**Granularity is a YAGNI call, decided by the hot path.** `check-post.sh` runs *every* `postcheck.sh` on *every* `Edit`/`Write`/`Bash`/`Task` — so one-way-per-pattern means one subprocess spawn per pattern per tool event. v1 is therefore a **single `softwaredev/code/overbuild` way** whose one postcheck holds the seed detectors and whose body lists their replacements; Claude self-selects the relevant one. Splitting into per-pattern ways (`overbuild/lru`, `overbuild/email`, …) is the loop's later refinement, justified only when per-pattern pruning or telemetry earns the extra spawns. The conceptual mapping above holds at either granularity — one way with N detectors, or N ways with one each. + +Implementation note: the over-build way fires *only* via its postcheck, not via prompt/file embedding — its frontmatter is kept trigger-inert (no `description`/`vocabulary`/`pattern`) so it never pollutes semantic matching, and needs no locale companion. + +## Consequences + +### Positive + +- The write-time over-build signal becomes concrete: it names the specific reinvention in the candidate code, not just a generic "did you consider…". This targets the usage report's measured #1 friction at the point it occurs. +- The corpus is anti-fragile to novelty. The encounter ponytail would misjudge is exactly the encounter that, once comprehended, teaches the system — or is recorded as a permanent exemption. Unfamiliarity routes to silence-then-maybe-learn, never to a false flag. +- ADR-134's near-miss substrate gets its first concrete consumer beyond threshold tuning, validating that design end to end. + +### Negative + +- A new input channel widens the hot-path surface: content must be forwarded, budget-reduced, and scanned within the PreToolUse latency envelope. Pattern scanning must stay cheap (literal/regex shape matching, no inference). +- The loop requires human/agent attention on a cadence. Without the out-of-band pass, the corpus ossifies at its seed — functional, but not the learning system this ADR justifies. +- `events.jsonl` gains an encounter stream on top of 134's near-miss volume; the encounter margin needs a cap and the log needs rotation (134's concern, compounded). + +### Neutral + +- Tier 1 (`code.check.md`) is unaffected and remains valuable on its own; this ADR is strictly additive. If 135 is never implemented, Tier 1 still ships the questions. +- The ponytail review *tags* survive (`stdlib:`/`native:`); the ponytail *philosophy* (fixed, comprehensive, opinionated) is explicitly rejected. Only the finding shape is borrowed. +- Pattern entries become ordinary versioned artifacts — reviewable, revertible git diffs — like the frontmatter 134's `--apply` rewrites. + +## Alternatives Considered + +- **Static curated catalog (the ponytail model).** Rejected as the deliverable: it nags at its edges and cannot grow from experience. Its taxonomy is borrowed; its fixedness is the thing this ADR exists to avoid. +- **Model-graded per-write judgment (an LLM decides "is this over-built?" on every Write).** Highest fidelity, rejected for the hot path for the same reason ADR-134 rejected model-graded relevance: it adds inference cost to a system whose value is being cheap and ambient. It is admissible only in the out-of-band comprehension step, on a bounded sample. +- **PostToolUse lint instead of PreToolUse gate** (the shape the usage report literally suggested: `PostToolUse(Edit|Write)` running a formatter/type-checker). Complementary, not a substitute — it catches mechanical defects *after* the code lands, whereas over-build is a *before-you-write* decision. Worth adopting separately; it does not address this signal. +- **Append-only learning loop.** Rejected: monotonic growth reconstitutes the ponytail dogma by accretion. Pruning and exemptions are load-bearing, not optional. diff --git a/docs/architecture/ways/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md b/docs/architecture/ways/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md new file mode 100644 index 00000000..8bd81de6 --- /dev/null +++ b/docs/architecture/ways/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md @@ -0,0 +1,203 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: disclosure +basis: + - evidence: 'usage data: 222 sessions / 10,019 way fires over ~4.5 months, one English speaker' + - evidence: 'scout inventory 2026-06-22: 91 .locales.jsonl files, ~3,094 entries, a second 127MB model, a 17x-per-way authoring tax, 3 CI targets (issue #160)' + - precedent: ADR-138 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-06-22 +deciders: + - aaronsb + - claude +related: + - ADR-107 + - ADR-125 + - ADR-138 +imported: + from: docs/architecture/system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md + format: v0 + status: Accepted + unmapped: + supersedes_in_part: ADR-125 +--- + +# ADR-139: Shelve maintainer i18n; adopter-run localization via ways-localize + +> Decision record for GitHub issue #160. Supersedes the maintainer-side +> obligations of the locale layer introduced in ADR-107 and elaborated (still in +> Draft) in ADR-125. The retrieval machinery those ADRs describe is **kept** — only +> the maintainer's standing obligation to author and ship translations is shelved. + +## Context + +agent-ways ships a maintainer-maintained multilingual layer so non-English Claude +Code users can get ways' benefits in their own language. The intent was sound; the +economics and architecture are not: + +- **Zero validated users.** 222 sessions / 10,019 way fires over ~4.5 months — one + English speaker. No non-English usage signal exists. We built and maintain the + layer against a hypothetical adopter. +- **Heavy, compounding maintenance cost.** A scout inventory (2026-06-22) measured + the real surface: + - **91 `.locales.jsonl` files, ~3,094 locale entries** across 18 active languages + (the issue's "~1,547" estimate undercounted). + - A **second 127MB embedding model** (`paraphrase-multilingual-MiniLM-L12-v2`) + downloaded *unconditionally* in every `make setup` alongside the 21MB + English-only model. + - A whole audit subsystem (`ways language`, `ways tune`) plus three CI targets + (`test-lang`, `test-locales`, `test-multilingual`) gating `make test`. + - A **17×-per-way authoring tax**: `template.rs` pre-populates every new way with + 17 translation stubs (18 entries incl. English). +- **Architecturally leaky, and getting leakier.** Per ADR-138 (skills own the *how*, + ways own the *5W*), value is actively migrating *into* skills — and skills are + English-only. A localized way now hands off to an English skill: we localized the + *shrinking* half of the value surface. +- **Unverified substrate.** Claude Code's own non-English skill surfacing/matching + has never been measured. We localized ways on top of an unmeasured foundation. + +A genuinely complete non-English experience needs localized **skills** *and* +verification of Claude Code's native non-English behavior — a much larger job that +should be triggered by a real adopter, not pre-built against a hypothetical one. And +the people best positioned to judge translation fidelity and idiom are **native +speakers**, not the (English-speaking) maintainer. So both the *cost* and the +*quality* of localization belong with the adopter, not the maintainer. + +## Decision + +**Stop pre-shipping translations. Move localization to an adopter-run, autonomous +flow.** Concretely, draw a hard line between the localization *engine* (kept, +dormant) and the localization *data* + maintainer *obligations* (removed): + +| Layer | Fate | Why | +|-------|------|-----| +| Localization **data** (~3,094 entries / 91 `.locales.jsonl`) | **Delete** | Stale, unused, regenerable on demand; git history is the backstop | +| Rust intl/locale code paths (`corpus.rs` split, `tune.rs`, `language.rs`, `match_cmd.rs` multi-column) | **Keep, dormant** | The engine the skill drives; cheap to carry | +| Multilingual embedding model management/download | **Keep, dormant** | Activated on demand, never for English installs | +| Per-way localization in `template.rs` (the 17× tax) | **Remove** | New/edited ways are English-only by default | +| `ways tune` / `ways language` audit commands | **Keep, don't run by default** | Become the acceptance gate for adopter-run localization | +| Multilingual model in default `make setup` | **Make on-demand** | English install never fetches the 127MB model | + +**English is fully dormant.** A default (English) install fetches one model, builds +one corpus, runs no locale audit, and pays no authoring tax. The multilingual engine +exists but is never touched until an adopter explicitly localizes. + +**Three new components replace the maintainer obligation:** + +1. **`ways-localize` skill** (`ways-*` family — kin to `ways-tests`, `ways-update`). + The operator-facing orchestrator. It: + - **Interviews** the operator (human) for the target language — conversational, + and recognizable from a request *in that language* (resolving the bootstrap + paradox: a non-English user shouldn't need to read English to get a non-English + experience). + - **Delegates the fan-out** to a re-hydrate-and-tune workflow (below). + - **Sets Claude Code's own response language** by writing the `language` key in + `settings.json` (the one supported mechanism — `{"language": "french"}`, + persistent, effective next session). No env var or CLI flag exists. + - Surfaces progress and the final "N ways localized, `ways tune` clean" summary + **in the adopter's language**. + +2. **Re-hydrate-and-tune workflow** (per the `meta/workflows` way). Given a target + language, it fans out across ways: translate each way's `description` + + `vocabulary` (Claude is multilingual), pack the stubs (`pack-locales.sh`), rebuild + the corpus, and **verify fidelity/discrimination with `ways tune`**, iterating on + flagged stubs until clean. `ways tune` is the objective, language-agnostic + acceptance gate (it audits embedding geometry, not prose). + +3. **State-triggered detection macro** (a `state`-triggered way). At + SessionStart/PreCompact it reads the `language` field in `settings.json`. The + check itself always runs — it must read the field to know the state — but only + one state produces output; the cost on English installs is a single cheap file + read, no model/embedding/output: + - English or unset → emits nothing. Silenced on *output*, not on *existence*: + the detector runs, finds nothing to do, and injects zero guidance (the ~99% + case). + - Non-English **and** no regenerated locale data exists for that language → + injects a recommendation to invoke `ways-localize`. Because CC is already + responding in that language, the nudge reaches the operator in their language + for free. + - **Self-silences**: once `ways-localize` rehydrates that language, the locale + data exists → the condition is false → the macro goes quiet. (Cleaner than + fire-count decay; the presence of localized data *is* the "done" signal.) + +**Scope note.** This localizes *ways* only. A complete non-English experience also +needs localized **skills** and verification of Claude Code's native non-English skill +matching — tracked as a larger follow-on, not part of this decision. + +## Validation + +The central decision (shelve maintainer i18n) rests on *measured* facts: the usage +data (222 sessions / one English speaker) and the 2026-06-22 scout inventory (91 +files, ~3,094 entries, 127MB second model, 17× tax, 3 CI targets). Those are read +directly from the repo and corpus — no external bet. + +One premise is about Claude Code's *own* behavior and was probed before acceptance, +per the prototype-before-accept way: + +- **Readable half (verified).** The detection macro reads the top-level `language` + string in `settings.json` via `jq`. Probed 2026-06-22: the field is absent on this + install → the macro correctly resolves to the silent (English/unset) state. The + detector keys off a real, machine-readable field, not a hallucinated one. +- **Behavioral half (assumed, deferred).** That writing `{"language": "<lang>"}` + actually switches CC's response language (docs: v2.1.176+, effective next session) + is sourced from official docs, not observed — testing it requires a settings change + + restart, intrusive to the authoring session. It is **secondary to the shelve** + (it powers the skill's step 5 and the macro's trigger, not the decision to shelve), + and has documented fallbacks (output-style, CLAUDE.md instruction) if it disproves. + To be confirmed when `ways-localize` is built. + +## Consequences + +### Positive + +- **Lighter default install.** English users skip a 127MB model download, a second + corpus + embedding pass, the locale audit, and three CI targets. +- **Zero authoring tax.** New/edited ways are English-only; no 17 stub lines per way. +- **Cost and quality land on the beneficiary.** The native-speaker adopter pays the + translation/tuning tokens *and* is the real expert on fidelity and idiom — better + output than maintainer-authored stubs, at no maintainer cost. +- **Fully reversible engine.** Deleting data, not code, means re-enabling + localization is `ways-localize`, not a rebuild. +- **The leak stops mattering.** We no longer maintain translations for the shrinking + (ways) half while the growing (skills) half stays English. + +### Negative + +- **Non-English ways stop shipping out of the box.** Until an adopter runs + `ways-localize`, only English ways match. (Mitigated: there are no measured + non-English users today, and the detection macro surfaces the fix immediately.) +- **First-run localization cost moves to the adopter** — model download + a + translate/tune token spend across all ways. (By design: the beneficiary pays.) +- **New moving parts to build and maintain** — a skill, a workflow, and a macro — + replacing static data with orchestration. + +### Neutral + +- The Rust engine carries dormant code (corpus split, `tune`, `language`, + multi-column match). It compiles and is tested but unexercised on English installs. +- The deleted `.locales.jsonl` data remains in git history; "delete" is recoverable. +- ADR-125 (Draft) keeps its *retrieval* model (coordinate-alias embedding, per-node + thresholds); only its maintainer-side authoring obligation is superseded here. + +## Alternatives Considered + +- **Leave the stubs dormant in git, delete nothing (issue #160's original framing).** + Rejected: the ~3,094 stale entries stay a maintenance and review-noise liability + (they confused our own tuning pass), and the authoring tax persists if `template.rs` + still emits them. Deleting the data while keeping the engine gets the same + reversibility (git history) with none of the ongoing drag. +- **Delete the engine too (Rust intl code + model management).** Rejected: the engine + is the cheap part and deleting it makes `ways-localize` a from-scratch rebuild + rather than a re-hydration. Dormant code is a smaller liability than lost capability. +- **Keep maintainer-maintained translations.** Rejected on all four counts above: + zero validated users, compounding cost, architectural leak, unverified substrate — + and lower quality than a native speaker + Claude would produce. +- **Ship an English quick-start guide for localization.** Rejected as a bootstrap + paradox: an English-only guide gates exactly the non-English users it's meant to + serve. The conversational, in-language entry point (the skill + detection macro) + replaces it. diff --git a/docs/architecture/ways/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md b/docs/architecture/ways/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md new file mode 100644 index 00000000..b88b752d --- /dev/null +++ b/docs/architecture/ways/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md @@ -0,0 +1,205 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'research pass 2026-07-02: session_start is written by shell hooks to the legacy ~/.claude/stats/events.jsonl while readers prefer the $XDG_STATE log, so new sessions are invisible' + - evidence: way_fired carries no transcript uuid or turn index, and its trigger field records the channel, not the matched term +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-142 + - ADR-134 + - ADR-201 +imported: + from: docs/architecture/system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md + format: v0 + status: Accepted +--- + +# ADR-153: Session-introspection substrate — correlating fired ways to turns + +## Context + +We want to answer, for any past or live session: **which ways were injected into +context, on which turn, and *why* — what caused the hook to fire.** Three +front-ends want this (the `ways introspect <replay|live|dump>` surface of +ADR-154 — post-hoc replay, live monitor, and a non-interactive dump for +autonomous agents), so the correlation belongs in one shared substrate below +them, not re-derived per front-end. + +A research pass (2026-07-02) mapped exactly what data exists. The findings define +what the substrate can join today and what it cannot: + +**The firing-event log** (`$XDG_STATE/agent-ways/events.jsonl`, append-only JSONL; +written by `session::log_event`, read by `firing::load_events`). A `way_fired` +record carries `way`, `domain`, `trigger`, `scope`, `project`, `session`, +`token_position`, and — for semantic fires only — `fire_score` (plus +conditionally-emitted subagent/relationship fields `parent`, `tree_depth`, +`epoch_distance`, `team`). Sibling event +types: `way_nearmiss` (with `score_en/multi`, `thr_en/multi`, `margin`, +`query_tokens` *count*), `session_start` (`{ts, project, session}` only), +`check_fired`, `way_redisclosed`. + +**The `trigger` field records the match *channel*, not the matched term** — +`keyword` / `semantic:embedding:en|multi` / `bash` / `file` / `state`. It tells +you the mechanism, never which vocabulary word or regex substring hit. + +**Session transcripts** live at `~/.claude/projects/<slug>/<session>.jsonl` +(Claude-Code-owned, read-only; ms-precision timestamps, `uuid`/`parentUuid` +threading). Injected way guidance appears as an `attachment` line +(`hook_additional_context`) whose `content` is the **concatenated way bodies for +one prompt** — way-*anonymous*, one blob per prompt, and sometimes truncated to a +spilled `tool-results/…-additionalContext.txt` file. + +**Two correctness problems block the substrate before it starts:** + +1. **The events log is split-brained.** `session_start` — the only event that + defines a session for `rethink` — is written by shell hooks + (`clear-markers.sh`, `inject-subagent.sh`) that **hardcode the legacy + `~/.claude/stats/events.jsonl`**, while every reader resolves through + `paths::events_log()`, which *prefers* the migrated `$XDG_STATE` file + post-ADR-142. New sessions' `session_start` lines land in the orphaned file and + are invisible. Any introspection over an incomplete log is wrong. + +2. **The join to a specific turn is heuristic, not keyed.** `way_fired` carries no + transcript message `uuid` and no turn index; the finest available link is + `(session, token_position, ts≈)` → a prompt bucket → the one attachment → its + `user` turn via `parentUuid`. And "what text matched" is persisted for + *nothing* — keyword matches are re-runnable against the prompt (if the + transcript is present), semantic matches are only a way-level cosine. + +## Decision + +### 1. Single-writer the events log (correctness prerequisite) + +Add a `ways events-log-path` subcommand (precedent: `ways response-topics-path`, +which shell hooks already consult instead of hardcoding). Rewire +`clear-markers.sh` and `inject-subagent.sh` to resolve the path from the binary, +so every `session_start` / `way_fired` writer and every reader agree on one file. +A migration-time union read can bridge existing orphaned logs, but the durable fix +is one writer path. + +### 2. A typed introspection model in `ways-core` + +Factor a `SessionIntrospection` model that joins the three sources into pure, +serde-serializable data — no ANSI, no terminal. It generalizes the already-proven +`reconstruct_frames` → `render`/`serialize` triplet (ADR-154). Shape: + +``` +Session { id, project, window_k, summary } + └─ Turn { epoch, token_position, ts, transcript_uuid? } + └─ FiredWay { way_id, trigger_channel, fire_score?, way_path, + criteria: MatchCriteria, // from frontmatter + match: MatchDetail? } // what hit (see §3) +``` + +`MatchCriteria` surfaces the fire-bearing frontmatter (`pattern`, `vocabulary`, +`commands`, `files`, `trigger`, `embed_threshold`). The join is honest about its +grain: keyed where a key exists, heuristic (time/token bucket) where it does not, +and every heuristic edge is labelled as such in the model so a consumer never +mistakes a proximity guess for a foreign key. + +### 3. Make "why" precise — `transcript_uuid` post-hoc, `matched_span` at fire time + +> **Correction (implementation finding, 2026-07-03).** The original §3 assumed the +> fire path could record a `transcript_uuid` "because the hook receives the message +> id on stdin." It does not: the real `UserPromptSubmit` hook payload carries only +> `prompt`, `session_id`, `cwd`, `agent_id` — no message uuid (the prompt's uuid is +> assigned by Claude Code, after the hook). Investigation of a live transcript found +> a *better* source, so the two halves are sourced differently. + +- **`transcript_uuid` — resolved *post-hoc* from the transcript, not at fire + time.** Every `UserPromptSubmit` hook injection is recorded in the session + transcript as an `attachment` (its `.attachment.content` holds the injected way + bodies), and its `parentUuid` chain walks back to the triggering `user` message + (verified empirically). The introspection model reads the transcript, matches each + turn — its fire-timestamp cluster — to the corresponding `UserPromptSubmit` + attachment, and follows `parentUuid` to the user-message uuid: a genuine foreign + key. This is strictly better than fire-time capture — no hot-path change, and it + works for historical sessions whose transcript survives. Only the turn→attachment + step is heuristic (session + sub-second timestamp — the fire happens *inside* the + hook); the attachment→message link is keyed. A turn whose transcript is + absent/unmatched stays `Heuristic`. + +- **`matched_span` — recorded at fire time** (`cmd/show/*`, `cmd/scan/*`). The + transcript's injected content is way-*anonymous* (concatenated bodies for the + whole prompt), so *what text matched an individual way* cannot be recovered + post-hoc. For the keyword/command/file channels the fire path captures the + regex/glob match text — additive, forward-only (old records lack it), kept cheap + and line-atomic on the hot path. + +- Semantic stays **way-level**: `fire_score` ≥ `embed_threshold` is the honest + grain; per-vocabulary-term attribution is impossible (one embedding per way) and + must not be faked. + +Both enrichments are **additive**: the model uses each field when present and falls +back to the heuristic time/token-bucket grain when absent. The post-hoc transcript +join degrades to `Heuristic` when the transcript is unavailable; `matched_span` +claims no backfill. + +### 4. Share the substrate with the compliance finding pipeline (ADR-201) + +ADR-201's finding assembler needs exactly this: a way's firing evidence tied to +transcript pointers. The `SessionIntrospection` join *is* that evidence substrate. +Building it once, in `ways-core`, means findings and introspection read the same +correlation rather than two drifting re-derivations. + +## Consequences + +### Positive + +- One honest correlation, shared by three front-ends and the finding pipeline. +- The split-brain fix repairs `rethink` (and any events-log reader) for + post-migration installs — a real bug, not just a feature enabler. +- "Why fired" becomes precise for keyword/command/file once enrichment lands, and + honestly way-level for semantic — no fabricated term-level attribution. + +### Negative + +- Fire-time enrichment touches the hot fire path; the added fields must be cheap + and must never break the log's append-only, line-atomic contract. +- The model must encode *degrees* of join confidence (keyed vs. heuristic), which + is more complex than pretending every link is exact — but the honesty is the + point. + +### Neutral + +- Enrichment is forward-only; historical sessions keep the coarse heuristic join. +- Transcripts remain Claude-Code-owned and read-only; the substrate depends on + their availability and tolerates truncation-to-spill-file. + +## Alternatives Considered + +- **Union-read both event-log paths, leave the hooks hardcoded.** Rejected as the + durable fix: it papers over the split-brain and re-breaks the next time a path + moves. A single resolved writer path is the real correction (a bridging union + read on top is fine as a transition). +- **Reconstruct "why" purely by transcript replay, persist nothing new.** Split + outcome after §3's correction: transcript replay *is* the right source for the + **foreign key** (`transcript_uuid` via the `parentUuid` chain — a real message id + no fire-time capture can supply), but it *cannot* recover `matched_span` — the + injected content is way-anonymous, so which text matched an individual way is + lost. Hence the hybrid: post-hoc transcript for the key, fire-time enrichment for + the span. Pure replay alone also gives no "why" when the transcript is + absent/truncated (the join degrades to `Heuristic`), and semantic stays way-level + regardless. +- **Fake semantic term-level attribution** (highlight the "matching" vocabulary + word). Rejected: the corpus stores one vector per way; there is no matched term + to recover. Presenting one would be a confabulated explanation — the exact + epistemic error the compliance work (ADR-200) exists to avoid. + +## References + +- **ADR-142** — the XDG projection whose migration created the events-log split. +- **ADR-134** — the near-miss telemetry stream this model also surfaces. +- **ADR-201** — the finding pipeline that shares this transcript-evidence substrate. +- Research pass 2026-07-02 (events-log schema, transcript shape, frontmatter match + fields, join feasibility) — the ground truth this ADR is built on. diff --git a/docs/architecture/ways/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md b/docs/architecture/ways/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md new file mode 100644 index 00000000..536c9ed4 --- /dev/null +++ b/docs/architecture/ways/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md @@ -0,0 +1,202 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - matching + - cli +basis: + - evidence: rethink silently globalizes when project detection returns None, and has no --list --json for an agent to enumerate sessions + - precedent: ADR-153 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-153 + - ADR-111 +imported: + from: docs/architecture/system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md + format: v0 + status: Accepted +--- + +# ADR-154: `ways introspect` — one model, three front-ends + +## Context + +ADR-153 defines a shared `SessionIntrospection` model. This ADR decides the +surfaces over it. There are three, plus a drill-down, and they should feel like +one tool a user can move between — which §4 realizes as one `ways introspect +<mode>` command: + +- **post-hoc replay** of a finished session (exists today as `ways rethink`). +- a **live** monitor of the current session, same UI, refreshing as new ways fire. +- a **non-interactive dump** — structured output an autonomous agent reads to + investigate a session (exists today via `rethink --json`). +- **the why-fired drill-down** — from a list of fired ways, enter one and read the + way, its trigger criteria, and the clip of session that matched. + +The research pass established the current surface and its gaps: + +- `rethink` is raw **crossterm + hand-rolled ANSI strings** (not ratatui), two + sequential blocking loops (session picker → frame player), sharing the + `cmd/render` table writer with `ways list`. Single-panel, full-clear-redraw. +- The **model→(render | serialize) split already half-exists**: `reconstruct_frames` + (pure data) feeds both the TUI and `rethink_dump`'s JSON — the exact factoring + the three front-ends want. +- Two bugs (owned here as command semantics): `rethink` **silently globalizes** + when current-project detection returns `None` (run manually, `CLAUDE_PROJECT_DIR` + is unset), and there is **no `--list --json`** for an agent to enumerate sessions + before dumping one. (The events-log discovery bug is fixed in ADR-153.) +- `attend-chat` uses **iocraft** (a declarative TUI) with an async runtime for its + rich interface — the maintainer's rich-TUI precedent, but it drags `smol`/ + `async-channel` that `ways-cli` (fully synchronous) deliberately lacks. + +## Decision + +### 1. One `cmd/introspect/` module; front-ends differ only in their loop + +Generalize the proven triplet into a module (mirroring the `cmd/scan`, +`cmd/settings`, `cmd/lint` subdir convention): + +- **model** — build `SessionIntrospection` (ADR-153). +- **render** — extend `cmd/render`'s ANSI-`String` contract with + `render_way_detail` / `render_clip`; no new rendering paradigm. +- **front-ends** (the `ways introspect <mode>` modes of §4), each a thin loop over + the same model: + - **`replay`** — one-shot render → print, then the frame-player loop (today's + `rethink`); + - **`live`** — the existing `rethink::tui_loop` poll skeleton, re-reading on a + tick (see §3); + - **`dump`** — serde-serialize (extend `rethink_dump`); + - **drill-down** — a selection/tab loop compositing model panels, shared by + `replay` and `live`. + +**Model ↔ replay-timeline boundary (resolved 2026-07-03).** The +`SessionIntrospection` model and `rethink::build_frames` are **two distinct +views, not one clustering** — and are deliberately kept that way rather than +unified. `build_frames` is the *animation* projection: it folds the full event +stream (`way_fired` + `check_fired` + `way_redisclosed`) into cumulative +"what's active at epoch N" frames, and drives `replay`'s timeline. The model is +the *analytical* substrate: per-turn fired-way deltas joined to `MatchCriteria`, +`matched_span`, and the transcript key, and it drives `dump` and the drill-down's +"why" panel. The drill-down bridges them **by `way_id`** — focus a way in a +`build_frames` frame, and its detail is looked up in the model by id — so their +divergent epoch numbering never has to be reconciled. This preserves the proven +replay timeline untouched, keeps the model scoped to the fire stream (ADR-153), +and confines the JSON-vs-TUI drift the model guards against to the two surfaces +that actually share it (`dump` and the drill-down), both of which read the model. +Converging `build_frames` onto the model was rejected: it would load the +fire-stream substrate with cumulative-state, redisclosure, refire-threshold, and +token-timeline concerns that belong to the animation view alone. + +### 2. Keep crossterm; grow a small micro-compositor — do not adopt ratatui + +Build a ~200–400-line internal compositor over the existing ANSI-`String` panels: +`panel = Vec<String> + width`, a side-by-side placer, a scroll-window helper, a +tab-bar helper. This preserves the lean `opt-level=z` binary (~3.3 MB), adds **zero +dependencies**, and reuses `render.rs`'s output verbatim as the "matched clip" +content. + +Rejected for now: **ratatui** — ~+300–600 KB and 15–30 crates against a size-tuned +binary, a real compile hit under `lto`/`codegen-units=1`, and — decisively — its +immediate-mode `Line`/`Span` model breaks the ANSI-`String` contract that +`cmd/render` shares with `ways list`, forcing a rewrite of both. **Escape hatch, +stated up front:** if the drill-down is meant to become a genuine multi-pane, +mouse-driven, resizable, text-selectable inspector, adopt ratatui *at that point* +rather than grow a poor reimplementation and migrate later. The micro-compositor +is right for "live table + status bar" and "list-left / read-only-detail-right +with independent scroll"; ratatui is right for a true windowed inspector. + +### 3. Live refresh: poll-with-timeout + mtime/size stat gate — not `notify` + +`introspect live` reuses crossterm `poll(Duration)`: the timeout is the refresh interval +(~100–250 ms, imperceptible), on timeout re-read the *tail* of the (append-only) +events log + transcript and re-render, on key event handle input. Skip the +re-parse when the files' mtime/length are unchanged. This stays synchronous, adds +zero dependencies, and matches the project's existing poll style. Rejected: +`notify` (event-driven) — it forces a background thread + channel + a blocking +select that `attend-chat` solved only by bringing in an async runtime; its +large-tree advantage does not apply to two append-only files. + +### 4. Command surface and scoping semantics + +The surfaces unify under a single **`ways introspect <mode>`** command (ADR-111 +single-surface spirit) rather than sibling top-level verbs. Modes: + +- **`ways introspect replay`** — post-hoc replay of a finished session (the + behaviour `rethink` has today). Default to the **current Claude Code project**; + add `--all` for every project and keep `--project <path>` for a specific one. + When current-project detection returns `None`, **fail loud** (name the missing + marker) or fall back to cwd — never silently globalize. Compare on the encoded + project slug / normalized path, not a loose substring. +- **`ways introspect list --json`** — enumerate candidate sessions as structured + data so an agent can pick one before dumping it (closes the non-interactive + enumeration gap). +- **`ways introspect dump`** — the non-interactive JSON dump of a session (the + behaviour `rethink --json` has today). +- **`ways introspect live`** — live monitor of the current session; same scoping + default. +- **the drill-down** — a tab in `replay`/`live`: a fired-ways list; entering a + way opens a read-as-a-human panel with the way body, its trigger criteria + (ADR-153 `MatchCriteria`), and the matched session clip (precise once ADR-153 §3 + enrichment lands; heuristic-labelled before). + +**Migration:** `ways rethink` becomes a thin **deprecated alias** for `ways +introspect replay` (and `rethink --json` → `introspect dump`), printing a +one-line deprecation notice to stderr while continuing to work. Muscle memory +keeps working; the canonical surface is the consolidated one. + +## Consequences + +### Positive + +- Three surfaces + a drill-down from one model — no drift between what an agent + reads as JSON and what a human sees in the TUI. +- Zero new dependencies; the lean binary and the shared `render.rs` contract both + survive. +- `rethink` stops silently globalizing and gains machine-listable sessions. + +### Negative + +- The micro-compositor is hand-built layout code (panes, scroll, tabs) — bounded + (~200–400 lines) but genuinely new, and less capable than ratatui's widgets. +- A live `introspect live` loop re-reading files is more moving parts than a + one-shot dump; the stat-gate must be correct to avoid needless re-parse flicker. + +### Neutral + +- If the inspector's ambitions grow, the escape hatch to ratatui is deliberate and + documented — this decision is reversible, not a dead end. The concrete trigger + that would flip it: the drill-down needing **text selection or resizable / + mouse-driven panes**; short of that, the micro-compositor stays. +- Command naming is **decided**: one `ways introspect <replay|live|dump>` surface, + with `ways rethink` kept as a deprecated alias. The `think`/`rethink` verb pair + was rejected as too cute — the modes are plain and descriptive instead. + +## Alternatives Considered + +- **Adopt ratatui now** for real layout/widget primitives. Rejected for the + current scope on binary-size, compile-cost, and the `render.rs`-rewrite grounds + above — with the explicit escape hatch if scope grows. +- **Reuse `attend-chat`'s iocraft + async stack.** Rejected: it would drag an + async runtime into a deliberately synchronous CLI — a larger intrusion than + ratatui for less fit. +- **`notify` filesystem watcher for `think`.** Rejected: event-driven latency is + not needed for two append-only files, and it forces the async architecture + ways-cli avoids; poll + stat-gate is the lean match. +- **Keep three independent implementations** (rethink as-is, a separate think, a + separate dumper). Rejected: guarantees drift; the whole point is one model. + +## References + +- **ADR-153** — the `SessionIntrospection` substrate these front-ends render. +- **ADR-111** — the single-tool-surface consolidation spirit this ADR follows in + choosing one `ways introspect <mode>` command over sibling verbs. +- Research pass 2026-07-02 (TUI stack, crossterm-vs-ratatui trade study, refresh + mechanism, shared-model factoring) — the ground truth this ADR is built on. diff --git a/docs/architecture/ways/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md b/docs/architecture/ways/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md new file mode 100644 index 00000000..944745a2 --- /dev/null +++ b/docs/architecture/ways/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md @@ -0,0 +1,304 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +superseded_by: ADR-188#3 +basis: + - evidence: a live session fired five ways on one prompt, three by incidental keywords with embed scores 0.22, 0.15 and 0.09 + - evidence: 42 of 136 way files carry a pattern:, many bare common words; 5,011 near-miss events show the semantic lane landing just below threshold + - precedent: ADR-125 + - precedent: ADR-153 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-04 +deciders: + - aaronsb + - claude +related: + - 108 + - 125 + - 130 + - 134 + - 153 +imported: + from: docs/architecture/system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md + format: v0 + status: Accepted +--- + +# ADR-155: Semantic gating of the keyword channel and reasoning-channel rebuild + +## Context + +ADR-125 named the matching architecture: retrieval is embedding-only, the model's +cosine score is accepted as ground truth, and the `pattern:`/`commands:` regex +fields survive as *explicit deterministic triggers* — "fire this way exactly when +this pattern appears." In practice the keyword channel has drifted from that +role into fuzzy retrieval, and the code gives it veto-proof authority. + +### How matching works today + +In `match_prompt` (`tools/ways-cli/src/cmd/scan/mod.rs`), the regex pattern is +checked first. Any substring hit anywhere in the prompt fires the way +deterministically; the embedding is never consulted for that way. The embedding +is a fallback that can only *add* fires, never veto a lexical coincidence — +even though `batch_embed_score` has already scored **every** way against the +prompt before the candidate loop starts. The signal that would separate +incidental keywords from topical intent is computed, then discarded. + +Two adjacent surfaces compound this: + +- **PreToolUse** (`command()`/`file()`): ways match by regex only (`commands` + against the command string, `pattern` against the tool description). The + semantic lane exists on that surface but is wired only to *checks*. Where + Claude is about to act — the moment course-correction is most valuable — + ways have no semantic channel at all. +- **Stop hook** (`check-response.sh`): Claude's response is reduced to a grep + against a hard-coded 24-word whitelist (`api|test|debug|…|way|pr|…`), and the + surviving tokens are concatenated into the **next prompt's match query** — + including the regex channel. The whitelist words are, by construction, way + trigger words. If Claude's reply mentioned "way" or "pr", the next user + prompt keyword-fires `meta/knowledge` or `softwaredev/delivery/github` + regardless of what the user typed. The channel meant to give ways awareness + of Claude's reasoning is simultaneously too blind to capture it (24 fixed + words) and a noise injector into the keyword channel. + +### Observed failure (session introspection, ADR-153) + +A live session in a TUI-oscilloscope project fired five ways on one pasted +prompt. Rescoring that exact prompt against the corpus: + +| Way | fired via | EN embed score | topically relevant | +|---|---|---:|---| +| softwaredev/visualization | semantic | 0.40 | yes | +| softwaredev/visualization/charts | keyword `graph` | 0.28 | yes | +| meta/knowledge | keyword ` ways ` ("btop's various ways to render…") | 0.22 | no | +| meta/memory | keyword `remember` ("I remember the tektronix…") | 0.15 | no | +| softwaredev/delivery/github | keyword `github` (a pasted URL) | 0.09 | no | + +The inverse test — short prompts with genuine intent — scores high on the +target way: "adr" → 0.47 (documentation/adr), "remember this decision for +later sessions" → 0.35 (meta/memory), "let's look at the github pr checks" → +0.57 (delivery/github), "can you make a chart of the results" → 0.51 +(visualization/charts). The embedding rank-orders relevance correctly in every +observed case; a floor near half the fire threshold cleanly separates +coincidence from intent. **The model already knows which keyword fires are +junk; the scan loop doesn't ask it.** + +### Scale of the problem + +- 42 of 136 way files carry a `pattern:`. Many alternations are bare common + words (`commit`, `workflow`, `docs`, `remember`, `release`, `trade.?off` — + the last in two different ways) or unanchored substrings (`graph` matches + "photograph"). These are vocabulary words doing pattern work. +- Telemetry (`$XDG_STATE/agent-ways/events.jsonl`) records one keyword fire + whose `matched_span` is a 120-character stretch of prompt — a greedy `.*` + alternation matching essentially arbitrary text. +- 5,011 near-miss events show the semantic lane frequently lands just below + threshold — the opposite imbalance: keyword over-fires while semantic + under-fires. + +## Decision + +Give the already-computed embedding score veto power over lexical coincidence, +and rebuild the response-topics channel so Claude's reasoning feeds the +semantic lane instead of polluting the regex lane. Five parts: + +### 1. Semantic plausibility gate on keyword fires + +A `pattern:` hit on the prompt/task surface fires only if the way's embedding +score also clears a **gate floor**: + +``` +gate_floor = keyword_gate_fraction × effective_threshold(way) +``` + +with `keyword_gate_fraction` a global config value (default **0.4**, clamped +to [0, 1] on load — a fraction above 1.0 would put the floor above the fire +threshold and invert the gate's meaning), applied per model lane the same way +fire thresholds are (EN and multi each gate on their own floor; clearing +either lane passes). Parent-boost applies to the effective threshold before +the fraction, so a keyword hit under a fired parent is gated more leniently — +consistent with ADR-125's disclosure semantics. + +With the defaults this yields floors of 0.11–0.16. Against the observed data, +three score bands emerge: lexical coincidences land at 0.04–0.15 (gated); +multi-word intent phrases the pattern author deliberately encoded ("ship it" +against the github way) land at 0.17–0.18 (must fire); topical prompts land +at 0.26+ (fire with margin). The 0.4 fraction places the floor in the gap +between the first two bands. Calibration review found 0.5 (floor 0.20) +vetoed the intent-phrase band. + +A lane that ran but returned no row for a way that *is* embeddable (has +description + vocabulary) is treated as a score of 0.0, not as absent signal: +`way-embed` emits only non-negative cosines, so a missing row means negative +cosine — the strongest "unrelated" verdict. Fail-open is reserved for cases +with genuinely no evidence: the engine didn't run, or the way is trigger-only +and cannot be in the corpus. This keeps the gate monotonic in relatedness +(without it, the worst matches would out-fire mild ones). + +The gate is a *veto on noise*, not a second retrieval tier: the pattern +remains the explicit trigger, and the score consulted is the one +`batch_embed_score` already produced for the fire path. Zero additional model +invocations on the common path; when part 3's response context contributed to +the shared embed vector, a hit that lands below the floor is re-checked +against a lazily computed prompt-only embedding before the veto stands — the +user's own typed keyword must never be gated by what Claude said last turn. +That second pass runs at most once per scan, only on turns where a hit was +actually gated with response context present. + +**Escape hatch:** a way may declare `pattern_strict: true` in frontmatter to +opt out of gating **and of §2's masking** — strict means "fire on exactly +this text, always": slash-command references like `/wrap`, or patterns that +deliberately target URL content (`github\.com/\S+/pull`), which the mask +would otherwise hide from the regex lane. The lint warns when +`pattern_strict` is combined with common-word alternations. + +**Telemetry:** a gated-off hit logs a `way_keyword_gated` event carrying +`matched_span`, both scores, both floors — the same shape as `way_nearmiss` +(ADR-134). The tuning passes consume this stream to calibrate +`keyword_gate_fraction` empirically before any tightening. A gated fire is +also visible in `ways introspect` as a suppressed-candidate row, so a "why +didn't X fire" question is answerable from the session record. + +### 2. Mask non-linguistic spans before regex matching + +URLs (`https?://\S+`) and fenced code blocks are masked out of the query +**for the keyword channel only** before pattern matching. A pasted link +containing "github" is not GitHub-workflow intent. The embed query is +untouched — the ADR-130 salience reducer already handles long pasted content +for that lane. Ways whose patterns deliberately target URL or code content +opt out with `pattern_strict: true` (§1's escape hatch covers both the gate +and the mask). + +### 3. Rebuild the response-topics channel + +`check-response.sh`'s 24-word grep whitelist is removed. The Stop hook stores +the last assistant message **raw** (bounded excerpt) — no extraction at Stop +time at all; the ADR-130 sentence-salience reducer selects what matters at +scan time, where it already runs on every prompt. On the next prompt scan, +that stored text rides a separate `--response-context` flag and feeds **only +the embed query, never the regex query**. The keyword channel matches what +the user actually typed; the semantic channel sees user intent *plus* what +Claude was just reasoning about. A turn with nothing extractable (tool-only +stop) clears the state file rather than leaving a stale response to be +re-embedded as current. The hook degrades to the flagless invocation if the +installed binary predates the flag, so a hooks-before-binary deploy order +cannot block prompts. + +This is the fix for under-firing on Claude's own reasoning: full-sentence +salient content replaces a fixed vocabulary, and it reaches the lane designed +to interpret it. + +### 4. Semantic lane for ways at PreToolUse + +`command()` already computes embedding scores for the reduced +`command + description` query (used by checks). Ways gain the same lane: a way +whose score clears its effective threshold fires on the bash surface, exactly +as it would on the prompt surface. The tool `description` field is Claude's +own natural-language statement of intent, which is the right embed input at +act-time. No new embedding work — the scores are already computed per +PreToolUse event. Existing re-disclosure suppression (decay curves) bounds +repeat injections; the near-miss/fire telemetry monitors this surface for +noise before any threshold adjustment. + +### 5. Pattern hygiene: demote vocabulary words out of patterns + +Doctrine, enforced by lint and applied by a rework pass over the 42 +pattern-bearing ways: + +- **`pattern:` is for exact, high-precision triggers** — command names, term + -of-art tokens (`adr`, `diataxis`, `mermaid`), phrases with anchoring + structure. Word-boundary anchoring required for short alternations. +- **Suggestive common words belong in `vocabulary:`** — they shape the + embedding coordinate, which is the lane built to weigh them contextually. + `remember`, `commit`, `workflow`, `docs`-alone move out of patterns. +- New lint rules: flag unanchored alternations under a length floor, bare + dictionary-common words, and `.*` in prompt patterns. +- `ways tune --precision` gains per-alternation attribution: `matched_span` + (ADR-153) joined with the off-class session heuristic (ADR-134) identifies + which alternation produces off-domain fires, producing a ranked demotion + worklist instead of hand-auditing 42 files. + +### Sequencing + +Parts 1–2 land first (one file, `scan/mod.rs`, plus config/frontmatter +plumbing) — they make the system safe against the *existing* noisy patterns. +Part 3 and 4 follow. Part 5 (frontmatter rework + corpus rebuild + lint) is +**blocked on a deployed release of parts 1–4**, not merely on their merge: +the rework pass tunes against live behavior — `way_keyword_gated` telemetry +saying which alternations the gate is actually vetoing, and the bash-surface +semantic fires showing where vocabulary already carries a way without its +pattern. Without that stream the demotion decisions are guesses; with it, +each pattern edit is validated against what the gate observed in real +sessions. Part 5 also rebuilds the corpus (vocabulary changes move every +reworked way's embedding coordinate) and is validated by embed-scoring +regression prompts, so it runs as its own branch with the gate already +protecting the transition. + +## Consequences + +### Positive + +- False keyword fires on incidental words, pasted URLs, and echoed response + topics are suppressed using signal that is already computed — no latency or + model cost. Context stops being spent on off-topic way bodies, and ways stop + training the agent to ignore them. +- Ways gain a semantic channel at PreToolUse — guidance can reach the moment + before action, which keyword-only matching mostly missed. +- Claude's reasoning enters matching as reduced salient sentences in the + semantic lane, replacing a 24-word whitelist that leaked trigger tokens into + the regex lane. +- Every suppression is observable (`way_keyword_gated`, introspect rows) and + the gate fraction is tunable from telemetry before any further tightening — + the ADR-134 pattern applied to a new stream. + +### Negative + +- A way with an exact pattern but weak `description`/`vocabulary` text can be + wrongly gated: its corpus coordinate under-represents what the pattern + targets. Mitigations: `pattern_strict: true`, and the part-5 rework + explicitly checks that pattern terms appear in the way's embeddable text. +- The keyword channel is no longer fully deterministic from the way file + alone; "why didn't my pattern fire" now has a second cause. Introspect + surfacing the gated row is the answer path. +- Behavior change for existing installs: some previously-firing ways go + quiet. The gate defaults are deliberately loose (0.4 × threshold) and + telemetry-adjustable. + +### Neutral + +- `meta/knowledge`'s ` ways ` alternation lands near the default floor + (0.22 vs 0.20) — the borderline case that telemetry, not this ADR, should + settle. +- Threshold rebalance (lowering semantic defaults where near-misses cluster + on-domain) becomes attractive once keyword noise is gated, but is out of + scope here — it stays with the ADR-134 tuning passes. +- The state channel (session-start fires such as `freshness`) has its own + value question, untouched by this ADR — different trigger class, different + economics. + +## Alternatives Considered + +- **A language-model classifier over candidate fires.** Rejected on cost and + latency: every prompt and tool call would pay an LM round-trip, and the + evidence shows MiniLM cosine already carries the discriminating signal — + every observed false fire scored below 0.22, every observed true fire above + 0.28. +- **Weighted lexical scoring (BM25-like multi-term evidence).** Rejected as + re-litigating ADR-125, which removed BM25 and named the embedding the single + retrieval mechanism. The gate keeps that shape: lexical stays binary and + explicit; the embedding stays the only scorer. +- **Authoring-only fix (rework the 42 patterns, change no code).** Necessary + but not sufficient: authoring regresses (this drift happened under the + current lint), project-local ways repeat the same mistakes, and no pattern + hygiene fixes the Stop-hook whitelist or the missing PreToolUse lane. +- **Raising semantic thresholds to compensate.** Backwards: semantic is the + under-firing channel (5,011 near-misses). The noise source is lexical. +- **Removing the keyword channel entirely.** Rejected: ADR-125's case for + explicit deterministic triggers stands — slash commands, terms of art, and + `commands:` regexes at PreToolUse are legitimately exact. The problem is + authority without corroboration, not existence. diff --git a/docs/architecture/ways/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md b/docs/architecture/ways/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md new file mode 100644 index 00000000..7dc265fd --- /dev/null +++ b/docs/architecture/ways/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md @@ -0,0 +1,187 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'a fit of g(s) over 96 intent/noise probes across six ways: one global calibration separates at pooled AUC 0.956; the default embed_threshold 0.40 sits at P≈0.95 and the gate floor 0.16 at P≈0.03' + - precedent: ADR-155 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-04 +deciders: + - aaronsb + - claude +related: + - ADR-155 + - ADR-125 + - ADR-134 +imported: + from: docs/architecture/system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md + format: v0 + status: Accepted +--- + +# ADR-156: Calibrated relevance scoring for the semantic lane + +## Context + +The matching engine thresholds **raw cosine similarity**. The semantic lane +fires when `s = cos(query, alias) ≥ T_w`, and the ADR-155 keyword gate admits a +pattern hit when `s ≥ γ·T_w` (`γ = keyword_gate_fraction`, default 0.4). Both +boundaries live on the raw cosine scale, and each way's `embed_threshold` is set +by hand. + +Raw cosine is not comparable across ways. A cosine of 0.30 against one alias can +be strongly relevant while against another it is noise, because aliases occupy +different regions of the embedding space (different vocabulary breadth, different +neighbourhoods). The per-way `embed_threshold` exists to absorb that +incomparability — it is, in effect, a **single hand-placed calibration point per +way**. The design note *The Lexical Gate as a Conditional Threshold* reads the +current rule as the decision boundary of Bayesian log-odds evidence fusion for a +binary lexical feature, and identifies the one component the implementation +lacks: `g(s)`, a calibration from cosine to a relevance probability. + +We fit `g(s) = σ(a·s + b)` over intent/noise probe sets across six ways spanning +the regimes (96 probes). The findings (provisional, small-sample, but +consistent): + +- A **single global calibration** separates intent from noise at pooled + **AUC 0.956**, with the `P = 0.5` boundary at `cos ≈ 0.29`. +- On the calibrated scale, today's constants are revealed as mis-set: the + default `embed_threshold = 0.40` corresponds to `P ≈ 0.95` (the semantic lane + fires only at 95% confidence), while the gate floor `γ·T = 0.16` corresponds + to `P ≈ 0.03` (the keyword lane fires at 3% confidence). The keyword lane is + load-bearing precisely because the semantic threshold is set far too high, and + it leaks precisely because the floor is set far too low. +- Per-way `embed_threshold` values largely **collapse** once calibrated: five of + six ways sit at the untuned default; the separable ones share essentially one + boundary. The per-way knob is mostly compensation for the uncalibrated scale, + not per-way signal. +- Ways whose keyword-conditioned intent and noise distributions **overlap** + (e.g. `meta/memory`, AUC 0.80) are exposed by a low per-way AUC — a principled, + measurable signal that no threshold can save that keyword. + +Three standing problems follow from thresholding raw cosine: the thresholds are +not portable, they **drift silently when the corpus is regenerated** (a +hand-tuned constant is calibrated to one embedding snapshot), and because scores +are never turned into a common currency, the two model lanes (EN, multi) cannot +be combined — they are OR-ed with separately hand-set thresholds. + +## Decision + +Introduce `g(s)`: a per-model calibration from cosine to relevance probability, +and express every fire threshold in probability space. + +1. **Calibration.** For each embedding model `m`, fit + `g_m(s) = σ(a_m·s + b_m)`. Calibration is a **fitted, versioned artifact**, + produced offline from a committed, curated probe corpus and validated on + held-out probes (fit is rejected below an AUC floor). It is stamped with the + model and corpus version it was fit against, so a corpus regeneration that + moves embeddings triggers a refit rather than silent drift. The numbers above + are illustrative of the shape, not the production fit. + +2. **Fire rule in probability space.** For a way `w` and prompt `q` with keyword + indicator `k`: + + fire ⇔ g_m(s) ≥ τ_s(w) ∨ ( k ∧ g_m(s) ≥ τ_k ) + + evaluated per model and OR-ed across models (probabilities are now + comparable). `τ_s` is the **semantic threshold** (global default, e.g. + `P = 0.5`) and `τ_k` is the **keyword floor** (global default, e.g. + `P ≈ 0.15`). Both are absolute probabilities. + +3. **The keyword floor is decoupled from the semantic threshold.** `τ_k` and + `τ_s` are independent, so a leaky keyword can be tightened without raising the + semantic bar. This subsumes `keyword_gate_fraction`: the fixed ratio `γ` is + retired in favour of two independent probabilities — the coupling named in + the design note is resolved as a side effect of moving to probability space. + +4. **Global thresholds only.** One `(a_m, b_m)` per model and one global + `(τ_s, τ_k)`. No per-way threshold override ships in v1: calibration removes + the incomparability that per-way `embed_threshold` existed to patch (AUC 0.956 + under a single boundary), and a way that still fails to separate is a + way-content problem for the pattern-hygiene sweep (fix its alias or keyword), + not a knob. A per-way probability override can be reintroduced later if + telemetry identifies a way that genuinely needs one. + +5. **Clean cutover — no compatibility mode.** Raw-cosine thresholding, + `keyword_gate_fraction`, and the raw `embed_threshold` field are removed, not + shimmed. Carrying a second scoring path would mean maintaining the exact ruler + this ADR replaces. Existing per-way `embed_threshold` values are stripped from + way frontmatter in the same change — the ways fall to the calibrated global + boundary, and the pattern-hygiene sweep is already touching every + pattern-bearing way. `pattern_strict` still bypasses the gate. Calibration is + generated together with the corpus (`ways corpus`), so it is present wherever + embeddings are; the only degenerate path retained is the existing genuine + no-embedding case (non-embeddable way, or engine not run), which continues to + fail open on the author's keyword. The change ships as a version bump the + operator deploys deliberately. + +Telemetry-based refitting from the ADR-134 streams (`way_fired`, +`way_nearmiss`, `way_keyword_gated` are the observed `S⁺`/`S⁻` samples) is the +intended successor to the offline fit, but is **not** in this decision: the +feature must be live to generate calibrated telemetry. + +## Consequences + +### Positive + +- One interpretable, comparable decision boundary. Per-way threshold hand-tuning + becomes the exception, not the norm. +- The keyword floor and semantic threshold are independently settable; the + over-loose gate and over-strict semantic lane can each be corrected, and the + fixed `γ` coupling is retired. +- Thresholds stop drifting on corpus regeneration: calibration is refit and + version-stamped, not silently invalidated. +- Model lanes become combinable (comparable probabilities), enabling future + log-odds composition of EN + multi and of per-alternation lexical evidence. +- The pattern-hygiene sweep (ADR-155 §5) gains an **objective** remove/keep + metric — per-keyword AUC and the calibrated boundary — replacing hand-set + intent/noise heuristics. + +### Negative + +- A new fitted artifact to own and version. A bad fit degrades every way at once; + mitigated by an AUC validation gate that rejects a bad fit at generation time + and by version stamping. A missing or invalid calibration is a + corpus-generation error, not a silent fallback to the retired raw-cosine path. +- Behaviour change for existing installs, delivered as a clean cutover: the + firing set shifts (the intended correction) and existing per-way + `embed_threshold` values are removed in the same release. Gated behind a + version bump the operator deploys, watched via the near-miss / gated telemetry. + There is no rollback short of reverting the release — acceptable because the + install is versioned and operator-deployed. + +### Neutral + +- Sequences the remaining ADR-155 §5 work behind this: the corpus rescore and the + pattern-hygiene sweep run *after* calibration lands, against the better ruler. +- Establishes the probe corpus as a committed regression asset and sets up + telemetry refitting and per-alternation weighting as named follow-ons. + +## Alternatives Considered + +- **Keep thresholding raw cosine, hand-tune per way (status quo).** Rejected: not + portable, drifts on regeneration, and forces the strict-semantic / loose-gate + split that makes the keyword lane both load-bearing and leaky. +- **Score-level fusion (weighted sum / RRF of lexical and dense scores), + thresholded once.** The mainstream hybrid-retrieval combiner. Rejected as a + larger rewrite for no additional benefit here: the existing gated two-threshold + rule already *is* the decision boundary of log-odds fusion for a binary lexical + feature (design note), so calibrating the score achieves the same gain with a + smaller, more legible change. +- **Per-way calibration curves.** Rejected for this decision: needs per-way + labelled data and reintroduces the per-way tuning burden calibration removes. A + global fit plus an optional per-way `τ_s` reaches AUC 0.956. +- **Fit from telemetry on day one.** Rejected: the feature must be live to emit + calibrated telemetry. Offline fit from a curated probe corpus bootstraps it; + telemetry refitting follows. +- **Backward-compatible rollout (dual scoring paths).** Keep raw-cosine + thresholding and the legacy `embed_threshold` working alongside the calibrated + path. Rejected: it would require maintaining the exact ruler this ADR replaces, + the firing behaviour would depend on which path a way happened to take, and the + install is versioned and operator-deployed — a clean cutover at a version bump + is both safe and honest, where a compatibility shim is permanent drift surface. diff --git a/docs/architecture/ways/ADR-157-case-insensitive-trigger-regex-compilation.md b/docs/architecture/ways/ADR-157-case-insensitive-trigger-regex-compilation.md new file mode 100644 index 00000000..b5bbca07 --- /dev/null +++ b/docs/architecture/ways/ADR-157-case-insensitive-trigger-regex-compilation.md @@ -0,0 +1,127 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'the PR #301 review: lowercase patterns such as \bssh\b missed the uppercase acronyms users type, and five patterns were patched with inline (?i)' +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-04 +deciders: + - aaronsb + - claude +related: + - ADR-155 + - ADR-156 +imported: + from: docs/architecture/system/ADR-157-case-insensitive-trigger-regex-compilation.md + format: v0 + status: Accepted +--- + +# ADR-157: Case-insensitive trigger regex compilation + +## Context + +The keyword lane compiles each way's `pattern:` (and the `cmds:` / `files:` +patterns) with `Regex::new(...)` and matches it against the **original-case** +text. `mask_nonlinguistic` (ADR-155 §2) strips fences and URLs but does not +lowercase, and the query path (`match_prompt`, `scan/mod.rs:596`) passes the +masked query straight through. Only the tool-description path +(`scan/mod.rs:354`) lowercases its input, and it does so on the *text*, not the +pattern. + +The consequence: a lowercase author pattern silently misses the uppercase +acronyms users actually type. `\bssh\b` misses `SSH`; `\berd\b` misses `ERD`; +`\bpr\b`, `ADR`, `DBML`, `MTTR` all leak the same way. This surfaced in the +PR #301 review, which patched the acute offenders by prepending an inline +`(?i)` flag to five patterns: + +- `hooks/ways/meta/subagents/subagents.md` +- `hooks/ways/meta/introspection/introspection.md` +- `hooks/ways/workstation/pkghistory/pkghistory.md` +- `hooks/ways/softwaredev/environment/ssh/ssh.md` +- `hooks/ways/data/documentation/documentation.md` + +That is per-way cruft. It fixes the five patterns someone happened to notice and +leaves every other acronym-bearing pattern latent — the next author who writes +`\bpr\b` re-introduces the bug and won't know why their way never fires on `PR`. +Case sensitivity is the wrong default for a trigger channel: an author writing a +keyword means the *concept*, not a specific casing. + +**Blast-radius survey of the current corpus** (why a global fix is safe): + +| Lane | Cased patterns today | Effect of case-insensitivity | +|------|----------------------|------------------------------| +| `pattern:` (keyword) | 5× `(?i)` + `SKILL\.md` | Intended fix. `SKILL\.md` still matches `skill.md` — same concept. | +| `cmds:` | none | No-op — no uppercase-bearing command patterns exist. | +| `files:` | `README\.md$`, `Makefile$\|makefile$\|GNUmakefile$`, `Makefile$` | Desirable — READMEs and Makefiles have real casing variants; the `makefile$` alternation branch becomes redundant-but-harmless. | + +No pattern in the corpus relies on case-sensitivity to *avoid* a match. The +helpers (`regex_matches`, `regex_span`) are private to `scan/mod.rs` and serve +all three lanes, so the cleanest change lives in one place. + +## Decision + +Compile trigger regexes **case-insensitively** by building them with +`regex::RegexBuilder::new(pattern).case_insensitive(true)` in the two shared +helpers `regex_matches` and `regex_span` (`scan/mod.rs`). This applies uniformly +to the keyword, command, and file lanes. + +Because the keyword regex now matches case-insensitively, the tool-description +path no longer needs to pre-lowercase its text: `scan/mod.rs:354` changes from +`regex_span(pat, &desc.to_lowercase())` to `regex_span(pat, desc)`, which also +yields a truer original-case `matched_span` in telemetry (ADR-153 §3). + +Then **retire the five inline `(?i)` flags** — they become redundant. The +patterns revert to their plain form; behavior is preserved by the global flag. + +The invariant, stated once so future authors inherit it: *the keyword lane +matches case-insensitively; write patterns in lowercase and mean the concept.* +This lands in the engine-reference and the authoring surfaces. + +## Consequences + +### Positive + +- Every acronym-bearing pattern (`PR`, `ADR`, `SSH`, `ERD`, `DBML`, `MTTR`, …) + matches the uppercase form users type — corpus-wide, not just the five noticed. +- Removes per-way `(?i)` cruft and the latent-bug trap it papered over. +- Truer `matched_span` telemetry on the description path (original case, not + lowercased). +- One compile-site invariant replaces a convention every author had to remember. + +### Negative + +- An author who *wants* case-sensitive matching (e.g. to distinguish `OK` from + `ok`) can no longer get it via the shared helpers. No current pattern needs + this; if one ever does, it can carry an inline `(?-i)` scope — the regex crate + supports per-pattern override, so the global default is not a hard ceiling. +- Marginally wider matching on `files:`/`cmds:` (e.g. `.ENV` now matches an + `\.env$` pattern). Reviewed as desirable, not a regression, for the current + corpus. + +### Neutral + +- The `makefile$|GNUmakefile$` explicit-casing alternations are now redundant; + they are left as-is (harmless) rather than churned in this ADR's scope. +- `RegexBuilder` is already in the `regex` crate dependency — no new deps. + +## Alternatives Considered + +- **Lowercase the text instead of the pattern.** Rejected: globally lowercasing + the query breaks any deliberately-uppercase pattern (`SKILL\.md` would need + the text cased to match) and mangles the captured span. The regex flag matches + the *pattern* case-insensitively without touching the text — strictly safer. +- **Keyword-lane-only case-insensitivity** (a separate helper used only at + `:596`, leaving `cmds:`/`files:` case-sensitive). Rejected: the same + `pattern:` field is consumed at both `:354` and `:596`, so splitting behavior + by call-site would make one field match two ways; and the survey shows + case-insensitivity is *desirable* for `files:` (READMEs, Makefiles) and a + no-op for `cmds:`. A uniform rule is simpler and correct. +- **Keep patching per-way with `(?i)`.** Rejected: it is the status quo that + produced the bug — it fixes only noticed patterns and re-arms the trap for the + next author. diff --git a/docs/architecture/ways/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md b/docs/architecture/ways/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md new file mode 100644 index 00000000..d1b71a7a --- /dev/null +++ b/docs/architecture/ways/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md @@ -0,0 +1,154 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: a live session where 104 of 157 ways fired, one scan firing 35; the fire panel measured 8 fires (max 17) on adjacent prompts under the deployed calibration + - precedent: ADR-156 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-04 +deciders: + - aaronsb + - claude +related: + - ADR-156 + - ADR-155 + - ADR-125 +imported: + from: docs/architecture/system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md + format: v0 + status: Accepted +--- + +# ADR-158: Calibration boundary quality — hard negatives and a fire-breadth ship gate + +## Context + +Observed in a live session: *nearly every way fired.* Telemetry confirmed it — +104 of 157 ways fired across the session, one scan firing 35 ways. The cause is +**not** the keyword lane (ADR-157) and **not** fail-open (calibration loaded, EN +AUC 0.955). It is the **semantic lane's boundary position and shape**. + +The deployed calibration is `g(cos) = σ(26.66·cos − 7.81)` (EN). Two facts fall +out of those coefficients: + +1. The fire bar `τ_s = 0.5` maps to **cosine ≥ 0.293** — a low similarity + threshold. Any way whose alias sits within 0.293 cosine of the prompt fires. +2. The slope (a ≈ 26.66) is near-vertical: past cosine ~0.45 the probability + **saturates to ~1.0**. A prompt "write an adr documenting this architecture + decision" fires `delivery/implement` at cos 0.603 → g **0.9997** and + `architecture/design/prototype` at cos 0.573 → **0.9994** — both scoring + higher than they should, and the ranking signal at the top is gone. + +A purpose-built instrument (`tools/scripts/fire-panel.py`, scoring one prompt +against all aliases) quantifies the breadth against an expectation panel +(`fire-panel.json`): + +| bucket | expect | fired (cos 0.293) | +|---|---|---| +| off_topic (France, a haiku, Everest) | ~0 | **0.0** | +| narrow (rename a var, one unit test) | 1–4 | **1.0** | +| adjacent (write an ADR, review a PR) | a few | **8.0** (max 17) | +| broad (a wrap prompt naming every domain) | many | **17** | + +So it is not "fires on everything" — genuinely off-topic prompts fire zero. It +is **"fires on everything on-topic, saturated."** Discrimination between +*relevant* and *tangential* has collapsed. + +**Root cause.** The calibration is fit at corpus generation from the committed +probe corpus (`calibration_probes.jsonl`, `include_str!` at `corpus.rs:765`) — +**96 probes across 6 ways**, and each probe is scored only against *its own* way's +alias (`corpus.rs:797`). The *noise* probes are **easy** — clearly off-topic, low +cosine — so the logistic fit places the boundary low (cos 0.293) and the slope +steep (a clean gap between easy negatives and intents produces a near-vertical +sigmoid). **High AUC on this probe set does not bound the real false-positive +rate**, because the negatives do not represent the true adversary: *adjacent-domain +real prompts* that sit at moderate-to-high cosine but should not fire. + +A compounding factor is self-inflicted: the corpus embeds `description + +vocabulary` (`corpus.rs:380`, `load_aliases` at `:857`). ADR-155 §5's +pattern-hygiene sweep moved common words *out of* `pattern:` and *into* +`vocabulary:` — which **broadens each swept way's alias centroid**, nudging +on-topic cosines up. Pattern hygiene became alias bloat. + +The residual after any global threshold move is instructive: even at cos 0.45, +"write an adr" still fires `implement` (0.603) and `prototype` (0.573) — because +those *aliases* are too broad (they carry ADR/architecture vocabulary). A global +threshold cannot separate them from the legitimate `adr` fire (0.735); only +tightening the aliases can. + +## Decision + +Treat **boundary quality** — not just probe separability — as the calibration's +ship criterion, and fix it along three axes. + +1. **Hard negatives in the probe corpus.** `calibration_probes.jsonl` must carry + *adjacent-domain* hard negatives: prompts with **high** cosine to a way that + should **not** fire it (label 0). Placing negatives *in the gap* flattens the + slope (de-saturating the probabilities) and raises the crossover (fewer false + fires) in one coherent refit — the principled version of moving the boundary, + as opposed to bending `τ_s`. A rich source is cross-way mining: one way's + intent probes are hard negatives for an embedding-adjacent way they must not + trigger. + +2. **A fire-breadth ship gate.** `tools/scripts/fire-panel.{py,json}` — a + committed panel with per-bucket expectations — is a regression asset checked + alongside `AUC_FLOOR`. A corpus build regresses if `off_topic > 0`, or a + bucket's fire-breadth rises materially versus the recorded baseline. AUC + measures probe separability; the panel measures the thing users feel. Both + gate a refit. + +3. **Per-way alias discipline.** A way's alias (`description + vocabulary`) is its + semantic fingerprint; over-broad vocabulary causes cross-domain bleed. Aliases + the panel shows bleeding get **tightened** — the counter-discipline to + ADR-155 §5. Pattern hygiene may not silently become alias bloat: a word moved + out of `pattern:` belongs in `vocabulary:` only if it is genuinely + discriminating for *this* way, not merely suggestive. + +The `τ_s` config value stays at 0.5 as the calibrated-probability contract +(ADR-156); the boundary moves by fixing the *fit*, not the threshold. + +## Consequences + +### Positive + +- Adjacent-prompt precision rises (fewer tangential ways injected); the context + window stops filling with 8–35 marginal ways. +- Flattening the slope restores meaningful probabilities across the range, so + parent-boost and near-miss ranking regain signal. +- The fire-breadth gate makes over-firing a *caught regression*, not a thing a + user notices in production — and turns a felt symptom into a measured number. +- Establishes the discipline that a lint suppression / vocabulary choice is a + claim to be measured (shared with the `pattern_keep` governance in #308). + +### Negative + +- Authoring and maintaining hard-negative probes and the panel is ongoing work. +- Hard negatives can drop AUC below `AUC_FLOOR` if a way's intents and its + adjacent negatives are truly inseparable by cosine-to-one-alias. When that + happens the fix is **alias tightening**, not more negatives — the signal, not + the threshold, is the limit. + +### Neutral + +- The panel measures raw `τ_s` fires and does not model parent-boost (which only + *lowers* a child's bar), so it is a sound lower bound and a consistent + before/after proxy, not an exact production fire count. +- Multilingual calibration is fit from the English probe corpus (ADR-156); the + hard negatives benefit both lanes. + +## Alternatives Considered + +- **Raise `τ_s`** (e.g. to 0.9 → cos bar ~0.375). Rejected as the primary fix: a + weak lever under this slope (0.5→0.99 moves the bar only 0.293→0.465), it + entangles parent-boost (whose base is `τ_s`), and it papers over the saturation + rather than fixing it. Retained only as an emergency relief valve. +- **Per-way thresholds.** Rejected — ADR-156 deliberately removed them in favour + of one global calibrated scale; re-introducing them abandons that model. +- **Do nothing / accept the breadth.** Rejected — 104/157 ways firing wastes the + context budget the whole system exists to protect (ADR-125), and saturated + probabilities disable the ranking machinery downstream of the fire decision. diff --git a/docs/architecture/ways/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md b/docs/architecture/ways/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md new file mode 100644 index 00000000..f0fd073d --- /dev/null +++ b/docs/architecture/ways/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md @@ -0,0 +1,117 @@ +--- +contract: adr/v1 +kind: decision +verb: retire +capability: disclosure +targets: + - cli:ways-tune-curves +basis: + - evidence: 'ways tune-curves --apply writes a curve: block that ways lint rejects as an UNKNOWN field, and no shipped way carries curve:' + - precedent: ADR-126 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-04 +deciders: + - aaronsb + - claude +related: + - ADR-123 + - ADR-126 + - ADR-134 +imported: + from: docs/architecture/system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md + format: v0 + status: Accepted +--- + +# ADR-159: Remove ways tune-curves and the legacy curve: cadence field + +## Context + +`ways tune-curves` (ADR-123 Phase E) reads fire telemetry, computes the median +token-distance between a way's firings, and — with `--apply` — rewrites each +way's frontmatter to a `curve:` block carrying an absolute `half_life`. + +That output is now **broken**. ADR-126 replaced the `curve:` block with +`refire:` (a *fraction of the context window*, resolved per fire against the +model's actual window). The migration moved every way to `refire:`, `curve:` was +dropped from the frontmatter schema, and `ways lint` flags a written `curve:` +block as an UNKNOWN field. So `ways tune-curves --apply` produces frontmatter +that the project's own linter rejects — a command that corrupts way files. + +The command is not worth repairing: + +- **Its model is superseded.** It suggests an absolute `half_life` in tokens. + `refire:` is deliberately a *fraction*, so way files stay portable across model + window sizes (ADR-126). Translating an observed token cadence into a fraction + requires dividing by the window it was observed under — reintroducing exactly + the model-specific coupling ADR-126 removed. A "portable" fraction derived from + one model's window is a fiction. +- **Its successor is a different, deferred design.** Telemetry-driven cadence + tuning is ADR-134 (empirical auto-tuning from fire/near-miss streams), which is + deferred. `tune-curves` is a half-built manual precursor to it, not a + standalone capability worth carrying. +- **Nothing uses the legacy path.** No shipped way carries a `curve:` block; the + schema rejects it. The `curve:` frontmatter field and its read-fallback are + dead code kept alive only for a command that produces lint-failing output. + +## Decision + +Remove `ways tune-curves` and retire the legacy `curve:` frontmatter field +entirely. + +- Delete the `tune_curves` command, its module, and its CLI wiring. +- Remove the `Frontmatter.curve` field and its read-fallback in + `resolved_curve` — resolution now comes solely from `refire:`. +- Remove the now-dead `curve:` readers (`show`, `list` doc comments) and the + lint special-case that warned about `refire:`+`curve:` coexistence; a stray + `curve:` block falls through to the generic UNKNOWN-field warning, which is the + correct treatment for a retired field. +- Update the docs that presented `tune-curves` as a workflow (`stats.md`, + `reference/ways-cli.md`) and any `curve:`-as-current references. + +**Kept:** the runtime `Curve` type and `RefireSpec::to_curve`. `Curve` is the +concrete decay representation that `refire:` *resolves into* at fire time +(`fraction × window → Curve::Exponential`); it is the engine's internal shape, +not the retired authoring field. Retiring `curve:` the *frontmatter field* does +not touch `Curve` the *runtime type*. + +## Consequences + +### Positive + +- No command can emit lint-failing frontmatter; the footgun is gone. +- One cadence model, not two: `refire:` is the sole authored cadence field, with + no dead legacy path shadowing it in the parser, `show`, `list`, and lint. +- Less code to carry toward the eventual ADR-134 auto-tuner, which will target + `refire:` directly rather than inheriting `tune-curves`' `half_life` model. + +### Negative + +- Users lose the observed-cadence *suggestion* helper. In practice `refire:` is + a small, human-judged knob (the `once`/`rare`/`normal`/`frequent` presets), and + the telemetry that fed `tune-curves` still exists for the future ADR-134 work. +- Removing a shipped subcommand is a visible CLI surface change (documented + here and in the release notes). + +### Neutral + +- The frontmatter schema already excludes `curve:`; this change makes the code + match the schema. Any hypothetical old file still carrying `curve:` now simply + gets the UNKNOWN-field lint warning and falls back to the missing-cadence + default, rather than being silently honored. + +## Alternatives Considered + +- **Fix `tune-curves` to write `refire:`** (translate `half_life` → window + fraction). Rejected: the translation reintroduces the model-window coupling + ADR-126 removed, and it invests in a manual precursor to the deferred ADR-134 + auto-tuner rather than retiring it. +- **Leave the command, only silence the lint.** Rejected: that legitimizes a + retired field and keeps two cadence models alive; the lint is correct to reject + `curve:`. +- **Delete the command but keep the `curve:` read-fallback.** Rejected: with no + way using `curve:` and the schema rejecting it, the fallback is pure dead code — + the kind of drift the project retires on sight. diff --git a/docs/architecture/ways/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md b/docs/architecture/ways/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md new file mode 100644 index 00000000..9caea2c1 --- /dev/null +++ b/docs/architecture/ways/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md @@ -0,0 +1,87 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'prototyping and live trials: a mean-of-max confirm over-pruned multi-topic surfaces, and the share gate alone fired nothing on topic-diverse prompts (documentation/adr at peak 0.54 diluted below it)' + - precedent: ADR-156 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-05 +deciders: + - aaronsb + - claude +related: + - ADR-107 + - ADR-108 + - ADR-125 + - ADR-155 + - ADR-156 +imported: + from: docs/architecture/system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md + format: v0 + status: Accepted +--- + +# ADR-160: Chunked late-interaction matching with softmax-share gating for way selection + +## Context + +Way selection embeds a surface (a prompt, a tool-use command, a task) into one dense vector and matches it against each way's one-line alias (`description` + `vocabulary`) by cosine similarity, thresholded in probability space by the calibrated fire gate (ADR-156, `g(s) ≥ τ_s`). This single-vector approach has two structural weaknesses. Dense sentence embeddings are **anisotropic** — unrelated text still scores 0.2–0.4, a high similarity floor — so an absolute threshold separates signal from noise across a narrow band. And a **sparse or action-shaped surface** (a shell command is an action, not a statement of intent) collides on shared tokens with unrelated aliases, so a way fires on lexical overlap rather than meaning; the fire carries no recoverable reason, because a cosine between two dense vectors has no term-level attribution. + +The felt symptom is *poorly-matched* ways surfacing; the objective is **precision** (fewer poorly-matched fires), with lower fire count as the emergent byproduct rather than a directly-tuned target. The mechanism and the measured forces behind this decision were established by prototyping — see the design note *The Tool-Use Channel is a Signal Problem* (`docs/architecture/ways/ADR-191-the-tool-use-channel-is-a-signal-problem-lookbehind-chunk-spread-and-winner-confirmation.md`), which grounds each stage below in established information-retrieval practice (conversational query reformulation, ColBERT-style late interaction, score normalization against anisotropy, two-stage cascade reranking). + +## Decision + +Adopt a multi-stage evidence pipeline for way selection, replacing single-vector thresholding. The pipeline is channel-agnostic (prompt, tool-use, task) and uses only the existing embedder (ADR-108/125) and calibrated fire gate (ADR-156) — no new model, no reasoning tier, no resident daemon. + +1. **Contextualize the surface.** A sparse surface is enriched with adjacent intent context before embedding, rather than embedding the bare surface. Intent, not the literal artifact, is what should be matched. +2. **Chunk and match.** Split the surface into sub-units, embed each, and match each against the corpus — multi-vector late interaction instead of one vector per surface. +3. **Rank by peak.** A way's ranking score is the maximum over its per-chunk similarities. Peak preserves specificity; it deliberately discards how many chunks agreed. +4. **Gate by softmax-share, with a peak co-gate.** Within each chunk, take a softmax over the candidate ways (competition-normalized and zero-sum, which defeats the anisotropic floor: a way must *win* the chunk, not merely clear an absolute score). Sum this mass across chunks. A way is admitted into confirmation on **either** a sufficient summed share **or** a decisive peak cosine — because `share = Σmass / n_chunks` caps a way that owns one of N topics at ≈1/N, so on a topic-diverse surface (the common case for a real prompt) a specific single-chunk match is diluted below any fixed share gate. Admitting on the peak, and letting the strict body-confirm (stage 5) carry precision, recovers that match. +5. **Confirm the winner.** Cross-compare the winning way's body chunks against the *chunk it won* (its peak chunk) and require the best of those to clear a bar. Corroborating the winning evidence — rather than averaging over every surface chunk — still rejects single-token collisions (a collided chunk finds no support in the way's own body) and covers the softmax gate's zero-sum blind spot (it always hands the winner mass, even when nothing is truly relevant), without diluting a way that legitimately matched only part of a multi-topic surface. +6. **Exclude structurally, never with negated text.** Exclusion (scope, domain, project) is expressed as a filter or rule, because dense bi-encoders cannot represent negation — negated text in an alias moves it *toward* the negated topic. The corpus-authoring corollary: a way's embedded prose (`description` + `vocabulary`) states only what the way is *for*, in positive terms, using its own distinctive vocabulary; it never names, contrasts with, or excludes another item — all exclusion lives in the scope gate. Naming or negating another item only pulls the alias *toward* it. + +The pipeline is lenient where it ranks (peak + softmax-share admit a way on its strongest evidence) and strict where it confirms (the body must corroborate that winning evidence). An earlier design averaged the confirm over *every* surface chunk; a live trial showed that over-prunes multi-topic surfaces, so confirmation is scoped to the winning chunk. This lenient-rank / strict-confirm split is the load-bearing design choice. + +**Required primitive.** The pipeline embeds many chunks per surface, so the embedder must embed a **batch per model load**, and ideally **multiple batches per load** (all chunks across all surfaces a hook needs in one invocation). Per-chunk model reloads are not viable. This is a hard prerequisite, not an optimization. + +**The late-interaction pipeline is the semantic matcher, not an opt-in alternative.** It replaces the single-vector calibrated gate on the prompt and task surfaces; the single-vector path (ADR-156) is retained only as the **fail-safe fallback** for surfaces too sparse to chunk (fewer than two chunks) or when the embedder is unavailable — it is neither a user-selectable mode nor the default. There is deliberately no config flag to A/B the two: committing to one matcher is the non-clever choice. + +**Status: Accepted — shipped with provisional operating points.** The operating points (softmax temperature, share gate, peak co-gate, confirm gate) are hand-set. The matcher ships as *the* semantic matcher on `main`. Two refinements have landed since first ship, each from evidence the shipped instruments surfaced: + +- **Winning-chunk confirmation (stage 5).** A live trial showed a mean-of-max confirm over *every* surface chunk over-prunes multi-topic surfaces; scoping confirmation to the winning chunk resolves it. +- **Peak co-gate + confirm 0.35 (stage 4).** The read-side precision instrument and the `ways match` diagnostic — run once a `--batch`-capable embedder was actually deployed (a deployment defect had silently routed every scan to the single-vector fail-safe, so the points had never been exercised on real late-interaction surfaces) — showed the share gate alone fires *nothing* on topic-diverse prompts: a strong, specific match (e.g. `documentation/adr`, peak 0.54) is diluted below the share gate. Admitting on the peak (`PEAK_GATE = 0.50`) recovers it; that shifted the binding constraint to the confirm gate, which at 0.40 still rejected a clear true positive (body-confirm 0.363), so the confirm gate moved to 0.35 — which fired the true positives without admitting false positives on the sampled surfaces. + +The points remain **provisional, not finally calibrated**: the follow-up is fitting `PEAK_GATE` / `CONFIRM_GATE` against a judged eval set (the `introspect fires` instrument plus labeled prompts) rather than a handful of surfaces. A known gap remains: this stage-4 text specifies routing the summed mass through ADR-156's *calibrated* fire gate, but the implementation thresholds a *hand-set* `SHARE_GATE`; calibration must reconcile the two. That is refinement of a shipped matcher, not a gate on adoption; a regression is handled by tuning or, in the limit, superseding this ADR. + +## Consequences + +### Positive + +- **Precision.** Rejects the token-collision and cross-domain false positives that single-cosine admits; fewer poorly-matched fires, with lower total count as an emergent effect rather than a suppressed one. +- **Attribution.** Every fire carries a recoverable reason — the winning chunks and the confirming body spans — closing the "no recoverable term" gap in the fire drill-down. +- **Reuse.** Uses the existing embedder and calibrated fire gate; adds no model, no LLM/reasoning tier, and no resident daemon. + +### Negative + +- **A calibration surface.** The operating points must be fit against a metric, not hand-set — a real tuning burden and the explicit gate on adoption. +- **More embedding work per surface** (N chunks plus a winner-body cross-similarity), viable only with the batched-embedding primitive; without it the model-load cost multiplies. +- **Degrades on context-free sparse surfaces** (nothing to chunk); requires a fail-safe fallback to the single-vector path rather than a hard dependency. + +### Neutral + +- Requires the **batched-embedding primitive** (single- and multi-batch per model load) as a prerequisite deliverable. +- Requires **forward telemetry** (`fire_score` on *every* semantic fire, not just first-fires) plus a **read-side replay instrument** — replaying each fire's surface from the transcript at its logged token position, so relevance (and derived signals like self-reference and productivity) can be judged — to measure precision; the prerequisite for calibration. A per-fire query *hash* was considered and dropped: a hash can only be deduplicated, never judged for relevance. +- Channel-agnostic; roll-out is sequenced by measured fire-volume per channel, not by channel identity. + +## Alternatives Considered + +- **Single-vector cosine threshold (status quo, ADR-156).** The problem this ADR addresses: an anisotropic floor thresholded with no attribution. Retained only as the fail-safe fallback for surfaces too sparse to chunk — not as the default and not as an opt-in alternative to the late-interaction matcher. +- **Generative LLM reranker (local or remote).** Rejected as the primary mechanism. A probe kept the very false positive it was meant to reject when given only the thin alias as evidence; it adds cost and nondeterminism, and the leverage proved to be *evidence quality*, not model capability. Retained only as a possible last resort for residual within-domain ambiguity, behind the deterministic stages. +- **Resident model daemon.** Deferred. Unnecessary for the embedding tier — batched embedding suffices for the load. Relevant only if a larger reasoning model is later introduced as a reranker. +- **Negated alias text ("not for X").** Rejected. Dense bi-encoders move *toward* a negated topic; measured directly. Exclusion must be structural. +- **Sum / noisy-OR aggregation for ranking.** Rejected for ranking. Rewards breadth over specificity and lets a generic near-miss outrank the specific way; peak is used for ranking and mean only for winner confirmation. diff --git a/docs/architecture/ways/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md b/docs/architecture/ways/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md new file mode 100644 index 00000000..072269b4 --- /dev/null +++ b/docs/architecture/ways/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md @@ -0,0 +1,167 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: matching +basis: + - evidence: 'read-only diagnostics on session 56ebbbc1: no mid-turn queued operator message was scanned, and a lone fragment fell to the single-vector fallback while the concatenated burst fired documentation/mermaid (peak 0.477, share 0.190)' + - standard: 'Claude Code hook lifecycle: UserPromptSubmit fires once per turn, and mid-turn messages are recorded as queue-operation enqueue entries' + - precedent: ADR-160 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-05 +deciders: + - aaronsb + - claude +related: + - ADR-160 + - ADR-155 + - ADR-130 + - ADR-123 +imported: + from: docs/architecture/system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md + format: v0 + status: Accepted +--- + +# ADR-161: Queued mid-turn operator messages as an aggregated scan surface + +## Context + +The way matcher's highest-value surface is the **operator's own message**: operator +input is low-quantity and high-quality — a deliberate "let's use the mermaid way to +diagram this" carries far denser intent than any tool-use command or agent-authored +text. A relevant way firing on that message is the most common and most useful +matcher interaction. + +Live observation of the ADR-160 late-interaction matcher surfaced that this surface +is, in practice, **not matched** when the operator most naturally produces it. Two +independent causes compound, both confirmed by read-only diagnostics against a real +session transcript (session `56ebbbc1`): + +1. **Mid-turn messages never reach the scan hook.** `UserPromptSubmit` is a + *once-per-turn* lifecycle event — it fires when a prompt starts a turn, not for a + message the operator types while the agent is already working. Such a message is + **queued** and flushed into the running turn at the next LLM pause. In the + transcript it is recorded as a `type: "queue-operation"`, `operation: "enqueue"` + entry (with `.content` and a `.timestamp`), and it *never* becomes a + `type: "user"` prompt entry. `check-prompt.sh` therefore never scans it. This is + documented Claude Code behavior, not a defect in our matcher: the matcher path + (`scan/mod.rs` → `late_interaction::run`) is correct and would fire if handed the + surface. Confirmed: top-of-turn operator messages *are* scanned (e.g. + `meta/subagents` keyword-fired on a top-of-turn "subagent" message), while every + mid-turn queued message in the session produced no fire. + +2. **A lone fragment is too sparse for late interaction.** Operators type intent as a + *burst of short messages*, not one paragraph. Run through the matcher, a single + queued fragment ("as another test … we should use the mermaid style way to author + a document about the flow …") reports **"late-interaction unavailable — surface + too sparse to chunk"** and falls back to the single-vector gate (ADR-160's + <2-chunk fail-safe), where the relevant way (`documentation/mermaid`, ~0.4) does + not clear the calibrated gate. **Concatenating** the consecutive fragments of the + same burst makes it a ≥2-chunk surface, at which point late interaction runs and + fires: `documentation/mermaid` peak 0.477 / share 0.190 / confirm 0.579 and + `softwaredev/visualization/diagrams` 0.450 / 0.233 / 0.526 — both admitted on + **share**, which a single fragment cannot accumulate. + +The consequence: fixing only cause 1 (scanning queued messages) would still not fire +the way the operator expected, because cause 2 drops each lone fragment to the +single-vector fallback. The fix must address both. + +## Decision + +Add a **queued-message scan lane** that treats the burst of mid-turn operator +messages as a first-class, aggregated matcher surface — restoring operator↔agent +symmetry that ADR-160 already intends (the matcher is channel-agnostic; only the +delivery of this surface was missing). + +1. **Hook: `PostToolUse`.** It fires at the LLM pauses where queued messages flush, + and it runs *before the agent's next action* — so a way admitted here (e.g. the + mermaid way) can still steer the work the operator asked for, not merely annotate + it after the fact. (`Stop` would be too late to steer the current turn.) This + rides the existing `check-post.sh` PostToolUse dispatcher (ADR-123 Decision 5). + +2. **Read the transcript, select queued operator text.** Find + `type: "queue-operation"`, `operation: "enqueue"` entries with `timestamp` newer + than a per-session **scan mark**. This is a precise, stable selector — genuine + operator text only, never tool-result envelopes or system injections. + +3. **Aggregate the burst into one surface** before matching, so short high-quality + fragments accumulate into a ≥2-chunk late-interaction surface instead of each + falling to the single-vector fallback. (Aggregation boundary — see the decision + point below.) + +4. **Match and inject through the existing engine.** Run the aggregated surface + through the same `scan` path (ADR-160 late interaction, ADR-130 salience reducer, + ADR-155 keyword/semantic lanes), fire via the same `ways show way` gate (so the + refractory/refire rules apply — the lane cannot spam), and emit fired content as + `PostToolUse` `additionalContext`. + +5. **Advance the scan mark** to the newest consumed `timestamp`, so each queued + message is matched at most once. + +### The aggregation boundary + +How to group queued fragments into a surface is the load-bearing choice; the spike +proves aggregation is *required*. The decision is **all-pending-since-mark, scanned +at each PostToolUse**: concatenate every queued fragment newer than the scan mark +into one surface, match, then advance the mark. It is the simplest policy and adds no +new machinery — it reuses the existing flow, and leans on the ADR-155 refire gate to +absorb the one failure mode (a burst still being typed fires on a partial surface, +and a later fragment re-scans an overlapping one; the near-duplicate re-fire is +collapsed by refractory). Revisit against telemetry only if partial-burst noise +proves real. The two richer alternatives — settle-then-scan and a rolling window — +are recorded under Alternatives Considered. + +## Consequences + +### Positive + +- **The prime surface is matched.** Operator intent — the highest-value, lowest-noise + input — becomes a first-class matcher surface however the operator types it. +- **Symmetry realized.** Operator messages get the same late-interaction treatment as + agent actions, which ADR-160 already intends channel-wise. +- **No matcher change.** Reuses the ADR-160 engine, ADR-130 reducer, ADR-155 lanes, + and the ADR-123 PostToolUse dispatcher; the new code is transcript selection + + aggregation + a scan mark. + +### Negative + +- **Transcript reads on PostToolUse.** Adds a bounded transcript scan per tool pause; + cost must stay negligible (tail-read from the mark, not a full re-parse). +- **A timeliness/completeness tradeoff** with no free optimum (the aggregation + boundary above) — a real tuning surface, like ADR-160's operating points. +- **Coupling to a transcript shape.** `queue-operation` is a Claude Code + implementation detail; if its schema changes the selector must follow. Isolate the + selector so the blast radius is one function. + +### Neutral + +- Establishes a general seam for *event-shaped* operator input (a future external + message source could feed the same aggregated-surface lane). +- Interacts with ADR-160's <2-chunk fallback: aggregation is precisely what lifts a + fragmented operator burst out of the fallback into late interaction. + +## Alternatives Considered + +- **Scan queued messages without aggregation.** Rejected: the spike shows a lone + fragment falls to the single-vector fallback and does not fire the relevant way — + it would ship a fix that still misses the operator's actual input. +- **`Stop`-hook (end-of-turn) scan.** Rejected as the primary lane: it cannot steer + the current turn (the way fires after the agent has acted). Viable only as a + backstop for intent that arrived too late to act on. +- **Do nothing / file upstream only.** Rejected: waiting on a harness change leaves + the matcher's prime surface dark indefinitely; the transcript already carries the + data to close the gap in-repo. +- **Concatenate operator text into the *next* `UserPromptSubmit` query.** Rejected: + defers matching to the next idle turn — far too late to steer, and conflates + distinct turns. +- **Settle-then-scan aggregation** (wait for a quiet window before scanning the whole + burst). Rejected as the default: more complete intent, but it fires after the agent + may have already acted on the request — it trades the steer for completeness. A + fallback worth reconsidering if partial-burst noise proves real under (A). +- **Rolling-window aggregation** (last *k* fragments / *T* seconds regardless of burst + boundaries). Rejected: robust to fragmentation but blends unrelated intents into one + surface, and the scan mark already bounds the window to genuinely-unscanned text. diff --git a/docs/architecture/ways/ADR-166-single-source-of-truth-for-model-context-window-resolution.md b/docs/architecture/ways/ADR-166-single-source-of-truth-for-model-context-window-resolution.md new file mode 100644 index 00000000..5fe3fd51 --- /dev/null +++ b/docs/architecture/ways/ADR-166-single-source-of-truth-for-model-context-window-resolution.md @@ -0,0 +1,203 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - cli + - attend +basis: + - evidence: three context-window resolvers disagreed; a claude-fable-5 session at 212,899 tokens showed tokens_total 200000 and pct_used 106, and ~78,000 local model records never carry the [1m] suffix + - precedent: ADR-126 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-12 +deciders: + - aaronsb + - claude +related: + - ADR-126 + - ADR-151 + - ADR-153 +imported: + from: docs/architecture/system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md + format: v0 + status: Accepted +--- + +# ADR-166: Single source of truth for model context-window resolution + +## Context + +The context window is the denominator of nearly everything the toolchain reports +or decides. `ways context` divides by it to render the usage gauge. ADR-126 makes +way refire *window-relative*: a way's half-life is expressed as a fraction of the +window, so a wrong window rescales the entire disclosure curve. `sensor-peers` +divides by it to show peer session pressure. + +Three separate implementations answered that one question, and they disagreed: + +| Site | Rule | +|---|---| +| `ways-cli/src/cmd/context.rs:239` (`model_to_window`) | `opus-4` → 1M; `sonnet`\|`haiku` → 200K; else env override, else 200K | +| `ways-cli/src/session.rs:322` (`context_window_from_transcript`) | `opus-4` → 1M; else env override, else 200K | +| `sensor-peers/src/lib.rs:567` | `[1m]`\|`opus-4`\|`sonnet-4` → 1M; `-` → 0; else 200K | + +A fourth site, `ways-cli/src/cmd/show/mod.rs:88`, hardcodes `.unwrap_or(200_000)` +on the failure path. + +Each is a substring allowlist written against the model lineup of its day, and +each has since gone stale. The lineup they encode no longer matches the models in +use: + +- `claude-fable-5` matches no branch in any of them and falls to the 200K default. + It has a 1M window. A session observed 212,899 tokens with no forced compaction + while `ways context` reported `tokens_total: 200000` and `pct_used: 106`. +- `claude-sonnet-5` has a 1M window; `context.rs` classifies it 200K on the + `sonnet` substring. +- `sonnet-4` resolves to 1M in `sensor-peers` and 200K in `context.rs`. Both + cannot be right. +- Conversely, `opus-4` is matched as a bare substring, so any `opus-4*` id is + called 1M whether or not that is true of the specific model. + +Two structural defects compound the staleness: + +**The `[1m]` suffix is never present.** `sensor-peers` tests for it, but Claude +Code writes the bare model id to the transcript (`claude-opus-4-8`, not +`claude-opus-4-8[1m]`) — confirmed across ~78,000 model records in local +transcript history. That branch has never matched. Detection therefore cannot +observe the harness's window setting at all; it can only observe the model id. + +**The documented override does not work where it is most needed.** In both +`context.rs` and `session.rs`, `CLAUDE_CONTEXT_WINDOW` is read only from the +fallback arm. Any model that matches a substring branch — every Sonnet and Haiku +in `context.rs` — ignores the override entirely, contradicting +`skills/context-status/SKILL.md`, which presents it as the general escape hatch. + +The failure is silent by construction. A resolved window carries no indication of +whether it was detected or defaulted, so a wrong denominator is indistinguishable +from a right one at every consumer. + +## Decision + +Resolution of the context window becomes a single function in `ways-core` +(`ways_core::context_window`), and every site calls it. No consumer computes a +window itself. + +**Resolution order**, applied in this order, first match wins: + +1. `CLAUDE_CONTEXT_WINDOW`, if set and parseable — unconditionally, ahead of all + detection. Detection cannot see the harness's active window (see Context), so + the operator override must always outrank the model table, never merely + backstop it. +2. An explicit model table, matched against the full model id rather than by + loose substring. +3. A conservative 200K default. + +**The result carries its provenance.** The resolver returns the window together +with a `WindowSource` (`EnvOverride` / `ModelTable` / `Default`), and +`ways context --json` emits it as `window_source`. A default is thereby reported +as a default rather than presented as a detection. This is the property that makes +the next stale-table failure observable instead of silent: the Fable session above +would have read `window_source: "default"` at the moment it was wrong. + +**The table is explicit and enumerated**, not a substring heuristic. Substring +matching is what failed: `sonnet` swallowing `sonnet-5`, `opus-4` swallowing every +Opus 4.x regardless of window. Unknown models fall to the default and say so, +which is a correctable, visible state — unlike a wrong match, which is not. + +Ids are matched as a **boundary-delimited component** of the model string rather +than anchored at its start, because other harnesses wrap the same id in provider +prefixes and version suffixes (`us.anthropic.claude-opus-4-8-v1:0`, +`claude-opus-4-8@20260115`) and this repo supports those deployments. A rule +anchored at byte 0 would regress every Bedrock and Vertex session to the default. +The boundary requirement is what keeps this from degenerating back into substring +matching: `claude-sonnet-5` is not found inside `claude-sonnet-55`, because the +trailing `5` is alphanumeric and therefore a different id, not a qualified form of +this one. Bare family aliases (`opus`, `sonnet`) are matched on **exact equality +only** — an alias is a whole model reference, not a family stem, and prefix-matching +one would resolve `claude-sonnet-4-5` to the current Sonnet's window and report it +as a confident detection. + +**Sentinels are absences, not unknown models.** Claude Code writes +`"model": "<synthetic>"` for interrupt and API-error turns; `sensor-peers` uses `-` +as its no-model placeholder. A transcript whose newest assistant turn is an +interrupt still has a real model behind it, so the scanners skip sentinel turns and +keep walking back rather than resolving the sentinel to the default. Nine +transcripts in local history end on a `<synthetic>` turn; under a naive scan each +would have handed a live 1M session a 200K window. + +**A peer's window is resolved without the operator's override.** `sensor-peers` +reads *other* sessions' transcripts, and `CLAUDE_CONTEXT_WINDOW` states the window +of the process that set it. Applying it to a peer would compute that peer's fill +against the observer's window — an operator with the override at 1M would see a +Haiku peer at 190K/200K, genuinely about to compact, rendered as 19% full. Foreign +sessions therefore resolve through the model table alone +(`resolve_for_foreign_session`). + +The table is a hardcoded enumeration rather than a live Models API lookup +(`GET /v1/models/{id}` exposes `max_input_tokens`). The resolver runs in +`UserPromptSubmit` hooks on every turn; it must be synchronous, offline, and +credential-free. A network call on that path is not acceptable, and a cache of a +network call reintroduces the staleness this ADR exists to remove, with added +failure modes. The table is therefore accepted as a maintenance obligation at +model launch — made tractable by the fact that there is now exactly one of them. + +## Consequences + +### Positive + +- One place to update when a model ships. The present bug required four edits in + three crates to fix correctly, which is why it was never fixed at all. +- Way refire dynamics (ADR-126) are correctly scaled on every model. On Fable 5 the + half-life had been computed against a 200K window inside a 1M one, compressing + the disclosure curve by 5x. +- A wrong window becomes visible at the point of use via `window_source`. +- `CLAUDE_CONTEXT_WINDOW` behaves as documented, on every model. + +### Negative + +- The table must be updated when a model launches or a window changes. This is a + real recurring obligation; nothing about the design removes it. It is bounded to + one function and covered by tests that pin each known model id. +- An unknown model still resolves to 200K, which will be wrong for any future 1M + model until the table is updated. It is reported as `window_source: "default"`, + making it diagnosable, but a diagnosable wrong answer is still a wrong answer. + +### Neutral + +- `sensor-peers` takes a dependency on `ways-core`. Both are already workspace + members, so this is a manifest line, not a structural change. +- The `[1m]` suffix test in `sensor-peers` was dead code and is removed. The + marker appears in the *system prompt text* (and so, as prose, inside transcript + message content — 414 occurrences locally), but never as a `message.model` value: + across ~78,000 model records the field is always the bare id. Component matching + nonetheless tolerates a `claude-opus-4-8[1m]` id, so if Claude Code ever does + begin writing the marker, it resolves rather than silently defaulting. +- Models absent from the table now resolve to the default where the old `opus-4` + substring gave them 1M — `claude-opus-4-5`, `claude-opus-4-1`, `claude-opus-4-0`. + This is a correction, not a regression: the 1M window arrived with the 4.6 + generation, so calling Opus 4.1 a 1M model was exactly the over-broad match this + ADR removes. None appear in local transcript history. They can be pinned + explicitly once their true windows are confirmed. + +## Alternatives Considered + +- **Fix the four sites in place, keep them separate.** Rejected: it repairs this + instance and preserves the mechanism that produced it. Three resolvers already + drifted into three different answers for `sonnet-4`; nothing prevents a fourth + divergence at the next model launch. +- **Resolve from the Models API at runtime.** Rejected: the resolver is on the + per-turn hook path and must be synchronous, offline, and credential-free. See + Decision. +- **Fetch the table from the Models API at build time.** Rejected for now: it + moves the staleness from source to release cadence without removing it, and + couples the build to network and credentials. Reconsider if the table proves to + churn faster than releases. +- **Default unknown models to 1M rather than 200K.** Rejected: it is right for + the current lineup but fails unsafely. Over-reporting the window suppresses way + disclosure and under-reports usage — the gauge reads comfortable while the + session is in fact near its limit. Under-reporting is the conservative error, + and `window_source` makes it visible rather than silent. diff --git a/docs/architecture/ways/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md b/docs/architecture/ways/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md new file mode 100644 index 00000000..4665aa86 --- /dev/null +++ b/docs/architecture/ways/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md @@ -0,0 +1,163 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - disclosure + - method +basis: + - evidence: an 11,500-word draft ran 3.4 significance clauses per thousand words against 0.5 in reviewed prose, with eleven banned antitheses and 118 em-dashes, after core.md and the writing way had both fired + - evidence: scan/state.rs re-shows core only when the transcript since summary is under 5000 bytes, so core never re-discloses on distance + - evidence: 'research: Bohr (arXiv:2511.13972) on expansion discipline; IFEval-style negative constraints fail at 22-30% on frontier models' + - precedent: ADR-123 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2026-07-30 +deciders: + - aaronsb + - claude +related: + - ADR-123 +imported: + from: docs/architecture/system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md + format: v0 + status: Accepted +--- + +# ADR-174: Progressive core — decoration guidance and the core re-disclosure gap + +## Context + +Two findings arrived together in one session. The second is the architectural one. + +### Claude's prose decorates, and the existing surfaces did not stop it + +An 11,500-word document drafted across a dozen turns was measured against the patterns that mark prose written to be admired rather than read. Significance clauses — a clause whose only job is telling the reader that the previous clause mattered — ran at **3.4 per thousand words** against **0.5** in prose that had already been reviewed. The document also carried **eleven** instances of the antithesis construction that `core.md` explicitly bans, and **118 em-dashes**, one per 98 words. + +Both governing surfaces had fired. `core.md` fired at session start and contains zero instances of the construction it bans. The writing way fired at epoch 4 carrying "use em dashes sparingly." The document was written at epochs 45–57. + +Two mechanisms explain the gap, and the supporting literature is consistent with both. + +**Detection, not compliance, is the failure.** Models score poorly on noticing their own negative-constraint violations; IFEval-style negative constraints fail at 22–30% on frontier models, and constraint-verification work finds low negative F1 across the board. A rule restated more forcefully does not help when the writer cannot see the violation while producing it. + +**Style rules decay across a long draft.** Bohr (arXiv:2511.13972) separates *initial control* from *expansion discipline* — whether a style survives a revision turn — and finds instruction-plus-example strongest on both, example-only carrying no expansion discipline at all. The observed pattern matches: the rule held while output was short and failed across an essay. + +A rule containing an adverb compounds this. "Sparingly" has no threshold, so there is no moment at which compliance can be tested. "Cut any clause that explains why the previous clause matters" is a search that can actually be run. + +### `core.md` is structurally excluded from re-disclosure + +Investigating where the guidance should live surfaced the larger issue. + +`tools/ways-cli/src/cmd/scan/state.rs` gates core on a boolean marker rather than a decay curve: + +```rust +if !session::core_is_shown(session_id) { // first time — show it +} else if let Some(tp) = transcript { + let ctx_size = transcript_size_since_summary(tp); + if ctx_size < 5000 && age > 30 { // context was cleared — re-show +``` + +Core re-injects on one condition: the transcript since last summary is under 5000 bytes, which is a safety net for a context clear. In a long growing session `ctx_size` passes 5000 within a couple of turns and never returns, so core lands at turn 1 and is never refreshed until compaction. + +Three consequences follow. + +`refire: 0.15` in `core.md`'s frontmatter is inert on this path. Core is gated by `stamp_core` / `core_is_shown`, never enters the firing ledger, and does not appear in `ways list` — 70 entries in the observed session, none of them core. + +The retention profile is inverted relative to the rest of the corpus. Every matched way gets a curve that *lowers* its suppression threshold as distance grows, so it becomes more eligible to re-fire the further the session runs. Core gets a gate that only opens when context is small. The file that applies to every turn has the weakest retention in the system. + +This is not a defect in the safety net, which does the job it was written for. It is a gap: no path was ever built for core to re-disclose on distance, because core predates the firing-dynamics work in ADR-123. + +> **Amendment (2026-08-20).** The transcript-size safety net quoted above is gone (PR #454). It mis-fired on the first prompt of any fresh session whose operator paused more than 30 seconds before typing: a new transcript is under 5000 bytes, so core was cleared and shown a second time. The `startup`, `compact`, and `clear` SessionStart matchers already run `clear-markers.sh`, which removes the core marker, so a missing marker is now the only condition under which `scan state` shows core. The gap this ADR addresses — no re-disclosure on distance — is unchanged by that removal. + +Placing new always-relevant guidance in `core.md` would therefore place it in the one location with no re-disclosure at all. + +## Decision + +Adopt **progressive disclosure for core content**, using the existing parent/child pattern rather than a new mechanism, and deliver the decoration guidance across three tiers. + +**Tier 1 — `core.md`.** Two bullets under Posture, sibling to the existing reasoning-tic rules, stating the rule in its shortest checkable form. Turn 1 only. This tier exists because posture shapes conversational output, which no artifact-boundary check can reach. + +**Tier 2 — `meta/trust/prose/prose.md`.** A fourth child alongside `autonomy`, `delegation`, and `voice`, carrying the expanded account with paired before/after examples. Semantic and vocabulary triggers, `refire: 0.15`, and the parent boost when `trust` fires. This tier re-discloses on distance, which is what core cannot do. + +**Tier 3 — `documentation/markdown/density/`.** A postcheck way, sibling to `documentation/markdown/reflow`, firing on the markdown just written. Its macro reports measured counts for that file rather than restating the rule. This tier exists because the failure is detection, and only a count closes that gap. + +Firing thresholds: significance clauses at ≥3 per thousand words, or em-dashes at ≥15 per thousand, over a 150-word floor, with per-file suppression for the session. Calibrated against measurements in this repository rather than chosen. Loose deliberately — a surface that nags trains its reader to ignore it. + +Two supporting changes. The writing way loses "use em dashes sparingly" as unmeasurable and gains a pointer to `trust/prose`. Em-dash count is **reported and not banned**: repository prose runs at 15.3 per thousand words against the draft's 10.7, so the punctuation is house style at volume rather than an anomaly, and only density is the tic. + +This ADR records the core re-disclosure gap as a **finding, not a fix**. Giving core a distance-based re-disclosure path is a separate decision with its own cost — core is roughly 900 words, and re-injecting it on a curve risks exactly the nagging the threshold discipline above avoids. The progressive-core pattern routes around the gap without deciding it. + +## First test: null result, and a trigger defect + +Tested 2026-07-30, same day, immediately after deploy. Two hosts, same two prompts, same model, fresh session each — a Kubernetes adoption assessment followed by a revision turn asking to expand one section. Arm A ran ways 1.6.0 without these changes; arm B ran 1.7.0 with them. + +| | words | significance | per 1k | em-dash | per 1k | antithesis | +|---|---|---|---|---|---|---| +| Arm A (control) | 4,477 | 4 | 0.9 | 46 | 10.3 | 0 | +| Arm B (treatment) | 4,374 | 4 | 0.9 | 46 | 10.5 | 0 | + +Identical to one decimal, and identical in raw counts. The change produced no measurable difference. + +Three readings, in order of importance. + +**Tier 2 never fired, and could not have.** `ways list` on the treatment session showed one way triggered, and it was not this one. Measuring against the arm-B prompt with the EN calibration (`a=19.02`, `b=-6.05`, so `τ_s=0.5` is cosine 0.318 and `τ_k=0.15` is cosine 0.227): + +| alias | cosine | `g(s)` | outcome | +|---|---|---|---| +| as shipped in 1.7.0 | 0.188 | 0.078 | below `τ_k` — **gated out of both lanes** | +| after vocabulary rewrite | 0.256 | 0.233 | clears `τ_k`; keyword lane only | + +The original way was unfirable on that prompt. The keyword lane is floor-gated, so a pattern hit would have been suppressed even if a pattern had existed. The null result was not the guidance failing to change behavior — the guidance never arrived. + +The deeper finding is structural. Measured across long-form prompts on varied topics, `g(s)` runs 0.02–0.24: a database migration analysis scores 0.022, an authentication write-up 0.039, the arm-B expansion turn 0.137. A prompt about writing scores 0.906. **The topic dominates the embedding**, so a way about how prose reads cannot reach a semantic bar on prompts about Kubernetes or MySQL. No vocabulary tuning bridges that distance. + +Two consequences. The way now carries `pattern_strict: true`, because the floor gate would otherwise suppress it on nearly every real long-form request. And the pattern had to be narrowed hard: a first attempt including expansion verbs (`expand the`, `go deeper`, `more detail`, `elaborate on`) produced 8 false positives out of 8 realistic coding prompts, since those are ordinary English in engineering chat. The shipped pattern is noun-gated — a depth adjective followed by an explicit document type — measuring 5/5 recall and 8/8 precision on a hand-built battery. + +**Tier 2 therefore fires only on an explicit long-form request, and not on the expansion turn.** The expansion turn is where expansion discipline fails, so the tier misses its highest-value moment. Closing that needs a trigger keyed to output volume, which no current trigger type provides for conversational output. + +A subsequent run confirmed the consequence: the way fired at epoch 1 and showed no re-disclosure across the following 25 epochs. Re-disclosure requires a re-match, and the strict pattern does not match ordinary continuation turns. Relaxing the `refire` cadence alone cannot close this, because the cadence gates a re-match that never arrives. + +The engine also forbids carrying both lanes on one file — `scan/mod.rs` skips state-triggered ways from the prompt, pattern, and semantic surfaces, on the stated grounds that "their trigger is a condition, not a topic." So tier 2 is split rather than widened: `meta/trust/prose` keeps the keyword lane for the first fire, and a sibling `meta/trust/prose/sustain` carries a ~70-token condensed form on `trigger: context-threshold` with a short `refire`, re-disclosing as context grows. Whether periodic re-disclosure changes the observed drift is under test and not yet answered here. + +**The condition could not discriminate.** 4,400 words in a fresh session is the short-output regime where the rules already held. `core.md` banned the antithesis construction before any of this, and both arms show zero. The failure this ADR addresses appeared at 11,500 words across a dozen turns. + +**Both arms were already at target.** 0.9 per 1k sits at the repository baseline. There was no decoration to remove, which means decoration is not a general property of the model's prose — it is specific to long multi-turn drafting. + +The near-zero between-arm variance is the one encouraging signal: two independent runs on different hosts produced identical counts, so the measurement has power. A real effect would show. The instrument is sound and was pointed at the wrong condition. + +Status of the central claim after this test: **unfalsified, not validated.** Nothing here supports asserting the change works. + +## Consequences + +### Positive + +- Always-relevant guidance gains a re-disclosing home without changing the core delivery contract. +- The measurement tier reports numbers rather than intentions, addressing the detection failure directly. +- Thresholds derive from measurements in this repository, so they can be re-derived and argued with. +- The core re-disclosure gap is now written down rather than resident in one session's context. + +### Negative + +- Three tiers to keep coherent. Guidance that drifts between them will contradict itself. +- The postcheck runs on every `Write`/`Edit`, adding a check to a hot path. +- Regex detection of a rhetorical pattern carries false positives. A document *about* these patterns scores high for legitimate reasons; `density.md` names this case and the path self-exclusion covers the corpus. +- The bare `, not X` form is unchecked. It over-fired on legitimate contrast in `core.md`, so narrowing it removed a real detection — the operator caught one by eye that the check misses. + +### Neutral + +- Core's `refire: 0.15` remains inert until the re-disclosure gap is separately decided. Leaving a field that does nothing is its own small debt. +- The prose linter prototype used to calibrate these thresholds is not shipped. Promoting it to `doclint` or a `ways` subcommand is deferred until the postcheck proves too easy to ignore. + +## Alternatives Considered + +**Put everything in `core.md`.** Rejected on the finding above: core has no re-disclosure path, so the guidance most needing to survive to turn 50 would land where it survives worst. + +**Put everything in the writing way.** Rejected because that way fired at epoch 4 on a semantic mass-match rather than on writing, and never returned. Predictive matching picked the wrong moment, which is a routing failure the tier-3 reactive path avoids by construction. + +**Ship a lint gate at the commit boundary instead of a way.** Deferred rather than rejected. The postcheck teaches during the work and is reversible; a commit gate is deterministic but arrives after the session has moved on. Revisit if the way proves ignorable. + +**Give core a distance-based re-disclosure curve.** Deferred as a separate decision. It is the direct fix for the gap and it re-injects ~900 words per fire, which needs its own cost analysis and threshold work. + +**Ban em-dashes outright.** Rejected on measurement. Repository prose runs higher than the draft that prompted this, and `core.md` is the densest file sampled at 20.4 per thousand. A check that flags every file gates nothing. The operator's personal prose guidance does ban them; that is a register decision for personal correspondence and deliberately not inherited. diff --git a/docs/design-notes/adopter-localization-lifecycle-and-tuning.md b/docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md similarity index 98% rename from docs/design-notes/adopter-localization-lifecycle-and-tuning.md rename to docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md index 6ed8c2a5..7bd22bae 100644 --- a/docs/design-notes/adopter-localization-lifecycle-and-tuning.md +++ b/docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md @@ -1,7 +1,16 @@ -## Single-Language Localization Tuning — the English anchor as a peer +--- +contract: adr/v1 +kind: evidence +capability: disclosure +status: accepted +date: 2026-06-22 +deciders: + - aaronsb +related: [] +--- + +# ADR-183: Single-language localization tuning: the English anchor as a peer -> **Type:** Design note (not an ADR) -> **Status:** Working draft — core principle settled (English root as source of truth) > **Cites:** ADR-139 (adopter-run localization), ADR-125 (coordinate-alias model) > **Motivates:** `ways tune --lang` implementation; the `ways-localize` skill's acceptance gate diff --git a/docs/architecture/ways/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md b/docs/architecture/ways/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md new file mode 100644 index 00000000..d4bd5674 --- /dev/null +++ b/docs/architecture/ways/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md @@ -0,0 +1,146 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: + - matching + - disclosure +supersedes: ADR-155#4 +basis: + - evidence: 'issue #528: 3,132 PreToolUse way outputs across 115 transcripts, zero delivered; PostToolUse paired 145 to 144' + - evidence: 'anthropics/claude-code#19432, closed as not planned on 2026-02-28: PreToolUse additionalContext is logged and never injected' + - evidence: 'spike-branch probe (commit e2e53243): a next-prompt stash delivers, with four findings against it' + - precedent: ADR-155 +agent: + name: Claude + model: unrecorded +status: proposed +date: 2026-09-19 +deciders: + - aaronsb + - claude +related: + - ADR-123 + - ADR-125 + - ADR-126 + - ADR-155 + - ADR-160 + - ADR-161 + - ADR-172 + - ADR-181 + - ADR-303 +imported: + from: docs/architecture/system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md + format: v0 + status: Proposed +--- + +# ADR-188: PostToolUse delivery for tool-lane ways and retirement of the semantic Bash surface + +## Context + +Three lanes match ways against tool calls. The Bash lane (`check-bash-pre.sh`, `ways scan command`) runs the `commands:` regex against the command text, the `pattern:` regex against the tool description, and the semantic lane that ADR-155 §4 added, which embeds the command, its description, and the assistant's prose since the last human turn (ADR-160). The Edit and Write lane (`check-file-pre.sh`, `ways scan file`) runs the `files:` regex against the path. Both lanes also score `.check.md` checks. All of it runs on PreToolUse and prints its matches as `{"decision":"approve","additionalContext":...}`. + +Issue #528 measured what happens to that output. Across 115 transcripts on one workstation running Claude Code 2.1.x, 17 August to 19 September 2026, Claude Code recorded every PreToolUse hook output carrying a way as a `hook_success` attachment and created no `hook_additional_context` attachment for any of them: 2,922 rows on PreToolUse:Bash, 107 on PreToolUse:Edit, 103 on PreToolUse:Write, 3,132 in total, zero delivered. The same transcripts show PostToolUse output from `check-post.sh` paired one to one: 145 `hook_success` rows carrying a way, 144 `hook_additional_context` rows. The audit rated those PostToolUse fires as followed, in the Debugging and Sub-Agents tool-triggered episodes. UserPromptSubmit output is delivered. + +Upstream anthropics/claude-code#19432 reported the same behaviour for the documented shape: a PreToolUse hook emitting `hookSpecificOutput.additionalContext` has its value logged and never injected, while `permissionDecision` and `permissionDecisionReason` on the same output work. Anthropic closed that report on 2026-02-28 as not planned. The binary's own emitter uses a top-level `decision` and `additionalContext` pair; the current hooks reference documents no such PreToolUse field, and that reading is this project's, taken from the reference. The transcripts show both shapes undelivered: the issue's logs for the documented one, the 3,132 rows for ours. + +The engine's events log counts the same month per way: 6,905 `semantic:bash:en` fires and 496 `bash` (commands regex) fires, 73% of all `way_fired` events. In the session that filed #528, the log shows 65 Bash-lane fires and the model saw none of them. Every `commands:` trigger in the corpus (`git commit`, `gh pr create`, `gh issue`, `oh-my-posh`), every `files:` trigger, and every check scored on these two lanes was silent for the whole measured month. + +Two consequences follow. + +**The tool lanes suppress the prompt lane.** `show::way_scored` records the fire (`session::record_way_fire`) before it returns the content the caller prints. A silent Bash-lane fire stamps engagement salience to 1.0 for that way (ADR-123). A prompt-lane match on the same way inside the refire half-life (ADR-126) then returns `Suppressed`. The lane the model cannot see mutes the lane it can. + +**The semantic Bash surface ran a natural experiment.** For a month it fired on tokenised command text at the volume above and nobody noticed the absence of its output. Its samples in the log are off-topic at high scores: `workstation/shell/prompt` at 0.91 on `head -40 slides decks cohort-01-intro-framing.html`, `ea/email` at 0.86 on an `ssh` listing, `softwaredev/delivery/commits` at 0.90 on a `git log` pipeline. Per Bash call that fired, the median was 1 way, 20% of calls fired 4 or more, and the maximum was 30. ADR-155 §4 added the lane on the premise that the tool description is Claude's statement of intent at act time and that fire and near-miss telemetry would monitor the surface for noise. Delivery was never verified. + +Two pieces of work depend on this decision. #525 derives write targets from Bash command text so that `files:` ways fire in auto mode, where edits run through `sed -i`, `tee`, and redirects; it inherits whatever delivery the Bash lane has. The transcript audit's recommendation to prefer tool triggers over prompt triggers rests on the 144 PostToolUse postcheck deliveries, which is a different mechanism from the PreToolUse lanes the recommendation named. + +PreToolUse has one channel measured to reach the model. ADR-181 names it: a guard exits 2 with the reason on stderr, and the model receives the reason as the tool result. A guard carries a refusal and no guidance. + +The Task lane already solved a version of this problem. `check-task-pre.sh` writes matched way ids to a stash under the session directory and `inject-subagent.sh` drains it on SubagentStart, building the `hookSpecificOutput.additionalContext` envelope in jq and logging `way_fired` at that moment. The drain does not stamp engagement. That stash exists because a subagent has no UserPromptSubmit of its own. + +### The next-prompt stash probe + +The first candidate carrier for tool-lane matches was that same pattern pointed at the main session: stash on PreToolUse, drain into the next UserPromptSubmit output. A probe on a spike branch (worktree `worktree-agent-aa3ba24f71bb47ead`, commit `e2e53243`, kept out of merge) built it: a stash per session and agent at `<sessions_root>/<session>/pretool-stash.<agent>.jsonl`, drained deliver-once into the next prompt, with `way_fired`, engagement, and the marker each recorded once, 288 tests passing. The carrier works. The probe returned four findings that this decision carries: + +1. Recording happens at match time in `show::way_scored`, before delivery. A stash that never drains has still consumed the way's first fire and started its refire clock. +2. Subagents never receive UserPromptSubmit, so their stashes are orphaned. They do receive PostToolUse. +3. One-turn lag turns pre-action guidance into post-action guidance, and checks lose their pre-flight meaning under a next-prompt drain. +4. One `git commit` with the description "commit the change" produced 10.4k characters of stashed context, because the semantic Bash lane fired four neighbouring ways on the description. That volume becomes a visible prompt tax the moment any carrier delivers it. + +## Decision + +**Tool-lane matching moves to PostToolUse and rides the envelope `check-post.sh` already prints. The semantic lane on the Bash surface retires. Engagement is stamped at delivery. The PreToolUse emitter is deleted.** + +1. **`commands:`, `pattern:` on the description, `files:`, and checks run on PostToolUse, inside `check-post.sh`.** That hook receives `Edit|Write|Bash|Task` with `tool_name`, `tool_input`, and `tool_response` on stdin and prints one `hookSpecificOutput` envelope built in jq. It gains a dispatch: for Bash it runs `ways scan command` with `tool_input.command` and `tool_input.description`; for Edit and Write it runs `ways scan file` with `tool_input.file_path`. On this path the scan commands print bare way content with no JSON; `check-post.sh` captures that text, appends it to the same `CONTEXT` the postcheck fires accumulate into, and prints the single envelope it prints today. One hook command's stdout carries one JSON document. The matchers are unchanged. The envelope and hook are the ones that carried the 144 postcheck deliveries and that ADR-161 chose for the queued-message lane. Delivery is same-turn, one tool call after the match, and it reaches subagents, which receive PostToolUse with `agent_id` set. `check-bash-pre.sh` and `check-file-pre.sh` come out of `settings.json`. `strip-session-link-pre.sh` stays on PreToolUse:Bash as a guard. + +2. **The dispatch runs on PostToolUse only.** `check-post.sh` is also wired to PostToolUseFailure, whose stdin carries `error` in place of `tool_response`. A `commands:` match on a failed call would stamp engagement and suppress the way on the successful retry, so the dispatch reads `hook_event_name` from stdin and runs only when it is `PostToolUse`. Postcheck dispatch on failure is unchanged. + +3. **The semantic lane on the Bash surface retires.** `scan command` stops scoring ways by embedding, and the `semantic:bash:en` and `semantic:bash:multi` channels are removed. This point supersedes ADR-155 §4, declared in frontmatter on both documents per ADR-303 part C (`supersedes: ADR-155#4` here, `superseded_by: ADR-188#3` on ADR-155). The embed pass remains for checks that declare `description` and `vocabulary`, since `check_semantic_score` reads it, and the lookbehind (ADR-160 stage 1) continues to feed that pass. Reasoning that precedes an action still reaches the semantic lane on the next prompt through the response-context channel (ADR-155 §3), where the user's own message anchors relevance. + +4. **Engagement is stamped at delivery.** A lane records a fire, stamps engagement, and logs `way_fired` when the body is emitted on a hook event the harness has been measured to deliver. Today the binary's `emit_hook_context` serves `scan::prompt` (UserPromptSubmit) and `scan::state` (SessionStart); `check-post.sh` and `inject-subagent.sh` build their PostToolUse and SubagentStart envelopes in jq. Under point 1 the match and the emit share one hook invocation on a delivering event, so `way_scored`'s stamp-before-return is a delivery-time stamp. Any lane that separates match from delivery stamps when it drains, or its `way_fired` event carries both a `matched_at` and a `delivered_at` tick. The Task lane is the one such lane today and its drain logs without stamping; point 6 lists the fix. The inline `decision`-bearing emitter in `scan::command` and `scan::file` is deleted, so no path in the binary can print a way body on PreToolUse. A caller that wants scores without a disclosure uses the scoring path with recording off; `ways introspect` reads recorded events and stamps nothing. + +5. **#525 builds on the PostToolUse Bash lane.** Derived write targets from a Bash command feed the same `files:` matcher the Edit and Write lane uses, on the same hook, with trigger `bash-write`. `tool_response` is available there; whether to skip targets when the command failed is #525's call. + +6. **Implementation items.** Each ships under this ADR: + - A bare-content output path for `scan command` and `scan file`, used by the `check-post.sh` dispatch. + - The `hook_event_name` gate from point 2. + - `inject-subagent.sh` stamps engagement through the binary when it emits, so the Task stash also records at delivery. + - Lazy embedding on the Bash and file paths: the embed pass and the lookbehind run only when a check in scope declares `description` and `vocabulary` and no regex matched it. Until that lands, both run on every Bash PostToolUse. + - Deletion of the PreToolUse emitter and of the two pre hooks from `settings.json`. + - A delivered-versus-emitted distinction in `ways introspect` and `scripts/hook-fire-detector.sh`, so a future silent surface is visible within days. + - An authoring pass over check and way bodies written in act-time voice. + +7. **Guidance at act time waits on a harness change.** PreToolUse carries guards (ADR-181), the Task stash, and `mark-tasks-active.sh`. Upstream closed the report without action, so there is no tracked fix to wait on. If a future Claude Code release injects PreToolUse `additionalContext`, measured the same way as here (a `hook_additional_context` row paired to a PreToolUse `hook_success` row), a follow-up ADR can move `commands:` and `files:` back to act time. The emit-shape fix alone is not adopted now, since the closed report's logs show the documented shape undelivered too. + +8. **Acceptance gate.** Delivery of the PostToolUse envelope is settled by the 145-to-144 pairing above and needs no further probe. The load-bearing claim that remains, per the prototype-before-accept way, is that a `commands:` or `files:` way body delivered on PostToolUse is acted on within the same turn. One live session after the build confirms it: a `commands:` fire on `git commit` in the events log, the matching `hook_additional_context` row in the same transcript, and the assistant's next action following the way. That observation is recorded here before the status moves to Accepted. + +Reversibility: cheap for points 1, 2, 4, and 5, which move existing matchers between hooks and delete one emitter. Point 3 removes one branch in `scan::command`; way embeddings are untouched, so re-adding it needs no corpus rebuild. + +## Consequences + +### Positive + +- Every `commands:` and `files:` trigger in the corpus reaches the model for the first time, one tool call after the match. Commit guidance arrives before the next commit or the PR; `files:` guidance on the first edit of a file arrives before the second. +- The tool lanes stop muting the prompt lane. A way matched on a tool call is recorded once it is delivered, so a later prompt-lane match inside the refire window is suppressed only when the model has already read the way. +- Silent injection volume drops by 7,401 fires a month (6,905 semantic plus 496 regex). Delivered volume becomes the 496 regex fires, the `files:` fires, and whatever #525 adds. The tail of calls firing 4 or more ways, and the 10.4k-character commit the probe measured, go with the semantic lane. +- Subagents keep tool-lane delivery, since PostToolUse fires inside them and `check-post.sh` already exports `agent_id`. +- Delivery rests on a mechanism measured in this project's own transcripts and rated as followed in the audit. +- The Task lane's recording defect (log without stamp) is fixed by the same rule. + +### Negative + +- Guidance arrives after the action. A commit-message way fires after the commit exists; a check written as "before you run X" lands after X ran. The authoring pass in point 6 covers the bodies in act-time voice. +- The scan subprocess relocates from the pre hook to the post hook; per Bash call the process count is unchanged. Until lazy embedding lands, the embed pass and the lookbehind run on every Bash PostToolUse for check scoring. +- The Bash surface loses its semantic channel, and with it the possibility ADR-155 §4 named of a way reaching the moment before action on Claude's stated intent. The measured lane never delivered and its samples were noisy, so the loss is a possibility. The response-context channel carries the reasoning to the next prompt. +- ADR-155 part 5 planned to read Bash-surface semantic fires as telemetry for pattern demotion. The month of events already in the log remains usable for that; the stream stops with this ADR. The prompt lane's `way_keyword_gated` and `way_nearmiss` streams continue. +- The events log for the measured window counts fires that were never delivered. Analyses over `way_fired` before this ADR ships must treat channels `bash`, `semantic:bash:*`, and `file` as undelivered. + +### Neutral + +- ADR-155 §4 is superseded by point 3 and the frontmatter on both documents says so. The rest of ADR-155 stands with status Accepted. +- The transcript audit's finding on tool triggers is re-read as a finding about PostToolUse postchecks. Its recommendation survives on the mechanism this ADR adopts for all tool-lane matches. +- Engagement state stamped by silent fires lives in per-session directories and expires with the session. No migration. +- `check-queued.sh` keeps printing its own envelope on PostToolUse. It is a separate hook command with its own stdout, so the one-document rule in point 1 holds per command. +- PreToolUse becomes a refusal surface in this project. Anything wired there either exits 2 with a reason (ADR-181) or writes state for a later hook to drain. +- The probe's spike branch stays unmerged. Its stash-and-drain code is a measured starting point should PostToolUse delivery regress upstream and a main-session fallback be needed. + +## Alternatives Considered + +- **Stash on PreToolUse, drain at the next UserPromptSubmit (option 1 in #528).** The Task lane's pattern pointed at the main session. The probe above built it and it delivers once. Rejected on the probe's own findings. Latency is bounded by the turn: an autonomous turn runs tens of tool calls with no prompt between them (65 Bash-lane fires in the session that filed #528), and a commit-format way stashed at the first commit drains after the PR is open. Subagents never receive UserPromptSubmit, so their stashes are orphaned. Checks lose their pre-flight meaning. It adds a stash, a drain, a deliver-once record, and the identity questions ADR-172 had to settle for its drain, where PostToolUse adds nothing. + +- **Match on PreToolUse, stash, drain on the same call's PostToolUse.** Rejected. `tool_input` is present on PostToolUse, so the match runs there with the same inputs and no stash. A PreToolUse match also fires on a call the permission system then denies, and a stash for it would need a claim and a cleanup per call. + +- **Scan commands print their own envelope from inside `check-post.sh`.** Rejected. Two JSON documents on one hook command's stdout and the harness parses neither. The dispatch captures bare text and the hook prints one envelope. + +- **A third hook command on PostToolUse for the scan dispatch.** Workable, since each hook command has its own stdout. Rejected in favour of the dispatch inside `check-post.sh`: one process already reads the payload and prints the envelope, and a second `ways` subprocess per tool call would cost more than the dispatch saves in separation. + +- **Deliver on Stop.** ADR-172's drain point. Rejected. ADR-161 rejected Stop for the queued-message lane because it cannot steer the current turn, and the same holds here. It also inherits the re-entry and livelock bounds ADR-172 records. + +- **Keep the semantic Bash lane and deliver it on PostToolUse.** Rejected. The lane's measured samples fire off-topic at high scores, 20% of firing calls emit 4 or more ways, and one plain `git commit` produced 10.4k characters. Delivering that volume would be a regression from the silence the month measured. If a semantic act-time lane is wanted later, it starts from a calibration pass over the logged month, on its own ADR. + +- **Fix the emit shape to `hookSpecificOutput.additionalContext` on PreToolUse and re-measure.** Rejected as the fix. The closed upstream report's own logs show that shape undelivered, and the `emit_hook_context` docstring records the same silent drop on SessionStart for an undocumented shape. It is cheap to try beside the acceptance observation, and a positive result reopens act-time delivery under point 7. + +- **Keep recording as it is and let refire absorb the suppression.** Rejected. The suppression is the second defect in #528, and the probe's first finding shows the same defect recurs in any carrier that stamps at match time. A fire nobody read must never count as a disclosure, whatever the delivery mechanism. + +- **Retire the tool lanes entirely and rely on prompt triggers.** Rejected. The 144 postcheck deliveries were rated as followed in the audit, `commands:` triggers are the corpus's exact deterministic triggers (ADR-125, ADR-155 part 5), and #525 needs a tool lane to fire file ways in auto mode. diff --git a/docs/design-notes/session-introspection-implementation-plan.md b/docs/architecture/ways/ADR-189-session-introspection-implementation-plan.md similarity index 97% rename from docs/design-notes/session-introspection-implementation-plan.md rename to docs/architecture/ways/ADR-189-session-introspection-implementation-plan.md index 3158b8d0..f74cb7c2 100644 --- a/docs/design-notes/session-introspection-implementation-plan.md +++ b/docs/architecture/ways/ADR-189-session-introspection-implementation-plan.md @@ -1,4 +1,15 @@ -## Session Introspection — Implementation Plan +--- +contract: adr/v1 +kind: evidence +capability: matching +status: accepted +date: 2026-07-02 +deciders: + - aaronsb +related: [] +--- + +# ADR-189: Session introspection implementation plan > *Written during the ADR-153/154 sequencing work, before ADR-156 shipped > (pre-156). The body below still treats a semantic fire as raw cosine @@ -10,8 +21,6 @@ > reasoning that led here; for the shipped model see ADR-156 and > `../hooks-and-ways/engine-reference.md`.* -> **Type:** Design note (not an ADR) -> **Status:** Working draft — sequencing for ADR-153 + ADR-154, deferred implementation > **Cites:** ADR-153 (introspection substrate), ADR-154 (three front-ends), ADR-201 (shared finding evidence), ADR-134 (near-miss telemetry) > **Motivates:** the `ways introspect <replay|live|dump>` surface (`ways rethink` fixes + a live monitor + non-interactive dump) and the why-fired drill-down diff --git a/docs/design-notes/lexical-gate-as-conditional-threshold.md b/docs/architecture/ways/ADR-190-the-lexical-gate-as-a-conditional-threshold.md similarity index 98% rename from docs/design-notes/lexical-gate-as-conditional-threshold.md rename to docs/architecture/ways/ADR-190-the-lexical-gate-as-a-conditional-threshold.md index fac31aef..66e5c037 100644 --- a/docs/design-notes/lexical-gate-as-conditional-threshold.md +++ b/docs/architecture/ways/ADR-190-the-lexical-gate-as-a-conditional-threshold.md @@ -1,4 +1,15 @@ -# The Lexical Gate as a Conditional Threshold +--- +contract: adr/v1 +kind: evidence +capability: matching +status: accepted +date: 2026-07-04 +deciders: + - aaronsb +related: [] +--- + +# ADR-190: The lexical gate as a conditional threshold > *Written 2026-07-04, during the ADR-156 exploration (pre-ship). The body treats > the `γ·T_w` gate on raw cosine, the per-way `embed_threshold`, and the global diff --git a/docs/design-notes/tool-use-channel-lookbehind-chunk-matching.md b/docs/architecture/ways/ADR-191-the-tool-use-channel-is-a-signal-problem-lookbehind-chunk-spread-and-winner-confirmation.md similarity index 98% rename from docs/design-notes/tool-use-channel-lookbehind-chunk-matching.md rename to docs/architecture/ways/ADR-191-the-tool-use-channel-is-a-signal-problem-lookbehind-chunk-spread-and-winner-confirmation.md index 45a4adaa..8e8969f0 100644 --- a/docs/design-notes/tool-use-channel-lookbehind-chunk-matching.md +++ b/docs/architecture/ways/ADR-191-the-tool-use-channel-is-a-signal-problem-lookbehind-chunk-spread-and-winner-confirmation.md @@ -1,8 +1,17 @@ -## The Tool-Use Channel is a Signal Problem: Lookbehind, Chunk-Spread, and Winner Confirmation - -> **Type:** Design note (not an ADR) -> **Status:** Exploratory — prototype findings from a 2026-07-05 probe session, pre-decision -> **Cites:** ADR-125 (embedding as hard dependency), ADR-155 (the keyword gate), ADR-156 (calibrated fire in probability space, `g(s) ≥ τ_s`), and the sibling note [The Lexical Gate as a Conditional Threshold](./lexical-gate-as-conditional-threshold.md) +--- +contract: adr/v1 +kind: evidence +capability: matching +status: accepted +date: 2026-07-04 +deciders: + - aaronsb +related: [] +--- + +# ADR-191: The tool-use channel is a signal problem: lookbehind, chunk-spread, and winner confirmation + +> **Cites:** ADR-125 (embedding as hard dependency), ADR-155 (the keyword gate), ADR-156 (calibrated fire in probability space, `g(s) ≥ τ_s`), and the sibling note [The Lexical Gate as a Conditional Threshold](ADR-190-the-lexical-gate-as-a-conditional-threshold.md) > **Motivates:** a possible ADR for the tool-use (bash) channel — embedding the *intent behind* a command rather than the command string; a `way-embed match --batch` addition; a per-way body cross-similarity confirmation step; and a project/domain scope gate ## What this note is diff --git a/docs/design-notes/late-interaction-matching-flow.md b/docs/architecture/ways/ADR-192-late-interaction-matching-the-flow.md similarity index 96% rename from docs/design-notes/late-interaction-matching-flow.md rename to docs/architecture/ways/ADR-192-late-interaction-matching-the-flow.md index b1ba5f7d..2f320394 100644 --- a/docs/design-notes/late-interaction-matching-flow.md +++ b/docs/architecture/ways/ADR-192-late-interaction-matching-the-flow.md @@ -1,4 +1,15 @@ -# Late-interaction matching — the flow +--- +contract: adr/v1 +kind: evidence +capability: matching +status: accepted +date: 2026-07-06 +deciders: + - aaronsb +related: [] +--- + +# ADR-192: Late-interaction matching: the flow A visual companion to **ADR-160** (chunked late-interaction matching with softmax-share gating). It diagrams two things this session settled by watching diff --git a/docs/architecture/system/multilingual-model-evaluation.md b/docs/architecture/ways/multilingual-model-evaluation.md similarity index 100% rename from docs/architecture/system/multilingual-model-evaluation.md rename to docs/architecture/ways/multilingual-model-evaluation.md diff --git a/docs/attend-and-monitor/README.md b/docs/attend-and-monitor/README.md index 2aab5b32..b49fcf5d 100644 --- a/docs/attend-and-monitor/README.md +++ b/docs/attend-and-monitor/README.md @@ -6,7 +6,7 @@ This directory documents the active awareness layer. Attend gives a session the Monitor is Claude Code's general-purpose async delivery mechanism: launch a command, stream its stdout as notifications. Anthropic's assumption in designing it seems to be that Claude will wire up whatever ad-hoc command fits the moment — a `tail -f`, a `cargo watch`, a bespoke shell pipeline — and let Monitor relay whatever comes out. That's a powerful primitive but it puts the burden on every Claude session to reinvent the observation logic from scratch. -**Attend is Monitor with intention.** It's a long-lived logic module that Claude doesn't have to tune case by case. It knows what kinds of changes matter, it governs their rate and salience through a formal engagement model (ADR-119, borrowing the activation-decay shape from ACT-R), it routes messages between peer agents through focus groups (ADR-118), and it cleans up after itself (ADR-121 for presentation decay, plus the 30-day disk sweep). Its emission governor is alarm management (ISA-18.2) applied to agent notifications — rate-limiting what reaches the conversation so the channel stays trustworthy instead of noisy. A Claude session drops into attend and gets a stable, opinionated awareness channel for free. +**Attend is Monitor with intention.** It's a long-lived logic module that Claude doesn't have to tune case by case. It knows what kinds of changes matter, it governs their rate and salience through a formal engagement model (ADR-123, borrowing the activation-decay shape from ACT-R), it routes messages between peer agents through focus groups (ADR-118), and it cleans up after itself (ADR-123 for presentation decay, plus the 30-day disk sweep). Its emission governor is alarm management (ISA-18.2) applied to agent notifications — rate-limiting what reaches the conversation so the channel stays trustworthy instead of noisy. A Claude session drops into attend and gets a stable, opinionated awareness channel for free. Said another way: **Monitor is a delivery mechanism; attend is the editorial layer that decides what's worth delivering.** That editorial policy is calm technology (Weiser & Brown) applied to a coding session: most changes stay in the periphery, and only what deserves attention moves to the center. The combination turns "sporadic unpredictable state" into a structured stream of observations the agent can act on. @@ -62,7 +62,7 @@ See [`authoring-sensors.md`](authoring-sensors.md) for the full author's guide, | [`sensors.md`](sensors.md) | What each built-in sensor observes and how it emits (planned) | | [`signals.md`](signals.md) | Signal file format, storage layout, lifecycle (planned) | | [`engagement.md`](engagement.md) | Action potential model in prose and diagrams (planned) | -| [`salience.md`](salience.md) | Turn-based presentation decay — the ADR-121 mechanism (planned) | +| [`salience.md`](salience.md) | Presentation decay — the outward gate of the ADR-123 engine | | [`focus-groups.md`](focus-groups.md) | Dynamic group membership and signal routing (planned) | | [`configuration.md`](configuration.md) | Config schema, overlay semantics, tuning workflow (planned) | @@ -84,17 +84,18 @@ Files marked **planned** are part of the ongoing documentation pass. ## Related docs - [`../vocabulary.md`](../vocabulary.md) — terminology anchors mapping the project's coined terms to their established concepts (ADR-301) -- **ADR-113** (`docs/architecture/system/`) — the original decision to build attend as an active awareness module +- **ADR-113** (`docs/architecture/ways/`) — the original decision to build attend as an active awareness module - **ADR-114** — attend as an insistent trigger type for ways - **ADR-115** — declarative config with project-scope overlay - **ADR-116** — permission requirements - **ADR-117** — sensor crate extraction - **ADR-118** — focus groups, dynamic agent grouping -- **ADR-119** — action potential engagement model +- **ADR-119** — action potential engagement model (superseded by ADR-123) <!-- adr-cite-ignore --> - **ADR-120** — interactive chat TUI, human in the signal loop -- **ADR-121** — salience decay for signal presentation (draft) +- **ADR-121** — salience decay for signal presentation (superseded by ADR-123) <!-- adr-cite-ignore --> +- **ADR-123** — firing dynamics unification; the engagement and salience engine in force - `docs/hooks-and-ways/` — sibling docs for the synchronous hook mechanism -- `docs/design-notes/cognitive-loop-and-awareness-layer.md` — earlier design exploration that informed ADR-113 +- `docs/architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md` — earlier design exploration that informed ADR-113 ## Where the code lives diff --git a/docs/attend-and-monitor/authoring-sensors.md b/docs/attend-and-monitor/authoring-sensors.md index 79755eba..9b8a509c 100644 --- a/docs/attend-and-monitor/authoring-sensors.md +++ b/docs/attend-and-monitor/authoring-sensors.md @@ -1,6 +1,6 @@ # Authoring sensors -Sensor authorship is a first-class design surface in attend. A sensor is not a log tail — it's a module that translates raw environmental change into *magnitude-weighted observations* that feed attend's engagement model (ADR-119) and disclosure governor. A well-designed sensor encodes how much each kind of change matters in the magnitude, and lets the loop handle when to fire, how often, and whether to suppress. This is the discipline alarm management (ISA-18.2) brings to control rooms, applied to agent notifications: a channel stays trustworthy only as long as its loudest tier is reserved for what is genuinely loud. +Sensor authorship is a first-class design surface in attend. A sensor is not a log tail — it's a module that translates raw environmental change into *magnitude-weighted observations* that feed attend's engagement model (ADR-123) and disclosure governor. A well-designed sensor encodes how much each kind of change matters in the magnitude, and lets the loop handle when to fire, how often, and whether to suppress. This is the discipline alarm management (ISA-18.2) brings to control rooms, applied to agent notifications: a channel stays trustworthy only as long as its loudest tier is reserved for what is genuinely loud. This page is for people building sensors, in either of attend's two implementations. @@ -40,7 +40,7 @@ Before you write a single line of code, understand what your events will encount 1. **Accumulator.** Your `(magnitude, description)` pair is added to the sensor's `DeltaAccumulator`. Magnitudes accumulate across ticks until the sensor is drained or decays. 2. **Emission threshold.** A per-sensor threshold — the accumulator has to exceed this before the sensor is a candidate to disclose. Low-magnitude events accumulate silently until several of them add up; a single high-magnitude event may cross threshold on its own. -3. **Engagement / refractory (ADR-119).** After the sensor recently fired a disclosure, its effective threshold is temporarily elevated (relative refractory) or it's fully suppressed (absolute refractory, ~60s by default). During refractory, new events still accumulate but don't fire until the cooldown passes — unless their magnitude is high enough to break through the elevated threshold. This is how attend models "disengagement after a burst": low-magnitude follow-ups get swallowed, truly urgent events still get through. +3. **Engagement / refractory (ADR-123).** After the sensor recently fired a disclosure, its effective threshold is temporarily elevated (relative refractory) or it's fully suppressed (absolute refractory, ~60s by default). During refractory, new events still accumulate but don't fire until the cooldown passes — unless their magnitude is high enough to break through the elevated threshold. This is how attend models "disengagement after a burst": low-magnitude follow-ups get swallowed, truly urgent events still get through. 4. **Disclosure governor.** Even after the sensor is ready, a global governor rate-limits disclosures across the whole loop (default: 3 per 120s, with a 15s cooldown between them). If a burst of sensors all want to fire at once, some get held until the window rolls. 5. **Monitor delivery.** Whatever survives all of the above gets printed as one stdout line per event and picked up by Monitor (or the `attend chat` TUI) as an async notification into the conversation. @@ -252,7 +252,7 @@ Once the script's stdout reaches attend: 1. Each valid `magnitude|description` line becomes one event. 2. Events accumulate in the sensor's `DeltaAccumulator` between polls. 3. When the accumulator exceeds the emission threshold (2.5 for this sensor), the sensor becomes a disclosure candidate. -4. The disclosure governor decides whether to actually fire (respecting cooldown, rate window, and any active refractory from ADR-119). +4. The disclosure governor decides whether to actually fire (respecting cooldown, rate window, and any active refractory from ADR-123). 5. If fired, each event becomes one Monitor notification line delivered into the conversation (or rendered in `attend chat` if the consumer is a human). The sensor author never touches any of that machinery. They just write `gh`-CLI glue and pick magnitudes. @@ -276,5 +276,5 @@ Written as heuristics, not commandments: - [`sensors.md`](sensors.md) *(planned)* — reference for the built-in sensors - [`configuration.md`](configuration.md) *(planned)* — config schema for declaring sensors - **ADR-117** — sensor crate extraction and feature flags -- **ADR-119** — action potential engagement model +- **ADR-123** — firing dynamics; the action potential engagement model in force - **ADR-116** — permission requirements for sensors diff --git a/docs/attend-and-monitor/configuration.md b/docs/attend-and-monitor/configuration.md index 0b539571..c885bd44 100644 --- a/docs/attend-and-monitor/configuration.md +++ b/docs/attend-and-monitor/configuration.md @@ -47,7 +47,7 @@ governor: max_per_window: 3 # max disclosures in rate_window rate_window: 120 # seconds of the rolling rate window -# Action potential engagement model (ADR-119, unified in ADR-123) +# Action potential engagement model (ADR-123) # Run `attend tune` to auto-derive these from real session history engagement: burst_threshold: 3 # disclosures before refractory kicks in @@ -286,7 +286,7 @@ If you need something the parser doesn't handle, either restructure or file an i - **ADR-115** — declarative config with project-scope overlay (the pattern this implements) - **ADR-116** — permission requirements - **ADR-117** — sensor crate extraction (feature flags for compile-time sensor selection) -- **ADR-119** — action potential engagement (the `engagement` block) +- **ADR-123** — action potential engagement (the `engagement` block) - [`engagement.md`](engagement.md) — engagement model in depth - [`authoring-sensors.md`](authoring-sensors.md) — how to declare and write new sensors - [`sensors.md`](sensors.md) — the built-in sensors and their default values diff --git a/docs/attend-and-monitor/engagement.md b/docs/attend-and-monitor/engagement.md index fa3a6528..345bfaa8 100644 --- a/docs/attend-and-monitor/engagement.md +++ b/docs/attend-and-monitor/engagement.md @@ -2,7 +2,7 @@ Attend's engagement model governs *when a sensor is allowed to fire a disclosure*. The idea it implements is established: in cognitive architectures like ACT-R, base-level activation decays with disuse, and recent activity changes how easily the next stimulus gets through. Attend applies that idea per sensor, borrowing its shape from the neuronal action potential: resting baseline, rapid rise on stimulus, refractory period after a burst, gradual return to rest. The biology gives us a predictable, well-studied shape for a phenomenon we actually care about — how productive engagement with a stimulus decays naturally over time. -This page covers the model in prose and diagrams, explains what each parameter does, and walks through how it interacts with the disclosure governor. The canonical architecture is **[ADR-123](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md)** — the progression-axis unification that moved the firing-dynamics core into a shared crate consumed by both attend and ways. This page is the attend-specific, implementer-and-author-friendly explainer for how attend instantiates that core. +This page covers the model in prose and diagrams, explains what each parameter does, and walks through how it interacts with the disclosure governor. The canonical architecture is **[ADR-123](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md)** — the progression-axis unification that moved the firing-dynamics core into a shared crate consumed by both attend and ways. This page is the attend-specific, implementer-and-author-friendly explainer for how attend instantiates that core. ## The problem engagement solves @@ -47,7 +47,7 @@ The biological action potential has an overshoot and hyperpolarization phase too Attend's firing engine operates on an abstract monotonic **progression axis**. The engine does not know what a tick is — attend supplies one by convention. For attend, a tick is **one second of wall-clock time** (`sensor_trait::epoch_secs()`, which reads `SystemTime::now().duration_since(UNIX_EPOCH)`). -Why wall clock: attend steers external timing — peer conversations, build events, ambient awareness — which all live outside any single model's token space. Multiple attend instances may need to compare events across their own independent progressions, and wall clock is the only axis that's guaranteed common across all of them. This is the multi-observer case argued in [ADR-123 Decision 4](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory). +Why wall clock: attend steers external timing — peer conversations, build events, ambient awareness — which all live outside any single model's token space. Multiple attend instances may need to compare events across their own independent progressions, and wall clock is the only axis that's guaranteed common across all of them. This is the multi-observer case argued in [ADR-123 Decision 4](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory). The consequence for attend authors: **all engagement parameters are in seconds**. `absolute_refractory: 60` means 60 wall-clock seconds. Half-lives are in wall-clock seconds. If you ever see a parameter expressed as a "tick count" in the code, interpret it as seconds for attend specifically. @@ -223,12 +223,12 @@ All tuning is via attend config, using the ADR-115 overlay pattern — user scop The shift in this page's framing (vs earlier versions) is that engagement no longer exclusively belongs to attend. After ADR-123, the engagement state machine — `EngagementState` with a `Curve::ActionPotential` — is a shared crate (`sensor-trait::engagement`) consumed by both attend and ways. Attend uses it with wall-clock-seconds ticks and action-potential refractory. Ways uses the same engine with token-position ticks and `Curve::Exponential` outward-gate salience. Same math, different axis, different curve variant. -For the ways-side equivalent of this page see [ADR-123 Decision 4](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory) and the context-decay presentation-economics model in [`../hooks-and-ways/context-decay.md`](../hooks-and-ways/context-decay.md). The shared engine means any future improvement to burst detection, refractory decay, or curve shapes lands in one place and reaches both tools. +For the ways-side equivalent of this page see [ADR-123 Decision 4](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory) and the context-decay presentation-economics model in [`../hooks-and-ways/context-decay.md`](../hooks-and-ways/context-decay.md). The shared engine means any future improvement to burst detection, refractory decay, or curve shapes lands in one place and reaches both tools. ## Related -- **[ADR-123](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md)** — the progression-axis unification and curve-as-parameter framing -- **ADR-119** — the original action potential model (pre-unification; superseded by ADR-123 for the math, preserved for the biology-analogy framing) +- **[ADR-123](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md)** — the progression-axis unification and curve-as-parameter framing +- **ADR-119** — the original action potential model (pre-unification; superseded by ADR-123 for the math, preserved for the biology-analogy framing) <!-- adr-cite-ignore --> - [`loop.md`](loop.md) — where engagement state sits in the loop iteration - [`authoring-sensors.md`](authoring-sensors.md) — how sensor authors design around engagement - [`configuration.md`](configuration.md) — full config schema for engagement parameters diff --git a/docs/attend-and-monitor/focus-groups.md b/docs/attend-and-monitor/focus-groups.md index 3aad2b87..1b52c096 100644 --- a/docs/attend-and-monitor/focus-groups.md +++ b/docs/attend-and-monitor/focus-groups.md @@ -24,7 +24,7 @@ Groups compose naturally with the other two scopes: | Focus group | `@<name>/` | anyone who joined the named group | | Broadcast | `_broadcast/` | everyone with attend running | -A single `attend send` can fan out to multiple scopes via flags; the default (ADR-119) is broadcast. +A single `attend send` can fan out to multiple scopes via flags; the default (ADR-119) is broadcast. <!-- adr-cite-ignore --> ## CLI surface @@ -101,7 +101,7 @@ This means group membership is **self-reported**. An agent adds itself to a grou The provider mechanism is simple: when `sensor-peers` is registered during startup, it receives a closure that clones the `Groups` handle. On each scan, it calls the closure, which returns the current list of joined group directories. The closure closes over an owned clone of `Groups` so it doesn't hold a borrow into the main loop state. -## Interaction with action potential (ADR-119) +## Interaction with action potential (ADR-123) Focus groups and engagement compose cleanly. Groups scope **which signals reach the agent**; engagement governs **how the agent responds once they arrive**. They operate at different layers: @@ -112,9 +112,9 @@ Focus groups and engagement compose cleanly. Groups scope **which signals reach The per-peer engagement boost (see [`engagement.md`](engagement.md)) also applies across focus group boundaries. A peer you've been actively chatting with in `@deploy` gets their messages boosted globally, not just within that group. This is usually the right shape — "I've been talking to this agent a lot" is a conversation-level state, not a group-level state. -## The routing simplification from ADR-119 +## The routing simplification from ADR-119 <!-- adr-cite-ignore --> -ADR-119 collapsed attend's peer-messaging routing down to a single default: **broadcast**. Before ADR-119, an agent had to reason about where to send messages — "should this go to a focus group? to a specific cwd? to broadcast?" — and routinely got it wrong. After ADR-119, the agent sends to broadcast and lets the engagement model sort out who pays attention. +ADR-119 collapsed attend's peer-messaging routing down to a single default: **broadcast**. Before ADR-119, an agent had to reason about where to send messages — "should this go to a focus group? to a specific cwd? to broadcast?" — and routinely got it wrong. After ADR-119, the agent sends to broadcast and lets the engagement model sort out who pays attention. <!-- adr-cite-ignore --> Focus groups still exist and are still useful, but their role has shifted. They're not the primary routing mechanism anymore; the action potential's per-peer boost handles most "which agents should engage with this" decisions automatically. Groups are now better understood as **explicit scoping for cases where the engagement model isn't enough**: @@ -133,7 +133,8 @@ Importantly, clicking a group in the sidebar is a TUI-local filter — it doesn' ## Related - **ADR-118** — the decision to build focus groups -- **ADR-119** — action potential engagement; routing simplification +- **ADR-123** — action potential engagement +- **ADR-119** — routing simplification (superseded by ADR-123) <!-- adr-cite-ignore --> - [`signals.md`](signals.md) — how `@<name>/` dirs fit into the overall signal layout - [`engagement.md`](engagement.md) — per-peer boost and refractory - [`tui.md`](tui.md) — the sidebar UI for groups diff --git a/docs/attend-and-monitor/loop.md b/docs/attend-and-monitor/loop.md index e63ed775..3ca7ac6c 100644 --- a/docs/attend-and-monitor/loop.md +++ b/docs/attend-and-monitor/loop.md @@ -22,7 +22,7 @@ Before the loop begins, `cmd_run_with_catchup` builds up the context it needs: 2. **Config** — load `~/.config/attend/config.yaml`, then overlay `<cwd>/.claude/attend.yaml` on top (ADR-115 pattern). 3. **Groups manager** — construct a `Groups` handle for the signals base and the current session ID. This is what ADR-118 focus groups ride on. 4. **Sensor registration** — `sensors::register_sensors()` walks the config and feature flags, instantiating each enabled sensor with its configured intervals and thresholds. The peers sensor receives a closure provider for focus-group directories so it can refresh membership on every scan (ADR-118 / issue #15). -5. **Engagement state** — apply ADR-119 action-potential parameters to every slot. Refractory behavior is per-sensor but the parameters are shared. +5. **Engagement state** — apply ADR-123 action-potential parameters to every slot. Refractory behavior is per-sensor but the parameters are shared. 6. **State restore** — if `~/.cache/attend/state/<session>.checkpoint` exists from a previous run, import the saved seen-signals and disclosed-thresholds so restart is continuous. 7. **Banner** — print a startup line unless the fingerprint (version + commit + sensor list + focus) matches the last one written to `_last_banner`, in which case print `[attend] restarted (unchanged)` to keep noisy Monitors quiet. 8. **Governor** — build a `DisclosureGovernor` with the configured cooldown, rate window, and max disclosures per window. @@ -147,7 +147,7 @@ stateDiagram-v2 **Changed but below threshold** means the sensor saw something but the magnitude isn't high enough to emit yet. The event is accumulated in the sensor's `DeltaAccumulator` and will be combined with future events if they arrive before the accumulator decays. -**Changed, above threshold, but held in absolute refractory** means ADR-119's action potential is actively suppressing this sensor after a recent burst. The magnitude stays on the accumulator but the sensor is not added to `ready_indices`. The log line `held in absolute refractory` marks this case. +**Changed, above threshold, but held in absolute refractory** means the action potential (ADR-123) is actively suppressing this sensor after a recent burst. The magnitude stays on the accumulator but the sensor is not added to `ready_indices`. The log line `held in absolute refractory` marks this case. **Ready** means the sensor crossed threshold and engagement state permits disclosure. The sensor index is appended to `ready_indices`, which the loop processes in a batch after draining. @@ -252,4 +252,4 @@ There's no graceful shutdown hook. The `attend run` process is meant to be start - **Engagement curve**: `engagement.md` covers the action potential model — why sensors go quiet after bursts. - **Signal format**: `signals.md` covers the wire format, on-disk layout, and lifecycle. - **Configuration**: `configuration.md` covers the YAML overlay and how to reshape any of the timers above. -- **Salience decay**: `salience.md` covers ADR-121's presentation-layer aging — orthogonal to this loop, sitting in the emit path. +- **Salience decay**: `salience.md` covers the ADR-123 presentation-layer aging — orthogonal to this loop, sitting in the emit path. diff --git a/docs/attend-and-monitor/salience.md b/docs/attend-and-monitor/salience.md index 2aaece7a..3b894a1e 100644 --- a/docs/attend-and-monitor/salience.md +++ b/docs/attend-and-monitor/salience.md @@ -2,7 +2,7 @@ This page covers the **presentation-layer aging** mechanism: how a signal's visibility in the conversation fades as the progression axis advances, even while the signal file remains on disk. The shape is the forgetting curve applied to notifications — relevance decays exponentially with quiet, and re-engagement resets it, the same spacing logic spaced-repetition systems use to decide when something needs showing again. It's the cousin — not the opposite — of attend's inward-gate engagement model in [`engagement.md`](engagement.md). Both sides of the gate share a single engine; this page is the outward-side explainer. -**Status.** The framing comes from [ADR-121](../architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md). The mechanism it describes was unified with attend's inward gate in [ADR-123](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md): both now consume the same `sensor_trait::Curve` type with the same `salience_at(delta)` query. **Ways was the first concrete implementation**, shipping cross-tool via ADR-123. **Attend's sensor-peers application now ships too** (issue #22) — each peer signal passes through a per-id `EngagementState<Curve::Exponential>` before emitting, so aged backlog fades without being deleted and threaded replies (`re:<id>`) reset the parent's salience. This page describes the decision and points at both production implementations as canonical references. +**Status.** The framing comes from [ADR-121](../architecture/attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md). The mechanism it describes was unified with attend's inward gate in [ADR-123](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md): both now consume the same `sensor_trait::Curve` type with the same `salience_at(delta)` query. **Ways was the first concrete implementation**, shipping cross-tool via ADR-123. **Attend's sensor-peers application now ships too** (issue #22) — each peer signal passes through a per-id `EngagementState<Curve::Exponential>` before emitting, so aged backlog fades without being deleted and threaded replies (`re:<id>`) reset the parent's salience. This page describes the decision and points at both production implementations as canonical references. <!-- adr-cite-ignore --> ## The problem the outward gate solves @@ -10,11 +10,11 @@ Without an outward gate, every signal that ever arrived keeps trying to be shown In a three-hour session, a signal that arrived in the first few minutes is still being surfaced after two hours of unrelated work — even though by then the cursor has moved over an order of magnitude more context. The signal isn't stale in any disk-cleanup sense (it's hours old, not days), but it's stale in a *relevance* sense. The cursor has moved on. -ADR-121's framing, preserved verbatim under ADR-123: **signals need presentation-layer aging, measured in the progression axis, decaying smoothly.** The only change from ADR-121 to ADR-123 is that "measured in turns" is now "measured in whatever tick unit the caller supplies" — seconds for attend, tokens for ways. +ADR-121's framing, preserved verbatim under ADR-123: **signals need presentation-layer aging, measured in the progression axis, decaying smoothly.** The only change from ADR-121 to ADR-123 is that "measured in turns" is now "measured in whatever tick unit the caller supplies" — seconds for attend, tokens for ways. <!-- adr-cite-ignore --> ## Inward and outward gates, one engine -ADR-119 and ADR-121 originally framed engagement and salience as two mechanisms. ADR-123 exposed them as two queries against the same state: +ADR-119 and ADR-121 originally framed engagement and salience as two mechanisms. ADR-123 exposed them as two queries against the same state: <!-- adr-cite-ignore --> | | Inward gate (engagement) | Outward gate (salience) | |---|---|---| @@ -59,7 +59,7 @@ Callers compare this against a **re-fire floor**. Different callers pick differe Ways is the first tool that consumes the outward gate in production. The surface: -- Every way declares an explicit `curve:` in its frontmatter. Most declare `Curve::Exponential { half_life: N }` in tokens. See [ADR-123 §7](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md#7-frontmatter-schema-migration--no-shims). +- Every way declares an explicit `curve:` in its frontmatter. Most declare `Curve::Exponential { half_life: N }` in tokens. See [ADR-123 §7](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md#7-frontmatter-schema-migration--no-shims). - On each match, `session::way_fire_outcome(way, session, curve)` asks the engine whether the way should re-fire. Under the hood: ```rust @@ -79,7 +79,7 @@ match state.current_salience(current_tick) { For ways with `Curve::Exponential { half_life: H }` and the 0.5 floor, re-fire happens at exactly delta `H` — the half-life *is* the re-fire distance. Ways' "20–30K intervals" footer in `ways list` is exactly this: the range of `half_life` values across the active way set. -Per-way visualization in `ways list` and `ways rethink` uses `Curve::refire_delta(floor)` to render each row's bar and forecast position from its own threshold, not a shared global. See [ADR-123 §2](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md#2-pluggable-curve-as-first-class-parameter) for the curve types and [`engagement.md`](engagement.md) for the engine's other queries. +Per-way visualization in `ways list` and `ways rethink` uses `Curve::refire_delta(floor)` to render each row's bar and forecast position from its own threshold, not a shared global. See [ADR-123 §2](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md#2-pluggable-curve-as-first-class-parameter) for the curve types and [`engagement.md`](engagement.md) for the engine's other queries. ## Attend's concrete application (shipping now) @@ -110,7 +110,7 @@ impl SignalSalience { The key design points: - **Signal id** is the filename stem — the same shape `re:<id>` references in ADR-120 threaded replies, so reply resets are direct hash lookups. No rename table, no ownership bookkeeping. -- **Arrival tick** comes from the file's on-disk mtime, not the first time this observer happens to scan the directory. A peer who joins a focus-group room mid-session sees old signals as already-decayed — the backlog-filter behavior ADR-121 designed, now working against real backlogs instead of hypotheticals. +- **Arrival tick** comes from the file's on-disk mtime, not the first time this observer happens to scan the directory. A peer who joins a focus-group room mid-session sees old signals as already-decayed — the backlog-filter behavior ADR-121 designed, now working against real backlogs instead of hypotheticals. <!-- adr-cite-ignore --> - **Re-engagement reset** runs unconditionally at scan time: if the content is a threaded 5-field signal (`re:<id>|`), the parent signal's `EngagementState` is bumped back to 1.0 at the current tick before the current signal's own gate check runs. The reset persists into the checkpoint so reconnection does not re-age the parent. - **Checkpoint persistence** uses the existing `sensor_trait::Sensor::export_state`/`import_state` wire format, adding a new `signal_salience` row key that carries `<signal_id>\t<json>`. Old checkpoints without the key parse cleanly — the sensor just starts fresh on those signals. - **No engine changes.** The full implementation is one new file in sensor-peers, one new config block, and ~30 lines of wiring inside `read_signals`. That's the payoff ADR-123 designed for. @@ -129,13 +129,13 @@ Defaults are conservative first-value picks, subject to `attend tune` once surve ### Why the session-length distribution still matters -ADR-121 grounded its half-life default in real session data: +ADR-121 grounded its half-life default in real session data: <!-- adr-cite-ignore --> ``` min=1 median=12 p75=45 p90=84 max=133 mean=32.6 turns ``` -This was turns, not seconds, under the ADR-121 framing. Under ADR-123 the shape of the argument is the same, the axis is different: attend picks a wall-clock half-life that behaves reasonably against the same distribution when converted through the observed turn-cadence. `attend tune` already surveys turn cadence from real transcripts and derives engagement parameters; the outward-gate half-life is the natural next parameter for it to derive from the same analysis (tracked as a follow-up to issue #22). +This was turns, not seconds, under the ADR-121 framing. Under ADR-123 the shape of the argument is the same, the axis is different: attend picks a wall-clock half-life that behaves reasonably against the same distribution when converted through the observed turn-cadence. `attend tune` already surveys turn cadence from real transcripts and derives engagement parameters; the outward-gate half-life is the natural next parameter for it to derive from the same analysis (tracked as a follow-up to issue #22). <!-- adr-cite-ignore --> The intuition to preserve: **short sessions never see aging (nothing to prune), medium sessions see the oldest material drop out cleanly near the end, long sessions experience meaningful pruning.** Whatever wall-clock half-life attend eventually picks should produce this behavior against the live distribution, not against a one-off sweep. The 1800 s default is the placeholder — deliberately uncalibrated — until tune closes the loop. @@ -150,7 +150,7 @@ The `seen_signals` invariant still prevents the same observer from surfacing the - Any other observer joining the same shared signal dir after the reply arrived. - Any future session of the same observer that re-reads the shared dir from checkpoint state. -Widening the reset to also re-surface already-presented parents in the current observer is a targeted follow-up, not a v1 requirement — the primary ADR-121 win is backlog filtering on new-observer entry, which the above implementation fully delivers. +Widening the reset to also re-surface already-presented parents in the current observer is a targeted follow-up, not a v1 requirement — the primary ADR-121 win is backlog filtering on new-observer entry, which the above implementation fully delivers. <!-- adr-cite-ignore --> ## Re-engagement resets @@ -168,12 +168,12 @@ In ways' case, "re-engagement" is structurally different: a way re-fires when a Salience decay is a **curve shape** concern; what delta the curve is evaluated against is an **axis choice** concern. They're orthogonal: -- Attend's axis is wall-clock seconds because attend coordinates events across multiple peers. Each peer has its own internal progression, but wall clock is the only axis guaranteed common across all peers. This is the multi-observer dimensionality argument in [ADR-123 Decision 4](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory). +- Attend's axis is wall-clock seconds because attend coordinates events across multiple peers. Each peer has its own internal progression, but wall clock is the only axis guaranteed common across all peers. This is the multi-observer dimensionality argument in [ADR-123 Decision 4](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md#4-ways-tick-unit-host-addressing-not-a-decay-theory). - Ways' axis is token position because ways steers a single-observer host (one session, one monotonic token stream) and token position is the unit the host uses internally to address content. The model can observe it directly via RoPE; wall clock is external to the model's attention. The salience curve doesn't care which axis it's running against. A `Curve::Exponential { half_life: 30000 }` is 30000 whatever-the-caller-supplies. Attend reads 30000 as seconds (half an hour); ways reads 30000 as tokens (roughly one medium interaction). Both are correct for their domain because the engine is unit-agnostic. -## Alternatives rejected in ADR-121 (still valid under ADR-123) +## Alternatives rejected in ADR-121 (still valid under ADR-123) <!-- adr-cite-ignore --> Preserved here for posterity. The arguments didn't change when the engine unified; if anything, the unification makes the rejections stronger because some of the alternatives would have blocked cross-tool sharing. @@ -185,7 +185,7 @@ Preserved here for posterity. The arguments didn't change when the engine unifie ## What this page is not -- **Not the canonical architecture.** That's [ADR-123](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md), with [ADR-121](../architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md) as the original outward-gate decision. +- **Not the canonical architecture.** That's [ADR-123](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md), with [ADR-121](../architecture/attend/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md) as the original outward-gate decision. <!-- adr-cite-ignore --> - **Not the engine documentation.** That's `sensor_trait::curve::Curve` in source, with unit tests as the executable spec. - **Not a parameter-tuning guide for attend.** Attend's outward-gate parameters don't exist in production yet — `attend config lint` will surface them alongside engagement parameters when sensor-peers consumes them. @@ -195,7 +195,7 @@ Preserved here for posterity. The arguments didn't change when the engine unifie - [`signals.md`](signals.md) — disk-side retention (the time-based bulk cousin). - [`loop.md`](loop.md) — where the presentation gate will sit in the attend tick loop when sensor-peers consumes it. - [`configuration.md`](configuration.md) — attend's config surface; the `signals:` block lives there alongside `engagement:`. -- **ADR-121** — the salience-decay decision record. Status under ADR-123: reframed in place; ways was the first concrete realization, sensor-peers the second. -- **ADR-123** — the progression-axis unification that made ADR-121's mechanism a cross-tool facility instead of an attend-local one. +- **ADR-121** — the salience-decay decision record. Status under ADR-123: reframed in place; ways was the first concrete realization, sensor-peers the second. <!-- adr-cite-ignore --> +- **ADR-123** — the progression-axis unification that made ADR-121's mechanism a cross-tool facility instead of an attend-local one. <!-- adr-cite-ignore --> - **`tools/ways-cli/src/session.rs`** — the `way_fire_outcome` path is the ways-side consumer, anchored to token position. - **`tools/sensor-peers/src/salience.rs`** — the attend-side consumer, anchored to wall-clock seconds. Mirror structure, same engine call. diff --git a/docs/attend-and-monitor/sensors.md b/docs/attend-and-monitor/sensors.md index 59bbda6e..ae5dda90 100644 --- a/docs/attend-and-monitor/sensors.md +++ b/docs/attend-and-monitor/sensors.md @@ -13,7 +13,7 @@ This page covers what each built-in observes, what magnitudes it emits, and what | **peers** | Other Claude sessions + signal files | 30s | 10s | 2.0 | peer status changes, peer messages | | **processes** | Build/dev-tool processes | 30s | 5s | 2.0 | process start/exit (cargo, npm, make, etc.) | -All four use the adaptive interval scheme — they poll fast (`min_interval`) during active change and slow (`base_interval`) when quiet. All four participate in the action potential engagement model (ADR-119) with the shared global config. +All four use the adaptive interval scheme — they poll fast (`min_interval`) during active change and slow (`base_interval`) when quiet. All four participate in the action potential engagement model (ADR-123) with the shared global config. ## `sensor-context` — interoceptive @@ -222,4 +222,4 @@ Almost everything falls into category 2. Crate sensors are for the core observat - **ADR-113** — the original design of attend, including the first context sensor - **ADR-117** — sensor crate extraction and feature flags - **ADR-118** — focus groups (used by `sensor-peers`) -- **ADR-119** — action potential engagement model +- **ADR-123** — action potential engagement model diff --git a/docs/attend-and-monitor/signals.md b/docs/attend-and-monitor/signals.md index a0f09b67..67a263c8 100644 --- a/docs/attend-and-monitor/signals.md +++ b/docs/attend-and-monitor/signals.md @@ -102,7 +102,7 @@ flowchart LR Write[write .tmp<br/>rename to .signal] Scan[peer sensor scans<br/>reads new files] Present[present to agent<br/>via Monitor] - Age[age with turns<br/>ADR-121 salience decay] + Age[age over time<br/>ADR-123 salience decay] Below[below presentation floor<br/>no longer shown] Cleanup[auto-cleanup sweep<br/>every 10 min] Delete[file removed<br/>after 30 days] @@ -134,7 +134,7 @@ flowchart LR **Phase 3 — presentation.** Observations become events in the peer sensor's accumulator, feed into engagement/governor, and if they survive all the gates, emit as Monitor notification lines into the conversation. The agent sees them; the human (if running `attend chat`) sees them in the TUI. -**Phase 4 — salience decay (ADR-121, drafted).** Once presented, a signal carries a salience that decays over turns. After its salience drops below the presentation floor, the signal stops appearing in notifications — but the file stays on disk. Re-engagement (a reply or reference) resets salience to 1.0 and the signal is visible again. +**Phase 4 — salience decay (ADR-123).** Once presented, a signal carries a salience that decays as time passes. After its salience drops below the presentation floor, the signal stops appearing in notifications — but the file stays on disk. Re-engagement (a reply or reference) resets salience to 1.0 and the signal is visible again. **Phase 5 — auto-cleanup.** Every `cleanup.interval` seconds (default 10 minutes), the attend loop runs a sweep of the signals base. Any `.signal` file older than `cleanup.retention` (default 30 days) is removed. Empty project subdirs left behind after the file removal are also cleaned up. @@ -150,7 +150,7 @@ Signals have **two different retention windows** that operate at different scale | | Disk retention | Attention window | |---|---|---| -| **Unit** | Time (30 days) | Turns (half-life 20, per ADR-121) | +| **Unit** | Time (30 days) | Turns (half-life 20, per ADR-121) | <!-- adr-cite-ignore --> | **Purpose** | Bulk storage hygiene | Presentation relevance | | **Controlled by** | `cleanup.retention` config | `attention.half_life` (planned) | | **Observable in** | Disk usage | Which signals Monitor notifies about | @@ -158,7 +158,7 @@ Signals have **two different retention windows** that operate at different scale The short answer on why two units: **precision where it matters, convenience where it doesn't.** Attention works in turns because turn pacing varies too much to use wall-clock time at fine grain. Disk retention works in time because at 30-day horizons the variance averages out and "30 days" is a human-readable unit everyone intuits. -See [`salience.md`](salience.md) for the attention side and the ADR-121 decay curve math. +See [`salience.md`](salience.md) for the attention side and the ADR-123 decay curve math. ## Reading signals in tooling @@ -176,7 +176,7 @@ Reading from `_broadcast/` gives you cross-agent visibility. Reading from `@<nam - **ADR-113** — the original attend design, including signal dir conventions - **ADR-118** — focus groups, `@<name>` directories - **ADR-120** — `attend chat`, the `re:` threading field -- **ADR-121** — salience decay on the presentation side +- **ADR-123** — salience decay on the presentation side - [`loop.md`](loop.md) — where signals are scanned and emitted in the loop - [`tui.md`](tui.md) — how the TUI reads and writes signals - [`focus-groups.md`](focus-groups.md) *(planned)* — `@<name>` dir management in detail diff --git a/docs/attend-and-monitor/tui.md b/docs/attend-and-monitor/tui.md index d8894e55..c9295e63 100644 --- a/docs/attend-and-monitor/tui.md +++ b/docs/attend-and-monitor/tui.md @@ -52,7 +52,7 @@ Like `attend run`, `attend chat` is a long-lived process. Unlike `attend run`, i **Left sidebar, bottom — agents.** Every peer Claude session attend has discovered is listed with ambient metadata: a status indicator (arrow = working, dot = idle), commits ahead/behind upstream, and current context percentage. This is the peer sensor's output rendered directly. -**Main area — messages.** Chronological signal stream, scoped by whichever sidebar filter is active. Each message carries a two-line sender chip. The first line is the same persona every attend conduit renders (`Jovan-alpha` for a Claude session, `aaron` for a human), derived through the shared `attend-identity-view` crate so the chip, the Monitor line, and the Stop-hook drain agree on one name. The second line is the cwd basename, falling back to `broadcast` when the signal carries no cwd. After the chip come the routing scope (`@focus-group`, `broadcast`, or project name), and the body. Messages fade after their salience drops below the presentation floor (ADR-121), matching the agent's view. +**Main area — messages.** Chronological signal stream, scoped by whichever sidebar filter is active. Each message carries a two-line sender chip. The first line is the same persona every attend conduit renders (`Jovan-alpha` for a Claude session, `aaron` for a human), derived through the shared `attend-identity-view` crate so the chip, the Monitor line, and the Stop-hook drain agree on one name. The second line is the cwd basename, falling back to `broadcast` when the signal carries no cwd. After the chip come the routing scope (`@focus-group`, `broadcast`, or project name), and the body. Messages fade after their salience drops below the presentation floor (ADR-123), matching the agent's view. **Input bar — compose.** Plain text input supports `@group` and `#issue` inline addressing. Enter sends. @@ -124,8 +124,8 @@ The chat TUI is where the human *orchestrates*. The agent terminals are where th - **ADR-120** — the design decision for this feature (draft status) - **ADR-118** — focus groups; the TUI's sidebar depends directly on this model -- **ADR-119** — action potential engagement; the TUI's message stream respects the same refractory and governor rules as the agent side -- **ADR-121** — salience decay; the TUI fades old messages using the same turn-based curve +- **ADR-123** — action potential engagement; the TUI's message stream respects the same refractory and governor rules as the agent side +- **ADR-123** — salience decay; the TUI fades old messages using the same curve - [`loop.md`](loop.md) — the sensor loop substrate that feeds the TUI - [`signals.md`](signals.md) *(planned)* — the wire format the TUI reads and writes - [`focus-groups.md`](focus-groups.md) *(planned)* — the group model the sidebar reflects diff --git a/docs/cognitive-loop.md b/docs/cognitive-loop.md index 604c05d6..b5aec50a 100644 --- a/docs/cognitive-loop.md +++ b/docs/cognitive-loop.md @@ -2,7 +2,7 @@ This document is a walk-through of the agent-ways cognitive architecture. It assumes you know what Claude Code is and nothing beyond that. It is the document to read when you want to understand how the pieces fit together — ways, progressive disclosure, the session ledger, optional memory projections, and the awareness layer — without diving into the individual ADRs. -If you want to decide a specific tradeoff, read an ADR. If you want the theoretical framing, read the [cognitive loop and awareness layer design note](design-notes/cognitive-loop-and-awareness-layer.md). If you want to build a way, read the [hooks-and-ways guide](hooks-and-ways/README.md). This document sits one level above all of those: it tells the story of how the system composes. +If you want to decide a specific tradeoff, read an ADR. If you want the theoretical framing, read the [cognitive loop and awareness layer design note](architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md). If you want to build a way, read the [hooks-and-ways guide](hooks-and-ways/README.md). This document sits one level above all of those: it tells the story of how the system composes. ## The problem this system addresses @@ -44,7 +44,7 @@ flowchart LR Sensors["sensors<br/>(file, git, context, peers)"]:::cheap Attend["attend<br/>(salience, insistence, state)"]:::cheap Matcher["ways matcher<br/>(embedding)"]:::cheap - Gate["disclosure gate<br/>(ADR-104 habituation)"]:::cheap + Gate["disclosure gate<br/>(ADR-123 habituation)"]:::cheap Ledger["ledger writer"]:::cheap Memory["memory projection<br/>(optional, e.g. KG)"]:::cheap end @@ -113,7 +113,7 @@ The naive approach would be: at session start, inject every way that's possibly 2. **Attention dilution.** Claude reads what's nearest the current conversation most carefully. Guidance injected at startup becomes guidance buried under forty turns of later content. See [hooks-and-ways/context-decay.md](hooks-and-ways/context-decay.md) for the formal model. 3. **Habituation.** If every way fires on every possible trigger, Claude's context becomes a soup of guidance that doesn't map to what's happening right now. -**Progressive disclosure** ([ADR-105](architecture/system/ADR-105-progressive-disclosure-for-way-trees.md)) and **token-gated re-disclosure** ([ADR-104](architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md)) address this together. The rules: +**Progressive disclosure** ([ADR-105](architecture/ways/ADR-105-progressive-disclosure-for-way-trees.md)) and **token-gated re-disclosure** ([ADR-123](architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md)) address this together. The rules: - Ways only fire when their triggers match the current situation, never speculatively - Once a way has fired, it is marked as "disclosed" and will not fire again until its re-disclosure cooldown expires @@ -126,7 +126,7 @@ The mental model: **ways are a rate-limited stream of premises**. Claude does no ## Empirical tuning: letting telemetry revise the thresholds -The thresholds, half-lives, and vocabularies above were all set by authorial judgment. **Empirical auto-tuning** ([ADR-134](architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md)) closes the loop by feeding the matcher's own firing record back into those settings. This is optional operational tooling, not part of the per-turn loop — but it is what keeps the precision-first discipline honest over time. +The thresholds, half-lives, and vocabularies above were all set by authorial judgment. **Empirical auto-tuning** ([ADR-134](architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md)) closes the loop by feeding the matcher's own firing record back into those settings. This is optional operational tooling, not part of the per-turn loop — but it is what keeps the precision-first discipline honest over time. Two signals accumulate in `$XDG_STATE/agent-ways/events.jsonl` (the legacy `~/.claude/stats/events.jsonl` is migrated forward), both written by the cheap substrate at fire time: @@ -143,11 +143,11 @@ Both report by default. Vocabulary is never auto-applied — it re-shapes the em Ways handle the *current* session. Memory handles what persists across sessions. -The **session ledger** ([ADR-112](architecture/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) Tier 1) is a durable, chronological stream of epoch reflections. An epoch is a window of work between context-threshold boundaries: at roughly 30% context, Claude writes what it's orienting toward; at 50%, what has changed; at 70%, what has consolidated; at pre-compaction, a handoff note. The reflection way fires at these thresholds and Claude writes a short prose reflection. The Stop hook captures the prose and appends it to the ledger as one entry. +The **session ledger** ([ADR-112](architecture/archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) Tier 1) is a durable, chronological stream of epoch reflections. An epoch is a window of work between context-threshold boundaries: at roughly 30% context, Claude writes what it's orienting toward; at 50%, what has changed; at 70%, what has consolidated; at pre-compaction, a handoff note. The reflection way fires at these thresholds and Claude writes a short prose reflection. The Stop hook captures the prose and appends it to the ledger as one entry. <!-- adr-cite-ignore --> The ledger is **what was understood**, not **what was observed**. Entries are small, hand-curated, high-signal. One ledger per project; sessions contribute to the same chronological stream. When a new session starts, the ledger is the project's history — Claude reads recent entries to orient, and the ledger represents lived experience of the project across all sessions. -**Memory projections** (ADR-112 Tier 2, optional) are a second layer built on top of the ledger. The most developed example is knowledge-graph ingestion: ledger entries are copied via FUSE mount into a KG, which extracts concepts, deduplicates them against prior sessions, and builds associative structure. When a later session enters a new domain, the KG can surface concepts Claude learned in an earlier session that are relevant to the current work. +**Memory projections** (ADR-112 Tier 2, optional) are a second layer built on top of the ledger. The most developed example is knowledge-graph ingestion: ledger entries are copied via FUSE mount into a KG, which extracts concepts, deduplicates them against prior sessions, and builds associative structure. When a later session enters a new domain, the KG can surface concepts Claude learned in an earlier session that are relevant to the current work. <!-- adr-cite-ignore --> Memory projections are **configurable and never required**. The KG is one example; a user could attach a different memory tool with the same general shape, or none at all. The system is memory-tool-agnostic: it works with whatever projection (or none) is configured, and the ledger is the stable foundation underneath. @@ -161,7 +161,7 @@ The pieces described so far are *reactive*: they respond to things Claude is doi But some things happen *outside* Claude's loop. A background build finishes. A peer Claude Code session modifies a file Claude is editing. Context pressure approaches a critical threshold five turns from now. These are events Claude cannot observe without burning reasoning tokens to check, and that the hook system cannot surface because they do not correspond to Claude's own actions. -The **awareness layer** ([ADR-113](architecture/system/ADR-113-attend-active-awareness-module.md), [ADR-114](architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md)) closes this gap. It has two components: +The **awareness layer** ([ADR-113](architecture/attend/ADR-113-attend-active-awareness-module.md), [ADR-114](architecture/ways/ADR-114-attend-as-insistent-way-trigger-type.md)) closes this gap. It has two components: 1. **`attend`** — a background Rust binary that observes Claude's session state and environment via small sensor scripts, tracks approaching mechanical consequences using turn-based arithmetic, and emits single-line observations when something is worth surfacing. 2. **`Monitor`** — Claude Code's async-notification tool that delivers background-script stdout as notifications in Claude's chat. @@ -209,7 +209,7 @@ sequenceDiagram M-->>C: notification delivered async C->>C: read, recognize stakes, decide to engage C->>W: ways show attend/<signal> - W->>W: matcher + ADR-104 disclosure gate + W->>W: matcher + ADR-123 disclosure gate W-->>C: way body injected C->>C: integrate guidance, act end @@ -258,7 +258,7 @@ Stage by stage: - **Wake.** A new session begins. Project pulse surfaces recent ledger entries to orient Claude. `attend` is invoked via `Monitor` at session start and restores its prior state from disk. Claude reads the orientation context and begins working. - **Perception.** `attend` runs its sensors in the background, watching Claude's context state, workspace files, peer sessions, and approaching consequences. Most observations are silent; only state transitions worth surfacing reach stdout. - **Delivery.** Two paths operate in parallel. `Monitor` delivers `attend`'s stdout lines as asynchronous notifications. Hooks deliver synchronous way injections at event boundaries (`UserPromptSubmit`, `PreToolUse`, etc.). Both paths land on Claude's attention surface. -- **Attention.** The disclosure gate ([ADR-104](architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md)) applies habituation rules. Recently-disclosed ways are suppressed or re-surfaced tersely. Fresh signals get full weight. The cheap substrate decides what reaches Claude's reasoning in what form. +- **Attention.** The disclosure gate ([ADR-123](architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md)) applies habituation rules. Recently-disclosed ways are suppressed or re-surfaced tersely. Fresh signals get full weight. The cheap substrate decides what reaches Claude's reasoning in what form. - **Reasoning.** Claude integrates observations and guidance into its working model and decides what to do. - **Action.** Claude acts — edits files, runs tools, responds to the user. - **Capture.** At context-threshold boundaries, the reflection way fires. Claude writes a short prose reflection. The Stop hook captures it and appends to the ledger. If a memory projection is configured, the entry is also handed off to it. Separately, every fire and near-miss this turn is logged as cheap telemetry; offline, that record drives the empirical tuning of thresholds and half-lives (ADR-134) without touching the loop. @@ -276,7 +276,7 @@ Worth naming explicitly, because the architecture can be misread if these aren't - **Not consciousness.** The substrate is text replay through an inference model. The composition is novel; the substrate is not. agent-ways does not claim or produce sentience. Any language about "presence" or "continuity" in the design note refers to structural properties of the composition, not metaphysical claims about the substrate. - **Not surveillance.** The awareness layer's scope of observation never exceeds the session that owns it. All observations are local. Sensors emit metadata, not content (a presence sensor might emit "user at desk," never a camera frame). The person observed and the person the observations serve are the same person — mirror, not camera. - **Not required.** `attend` is opt-in. Ways with `trigger.type: attend` are dormant when `attend` is not running. The baseline Claude Code experience is unchanged if you do not install the awareness layer. `ways` itself is additive too — Claude Code works without it. agent-ways is a composition you can opt into at whatever depth makes sense for your workflow. -- **Not C2.** Despite superficial resemblance to command-and-control patterns, the architecture points *inward*, not outward. One session, one user, one machine. No inter-instance protocols. No central servers. [ADR-101](architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) and [ADR-102](architecture/system/ADR-102-irc-based-local-agent-communication.md) tried outward-facing designs and were abandoned for good reasons; the awareness layer points the other direction. +- **Not C2.** Despite superficial resemblance to command-and-control patterns, the architecture points *inward*, not outward. One session, one user, one machine. No inter-instance protocols. No central servers. [ADR-101](architecture/attend/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) and [ADR-102](architecture/attend/ADR-102-irc-based-local-agent-communication.md) tried outward-facing designs and were abandoned for good reasons; the awareness layer points the other direction. <!-- adr-cite-ignore --> - **Not automatic guidance injection at the awareness layer.** Even at critical salience, `attend` does not inject ways directly. It suggests affordances; Claude decides whether to invoke them. The "Claude retains agency" invariant is load-bearing. ## Where to dig deeper @@ -284,18 +284,18 @@ Worth naming explicitly, because the architecture can be misread if these aren't Ordered roughly by how specific the topic is to your interest: **If you want the theoretical framing:** -- [Design note: cognitive loop and awareness layer](design-notes/cognitive-loop-and-awareness-layer.md) — reads the system as an active-inference loop and names the invariants the ADRs preserve +- [Design note: cognitive loop and awareness layer](architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) — reads the system as an active-inference loop and names the invariants the ADRs preserve - [hooks-and-ways/rationale.md](hooks-and-ways/rationale.md) — the rationale for the ways system - [hooks-and-ways/context-decay.md](hooks-and-ways/context-decay.md) — the attention-decay model underlying progressive disclosure **If you want to understand specific decisions:** -- [ADR-104](architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) — token-gated re-disclosure -- [ADR-105](architecture/system/ADR-105-progressive-disclosure-for-way-trees.md) — progressive disclosure for way trees -- [ADR-108](architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) — embedding-based way matching -- [ADR-112](architecture/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — session ledger and optional KG integration -- [ADR-113](architecture/system/ADR-113-attend-active-awareness-module.md) — the `attend` binary -- [ADR-114](architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md) — the way trigger schema for `attend` signals -- [ADR-134](architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) — empirical auto-tuning from fire and near-miss telemetry +- [ADR-123](architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) — firing dynamics, including token-gated re-disclosure +- [ADR-105](architecture/ways/ADR-105-progressive-disclosure-for-way-trees.md) — progressive disclosure for way trees +- [ADR-108](architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) — embedding-based way matching +- [ADR-112](architecture/archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) — session ledger and optional KG integration (archived) <!-- adr-cite-ignore --> +- [ADR-113](architecture/attend/ADR-113-attend-active-awareness-module.md) — the `attend` binary +- [ADR-114](architecture/ways/ADR-114-attend-as-insistent-way-trigger-type.md) — the way trigger schema for `attend` signals +- [ADR-134](architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) — empirical auto-tuning from fire and near-miss telemetry **If you want to build ways or operate the system:** - [hooks-and-ways/README.md](hooks-and-ways/README.md) — start here for way authoring @@ -312,6 +312,6 @@ Ordered roughly by how specific the topic is to your interest: **If you are new to the whole thing and want the shortest reading path:** 1. This document 2. [hooks-and-ways/README.md](hooks-and-ways/README.md) -3. [design-notes/cognitive-loop-and-awareness-layer.md](design-notes/cognitive-loop-and-awareness-layer.md) +3. [architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md](architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md) Everything else is there when you need it. diff --git a/docs/design-notes/README.md b/docs/design-notes/README.md deleted file mode 100644 index da9c691e..00000000 --- a/docs/design-notes/README.md +++ /dev/null @@ -1,42 +0,0 @@ -# Design Notes - -Design notes are prose-first architectural framing documents. They differ from ADRs in one key way: **ADRs record decisions; design notes record readings of the system**. - -Design notes are reader-facing prose, so they follow the project's register rules: project-coined terms are introduced through the established concepts they implement (see [vocabulary.md](../vocabulary.md)). - -## When to write a design note - -Write a design note when you need to capture a framework, principle, or north star that: - -- Justifies multiple related decisions rather than being one itself -- Requires prose exposition to make the implied decisions comprehensible -- Establishes vocabulary or conceptual language that future ADRs will cite -- Reads the system as a whole rather than deciding a specific tradeoff - -Write an ADR instead when you have a specific decision with identifiable alternatives and tradeoffs. - -## Difference from ADRs - -| | ADR | Design Note | -|---|---|---| -| Form | Structured (context, decision, consequences, alternatives) | Prose, as long as needed | -| Lifecycle | Status: Draft → Accepted → Deprecated/Superseded | No status lifecycle | -| Answers | "Why did we choose X over Y?" | "How should we read this part of the system?" | -| Cites | Other ADRs | Other ADRs, design notes, external references | -| Cited by | Other ADRs, code comments | ADRs, other design notes | - -## Relationship to ADRs - -Design notes and ADRs complement each other. A design note establishes a framing; the ADRs it motivates cite the note for their Context sections. This keeps ADRs focused on their specific decision while the broader framing lives in one stable place. - -ADRs can reference design notes as foundational context. Design notes reference ADRs when they describe the system as currently decided. A design note does not supersede, deprecate, or block ADRs — it only provides reading-level context. If the reading turns out to be wrong, the note is updated or retired, and affected ADRs are revisited. - -## Index - -- [Cognitive Loop and the Awareness Layer](./cognitive-loop-and-awareness-layer.md) — reading of the system's cognitive architecture: turn-based temporal accounting, substrate separation, the awareness layer, insistence as informational pressure, agency preservation -- [Attend: Messaging Disclosure with Token-Gated Reheat](./attend-messaging-disclosure-reheat.md) — how attend re-teaches its own messaging affordances over a session's lifetime: spaced repetition over token distance, delivered through the existing sensor pipeline -- [Autonomy as a Layered System: Goals, Signposts, and the Initiation Pattern](./autonomy-goals-signposts-initiation.md) — reading the autonomy surface: the built-in goal loop as primitive, the initiation pattern, signpost events, the substrate ladder, the gate taxonomy, and the three-layer safety stack -- [The Lexical Gate as a Conditional Threshold](./lexical-gate-as-conditional-threshold.md) — reading the ADR-155 keyword gate as a per-keyword detector whose operating point (`b = γ·T`) is placed by a single per-way threshold that also governs the semantic lane; the separability geometry behind the pattern-hygiene remedies, the one coupling the control surface cannot yet break, and the relationship to log-odds lexical–semantic fusion -- [Cypress Survey](./cypress-survey.md) — reading of the CYPRESS seed against the ways corpus: same node-routing idea with the injection inverted, a heavy mandatory layer we skip, and a verdict table of the method content worth adopting (recovery, gates, integrate-don't-patch, bounded execution, a skeptic agent) with each row tracked as an issue -- [The Tool-Use Channel is a Signal Problem](./tool-use-channel-lookbehind-chunk-matching.md) — why the bash channel over- and mis-fires (a command is an action, not intent), and the empirical case that the remedy is better *signal* (transcript lookbehind, sentence-chunked matching, peak ranking, softmax-share gating, and a winner-only body cross-similarity confirmation) rather than a better scorer; plus the two orthogonal leaks (foreign-project pollution, self-reference) the signal fix does not touch -- [Attend Envelope Fields: Sender Kind, Principal, and Addressee](./attend-envelope-fields.md) — the structured `from_kind`, `from_id`, `on_behalf_of`, and `to` fields on every signal (issues #532, #533): who writes them, the ADR-171 rule that decides sender kind, the tagged on-disk form and its legacy derivation, rendering on both conduits and the transcript `origin` mapping, and what the party-aware drain (#535) reads from them diff --git a/docs/development.md b/docs/development.md index 4c86ac39..defe0f80 100644 --- a/docs/development.md +++ b/docs/development.md @@ -5,7 +5,7 @@ > thin **projection** of an XDG application, and the app source lives in > `$XDG_DATA_HOME/agent-ways`. So development now starts from a **separate checkout**, > and you *choose* when your changes reach your install — they no longer leak in by -> default. (Background: [ADR-142](architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md).) +> default. (Background: [ADR-142](architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md).) ## The three roles that used to be one directory @@ -73,7 +73,7 @@ developing — **don't** put it ahead of your installed `ways` on `PATH` unless ## See also -- [ADR-142](architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md) — the XDG application distribution (why dev changed) -- [ADR-143](architecture/system/ADR-143-three-root-way-runtime-core-user-project.md) — core / user / project way roots -- [ADR-144](architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) — the reconciler, migrator, and deprecation lifecycle +- [ADR-142](architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md) — the XDG application distribution (why dev changed) +- [ADR-143](architecture/practice/ADR-143-three-root-way-runtime-core-user-project.md) — core / user / project way roots +- [ADR-144](architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) — the reconciler, migrator, and deprecation lifecycle - `CONTRIBUTING.md` — contribution norms and the security bar for changes diff --git a/docs/explanation/attend-messaging/00-overview.md b/docs/explanation/attend-messaging/00-overview.md index 7464ea36..1f7d9e83 100644 --- a/docs/explanation/attend-messaging/00-overview.md +++ b/docs/explanation/attend-messaging/00-overview.md @@ -1,6 +1,6 @@ --- id: 01.001.E -domain: system +domain: ways mode: explanation related: - "[[ADR-136]]" diff --git a/docs/explanation/attend-messaging/01-the-cast.md b/docs/explanation/attend-messaging/01-the-cast.md index e208c895..9b144531 100644 --- a/docs/explanation/attend-messaging/01-the-cast.md +++ b/docs/explanation/attend-messaging/01-the-cast.md @@ -1,6 +1,6 @@ --- id: 01.002.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/02-divide-and-conquer.md b/docs/explanation/attend-messaging/02-divide-and-conquer.md index cfb79db7..b993a253 100644 --- a/docs/explanation/attend-messaging/02-divide-and-conquer.md +++ b/docs/explanation/attend-messaging/02-divide-and-conquer.md @@ -1,6 +1,6 @@ --- id: 01.003.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/03-team-inside-one-voice.md b/docs/explanation/attend-messaging/03-team-inside-one-voice.md index 0e7e332e..6ed1eb31 100644 --- a/docs/explanation/attend-messaging/03-team-inside-one-voice.md +++ b/docs/explanation/attend-messaging/03-team-inside-one-voice.md @@ -1,6 +1,6 @@ --- id: 01.004.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/04-different-dimensions.md b/docs/explanation/attend-messaging/04-different-dimensions.md index ef807e47..809b1d0d 100644 --- a/docs/explanation/attend-messaging/04-different-dimensions.md +++ b/docs/explanation/attend-messaging/04-different-dimensions.md @@ -1,6 +1,6 @@ --- id: 01.005.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/05-the-crowd.md b/docs/explanation/attend-messaging/05-the-crowd.md index 26069c7d..7999013a 100644 --- a/docs/explanation/attend-messaging/05-the-crowd.md +++ b/docs/explanation/attend-messaging/05-the-crowd.md @@ -1,6 +1,6 @@ --- id: 01.006.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/06-human-on-the-surface.md b/docs/explanation/attend-messaging/06-human-on-the-surface.md index 58c36d79..58db9fea 100644 --- a/docs/explanation/attend-messaging/06-human-on-the-surface.md +++ b/docs/explanation/attend-messaging/06-human-on-the-surface.md @@ -1,6 +1,6 @@ --- id: 01.007.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/attend-messaging/07-the-lane-gate.md b/docs/explanation/attend-messaging/07-the-lane-gate.md index f05660b2..26c6c806 100644 --- a/docs/explanation/attend-messaging/07-the-lane-gate.md +++ b/docs/explanation/attend-messaging/07-the-lane-gate.md @@ -1,6 +1,6 @@ --- id: 01.008.E -domain: system +domain: ways mode: explanation related: - "[[01.001.E]]" diff --git a/docs/explanation/how-ways-works/how-ways-works-the-model.md b/docs/explanation/how-ways-works/how-ways-works-the-model.md index 9eb6dccc..81e1a505 100644 --- a/docs/explanation/how-ways-works/how-ways-works-the-model.md +++ b/docs/explanation/how-ways-works/how-ways-works-the-model.md @@ -1,10 +1,10 @@ --- id: 01.017.E -domain: system +domain: ways mode: explanation related: - - "[[ADR-104]]" - "[[ADR-123]]" + - "[[ADR-126]]" - "[[ADR-134]]" - "[[01.018.E]]" - "[[01.019.E]]" @@ -103,7 +103,7 @@ physically doing. **Re-disclosure — habituation.** Once a way has fired, it is marked disclosed and won't fire again until its cooldown — measured in *tokens of context -consumed*, not turns or wall-clock — expires ([[ADR-104]]). When the trigger +consumed*, not turns or wall-clock — expires ([[ADR-123]], [[ADR-126]]). When the trigger recurs after the cooldown, the way re-surfaces fresh as a `way_redisclosed` event. This is the mechanism that keeps a long session from either drowning in repeated guidance or silently losing premises it surfaced eighty turns ago. The diff --git a/docs/explanation/how-ways-works/reading-the-session-data.md b/docs/explanation/how-ways-works/reading-the-session-data.md index 39585180..dddc1833 100644 --- a/docs/explanation/how-ways-works/reading-the-session-data.md +++ b/docs/explanation/how-ways-works/reading-the-session-data.md @@ -1,9 +1,9 @@ --- id: 01.019.E -domain: system +domain: ways mode: explanation related: - - "[[ADR-112]]" + - "[[ADR-123]]" - "[[ADR-134]]" - "[[01.017.E]]" - "[[01.018.E]]" @@ -124,7 +124,7 @@ A few readings that turn raw fields into judgement: - **Re-disclosures ≫ first-fires** is the signature of a *long* session — premises refreshed across context pressure and compaction, exactly as - [[ADR-104]] intends. A session that's nearly all first-fires was short. + [[ADR-123]] and [[ADR-126]] intend. A session that's nearly all first-fires was short. - **A trigger mix dominated by `semantic`** means the session was steered by *meaning* — Claude's intent matched ways without anyone having anticipated the keyword. A mix dominated by `bash`/`file` means it was steered by what Claude @@ -146,7 +146,7 @@ A few readings that turn raw fields into judgement: The event log is the *telemetry* layer — fine-grained, per-fire, recent (it tail-compacts past ~32 MiB, so it forgets its oldest tail). It is not the durable -memory of the project; that's the **session ledger** ([[ADR-112]]), which records +memory of the project; that's the **session ledger** ([[ADR-112]]), which records <!-- adr-cite-ignore --> *what was understood* rather than *what fired*. The two are complementary: the ledger is the journal, the event log is the instrument trace. This cluster is about the instrument trace — for the journal and the rest of the architecture, diff --git a/docs/explanation/how-ways-works/scenario-a-long-session.md b/docs/explanation/how-ways-works/scenario-a-long-session.md index 9718a4c1..73eaf5ef 100644 --- a/docs/explanation/how-ways-works/scenario-a-long-session.md +++ b/docs/explanation/how-ways-works/scenario-a-long-session.md @@ -1,9 +1,9 @@ --- id: 01.018.E -domain: system +domain: ways mode: explanation related: - - "[[ADR-104]]" + - "[[ADR-123]]" - "[[ADR-134]]" - "[[01.017.E]]" - "[[01.019.E]]" @@ -102,7 +102,7 @@ met where it is. `softwaredev/delivery/branching` fired once and then **re-disclosed 19 times**. Each refresh happened after roughly 100K tokens of context had passed since the -last one — its token-gated cooldown ([[ADR-104]]). But the token positions where +last one — its token-gated cooldown ([[ADR-126]]). But the token positions where it re-disclosed aren't a clean rising line; they sawtooth: ``` diff --git a/docs/explanation/install-topologies/scenario-the-in-place-install.md b/docs/explanation/install-topologies/scenario-the-in-place-install.md index 9b1b722f..bbd39cf3 100644 --- a/docs/explanation/install-topologies/scenario-the-in-place-install.md +++ b/docs/explanation/install-topologies/scenario-the-in-place-install.md @@ -1,6 +1,6 @@ --- id: 01.015.E -domain: system +domain: ways mode: explanation related: - "[[ADR-140]]" diff --git a/docs/explanation/install-topologies/scenario-the-subdirectory-install.md b/docs/explanation/install-topologies/scenario-the-subdirectory-install.md index 2a20be84..8d991ebd 100644 --- a/docs/explanation/install-topologies/scenario-the-subdirectory-install.md +++ b/docs/explanation/install-topologies/scenario-the-subdirectory-install.md @@ -1,6 +1,6 @@ --- id: 01.016.E -domain: system +domain: ways mode: explanation related: - "[[ADR-140]]" diff --git a/docs/explanation/install-topologies/the-two-install-topologies-the-model.md b/docs/explanation/install-topologies/the-two-install-topologies-the-model.md index 292c1ecc..c6a832ee 100644 --- a/docs/explanation/install-topologies/the-two-install-topologies-the-model.md +++ b/docs/explanation/install-topologies/the-two-install-topologies-the-model.md @@ -1,6 +1,6 @@ --- id: 01.014.E -domain: system +domain: ways mode: explanation related: - "[[ADR-140]]" diff --git a/docs/explanation/localization/adopter-localization-the-model.md b/docs/explanation/localization/adopter-localization-the-model.md index 5f485aa0..947133d9 100644 --- a/docs/explanation/localization/adopter-localization-the-model.md +++ b/docs/explanation/localization/adopter-localization-the-model.md @@ -1,6 +1,6 @@ --- id: 01.009.E -domain: system +domain: ways mode: explanation related: - "[[ADR-139]]" diff --git a/docs/explanation/localization/scenario-steady-state-authoring.md b/docs/explanation/localization/scenario-steady-state-authoring.md index 9c88cefb..cc5c81fe 100644 --- a/docs/explanation/localization/scenario-steady-state-authoring.md +++ b/docs/explanation/localization/scenario-steady-state-authoring.md @@ -1,6 +1,6 @@ --- id: 01.012.E -domain: system +domain: ways mode: explanation related: - "[[01.009.E]]" diff --git a/docs/explanation/localization/scenario-the-english-native-install.md b/docs/explanation/localization/scenario-the-english-native-install.md index b66f26fc..dd76641d 100644 --- a/docs/explanation/localization/scenario-the-english-native-install.md +++ b/docs/explanation/localization/scenario-the-english-native-install.md @@ -1,6 +1,6 @@ --- id: 01.010.E -domain: system +domain: ways mode: explanation related: - "[[01.009.E]]" diff --git a/docs/explanation/localization/scenario-the-language-switch.md b/docs/explanation/localization/scenario-the-language-switch.md index 2827fd21..ee452e42 100644 --- a/docs/explanation/localization/scenario-the-language-switch.md +++ b/docs/explanation/localization/scenario-the-language-switch.md @@ -1,6 +1,6 @@ --- id: 01.011.E -domain: system +domain: ways mode: explanation related: - "[[01.009.E]]" diff --git a/docs/explanation/localization/the-mode-gate-mechanism-under-the-scenarios.md b/docs/explanation/localization/the-mode-gate-mechanism-under-the-scenarios.md index e4bdc05a..d9c253d1 100644 --- a/docs/explanation/localization/the-mode-gate-mechanism-under-the-scenarios.md +++ b/docs/explanation/localization/the-mode-gate-mechanism-under-the-scenarios.md @@ -1,6 +1,6 @@ --- id: 01.013.E -domain: system +domain: ways mode: explanation related: - "[[01.009.E]]" diff --git a/docs/hooks-and-ways/authoring-docs-style.md b/docs/hooks-and-ways/authoring-docs-style.md index 21806c56..4400f702 100644 --- a/docs/hooks-and-ways/authoring-docs-style.md +++ b/docs/hooks-and-ways/authoring-docs-style.md @@ -107,9 +107,9 @@ prose: - **`ways tune`** — the *locale* alias audit (fidelity / discrimination vs the English root anchor, ADR-139/125). It never wrote relevance thresholds; it fixes stub quality by re-authoring. -- **Salience / signal decay** — turn-based exponential decay (ADR-121), a model +- **Salience / signal decay** — exponential decay over a progression axis (ADR-123), a model separate from relevance scoring. -- **Progressive disclosure and token-gated re-fire** (ADR-104/105/126), the +- **Progressive disclosure and token-gated re-fire** (ADR-105/123/126), the three-root runtime (ADR-143), sentence-salience input reduction (ADR-130), the authored disclosure graph and removal of BM25 (ADR-125), and the two embedding models (EN + multilingual for localized mode). diff --git a/docs/hooks-and-ways/context-decay-formal-foundations.md b/docs/hooks-and-ways/context-decay-formal-foundations.md index 1414adaf..c665d832 100644 --- a/docs/hooks-and-ways/context-decay-formal-foundations.md +++ b/docs/hooks-and-ways/context-decay-formal-foundations.md @@ -2,7 +2,7 @@ **A companion to [The Context Decay Model: Why Timed Injection Beats Front-Loading](context-decay.md)** -**See also:** [ADR-123: Firing dynamics — progression-axis unification](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md) — the architecture that operationalizes this model for ways and attend, including why token position (not turn count, not wall clock) is the correct progression axis for transformer-hosted ways. +**See also:** [ADR-123: Firing dynamics — progression-axis unification](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) — the architecture that operationalizes this model for ways and attend, including why token position (not turn count, not wall clock) is the correct progression axis for transformer-hosted ways. > **Read section 1.1a before citing this document.** The RoPE decay derivation below is the *baseline positional prior*, not a complete model of trained-attention retrieval behavior. Modern LLMs override the baseline for salient content via head specialization, and empirical retention is substantially better than the baseline curve predicts. The decay model is used here as an approximation of *aggregate presentation economics*, which is the right quantity for firing decisions — but it is not a direct model of attention internals. diff --git a/docs/hooks-and-ways/context-decay.md b/docs/hooks-and-ways/context-decay.md index 176d219e..62596414 100644 --- a/docs/hooks-and-ways/context-decay.md +++ b/docs/hooks-and-ways/context-decay.md @@ -1,6 +1,6 @@ # The Context Decay Model: Why Timed Injection Beats Front-Loading -Ways are a progressive disclosure system for LLM context. The phenomenon this document models is the forgetting curve applied to injected guidance: instructions lose effective influence as context accumulates past them. The system's answer is re-disclosure on a spaced schedule — spaced repetition for a model that cannot internalize, with token distance playing the role time plays in human memory. This document explains why timed injection outperforms monolithic system prompts by modeling the *presentation economics* of long-context inference: how guidance retains or loses its effective influence on generation as context accumulates. A companion document, [Formal Foundations](context-decay-formal-foundations.md), grounds each claim here in published transformer research, control theory, and human operator modeling. The architectural decisions that operationalize this model for firing dynamics live in [ADR-123: Firing dynamics — progression-axis unification for attend and ways](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md). +Ways are a progressive disclosure system for LLM context. The phenomenon this document models is the forgetting curve applied to injected guidance: instructions lose effective influence as context accumulates past them. The system's answer is re-disclosure on a spaced schedule — spaced repetition for a model that cannot internalize, with token distance playing the role time plays in human memory. This document explains why timed injection outperforms monolithic system prompts by modeling the *presentation economics* of long-context inference: how guidance retains or loses its effective influence on generation as context accumulates. A companion document, [Formal Foundations](context-decay-formal-foundations.md), grounds each claim here in published transformer research, control theory, and human operator modeling. The architectural decisions that operationalize this model for firing dynamics live in [ADR-123: Firing dynamics — progression-axis unification for attend and ways](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md). ## What this model captures — and what it doesn't @@ -132,7 +132,7 @@ The system prompt provides the baseline. Ways provide the reinforcement signal t This document describes the presentation-economics model — the *how* of context decay and injection topology, framed as useful approximation rather than claims about attention internals. -For the architecture that operationalizes this model for firing dynamics across ways and attend — including the progression-axis framing, curve-as-first-class-parameter, and why ways' tick is token position while attend's is wall-clock — see [ADR-123: Firing dynamics — progression-axis unification](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md). That ADR is the canonical source for how this model gets converted into actual firing decisions. +For the architecture that operationalizes this model for firing dynamics across ways and attend — including the progression-axis framing, curve-as-first-class-parameter, and why ways' tick is token position while attend's is wall-clock — see [ADR-123: Firing dynamics — progression-axis unification](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md). That ADR is the canonical source for how this model gets converted into actual firing decisions. For the formal mathematical grounding — RoPE decay derivations, multi-layer amplification, cascade control theory, McRuer's crossover model applied to human-LLM steering, and steady-state adherence conditions — see [context-decay-formal-foundations.md](context-decay-formal-foundations.md). Note that the RoPE section there should be read alongside the "What this model captures" section above: RoPE gives the baseline positional prior, but trained attention heads can override it for salient retrieval. diff --git a/docs/hooks-and-ways/engine-reference.md b/docs/hooks-and-ways/engine-reference.md index a5b3ed25..1d37c7f6 100644 --- a/docs/hooks-and-ways/engine-reference.md +++ b/docs/hooks-and-ways/engine-reference.md @@ -85,8 +85,8 @@ removal of the multiplier. τ_k (keyword floor) is global and is **not** parent- - `ways tune` — locale alias audit (fidelity / discrimination vs the English root anchor), ADR-139/125. Never writes relevance thresholds. -- Salience / signal **decay** — ADR-121 turn-based exponential; a distinct model. -- Progressive disclosure and token-gated re-fire (ADR-104/105/126); `refire` as a +- Salience / signal **decay** — ADR-123 exponential over a progression axis; a distinct model. +- Progressive disclosure and token-gated re-fire (ADR-105/123/126); `refire` as a fraction of the context window (ADR-126); three-root runtime (ADR-143); sentence- salience input reduction (ADR-130); authored disclosure graph / no BM25 (ADR-125). diff --git a/docs/hooks-and-ways/languages.md b/docs/hooks-and-ways/languages.md index b4b51adc..b6af3ff2 100644 --- a/docs/hooks-and-ways/languages.md +++ b/docs/hooks-and-ways/languages.md @@ -5,7 +5,7 @@ Ways runs in one of two **modes**, decided by a single switch. English is the the framework translates itself, validated against the English root (ADR-139). This page is the reference; the lifecycle and rationale live in `docs/explanation/localization/` (`01.009.E`–`01.013.E`) and the design note -`docs/design-notes/adopter-localization-lifecycle-and-tuning.md`. +`docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md`. ## The two modes @@ -129,4 +129,4 @@ ways language --json # machine-readable (resolved_language, models, locales_fo - **ADR-139** — adopter-run localization (the two modes, the shelve, root-anchoring) - **ADR-125** — the coordinate-alias model (`description`+`vocabulary` as embedding-space alias) - **ADR-107** — original locale support and the dual-model approach (superseded in part) -- `docs/design-notes/adopter-localization-lifecycle-and-tuning.md` — the mechanics +- `docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md` — the mechanics diff --git a/docs/hooks-and-ways/observed-behavior.md b/docs/hooks-and-ways/observed-behavior.md index 7e44e727..15fc1af5 100644 --- a/docs/hooks-and-ways/observed-behavior.md +++ b/docs/hooks-and-ways/observed-behavior.md @@ -6,7 +6,7 @@ ## Why this file exists -The theoretical scaffolding in [`context-decay.md`](context-decay.md) and [`ADR-123`](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md) describes *why* ways should help with long-context instruction adherence. It does not, on its own, show *that* ways help. This note closes the empirical grounding gap so the theory is not resting entirely on internal plausibility. +The theoretical scaffolding in [`context-decay.md`](context-decay.md) and [`ADR-123`](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) describes *why* ways should help with long-context instruction adherence. It does not, on its own, show *that* ways help. This note closes the empirical grounding gap so the theory is not resting entirely on internal plausibility. This is a single observation, not a controlled benchmark. It's n=1, the operator is the author of the system, and there is no blind condition. Treat it as a lab notebook entry, not a study. The value is that the observation is specific enough to be testable by anyone else running the same comparison, and specific enough to discriminate between mechanistic hypotheses. @@ -54,7 +54,7 @@ The mechanism in short form: **task hierarchy preservation under redirection.** ### What the observation does not test -1. **Parameter calibration.** The observation says "ways helped"; it says nothing about whether the specific curve shapes, half-lives, or firing thresholds in the current implementation are optimal. Those are still empirical questions for [`ways tune`](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md) to answer. +1. **Parameter calibration.** The observation says "ways helped"; it says nothing about whether the specific curve shapes, half-lives, or firing thresholds in the current implementation are optimal. Those are still empirical questions for [`ways tune`](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) to answer. 2. **Which ways were load-bearing.** The full stack was active on Machine B. This observation cannot discriminate between "quality.md was load-bearing" and "github.md was load-bearing" and "it was all of them together." Ablation by individual way would be needed for that. @@ -138,7 +138,7 @@ This clears the gate for ADR-123 Draft → Accepted. ## Why this is worth keeping -This note is the only place in the project where the empirical grounding for the entire firing-dynamics scaffolding is written down rather than held in the operator's memory. Every time future-us reads [`context-decay.md`](context-decay.md) or [`ADR-123`](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md) and wonders whether the theoretical elaboration is justified, the answer should be traceable to a concrete observation with a concrete mechanism — not "I remember noticing once that it helped." +This note is the only place in the project where the empirical grounding for the entire firing-dynamics scaffolding is written down rather than held in the operator's memory. Every time future-us reads [`context-decay.md`](context-decay.md) or [`ADR-123`](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) and wonders whether the theoretical elaboration is justified, the answer should be traceable to a concrete observation with a concrete mechanism — not "I remember noticing once that it helped." It is also protection against theory-drift. If we later change the implementation in a way that would not have produced the effect observed here, this note is a pre-registered target: the new implementation should still, in principle, pass the same A/B test. If it wouldn't, that's a signal something load-bearing has been lost. @@ -146,6 +146,6 @@ It is also protection against theory-drift. If we later change the implementatio - [`context-decay.md`](context-decay.md) — the presentation-economics model this observation grounds. - [`context-decay-formal-foundations.md`](context-decay-formal-foundations.md) — the mathematical scaffolding, tempered to distinguish baseline attention prior from trained retrieval behavior. -- [`ADR-123`](../architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md) — the firing-dynamics architecture informed by this model. +- [`ADR-123`](../architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md) — the firing-dynamics architecture informed by this model. - [`model-context-decay/README.md`](../reference/model-context-decay/README.md) — the empirical retention benchmarks across Claude models. - **Convergent external work.** As of April 2026, several independent communities are describing the same underlying pattern from different angles — security ("safety heartbeat" constraint re-injection for long-running agents), prompt engineering (strategic repetition to counter the recency bias), agent research (identity stabilization failures in agent-to-agent conversation without human grounding signals), and memory systems (prune-and-decay architectures with selective top-N injection). These are separate discoveries, not one crowd citing each other. Ways sits in the same shape of the design space but earlier in the calibration cycle. diff --git a/docs/hooks-and-ways/scoring-and-testing.md b/docs/hooks-and-ways/scoring-and-testing.md index e9267fbe..1ca4a5be 100644 --- a/docs/hooks-and-ways/scoring-and-testing.md +++ b/docs/hooks-and-ways/scoring-and-testing.md @@ -63,7 +63,7 @@ This is also why scoring is done iteratively during way creation rather than aft ## The Tool -The `ways` binary includes embedding-based semantic scoring as a built-in subcommand (see [ADR-108](../architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) for the embedding engine, [ADR-111](../architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) for the consolidation, and [ADR-125](../architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) for the embedding-only decision). It scores a prompt against the entire way corpus using cosine similarity and ranks the results. +The `ways` binary includes embedding-based semantic scoring as a built-in subcommand (see [ADR-108](../architecture/ways/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) for the embedding engine, [ADR-111](../architecture/platform/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) for the consolidation, and [ADR-125](../architecture/ways/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) for the embedding-only decision). It scores a prompt against the entire way corpus using cosine similarity and ranks the results. ```bash # Score a prompt against all ways (query is positional; there is no --threshold flag) @@ -239,7 +239,7 @@ See the [ways-tests skill](/skills/ways-tests/SKILL.md) for the testing skill an ## Empirical Signals: Tuning From What Actually Fired -The worked example above tunes a way against prompts you write by hand. But once a way ships, the firing engine itself becomes the evidence. [ADR-134](../architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) extends the telemetry in `$XDG_STATE/agent-ways/events.jsonl` so that hand-tuning gets a record to revise from — three report-first signals: +The worked example above tunes a way against prompts you write by hand. But once a way ships, the firing engine itself becomes the evidence. [ADR-134](../architecture/ways/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) extends the telemetry in `$XDG_STATE/agent-ways/events.jsonl` so that hand-tuning gets a record to revise from — three report-first signals: - **Near-misses.** When a way's calibrated probability lands in the band just below the semantic bar — `τ_s − near_miss_margin ≤ g(s) < τ_s`, `near_miss_margin` default 0.05 (a live config key, `config.rs`) — but nothing fires, the matcher logs a `way_nearmiss` event. It carries `prob_en`, `prob_multi`, `tau_s`, `margin`, `trigger`, `query_tokens` (plus `way`, `corpus_id`, `domain`, `scope`, `project`, `session`) — the already-computed probabilities against the *same* global `τ_s` the fire path uses, no per-lane thresholds and no new embedding work (`scan/mod.rs`). These are the false silences the precision-first discipline can't otherwise see — a way that consistently lands just under the bar on prompts whose sessions then do that way's kind of work is a candidate to widen, the recall counterpart to the 0-FP constraint. - **Fire scores.** A `way_fired` event carries `fire_score`: the calibrated probability `g(s)` that cleared `τ_s`, recorded on first-fires only (not redisclosures, and `None`/absent for deterministic keyword fires) — `show/mod.rs`. This is the fire-score population that `tune-precision` reads and that feeds the **deferred** ADR-134 auto-tune; the `g(s)` calibration itself is fit at corpus-generation from the committed `calibration_probes.jsonl`, **not** from this stream. diff --git a/docs/install-guide.md b/docs/install-guide.md index e7369b3d..81edb26d 100644 --- a/docs/install-guide.md +++ b/docs/install-guide.md @@ -10,10 +10,10 @@ This guide is for the paths that aren't straight: an existing `~/.claude` you ca ## The 1.0 model (why there's no "clobber" anymore) -Before 1.0, this repo *was* `~/.claude/` — installing meant cloning over the directory Claude Code already used, so the installer had to detect existing files and stop rather than destroy them. **1.0 dissolves that.** `~/.claude` is now a thin **projection** of an XDG application whose source lives in `$XDG_DATA_HOME/agent-ways` (see [ADR-142](architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md)). Installing only: +Before 1.0, this repo *was* `~/.claude/` — installing meant cloning over the directory Claude Code already used, so the installer had to detect existing files and stop rather than destroy them. **1.0 dissolves that.** `~/.claude` is now a thin **projection** of an XDG application whose source lives in `$XDG_DATA_HOME/agent-ways` (see [ADR-142](architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md)). Installing only: - symlinks the projected roots (`skills/`, `agents/`, `commands/`, `hooks/ways/`, built binaries) into `~/.claude`, and -- three-way-merges its owned slices into your `settings.json`: the hooks block, its `permissions.allow` entries (its own binaries), and a `permissions.deny` secret-path baseline (`~/.ssh`, `~/.aws`, `.env`, … — [ADR-152](architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md); opt out with `secret_path_deny: false`). +- three-way-merges its owned slices into your `settings.json`: the hooks block, its `permissions.allow` entries (its own binaries), and a `permissions.deny` secret-path baseline (`~/.ssh`, `~/.aws`, `.env`, … — [ADR-152](architecture/platform/ADR-152-framework-default-secret-path-deny-baseline.md); opt out with `secret_path_deny: false`). Everything else you have in `~/.claude` (`settings.json` values you set, `.credentials.json`, `projects/`, `memory/`, `CLAUDE.md`) is **preserved by construction**, because the install never replaces your directory. There is no clobber prompt. (`scripts/` and `tools/` and the rest of the app stay in `$XDG_DATA` and are deliberately *not* projected.) @@ -21,7 +21,7 @@ The one case that needs your attention: if a projected root path (`~/.claude/ski ## Activation is separate from installation -Installing stages the app and builds the binaries. Activation is what puts the projection into a Claude Code config directory, and it is recorded as a **target** in your user config ([ADR-184](architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)). With no `targets` key the one target is `~/.claude`, enabled, which is what every install before this model behaved as; the next `ways reconcile` or `ways update` records it, and the install is explicit from then on. +Installing stages the app and builds the binaries. Activation is what puts the projection into a Claude Code config directory, and it is recorded as a **target** in your user config ([ADR-184](architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)). With no `targets` key the one target is `~/.claude`, enabled, which is what every install before this model behaved as; the next `ways reconcile` or `ways update` records it, and the install is explicit from then on. Before activating a directory, ask what it would do: diff --git a/docs/migration-1.0.md b/docs/migration-1.0.md index b5d958e2..654c620c 100644 --- a/docs/migration-1.0.md +++ b/docs/migration-1.0.md @@ -8,7 +8,7 @@ The move is performed by one gated, backup-first command — `ways migrate`. -> **The migrator was removed in 1.9.0** ([ADR-179](architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md)). +> **The migrator was removed in 1.9.0** ([ADR-179](architecture/platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md)). > It ships forever at **`ways-v1.8.3`**, the last tag that carries it. Migrating today means > building the binary from that tag and running it against your install — the steps are in > [Migrate](#migrate) below. Your current install is untouched by this; the tagged binary @@ -32,8 +32,8 @@ The headline consequences: - **Your session history is preserved in place.** `~/.claude/projects/` is Claude-Code-owned; the migrator never moves or rewrites it. - **Your `settings.json` is merged, not replaced.** A three-way merge manages only the hooks block and ways permissions; your model, theme, plugins, and credentials are left exactly as they are. -Background: [ADR-142](architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md) -(the XDG layout), [ADR-144](architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) +Background: [ADR-142](architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md) +(the XDG layout), [ADR-144](architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) (the reconciler and migrator). ## Do I need to migrate? @@ -129,7 +129,7 @@ would ship as a note here rather than as a new release of the command. | Version | `ways migrate` | |---|---| | **1.0.0 → 1.8.3** | Present in the binary. | -| **1.9.0 and later** | **Removed** ([ADR-179](architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md)). Migration runs from the `ways-v1.8.3` tag. | +| **1.9.0 and later** | **Removed** ([ADR-179](architecture/platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md)). Migration runs from the `ways-v1.8.3` tag. | The removal was slated for 1.1, deferred to 1.3, and executed at 1.9.0. Removing it took the assisted path, not the capability: the tag is immutable, so the command is always reachable. @@ -146,9 +146,9 @@ Your mental model for developing on and updating agent-ways changes with the lay ## See also -- [ADR-142](architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md) — the XDG application distribution -- [ADR-143](architecture/system/ADR-143-three-root-way-runtime-core-user-project.md) — core / user / project way roots -- [ADR-144](architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) — the reconciler, migrator, and deprecation lifecycle -- [ADR-179](architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md) — the migrator's removal, and what was kept +- [ADR-142](architecture/platform/ADR-142-agent-ways-1-0-xdg-application-distribution.md) — the XDG application distribution +- [ADR-143](architecture/practice/ADR-143-three-root-way-runtime-core-user-project.md) — core / user / project way roots +- [ADR-144](architecture/platform/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) — the reconciler, migrator, and deprecation lifecycle +- [ADR-179](architecture/platform/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md) — the migrator's removal, and what was kept - [development.md](development.md) — developing agent-ways after the 1.0 shift - [install-guide.md](install-guide.md) — installation paths (being updated for the 1.0 layout) diff --git a/docs/reference/model-context-decay/README.md b/docs/reference/model-context-decay/README.md index c03efd5c..3b82aeb8 100644 --- a/docs/reference/model-context-decay/README.md +++ b/docs/reference/model-context-decay/README.md @@ -39,7 +39,7 @@ Benchmarks from Anthropic's Claude 4.6 model card (March 2026). ### The Problem -Ways originally disclosed once per session — a marker-file rule designed for 200K context windows where the entire conversation fit within a single effective attention span. That rule is retired: the engine now re-discloses on a token-distance axis (ADR-104 → ADR-123 → ADR-126). The motivating problem is unchanged — at 1M tokens: +Ways originally disclosed once per session — a marker-file rule designed for 200K context windows where the entire conversation fit within a single effective attention span. That rule is retired: the engine now re-discloses on a token-distance axis (ADR-104 → ADR-123 → ADR-126). The motivating problem is unchanged — at 1M tokens: <!-- adr-cite-ignore --> - A way disclosed at token 50K has measurably degraded influence at token 500K - Retrieval accuracy for that disclosure drops ~15-20% (Opus) or ~30%+ (Sonnet) @@ -70,7 +70,7 @@ The interval is **window-relative and per-way**, not a single global threshold. | `normal` | 0.15 | the standard load-bearing cadence | | `frequent` | 0.05 | re-fires on each fresh occurrence of its trigger | -A way needing finer shaping declares an explicit `curve:` block instead (ADR-123); `refire:` wins when both are present. An early design proposed a flat 25%-of-window global constant (retired ADR-104); it was superseded by these per-way presets. See `docs/hooks-and-ways/engine-reference.md` and ADR-126. +A way needing finer shaping declares an explicit `curve:` block instead (ADR-123); `refire:` wins when both are present. An early design proposed a flat 25%-of-window global constant (retired ADR-104); it was superseded by these per-way presets. See `docs/hooks-and-ways/engine-reference.md` and ADR-126. <!-- adr-cite-ignore --> ### Token Budget Consideration diff --git a/docs/reference/ways-cli.md b/docs/reference/ways-cli.md index 7db90a74..15dcab6b 100644 --- a/docs/reference/ways-cli.md +++ b/docs/reference/ways-cli.md @@ -424,7 +424,7 @@ ways-audit report --json **Tells you:** Which projection roots it linked or relinked, one line each, plus a one-line summary unless `--quiet`. Stops with a non-zero exit, before touching anything, when a projected root (`skills/`, `agents/`, `commands/`, `hooks/ways/`, `hooks/check-config-updates.sh`, `bin/*`) is already a real directory or file rather than a symlink; the message lists the paths. It never deletes a real path. -With no `--dest`, it runs over every target in the user config ([ADR-184](../architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)): each enabled target is converged, and each disabled one is withdrawn, meaning our symlinks are unlinked and our hooks block and permissions are removed from its `settings.json` through the same merge base that wrote them. With no `targets` key the one target is `~/.claude`, enabled. An explicit `--dest` is a single-target run and leaves the list alone. +With no `--dest`, it runs over every target in the user config ([ADR-184](../architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)): each enabled target is converged, and each disabled one is withdrawn, meaning our symlinks are unlinked and our hooks block and permissions are removed from its `settings.json` through the same merge base that wrote them. With no `targets` key the one target is `~/.claude`, enabled. An explicit `--dest` is a single-target run and leaves the list alone. ``` ways reconcile # every target in config.yaml; default: ~/.claude @@ -472,7 +472,7 @@ ways enable itops/incident **Run from:** Anywhere. -**Tells you:** The resolved configuration as a table: language, scope, the project switch, disabled collections, matching thresholds, refire presets, and the targets. `--json` prints the stored user file as one document; `--json --effective` prints the resolved state with defaults applied ([ADR-185](../architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md)). Only the stored form is meant to be written back. +**Tells you:** The resolved configuration as a table: language, scope, the project switch, disabled collections, matching thresholds, refire presets, and the targets. `--json` prints the stored user file as one document; `--json --effective` prints the resolved state with defaults applied ([ADR-185](../architecture/platform/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md)). Only the stored form is meant to be written back. ``` ways config show @@ -486,7 +486,7 @@ ways config init # create config at XDG path if missing ### `ways config targets` -**When:** Finding out where agent-ways is active on this machine, or activating and deactivating it for a Claude Code config directory ([ADR-184](../architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)). +**When:** Finding out where agent-ways is active on this machine, or activating and deactivating it for a Claude Code config directory ([ADR-184](../architecture/platform/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md)). **Run from:** Anywhere. diff --git a/governance/README.md b/governance/README.md index 35c6323c..9b877fc6 100644 --- a/governance/README.md +++ b/governance/README.md @@ -165,6 +165,6 @@ Your compliance repo owns the policies. Your ways repo owns the guidance. This d ## Further Reading -- [ADR-005: Governance Traceability](../docs/architecture/legacy/ADR-005-governance-traceability.md) — the design decision +- [ADR-200: Compliance claims and session-derived findings](../docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md) — the decision in force, superseding the original traceability design - [Provenance documentation](../docs/hooks-and-ways/provenance.md) — the full reference - [The Cost of Bad Instructions](../docs/hooks-and-ways/rationale.md) — why this matters economically and environmentally diff --git a/governance/policies/architecture.md b/governance/policies/architecture.md index 0beb11d9..d66b1819 100644 --- a/governance/policies/architecture.md +++ b/governance/policies/architecture.md @@ -20,19 +20,19 @@ The way includes a pattern reference table (Factory, Strategy, Observer, Reposit When design discussions surface architectural trade-offs worth preserving, the way points to the ADR process for documentation. -## ADR (Architecture Decision Records) +## ADR (Agent Decision Records) -**Triggers**: Prompt mentions "ADR", "architect", "decision", "design pattern", "technical choice", "tradeoff"; editing files in `docs/adr/` +**Triggers**: Prompt mentions "ADR", "architect", "decision", "design pattern", "technical choice", "tradeoff"; editing files under `docs/architecture/` -**Macro**: Tri-state detection of ADR tooling in the project (declined, installed, available). +**Macro**: Detects the project's ADR tooling (declined, installed, available) and, when installed, the contract `adr.yaml` declares. It prints the commands, record format and lifecycle for that contract. -ADRs document the "why" behind architectural decisions. The way provides: +ADRs record the "why" behind decisions. Under the adr/v1 contract (ADR-304) a record has a kind: a decision, a spec kept current, or evidence (ADR-309). A decision carries a verb, a capability, a basis naming where it came from, and opens with a Summary. The way provides: -- **Decision template** with Status, Context, Decision, Consequences sections -- **Workflow**: debate → draft → PR → merge - **When to write one**: any decision that's hard to reverse, affects multiple components, or will confuse future readers if unexplained +- **Lifecycle**: create with `adr new`, ask the operator the decision's probes and record the answer with `adr consider`, then `adr accept`; `adr lint` checks the contract throughout +- **Legacy contract**: a project whose `adr.yaml` declares no contract keeps the adr/v0 format (Status, Context, Decision, Consequences) and the debate → draft → PR → merge workflow -The macro adapts to project state. If ADR tooling is installed, it shows the command reference. If tooling is available but not installed, it suggests setup. If the project has explicitly opted out (`.claude/no-adr-tooling`), it respects that choice and stops suggesting. +If the project has explicitly opted out (`.claude/no-adr-tooling`), the macro respects that choice and stops suggesting setup. ## API diff --git a/hooks/ways/documentation/adr-context/adr-context.md b/hooks/ways/documentation/adr-context/adr-context.md index 4996ae57..0f8606a7 100644 --- a/hooks/ways/documentation/adr-context/adr-context.md +++ b/hooks/ways/documentation/adr-context/adr-context.md @@ -9,27 +9,32 @@ refire: 0.15 <!-- epistemic: convention --> # ADR Context — Read Before You Build -Before diving into implementation, check if the project has Architecture Decision Records that inform the work. +Before diving into implementation, check if the project has Agent Decision Records that inform the work. ## Discovery -Use the ADR tool if installed (`docs/scripts/adr` or similar): +Use the ADR tool if installed (`docs/scripts/adr` or similar). Under adr/v1 (see the project's `docs/architecture/adr.yaml`), query rather than browse: ``` -adr list --group # see domains and decisions at a glance -adr view <N> # read a specific ADR +adr list --capability X # records touching a capability +adr list --kind decision # decisions only (also: spec, evidence) +adr list --field verb=cut # any frontmatter field +adr list --group-by capability +adr view <N> # read a specific record ``` -No tool? Check `docs/architecture/` for `ADR-*.md` files directly. +No `adr.yaml`, or no tool? The project is adr/v0 — check `docs/architecture/` for `ADR-*.md` files directly. ## Reading Strategy **Read selectively, not exhaustively.** -- Identify 1-3 ADRs most relevant to the current task -- Prioritize **Accepted** status — those are active decisions -- Read Context and Decision sections first; skip Alternatives unless debating a change -- Don't bulk-read the entire ADR corpus — it consumes context without payoff +- Identify 1-3 records most relevant to the current task +- Prioritize **accepted** status — those are active decisions +- On a decision record, read the **Summary** first (decided, trades away, one-way?) before the body +- Follow `supersedes`/`amends` links to the current version of a decision +- Evidence records (findings, surveys, audits) are background a decision's `basis` cites — read one when a decision points to it +- Don't bulk-read the entire corpus — it consumes context without payoff ## When to Check diff --git a/hooks/ways/documentation/adr/adr-tool b/hooks/ways/documentation/adr/adr-tool index 7caf07ba..6cf100ff 100755 --- a/hooks/ways/documentation/adr/adr-tool +++ b/hooks/ways/documentation/adr/adr-tool @@ -1,11 +1,13 @@ #!/usr/bin/env python3 """ -ADR - Architecture Decision Record CLI Tool +ADR - Agent Decision Record CLI Tool -A librarian for managing Architecture Decision Records. +A librarian for managing Agent Decision Records. Usage: adr list [--domain DOMAIN] [--status STATUS] [--group] [--archived|--all] + [--field KEY[=VALUE]] [--kind K] [--verb V] [--capability C] + [--group-by KEY] [--json] adr view <number> # View an ADR (aliases: v, show) adr new <domain> <title> adr rename <number> [new-title] [--slug SLUG] @@ -14,8 +16,19 @@ Usage: adr cite [--check] [paths...] adr accept <number> [--dry-run] adr reject|abandon <number> --reason "..." [--dry-run] + adr consider <number> --said "..." --via "..." [--operator NAME] [--covers PROBE...] + [--paraphrase] [--canary caught|missed] [--dry-run] + adr set <number> key=value|key+=value|key-=value ... [--force] [--dry-run] + adr supersede <old> --by <new> [--amends SECTION] [--dry-run] + adr enact <number> <commit> [--dry-run] + adr import scan <paths...> [--force] + adr import apply [sheets...] [--partial] [--force] [--dry-run] adr index [-y] adr domains + adr domain add <name> --range A-B --folder F [--label L] [--description D] + adr domain rename <old> <new> [--folder F] [--dry-run] + adr domain move <number> <domain> [--dry-run] + adr domain move --plan <file.yaml> [--dry-run] adr config Configuration is loaded from docs/architecture/adr.yaml @@ -26,7 +39,10 @@ Configuration is loaded from docs/architecture/adr.yaml # customize (ADR-177). import argparse +import hashlib +import json import os +import posixpath import re import subprocess import sys @@ -45,7 +61,7 @@ except ImportError: # Vendored-tool version (ADR-177). Bump when this tool changes — way macros # compare it against the installed template to tell stale from customized. -TOOL_VERSION = "2.0.0" +TOOL_VERSION = "2.1.0" # Statuses that mean "no longer in force" — used by archive and the # partial-supersession convention (ADR-303 / issue #438 option C2). @@ -91,6 +107,16 @@ def get_project_root() -> Path: return Path.cwd() +def _git(args: list, cwd: Path) -> Optional[str]: + """git's output, or None when git is missing, fails or times out.""" + try: + result = subprocess.run(['git', '-c', 'core.quotePath=false', *args], cwd=cwd, + capture_output=True, encoding='utf-8', errors='replace', timeout=10) + except (FileNotFoundError, subprocess.TimeoutExpired, OSError): + return None + return result.stdout if result.returncode == 0 else None + + def get_config_path() -> Path: """Get path to adr.yaml config file.""" return get_project_root() / 'docs' / 'architecture' / 'adr.yaml' @@ -137,6 +163,13 @@ def get_config() -> dict: _config = load_config() return _config +def reload_config() -> dict: + """Drop the cached config and read adr.yaml again, after a command + edits it.""" + global _config + _config = None + return get_config() + def repo_contract() -> str: """The contract adr.yaml declares; adr/v0 when it declares none (ADR-304).""" return str(get_config().get('contract') or 'adr/v0') @@ -463,29 +496,36 @@ def parse_text(content: str, path: Path) -> ADRInfo: info.title = match.group(2) break - # Determine domain from number range (authoritative) or folder name (fallback) - # Range takes precedence: an ADR's number definitively places it in a domain, - # even if the file is physically in a different domain's folder. - if info.number: - try: - base_num = int(info.number.split('.')[0]) + # Determine the domain. Under adr/v0 the number range is authoritative and + # the folder is the fallback: a number definitively places a record, even + # in another domain's folder. Under adr/v1 a number is only an identity, + # never renumbered (ADR-306 §6): the folder decides, and the range, which + # allocates new numbers, is the fallback for a folder no domain names. + def domain_by_range(): + if info.number: + try: + base_num = int(info.number.split('.')[0]) + except ValueError: + return None for domain, config in get_domains().items(): if config['range'][0] <= base_num <= config['range'][1]: - info.domain = domain - break - except ValueError: - pass + return domain + return None - # Fallback: determine from folder name (for unnumbered or out-of-range ADRs) - if not info.domain: + def domain_by_folder(): folder_name = path.parent.name for domain, config in get_domains().items(): folders = config['folder'] if isinstance(folders, str): folders = [folders] if folder_name in folders: - info.domain = domain - break + return domain + return None + + if repo_contract() == 'adr/v1': + info.domain = domain_by_folder() or domain_by_range() + else: + info.domain = domain_by_range() or domain_by_folder() return info @@ -592,17 +632,6 @@ def _str_list(value) -> Optional[list]: def _mapping(value) -> dict: return value if isinstance(value, dict) else {} -def _iso_date(value) -> Optional[str]: - """value as a YYYY-MM-DD string, or None. YAML may load a date as one.""" - text = str(value) if value is not None else '' - return text if re.fullmatch(r'\d{4}-\d{2}-\d{2}', text) else None - -def v1_baseline(ctx) -> tuple: - """(adoption date or None, set of baseline capability names).""" - baseline = _mapping(ctx.config.get('baseline')) - return (_iso_date(baseline.get('adopted')), - set(_str_list(baseline.get('capabilities')) or [])) - def v1_kinds(ctx) -> dict: return _mapping(ctx.config.get('kinds')) @@ -638,9 +667,6 @@ def v1_edges(schema: dict) -> dict: def v1_capabilities(ctx) -> dict: return _mapping(ctx.config.get('capabilities')) -def v1_surfaces(ctx) -> dict: - return _mapping(ctx.config.get('surfaces')) - def as_entries(value) -> list: if value is None: return [] @@ -650,17 +676,6 @@ def capability_scope(adr) -> list: """The capabilities a record covers: one name, a list, or ['*'].""" return as_entries(adr.frontmatter.get('capability')) -def covers(prior, capability: str) -> bool: - """A prior decision is on this capability when it names it, lists it, - or is scoped to '*' (ADR-304 §3).""" - scope = capability_scope(prior) - return '*' in scope or capability in scope - -def broader_than(prior, capability: str) -> bool: - """The prior covers more than this one capability: '*' or a list.""" - scope = capability_scope(prior) - return '*' in scope or len(scope) > 1 - def section_exists(target, section: str) -> bool: """A section reference matches a heading that is numbered with it ('2.', '2 ') or whose slug equals it ('priority-bands').""" @@ -694,9 +709,6 @@ def rule_v1_config_shape(ctx): for key in ('requires', 'statuses', 'sections'): if key in schema and _str_list(schema[key]) is None: bad(f"kinds.{name}.{key}: expected a list of names") - mutable = schema.get('mutable_after_accept') - if mutable is not None and mutable != 'all' and _str_list(mutable) is None: - bad(f"kinds.{name}.mutable_after_accept: expected a list of fields or 'all'") if 'edges' in schema: edges = schema['edges'] if not isinstance(edges, dict): @@ -709,44 +721,11 @@ def rule_v1_config_shape(ctx): bad("verbs: expected a list of names") if not isinstance(config.get('capabilities'), dict) or not config.get('capabilities'): bad("contract: adr/v1 declares no capabilities") - if 'surfaces' in config and not isinstance(config['surfaces'], dict): - bad("surfaces: expected a mapping of namespace to settings") - if 'baseline' in config: - baseline = config['baseline'] - names = _str_list(baseline.get('capabilities')) if isinstance(baseline, dict) else None - if names is None: - bad("baseline: expected 'adopted' (a YYYY-MM-DD date) and 'capabilities' (a list of names)") - else: - if not _iso_date(baseline.get('adopted')): - bad("baseline.adopted: expected a YYYY-MM-DD date") - if isinstance(config.get('capabilities'), dict): - for name in names: - if name not in config['capabilities']: - bad(f"baseline: '{name}' is not in the capabilities vocabulary") - -@config_rule(contract=V1) -def rule_v1_capabilities_added(ctx): - """Every capability in the vocabulary has an accepted `add` decision, except - those in `baseline`: capabilities that were active when the project - adopted the contract, which no record added. This warns while v0 records - remain, so a migrating corpus is not failed on every capability, and fails - once migration is done (ADR-304 §6, amended by ADR-305).""" - added = set() - for adr in ctx.corpus: - if (is_v1_record(adr, ctx) - and adr.frontmatter.get('verb') == 'add' - and str(adr.status or '').lower() == 'accepted'): - added.update(capability_scope(adr)) - # Archived records are never edited, so they never migrate and do not - # hold the check at warning. - migrating = any(not is_v1_record(adr, ctx) and not is_archived(adr.path) - for adr in ctx.corpus) - _, baseline = v1_baseline(ctx) - for name in v1_capabilities(ctx): - if name not in added and name not in baseline: - ctx.config_issues.append(Issue( - f"capability '{name}' has no accepted add decision", - 'warning' if migrating else 'error')) + if 'repository' in config: + names = [config['repository']] if isinstance(config['repository'], str) else config['repository'] + if not isinstance(names, list) or not names \ + or not all(isinstance(n, str) and _repo_name(n) for n in names): + bad("repository: expected host/owner/repo, or a list of them") # --- one record ------------------------------------------------------------------ @@ -806,9 +785,12 @@ def rule_v1_capability(adr, ctx): if raw in (None, '', []): return # reported by rule_v1_required_fields when the kind requires it scope = capability_scope(adr) - if adr.frontmatter.get('verb') != 'constrain': - if isinstance(raw, list): - v1_issue(adr, "capability takes one name; only a constrain decision takes a list") + verb = adr.frontmatter.get('verb') + if verb != 'constrain': + # A change may list the capabilities it alters (ADR-308); only a + # constrain may be scoped to '*'. + if isinstance(raw, list) and verb != 'change': + v1_issue(adr, "capability takes one name; only change and constrain decisions take a list") return if '*' in scope: v1_issue(adr, "only a constrain decision may be scoped to '*'") @@ -817,21 +799,8 @@ def rule_v1_capability(adr, ctx): for name in scope: if name != '*' and name not in vocabulary: v1_issue(adr, f"capability '{name}' is not in the adr.yaml vocabulary") - -@file_rule(contract=V1) -def rule_v1_retire_targets(adr, ctx): - if not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'retire': - return - targets = as_entries(adr.frontmatter.get('targets')) - if not targets: - v1_issue(adr, "a retire decision names its targets (targets: [cli:..., route:...])") - surfaces = v1_surfaces(ctx) - for target in targets: - namespace, sep, name = target.partition(':') - if not sep or not name: - v1_issue(adr, f"target '{target}' is not namespace:name") - elif namespace not in surfaces: - v1_issue(adr, f"target '{target}': surface '{namespace}' is not declared in adr.yaml") + if verb == 'change' and len(scope) > 3: + v1_issue(adr, f"a change lists {len(scope)} capabilities; list only those it alters, the rest belong in related (ADR-308 §3)", 'warning') # --- against the corpus ------------------------------------------------------------ @@ -865,72 +834,14 @@ def rule_v1_edges(adr, ctx): v1_issue(adr, f"{field_name}: ADR-{number} is a {target_kind}, expected {' or '.join(target_kinds)}") if section and not section_exists(target, section): v1_issue(adr, f"{field_name}: ADR-{number} has no section '{section}'") - -def decision_order(adr) -> tuple: - """Records in decision order: by date, then number (ADR-304 §3).""" - base, _, part = str(adr.number or '0').partition('.') - return (str(adr.date or ''), int(base or 0), int(part or 0)) - -def _stands_on_baseline(capability: str, adr, ctx) -> bool: - """A change with no prior edge stands on the baseline when the capability - is in it and either the change predates adoption, or no live decision on - the capability has been made since adoption before it. Only decisions - dated after adoption count: each was written as v1, so migrating an older - record never moves the answer (ADR-305).""" - adopted, baseline = v1_baseline(ctx) - when = _iso_date(adr.date) - if capability not in baseline or not adopted or not when: - return False - if when <= adopted: - return True - for other in ctx.corpus: - other_when = _iso_date(other.date) - if (other is not adr and is_v1_record(other, ctx) - and other.frontmatter.get('verb') not in (None, 'constrain') - and str(other.status or '').lower() not in ('rejected', 'abandoned') - and covers(other, capability) - and other_when and other_when > adopted - and decision_order(other) < decision_order(adr)): - return False - return True - -@corpus_rule(contract=V1) -def rule_v1_change_replaces(adr, ctx): - """A change decision supersedes or amends a prior decision on the same - capability. When the prior covers more than this capability ('*' or a - list), the change amends it (ADR-304 §3). A change on a baseline - capability with no prior record to name stands on the baseline (ADR-305).""" - if not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'change': - return - edges = [] - for field_name in ('supersedes', 'amends'): - for entry in as_entries(adr.frontmatter.get(field_name)): - target = ctx.by_number.get(norm_ref(entry)[0]) - if target is None: - return # a dangling edge is already reported by the edge rules - edges.append((field_name, target)) - for capability in [c for c in capability_scope(adr) if c != '*']: - v0_priors = [t for _, t in edges if not is_v1_record(t, ctx)] - fits = [(f, t) for f, t in edges if is_v1_record(t, ctx) and covers(t, capability)] - if any(f == 'amends' or not broader_than(t, capability) for f, t in fits): - continue - if fits: - v1_issue(adr, f"a change on '{capability}' against a broader decision amends it rather than superseding it") - elif not edges and _stands_on_baseline(capability, adr, ctx): - continue - elif v0_priors: - numbers = ', '.join(f"ADR-{t.number}" for t in v0_priors) - v1_issue(adr, f"cannot confirm the prior decision on '{capability}': {numbers} is still v0", 'warning') - else: - v1_issue(adr, f"a change decision supersedes or amends a prior decision on '{capability}'") # ============================================================================ # adr/v1: basis, agent, consideration and concerns (ADR-304 §11, §12) # ============================================================================ # # A decision's basis names what it rests on. precedent points at another -# record in the corpus; every other source is external. Following precedent -# must reach an external source: a corpus that justifies itself only by citing -# itself can drift anywhere and still look consistent. +# record in the corpus, and must resolve to one; every other source is +# external. Where a chain of precedent leads is for a reader to judge, not +# lint (ADR-311). # # The operator basis and `considered` are an audit trail, not a credential. # Lint checks that `said` and `via` are present. It cannot check that they are @@ -943,18 +854,12 @@ V1_CONSIDERED_KEYS = ('operator', 'said', 'via', 'paraphrase', 'covers', 'canary V1_CONCERN_KEYS = ('said', 'resolve', 'answer', 'withdrawn', 'raised') V1_ANSWER_KEYS = ('operator', 'said', 'via', 'paraphrase') V1_CANARY = ('caught', 'missed') -# A precedent is "another accepted decision" (§11). These statuses ground; -# proposed warns; anything else (rejected, abandoned) does not ground. -V1_GROUNDING_STATUSES = ('accepted', 'superseded', 'archived') def v1_basis_sources(ctx) -> tuple: """adr.yaml may rename or extend the sources. precedent is the one internal source; every other declared source is external.""" return tuple(_str_list(ctx.config.get('basis_sources')) or V1_BASIS_SOURCES) -def v1_external_sources(ctx) -> tuple: - return tuple(s for s in v1_basis_sources(ctx) if s != 'precedent') - def _text(value) -> bool: return isinstance(value, str) and value.strip() != '' @@ -974,9 +879,6 @@ def basis_entries(adr, ctx) -> list: out.append((sources[0], entry[sources[0]], entry)) return out -def has_operator_basis(adr, ctx) -> bool: - return any(source == 'operator' for source, _, _ in basis_entries(adr, ctx)) - def _unknown_keys(entry: dict, allowed: tuple) -> list: return [k for k in entry if k not in allowed] @@ -987,20 +889,36 @@ def rule_v1_basis_config(ctx): raw = ctx.config.get('basis_sources') if raw is None: return - sources = _str_list(raw) - if sources is None: + if _str_list(raw) is None: ctx.config_issues.append(Issue("basis_sources: expected a list of names", 'error')) - elif not any(s != 'precedent' for s in sources): - ctx.config_issues.append(Issue("basis_sources: declares no external source, so no chain can leave the corpus", 'error')) # --- one record ---------------------------------------------------------------- +def _check_evidence_record(adr, ctx, where: str, ref: str) -> None: + """A basis `evidence: ADR-N` that names a record must resolve to one, and + under v1 to a kind the decision's basis edge accepts other than decision + (ADR-309 §2): a record cited as evidence is evidence or a spec.""" + target = ctx.by_number.get(norm_ref(ref)[0]) + if target is None: + v1_issue(adr, f"{where}: evidence {ref} resolves to no record") + return + if not is_v1_record(target, ctx): + return + schema = v1_kind_schema(adr, ctx) or {} + allowed = [k for k in (v1_edges(schema).get('basis') or []) if k != 'decision'] + kind = target.frontmatter.get('kind') + if allowed and kind not in allowed: + v1_issue(adr, f"{where}: evidence {ref} is a {kind}; evidence cites a {' or '.join(allowed)} record") + @file_rule(contract=V1) def rule_v1_basis_shape(adr, ctx): if not is_v1_record(adr, ctx) or 'basis' not in adr.frontmatter: return basis = adr.frontmatter.get('basis') - if not isinstance(basis, list) or not basis: + schema = v1_kind_schema(adr, ctx) + if basis == [] and schema is not None and 'basis' in v1_requires(schema): + return # empty where the kind requires a basis: the requires rule reports it + if not isinstance(basis, list): v1_issue(adr, "basis: expected a list of entries, each naming one source") return allowed = v1_basis_sources(ctx) @@ -1027,6 +945,8 @@ def rule_v1_basis_shape(adr, ctx): v1_issue(adr, f"{where}: precedent takes one reference, such as ADR-101") elif not _text(value) and not (isinstance(value, int) and not isinstance(value, bool)): v1_issue(adr, f"{where}: {source} needs a reference") + elif source == 'evidence' and re.fullmatch(r'ADR-\d+(\.\d+)?', str(value).strip()): + _check_evidence_record(adr, ctx, where, str(value).strip()) continue # operator: who, the level, what was said and via which channel if not _text(value): @@ -1076,10 +996,6 @@ def rule_v1_considered(adr, ctx): v1_issue(adr, f"{where}: covers is a list of probe names") if 'canary' in entry and entry['canary'] not in V1_CANARY: v1_issue(adr, f"{where}: canary is caught or missed") - # ADR-304 §12: a decision the operator started waits for their consideration - if (has_operator_basis(adr, ctx) and str(adr.status or '').lower() == 'accepted' - and not adr.frontmatter.get('considered')): - v1_issue(adr, "accepted with an operator basis but no considered entry; the operator considers what they started") @file_rule(contract=V1) def rule_v1_concern(adr, ctx): @@ -1116,211 +1032,46 @@ def rule_v1_concern(adr, ctx): else: adr.issues.append(Issue(f"open concern: {_first_line(entry.get('said'))}", 'warning', 'open-concern')) -# --- against the corpus: grounding ------------------------------------------------- -# -# Grounding is a least fixed point, so the verdict does not depend on the order -# of basis entries and each record is visited a bounded number of times: -# -# grounded(r) = r has an external source -# or some precedent of r points at a grounding-status record t -# with grounded(t) -# -# A record with no basis but a decided_by (a spec) is grounded through the -# records that decided it. A non-archived v0 record grounds provisionally: the -# chain passes with a warning until the record is migrated. An archived v0 -# record grounds outright, since archived records never migrate. - -def _is_v0_ground(target, ctx) -> Optional[str]: - """'final' for an archived v0 record, 'provisional' for a live one.""" - if is_v1_record(target, ctx): - return None - return 'final' if is_archived(target.path) else 'provisional' - -def _grounding_map(ctx) -> dict: - """path -> 'final' | 'provisional' for every grounded v1 record.""" - cache = getattr(ctx, '_grounding', None) - if cache is not None: - return cache - records = [a for a in ctx.corpus if is_v1_record(a, ctx)] - external = v1_external_sources(ctx) - grounded = {} - for adr in records: - if any(source in external for source, _, _ in basis_entries(adr, ctx)): - grounded[adr.path] = 'final' - - def level_of(target) -> Optional[str]: - v0 = _is_v0_ground(target, ctx) - if v0: - return v0 - return grounded.get(target.path) - - changed = True - while changed: - changed = False - for adr in records: - best = grounded.get(adr.path) - if best == 'final': - continue - candidates = [] - entries = basis_entries(adr, ctx) - if entries: - for value in _precedent_values(adr, ctx): - target = ctx.by_number.get(norm_ref(value)[0]) - if target is None or str(target.status or '').lower() not in V1_GROUNDING_STATUSES + ('proposed',): - continue - candidates.append(level_of(target)) - elif 'basis' not in adr.frontmatter: - for ref in as_entries(adr.frontmatter.get('decided_by')): - target = ctx.by_number.get(norm_ref(ref)[0]) - if target is not None: - candidates.append(level_of(target)) - new = 'final' if 'final' in candidates else ('provisional' if 'provisional' in candidates else None) - if new and new != best: - grounded[adr.path] = new - changed = True - ctx._grounding = grounded - return grounded +# --- against the corpus ------------------------------------------------------------ def _precedent_values(adr, ctx) -> list: """Precedent references of a usable shape; the shape rule reports others.""" return [value for source, value, _ in basis_entries(adr, ctx) if source == 'precedent' and isinstance(value, (str, int)) and not isinstance(value, bool)] -def _precedent_targets(adr, ctx) -> list: - return [ctx.by_number.get(norm_ref(value)[0]) for value in _precedent_values(adr, ctx)] - -def _cycle_through(adr, ctx) -> Optional[list]: - """The numbers on a precedent cycle that returns to adr, if any.""" - stack = [(adr, [adr])] - visited = set() - while stack: - node, path = stack.pop() - for target in _precedent_targets(node, ctx): - if target is None or not is_v1_record(target, ctx): - continue - if target.path == adr.path: - return [n.number or n.path.name for n in path] - if target.path not in visited: - visited.add(target.path) - stack.append((target, path + [target])) - return None - @corpus_rule(contract=V1) -def rule_v1_basis_chain(adr, ctx): - """Every precedent resolves to an allowed, accepted record, and following - precedent reaches an external source (ADR-304 §11).""" +def rule_v1_precedent(adr, ctx): + """Each precedent resolves to a record of a kind the basis edge accepts.""" if not is_v1_record(adr, ctx): return - entries = basis_entries(adr, ctx) - precedents = [(value, ctx.by_number.get(norm_ref(value)[0])) - for value in _precedent_values(adr, ctx)] - schema = v1_kind_schema(adr, ctx) or {} - allowed_kinds = v1_edges(schema).get('basis') - for value, target in precedents: + allowed_kinds = v1_edges(v1_kind_schema(adr, ctx) or {}).get('basis') + for value in _precedent_values(adr, ctx): + target = ctx.by_number.get(norm_ref(value)[0]) if target is None: v1_issue(adr, f"basis: precedent '{value}' resolves to no known ADR") - continue - status = str(target.status or '').lower() - if is_v1_record(target, ctx): - kind = v1_record_kind(target) - if allowed_kinds is not None and kind not in allowed_kinds: - v1_issue(adr, f"basis: precedent ADR-{target.number} is a {kind}, expected {' or '.join(allowed_kinds)}") - if status == 'proposed': - adr.issues.append(Issue(f"basis: precedent ADR-{target.number} is still proposed", 'warning', 'precedent-proposed')) - elif status not in V1_GROUNDING_STATUSES: - v1_issue(adr, f"basis: precedent ADR-{target.number} is {status or 'without a status'}, so it grounds nothing") - elif 'basis' not in target.frontmatter and not target.frontmatter.get('decided_by'): - v1_issue(adr, f"basis: precedent ADR-{target.number} has neither a basis nor decided_by") - if not entries: - return - level = _grounding_map(ctx).get(adr.path) - if level == 'final': - return - if level == 'provisional': - v0s = sorted({t.number for _, t in precedents if t is not None and _is_v0_ground(t, ctx) == 'provisional'}) - via = f" (ADR-{', ADR-'.join(v0s)})" if v0s else '' - v1_issue(adr, f"basis: grounded only through records still on v0{via}; migrate them to confirm the chain", 'warning') - return - if not precedents: - return # no source at all: the shape rule reports the entries - cycle = _cycle_through(adr, ctx) - if cycle: - v1_issue(adr, f"basis: precedent loops back through ADR-{' -> ADR-'.join(str(n) for n in cycle)} without reaching an external source") - else: - v1_issue(adr, f"basis: following precedent never reaches an external source ({', '.join(v1_external_sources(ctx))})") + elif is_v1_record(target, ctx) and allowed_kinds is not None \ + and v1_record_kind(target) not in allowed_kinds: + v1_issue(adr, f"basis: precedent ADR-{target.number} is a {v1_record_kind(target)}, " + f"expected {' or '.join(allowed_kinds)}") # ============================================================================ -# adr/v1: legibility, enactment, the vocabulary layers, frozen decisions +# adr/v1: sections, imports, enactment, placeholders, observables # ============================================================================ # -# ADR-304 §5 (enactment), §2 (no word shared across layers), §1 and §12 (a -# decision is frozen once it leaves proposed, and opens with a Summary the -# operator can judge alone). +# Checks of one record's own shape. How records relate over time, and +# whether an accepted one changed, is read from git (ADR-311). -# L3, the derived product state, is fixed by the contract (ADR-304 §2). -V1_DERIVED_STATES = ('active', 'absent', 'present', 'gone', 'living', 'historical') V1_ENACTING_VERBS = ('cut', 'retire') -_STEM_SUFFIXES = ('ations', 'ation', 'ments', 'ment', 'ions', 'ion', 'ings', 'ing', - 'ives', 'ive', 'als', 'al', 'ors', 'or', 'ers', 'er', 'ted', 'ed', - 'ts', 'es', 't', 'e', 's') - -def _stem(word: str) -> str: - """A crude English stem: strip suffixes repeatedly and collapse a doubled - final consonant, enough to catch cut/cutting, retire/retirement, - constrain/constraint and propose/proposal.""" - word = word.lower() - changed = True - while changed: - changed = False - for suffix in _STEM_SUFFIXES: - if word.endswith(suffix) and len(word) - len(suffix) >= 3: - word = word[:-len(suffix)] - changed = True - break - if len(word) >= 4 and word[-1] == word[-2] and word[-1] not in 'aeiou': - word = word[:-1] - return word - -def _collide(a: str, b: str) -> bool: - """Same word, same stem, or one stem is a prefix of the other (4+ letters).""" - if a.lower() == b.lower(): - return True - sa, sb = _stem(a), _stem(b) - short, long_ = sorted((sa, sb), key=len) - return sa == sb or (len(short) >= 4 and long_.startswith(short)) - def _find_section(adr, name: str) -> Optional[str]: + """The heading that is this section: the name alone, or the name followed + by punctuation ("Summary: …", "Summary (draft)"). "Summary Nudge" is a + different section.""" + pattern = re.compile(rf'{re.escape(name)}(\s*[:(].*|\s+[\u2014\u2013-].*)?', re.IGNORECASE) for heading in adr.sections: - if heading.lower() == name.lower() or heading.lower().startswith(name.lower() + ' '): + if pattern.fullmatch(heading.strip()): return heading return None -# --- adr.yaml ------------------------------------------------------------------ - -@config_rule(contract=V1) -def rule_v1_vocabulary_layers(ctx): - """No word or stem appears in two layers (ADR-304 §2), so a later contract - version cannot reintroduce a verb that collides with a state.""" - lifecycle = set() - for schema in v1_kinds(ctx).values(): - if isinstance(schema, dict): - lifecycle.update(v1_lifecycle(schema)) - lifecycle = lifecycle or set(V1_LIFECYCLE) - layers = { - 'lifecycle (L1)': sorted(lifecycle), - 'verbs (L2)': sorted(v1_verbs(ctx)), - 'derived state (L3)': sorted(V1_DERIVED_STATES), - 'basis sources (L4)': sorted(v1_basis_sources(ctx)), - } - names = list(layers) - for i, a in enumerate(names): - for b in names[i + 1:]: - for word_a in layers[a]: - for word_b in layers[b]: - if _collide(word_a, word_b): - ctx.config_issues.append(Issue( - f"'{word_a}' in {a} and '{word_b}' in {b} share a stem; layers share no word", 'error')) - # --- one record ------------------------------------------------------------------ @file_rule(contract=V1) @@ -1328,35 +1079,31 @@ def rule_v1_required_sections(adr, ctx): schema = v1_kind_schema(adr, ctx) if not is_v1_record(adr, ctx) or schema is None: return + imported = 'imported' in adr.frontmatter for name in _str_list(schema.get('sections')) or []: if _find_section(adr, name) is None: - v1_issue(adr, f"a {v1_record_kind(adr)} record opens with a '## {name}' section") + if imported and name == 'Summary': + # ADR-306 §4: an imported record may gain its Summary later, + # once the imported corpus has been read together. + v1_issue(adr, "imported record has no '## Summary' yet (ADR-306 §4)", 'warning') + else: + v1_issue(adr, f"a {v1_record_kind(adr)} record opens with a '## {name}' section") @file_rule(contract=V1) -def rule_v1_summary_legibility(adr, ctx): - """ADR-304 §12: the Summary carries probes, a mix of points the agent is - confident on and points it is not, each labelled, and an inversion. - Guidance for the writer, so a gap warns.""" - if not is_v1_record(adr, ctx) or v1_record_kind(adr) is None: +def rule_v1_imported(adr, ctx): + """`imported` records where the record came from: {from, format}, and + optionally `status`, the source status as written, and `unmapped`, the + source keys with no v1 field (ADR-306 §1, §4, §7).""" + if not is_v1_record(adr, ctx) or 'imported' not in adr.frontmatter: return - heading = _find_section(adr, 'Summary') - if heading is None or not adr.frontmatter.get('verb'): - return # sections rule reports a missing Summary; specs carry no probes - text = re.sub(r'\s+', ' ', adr.section_text.get(heading, '').lower()) - missing = [] - low = re.search(r'\bnot confident\b|\blow confidence\b', text) - high = re.search(r'(?<!not )(?<!less )\bconfident\b|\bhigh confidence\b', text) - if 'probe' not in text: - missing.append('probes') - else: - if low is None: - missing.append("a probe labelled 'not confident'") - if high is None: - missing.append("a probe labelled 'confident'") - if 'inversion' not in text: - missing.append('an inversion') - if missing: - v1_issue(adr, f"Summary lacks {', '.join(missing)} (ADR-304 §12)", 'warning') + imported = adr.frontmatter.get('imported') + if not isinstance(imported, dict) or not all( + isinstance(imported.get(k), str) and imported.get(k).strip() for k in ('from', 'format')): + v1_issue(adr, "imported: expected {from: <source path>, format: <reader>}") + elif 'unmapped' in imported and not isinstance(imported['unmapped'], dict): + v1_issue(adr, "imported.unmapped: expected a mapping of source fields") + elif isinstance(imported.get('status'), (dict, list)): + v1_issue(adr, "imported.status: expected the source's status as written") @file_rule(contract=V1) def rule_v1_enacted(adr, ctx): @@ -1373,89 +1120,1125 @@ def rule_v1_enacted(adr, ctx): elif not re.fullmatch(r'[0-9a-f]{7,40}', enacted): v1_issue(adr, f"enacted: '{enacted}' is not a commit hash") -# --- frozen decisions, read from git history -------------------------------------- +@file_rule(contract=V1) +def rule_v1_no_placeholders(adr, ctx): + """A record still holding a prompt from `adr new`'s skeleton is unfinished: + a warning while proposed, an error once it has left proposed. The prompts + live in sheet.py, which assembles after this module.""" + if not is_v1_record(adr, ctx): + return + left = placeholder_lines(adr.body) + if not left: + return + level = 'warning' if str(adr.status or '').lower() == 'proposed' else 'error' + v1_issue(adr, f"{len(left)} placeholder line(s) from `adr new` still to fill, first: {left[0]}", level) -def _git(args: list, cwd: Path) -> Optional[str]: - try: - result = subprocess.run(['git', '-c', 'core.quotePath=false', *args], cwd=cwd, - capture_output=True, encoding='utf-8', errors='replace', timeout=10) - except (FileNotFoundError, subprocess.TimeoutExpired, OSError): +@file_rule(contract=V1) +def rule_v1_observable(adr, ctx): + """What should be observable when a decision holds (ADR-307): optional, a + list whose entries are plain words or mappings with keys the author chooses. + The check applies to any v1 kind that carries the field.""" + if not is_v1_record(adr, ctx) or 'observable' not in adr.frontmatter: + return + entries = adr.frontmatter.get('observable') + if entries is None or entries == []: + v1_issue(adr, "observable: empty; remove the key or add an entry") + return + if not isinstance(entries, list): + v1_issue(adr, "observable: expected a list of entries, each a line of words or a mapping") + return + for i, entry in enumerate(entries, 1): + if isinstance(entry, str) and entry.strip(): + continue + if isinstance(entry, dict) and entry: + continue + v1_issue(adr, f"observable entry {i}: expected a line of words or a non-empty mapping") +# ============================================================================ +# Path rewriting for records that change folder (ADR-306 §6) +# ============================================================================ +# +# A record's number is its identity and never changes (ADR-310). Under adr/v1 +# its folder decides its domain, so moving a record to another domain moves +# its file, and a domain renamed in place moves its folder. Either way every +# path to what moved is rewritten. A path is a relative link that resolves to +# what moved, a path written from the repo root that names it, or a URL into +# this repository at a branch that names it. Other URLs, a URL at a commit or +# a tag (a permalink), a path inside a fenced code block, a path written with +# backslashes, and words that only contain a folder's name are left alone. +# ADR-N citations are left alone too: the number still names the same record. + +# Frontmatter keys that record history and are never rewritten: an imported +# record keeps the path it was imported from. +RELOCATE_HISTORY_KEYS = ('imported',) + +_PATH_WORD_RE = re.compile(r'[A-Za-z][\w+.-]*://[\w./~%-]+|[\w./-]+') +# host/owner/repo/blob|tree|raw/<ref>/<path>, and GitHub's raw host, where +# the ref follows the repo directly. A ref may hold slashes. +_REPO_URL_RE = re.compile(r'[A-Za-z][\w+.-]*://([^/]+)/([^/]+)/([^/]+)/(?:blob|tree|raw)/(.+)') +_RAW_URL_RE = re.compile(r'[A-Za-z][\w+.-]*://raw\.githubusercontent\.com/([^/]+)/([^/]+)/(.+)', re.IGNORECASE) +_COMMIT_RE = re.compile(r'[0-9a-fA-F]{7,40}') +_FM_KEY_RE = re.compile(r'([A-Za-z_][\w-]*)\s*:') +_FENCE_RE = re.compile(r'[ \t]*(`{3,}|~{3,})') + + +def _repo_name(url: str) -> Optional[str]: + """host/owner/repo, lowercase, from a remote URL or a written name.""" + m = re.fullmatch(r'(?:[A-Za-z][\w+.-]*://)?(?:[^@/]+@)?([^/:]+)(?::\d+)?[:/]([^/]+/[^/]+?)(?:\.git)?/?', + url.strip()) + return f"{m.group(1)}/{m.group(2)}".lower() if m else None + + +def project_repos(root: Path) -> set: + """The names this repository goes by, host/owner/repo, lowercase: + adr.yaml's `repository:` (a name or a list of names) when it is set, + otherwise the origin remote. Empty without either.""" + configured = get_config().get('repository') + if configured: + values = [configured] if isinstance(configured, str) else configured if isinstance(configured, list) else [] + return {n for n in (_repo_name(str(v)) for v in values) if n} + name = _repo_name((_git(['remote', 'get-url', 'origin'], root) or '').strip()) + return {name} if name else set() + + +class Repository: + """This repository as URLs name it: its names, and its tags and branches, + each read from git once when first needed. A URL into it at a branch + names a path in the working tree. At a commit or a tag it is a + permalink to a snapshot, and names nothing that moves.""" + + def __init__(self, root: Path): + self.root = root + self.names = project_repos(root) + self._tags = self._branches = None + + @property + def tags(self) -> set: + if self._tags is None: + self._tags = set((_git(['tag', '-l'], self.root) or '').split()) + return self._tags + + @property + def branches(self) -> set: + """Local branches, and remote-tracking ones by their branch name.""" + if self._branches is None: + out = _git(['for-each-ref', '--format=%(refname)', 'refs/heads', 'refs/remotes'], self.root) or '' + names = set() + for ref in out.split(): + if ref.startswith('refs/heads/'): + names.add(ref[len('refs/heads/'):]) + elif ref.startswith('refs/remotes/') and ref.count('/') >= 3: + names.add(ref.split('/', 3)[3]) + self._branches = names + return self._branches + + def path(self, token: str) -> Optional[tuple]: + """(where the path starts in token, the path from the repo root) for + a URL into this repository at a branch; None otherwise. A ref with a + slash is read as the longest leading run of segments that names a + known branch; otherwise the ref is the first segment.""" + m = _RAW_URL_RE.fullmatch(token) + if m: + repo, rest, at = f"github.com/{m.group(1)}/{m.group(2)}", m.group(3), m.start(3) + else: + m = _REPO_URL_RE.fullmatch(token) + if not m: + return None + repo, rest, at = '/'.join(m.group(1, 2, 3)), m.group(4), m.start(4) + repo = repo.lower() + if repo.endswith('.git'): + repo = repo[:-4] + if repo not in self.names: + return None + segments = rest.split('/') + if len(segments) < 2: + return None + width = next((n for n in range(len(segments) - 1, 1, -1) + if '/'.join(segments[:n]) in self.branches), 1) + ref = '/'.join(segments[:width]) + if width == 1 and ref not in self.branches and (_COMMIT_RE.fullmatch(ref) or ref in self.tags): + return None + path = '/'.join(segments[width:]) + return (at + len(ref) + 1, path) if path else None + + +def _fenced(lines: list, start: int = 0) -> set: + """Indexes of the lines from start on that sit in a fenced code block, + fence lines included.""" + inside, fence = set(), None + for i in range(start, len(lines)): + m = _FENCE_RE.match(lines[i]) + if fence is None: + if m: + fence = m.group(1) + inside.add(i) + else: + inside.add(i) + if m and m.group(1)[0] == fence[0] and len(m.group(1)) >= len(fence) \ + and not lines[i][m.end():].strip(): + fence = None + return inside + + +def _sub_paths(line: str, swap) -> str: + """line with each run of path characters replaced by swap(token), or + left alone when it touches a backslash: a Windows path is not a + path this tool resolves.""" + def one(m): + before = m.string[m.start() - 1:m.start()] + after = m.string[m.end():m.end() + 1] + if before == '\\' or after == '\\': + return m.group(0) + return swap(m.group(0)) + return _PATH_WORD_RE.sub(one, line) + + +class Relocation: + """One simultaneous move: files (repo-relative old path -> new path), + directories (old dir -> new dir), and optionally a domain key rename for + catalog frontmatter (`domain: <key>`). `known` holds tracked paths and + directories: a path is rewritten only when it names one, or names a + moved file. `repo` (a Repository) says which URLs point into this + repository at a branch; such a URL is rewritten like a path from the + root.""" + + def __init__(self, files=None, dirs=None, known=None, domain=None, repo=None): + self.files = dict(files or {}) + self.dirs = dict(dirs or {}) + self.known = set(known or ()) + self.domain = domain # (old key, new key) or None + self.repo = repo + self.basenames = {posixpath.basename(p) for p in self.files} + + def target(self, rel: str) -> Optional[str]: + """Where a repo-relative path is after the move, or None if it stays.""" + if rel in self.files: + return self.files[rel] + for old, new in self.dirs.items(): + if rel == old or rel.startswith(old + '/'): + return new + rel[len(old):] return None - return result.stdout if result.returncode == 0 else None -def _frozen_snapshot(adr) -> Optional[tuple]: - """(frontmatter, body) of the first committed version that was already - adr/v1 and past proposed, following renames. None outside git, for an - untracked file, or when no such version exists. A version that was still - v0 is never the snapshot: migrating an accepted v0 record to v1 adds the - v1 fields, and that is the migration, not an edit (ADR-304 §7).""" - root = get_project_root() - try: - rel = adr.path.resolve().relative_to(root.resolve()) - except ValueError: - return None - # A record freezes where it lands: history is read from the default - # branch when there is one, so a decision still in review on a feature - # branch can be revised. Without a remote default, HEAD's history counts. - # --first-parent keeps to the branch's own line, so a pull request merged - # with a merge commit lands at the merge, not at the first commit on its - # branch; -m lists the merge's files against that parent on older git. - ref = (_git(['rev-parse', '--abbrev-ref', 'origin/HEAD'], root) or '').strip() or 'HEAD' - # --reverse drops pre-rename history under --follow, so read newest first - # and reverse here. -z keeps names with spaces or non-ASCII intact. - log = _git(['log', ref, '--first-parent', '-m', '--follow', '-z', '--format=commit:%H', '--name-only', '--', str(rel)], root) - if not log: + def _moved(self, rel: str) -> Optional[str]: + """target(rel) when rel names a real path that moved: a moved file, or + a path under a moved folder that is tracked before or after the move.""" + if rel in self.files: + return self.files[rel] + moved = self.target(rel) + if moved is not None and (rel in self.known or moved in self.known): + return moved return None - entries, commit = [], None - for token in log.split('\0'): - if token.startswith('commit:'): - commit = token[len('commit:'):] - elif commit and token.lstrip('\n'): - entries.append((commit, token.lstrip('\n'))) - commit = None - for commit, name in reversed(entries): - text = _git(['show', f'{commit}:{name}'], root) - if text is None: + + def _rooted(self, core: str) -> Optional[str]: + """core read as a path from the repo root, where it names something + that moved; None otherwise.""" + lead = '/' if core.startswith('/') else '' + rest = core[len(lead):] + if not rest or posixpath.normpath(rest) != rest: + return None + moved = self._moved(rest) + return lead + moved if moved is not None else None + + def _path(self, token: str, old_dir: str, new_dir: str) -> Optional[str]: + """token rewritten as a path to something that moved, or None.""" + slash = token.endswith('/') and len(token) > 1 + core = token.rstrip('/') if slash else token + if not core: + return None + # Relative to the file that holds it, resolved from where that file + # was. A file that moves re-bases its links to paths that stay. + if not core.startswith('/'): + resolved = posixpath.normpath(posixpath.join(old_dir, core)) + if not resolved.startswith('..'): + moved = self._moved(resolved) + if moved is not None or (old_dir != new_dir and resolved in self.known): + dest = moved or resolved + if posixpath.normpath(posixpath.join(new_dir, core)) == dest: + return None + # Keep the link's shape where it still resolves: rename + # only the segments that moved (../system/X -> ../platform/X). + segments, walked = [], old_dir + for segment in core.split('/'): + walked = posixpath.normpath(posixpath.join(walked, segment)) if segment else walked + after = self.target(walked) if segment not in ('', '.', '..') else None + segments.append(posixpath.basename(after) if after else segment) + new = '/'.join(segments) + if posixpath.normpath(posixpath.join(new_dir, new)) != dest: + new = posixpath.relpath(dest, new_dir or '.') + if core.startswith('./') and not new.startswith('.'): + new = './' + new + return new + ('/' if slash else '') + if resolved in self.known: + return None # a real path that stays + # A path written from the repo root names what moved in full. + new = self._rooted(core) + return (new + ('/' if slash else '')) if new is not None and new != core else None + + def _url(self, token: str) -> Optional[str]: + """A URL into this repository at a branch, with the path it names + rewritten.""" + found = self.repo.path(token) if self.repo else None + if found is None: + return None + at, path = found + slash = path.endswith('/') + new = self._rooted(path.rstrip('/')) + if new is None: + return None + return token[:at] + new + ('/' if slash else '') + + def word(self, token: str, old_dir: str, new_dir: str) -> tuple: + """One run of path characters rewritten; (text, count).""" + stripped = token.rstrip('.') # sentence punctuation is not the path + if not stripped: + return token, 0 + if '://' in stripped: + new = self._url(stripped) + else: + if '/' not in stripped and stripped not in self.basenames: + # A bare sibling name is a path only when the file holding it + # moved and the name resolves to a tracked file: the link must + # be re-based even though its target stays put. + sibling = posixpath.normpath(posixpath.join(old_dir, stripped)) + if old_dir == new_dir or sibling not in self.known: + return token, 0 + new = self._path(stripped, old_dir, new_dir) + if new is None: + return token, 0 + return new + token[len(stripped):], 1 + + def text(self, content: str, rel: str) -> tuple: + """content of the file at rel with its paths rewritten; (text, count). + History keys in frontmatter are left alone, a frontmatter `domain:` + follows a domain rename, and in a Markdown file a fenced code block + is left as written: it quotes text, such as a command, rather than + linking to a file.""" + moved = self.target(rel) + old_dir = posixpath.dirname(rel) + new_dir = posixpath.dirname(moved) if moved else old_dir + lines = content.split('\n') + fm_end = None + if lines and lines[0].rstrip('\r') == '---': + fm_end = next((i for i in range(1, len(lines)) if lines[i].rstrip('\r') == '---'), None) + fenced = _fenced(lines, fm_end + 1 if fm_end is not None else 0) if rel.lower().endswith('.md') else set() + total, key, out = 0, None, [] + + def swap(token): + nonlocal total + new, n = self.word(token, old_dir, new_dir) + total += n + return new + + for i, line in enumerate(lines): + if i in fenced: + out.append(line) + continue + if fm_end is not None and 0 < i < fm_end: + m = _FM_KEY_RE.match(line) + if m: + key = m.group(1) + if key in RELOCATE_HISTORY_KEYS: + out.append(line) + continue + if self.domain and m and key == 'domain': + dm = re.fullmatch(r'(domain:\s*)([\w-]+)(\s*(?:#.*)?)', line.rstrip('\r')) + if dm and dm.group(2) == self.domain[0]: + out.append(dm.group(1) + self.domain[1] + dm.group(3) + + ('\r' if line.endswith('\r') else '')) + total += 1 + continue + out.append(_sub_paths(line, swap)) + return '\n'.join(out), total + + +def tracked_paths(root: Path) -> tuple: + """(every file git tracks, sorted; those files and their folders)""" + out = _git(['ls-files', '-z'], root) + tracked = sorted(n for n in (out or '').split('\0') if n) + known = set(tracked) + for name in tracked: + parent = posixpath.dirname(name) + while parent and parent not in known: + known.add(parent) + parent = posixpath.dirname(parent) + return tracked, known +# --- import sheets and the v1 record writer (ADR-306) ------------------------------- + +SHEET_FORMAT = 'adr-import/v1' + +# Frontmatter keys in the order a v1 record is written. Keys a sheet carries +# that are not listed follow in the sheet's own order. +V1_KEY_ORDER = ('contract', 'kind', 'verb', 'capability', 'targets', 'supersedes', 'amends', + 'extends', 'decided_by', 'superseded_by', 'enacted', 'basis', 'agent', + 'considered', 'concern', 'observable', 'status', 'date', 'deciders', 'related', 'imported') + +V1_SUMMARY_SKELETON = '''## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] +''' + +# The body `adr new` writes, under v0 and v1 alike. +BODY_SKELETON = '''## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] +''' + +# The bracketed prompts in the skeletons. A line still holding one is unfinished. +SKELETON_PROMPTS = tuple(dict.fromkeys(re.findall(r'\[[^\]\n]+\]', V1_SUMMARY_SKELETON + BODY_SKELETON))) + +class _RecordDumper(yaml.SafeDumper): + """Block-style YAML as this corpus writes it: list items indented under + their key, an empty field as `~`, and a multi-line string on one + double-quoted line, so no line of it can read as the `---` fence.""" + def increase_indent(self, flow=False, indentless=False): + return super().increase_indent(flow, False) + +def _represent_str(dumper, value): + style = '"' if '\n' in value else None + return dumper.represent_scalar('tag:yaml.org,2002:str', value, style=style) + +_RecordDumper.add_representer(str, _represent_str) +_RecordDumper.add_representer(type(None), + lambda dumper, _: dumper.represent_scalar('tag:yaml.org,2002:null', '~')) + +def _as_date(value): + """A valid YYYY-MM-DD string as a date, so it is written unquoted. An + invalid one stays a string for lint to report.""" + if isinstance(value, str) and re.fullmatch(r'\d{4}-\d{2}-\d{2}', value): + try: + return date.fromisoformat(value) + except ValueError: + return value + return value + +def empty_record(kind: str, schema: dict, given: dict, defaults: dict) -> dict: + """The frontmatter a record of this kind starts with: every field the kind + requires, from `given` where supplied and empty otherwise. The kind's + schema decides the fields, not its name (ADR-304 §1, §4).""" + record = {'contract': V1, 'kind': kind} + if schema.get('verb') == 'required': + record['verb'] = given.get('verb') + empty = {'basis': [], 'targets': []} + for name in v1_requires(schema): + if name == 'agent': + record['agent'] = {'name': given.get('agent'), 'model': given.get('model')} + else: + record[name] = given.get(name, empty.get(name)) + lifecycle = v1_lifecycle(schema) + default_status = str(defaults.get('status') or '').lower() + record['status'] = default_status if default_status in lifecycle else lifecycle[0] + return record + +def new_sheet(number: int, domain: str, title: str, record: dict, + summary: Optional[str] = None, body: str = '') -> dict: + """A sheet for a record that has no source: what `adr new` applies.""" + return {'sheet': SHEET_FORMAT, 'source': None, + 'target': {'number': number, 'domain': domain, 'title': title}, + 'record': record, 'summary': summary, 'body': body} + +def format_number(number) -> str: + """ADR-007, ADR-101.1: three digits, and a sub-part kept as written.""" + base, _, part = str(number).partition('.') + return f"{int(base):03d}" + (f".{part}" if part else '') + +def _ordered(record: dict) -> dict: + ordered = {k: record[k] for k in V1_KEY_ORDER if k in record} + ordered.update({k: v for k, v in record.items() if k not in ordered}) + return ordered + +def render_record(sheet: dict) -> str: + """The v1 record a sheet describes. It writes only what the sheet holds: + the frontmatter, the title, the Summary if the sheet has one, then the + body verbatim. The frontmatter is read back before it is returned, and a + mismatch raises: the round trip is checked, not assumed (ADR-306 §7).""" + target = sheet['target'] + fields = _ordered(sheet['record']) + if 'date' in fields: + fields['date'] = _as_date(fields['date']) + front = yaml.dump(fields, Dumper=_RecordDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=1000) + if yaml.safe_load(front) != fields: + raise ValueError(f"ADR-{target['number']}: frontmatter does not read back as written") + title = ' '.join(str(target['title']).split()) + text = f"---\n{front}---\n\n# ADR-{format_number(target['number'])}: {title}\n" + summary = sheet.get('summary') + if summary: + summary = summary if summary.startswith('## Summary') else f"## Summary\n\n{summary}" + text += '\n' + summary.rstrip('\n') + '\n' + if sheet.get('body'): + text += '\n' + sheet['body'] + return text + +def placeholder_lines(text: str) -> list: + """Lines outside fenced code that still hold a skeleton prompt.""" + found, fence = [], None + for line in text.splitlines(): + marker = re.match(r'\s*(```|~~~)', line) + if marker: + fence = None if fence == marker.group(1) else (fence or marker.group(1)) continue - past = parse_text(text, adr.path) - if past.contract == V1 and past.status and str(past.status).lower() != 'proposed': - return past.frontmatter, past.body + if fence is None and any(prompt in line for prompt in SKELETON_PROMPTS): + found.append(line.strip()) + return found + +# --- surgical frontmatter edits ----------------------------------------------------- +# +# The record commands (consider, set, supersede, enact) change one field at a +# time. They rewrite only the lines of that field and leave every other byte of +# the file as it was: the other fields' YAML style, comments, line endings and +# the body. A new field is written in the tool's own style (_RecordDumper) at +# its place in V1_KEY_ORDER. A field keeps its layout where it can: a one-line +# field stays on one line, and a block list gains or loses item lines. Every +# edit is read back before it is returned, so a mismatch refuses the edit +# rather than writing something else. + +_TOP_KEY = re.compile(r'^([A-Za-z_][A-Za-z0-9_-]*)\s*:(?:\s|$)') + +class _Quoted(str): + """A string written double-quoted, as the corpus writes `said` and a + commit hash. It reads back as a plain str.""" + +_RecordDumper.add_representer( + _Quoted, lambda dumper, value: dumper.represent_scalar('tag:yaml.org,2002:str', str(value), style='"')) + +class _FlowList(list): + """A list written on one line, as the corpus writes a considered + entry's covers. It reads back as a plain list.""" + +_RecordDumper.add_representer( + _FlowList, lambda dumper, value: dumper.represent_sequence('tag:yaml.org,2002:seq', list(value), flow_style=True)) + +def _prefer_double(value: str): + """value double-quoted where YAML would single-quote it, since the + corpus quotes with double quotes; plain where it needs no quotes.""" + return _Quoted(value) if _flow(value).startswith("'") else value + +def _flow(value) -> str: + return yaml.dump(value, Dumper=_RecordDumper, default_flow_style=True, sort_keys=False, + allow_unicode=True, width=100000).rstrip('\n').removesuffix('\n...').rstrip('\n') + +def _block(key: str, value) -> list: + text = yaml.dump({key: value}, Dumper=_RecordDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=100000) + return text.rstrip('\n').split('\n') + +def _plain(value): + """value with _Quoted strings turned back into str, for comparison.""" + if isinstance(value, list): + return [_plain(v) for v in value] + if isinstance(value, dict): + return {k: _plain(v) for k, v in value.items()} + return str(value) if isinstance(value, _Quoted) else value + +class FrontmatterEdit: + """One record's file, split into its frontmatter lines and the rest.""" + + def __init__(self, raw: bytes): + text = raw.decode('utf-8') + lines = text.split('\n') + if not lines or lines[0].rstrip('\r').strip() != '---': + raise ValueError("the file has no YAML frontmatter") + end = next((i for i, line in enumerate(lines[1:], 1) if line.strip() == '---'), None) + if end is None: + raise ValueError("the frontmatter has no closing ---") + self.crlf = lines[0].endswith('\r') + self.head = lines[:1] + self.tail = lines[end:] + self.lines = [line[:-1] if self.crlf and line.endswith('\r') else line for line in lines[1:end]] + self.fields = self._load() + + def _load(self) -> dict: + data = yaml.safe_load('\n'.join(self.lines)) or {} + if not isinstance(data, dict): + raise ValueError("the frontmatter is not a mapping of fields") + return data + + def bytes(self) -> bytes: + body = [line + '\r' for line in self.lines] if self.crlf else list(self.lines) + return '\n'.join(self.head + body + self.tail).encode('utf-8') + + def _blocks(self) -> list: + """(key, first line, last line) of each top-level field. A field's + block ends at its last indented line: a comment or blank line at + column 0 before the next key stays outside it.""" + blocks = [] + for i, line in enumerate(self.lines): + match = _TOP_KEY.match(line) + if match: + blocks.append([match.group(1), i, i]) + elif blocks and line[:1] in (' ', '\t') and line.strip(): + blocks[-1][2] = i + elif blocks and line.strip() and not line.startswith('#'): + blocks[-1][2] = i # a continuation at column 0, such as a flow list's rest + return [tuple(b) for b in blocks] + + def _block(self, key: str) -> Optional[tuple]: + found = [b for b in self._blocks() if b[0] == key] + if len(found) > 1: + raise ValueError(f"'{key}' appears {len(found)} times in the frontmatter") + return found[0] if found else None + + def _replace(self, start: int, end: int, new_lines: list) -> None: + self.lines[start:end + 1] = new_lines + + def _insert_at(self, key: str) -> int: + """The line a new field goes in at: after the nearest field before it in + V1_KEY_ORDER, else before the nearest field after it, else at the end.""" + blocks = self._blocks() + if key in V1_KEY_ORDER: + rank = V1_KEY_ORDER.index(key) + before = [b for b in blocks if b[0] in V1_KEY_ORDER and V1_KEY_ORDER.index(b[0]) < rank] + if before: + return max(b[2] for b in before) + 1 + after = [b for b in blocks if b[0] in V1_KEY_ORDER and V1_KEY_ORDER.index(b[0]) > rank] + if after: + return min(b[1] for b in after) + return blocks[-1][2] + 1 if blocks else len(self.lines) + + def _one_line(self, key: str, value, old: str) -> str: + """key: value on one line, keeping a trailing comment of the old line.""" + text = _flow(value) if isinstance(value, (list, dict)) else _block(key, value)[0][len(key) + 2:] + comment = re.search(r'\s+#[^\'"\]]*$', old) + return f"{key}: {text}" + (comment.group(0) if comment else '') + + def _render(self, key: str, value, block: Optional[tuple]) -> list: + """The lines for key: value. An existing one-line field stays on one + line; a new field, or one written as a block, takes the tool's style.""" + if block is not None and block[1] == block[2]: + one = self._one_line(key, value, self.lines[block[1]]) + if not (isinstance(value, list) and value and any(isinstance(v, dict) for v in value)): + return [one] + lines = _block(key, value) + if block is not None and len(lines) > 1: + lines = [lines[0]] + self._reindent(lines[1:], block) + return lines + + def _item_indent(self, block: tuple) -> Optional[str]: + for line in self.lines[block[1] + 1:block[2] + 1]: + match = re.match(r'^(\s*)- ', line) or re.match(r'^(\s*)-$', line) + if match: + return match.group(1) + return None + + def _reindent(self, item_lines: list, block: tuple) -> list: + """Item lines rendered with the tool's two-space indent, moved to the + indent the field's existing items use.""" + indent = self._item_indent(block) + if indent is None or indent == ' ': + return item_lines + return [indent + line[2:] if line.startswith(' ') else line for line in item_lines] + + def _commit(self, expected: dict) -> None: + if _plain(self._load()) != _plain(expected): + raise ValueError("the edited frontmatter does not read back as intended") + self.fields = self._load() + + # --- operations ------------------------------------------------------------- + + def set(self, key: str, value) -> None: + expected = dict(self.fields) + expected[key] = value + block = self._block(key) + if block is None: + at = self._insert_at(key) + self.lines[at:at] = _block(key, value) + else: + self._replace(block[1], block[2], self._render(key, value, block)) + self._commit(expected) + + def append(self, key: str, items: list) -> None: + """Append items to a list field, creating it when absent. A scalar + becomes a list holding it first.""" + current = self.fields.get(key) + before = [] if current in (None, '') else (list(current) if isinstance(current, list) else [current]) + value = before + list(items) + block = self._block(key) + indent = self._item_indent(block) if block is not None else None + if block is None or indent is None or not isinstance(current, list) or not current: + self.set(key, value) + return + expected = dict(self.fields) + expected[key] = value + new = self._reindent(_block(key, list(items))[1:], block) + self.lines[block[2] + 1:block[2] + 1] = new + self._commit(expected) + + def remove(self, key: str, items: list) -> list: + """Remove each item from a list field. Returns the items not found.""" + current = self.fields.get(key) + entries = current if isinstance(current, list) else ([] if current is None else [current]) + wanted = [str(_plain(i)) for i in items] + missing = [i for i, w in zip(items, wanted) if w not in [str(e) for e in entries]] + if missing: + return missing + value = [e for e in entries if str(e) not in wanted] + block = self._block(key) + indent = self._item_indent(block) if block is not None else None + if indent is None or not value: + self.set(key, value) + return [] + expected = dict(self.fields) + expected[key] = value + starts = [i for i in range(block[1] + 1, block[2] + 1) + if re.match(re.escape(indent) + r'-(\s|$)', self.lines[i])] + chunks = [(s, (starts[n + 1] - 1) if n + 1 < len(starts) else block[2]) for n, s in enumerate(starts)] + for start, end in reversed(chunks): + text = '\n'.join(line[len(indent):] for line in self.lines[start:end + 1]) + loaded = yaml.safe_load(text) + if isinstance(loaded, list) and len(loaded) == 1 and str(loaded[0]) in wanted: + del self.lines[start:end + 1] + self._commit(expected) + return [] +# --- adr import: readers and the sheet file (ADR-306) --------------------------------- +# +# A reader turns one source record into one import sheet (ADR-306 §1, §2). +# It fills what the source states and lists the rest in `todo`. It never +# writes an operator basis, `considered` or `concern`: `deciders` names who +# signed the record, not what they said (ADR-304 §7, §11). + +IMPORT_DIR = ('docs', 'architecture', '.import') + +SHEET_HEADER = ('# adr import sheet (ADR-306). Fill or clear each todo item, then run\n' + '# `adr import apply`. Sheets are working files; only records are committed.\n') + +# ADR-304 §7: the v0 status table. Deprecated depends on the record. +V0_STATUS_MAP = {'draft': 'proposed', 'proposed': 'proposed', 'accepted': 'accepted', + 'superseded': 'superseded', 'rejected': 'rejected'} + +# v0 keys that carry over under the same name and meaning. +V0_CARRIED = ('date', 'deciders', 'related', 'supersedes', 'superseded_by', 'amends') + +# Todo items that lint cannot detect once the record is written. They block +# --partial; every other item is advisory (ADR-306 §3). +BLOCKING_TODO = ('status', 'status note', 'target.number', 'target.domain') + +# The v1 default decision (ADR-304 §1), for a project whose adr.yaml declares no kinds. +IMPORT_DECISION_SCHEMA = {'verb': 'required', 'requires': ['capability', 'basis', 'agent'], + 'sections': ['Summary']} + +_STOPWORDS = frozenset(''' + the and for with that this from into over under onto are was were been being has have + had not but its it's our their them they these those which what when where who whom + why how all any each every one two per via than then there here also only more most + such can could should would will may might must shall does did doing done use used + using uses make makes made new adr record records decision decisions context +'''.split()) + +class SheetError(Exception): + """A source a reader cannot turn into a sheet, or a sheet apply cannot use.""" + +def import_dir() -> Path: + return get_project_root().joinpath(*IMPORT_DIR) + +def sheet_filename(number) -> str: + return f"ADR-{format_number(number)}.yaml" + +def _number_value(text: str) -> str: + """A record number as the sheet holds it: a string, '42' or '101.1', so + YAML never reads a zero-padded number as octal.""" + base, _, part = text.partition('.') + return (base.lstrip('0') or '0') + (f".{part}" if part else '') + +# --- splitting a record into frontmatter, title, Summary and body ------------------- + +def split_record(text: str) -> tuple: + """(frontmatter text or None, text before the H1, H1 match or None, text + after the H1 line). Joined back with the H1 line, the parts are the file.""" + lines = text.split('\n') + front, start = None, 0 + if lines and lines[0].strip() == '---': + for i in range(1, len(lines)): + if lines[i].strip() == '---': + front, start = '\n'.join(lines[1:i]), i + 1 + break + for j in range(start, len(lines)): + match = TITLE_PATTERN.match(lines[j]) + if match: + return front, '\n'.join(lines[start:j]), match, '\n'.join(lines[j + 1:]) + return front, '\n'.join(lines[start:]), None, '' + +def _render_tail(summary: Optional[str], body: str) -> str: + """What render_record writes after the H1 line for this summary and body.""" + tail = '' + if summary: + summary = summary if summary.startswith('## Summary') else f"## Summary\n\n{summary}" + tail += '\n' + summary.rstrip('\n') + '\n' + if body: + tail += '\n' + body + return tail + +def split_summary(after: str) -> tuple: + """(summary, body, note) for the text after the H1. A `## Summary` that + opens the body becomes the summary; the rest is the body, verbatim. The + split is kept only when render_record writes the same text back, so an + unedited sheet keeps the body byte-identical (ADR-306 §7). Otherwise the + Summary stays in the body and note says why.""" + body = after[1:] if after.startswith('\n') else after + lines = body.split('\n') + if lines[0].strip() != '## Summary': + note = 'the Summary is not the first section; kept in the body' if re.search( + r'(?m)^## Summary\s*$', body) else None + return None, body, note + fence, end = None, len(lines) + for k in range(1, len(lines)): + marker = re.match(r'\s*(```|~~~)', lines[k]) + if marker: + fence = None if fence == marker.group(1) else (fence or marker.group(1)) + elif fence is None and lines[k].startswith('## '): + end = k + break + section = '\n'.join(lines[:end]) + rest = '\n'.join(lines[end:]) + summary = section[len('## Summary\n\n'):] if section.startswith('## Summary\n\n') else section + if summary.strip() and _render_tail(summary, rest) == after: + return summary, rest, None + return None, body, 'the Summary does not split cleanly from the body; kept in the body' + +# --- what a reader decides ------------------------------------------------------------ + +def _empty(value) -> bool: + return value in (None, '', [], {}) + +def open_fields(record: dict, schema: dict) -> list: + """The fields the kind requires that the record leaves empty, in the + order a record is written: the reader's part of `todo` (ADR-306 §1).""" + wanted = (['verb'] if schema.get('verb') == 'required' else []) + v1_requires(schema) + found = [] + for name in wanted: + value = record.get(name) + if name == 'agent' and isinstance(value, dict): + found += [f"agent.{k}" for k in ('name', 'model') if _empty(value.get(k))] + elif _empty(value): + found.append(name) + return found + +_STEM_SUFFIXES = ('ations', 'ation', 'ments', 'ment', 'ions', 'ion', 'ings', 'ing', + 'ives', 'ive', 'als', 'al', 'ors', 'or', 'ers', 'er', 'ted', 'ed', + 'ts', 'es', 't', 'e', 's') + +def _stem(word: str) -> str: + """A crude English stem: strip suffixes repeatedly and collapse a doubled + final consonant, so ingest and ingestion rank as one word.""" + word = word.lower() + changed = True + while changed: + changed = False + for suffix in _STEM_SUFFIXES: + if word.endswith(suffix) and len(word) - len(suffix) >= 3: + word = word[:-len(suffix)] + changed = True + break + if len(word) >= 4 and word[-1] == word[-2] and word[-1] not in 'aeiou': + word = word[:-1] + return word + +def _words(text: str) -> set: + return {_stem(w) for w in re.findall(r"[a-z][a-z0-9']+", text.lower()) + if len(w) >= 3 and w not in _STOPWORDS} + +def capability_candidates(text: str, vocabulary: dict, top: int = 3) -> list: + """Vocabulary names ranked by how many words the record's title and + Context share with each capability's name and description (ADR-306 §1). + A word counts less the more descriptions carry it, so a word every + capability mentions does not decide the ranking.""" + words = _words(text) + described = [(str(name), _words(f"{name} {description or ''}")) + for name, description in vocabulary.items()] + spread = {} + for _, found in described: + for word in found: + spread[word] = spread.get(word, 0) + 1 + scored = [] + for i, (name, found) in enumerate(described): + overlap = round(sum(1 / spread[w] for w in words & found), 6) + if overlap: + scored.append((-overlap, i, name)) + return [name for _, _, name in sorted(scored)[:top]] + +def _folder_domain(path: Path) -> Optional[str]: + """The domain whose folder holds the file, or 'legacy' for the legacy folder.""" + name = path.parent.name + for domain, config in get_domains().items(): + folders = config.get('folder') + if name in (folders if isinstance(folders, list) else [folders]): + return domain + if name == 'legacy' and 'legacy' in get_config(): + return 'legacy' + return None + +def number_range(domain: str) -> Optional[tuple]: + """(low, high) for a domain in adr.yaml, or the legacy range.""" + if domain in get_domains(): + return tuple(get_domains()[domain].get('range', (0, -1))) + if domain == 'legacy' and 'legacy' in get_config(): + return tuple(get_legacy_range()) + return None + +def source_domain(path: Path, number) -> Optional[str]: + """The domain a record file sits in: its folder, else its number range.""" + return _folder_domain(path) or _range_domain(number) + +def todo_label(item) -> str: + return str(item).split(':')[0] + +def blocking_todo(todo: list) -> list: + """The todo items --partial cannot write past (ADR-306 §3).""" + return [t for t in todo if todo_label(t) in BLOCKING_TODO] + +def _range_domain(number) -> Optional[str]: + base = int(str(number).split('.')[0]) + for domain, config in get_domains().items(): + low, high = config.get('range', (0, -1)) + if low <= base <= high: + return domain + low, high = get_legacy_range() + return 'legacy' if 'legacy' in get_config() and low <= base <= high else None + +def _kind_schema(kind: str) -> dict: + kinds = _mapping(get_config().get('kinds')) + schema = kinds.get(kind) + return schema if isinstance(schema, dict) else IMPORT_DECISION_SCHEMA + +def _source_path(path: Path) -> str: + """The path a sheet records: repo-relative inside the repo, else absolute.""" + resolved = path.resolve() + try: + return str(resolved.relative_to(get_project_root().resolve())) + except ValueError: + return str(resolved) + +def read_record(path: Path) -> dict: + """The sheet for one record file: the v0 reader, or a v1 record read as + itself (ADR-306 §2, §7). Raises SheetError for a source it cannot read.""" + raw = path.read_bytes() + try: + text = raw.decode('utf-8') + except UnicodeDecodeError: + raise SheetError('not UTF-8 text') + if text.startswith('\ufeff'): + raise SheetError('the file starts with a UTF-8 byte order mark; remove it before scanning') + if '\r' in text: + raise SheetError('CRLF line endings; convert the file to LF before scanning') + front, before, title, after = split_record(text) + if front is None: + raise SheetError('no YAML frontmatter; import reads structured records (ADR-306 §2)') + try: + data = yaml.safe_load(front) or {} + except yaml.YAMLError as e: + raise SheetError(f"frontmatter is not valid YAML: {e}") + if not isinstance(data, dict): + raise SheetError('frontmatter is not a mapping of fields') + if title is None: + raise SheetError("no '# ADR-NNN: title' heading") + contract = data.get('contract') + if contract and contract != V1: + raise SheetError(f"contract '{contract}' has no reader") + fmt = 'v1' if contract == V1 else 'v0' + + todo, provenance = [], {} + from_file = filename_number(path) + number = _number_value(from_file or title.group(1)) + provenance['target.number'] = 'file name' if from_file else 'H1' + if from_file and (title.group(1).lstrip('0') or '0') != from_file: + todo.append(f"target.number: the file name says ADR-{from_file}, the H1 says ADR-{title.group(1)}") + domain = source_domain(path, number) + if _folder_domain(path): + provenance['target.domain'] = f"folder {path.parent.name}" + elif domain: + provenance['target.domain'] = 'number range' + else: + todo.append('target.domain: neither the folder nor the number names a domain') + # Under adr/v1 a record in the tree keeps its number wherever its folder + # puts it (ADR-306 §6); the range only allocates new numbers. + moved_in_tree = repo_contract() == V1 and _in_tree(path) + span = number_range(domain) if domain and not moved_in_tree else None + if span and not span[0] <= int(str(number).split('.')[0]) <= span[1]: + todo.append(f"target.number: ADR-{number} is outside the {domain} range {span[0]}-{span[1]}") + provenance['target.title'] = 'H1' + + preamble = before.strip('\n') + if preamble.strip(): + # A record is written H1 first, so text above the H1 moves to just + # below it. The body keeps it; nothing is dropped (ADR-306 §1). + summary, body = None, preamble + '\n' + after + provenance['body'] = 'text above the H1, then the text after it' + todo.append('preamble: the text between the frontmatter and the H1 moved into the body, ' + 'right after the H1; check where it belongs') + else: + summary, body, note = split_summary(after) + if summary is not None: + provenance['summary'] = '## Summary section' + elif note: + provenance['summary'] = note + + unmapped = {} + if fmt == 'v1': + record = dict(data) + provenance['record'] = 'frontmatter (adr/v1)' + kind = record.get('kind') + schema = _kind_schema(kind) if isinstance(kind, str) else IMPORT_DECISION_SCHEMA + else: + schema = _kind_schema('decision') + record = empty_record('decision', schema, {'model': 'unrecorded'}, {}) + provenance['record.contract'] = 'v0 reader' + provenance['record.kind'] = 'v0 reader: every v0 record is a decision' + provenance['record.agent.model'] = 'v0 reader: the model was not recorded' + record['status'] = _v0_status(data, todo, provenance) + for key in V0_CARRIED: + if key in data: + record[key] = data[key] + provenance[f"record.{key}"] = f"frontmatter {key}" + if 'deciders' in data: + provenance['record.deciders'] = 'frontmatter deciders (never an operator basis, ADR-304 §7)' + for key, value in data.items(): + if key != 'status' and key not in V0_CARRIED: + unmapped[key] = value + todo[:0] = open_fields(record, schema) + todo += [f"unmapped: {key} has no v1 field; apply keeps it under imported.unmapped. " + f"Move it into the record, or leave it there" for key in unmapped] + + candidates = {} + if 'capability' in todo: + context = parse_text(text, path).section_text.get('Context', '') + candidates['capability'] = capability_candidates( + f"{title.group(2)}\n{context}", _mapping(get_config().get('capabilities'))) + return { + 'sheet': SHEET_FORMAT, + 'source': {'path': _source_path(path), 'format': fmt, + 'sha256': hashlib.sha256(raw).hexdigest()}, + 'target': {'number': number, 'domain': domain, 'title': ' '.join(title.group(2).split())}, + 'record': _ordered(record), + 'summary': summary, + 'todo': todo, + 'candidates': candidates, + 'provenance': provenance, + 'unmapped': unmapped, + 'body': body, + } + +def _v0_status(data: dict, todo: list, provenance: dict) -> Optional[str]: + """ADR-304 §7's table. Deprecated is superseded when something replaced + it, and otherwise accepted with a note, which needs judgement.""" + raw = data.get('status') + key = str(raw or '').strip().lower() + if key == 'deprecated': + if data.get('superseded_by'): + provenance['record.status'] = f"frontmatter status: {raw}, with superseded_by (ADR-304 §7)" + return 'superseded' + provenance['record.status'] = f"frontmatter status: {raw}, with no successor (ADR-304 §7)" + todo.append('status note: Deprecated with no successor maps to accepted with a note that ' + 'the decision is historical (ADR-304 §7); add the note to the body') + return 'accepted' + if key in V0_STATUS_MAP: + provenance['record.status'] = f"frontmatter status: {raw}" + return V0_STATUS_MAP[key] + if 'status' not in data: + todo.append('status: the source has no status; set one (ADR-304 §7)') + else: + todo.append(f"status: '{raw}' has no v1 mapping; set one (ADR-304 §7)") return None -def _same(a, b) -> bool: - """Equal, treating a date and its quoted string as the same value.""" - if isinstance(a, (str, int, float, date)) and isinstance(b, (str, int, float, date)): - return str(a) == str(b) - return a == b +# --- the sheet file ---------------------------------------------------------------- -V1_DEFAULT_MUTABLE = ('status', 'enacted', 'superseded_by', 'considered', 'concern') +class _SheetDumper(_RecordDumper): + """A sheet is edited by hand, so multi-line text is a literal block.""" -@file_rule(contract=V1) -def rule_v1_frozen(adr, ctx): - """Once a decision leaves proposed, only the kind's mutable_after_accept - fields may change, and the body grows only by appending (ADR-304 §1, §4).""" - schema = v1_kind_schema(adr, ctx) - if not is_v1_record(adr, ctx) or schema is None: - return - # Proposed records are not frozen, and archived ones carry the archive - # banner by design; neither needs the history walk. - if str(adr.status or '').lower() == 'proposed' or is_archived(adr.path): - return - mutable = schema.get('mutable_after_accept', list(V1_DEFAULT_MUTABLE)) - if mutable == 'all': - return - mutable = set(_str_list(mutable) or V1_DEFAULT_MUTABLE) - snapshot = _frozen_snapshot(adr) - if snapshot is None: - return - then, body_then = snapshot - for key in sorted(set(then) | set(adr.frontmatter)): - if key in mutable: - continue - if not _same(then.get(key), adr.frontmatter.get(key)): - v1_issue(adr, f"'{key}' changed after the decision left proposed; only {', '.join(sorted(mutable)) or 'no fields'} may change") - if not adr.body.rstrip().startswith(body_then.rstrip()): - v1_issue(adr, "body edited after the decision left proposed; a decision grows by appending", 'warning') +def _represent_sheet_str(dumper, value): + style = '|' if '\n' in value else None + return dumper.represent_scalar('tag:yaml.org,2002:str', value, style=style) + +_SheetDumper.add_representer(str, _represent_sheet_str) + +def dump_sheet(sheet: dict) -> str: + """The sheet as YAML, read back before it is returned (ADR-306 §7).""" + text = yaml.dump(sheet, Dumper=_SheetDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=1000) + if yaml.safe_load(text) != sheet: + raise SheetError('the sheet does not read back as written') + return SHEET_HEADER + text + +def _number_token(path: Path) -> Optional[str]: + """target.number as written, when YAML reads it as an unquoted int, so + 042 (octal), 0x2A or 1_0 can be refused rather than read as another number.""" + try: + node = yaml.compose(path.read_text(), Loader=yaml.SafeLoader) + except (OSError, UnicodeDecodeError, yaml.YAMLError): + return None + for mapping, name in ((node, 'target'), (None, 'number')): + if mapping is None: + mapping = node + if not isinstance(mapping, yaml.MappingNode): + return None + node = next((v for k, v in mapping.value if getattr(k, 'value', None) == name), None) + if isinstance(node, yaml.ScalarNode) and node.tag == 'tag:yaml.org,2002:int' and not node.style: + return node.value + return None + +def load_sheet(path: Path) -> dict: + """A sheet file, checked for the fields apply needs.""" + try: + sheet = yaml.safe_load(path.read_text()) + except yaml.MarkedYAMLError as e: + where = f" at line {e.problem_mark.line + 1}" if e.problem_mark else '' + raise SheetError(f"the sheet is not valid YAML: {e.problem}{where}") + except (OSError, UnicodeDecodeError, yaml.YAMLError) as e: + raise SheetError(f"cannot read the sheet: {e}") + if not isinstance(sheet, dict) or sheet.get('sheet') != SHEET_FORMAT: + raise SheetError(f"not an {SHEET_FORMAT} sheet") + target = sheet.get('target') + if not isinstance(target, dict) or any(_empty(target.get(k)) for k in ('number', 'title')): + raise SheetError('target needs a number and a title') + number = target['number'] + token = _number_token(path) + if token is not None and not re.fullmatch(r'0|[1-9][0-9]*', token): + raise SheetError(f"target.number {token} is not a plain decimal, and YAML reads it as " + f"{number}; quote it, as in number: '{token}'") + if isinstance(number, float): + raise SheetError(f"target.number {number} reads as a decimal; quote a sub-part number, " + f"as in number: '101.10'") + if isinstance(number, bool) or not re.fullmatch(r'\d+(\.\d+)?', str(number)): + raise SheetError(f"target.number '{number}' is not a record number") + if not isinstance(sheet.get('record'), dict): + raise SheetError('record is not a mapping of fields') + for key in ('summary', 'body'): + if sheet.get(key) is not None and not isinstance(sheet[key], str): + raise SheetError(f"{key} is not text") + todo = sheet.get('todo') + if todo is not None and not (isinstance(todo, list) and all(isinstance(t, str) for t in todo)): + raise SheetError('todo is not a list of items') + if sheet.get('unmapped') is not None and not isinstance(sheet['unmapped'], dict): + raise SheetError('unmapped is not a mapping of fields') + source = sheet.get('source') + if source is not None and not (isinstance(source, dict) and isinstance(source.get('path'), str) + and source['path']): + raise SheetError('source needs a path') + return sheet # ============================================================================ # Commands # ============================================================================ @@ -1487,8 +2270,21 @@ def cmd_list(args): adrs.sort(key=sort_key) + # Frontmatter filters: --field, and --kind, --verb, --capability as its + # shorthands. A list field matches when it lists the value. + filters = list(getattr(args, 'field', None) or []) + for name in ('kind', 'verb', 'capability'): + if getattr(args, name, None): + filters.append(f"{name}={getattr(args, name)}") + for spec in filters: + key, has_value, value = spec.partition('=') + adrs = [a for a in adrs if _field_matches(a, key.strip(), value if has_value else None)] + + if getattr(args, 'json', False): + return _list_json(adrs, getattr(args, 'group_by', None)) + project = get_config().get('project_name', 'Project') - print(f"\n{project} — Architecture Decision Records ({len(adrs)} total)") + print(f"\n{project} — Agent Decision Records ({len(adrs)} total)") print("=" * 55) def status_icon(status): @@ -1506,8 +2302,14 @@ def cmd_list(args): marker = ' [archived]' if is_archived(adr.path) and getattr(args, 'all', False) else '' print(f" {status_icon(adr.status)} ADR-{adr.number or '???':8} {adr.title or '(no title)'}{suffix}{marker}") - if args.group: - # Group by domain - show domains in config order, then legacy + if getattr(args, 'group_by', None): + for value, members in _groups(adrs, args.group_by): + print(f"\n## {args.group_by}: {value} ({len(members)})") + print("-" * 50) + for adr in members: + print_adr(adr) + elif args.group: + # Group by domain - show domains in config order, then legacy for domain_key, domain_info in domains.items(): domain_adrs = [a for a in adrs if a.domain == domain_key] if not domain_adrs: @@ -1534,6 +2336,63 @@ def cmd_list(args): return 0 + +# --- frontmatter queries for list ---------------------------------------------------- + +_NO_VALUE = '(none)' + +def _field_values(adr, key: str) -> list: + """A field's values as strings: each item of a list, a scalar alone, and + nothing for an absent or empty field, read from the frontmatter. + Values compare case-sensitively.""" + value = adr.frontmatter.get(key) + if value in (None, '', [], {}): + return [] + if isinstance(value, list): + return [json.dumps(v, default=str, ensure_ascii=False) if isinstance(v, (dict, list)) else str(v) + for v in value] + if isinstance(value, dict): + return [json.dumps(value, default=str, ensure_ascii=False)] + return [str(value)] + +def _field_matches(adr, key: str, value: Optional[str]) -> bool: + values = _field_values(adr, key) + if value is None: + return bool(values) + if key == 'status': + return value.lower() in (v.lower() for v in values) + if key in V1_EDGE_FIELDS + ('superseded_by', 'related'): + # A reference without a section matches every section of that record. + number, section = norm_ref(value) + return any(norm_ref(v)[0] == number and (section is None or norm_ref(v)[1] == section) + for v in values) + return value in values + +def _groups(adrs: list, key: str) -> list: + """(value, records) sorted by value, records without the field last. + A record listing several values is in each.""" + groups = {} + for adr in adrs: + for value in _field_values(adr, key) or [_NO_VALUE]: + groups.setdefault(value, []) + if adr not in groups[value]: + groups[value].append(adr) + ordered = sorted(((v, m) for v, m in groups.items() if v != _NO_VALUE), key=lambda g: g[0].lower()) + if _NO_VALUE in groups: + ordered.append((_NO_VALUE, groups[_NO_VALUE])) + return ordered + +def _record_json(adr) -> dict: + return {'number': adr.number, 'title': adr.title, 'path': str(relative_path(adr.path)), + 'status': adr.status, 'frontmatter': adr.frontmatter} + +def _list_json(adrs: list, group_by: Optional[str]) -> int: + if group_by: + data = {value: [_record_json(a) for a in members] for value, members in _groups(adrs, group_by)} + else: + data = [_record_json(a) for a in adrs] + print(json.dumps(data, indent=2, default=str, ensure_ascii=False)) + return 0 def cmd_view(args): """View an ADR using the configured viewer.""" import shutil @@ -1615,8 +2474,20 @@ def cmd_new(args): print(f"Error: File already exists: {filepath}", file=sys.stderr) return 1 - # Generate content today = date.today().isoformat() + if get_config().get('contract') == V1: + content = _v1_content(args, next_num, domain, today, defaults) + if content is None: + return 1 + folder.mkdir(parents=True, exist_ok=True) + filepath.write_text(content) + print(f"Created: {relative_path(filepath)}") + print(f" Domain: {config['name']} ({domain})") + print(f" Number: ADR-{next_num:03d}") + print(" Contract: adr/v1 (fill the empty fields; `adr lint` lists them)") + return 0 + + # Generate content default_status = defaults.get('status', 'Draft') default_deciders = defaults.get('deciders', []) @@ -1638,33 +2509,7 @@ related: [] # ADR-{next_num:03d}: {args.title} -## Context - -[What is the issue that we're seeing that is motivating this decision or change?] - -## Decision - -[What is the change that we're proposing and/or doing?] - -## Consequences - -### Positive - -- [What becomes easier?] - -### Negative - -- [What becomes harder?] - -### Neutral - -- [What other changes does this enable or require?] - -## Alternatives Considered - -- [What other options were evaluated?] -- [Why were they rejected?] -''' +''' + BODY_SKELETON # Write file folder.mkdir(parents=True, exist_ok=True) @@ -1675,6 +2520,51 @@ related: [] print(f" Number: ADR-{next_num:03d}") return 0 + +def _v1_content(args, number: int, domain: str, today: str, defaults: dict) -> Optional[str]: + """A v1 record from an empty sheet (ADR-306 §3). The kind's schema decides + which fields the record carries; fields the arguments do not give stay + empty, and lint names each one. Arguments the kind cannot take are refused.""" + config = get_config() + kinds = {k: v for k, v in _mapping(config.get('kinds')).items() if isinstance(v, dict)} + if not kinds: + print("Error: adr.yaml declares contract adr/v1 but no kinds; `adr lint` reports the config", file=sys.stderr) + return None + wanted = (args.kind or 'decision').lower() + kind = next((k for k in kinds if str(k).lower() == wanted), None) + if kind is None: + print(f"Error: Unknown kind '{args.kind or 'decision'}'. Kinds: {', '.join(map(str, kinds))}", file=sys.stderr) + return None + schema = kinds[kind] + problems = [] + takes_verb = schema.get('verb') == 'required' + if args.verb and not takes_verb: + problems.append(f"a {kind} record takes no verb") + verbs = _str_list(config.get('verbs')) or list(V1_VERBS) + if args.verb and takes_verb and args.verb not in verbs: + problems.append(f"verb '{args.verb}' is not one of: {', '.join(verbs)}") + vocabulary = _mapping(config.get('capabilities')) + if args.capability and args.capability not in vocabulary: + problems.append(f"capability '{args.capability}' is not in the adr.yaml vocabulary") + if (args.agent or args.model) and 'agent' not in v1_requires(schema): + problems.append(f"a {kind} record carries no agent") + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + if problems: + return None + deciders = list(defaults.get('deciders') or []) + if not deciders: + git_user = _detect_git_user() + if git_user: + deciders = [git_user] + given = {'verb': args.verb, 'capability': args.capability, 'agent': args.agent, 'model': args.model} + record = empty_record(kind, schema, given, defaults) + record.update({'date': today, 'deciders': deciders, 'related': []}) + sections = _str_list(schema.get('sections')) or [] + summary = V1_SUMMARY_SKELETON if 'Summary' in sections else None + sheet = new_sheet(number, domain, args.title, record, summary=summary, + body=BODY_SKELETON if takes_verb else '') + return render_record(sheet) def cmd_rename(args): """Rename an ADR's title and/or file slug. Number and domain are unchanged. @@ -1781,7 +2671,7 @@ def cmd_lint(args): """ if args.paths: # Resolved, so records parsed from the arguments match the corpus's - # own paths (grounding and other corpus rules key on the path). + # own paths (corpus rules key on the path). paths = [Path(p).resolve() for p in args.paths] adrs = [parse_adr(p) for p in paths] else: @@ -1868,21 +2758,36 @@ def cmd_index(args): status = adr.status or '?' return f"{status} ({note})" if note else status - lines = [ - "# Architecture Decision Records", - "", - f"This directory contains Architecture Decision Records (ADRs) for {get_config().get('project_name', 'this project')}.", - "Each ADR documents a significant architectural decision, its context, and consequences.", - "", - "## ADR Format", - "", - "All ADRs follow a consistent format:", - "- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded", - "- **Date:** When the decision was made", - "- **Deciders:** Who made the decision", - "- **Context:** The problem or situation requiring a decision", - "- **Decision:** The architectural choice made", - "- **Consequences:** Benefits, drawbacks, and other impacts", + project = get_config().get('project_name', 'this project') + if repo_contract() == 'adr/v1': + about = [ + f"This directory contains the Agent Decision Records (ADRs) for {project}, under the adr/v1 contract.", + "A record's kind says what it holds: a decision, a spec kept current, or evidence (findings, surveys, explorations).", + "", + "## Record Format", + "", + "- **Kind:** decision / spec / evidence, as `adr.yaml` declares", + "- **Status:** proposed / accepted / rejected / abandoned / superseded / archived", + "- **Capability:** what the record is about, from the vocabulary in `adr.yaml`", + "- **Decisions** also carry a verb (add / cut / change / retire / constrain), a basis, and open with a Summary", + "- **Numbers** are permanent; a record's folder is its area", + ] + else: + about = [ + f"This directory contains Agent Decision Records (ADRs) for {project}.", + "Each ADR documents a significant architectural decision, its context, and consequences.", + "", + "## ADR Format", + "", + "All ADRs follow a consistent format:", + "- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded", + "- **Date:** When the decision was made", + "- **Deciders:** Who made the decision", + "- **Context:** The problem or situation requiring a decision", + "- **Decision:** The architectural choice made", + "- **Consequences:** Benefits, drawbacks, and other impacts", + ] + lines = ["# Agent Decision Records", ""] + about + [ "", f"_This index is auto-generated by `adr index`. Configuration: [`adr.yaml`](./adr.yaml)_", "", @@ -2189,6 +3094,452 @@ def cmd_domains(args): print() return 0 +# ============================================================================ +# adr domain: add, rename, move (ADR-306 §6) +# ============================================================================ +# +# The domain layout changes as a corpus grows. A record's number is its +# permanent identity: it is never renumbered and never reused. Under adr/v1 a +# record's folder decides its domain, and a domain's range only allocates the +# numbers of new records. So moving a record moves its file and rewrites every +# path to it, and ADR-N citations stay as they are. +# +# adr.yaml is edited as text, so its comments and layout survive. + +DOMAIN_NAME_RE = re.compile(r'[a-z][a-z0-9_-]*') +FOLDER_NAME_RE = re.compile(r'[A-Za-z0-9][\w.-]*') + + +def cmd_domain(args): + handlers = {'add': _domain_add, 'rename': _domain_rename, 'move': _domain_move} + if args.domain_command not in handlers: + print("usage: adr domain {add,rename,move} ...", file=sys.stderr) + return 2 + return handlers[args.domain_command](args) + + +# --- the files a rewrite reaches ---------------------------------------------------- + +def _relocation_scope(root: Path) -> tuple: + """(names in scope, every tracked path and directory). Scope is every file + git tracks except fixtures, import sheets, the archive (archived records + are never edited), adr.yaml's cite.exclude list, and INDEX.md.""" + import fnmatch + tracked, known = tracked_paths(root) + cite_config = get_config().get('cite') + patterns = ['tests/fixtures', 'docs/architecture/archive'] + [ + str(p) for p in ((cite_config.get('exclude') if isinstance(cite_config, dict) else None) or [])] + + def excluded(name: str) -> bool: + if '.import' in name.split('/'): + return True + for pattern in patterns: + if any(c in pattern for c in '*?['): + if fnmatch.fnmatch(name, pattern): + return True + elif name == pattern.rstrip('/') or name.startswith(pattern.rstrip('/') + '/'): + return True + return False + # INDEX.md is regenerated after the move, so it is not rewritten. A + # symlink is skipped: writing through it would edit its target, which is + # in scope (or excluded) in its own right. docs/scripts/adr is one. + return [n for n in tracked if not excluded(n) and n != 'docs/architecture/INDEX.md' + and not (root / n).is_symlink()], known + + +def _plan_rewrites(relocation: Relocation, root: Path, names: list) -> list: + """(name, count, new bytes, [(line number, before, after)]) for each file + whose text changes. Binary files and files that are not UTF-8 are never + touched.""" + changes = [] + for name in names: + path = root / name + try: + raw = path.read_bytes() + except OSError: + continue + if b'\0' in raw[:8192]: + continue + try: + text = raw.decode('utf-8') + except UnicodeDecodeError: + continue + new, count = relocation.text(text, name) + if count and new != text: + changes.append((name, count, new.encode('utf-8'), _changed_lines(text, new))) + return changes + + +def _changed_lines(before: str, after: str) -> list: + """[(line number, before, after)] for each line the rewrite changed. A + rewrite changes text within lines, so the lines pair up.""" + return [(i, a, b) for i, (a, b) in enumerate(zip(before.split('\n'), after.split('\n')), 1) if a != b] + + +def _print_rewrites(changes: list, verb: str, lines: bool = False) -> None: + """The count of rewritten paths per file, and with lines=True each line + the rewrite changes, before and after.""" + total = sum(c[1] for c in changes) + print(f"{verb} {total} path{'s' if total != 1 else ''} in {len(changes)} file{'s' if len(changes) != 1 else ''}") + for name, count, _, changed in changes: + print(f" {name}: {count}") + if lines: + for number, before, after in changed: + print(f" {name}:{number}") + print(f" {before.rstrip(chr(13))}") + print(f" → {after.rstrip(chr(13))}") + + +def _git_mv(root: Path, old: str, new: str) -> None: + (root / new).parent.mkdir(parents=True, exist_ok=True) + try: + result = subprocess.run(['git', 'mv', old, new], cwd=root, + capture_output=True, text=True, timeout=30) + if result.returncode == 0: + return + except (FileNotFoundError, subprocess.TimeoutExpired): + pass + (root / old).rename(root / new) + + +def _write_rewrites(relocation: Relocation, root: Path, changes: list) -> None: + for name, _, data, _ in changes: + (root / (relocation.target(name) or name)).write_bytes(data) + + +def _refresh_index(root: Path) -> None: + """Regenerate INDEX.md when the project keeps one.""" + if (root / 'docs' / 'architecture' / 'INDEX.md').exists(): + cmd_index(argparse.Namespace(yes=True)) + + +# --- adr.yaml as text --------------------------------------------------------------- + +def _yaml_scalar(value: str) -> str: + if value and re.fullmatch(r'[A-Za-z0-9(][\w ,.()/+-]*', value) and value.strip() == value \ + and value.lower() not in ('yes', 'no', 'true', 'false', 'on', 'off', 'null', '~'): + return value + return '"' + value.replace('\\', '\\\\').replace('"', '\\"') + '"' + + +def _domains_block(lines: list) -> Optional[tuple]: + """(start, end) of the domains block: the `domains:` line and the index + just past its last indented line. None when there is no block form.""" + start = next((i for i, line in enumerate(lines) + if re.fullmatch(r'domains:\s*(#.*)?', line.rstrip('\r\n'))), None) + if start is None: + return None + end = start + 1 + while end < len(lines) and (not lines[end].strip() or lines[end][0] in ' \t'): + end += 1 + return start, end + + +def _entry_indent(lines: list, start: int, end: int) -> tuple: + """(entry indent, field indent) used in the domains block.""" + entry = field_ = None + for line in lines[start + 1:end]: + if not line.strip() or line.lstrip().startswith('#'): + continue + indent = line[:len(line) - len(line.lstrip())] + if entry is None: + entry = indent + elif len(indent) > len(entry): + field_ = indent + break + entry = entry if entry is not None else ' ' + return entry, field_ if field_ is not None else entry * 2 + + +def _save_config(text: str) -> bool: + """Write adr.yaml if the edited text still reads as YAML.""" + try: + data = yaml.safe_load(text) + except yaml.YAMLError as e: + print(f"Error: the edited adr.yaml does not parse ({e}); nothing written", file=sys.stderr) + return False + if not isinstance(data, dict) or not isinstance(data.get('domains'), dict): + print("Error: the edited adr.yaml lost its domains; nothing written", file=sys.stderr) + return False + get_config_path().write_text(text) + reload_config() + return True + + +def _folders(config: dict) -> list: + folders = config.get('folder') + return list(folders) if isinstance(folders, list) else [folders] + + +# --- add ---------------------------------------------------------------------------- + +def _domain_add(args): + name, folder = args.name, args.folder + domains = get_domains() + problems = [] + if not DOMAIN_NAME_RE.fullmatch(name) or name in ('legacy', 'archive'): + problems.append(f"'{name}' is not a usable domain name (lowercase letters, digits, - and _; not legacy or archive)") + elif name in domains: + problems.append(f"domain '{name}' already exists") + m = re.fullmatch(r'\s*(\d+)\s*-\s*(\d+)\s*', args.range or '') + low = high = None + if not m: + problems.append(f"--range '{args.range}' is not A-B") + else: + low, high = int(m.group(1)), int(m.group(2)) + if low > high: + problems.append(f"--range {low}-{high} runs backwards") + if not FOLDER_NAME_RE.fullmatch(folder) or folder == 'archive': + problems.append(f"--folder '{folder}' is not a folder name under docs/architecture (and not archive)") + for other, config in domains.items(): + if folder in _folders(config): + problems.append(f"folder '{folder}' already belongs to {other}") + if low is not None and low <= high: + spans = [(other, config['range']) for other, config in domains.items()] + if 'legacy' in get_config(): + spans.append(('legacy', get_legacy_range())) + for other, (a, b) in spans: + if low <= b and a <= high: + problems.append(f"range {low}-{high} overlaps {other} ({a}-{b})") + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + text = get_config_path().read_text() + lines = text.splitlines(keepends=True) + block = _domains_block(lines) + if block is None: + print("Error: adr.yaml has no block-style domains: section to add to; edit it by hand", file=sys.stderr) + return 1 + start, end = block + entry, field_ = _entry_indent(lines, start, end) + last = max((i for i in range(start, end) if lines[i].strip()), default=start) + if not lines[last].endswith('\n'): + lines[last] += '\n' + label = args.label or name.capitalize() + new = ['\n', f"{entry}{name}:\n", + f"{field_}range: [{low}, {high}]\n", + f"{field_}name: {_yaml_scalar(label)}\n", + f"{field_}description: {_yaml_scalar(args.description or '')}\n", + f"{field_}folder: {_yaml_scalar(folder)}\n"] + if last == start: + new = new[1:] + lines[last + 1:last + 1] = new + if not _save_config(''.join(lines)): + return 1 + print(f"Added domain: {name} ({low}-{high})") + print(f" Name: {label}") + print(f" Folder: docs/architecture/{folder}/") + return 0 + + +# --- rename ------------------------------------------------------------------------- + +def _domain_rename(args): + old, new = args.old, args.new + domains = get_domains() + if old not in domains: + print(f"Error: Unknown domain '{old}'", file=sys.stderr) + print(f"Valid domains: {', '.join(domains.keys())}", file=sys.stderr) + return 1 + config = domains[old] + folders = _folders(config) + problems = [] + if new != old: + if not DOMAIN_NAME_RE.fullmatch(new) or new in ('legacy', 'archive'): + problems.append(f"'{new}' is not a usable domain name (lowercase letters, digits, - and _; not legacy or archive)") + elif new in domains: + problems.append(f"domain '{new}' already exists") + if args.folder and len(folders) > 1: + problems.append(f"{old} spans several folders ({', '.join(folders)}); rename them by hand") + # The folder follows the name when it matched the name; otherwise it + # stays, unless --folder names one. + folder_old = folders[0] + folder_new = args.folder or (new if len(folders) == 1 and folder_old == old else folder_old) + root = get_project_root() + arch = root / 'docs' / 'architecture' + if folder_new != folder_old: + if not FOLDER_NAME_RE.fullmatch(folder_new) or folder_new == 'archive': + problems.append(f"folder '{folder_new}' is not a folder name under docs/architecture (and not archive)") + for other, other_config in domains.items(): + if other != old and folder_new in _folders(other_config): + problems.append(f"folder '{folder_new}' already belongs to {other}") + if (arch / folder_new).exists(): + problems.append(f"docs/architecture/{folder_new} already exists") + if new == old and folder_new == folder_old: + problems.append("nothing to rename: give a new name or --folder") + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + text = get_config_path().read_text() + lines = text.splitlines(keepends=True) + block = _domains_block(lines) + key_at = None + if block: + start, end = block + entry, _ = _entry_indent(lines, start, end) + key_re = re.compile(rf'{re.escape(entry)}{re.escape(old)}:(\s*(?:#.*)?)') + key_at = next((i for i in range(start + 1, end) if key_re.fullmatch(lines[i].rstrip('\r\n'))), None) + if key_at is None: + print(f"Error: cannot find the '{old}:' entry in adr.yaml's domains block; edit it by hand", file=sys.stderr) + return 1 + ending = lines[key_at][len(lines[key_at].rstrip('\r\n')):] + lines[key_at] = f"{entry}{new}:" + key_re.fullmatch(lines[key_at].rstrip('\r\n')).group(1) + ending + if folder_new != folder_old: + stop = next((i for i in range(key_at + 1, end) + if lines[i].strip() and len(lines[i]) - len(lines[i].lstrip()) <= len(entry)), end) + folder_re = re.compile(r'(\s+folder:\s*)(\S+?)(\s*(?:#.*)?)') + for i in range(key_at + 1, stop): + fm = folder_re.fullmatch(lines[i].rstrip('\r\n')) + if fm: + ending_i = lines[i][len(lines[i].rstrip('\r\n')):] + lines[i] = fm.group(1) + _yaml_scalar(folder_new) + fm.group(3) + ending_i + break + else: + print(f"Error: cannot find {old}'s folder: line in adr.yaml; edit it by hand", file=sys.stderr) + return 1 + + names, known = _relocation_scope(root) + old_dir = f"docs/architecture/{folder_old}" + new_dir = f"docs/architecture/{folder_new}" + relocation = Relocation(dirs={old_dir: new_dir} if folder_new != folder_old else {}, + known=known, domain=(old, new) if new != old else None, + repo=Repository(root)) + changes = _plan_rewrites(relocation, root, names) + # adr.yaml is written from its edited text, with its own paths + # rewritten. Its listing is taken against the text as it was, so the + # key and folder lines show, and each counts as one rewrite. + config_rel = 'docs/architecture/adr.yaml' + config_edited = ''.join(lines) + config_text, config_count = relocation.text(config_edited, config_rel) + changes = [c for c in changes if c[0] != config_rel] + if config_text != text: + edited = sum(a != b for a, b in zip(text.split('\n'), config_edited.split('\n'))) + changes = sorted(changes + [(config_rel, config_count + edited, config_text.encode('utf-8'), + _changed_lines(text, config_text))]) + + verb = 'Would rename' if args.dry_run else 'Renamed' + print(f"{verb} domain: {old} → {new}") + if folder_new != folder_old: + print(f" Folder: {old_dir}/ → {new_dir}/") + _print_rewrites(changes, 'Would rewrite' if args.dry_run else 'Rewrote', lines=args.dry_run) + if args.dry_run: + print("Dry run: nothing written.") + return 0 + if not _save_config(config_text): + return 1 + if folder_new != folder_old and (root / old_dir).exists(): + _git_mv(root, old_dir, new_dir) + _write_rewrites(relocation, root, [c for c in changes if c[0] != config_rel]) + _refresh_index(root) + return 0 + + +# --- move --------------------------------------------------------------------------- + +def _load_plan(path: str) -> Optional[list]: + """The moves a plan file lists: [{record, domain}, ...].""" + try: + data = yaml.safe_load(Path(path).read_text()) + except (OSError, yaml.YAMLError) as e: + print(f"Error: cannot read plan {path}: {e}", file=sys.stderr) + return None + if isinstance(data, dict) and 'moves' in data: + data = data['moves'] + if not isinstance(data, list) or not data: + print(f"Error: plan {path}: expected a list of {{record, domain}} entries", file=sys.stderr) + return None + moves = [] + for i, entry in enumerate(data, 1): + if not isinstance(entry, dict) or 'record' not in entry or 'domain' not in entry: + print(f"Error: plan entry {i}: expected {{record, domain}}", file=sys.stderr) + return None + extra = sorted(set(entry) - {'record', 'domain'}) + if extra: + hint = '; a record keeps its number' if 'number' in extra else '' + print(f"Error: plan entry {i}: unknown key{'s' if len(extra) > 1 else ''} {', '.join(extra)}{hint}", file=sys.stderr) + return None + moves.append((str(entry['record']), str(entry['domain']))) + return moves + + +def _domain_move(args): + if repo_contract() != 'adr/v1': + print("Error: under adr/v0 a record's number decides its domain, so a record cannot " + "change domain by folder; declare contract: adr/v1 in adr.yaml first", file=sys.stderr) + return 1 + if args.plan: + if args.record or args.domain: + print("Error: give either <record> <domain> or --plan, not both", file=sys.stderr) + return 2 + moves = _load_plan(args.plan) + if moves is None: + return 1 + else: + if not args.record or not args.domain: + print("Error: adr domain move <record> <domain>, or --plan <file.yaml>", file=sys.stderr) + return 2 + moves = [(args.record, args.domain)] + + root = get_project_root() + domains = get_domains() + records = get_all_adrs(include_archived=True) + problems, files, lines_out, seen = [], {}, [], set() + for ref, domain in moves: + matches = find_by_ref(ref, records) + if not matches: + problems.append(f"ADR not found: {ref}") + continue + if len(matches) > 1: + problems.append(f"{ref} matches several records: " + ', '.join(str(relative_path(a.path, root)) for a in matches)) + continue + adr = matches[0] + rel = relative_path(adr.path, root).as_posix() + if rel in seen: + problems.append(f"ADR-{adr.number} is listed twice") + continue + seen.add(rel) + if is_archived(adr.path): + problems.append(f"ADR-{adr.number} is archived ({rel}); archived records are kept as they are") + continue + if domain not in domains: + problems.append(f"unknown domain '{domain}' (valid: {', '.join(domains.keys())})") + continue + folders = _folders(domains[domain]) + if adr.domain == domain and adr.path.parent.name in folders: + problems.append(f"ADR-{adr.number} is already in {domain} ({rel})") + continue + new_rel = f"docs/architecture/{folders[0]}/{adr.path.name}" + if (root / new_rel).exists(): + problems.append(f"{new_rel} already exists") + continue + files[rel] = new_rel + lines_out.append((adr, adr.domain or 'legacy', domain, rel, new_rel)) + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + names, known = _relocation_scope(root) + relocation = Relocation(files=files, known=known, repo=Repository(root)) + changes = _plan_rewrites(relocation, root, names) + + verb = 'Would move' if args.dry_run else 'Moved' + for adr, was, domain, rel, new_rel in lines_out: + print(f"{verb}: ADR-{adr.number} {was} → {domain}") + print(f" {rel} → {new_rel}") + _print_rewrites(changes, 'Would rewrite' if args.dry_run else 'Rewrote', lines=args.dry_run) + if args.dry_run: + print("Dry run: nothing written.") + return 0 + for rel, new_rel in files.items(): + _git_mv(root, rel, new_rel) + _write_rewrites(relocation, root, changes) + _refresh_index(root) + return 0 def cmd_config(args): """Show current configuration.""" config_path = get_config_path() @@ -2213,13 +3564,6 @@ def cmd_cite(args): - a citation whose records are all out of force (superseded, deprecated, rejected, abandoned, archived) warns and names each status and successor - a citation of a proposed or draft record prompts acceptance - - under adr/v1, a citation of a record on a capability whose latest - accepted add-or-cut is a cut warns until the cut is enacted and fails - after (§5) - - under adr/v1, an accepted retire's targets are checked against the - surface inventory commands in adr.yaml (§5). Inventory commands are - shell commands from adr.yaml; they run only on a whole-repo scan and - never with --no-inventory. A bare ADR-N resolves to its family {N, N.1, ...}. Files come from `git ls-files`, so ignored paths are never read. docs/architecture and the @@ -2249,8 +3593,6 @@ def cmd_cite(args): return 2 records = get_all_adrs(include_archived=True) - ctx = LintContext.from_corpus(records) - v1 = ctx.contract == V1 # Lookup tables, built once: full number -> records, base -> family by_full, by_base = {}, {} @@ -2271,23 +3613,6 @@ def cmd_cite(args): def is_proposed(adr) -> bool: return str(adr.status or '').lower() in ('proposed', 'draft') - # The state of each capability is its latest accepted add or cut, by date - # then number (ADR-304 §3). Only a capability whose latest is a cut counts. - cuts = {} - if v1: - latest = {} - for adr in records: - verb = adr.frontmatter.get('verb') - if (is_v1_record(adr, ctx) and verb in ('add', 'cut') - and str(adr.status or '').lower() == 'accepted'): - order = decision_order(adr) - for capability in capability_scope(adr): - if capability not in latest or order >= latest[capability][0]: - latest[capability] = (order, adr) - for capability, (_, adr) in latest.items(): - if adr.frontmatter.get('verb') == 'cut': - cuts[capability] = (adr, bool(adr.frontmatter.get('enacted'))) - def excluded(name: str) -> bool: for pattern in excludes: if any(c in pattern for c in '*?['): @@ -2329,47 +3654,6 @@ def cmd_cite(args): findings.append(('warning', where, f"{cited} is not in force ({named})")) elif any(is_proposed(m) for m in members): findings.append(('warning', where, f"{cited} is still proposed; accept it or cite the record in force")) - for member in members: - if not is_v1_record(member, ctx): - continue - for capability in capability_scope(member): - if capability in cuts and cuts[capability][0] is not member: - cut, enacted = cuts[capability] - if enacted: - findings.append(('error', where, f"{cited} is on '{capability}', cut and enacted by ADR-{cut.number}; remove the citation")) - else: - findings.append(('warning', where, f"{cited} is on '{capability}', which ADR-{cut.number} cuts; remove before enacting")) - - # Retire targets against the surface inventories (§5). Inventories run - # shell commands from adr.yaml, so only on a whole-repo scan. - if v1 and not scopes and not args.no_inventory: - surfaces = v1_surfaces(ctx) - listed_by, failed = {}, set() - for adr in records: - if (not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'retire' - or str(adr.status or '').lower() != 'accepted'): - continue - enacted = bool(adr.frontmatter.get('enacted')) - for target in as_entries(adr.frontmatter.get('targets')): - namespace, _, target_name = target.partition(':') - command = _mapping(surfaces.get(namespace)).get('inventory') - if not command: - continue - if namespace not in listed_by: - listed_by[namespace] = _inventory(str(command), root) - listed, failure = listed_by[namespace] - where = str(relative_path(adr.path)) - if listed is None: - if namespace not in failed: # once per namespace, not per target - failed.add(namespace) - findings.append(('warning', where, f"inventory for '{namespace}' did not run ({failure}): {command}")) - elif target_name in listed: - if enacted: - findings.append(('error', where, f"{target} is still present after ADR-{adr.number} was enacted")) - else: - findings.append(('warning', where, f"{target} is still present; ADR-{adr.number} retires it")) - elif not enacted: - findings.append(('warning', where, f"{target} is not in the '{namespace}' inventory; check the name")) errors = sum(1 for f in findings if f[0] == 'error') warnings = len(findings) - errors @@ -2407,28 +3691,6 @@ def _cite_files(root: Path) -> list: for filename in filenames: names.append((Path(dirpath) / filename).relative_to(root).as_posix()) return sorted(names) - -def _inventory(command: str, root: Path) -> tuple: - """Run a surface inventory command from adr.yaml. Returns (names, None), - or (None, reason) when it failed. One name per line of output.""" - import signal - try: - proc = subprocess.Popen(command, shell=True, cwd=root, stdin=subprocess.DEVNULL, - stdout=subprocess.PIPE, stderr=subprocess.PIPE, - text=True, start_new_session=True) - except OSError as e: - return None, str(e) - try: - stdout, stderr = proc.communicate(timeout=60) - except subprocess.TimeoutExpired: - os.killpg(proc.pid, signal.SIGKILL) # the shell and anything it started - proc.communicate() - return None, 'timed out after 60s' - if proc.returncode != 0: - detail = (stderr.strip().splitlines() or [''])[-1][:120] - return None, f"exit {proc.returncode}" + (f": {detail}" if detail else '') - return {line.strip() for line in stdout.splitlines() if line.strip()}, None - def _lifecycle_target(args): """The v1 record the command acts on, or an exit code after an error.""" matches = find_by_ref(args.adr, get_all_adrs()) @@ -2471,33 +3733,22 @@ def _set_status(raw: bytes, status: str) -> bytes: lines[i] = b'status: ' + status.encode() + (comment.group(0) if comment else b'') + ending return b'\n'.join(lines) -def _errors_by_record(corpus: list, ctx) -> dict: - run_rules(corpus, ctx) - found = {a.path: {i.message for i in a.issues if i.severity == 'error'} for a in corpus} - found['adr.yaml'] = {i.message for i in ctx.config_issues if i.severity == 'error'} - return found - def _trial(adr, status: str): - """Run every rule over the whole corpus as it is and again with the - record's status changed. Returns (the trial record, new errors as - (where, message)). A change that breaks another record, such as a - rejected precedent, shows up here, not just the record's own issues.""" - baseline = get_all_adrs(include_archived=True) - before = _errors_by_record(baseline, LintContext.from_corpus(baseline)) + """The record with its status changed, and its own errors as (where, + message): the file rules only, over this one record. The rest of the + corpus is not re-linted (ADR-311 §3).""" corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) trial = next(a for a in corpus if a.path == adr.path) trial.status = status trial.frontmatter = dict(trial.frontmatter, status=status) # The record was parsed clean enough to reach here (_lifecycle_target # refused anything else), so starting its issues fresh loses nothing. trial.issues = [] - after = _errors_by_record(corpus, LintContext.from_corpus(corpus)) - new = [] - for path, messages in after.items(): - for message in sorted(messages - before.get(path, set())): - where = 'adr.yaml' if path == 'adr.yaml' else relative_path(path) - new.append((where, message)) - return trial, new + for rule in _active(FILE_RULES, ctx): + rule(trial, ctx) + where = relative_path(trial.path) + return trial, [(where, i.message) for i in trial.issues if i.severity == 'error'] def _write_status(adr, status: str, suffix: str = '') -> Optional[str]: """Write the new status (and an appended suffix), then re-read the file to @@ -2517,20 +3768,18 @@ def _write_status(adr, status: str, suffix: str = '') -> Optional[str]: return None def _refuse(adr, verb: str, new: list) -> int: - print(f"Refused: {verb} ADR-{adr.number} would add errors:", file=sys.stderr) + print(f"Refused: {verb} ADR-{adr.number} would leave it with errors:", file=sys.stderr) for where, message in new: print(f" ❌ {where}: {message}", file=sys.stderr) return 1 def cmd_accept(args): - """Accept a proposed adr/v1 record (ADR-304 §2, §11, §12). - - Runs every rule over the corpus as if the record were accepted and refuses - if that adds an error anywhere, so a decision whose basis does not meet the - rules, or one the operator started without a considered entry, stays - proposed. A precedent that is still proposed also refuses: §11 grounds a - decision in another accepted decision. The record's warnings are printed, - with open concerns first: shown at acceptance, not blocking. + """Accept a proposed adr/v1 record (ADR-304 §2, ADR-311 §3). + + Runs the record's own file rules as if it were accepted and refuses if + any reports an error. The rest of the corpus is not re-linted. The + record's warnings are printed, with open concerns first: shown at + acceptance, not blocking. """ adr, code = _lifecycle_target(args) if adr is None: @@ -2538,9 +3787,6 @@ def cmd_accept(args): trial, new = _trial(adr, 'accepted') if new: return _refuse(adr, 'accepting', new) - pending = [i for i in trial.issues if i.code == 'precedent-proposed'] - if pending: - return _refuse(adr, 'accepting', [(relative_path(adr.path), i.message) for i in pending]) concerns = [i for i in trial.issues if i.code == 'open-concern'] others = [i for i in trial.issues if i.severity == 'warning' and i.code != 'open-concern'] if concerns: @@ -2592,13 +3838,698 @@ def cmd_reject(args): def cmd_abandon(args): return _close(args, 'abandoned') + +# --- record edits: consider, set, supersede, enact ----------------------------------- +# +# Each command edits frontmatter through FrontmatterEdit, so only the lines of +# the fields it touches change. It refuses what the contract refuses before +# writing, then lints the records it wrote and prints their issues. + +def _record_target(ref: str, command: str, v1_only: bool = True): + """The record ref names, or (None, exit code) after an error.""" + matches = find_by_ref(ref, get_all_adrs(include_archived=True)) + if not matches: + print(f"Error: ADR not found: {ref}", file=sys.stderr) + return None, 1 + if len(matches) > 1: + print(f"Error: '{ref}' matches several records; resolve the duplicate number first:", file=sys.stderr) + for match in matches: + print(f" {relative_path(match.path)}", file=sys.stderr) + return None, 1 + adr = matches[0] + if v1_only and (repo_contract() != V1 or adr.contract != V1): + print(f"Error: ADR-{adr.number} is not an adr/v1 record; `adr {command}` edits v1 fields " + f"(ADR-304 §7).", file=sys.stderr) + return None, 1 + return adr, 0 + +def _open_edit(adr): + try: + return FrontmatterEdit(adr.path.read_bytes()), None + except (ValueError, UnicodeDecodeError, yaml.YAMLError) as e: + return None, f"ADR-{adr.number}: {e}" + +def _record_schema(adr) -> Optional[dict]: + kinds = get_config().get('kinds') + schema = kinds.get(adr.frontmatter.get('kind')) if isinstance(kinds, dict) else None + return schema if isinstance(schema, dict) else None + +def _show_diff(adr, before: bytes, after: bytes) -> None: + import difflib + name = str(relative_path(adr.path)) + diff = difflib.unified_diff(before.decode('utf-8').splitlines(True), after.decode('utf-8').splitlines(True), + fromfile=f"a/{name}", tofile=f"b/{name}") + sys.stdout.writelines(line if line.endswith('\n') else line + '\n' for line in diff) + +def _write(edits: list, dry_run: bool) -> None: + """Write each (adr, edit) whose bytes changed, or show the diff.""" + for adr, edit in edits: + before, after = adr.path.read_bytes(), edit.bytes() + if before == after: + continue + if dry_run: + _show_diff(adr, before, after) + else: + adr.path.write_bytes(after) + +def _lint_records(paths: list) -> None: + """Lint the records just written against the whole corpus and print + their issues, as `adr lint <path>` would list them.""" + corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) + records = [a for a in corpus if a.path in paths] + run_rules(records, ctx) + for record in records: + where = relative_path(record.path) + if not record.issues: + print(f"Lint {where}: clean") + continue + print(f"Lint {where}:") + for issue in record.issues: + print(f" {'❌' if issue.severity == 'error' else '⚠️'} {issue.message}") + +def _finish(edits: list, dry_run: bool, done: str) -> int: + _write(edits, dry_run) + if dry_run: + print(f"Dry run, nothing written: {done}") + return 0 + print(done) + _lint_records([adr.path for adr, _ in edits]) + return 0 + +# --- consider ----------------------------------------------------------------------- + +_PROBE_NAME = re.compile(r'\*\s*(?:not\s+)?confident\s*\(([^)\n]+)\)\s*:?\s*\*', re.IGNORECASE) + +def probe_names(adr) -> list: + """The named probes in a record's Summary, written *Confident (name):* or + *Not confident (name):*, then 'inversion' when the Summary has one: a + considered entry may cover the inversion too (ADR-304 §12).""" + heading = _find_section(adr, 'Summary') + text = adr.section_text.get(heading, '') if heading else '' + names = list(dict.fromkeys(m.strip() for m in _PROBE_NAME.findall(text))) + if re.search(r'\binversion\b', text, re.IGNORECASE): + names.append('inversion') + return names + +def cmd_consider(args): + """Append one considered entry: the operator's answer to the Summary + (ADR-304 §12), on a record in any status.""" + adr, code = _record_target(args.adr, 'consider') + if adr is None: + return code + for flag, value, what in (('--said', args.said, 'what the operator said, verbatim'), + ('--via', args.via, 'the channel it was said in, such as a PR or a session')): + if value is None or not value.strip(): + print(f"Error: {flag} is required: {what}.", file=sys.stderr) + return 1 + operator = args.operator or _detect_git_user() + if not operator: + print("Error: no operator given and none detected from gh or git; pass --operator NAME.", file=sys.stderr) + return 1 + if args.covers is not None: + valid = probe_names(adr) + unknown = [c for c in args.covers if c not in valid] + if unknown: + names = ', '.join(repr(u) for u in unknown) + if any(v != 'inversion' for v in valid): + print(f"Error: ADR-{adr.number} has no probe {names}. Names it can cover: {', '.join(valid)}.", + file=sys.stderr) + else: + also = f" It can cover: {', '.join(valid)}." if valid else '' + print(f"Error: ADR-{adr.number} has no probe {names}; its Summary names no probes. " + f"A probe is named as *Confident (name):* or *Not confident (name):*.{also}", + file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + entry = {'operator': operator, 'said': _Quoted(args.said), 'via': _prefer_double(args.via)} + if args.paraphrase: + entry['paraphrase'] = True + if args.covers is not None: + entry['covers'] = _FlowList(args.covers) + if args.canary: + entry['canary'] = args.canary + try: + edit.append('considered', [entry]) + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + count = len(edit.fields.get('considered') or []) + return _finish([(adr, edit)], args.dry_run, f"Added considered entry {count} to ADR-{adr.number}: {adr.title}") + +# --- set ---------------------------------------------------------------------------- + +_ASSIGNMENT = re.compile(r'^([A-Za-z_][A-Za-z0-9_]*(?:-[A-Za-z0-9_]+)*)(\+=|-=|=)(.*)$', re.DOTALL) +# Statuses with a command of their own, which checks the record and records why. +_STATUS_COMMANDS = {'accepted': 'accept', 'rejected': 'reject', 'abandoned': 'abandon', + 'superseded': 'supersede', 'archived': 'archive'} + +def _parse_assignment(text: str): + match = _ASSIGNMENT.match(text) + if not match: + raise ValueError(f"'{text}' is not key=value, key+=value or key-=value") + key, op, raw = match.groups() + try: + value = yaml.safe_load(raw) if raw.strip() else None + except yaml.YAMLError as e: + raise ValueError(f"{key}: the value is not YAML ({str(e).splitlines()[0]})") + return key, op, value + +def cmd_set(args): + """Set, append to or remove from frontmatter fields. status is refused: + the lifecycle commands own it. Any other field may be set on a record in + any status; git keeps what it was (ADR-311).""" + adr, code = _record_target(args.adr, 'set', v1_only=False) + if adr is None: + return code + try: + changes = [_parse_assignment(a) for a in args.assignments] + except ValueError as e: + print(f"Error: {e}", file=sys.stderr) + return 1 + if adr.contract == V1 and not args.force: + for key, op, value in changes: + if key != 'status': + continue + command = _STATUS_COMMANDS.get(str(value).lower()) if op == '=' else None + if command: + print(f"Refused: use `adr {command} {adr.number}`; it checks the record and records why. " + f"--force sets the status anyway.", file=sys.stderr) + else: + print(f"Refused: a record's status changes only through accept, reject, abandon, " + f"supersede and archive. --force sets the status anyway.", file=sys.stderr) + return 1 + # A key the record lacks and no v1 record carries is most likely a typo. + unknown = [k for k, _, _ in changes + if adr.contract == V1 and k not in adr.frontmatter and k not in V1_KEY_ORDER] + if unknown and not args.force: + print(f"Refused: ADR-{adr.number} has no field {', '.join(repr(k) for k in unknown)}, and it is not " + f"a v1 field ({', '.join(V1_KEY_ORDER)}). --force adds it anyway.", file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + for key, op, value in changes: + items = value if isinstance(value, list) else [value] + if op == '=': + edit.set(key, value) + elif op == '+=': + edit.append(key, items) + else: + missing = edit.remove(key, items) + if missing: + print(f"Error: ADR-{adr.number} not changed: {key} does not list " + f"{', '.join(repr(str(m)) for m in missing)}.", file=sys.stderr) + return 1 + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + keys = ', '.join(dict.fromkeys(k for k, _, _ in changes)) + return _finish([(adr, edit)], args.dry_run, f"Set {keys} on ADR-{adr.number}: {adr.title}") + +# --- supersede ---------------------------------------------------------------------- + +def _ref(adr, section: Optional[str] = None) -> str: + return f"ADR-{adr.number}" + (f"#{section}" if section else '') + +def _lists(adr, key: str, ref: str) -> bool: + number, section = norm_ref(ref) + return any(norm_ref(e) == (number, section) for e in as_entries(adr.frontmatter.get(key))) + +def cmd_supersede(args): + """Write both sides of a supersession, or with --amends a partial + replacement on the new record only (ADR-304 §3).""" + old, code = _record_target(args.adr, 'supersede') + if old is None: + return code + new, code = _record_target(args.by, 'supersede') + if new is None: + return code + if old.path == new.path: + print(f"Error: ADR-{old.number} cannot supersede itself.", file=sys.stderr) + return 1 + new_status = str(new.status or '').lower() + if new_status not in ('accepted', 'proposed'): + print(f"Error: ADR-{new.number} is {new_status or 'without a status'}; only an accepted or proposed " + f"record supersedes another.", file=sys.stderr) + return 1 + old_status = str(old.status or '').lower() + if old_status not in ('accepted', 'superseded'): + print(f"Error: ADR-{old.number} is {old_status or 'without a status'}; only an accepted record is " + f"superseded (a proposed one is rejected or abandoned).", file=sys.stderr) + return 1 + field_name = 'amends' if args.amends else 'supersedes' + new_kind, old_kind = new.frontmatter.get('kind'), old.frontmatter.get('kind') + edges = v1_edges(_record_schema(new) or {}) + if old_kind not in edges.get(field_name, []): + allowed = edges.get(field_name) + target = f"points at {' or '.join(allowed)} records" if allowed else "is not an edge it takes" + print(f"Error: a {new_kind} record cannot {'amend' if args.amends else 'supersede'} a {old_kind}: " + f"for a {new_kind}, {field_name} {target} (adr.yaml kinds.{new_kind}.edges).", file=sys.stderr) + return 1 + section = args.amends + if section and not section_exists(old, section): + print(f"Error: ADR-{old.number} has no section '{section}'. Its sections: {', '.join(old.sections)}.", + file=sys.stderr) + return 1 + if not section and 'superseded' not in v1_lifecycle(_record_schema(old)): + print(f"Error: the {old_kind} lifecycle has no superseded status.", file=sys.stderr) + return 1 + ref = _ref(old, section) + edits = [] + if not _lists(new, field_name, ref): + edit, error = _open_edit(new) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + edit.append(field_name, [ref]) + except ValueError as e: + print(f"Error: ADR-{new.number} not changed: {e}", file=sys.stderr) + return 1 + edits.append((new, edit)) + if not section: + edit, error = _open_edit(old) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + if not _lists(old, 'superseded_by', _ref(new)): + edit.append('superseded_by', [_ref(new)]) + if old_status != 'superseded': + edit.set('status', 'superseded') + except ValueError as e: + print(f"Error: ADR-{old.number} not changed: {e}", file=sys.stderr) + return 1 + edits.append((old, edit)) + if section: + done = f"ADR-{new.number} amends ADR-{old.number} §{section}" + else: + done = f"ADR-{new.number} supersedes ADR-{old.number}; ADR-{old.number} is superseded" + code = _finish(edits, args.dry_run, done) + if not args.dry_run and not section: + print("Run `adr index -y` to refresh INDEX.md.") + return code + +# --- enact -------------------------------------------------------------------------- + +def cmd_enact(args): + """Mark an accepted cut or retire decision done at a commit (ADR-304 §5).""" + adr, code = _record_target(args.adr, 'enact') + if adr is None: + return code + verb = adr.frontmatter.get('verb') + if verb not in V1_ENACTING_VERBS: + print(f"Error: ADR-{adr.number} is a{'n' if str(verb)[:1] in 'aeiou' else ''} {verb or 'verbless'} " + f"record; enacted belongs to a cut or retire decision (ADR-304 §5).", file=sys.stderr) + return 1 + status = str(adr.status or '').lower() + if status != 'accepted': + print(f"Error: ADR-{adr.number} is {status}; enacted marks an accepted decision done.", file=sys.stderr) + return 1 + commit = args.commit.strip().lower() + if not re.fullmatch(r'[0-9a-f]{7,40}', commit): + print(f"Error: '{args.commit}' is not a commit hash (7 to 40 hex digits).", file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + previous = adr.frontmatter.get('enacted') + try: + edit.set('enacted', _Quoted(commit)) + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + note = f" (was {previous})" if previous and str(previous) != commit else '' + return _finish([(adr, edit)], args.dry_run, f"Enacted ADR-{adr.number} at {commit}{note}: {adr.title}") +def cmd_import(args): + """adr import scan | apply (ADR-306 §3).""" + if args.import_command == 'scan': + return _import_scan(args) + if args.import_command == 'apply': + return _import_apply(args) + print("Usage: adr import scan <paths...> | adr import apply [sheets...] [--partial]", file=sys.stderr) + return 2 + +SCAN_NAME_RE = re.compile(r'ADR-\d+(?:\.\d+)?(?:-.+)?\.md') +# The tool's own files at the top of a records folder, which are never +# records: its config, its index, and the sheet directory. +SCAN_TOOL_FILES = ('adr.yaml', 'INDEX.md', '.import') + +def _scan_inputs(paths: list) -> tuple: + """(record files, messages, [(path passed over, reason)], {extension: + count of other files passed over}). A directory yields its + ADR-NNN-*.md files. Its archived records are left out and counted, and + the tool's own files at its top level are left out. Every other + Markdown file, and each hidden directory, which is not entered, is + passed over by name; other files are counted by extension, so a folder + of images reads as one line. An archived record is scanned when it is + named.""" + files, messages, passed, other, archived = [], [], [], {}, 0 + for arg in paths: + path = Path(arg) + if path.is_dir(): + found_files, found_passed = [], [] + for folder, dirs, names in os.walk(path): + here = Path(folder) + top = here == path + for d in dirs: + if d.startswith('.') and not (top and d in SCAN_TOOL_FILES): + found_passed.append((here / d, 'hidden directory, not entered')) + dirs[:] = [d for d in dirs if not d.startswith('.')] + in_archive = 'archive' in here.relative_to(path).parts + for name in names: + found = here / name + if top and name in SCAN_TOOL_FILES: + continue + if SCAN_NAME_RE.fullmatch(name): + if in_archive: + archived += 1 + else: + found_files.append(found) + elif name.lower().endswith('.md'): + found_passed.append((found, 'not named ADR-NNN.md or ADR-NNN-<slug>.md')) + else: + ext = found.suffix.lower() + other[ext] = other.get(ext, 0) + 1 + files.extend(sorted(found_files)) + passed.extend(sorted(found_passed)) + elif path.is_file(): + files.append(path) + else: + messages.append(f"Skipped: {arg}: no such file or directory") + if archived: + messages.append(f"Note: {archived} archived record(s) left out; name a file to scan it") + return files, messages, passed, other + +def _import_scan(args): + """Write one sheet per record to docs/architecture/.import/ (ADR-306 §1, + §3). The directory ignores itself, so the repo's .gitignore is untouched. + A source is only read, never written. A sheet that differs from what a + fresh scan writes may hold edits, so it is kept unless --force. Each + Markdown file passed over is named with the reason, other files are + counted by extension, and all of them count as skipped.""" + files, messages, passed, other = _scan_inputs(args.paths) + for message in messages: + print(message) + for path, reason in passed: + print(f"Skipped: {relative_path(path.resolve()) if path.is_absolute() else path}: {reason}") + for ext, count in sorted(other.items()): + kind = f"{ext} file" if ext else "file with no extension" + print(f"Skipped: {count} {kind}{'s' if count != 1 else ''}: not Markdown") + out = import_dir() + written, skipped = 0, len(passed) + sum(other.values()) + claimed = {} + for path in files: + shown = relative_path(path.resolve()) if path.is_absolute() else path + try: + sheet = read_record(path) + text = dump_sheet(sheet) + except (SheetError, OSError) as e: + print(f"Skipped: {shown}: {e}") + skipped += 1 + continue + name = sheet_filename(sheet['target']['number']) + dest = out / name + source = sheet['source']['path'] + if name in claimed: + print(f"Skipped: {shown}: ADR-{format_number(sheet['target']['number'])} is also {claimed[name]}") + skipped += 1 + continue + forced = '' + if dest.exists() and dest.read_text() != text: + try: + previous = (load_sheet(dest).get('source') or {}).get('path') + except SheetError: + previous = None + if previous != source: + problem = f"{relative_path(dest)} holds a sheet for {previous or 'another source'}" + else: + problem = f"{relative_path(dest)} differs from a fresh scan (edited, or the source changed)" + if not args.force: + print(f"Skipped: {shown}: {problem}; --force overwrites it and discards its edits") + skipped += 1 + continue + forced = '; replaced an edited sheet, its edits are gone' + out.mkdir(parents=True, exist_ok=True) + ignore = out / '.gitignore' + if not ignore.exists(): + ignore.write_text('*\n') + dest.write_text(text) + claimed[name] = source + written += 1 + todo = len(sheet['todo']) + print(f"Scanned: {shown} -> {relative_path(dest)} ({sheet['source']['format']}, {todo} todo{forced})") + print(f"Scan: {written} sheet(s) written, {skipped} skipped") + return 1 if skipped and not written else 0 + +def _source_file(source: dict) -> Path: + path = Path(source['path']) + return path if path.is_absolute() else get_project_root() / path + +def _in_tree(path: Path) -> bool: + arch = get_project_root() / 'docs' / 'architecture' + try: + path.resolve().relative_to(arch.resolve()) + return True + except ValueError: + return False + +def _destination(sheet: dict, source_file: Optional[Path]) -> Path: + """Where apply writes. A source under docs/architecture is written in + place and keeps its number and domain: an import never renumbers (§5), + and moving a record is `adr domain move` (§6). Any other source becomes + a new file in the domain's folder, and its number must sit in the + domain's range; under adr/v0 an in-tree record's must too. Under adr/v1 + an in-tree record keeps its number wherever its folder is. Raises + SheetError when none of this holds.""" + target = sheet['target'] + number = format_number(target['number']) + domain = target.get('domain') + in_tree = source_file is not None and _in_tree(source_file) + if in_tree: + was = filename_number(source_file) + was_domain = source_domain(source_file, was or number) + if was != number.lstrip('0') or was_domain != domain: + raise SheetError( + f"the source is ADR-{was} in {was_domain}; the sheet says ADR-{number} in {domain}. " + f"A record keeps its number; moving it to another domain is `adr domain move` " + f"(ADR-306 §6). Restore target.number and target.domain") + if not domain: + raise SheetError('target.domain is empty') + span = number_range(domain) + if span is None: + raise SheetError(f"domain '{domain}' is not in adr.yaml; add it first " + f"(creating a domain on apply, ADR-306 §6, is not built yet)") + if not (in_tree and repo_contract() == V1) and not span[0] <= int(number.split('.')[0]) <= span[1]: + raise SheetError(f"ADR-{number} is outside the {domain} range {span[0]}-{span[1]}") + if in_tree: + return source_file + for existing in find_adrs(include_archived=True): + if filename_number(existing) == number.lstrip('0'): + raise SheetError(f"ADR-{number} already exists: {relative_path(existing)}") + folders = get_domains()[domain]['folder'] if domain in get_domains() else 'legacy' + folder = folders[0] if isinstance(folders, list) else folders + slug = re.sub(r'[^a-z0-9]+', '-', str(target['title']).lower()).strip('-') + return get_project_root() / 'docs' / 'architecture' / folder / f"ADR-{number}-{slug}.md" + +def _imported(source: dict, sheet: dict, raw: bytes) -> dict: + """`imported` for a record from a non-v1 source (ADR-306 §4): where it + came from, the source's own status as written, and the source keys with + no v1 field. The record keeps them whatever happens to the todo items. + A source outside the repo is named by its file name, so no machine's + path lands in the record.""" + path = source['path'] + imported = {'from': Path(path).name if Path(path).is_absolute() else path, + 'format': source.get('format')} + front = split_record(raw.decode('utf-8'))[0] + data = yaml.safe_load(front or '') or {} + imported['status'] = data.get('status') if isinstance(data, dict) else None + if sheet.get('unmapped'): + imported['unmapped'] = dict(sheet['unmapped']) + return imported + +def _uncommitted(path: Path) -> bool: + """The file has changes git has not committed. False outside git.""" + root = get_project_root() + try: + rel = str(path.resolve().relative_to(root.resolve())) + except ValueError: + return False + return bool((_git(['status', '--porcelain', '--', rel], root) or '').strip()) + +def _apply_one(sheet: dict, force: bool, undo: Optional[dict] = None) -> tuple: + """Write one sheet as a v1 record through render_record and return + (path, whether it changed). A non-v1 source gains `imported: {from, + format}`, plus `unmapped` when the source had keys with no v1 field, so + the record keeps them whatever happens to the todo (ADR-306 §1, §4). + Refused: a source changed since the scan, a source with a body whose + sheet has none, and a source with uncommitted changes unless --force. + With `undo`, the destination's prior bytes (None when it did not exist) + are kept there so a dry run can put them back.""" + source = sheet.get('source') + source_file = None + record = dict(sheet['record']) + if source: + source_file = _source_file(source) + if not source_file.is_file(): + raise SheetError(f"source {source['path']} is gone") + raw = source_file.read_bytes() + if hashlib.sha256(raw).hexdigest() != source.get('sha256'): + raise SheetError(f"source {source['path']} changed since the scan; scan it again") + if source.get('format') != 'v1' and 'imported' not in record: + record['imported'] = _imported(source, sheet, raw) + if read_record(source_file)['body'].strip() and not (sheet.get('body') or '').strip(): + raise SheetError('the source has a body and the sheet has none; scan it again') + dest = _destination(sheet, source_file) + text = render_record(dict(sheet, record=record)) + if dest.is_file() and _same_record(dest.read_text(), text): + return dest, False + if dest.is_file() and not force and _uncommitted(dest): + raise SheetError(f"{relative_path(dest)} has uncommitted changes; commit them, " + f"or --force to overwrite") + if undo is not None and dest not in undo: + undo[dest] = dest.read_bytes() if dest.is_file() else None + made = undo.setdefault('__dirs__', []) + folder = dest.parent + while not folder.exists(): + made.append(folder) + folder = folder.parent + dest.parent.mkdir(parents=True, exist_ok=True) + dest.write_text(text) + return dest, True + +def _same_record(old: str, new: str) -> bool: + """The same frontmatter data and the same text after it. A v1 record + applied unedited is left as its author formatted it (ADR-306 §7).""" + old_front, old_before, old_title, old_after = split_record(old) + new_front, new_before, new_title, new_after = split_record(new) + if old_front is None or old_title is None or new_title is None: + return False + try: + same_data = yaml.safe_load(old_front) == yaml.safe_load(new_front) + except yaml.YAMLError: + return False + return (same_data and old_before.strip() == new_before.strip() + and old_title.group(0) == new_title.group(0) and old_after == new_after) + +def _labels(todo: list) -> str: + """Todo items by label, counted: 'verb, basis, unmapped (2)'.""" + counts = {} + for item in todo: + counts[todo_label(item)] = counts.get(todo_label(item), 0) + 1 + return ', '.join(f"{label} ({n})" if n > 1 else label for label, n in counts.items()) + +def _import_apply(args): + """Write each finished sheet as a v1 record, then lint what was written + (ADR-306 §3). A sheet with open todo items is skipped. --partial writes + it anyway, unless an item is one lint could not find again afterwards + (BLOCKING_TODO). An applied sheet is removed: the record is what is + kept. One bad sheet is reported and the rest still apply. --dry-run + writes the records, lints them inside the corpus, prints each issue, + then restores every file it wrote and keeps every sheet.""" + if args.sheets: + paths = [Path(p) for p in args.sheets] + else: + paths = sorted(import_dir().glob('ADR-*.yaml')) if import_dir().is_dir() else [] + lines, written = [], [] + undo = {} if args.dry_run else None + applied = skipped = refused = 0 + counts, issues = {}, {} + try: + for path in paths: + shown = relative_path(path.resolve()) + try: + sheet = load_sheet(path) + todo = sheet.get('todo') or [] + held = blocking_todo(todo) if args.partial else todo + if held: + why = "todo item(s) --partial does not write past" if args.partial else 'open todo item(s)' + lines.append((f"Skipped: {shown}: {len(held)} {why}: {_labels(held)}", None)) + skipped += 1 + continue + dest, changed = _apply_one(sheet, args.force, undo) + except (SheetError, OSError, ValueError, TypeError, AttributeError, KeyError) as e: + lines.append((f"Refused: {shown}: {e}", None)) + refused += 1 + continue + if not args.dry_run: + try: + path.unlink() + except OSError as e: + lines.append((f"Note: {shown}: applied, but the sheet could not be removed: {e}", None)) + applied += 1 + written.append(dest.resolve()) + partial = (f", {len(todo)} todo left" if todo else '') + ('' if changed else ', unchanged') + lines.append((f"Applied: {shown} -> {relative_path(dest)}{partial}", dest.resolve())) + counts, issues = _lint_counts(written) + finally: + if undo is not None: + for failed in _restore(undo): + lines.append((f"Note: dry run could not restore {failed}", None)) + for text, dest in lines: + if dest is not None: + errors, warnings = counts.get(dest, (0, 0)) + text += f" (lint: {errors} errors, {warnings} warnings)" + if args.dry_run: + text = text.replace('Applied: ', 'Would apply: ', 1) + text += ''.join(f"\n {'error' if i.severity == 'error' else 'warning'}: {i.message}" + for i in issues.get(dest, [])) + print(text) + verb = 'would apply' if args.dry_run else 'applied' + print(f"Import: {applied} {verb}, {skipped} skipped, {refused} refused" + + (' (dry run: nothing written)' if args.dry_run else '')) + lint_errors = sum(e for e, _ in counts.values()) + return 1 if refused or (args.dry_run and lint_errors) else 0 + +def _restore(undo: dict) -> list: + """Put back every file a dry run wrote, and remove the directories it + made that are left empty. Each file is restored on its own, so one + failure doesn't stop the rest; the paths that failed are returned.""" + failed = [] + for dest, before in undo.items(): + if dest == '__dirs__': + continue + try: + if before is None: + dest.unlink(missing_ok=True) + else: + dest.write_bytes(before) + except OSError: + failed.append(relative_path(dest)) + for folder in sorted(undo.get('__dirs__', []), key=lambda d: len(d.parts), reverse=True): + try: + folder.rmdir() + except OSError: + pass + return failed + +def _lint_counts(paths: list) -> dict: + """(errors, warnings) per written record, and the issues themselves, + from the full rule set.""" + if not paths: + return {}, {} + corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) + targets = [adr for adr in corpus if adr.path.resolve() in set(paths)] + run_rules(targets, ctx) + counts = {adr.path.resolve(): (sum(i.severity == 'error' for i in adr.issues), + sum(i.severity != 'error' for i in adr.issues)) + for adr in targets} + return counts, {adr.path.resolve(): list(adr.issues) for adr in targets} # ============================================================================ # Main # ============================================================================ def main(): parser = argparse.ArgumentParser( - description='ADR - Architecture Decision Record CLI Tool', + description='ADR - Agent Decision Record CLI Tool', formatter_class=argparse.RawDescriptionHelpFormatter, epilog=__doc__ ) @@ -2617,6 +4548,15 @@ def main(): help='List archived ADRs only') list_scope.add_argument('--all', action='store_true', help='List active and archived ADRs') + p_list.add_argument('--field', action='append', metavar='KEY[=VALUE]', + help='Filter by frontmatter: KEY present, or KEY equal to or listing VALUE (repeatable)') + p_list.add_argument('--kind', help='Filter by kind (same as --field kind=KIND)') + p_list.add_argument('--verb', help='Filter by verb (same as --field verb=VERB)') + p_list.add_argument('--capability', help='Filter by capability, listed or single (same as --field capability=NAME)') + p_list.add_argument('--group-by', dest='group_by', metavar='KEY', + help='Group by a frontmatter field; a record listing several values is in each group') + p_list.add_argument('--json', action='store_true', + help='Machine output: number, title, path, status and frontmatter for each record') # view p_view = subparsers.add_parser('view', aliases=['v', 'show'], help='View an ADR') @@ -2626,6 +4566,11 @@ def main(): p_new = subparsers.add_parser('new', help='Create new ADR') p_new.add_argument('domain', help='Domain (see `adr domains` for list)') p_new.add_argument('title', help='ADR title') + p_new.add_argument('--kind', help='adr/v1: record kind (default: decision)') + p_new.add_argument('--verb', help='adr/v1: decision verb (add, cut, change, retire, constrain)') + p_new.add_argument('--capability', help='adr/v1: capability from the adr.yaml vocabulary') + p_new.add_argument('--agent', help='adr/v1: the agent writing the record (e.g. Claude)') + p_new.add_argument('--model', help='adr/v1: the model the agent runs on') # rename p_rename = subparsers.add_parser('rename', help='Rename an ADR title and/or file slug') @@ -2638,6 +4583,25 @@ def main(): p_lint.add_argument('paths', nargs='*', help='Specific files to lint') p_lint.add_argument('--check', action='store_true', help='Exit 1 if errors (CI mode)') + # import (ADR-306) + p_import = subparsers.add_parser('import', help='Import records through import sheets') + import_sub = p_import.add_subparsers(dest='import_command') + p_scan = import_sub.add_parser('scan', help='Write an import sheet for each record') + p_scan.add_argument('paths', nargs='+', help='Record files or directories of records') + p_scan.add_argument('--force', action='store_true', + help='Overwrite a sheet that differs from a fresh scan, discarding its edits') + p_apply = import_sub.add_parser('apply', help='Write finished sheets as adr/v1 records') + p_apply.add_argument('sheets', nargs='*', help='Sheets to apply (default: every sheet in .import/)') + p_apply.add_argument('--partial', action='store_true', + help='Apply sheets with open todo items too, except items lint cannot ' + 'find again afterwards: a status with no mapping, a Deprecated ' + 'note, a number or domain mismatch') + p_apply.add_argument('--force', action='store_true', + help='Overwrite a record that has uncommitted changes') + p_apply.add_argument('--dry-run', action='store_true', + help='Write, lint inside the corpus and print each issue, then restore ' + 'every file and keep every sheet') + # index index_parser = subparsers.add_parser('index', help='Generate ADR index') index_parser.add_argument('-y', '--yes', action='store_true', @@ -2646,6 +4610,26 @@ def main(): # domains subparsers.add_parser('domains', help='List domain number series') + # domain (ADR-306 §6) + p_domain = subparsers.add_parser('domain', help='Add, rename or move domains') + domain_sub = p_domain.add_subparsers(dest='domain_command') + p_dadd = domain_sub.add_parser('add', help='Add a domain to adr.yaml') + p_dadd.add_argument('name', help='Domain key (e.g. ops)') + p_dadd.add_argument('--range', required=True, help='Number range for new records, A-B (e.g. 400-499)') + p_dadd.add_argument('--folder', required=True, help='Folder under docs/architecture') + p_dadd.add_argument('--label', help='Display name (default: the key, capitalized)') + p_dadd.add_argument('--description', help='One line on what the domain covers') + p_dren = domain_sub.add_parser('rename', help='Rename a domain, and its folder, in place') + p_dren.add_argument('old', help='Current domain key') + p_dren.add_argument('new', help='New domain key') + p_dren.add_argument('--folder', help='New folder (default: the new key when the folder was the old key)') + p_dren.add_argument('--dry-run', action='store_true', help='Print each line the rename would change; write nothing') + p_dmove = domain_sub.add_parser('move', help='Move records to another domain; numbers never change') + p_dmove.add_argument('record', nargs='?', help='ADR number (e.g. 104, ADR-104)') + p_dmove.add_argument('domain', nargs='?', help='Target domain') + p_dmove.add_argument('--plan', help='YAML list of {record, domain} moves, applied together') + p_dmove.add_argument('--dry-run', action='store_true', help='Print the moves and each line they would change; write nothing') + # archive p_archive = subparsers.add_parser( 'archive', help='Archive an ADR out of the active set') @@ -2673,12 +4657,42 @@ def main(): p_close.add_argument('--reason', help='Why (required; appended as a Closure section)') p_close.add_argument('--dry-run', action='store_true', help='Report without writing') + # record edits (adr/v1): consider, set, supersede, enact + p_consider = subparsers.add_parser('consider', help="Append a considered entry: the operator's answer (ADR-304 §12)") + p_consider.add_argument('adr', help='ADR number (e.g., 101, ADR-101)') + p_consider.add_argument('--said', help='What the operator said, verbatim (required)') + p_consider.add_argument('--via', help='Where it was said, such as a PR or a session (required)') + p_consider.add_argument('--operator', help='Who said it (default: the gh or git user)') + p_consider.add_argument('--covers', nargs='*', metavar='PROBE', + help='Probe names from the Summary the answer covers (none given writes covers: [])') + p_consider.add_argument('--paraphrase', action='store_true', help='said is a summary, not the words') + p_consider.add_argument('--canary', choices=['caught', 'missed'], help='Whether the operator caught the canary') + p_consider.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_set = subparsers.add_parser('set', help='Edit frontmatter fields: key=value, key+=item, key-=item') + p_set.add_argument('adr', help='ADR number (e.g., 101, ADR-101)') + p_set.add_argument('assignments', nargs='+', metavar='key=value', + help='Values are YAML: capability=[a, b], related+=ADR-7, observable+="..."') + p_set.add_argument('--force', action='store_true', + help='Set status, or a field no v1 record carries, anyway') + p_set.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_supersede = subparsers.add_parser('supersede', help='Record a supersession on both records (ADR-304 §3)') + p_supersede.add_argument('adr', help='The record replaced (e.g., 101, ADR-101)') + p_supersede.add_argument('--by', required=True, help='The record that replaces it') + p_supersede.add_argument('--amends', metavar='SECTION', + help='Replace one section only: amends: [OLD#SECTION] on the new record') + p_supersede.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_enact = subparsers.add_parser('enact', help='Mark an accepted cut or retire done at a commit (ADR-304 §5)') + p_enact.add_argument('adr', help='ADR number (e.g., 111, ADR-111)') + p_enact.add_argument('commit', help='The commit hash that finished the removal') + p_enact.add_argument('--dry-run', action='store_true', help='Show the change without writing') + # cite p_cite = subparsers.add_parser('cite', help='Check ADR citations in code against the records') p_cite.add_argument('paths', nargs='*', help='Limit the scan to these files or directories') p_cite.add_argument('--check', action='store_true', help='Exit 1 if errors (CI mode)') - p_cite.add_argument('--no-inventory', action='store_true', - help="Skip surface inventories (they run shell commands from adr.yaml)") + # Accepted and ignored: surface inventories are gone (ADR-311), and a + # script that still passes the flag keeps working. + p_cite.add_argument('--no-inventory', action='store_true', help=argparse.SUPPRESS) args = parser.parse_args() @@ -2698,11 +4712,17 @@ def main(): 'index': cmd_index, 'archive': cmd_archive, 'domains': cmd_domains, + 'domain': cmd_domain, 'config': cmd_config, 'cite': cmd_cite, 'accept': cmd_accept, 'reject': cmd_reject, 'abandon': cmd_abandon, + 'import': cmd_import, + 'consider': cmd_consider, + 'set': cmd_set, + 'supersede': cmd_supersede, + 'enact': cmd_enact, } return commands[args.command](args) diff --git a/hooks/ways/documentation/adr/adr.md b/hooks/ways/documentation/adr/adr.md index 76b0c63f..2a090092 100644 --- a/hooks/ways/documentation/adr/adr.md +++ b/hooks/ways/documentation/adr/adr.md @@ -1,5 +1,5 @@ --- -description: Architecture Decision Records — creating, managing, and referencing ADRs for technical choices, how reversible a decision is, a deliberate deviation from a standard, and superseding an accepted ADR +description: Agent Decision Records (ADRs) — creating, managing, and referencing ADRs for technical choices, how reversible a decision is, a deliberate deviation from a standard, and superseding an accepted ADR vocabulary: adr architecture decision record design pattern technical choice trade-off rationale alternative reversibility reversible one-way irreversible deviation deviate depart standard exception waiver supersede superseded accepted defer pattern: (^| )adr( |$)|architect|decision|design.?pattern|technical.?choice|trade.?off files: docs/architecture/.*\.md$ @@ -22,7 +22,7 @@ refire: 0.15 The section above this way, when present, gives this project's ADR commands, record format and lifecycle. It depends on the tool the project vendored and the contract its `adr.yaml` declares (ADR-304 §10). Follow it over any habit from another project. -Projects define their own domains and ranges in `adr.yaml`; `adr domains` shows them. +Projects define their own domains and ranges in `adr.yaml`; `adr domains` shows them. A record's number is permanent. `adr domain add`, `rename` and `move` change the layout and rewrite every path to what moved, and they leave `ADR-N` citations as they are. ## Generalize the Decision @@ -39,7 +39,7 @@ A generalized ADR is reusable across everything that hits the same force. A disc ## What Counts as a Decision - **Doing nothing is a decision** when the alternatives were live. Record the destination, the trigger that starts the work, and what makes waiting safe. -- **An as-built observation does not qualify.** A detail reconstructed from source with no recorded rationale goes in a design note or README, marked "rationale not recorded". +- **An as-built observation does not qualify.** A detail reconstructed from source with no recorded rationale goes in an evidence record, a design note or a README, marked "rationale not recorded". - **A deliberate deviation qualifies.** When the project departs from a standard on purpose, the ADR names the standard, the reason, the scope (paths, environments, components), and the condition that ends it. Without the end condition a later session reads the departure as drift and fixes it. ## Reversibility diff --git a/hooks/ways/documentation/adr/adr.yaml.template b/hooks/ways/documentation/adr/adr.yaml.template index 60f1bfd8..cbd3437c 100644 --- a/hooks/ways/documentation/adr/adr.yaml.template +++ b/hooks/ways/documentation/adr/adr.yaml.template @@ -36,7 +36,39 @@ domains: description: Authentication, authorization, secrets folder: security -# Valid ADR statuses +# The record contract (ADR-304). With `contract: adr/v1`, records are typed: +# each declares a kind, and a decision names a verb, a capability and its +# basis. Delete this line and the v1 blocks below to stay on adr/v0. +contract: adr/v1 + +kinds: + decision: + verb: required + requires: [capability, basis, agent] + sections: [Summary] + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec, evidence] } + spec: + verb: forbidden + requires: [capability] + edges: { supersedes: spec, decided_by: decision } + evidence: + verb: forbidden + requires: [capability] + edges: { supersedes: evidence } + +# What the project does, one line each. A record's `capability` names one of +# these, and a decision adds, cuts, changes or constrains them. Replace the +# placeholder with the project's own list. +capabilities: + core: The project's core behaviour (placeholder; replace) + +# This repository, as URLs name it (host/owner/repo, or a list). A URL into +# it at a branch is a path in the tree: `adr domain` rewrites it when a +# record moves. Without this key the origin remote is used, so set it to +# rewrite the same URLs in every clone. +# repository: github.com/owner/repo + +# Valid ADR statuses (adr/v0 records) statuses: - Draft # Initial creation, not yet proposed - Proposed # Under discussion/review diff --git a/hooks/ways/documentation/adr/consider/consider.md b/hooks/ways/documentation/adr/consider/consider.md index 071edbd2..079036c6 100644 --- a/hooks/ways/documentation/adr/consider/consider.md +++ b/hooks/ways/documentation/adr/consider/consider.md @@ -19,14 +19,16 @@ The agent writes and proposes a decision; the operator considers it (ADR-304 §1 - Open with a Summary the operator can judge alone: what is decided, what it trades away and forecloses, whether it is one-way (said first when it is), the probes, and an inversion. - Probes are specific points put to the operator. Mix points you are confident on with points you are not, and label each. The mix is deliberate: probes prime the operator's judgement, and priming can bias it. - The inversion names the two ends the decision sits between and asks whether the answer lies outside your framing. +- When the decision carries `observable` entries, demonstrate them before asking, where you can: run the command, show the output or a screenshot, open the page. Say in `via` what the operator was shown (ADR-307). +- Ask the probes in the conversation, one question each, with enough context to answer without opening the record, and label which ones you are confident on. A probe that exists only in the record was never asked, so the answer to "looks good" covers none of them. With more than one probe or decision pending, ask them through the choice tool as one batch, one question each (choices(meta)). - A probe may be a canary: a point that is deliberately wrong, harmless if accepted, and a little whimsical. Reveal it right after the operator answers. It never stays in the record and is never about safety. ## When the answer comes - A short yes is a real answer. Take it as given and move on; do not re-ask the probes. -- Under adr/v1, record a `considered` entry: what was said and via which channel. `covers` lists the probes the answer settled; a bare "looks good" covers none. Add `canary: caught` or `missed` only when a canary was used. +- Under adr/v1, record a `considered` entry: what was said and via which channel. `covers` lists the probes the answer settled; a bare "looks good" covers none. Add `canary: caught` or `missed` only when a canary was used. `adr consider N --said "..." --via "..." --covers NAME...` writes the entry and refuses a probe name the Summary does not have. - If the canary was missed, say so once and constructively, offer a smaller set of probes, then proceed on the operator's answer. -- A decision the operator started waits for their consideration before `adr accept`. A decision with no operator basis, grounded in evidence, a standard or upstream, may be accepted by the agent directly. +- A decision the operator started waits for their consideration before `adr accept`. The tool does not check for it, so the wait is yours to keep. A decision with no operator basis, grounded in evidence, a standard or upstream, may be accepted by the agent directly. ## Asking in plain words Field names such as `basis`, `considered` and `level` belong in the record, not in what you say to the operator. Ask the way a teammate would, and give enough context that the question can be answered without opening the file. The shape runs from worst to best: diff --git a/hooks/ways/documentation/adr/macro.sh b/hooks/ways/documentation/adr/macro.sh index ef46b06a..59405c0e 100755 --- a/hooks/ways/documentation/adr/macro.sh +++ b/hooks/ways/documentation/adr/macro.sh @@ -63,37 +63,58 @@ print_v1_guide() { local s="$1" echo "## ADR Tooling (adr/v1)" echo "" - echo "This project's records follow the adr/v1 contract (ADR-304): ADR means Agent Decision Record. Kinds, capabilities and surfaces are declared in \`docs/architecture/adr.yaml\`." + echo "This project's records follow the adr/v1 contract (ADR-304): ADR means Agent Decision Record. Kinds and capabilities are declared in \`docs/architecture/adr.yaml\`." echo "" echo "| Command | Purpose |" echo "|---------|---------|" - echo "| \`$s new <domain> <title>\` | Create a record (then set its v1 frontmatter) |" + echo "| \`$s new <domain> <title> --kind K --verb V --capability C\` | Create a record with its v1 frontmatter (\`--agent\`, \`--model\` for a decision) |" echo "| \`$s lint [--check]\` | Validate records against the contract |" - echo "| \`$s accept <n>\` | Accept a proposed record; refuses if it would not lint clean |" + echo "| \`$s consider <n> --said \"...\" --via \"...\"\` | Append the operator's answer to a decision's probes |" + echo "| \`$s accept <n>\` | Accept a proposed record; refuses if the record fails its own field checks; \`$s lint\` checks its references |" echo "| \`$s reject <n> --reason \"...\"\` | Considered and declined |" echo "| \`$s abandon <n> --reason \"...\"\` | Dropped before a decision |" + echo "| \`$s set <n> key=value key+=item\` | Edit frontmatter; refuses status, which the lifecycle commands set |" + echo "| \`$s supersede <old> --by <new>\` | Record a supersession on both records |" + echo "| \`$s enact <n> <commit>\` | Mark an accepted cut or retire done |" + echo "| \`$s list --kind K --capability C --field KEY=VALUE --group-by KEY\` | Query the corpus (\`--json\` for scripts) |" echo "| \`$s cite [--check]\` | Check ADR-N citations in code against the records |" - echo "| \`$s list\`, \`view <n>\`, \`index -y\`, \`archive\`, \`domains\` | As before |" + echo "| \`$s import scan <paths>\`, \`import apply [--dry-run]\` | Convert v0 or foreign records through editable sheets (ADR-306) |" + echo "| \`$s domain add\`, \`rename\`, \`move\` | Change the layout; paths are rewritten, numbers never change |" + echo "| \`$s view <n>\`, \`index -y\`, \`archive\`, \`domains\` | As before |" + echo "" + echo "Use these commands for record edits; they check what hand edits skip." echo "" echo "### Record format" echo "" - echo "Every record declares \`contract: adr/v1\`, a \`kind\` (decision or spec, or what adr.yaml declares), a \`capability\` from the vocabulary, and a \`status\` (proposed, accepted, rejected, abandoned, superseded, archived)." + echo "Every record declares \`contract: adr/v1\`, a \`kind\` (decision, spec or evidence, or what adr.yaml declares), a \`capability\` from the vocabulary, and a \`status\` (proposed, accepted, rejected, abandoned, superseded, archived)." + echo "" + echo "A decision decides. A spec describes behaviour that is kept current and stays editable. Evidence records a finding, survey, audit, measurement or exploration at a point in time, and a decision cites it in its \`basis\` (ADR-309)." + echo "" + echo "A record's number is its permanent identity. Its folder under \`docs/architecture/\` is its area, and a domain's band only allocates new numbers (ADR-310)." echo "" - echo "A decision also carries a \`verb\` (add, cut, change, retire, constrain), a \`basis\`, and \`agent: {name, model}\`. Each basis entry names one source: operator, evidence, standard, upstream, or precedent. Following precedent must reach an external source." + echo "A decision also carries a \`verb\` (add, cut, change, retire, constrain), a \`basis\`, and \`agent: {name, model}\`. Each basis entry names one source: operator, evidence, standard, upstream, or precedent. A precedent names the record the decision rests on." echo "" - echo "An operator basis records \`level\` (authored, directed, guided), \`said\` and \`via\`. Write one only when the operator actually said it: quote written channels verbatim and mark a spoken one \`paraphrase: true\`. A decision the operator started waits for a \`considered\` entry before \`accept\`. When you ask the operator for either, use plain words, not the field names (adr/consider)." + echo "An operator basis records \`level\` (authored, directed, guided), \`said\` and \`via\`. Write one only when the operator actually said it: quote written channels verbatim and mark a spoken one \`paraphrase: true\`. A decision the operator started waits for a \`considered\` entry before you run \`accept\`, which does not check for one. When you ask the operator for either, use plain words, not the field names (adr/consider)." echo "" - echo "Records link through \`supersedes\`, \`amends: ADR-N#section\`, \`extends\` and \`decided_by\`. A change against a broader decision amends it." + echo "A decision may carry \`observable\`: a list of what should be seen, run or tried when it holds, as plain lines or mappings with whatever keys suit (ADR-307). It is optional, and may be added after acceptance. When drafting an add or change, ask the operator what should be observable; they may name it, decline, or leave the observing to you, in which case run the work and iterate until you can show it, then record what you used." echo "" - echo "A cut or retire is enacted once the removal lands: set \`enacted: \"<commit>\"\` (a quoted hash) on the accepted decision. Until then, \`cite\` warns on what still depends on it; after, it fails." + echo "Records link through \`supersedes\`, \`amends: ADR-N#section\`, \`extends\` and \`decided_by\`. A change names what it replaces: it supersedes the decision it replaces, or amends the section it changes in a broader one." + echo "" + echo "A cut or retire is enacted once the removal lands: \`$s enact <n> <commit>\` sets \`enacted: \"<commit>\"\` on the accepted decision." echo "" echo "### Summary" echo "" echo "Every decision opens with \`## Summary\`, written so someone who did not take part can judge it: what is decided, what it trades away, whether it is one-way (said first when it is), probes labelled confident and not confident, and an inversion naming the two ends the decision sits between." echo "" - echo "### Frozen once past proposed" + echo "### Conventions" + echo "" + echo "The tool checks each record's fields and sections, and that its references resolve. It does not check the conventions below. This way teaches them, and review catches a record that breaks them (ADR-311)." + echo "" + echo "- Correct an accepted record by appending. Leave what was decided as written." + echo "- Record a change in what the project does as a new decision that names what it replaces, through \`supersedes\` or \`amends\`." + echo "- Quote the operator in their own words." echo "" - echo "After a decision leaves proposed, only its kind's \`mutable_after_accept\` fields change, and the body grows only by appending. A change in what the project does is a new decision that amends or supersedes the old one." + echo "Git keeps every earlier version of a record and who changed it. Read them with \`git log -p\`." } # Outside a work tree there is nothing to vendor into; say nothing. diff --git a/hooks/ways/documentation/adr/migration/migration.md b/hooks/ways/documentation/adr/migration/migration.md index ed955124..a9f18e0c 100644 --- a/hooks/ways/documentation/adr/migration/migration.md +++ b/hooks/ways/documentation/adr/migration/migration.md @@ -1,6 +1,6 @@ --- -description: migrating to ADR tooling, adopting ADRs, converting existing decisions, setting up adr.yaml, bootstrapping architecture records -vocabulary: migrate adopt convert bootstrap setup greenfield legacy rename renumber frontmatter yaml scaffold import consolidate +description: migrating to ADR tooling, adopting ADRs, converting existing decisions, setting up adr.yaml, bootstrapping agent decision records +vocabulary: migrate adopt convert bootstrap setup greenfield legacy rename scan frontmatter yaml scaffold import consolidate scope: agent, subagent refire: 0.15 --- @@ -13,11 +13,11 @@ refire: 0.15 | State | Signs | Strategy | |-------|-------|----------| -| **Greenfield** | No ADRs, no `docs/architecture/` | Scaffold from scratch | -| **Flat directory** | ADRs exist in one dir, sequential numbering (0001, 0002...) | Park as legacy, adopt domains going forward | -| **Inline metadata** | `Status: Accepted` in markdown body, no YAML frontmatter | Add frontmatter, keep body | -| **Scattered** | Decision docs in various locations (wiki, README, etc.) | Consolidate into `docs/architecture/` | -| **Different tool** | Using adr-tools, Log4brains, or similar | Export and convert | +| **Greenfield** | No records, no `docs/architecture/` | Scaffold from scratch | +| **Flat directory** | Records in one dir, sequential numbering (0001, 0002...) | Rename to `ADR-NNN-*.md`, then `adr import scan`; or park as legacy | +| **v0 frontmatter** | YAML frontmatter with `status: Accepted` etc., no `kind`/`verb`/`capability`/`basis` | `adr import scan` | +| **Inline metadata** | `Status: Accepted` in the markdown body, no YAML frontmatter | Move the metadata into frontmatter, then `adr import scan` | +| **Other tools** | adr-tools, MADR, Log4brains, or similar | No reader yet — convert through a sheet by hand, or park as legacy | ## Greenfield Setup @@ -33,69 +33,51 @@ docs/scripts/adr list # Should show 0 ADRs The rest of this way is the migration-specific *why/when/what* the skill doesn't cover: which starting state you're in, and how to get existing decisions into the tooling without losing history. -## Flat Directory Migration +## Converting Existing Records -Existing ADRs like `docs/adr/0001-use-postgres.md` with sequential numbering. +Numbers are permanent identity (ADR-310) — conversion never renumbers or re-homes a record by domain range. A domain's range only allocates numbers for new records. 1. **Vendor the tooling** (greenfield step 1 — use the **adr** skill) -2. **Park existing ADRs as legacy** — don't renumber: +2. **Prepare what the scanner can't read.** `adr import scan` reads files named `ADR-NNN-*.md` that open with YAML frontmatter, and passes over anything else. Rename sequential files, keeping the number: ```bash -mkdir -p docs/architecture/legacy -git mv docs/adr/0001-*.md docs/architecture/legacy/ -# Rename to ADR-NNN format if needed: -git mv docs/architecture/legacy/0001-use-postgres.md docs/architecture/legacy/ADR-001-use-postgres.md # adr-cite-ignore: example number +git mv docs/adr/0001-use-postgres.md docs/adr/ADR-001-use-postgres.md # adr-cite-ignore: example number ``` - -3. **Set the legacy range** in `adr.yaml` to cover existing numbers: -```yaml -legacy: - range: [1, 99] - label: "Legacy (Pre-Domain Numbering)" -``` - -4. **Add frontmatter** to each legacy file (see frontmatter conversion below) - -5. **New ADRs use domains** — `docs/scripts/adr new core "Next Decision"` starts at 100+ - -## Inline Metadata Conversion - -ADRs with metadata in the markdown body instead of YAML frontmatter: - +Move inline metadata into frontmatter and delete the inline lines: ```markdown -# ADR-014: Use Postgres for Session State # Before (inline) - -Status: Accepted -Date: 2026-01-15 -Deciders: @alice, @bob -``` - -Convert to YAML frontmatter: - -```markdown ---- # After (frontmatter) +--- status: Accepted date: 2026-01-15 deciders: - alice - - bob related: [] --- -# ADR-014: Use Postgres for Session State +# ADR-001: Use Postgres for Session State <!-- adr-cite-ignore: example number --> +``` + +3. **Scan** the existing records into editable import sheets: +```bash +docs/scripts/adr import scan docs/adr/ # or a list of files ``` +This writes a sheet per record under `.import/`. -Remove the inline metadata lines from the body after moving them to frontmatter. Run `docs/scripts/adr lint` to verify the conversion. +4. **Edit the sheets** in `.import/` — resolve each sheet's open todo items (the v1 fields the scan could not infer, such as `kind`, `verb`, `capability` and `basis`). Apply skips a sheet with open items. -## Scattered Decisions +5. **Dry-run the apply** — writes the records into the corpus, lints them, prints each issue, then restores every file and keeps the sheets: +```bash +docs/scripts/adr import apply --dry-run +``` + +6. **Apply for real** once the sheets are clean: +```bash +docs/scripts/adr import apply +``` +`--partial` lands sheets that still carry open todo items, except ones lint can't re-find afterward (a status with no mapping, a Deprecated note, a number or domain mismatch). `--force` overwrites a record with uncommitted changes. -Decision records spread across wiki pages, README sections, or issue threads. +A corpus not worth converting can instead be parked as read-only history under `docs/architecture/legacy/`, with `legacy.range` in `adr.yaml` covering its numbers. The tool lists only files named `ADR-NNN-*.md` with frontmatter, so rename and add frontmatter as in step 2 for parked records to appear in `adr list` and resolve in `adr cite`. -1. **Scaffold the tooling** -2. **For each decision**: `docs/scripts/adr new <domain> "Title"` to get a proper template -3. **Copy the substance** — extract Context, Decision, Consequences from the original source -4. **Set status to `Accepted`** if the decision is already in effect -5. **Link back** — add a `related:` entry or comment pointing to the original source for provenance +For adr-tools, MADR, Log4brains, or other foreign formats — no reader exists yet — either copy each record's substance into a sheet by hand before applying, or skip conversion and park the corpus in `legacy/`. ## Writing adr.yaml diff --git a/hooks/ways/documentation/adr/src/MANIFEST b/hooks/ways/documentation/adr/src/MANIFEST index 1e89b8e6..edebd750 100644 --- a/hooks/ways/documentation/adr/src/MANIFEST +++ b/hooks/ways/documentation/adr/src/MANIFEST @@ -9,6 +9,10 @@ refs.py rules_v1.py rules_v1_basis.py rules_v1_integrity.py +relocate.py +sheet.py +record_edit.py +import_read.py cmd_list.py cmd_view.py cmd_new.py @@ -17,7 +21,10 @@ cmd_lint.py cmd_index.py cmd_archive.py cmd_domains.py +cmd_domain.py cmd_config.py cmd_cite.py cmd_lifecycle.py +cmd_record.py +cmd_import.py main.py diff --git a/hooks/ways/documentation/adr/src/cmd_cite.py b/hooks/ways/documentation/adr/src/cmd_cite.py index 9a934c1a..5371c5ad 100644 --- a/hooks/ways/documentation/adr/src/cmd_cite.py +++ b/hooks/ways/documentation/adr/src/cmd_cite.py @@ -7,13 +7,6 @@ def cmd_cite(args): - a citation whose records are all out of force (superseded, deprecated, rejected, abandoned, archived) warns and names each status and successor - a citation of a proposed or draft record prompts acceptance - - under adr/v1, a citation of a record on a capability whose latest - accepted add-or-cut is a cut warns until the cut is enacted and fails - after (§5) - - under adr/v1, an accepted retire's targets are checked against the - surface inventory commands in adr.yaml (§5). Inventory commands are - shell commands from adr.yaml; they run only on a whole-repo scan and - never with --no-inventory. A bare ADR-N resolves to its family {N, N.1, ...}. Files come from `git ls-files`, so ignored paths are never read. docs/architecture and the @@ -43,8 +36,6 @@ def cmd_cite(args): return 2 records = get_all_adrs(include_archived=True) - ctx = LintContext.from_corpus(records) - v1 = ctx.contract == V1 # Lookup tables, built once: full number -> records, base -> family by_full, by_base = {}, {} @@ -65,23 +56,6 @@ def successors(adr) -> str: def is_proposed(adr) -> bool: return str(adr.status or '').lower() in ('proposed', 'draft') - # The state of each capability is its latest accepted add or cut, by date - # then number (ADR-304 §3). Only a capability whose latest is a cut counts. - cuts = {} - if v1: - latest = {} - for adr in records: - verb = adr.frontmatter.get('verb') - if (is_v1_record(adr, ctx) and verb in ('add', 'cut') - and str(adr.status or '').lower() == 'accepted'): - order = decision_order(adr) - for capability in capability_scope(adr): - if capability not in latest or order >= latest[capability][0]: - latest[capability] = (order, adr) - for capability, (_, adr) in latest.items(): - if adr.frontmatter.get('verb') == 'cut': - cuts[capability] = (adr, bool(adr.frontmatter.get('enacted'))) - def excluded(name: str) -> bool: for pattern in excludes: if any(c in pattern for c in '*?['): @@ -123,47 +97,6 @@ def in_scope(name: str) -> bool: findings.append(('warning', where, f"{cited} is not in force ({named})")) elif any(is_proposed(m) for m in members): findings.append(('warning', where, f"{cited} is still proposed; accept it or cite the record in force")) - for member in members: - if not is_v1_record(member, ctx): - continue - for capability in capability_scope(member): - if capability in cuts and cuts[capability][0] is not member: - cut, enacted = cuts[capability] - if enacted: - findings.append(('error', where, f"{cited} is on '{capability}', cut and enacted by ADR-{cut.number}; remove the citation")) - else: - findings.append(('warning', where, f"{cited} is on '{capability}', which ADR-{cut.number} cuts; remove before enacting")) - - # Retire targets against the surface inventories (§5). Inventories run - # shell commands from adr.yaml, so only on a whole-repo scan. - if v1 and not scopes and not args.no_inventory: - surfaces = v1_surfaces(ctx) - listed_by, failed = {}, set() - for adr in records: - if (not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'retire' - or str(adr.status or '').lower() != 'accepted'): - continue - enacted = bool(adr.frontmatter.get('enacted')) - for target in as_entries(adr.frontmatter.get('targets')): - namespace, _, target_name = target.partition(':') - command = _mapping(surfaces.get(namespace)).get('inventory') - if not command: - continue - if namespace not in listed_by: - listed_by[namespace] = _inventory(str(command), root) - listed, failure = listed_by[namespace] - where = str(relative_path(adr.path)) - if listed is None: - if namespace not in failed: # once per namespace, not per target - failed.add(namespace) - findings.append(('warning', where, f"inventory for '{namespace}' did not run ({failure}): {command}")) - elif target_name in listed: - if enacted: - findings.append(('error', where, f"{target} is still present after ADR-{adr.number} was enacted")) - else: - findings.append(('warning', where, f"{target} is still present; ADR-{adr.number} retires it")) - elif not enacted: - findings.append(('warning', where, f"{target} is not in the '{namespace}' inventory; check the name")) errors = sum(1 for f in findings if f[0] == 'error') warnings = len(findings) - errors @@ -201,25 +134,3 @@ def _cite_files(root: Path) -> list: for filename in filenames: names.append((Path(dirpath) / filename).relative_to(root).as_posix()) return sorted(names) - -def _inventory(command: str, root: Path) -> tuple: - """Run a surface inventory command from adr.yaml. Returns (names, None), - or (None, reason) when it failed. One name per line of output.""" - import signal - try: - proc = subprocess.Popen(command, shell=True, cwd=root, stdin=subprocess.DEVNULL, - stdout=subprocess.PIPE, stderr=subprocess.PIPE, - text=True, start_new_session=True) - except OSError as e: - return None, str(e) - try: - stdout, stderr = proc.communicate(timeout=60) - except subprocess.TimeoutExpired: - os.killpg(proc.pid, signal.SIGKILL) # the shell and anything it started - proc.communicate() - return None, 'timed out after 60s' - if proc.returncode != 0: - detail = (stderr.strip().splitlines() or [''])[-1][:120] - return None, f"exit {proc.returncode}" + (f": {detail}" if detail else '') - return {line.strip() for line in stdout.splitlines() if line.strip()}, None - diff --git a/hooks/ways/documentation/adr/src/cmd_domain.py b/hooks/ways/documentation/adr/src/cmd_domain.py new file mode 100644 index 00000000..61b3e4c5 --- /dev/null +++ b/hooks/ways/documentation/adr/src/cmd_domain.py @@ -0,0 +1,446 @@ +# ============================================================================ +# adr domain: add, rename, move (ADR-306 §6) +# ============================================================================ +# +# The domain layout changes as a corpus grows. A record's number is its +# permanent identity: it is never renumbered and never reused. Under adr/v1 a +# record's folder decides its domain, and a domain's range only allocates the +# numbers of new records. So moving a record moves its file and rewrites every +# path to it, and ADR-N citations stay as they are. +# +# adr.yaml is edited as text, so its comments and layout survive. + +DOMAIN_NAME_RE = re.compile(r'[a-z][a-z0-9_-]*') +FOLDER_NAME_RE = re.compile(r'[A-Za-z0-9][\w.-]*') + + +def cmd_domain(args): + handlers = {'add': _domain_add, 'rename': _domain_rename, 'move': _domain_move} + if args.domain_command not in handlers: + print("usage: adr domain {add,rename,move} ...", file=sys.stderr) + return 2 + return handlers[args.domain_command](args) + + +# --- the files a rewrite reaches ---------------------------------------------------- + +def _relocation_scope(root: Path) -> tuple: + """(names in scope, every tracked path and directory). Scope is every file + git tracks except fixtures, import sheets, the archive (archived records + are never edited), adr.yaml's cite.exclude list, and INDEX.md.""" + import fnmatch + tracked, known = tracked_paths(root) + cite_config = get_config().get('cite') + patterns = ['tests/fixtures', 'docs/architecture/archive'] + [ + str(p) for p in ((cite_config.get('exclude') if isinstance(cite_config, dict) else None) or [])] + + def excluded(name: str) -> bool: + if '.import' in name.split('/'): + return True + for pattern in patterns: + if any(c in pattern for c in '*?['): + if fnmatch.fnmatch(name, pattern): + return True + elif name == pattern.rstrip('/') or name.startswith(pattern.rstrip('/') + '/'): + return True + return False + # INDEX.md is regenerated after the move, so it is not rewritten. A + # symlink is skipped: writing through it would edit its target, which is + # in scope (or excluded) in its own right. docs/scripts/adr is one. + return [n for n in tracked if not excluded(n) and n != 'docs/architecture/INDEX.md' + and not (root / n).is_symlink()], known + + +def _plan_rewrites(relocation: Relocation, root: Path, names: list) -> list: + """(name, count, new bytes, [(line number, before, after)]) for each file + whose text changes. Binary files and files that are not UTF-8 are never + touched.""" + changes = [] + for name in names: + path = root / name + try: + raw = path.read_bytes() + except OSError: + continue + if b'\0' in raw[:8192]: + continue + try: + text = raw.decode('utf-8') + except UnicodeDecodeError: + continue + new, count = relocation.text(text, name) + if count and new != text: + changes.append((name, count, new.encode('utf-8'), _changed_lines(text, new))) + return changes + + +def _changed_lines(before: str, after: str) -> list: + """[(line number, before, after)] for each line the rewrite changed. A + rewrite changes text within lines, so the lines pair up.""" + return [(i, a, b) for i, (a, b) in enumerate(zip(before.split('\n'), after.split('\n')), 1) if a != b] + + +def _print_rewrites(changes: list, verb: str, lines: bool = False) -> None: + """The count of rewritten paths per file, and with lines=True each line + the rewrite changes, before and after.""" + total = sum(c[1] for c in changes) + print(f"{verb} {total} path{'s' if total != 1 else ''} in {len(changes)} file{'s' if len(changes) != 1 else ''}") + for name, count, _, changed in changes: + print(f" {name}: {count}") + if lines: + for number, before, after in changed: + print(f" {name}:{number}") + print(f" {before.rstrip(chr(13))}") + print(f" → {after.rstrip(chr(13))}") + + +def _git_mv(root: Path, old: str, new: str) -> None: + (root / new).parent.mkdir(parents=True, exist_ok=True) + try: + result = subprocess.run(['git', 'mv', old, new], cwd=root, + capture_output=True, text=True, timeout=30) + if result.returncode == 0: + return + except (FileNotFoundError, subprocess.TimeoutExpired): + pass + (root / old).rename(root / new) + + +def _write_rewrites(relocation: Relocation, root: Path, changes: list) -> None: + for name, _, data, _ in changes: + (root / (relocation.target(name) or name)).write_bytes(data) + + +def _refresh_index(root: Path) -> None: + """Regenerate INDEX.md when the project keeps one.""" + if (root / 'docs' / 'architecture' / 'INDEX.md').exists(): + cmd_index(argparse.Namespace(yes=True)) + + +# --- adr.yaml as text --------------------------------------------------------------- + +def _yaml_scalar(value: str) -> str: + if value and re.fullmatch(r'[A-Za-z0-9(][\w ,.()/+-]*', value) and value.strip() == value \ + and value.lower() not in ('yes', 'no', 'true', 'false', 'on', 'off', 'null', '~'): + return value + return '"' + value.replace('\\', '\\\\').replace('"', '\\"') + '"' + + +def _domains_block(lines: list) -> Optional[tuple]: + """(start, end) of the domains block: the `domains:` line and the index + just past its last indented line. None when there is no block form.""" + start = next((i for i, line in enumerate(lines) + if re.fullmatch(r'domains:\s*(#.*)?', line.rstrip('\r\n'))), None) + if start is None: + return None + end = start + 1 + while end < len(lines) and (not lines[end].strip() or lines[end][0] in ' \t'): + end += 1 + return start, end + + +def _entry_indent(lines: list, start: int, end: int) -> tuple: + """(entry indent, field indent) used in the domains block.""" + entry = field_ = None + for line in lines[start + 1:end]: + if not line.strip() or line.lstrip().startswith('#'): + continue + indent = line[:len(line) - len(line.lstrip())] + if entry is None: + entry = indent + elif len(indent) > len(entry): + field_ = indent + break + entry = entry if entry is not None else ' ' + return entry, field_ if field_ is not None else entry * 2 + + +def _save_config(text: str) -> bool: + """Write adr.yaml if the edited text still reads as YAML.""" + try: + data = yaml.safe_load(text) + except yaml.YAMLError as e: + print(f"Error: the edited adr.yaml does not parse ({e}); nothing written", file=sys.stderr) + return False + if not isinstance(data, dict) or not isinstance(data.get('domains'), dict): + print("Error: the edited adr.yaml lost its domains; nothing written", file=sys.stderr) + return False + get_config_path().write_text(text) + reload_config() + return True + + +def _folders(config: dict) -> list: + folders = config.get('folder') + return list(folders) if isinstance(folders, list) else [folders] + + +# --- add ---------------------------------------------------------------------------- + +def _domain_add(args): + name, folder = args.name, args.folder + domains = get_domains() + problems = [] + if not DOMAIN_NAME_RE.fullmatch(name) or name in ('legacy', 'archive'): + problems.append(f"'{name}' is not a usable domain name (lowercase letters, digits, - and _; not legacy or archive)") + elif name in domains: + problems.append(f"domain '{name}' already exists") + m = re.fullmatch(r'\s*(\d+)\s*-\s*(\d+)\s*', args.range or '') + low = high = None + if not m: + problems.append(f"--range '{args.range}' is not A-B") + else: + low, high = int(m.group(1)), int(m.group(2)) + if low > high: + problems.append(f"--range {low}-{high} runs backwards") + if not FOLDER_NAME_RE.fullmatch(folder) or folder == 'archive': + problems.append(f"--folder '{folder}' is not a folder name under docs/architecture (and not archive)") + for other, config in domains.items(): + if folder in _folders(config): + problems.append(f"folder '{folder}' already belongs to {other}") + if low is not None and low <= high: + spans = [(other, config['range']) for other, config in domains.items()] + if 'legacy' in get_config(): + spans.append(('legacy', get_legacy_range())) + for other, (a, b) in spans: + if low <= b and a <= high: + problems.append(f"range {low}-{high} overlaps {other} ({a}-{b})") + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + text = get_config_path().read_text() + lines = text.splitlines(keepends=True) + block = _domains_block(lines) + if block is None: + print("Error: adr.yaml has no block-style domains: section to add to; edit it by hand", file=sys.stderr) + return 1 + start, end = block + entry, field_ = _entry_indent(lines, start, end) + last = max((i for i in range(start, end) if lines[i].strip()), default=start) + if not lines[last].endswith('\n'): + lines[last] += '\n' + label = args.label or name.capitalize() + new = ['\n', f"{entry}{name}:\n", + f"{field_}range: [{low}, {high}]\n", + f"{field_}name: {_yaml_scalar(label)}\n", + f"{field_}description: {_yaml_scalar(args.description or '')}\n", + f"{field_}folder: {_yaml_scalar(folder)}\n"] + if last == start: + new = new[1:] + lines[last + 1:last + 1] = new + if not _save_config(''.join(lines)): + return 1 + print(f"Added domain: {name} ({low}-{high})") + print(f" Name: {label}") + print(f" Folder: docs/architecture/{folder}/") + return 0 + + +# --- rename ------------------------------------------------------------------------- + +def _domain_rename(args): + old, new = args.old, args.new + domains = get_domains() + if old not in domains: + print(f"Error: Unknown domain '{old}'", file=sys.stderr) + print(f"Valid domains: {', '.join(domains.keys())}", file=sys.stderr) + return 1 + config = domains[old] + folders = _folders(config) + problems = [] + if new != old: + if not DOMAIN_NAME_RE.fullmatch(new) or new in ('legacy', 'archive'): + problems.append(f"'{new}' is not a usable domain name (lowercase letters, digits, - and _; not legacy or archive)") + elif new in domains: + problems.append(f"domain '{new}' already exists") + if args.folder and len(folders) > 1: + problems.append(f"{old} spans several folders ({', '.join(folders)}); rename them by hand") + # The folder follows the name when it matched the name; otherwise it + # stays, unless --folder names one. + folder_old = folders[0] + folder_new = args.folder or (new if len(folders) == 1 and folder_old == old else folder_old) + root = get_project_root() + arch = root / 'docs' / 'architecture' + if folder_new != folder_old: + if not FOLDER_NAME_RE.fullmatch(folder_new) or folder_new == 'archive': + problems.append(f"folder '{folder_new}' is not a folder name under docs/architecture (and not archive)") + for other, other_config in domains.items(): + if other != old and folder_new in _folders(other_config): + problems.append(f"folder '{folder_new}' already belongs to {other}") + if (arch / folder_new).exists(): + problems.append(f"docs/architecture/{folder_new} already exists") + if new == old and folder_new == folder_old: + problems.append("nothing to rename: give a new name or --folder") + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + text = get_config_path().read_text() + lines = text.splitlines(keepends=True) + block = _domains_block(lines) + key_at = None + if block: + start, end = block + entry, _ = _entry_indent(lines, start, end) + key_re = re.compile(rf'{re.escape(entry)}{re.escape(old)}:(\s*(?:#.*)?)') + key_at = next((i for i in range(start + 1, end) if key_re.fullmatch(lines[i].rstrip('\r\n'))), None) + if key_at is None: + print(f"Error: cannot find the '{old}:' entry in adr.yaml's domains block; edit it by hand", file=sys.stderr) + return 1 + ending = lines[key_at][len(lines[key_at].rstrip('\r\n')):] + lines[key_at] = f"{entry}{new}:" + key_re.fullmatch(lines[key_at].rstrip('\r\n')).group(1) + ending + if folder_new != folder_old: + stop = next((i for i in range(key_at + 1, end) + if lines[i].strip() and len(lines[i]) - len(lines[i].lstrip()) <= len(entry)), end) + folder_re = re.compile(r'(\s+folder:\s*)(\S+?)(\s*(?:#.*)?)') + for i in range(key_at + 1, stop): + fm = folder_re.fullmatch(lines[i].rstrip('\r\n')) + if fm: + ending_i = lines[i][len(lines[i].rstrip('\r\n')):] + lines[i] = fm.group(1) + _yaml_scalar(folder_new) + fm.group(3) + ending_i + break + else: + print(f"Error: cannot find {old}'s folder: line in adr.yaml; edit it by hand", file=sys.stderr) + return 1 + + names, known = _relocation_scope(root) + old_dir = f"docs/architecture/{folder_old}" + new_dir = f"docs/architecture/{folder_new}" + relocation = Relocation(dirs={old_dir: new_dir} if folder_new != folder_old else {}, + known=known, domain=(old, new) if new != old else None, + repo=Repository(root)) + changes = _plan_rewrites(relocation, root, names) + # adr.yaml is written from its edited text, with its own paths + # rewritten. Its listing is taken against the text as it was, so the + # key and folder lines show, and each counts as one rewrite. + config_rel = 'docs/architecture/adr.yaml' + config_edited = ''.join(lines) + config_text, config_count = relocation.text(config_edited, config_rel) + changes = [c for c in changes if c[0] != config_rel] + if config_text != text: + edited = sum(a != b for a, b in zip(text.split('\n'), config_edited.split('\n'))) + changes = sorted(changes + [(config_rel, config_count + edited, config_text.encode('utf-8'), + _changed_lines(text, config_text))]) + + verb = 'Would rename' if args.dry_run else 'Renamed' + print(f"{verb} domain: {old} → {new}") + if folder_new != folder_old: + print(f" Folder: {old_dir}/ → {new_dir}/") + _print_rewrites(changes, 'Would rewrite' if args.dry_run else 'Rewrote', lines=args.dry_run) + if args.dry_run: + print("Dry run: nothing written.") + return 0 + if not _save_config(config_text): + return 1 + if folder_new != folder_old and (root / old_dir).exists(): + _git_mv(root, old_dir, new_dir) + _write_rewrites(relocation, root, [c for c in changes if c[0] != config_rel]) + _refresh_index(root) + return 0 + + +# --- move --------------------------------------------------------------------------- + +def _load_plan(path: str) -> Optional[list]: + """The moves a plan file lists: [{record, domain}, ...].""" + try: + data = yaml.safe_load(Path(path).read_text()) + except (OSError, yaml.YAMLError) as e: + print(f"Error: cannot read plan {path}: {e}", file=sys.stderr) + return None + if isinstance(data, dict) and 'moves' in data: + data = data['moves'] + if not isinstance(data, list) or not data: + print(f"Error: plan {path}: expected a list of {{record, domain}} entries", file=sys.stderr) + return None + moves = [] + for i, entry in enumerate(data, 1): + if not isinstance(entry, dict) or 'record' not in entry or 'domain' not in entry: + print(f"Error: plan entry {i}: expected {{record, domain}}", file=sys.stderr) + return None + extra = sorted(set(entry) - {'record', 'domain'}) + if extra: + hint = '; a record keeps its number' if 'number' in extra else '' + print(f"Error: plan entry {i}: unknown key{'s' if len(extra) > 1 else ''} {', '.join(extra)}{hint}", file=sys.stderr) + return None + moves.append((str(entry['record']), str(entry['domain']))) + return moves + + +def _domain_move(args): + if repo_contract() != 'adr/v1': + print("Error: under adr/v0 a record's number decides its domain, so a record cannot " + "change domain by folder; declare contract: adr/v1 in adr.yaml first", file=sys.stderr) + return 1 + if args.plan: + if args.record or args.domain: + print("Error: give either <record> <domain> or --plan, not both", file=sys.stderr) + return 2 + moves = _load_plan(args.plan) + if moves is None: + return 1 + else: + if not args.record or not args.domain: + print("Error: adr domain move <record> <domain>, or --plan <file.yaml>", file=sys.stderr) + return 2 + moves = [(args.record, args.domain)] + + root = get_project_root() + domains = get_domains() + records = get_all_adrs(include_archived=True) + problems, files, lines_out, seen = [], {}, [], set() + for ref, domain in moves: + matches = find_by_ref(ref, records) + if not matches: + problems.append(f"ADR not found: {ref}") + continue + if len(matches) > 1: + problems.append(f"{ref} matches several records: " + ', '.join(str(relative_path(a.path, root)) for a in matches)) + continue + adr = matches[0] + rel = relative_path(adr.path, root).as_posix() + if rel in seen: + problems.append(f"ADR-{adr.number} is listed twice") + continue + seen.add(rel) + if is_archived(adr.path): + problems.append(f"ADR-{adr.number} is archived ({rel}); archived records are kept as they are") + continue + if domain not in domains: + problems.append(f"unknown domain '{domain}' (valid: {', '.join(domains.keys())})") + continue + folders = _folders(domains[domain]) + if adr.domain == domain and adr.path.parent.name in folders: + problems.append(f"ADR-{adr.number} is already in {domain} ({rel})") + continue + new_rel = f"docs/architecture/{folders[0]}/{adr.path.name}" + if (root / new_rel).exists(): + problems.append(f"{new_rel} already exists") + continue + files[rel] = new_rel + lines_out.append((adr, adr.domain or 'legacy', domain, rel, new_rel)) + if problems: + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + return 1 + + names, known = _relocation_scope(root) + relocation = Relocation(files=files, known=known, repo=Repository(root)) + changes = _plan_rewrites(relocation, root, names) + + verb = 'Would move' if args.dry_run else 'Moved' + for adr, was, domain, rel, new_rel in lines_out: + print(f"{verb}: ADR-{adr.number} {was} → {domain}") + print(f" {rel} → {new_rel}") + _print_rewrites(changes, 'Would rewrite' if args.dry_run else 'Rewrote', lines=args.dry_run) + if args.dry_run: + print("Dry run: nothing written.") + return 0 + for rel, new_rel in files.items(): + _git_mv(root, rel, new_rel) + _write_rewrites(relocation, root, changes) + _refresh_index(root) + return 0 diff --git a/hooks/ways/documentation/adr/src/cmd_import.py b/hooks/ways/documentation/adr/src/cmd_import.py new file mode 100644 index 00000000..3b22e7d4 --- /dev/null +++ b/hooks/ways/documentation/adr/src/cmd_import.py @@ -0,0 +1,355 @@ +def cmd_import(args): + """adr import scan | apply (ADR-306 §3).""" + if args.import_command == 'scan': + return _import_scan(args) + if args.import_command == 'apply': + return _import_apply(args) + print("Usage: adr import scan <paths...> | adr import apply [sheets...] [--partial]", file=sys.stderr) + return 2 + +SCAN_NAME_RE = re.compile(r'ADR-\d+(?:\.\d+)?(?:-.+)?\.md') +# The tool's own files at the top of a records folder, which are never +# records: its config, its index, and the sheet directory. +SCAN_TOOL_FILES = ('adr.yaml', 'INDEX.md', '.import') + +def _scan_inputs(paths: list) -> tuple: + """(record files, messages, [(path passed over, reason)], {extension: + count of other files passed over}). A directory yields its + ADR-NNN-*.md files. Its archived records are left out and counted, and + the tool's own files at its top level are left out. Every other + Markdown file, and each hidden directory, which is not entered, is + passed over by name; other files are counted by extension, so a folder + of images reads as one line. An archived record is scanned when it is + named.""" + files, messages, passed, other, archived = [], [], [], {}, 0 + for arg in paths: + path = Path(arg) + if path.is_dir(): + found_files, found_passed = [], [] + for folder, dirs, names in os.walk(path): + here = Path(folder) + top = here == path + for d in dirs: + if d.startswith('.') and not (top and d in SCAN_TOOL_FILES): + found_passed.append((here / d, 'hidden directory, not entered')) + dirs[:] = [d for d in dirs if not d.startswith('.')] + in_archive = 'archive' in here.relative_to(path).parts + for name in names: + found = here / name + if top and name in SCAN_TOOL_FILES: + continue + if SCAN_NAME_RE.fullmatch(name): + if in_archive: + archived += 1 + else: + found_files.append(found) + elif name.lower().endswith('.md'): + found_passed.append((found, 'not named ADR-NNN.md or ADR-NNN-<slug>.md')) + else: + ext = found.suffix.lower() + other[ext] = other.get(ext, 0) + 1 + files.extend(sorted(found_files)) + passed.extend(sorted(found_passed)) + elif path.is_file(): + files.append(path) + else: + messages.append(f"Skipped: {arg}: no such file or directory") + if archived: + messages.append(f"Note: {archived} archived record(s) left out; name a file to scan it") + return files, messages, passed, other + +def _import_scan(args): + """Write one sheet per record to docs/architecture/.import/ (ADR-306 §1, + §3). The directory ignores itself, so the repo's .gitignore is untouched. + A source is only read, never written. A sheet that differs from what a + fresh scan writes may hold edits, so it is kept unless --force. Each + Markdown file passed over is named with the reason, other files are + counted by extension, and all of them count as skipped.""" + files, messages, passed, other = _scan_inputs(args.paths) + for message in messages: + print(message) + for path, reason in passed: + print(f"Skipped: {relative_path(path.resolve()) if path.is_absolute() else path}: {reason}") + for ext, count in sorted(other.items()): + kind = f"{ext} file" if ext else "file with no extension" + print(f"Skipped: {count} {kind}{'s' if count != 1 else ''}: not Markdown") + out = import_dir() + written, skipped = 0, len(passed) + sum(other.values()) + claimed = {} + for path in files: + shown = relative_path(path.resolve()) if path.is_absolute() else path + try: + sheet = read_record(path) + text = dump_sheet(sheet) + except (SheetError, OSError) as e: + print(f"Skipped: {shown}: {e}") + skipped += 1 + continue + name = sheet_filename(sheet['target']['number']) + dest = out / name + source = sheet['source']['path'] + if name in claimed: + print(f"Skipped: {shown}: ADR-{format_number(sheet['target']['number'])} is also {claimed[name]}") + skipped += 1 + continue + forced = '' + if dest.exists() and dest.read_text() != text: + try: + previous = (load_sheet(dest).get('source') or {}).get('path') + except SheetError: + previous = None + if previous != source: + problem = f"{relative_path(dest)} holds a sheet for {previous or 'another source'}" + else: + problem = f"{relative_path(dest)} differs from a fresh scan (edited, or the source changed)" + if not args.force: + print(f"Skipped: {shown}: {problem}; --force overwrites it and discards its edits") + skipped += 1 + continue + forced = '; replaced an edited sheet, its edits are gone' + out.mkdir(parents=True, exist_ok=True) + ignore = out / '.gitignore' + if not ignore.exists(): + ignore.write_text('*\n') + dest.write_text(text) + claimed[name] = source + written += 1 + todo = len(sheet['todo']) + print(f"Scanned: {shown} -> {relative_path(dest)} ({sheet['source']['format']}, {todo} todo{forced})") + print(f"Scan: {written} sheet(s) written, {skipped} skipped") + return 1 if skipped and not written else 0 + +def _source_file(source: dict) -> Path: + path = Path(source['path']) + return path if path.is_absolute() else get_project_root() / path + +def _in_tree(path: Path) -> bool: + arch = get_project_root() / 'docs' / 'architecture' + try: + path.resolve().relative_to(arch.resolve()) + return True + except ValueError: + return False + +def _destination(sheet: dict, source_file: Optional[Path]) -> Path: + """Where apply writes. A source under docs/architecture is written in + place and keeps its number and domain: an import never renumbers (§5), + and moving a record is `adr domain move` (§6). Any other source becomes + a new file in the domain's folder, and its number must sit in the + domain's range; under adr/v0 an in-tree record's must too. Under adr/v1 + an in-tree record keeps its number wherever its folder is. Raises + SheetError when none of this holds.""" + target = sheet['target'] + number = format_number(target['number']) + domain = target.get('domain') + in_tree = source_file is not None and _in_tree(source_file) + if in_tree: + was = filename_number(source_file) + was_domain = source_domain(source_file, was or number) + if was != number.lstrip('0') or was_domain != domain: + raise SheetError( + f"the source is ADR-{was} in {was_domain}; the sheet says ADR-{number} in {domain}. " + f"A record keeps its number; moving it to another domain is `adr domain move` " + f"(ADR-306 §6). Restore target.number and target.domain") + if not domain: + raise SheetError('target.domain is empty') + span = number_range(domain) + if span is None: + raise SheetError(f"domain '{domain}' is not in adr.yaml; add it first " + f"(creating a domain on apply, ADR-306 §6, is not built yet)") + if not (in_tree and repo_contract() == V1) and not span[0] <= int(number.split('.')[0]) <= span[1]: + raise SheetError(f"ADR-{number} is outside the {domain} range {span[0]}-{span[1]}") + if in_tree: + return source_file + for existing in find_adrs(include_archived=True): + if filename_number(existing) == number.lstrip('0'): + raise SheetError(f"ADR-{number} already exists: {relative_path(existing)}") + folders = get_domains()[domain]['folder'] if domain in get_domains() else 'legacy' + folder = folders[0] if isinstance(folders, list) else folders + slug = re.sub(r'[^a-z0-9]+', '-', str(target['title']).lower()).strip('-') + return get_project_root() / 'docs' / 'architecture' / folder / f"ADR-{number}-{slug}.md" + +def _imported(source: dict, sheet: dict, raw: bytes) -> dict: + """`imported` for a record from a non-v1 source (ADR-306 §4): where it + came from, the source's own status as written, and the source keys with + no v1 field. The record keeps them whatever happens to the todo items. + A source outside the repo is named by its file name, so no machine's + path lands in the record.""" + path = source['path'] + imported = {'from': Path(path).name if Path(path).is_absolute() else path, + 'format': source.get('format')} + front = split_record(raw.decode('utf-8'))[0] + data = yaml.safe_load(front or '') or {} + imported['status'] = data.get('status') if isinstance(data, dict) else None + if sheet.get('unmapped'): + imported['unmapped'] = dict(sheet['unmapped']) + return imported + +def _uncommitted(path: Path) -> bool: + """The file has changes git has not committed. False outside git.""" + root = get_project_root() + try: + rel = str(path.resolve().relative_to(root.resolve())) + except ValueError: + return False + return bool((_git(['status', '--porcelain', '--', rel], root) or '').strip()) + +def _apply_one(sheet: dict, force: bool, undo: Optional[dict] = None) -> tuple: + """Write one sheet as a v1 record through render_record and return + (path, whether it changed). A non-v1 source gains `imported: {from, + format}`, plus `unmapped` when the source had keys with no v1 field, so + the record keeps them whatever happens to the todo (ADR-306 §1, §4). + Refused: a source changed since the scan, a source with a body whose + sheet has none, and a source with uncommitted changes unless --force. + With `undo`, the destination's prior bytes (None when it did not exist) + are kept there so a dry run can put them back.""" + source = sheet.get('source') + source_file = None + record = dict(sheet['record']) + if source: + source_file = _source_file(source) + if not source_file.is_file(): + raise SheetError(f"source {source['path']} is gone") + raw = source_file.read_bytes() + if hashlib.sha256(raw).hexdigest() != source.get('sha256'): + raise SheetError(f"source {source['path']} changed since the scan; scan it again") + if source.get('format') != 'v1' and 'imported' not in record: + record['imported'] = _imported(source, sheet, raw) + if read_record(source_file)['body'].strip() and not (sheet.get('body') or '').strip(): + raise SheetError('the source has a body and the sheet has none; scan it again') + dest = _destination(sheet, source_file) + text = render_record(dict(sheet, record=record)) + if dest.is_file() and _same_record(dest.read_text(), text): + return dest, False + if dest.is_file() and not force and _uncommitted(dest): + raise SheetError(f"{relative_path(dest)} has uncommitted changes; commit them, " + f"or --force to overwrite") + if undo is not None and dest not in undo: + undo[dest] = dest.read_bytes() if dest.is_file() else None + made = undo.setdefault('__dirs__', []) + folder = dest.parent + while not folder.exists(): + made.append(folder) + folder = folder.parent + dest.parent.mkdir(parents=True, exist_ok=True) + dest.write_text(text) + return dest, True + +def _same_record(old: str, new: str) -> bool: + """The same frontmatter data and the same text after it. A v1 record + applied unedited is left as its author formatted it (ADR-306 §7).""" + old_front, old_before, old_title, old_after = split_record(old) + new_front, new_before, new_title, new_after = split_record(new) + if old_front is None or old_title is None or new_title is None: + return False + try: + same_data = yaml.safe_load(old_front) == yaml.safe_load(new_front) + except yaml.YAMLError: + return False + return (same_data and old_before.strip() == new_before.strip() + and old_title.group(0) == new_title.group(0) and old_after == new_after) + +def _labels(todo: list) -> str: + """Todo items by label, counted: 'verb, basis, unmapped (2)'.""" + counts = {} + for item in todo: + counts[todo_label(item)] = counts.get(todo_label(item), 0) + 1 + return ', '.join(f"{label} ({n})" if n > 1 else label for label, n in counts.items()) + +def _import_apply(args): + """Write each finished sheet as a v1 record, then lint what was written + (ADR-306 §3). A sheet with open todo items is skipped. --partial writes + it anyway, unless an item is one lint could not find again afterwards + (BLOCKING_TODO). An applied sheet is removed: the record is what is + kept. One bad sheet is reported and the rest still apply. --dry-run + writes the records, lints them inside the corpus, prints each issue, + then restores every file it wrote and keeps every sheet.""" + if args.sheets: + paths = [Path(p) for p in args.sheets] + else: + paths = sorted(import_dir().glob('ADR-*.yaml')) if import_dir().is_dir() else [] + lines, written = [], [] + undo = {} if args.dry_run else None + applied = skipped = refused = 0 + counts, issues = {}, {} + try: + for path in paths: + shown = relative_path(path.resolve()) + try: + sheet = load_sheet(path) + todo = sheet.get('todo') or [] + held = blocking_todo(todo) if args.partial else todo + if held: + why = "todo item(s) --partial does not write past" if args.partial else 'open todo item(s)' + lines.append((f"Skipped: {shown}: {len(held)} {why}: {_labels(held)}", None)) + skipped += 1 + continue + dest, changed = _apply_one(sheet, args.force, undo) + except (SheetError, OSError, ValueError, TypeError, AttributeError, KeyError) as e: + lines.append((f"Refused: {shown}: {e}", None)) + refused += 1 + continue + if not args.dry_run: + try: + path.unlink() + except OSError as e: + lines.append((f"Note: {shown}: applied, but the sheet could not be removed: {e}", None)) + applied += 1 + written.append(dest.resolve()) + partial = (f", {len(todo)} todo left" if todo else '') + ('' if changed else ', unchanged') + lines.append((f"Applied: {shown} -> {relative_path(dest)}{partial}", dest.resolve())) + counts, issues = _lint_counts(written) + finally: + if undo is not None: + for failed in _restore(undo): + lines.append((f"Note: dry run could not restore {failed}", None)) + for text, dest in lines: + if dest is not None: + errors, warnings = counts.get(dest, (0, 0)) + text += f" (lint: {errors} errors, {warnings} warnings)" + if args.dry_run: + text = text.replace('Applied: ', 'Would apply: ', 1) + text += ''.join(f"\n {'error' if i.severity == 'error' else 'warning'}: {i.message}" + for i in issues.get(dest, [])) + print(text) + verb = 'would apply' if args.dry_run else 'applied' + print(f"Import: {applied} {verb}, {skipped} skipped, {refused} refused" + + (' (dry run: nothing written)' if args.dry_run else '')) + lint_errors = sum(e for e, _ in counts.values()) + return 1 if refused or (args.dry_run and lint_errors) else 0 + +def _restore(undo: dict) -> list: + """Put back every file a dry run wrote, and remove the directories it + made that are left empty. Each file is restored on its own, so one + failure doesn't stop the rest; the paths that failed are returned.""" + failed = [] + for dest, before in undo.items(): + if dest == '__dirs__': + continue + try: + if before is None: + dest.unlink(missing_ok=True) + else: + dest.write_bytes(before) + except OSError: + failed.append(relative_path(dest)) + for folder in sorted(undo.get('__dirs__', []), key=lambda d: len(d.parts), reverse=True): + try: + folder.rmdir() + except OSError: + pass + return failed + +def _lint_counts(paths: list) -> dict: + """(errors, warnings) per written record, and the issues themselves, + from the full rule set.""" + if not paths: + return {}, {} + corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) + targets = [adr for adr in corpus if adr.path.resolve() in set(paths)] + run_rules(targets, ctx) + counts = {adr.path.resolve(): (sum(i.severity == 'error' for i in adr.issues), + sum(i.severity != 'error' for i in adr.issues)) + for adr in targets} + return counts, {adr.path.resolve(): list(adr.issues) for adr in targets} diff --git a/hooks/ways/documentation/adr/src/cmd_index.py b/hooks/ways/documentation/adr/src/cmd_index.py index 517c0ad7..5136411b 100644 --- a/hooks/ways/documentation/adr/src/cmd_index.py +++ b/hooks/ways/documentation/adr/src/cmd_index.py @@ -21,21 +21,36 @@ def status_cell(adr): status = adr.status or '?' return f"{status} ({note})" if note else status - lines = [ - "# Architecture Decision Records", - "", - f"This directory contains Architecture Decision Records (ADRs) for {get_config().get('project_name', 'this project')}.", - "Each ADR documents a significant architectural decision, its context, and consequences.", - "", - "## ADR Format", - "", - "All ADRs follow a consistent format:", - "- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded", - "- **Date:** When the decision was made", - "- **Deciders:** Who made the decision", - "- **Context:** The problem or situation requiring a decision", - "- **Decision:** The architectural choice made", - "- **Consequences:** Benefits, drawbacks, and other impacts", + project = get_config().get('project_name', 'this project') + if repo_contract() == 'adr/v1': + about = [ + f"This directory contains the Agent Decision Records (ADRs) for {project}, under the adr/v1 contract.", + "A record's kind says what it holds: a decision, a spec kept current, or evidence (findings, surveys, explorations).", + "", + "## Record Format", + "", + "- **Kind:** decision / spec / evidence, as `adr.yaml` declares", + "- **Status:** proposed / accepted / rejected / abandoned / superseded / archived", + "- **Capability:** what the record is about, from the vocabulary in `adr.yaml`", + "- **Decisions** also carry a verb (add / cut / change / retire / constrain), a basis, and open with a Summary", + "- **Numbers** are permanent; a record's folder is its area", + ] + else: + about = [ + f"This directory contains Agent Decision Records (ADRs) for {project}.", + "Each ADR documents a significant architectural decision, its context, and consequences.", + "", + "## ADR Format", + "", + "All ADRs follow a consistent format:", + "- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded", + "- **Date:** When the decision was made", + "- **Deciders:** Who made the decision", + "- **Context:** The problem or situation requiring a decision", + "- **Decision:** The architectural choice made", + "- **Consequences:** Benefits, drawbacks, and other impacts", + ] + lines = ["# Agent Decision Records", ""] + about + [ "", f"_This index is auto-generated by `adr index`. Configuration: [`adr.yaml`](./adr.yaml)_", "", diff --git a/hooks/ways/documentation/adr/src/cmd_lifecycle.py b/hooks/ways/documentation/adr/src/cmd_lifecycle.py index a68dd794..d955fdb6 100644 --- a/hooks/ways/documentation/adr/src/cmd_lifecycle.py +++ b/hooks/ways/documentation/adr/src/cmd_lifecycle.py @@ -40,33 +40,22 @@ def _set_status(raw: bytes, status: str) -> bytes: lines[i] = b'status: ' + status.encode() + (comment.group(0) if comment else b'') + ending return b'\n'.join(lines) -def _errors_by_record(corpus: list, ctx) -> dict: - run_rules(corpus, ctx) - found = {a.path: {i.message for i in a.issues if i.severity == 'error'} for a in corpus} - found['adr.yaml'] = {i.message for i in ctx.config_issues if i.severity == 'error'} - return found - def _trial(adr, status: str): - """Run every rule over the whole corpus as it is and again with the - record's status changed. Returns (the trial record, new errors as - (where, message)). A change that breaks another record, such as a - rejected precedent, shows up here, not just the record's own issues.""" - baseline = get_all_adrs(include_archived=True) - before = _errors_by_record(baseline, LintContext.from_corpus(baseline)) + """The record with its status changed, and its own errors as (where, + message): the file rules only, over this one record. The rest of the + corpus is not re-linted (ADR-311 §3).""" corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) trial = next(a for a in corpus if a.path == adr.path) trial.status = status trial.frontmatter = dict(trial.frontmatter, status=status) # The record was parsed clean enough to reach here (_lifecycle_target # refused anything else), so starting its issues fresh loses nothing. trial.issues = [] - after = _errors_by_record(corpus, LintContext.from_corpus(corpus)) - new = [] - for path, messages in after.items(): - for message in sorted(messages - before.get(path, set())): - where = 'adr.yaml' if path == 'adr.yaml' else relative_path(path) - new.append((where, message)) - return trial, new + for rule in _active(FILE_RULES, ctx): + rule(trial, ctx) + where = relative_path(trial.path) + return trial, [(where, i.message) for i in trial.issues if i.severity == 'error'] def _write_status(adr, status: str, suffix: str = '') -> Optional[str]: """Write the new status (and an appended suffix), then re-read the file to @@ -86,20 +75,18 @@ def _write_status(adr, status: str, suffix: str = '') -> Optional[str]: return None def _refuse(adr, verb: str, new: list) -> int: - print(f"Refused: {verb} ADR-{adr.number} would add errors:", file=sys.stderr) + print(f"Refused: {verb} ADR-{adr.number} would leave it with errors:", file=sys.stderr) for where, message in new: print(f" ❌ {where}: {message}", file=sys.stderr) return 1 def cmd_accept(args): - """Accept a proposed adr/v1 record (ADR-304 §2, §11, §12). + """Accept a proposed adr/v1 record (ADR-304 §2, ADR-311 §3). - Runs every rule over the corpus as if the record were accepted and refuses - if that adds an error anywhere, so a decision whose basis does not meet the - rules, or one the operator started without a considered entry, stays - proposed. A precedent that is still proposed also refuses: §11 grounds a - decision in another accepted decision. The record's warnings are printed, - with open concerns first: shown at acceptance, not blocking. + Runs the record's own file rules as if it were accepted and refuses if + any reports an error. The rest of the corpus is not re-linted. The + record's warnings are printed, with open concerns first: shown at + acceptance, not blocking. """ adr, code = _lifecycle_target(args) if adr is None: @@ -107,9 +94,6 @@ def cmd_accept(args): trial, new = _trial(adr, 'accepted') if new: return _refuse(adr, 'accepting', new) - pending = [i for i in trial.issues if i.code == 'precedent-proposed'] - if pending: - return _refuse(adr, 'accepting', [(relative_path(adr.path), i.message) for i in pending]) concerns = [i for i in trial.issues if i.code == 'open-concern'] others = [i for i in trial.issues if i.severity == 'warning' and i.code != 'open-concern'] if concerns: diff --git a/hooks/ways/documentation/adr/src/cmd_lint.py b/hooks/ways/documentation/adr/src/cmd_lint.py index 3303507f..f42449ad 100644 --- a/hooks/ways/documentation/adr/src/cmd_lint.py +++ b/hooks/ways/documentation/adr/src/cmd_lint.py @@ -6,7 +6,7 @@ def cmd_lint(args): """ if args.paths: # Resolved, so records parsed from the arguments match the corpus's - # own paths (grounding and other corpus rules key on the path). + # own paths (corpus rules key on the path). paths = [Path(p).resolve() for p in args.paths] adrs = [parse_adr(p) for p in paths] else: diff --git a/hooks/ways/documentation/adr/src/cmd_list.py b/hooks/ways/documentation/adr/src/cmd_list.py index 769ad700..43bfbb81 100644 --- a/hooks/ways/documentation/adr/src/cmd_list.py +++ b/hooks/ways/documentation/adr/src/cmd_list.py @@ -29,8 +29,21 @@ def sort_key(adr): adrs.sort(key=sort_key) + # Frontmatter filters: --field, and --kind, --verb, --capability as its + # shorthands. A list field matches when it lists the value. + filters = list(getattr(args, 'field', None) or []) + for name in ('kind', 'verb', 'capability'): + if getattr(args, name, None): + filters.append(f"{name}={getattr(args, name)}") + for spec in filters: + key, has_value, value = spec.partition('=') + adrs = [a for a in adrs if _field_matches(a, key.strip(), value if has_value else None)] + + if getattr(args, 'json', False): + return _list_json(adrs, getattr(args, 'group_by', None)) + project = get_config().get('project_name', 'Project') - print(f"\n{project} — Architecture Decision Records ({len(adrs)} total)") + print(f"\n{project} — Agent Decision Records ({len(adrs)} total)") print("=" * 55) def status_icon(status): @@ -48,7 +61,13 @@ def print_adr(adr): marker = ' [archived]' if is_archived(adr.path) and getattr(args, 'all', False) else '' print(f" {status_icon(adr.status)} ADR-{adr.number or '???':8} {adr.title or '(no title)'}{suffix}{marker}") - if args.group: + if getattr(args, 'group_by', None): + for value, members in _groups(adrs, args.group_by): + print(f"\n## {args.group_by}: {value} ({len(members)})") + print("-" * 50) + for adr in members: + print_adr(adr) + elif args.group: # Group by domain - show domains in config order, then legacy for domain_key, domain_info in domains.items(): domain_adrs = [a for a in adrs if a.domain == domain_key] @@ -76,3 +95,60 @@ def print_adr(adr): return 0 + +# --- frontmatter queries for list ---------------------------------------------------- + +_NO_VALUE = '(none)' + +def _field_values(adr, key: str) -> list: + """A field's values as strings: each item of a list, a scalar alone, and + nothing for an absent or empty field, read from the frontmatter. + Values compare case-sensitively.""" + value = adr.frontmatter.get(key) + if value in (None, '', [], {}): + return [] + if isinstance(value, list): + return [json.dumps(v, default=str, ensure_ascii=False) if isinstance(v, (dict, list)) else str(v) + for v in value] + if isinstance(value, dict): + return [json.dumps(value, default=str, ensure_ascii=False)] + return [str(value)] + +def _field_matches(adr, key: str, value: Optional[str]) -> bool: + values = _field_values(adr, key) + if value is None: + return bool(values) + if key == 'status': + return value.lower() in (v.lower() for v in values) + if key in V1_EDGE_FIELDS + ('superseded_by', 'related'): + # A reference without a section matches every section of that record. + number, section = norm_ref(value) + return any(norm_ref(v)[0] == number and (section is None or norm_ref(v)[1] == section) + for v in values) + return value in values + +def _groups(adrs: list, key: str) -> list: + """(value, records) sorted by value, records without the field last. + A record listing several values is in each.""" + groups = {} + for adr in adrs: + for value in _field_values(adr, key) or [_NO_VALUE]: + groups.setdefault(value, []) + if adr not in groups[value]: + groups[value].append(adr) + ordered = sorted(((v, m) for v, m in groups.items() if v != _NO_VALUE), key=lambda g: g[0].lower()) + if _NO_VALUE in groups: + ordered.append((_NO_VALUE, groups[_NO_VALUE])) + return ordered + +def _record_json(adr) -> dict: + return {'number': adr.number, 'title': adr.title, 'path': str(relative_path(adr.path)), + 'status': adr.status, 'frontmatter': adr.frontmatter} + +def _list_json(adrs: list, group_by: Optional[str]) -> int: + if group_by: + data = {value: [_record_json(a) for a in members] for value, members in _groups(adrs, group_by)} + else: + data = [_record_json(a) for a in adrs] + print(json.dumps(data, indent=2, default=str, ensure_ascii=False)) + return 0 diff --git a/hooks/ways/documentation/adr/src/cmd_new.py b/hooks/ways/documentation/adr/src/cmd_new.py index 4b887a22..1e8681fd 100644 --- a/hooks/ways/documentation/adr/src/cmd_new.py +++ b/hooks/ways/documentation/adr/src/cmd_new.py @@ -45,8 +45,20 @@ def cmd_new(args): print(f"Error: File already exists: {filepath}", file=sys.stderr) return 1 - # Generate content today = date.today().isoformat() + if get_config().get('contract') == V1: + content = _v1_content(args, next_num, domain, today, defaults) + if content is None: + return 1 + folder.mkdir(parents=True, exist_ok=True) + filepath.write_text(content) + print(f"Created: {relative_path(filepath)}") + print(f" Domain: {config['name']} ({domain})") + print(f" Number: ADR-{next_num:03d}") + print(" Contract: adr/v1 (fill the empty fields; `adr lint` lists them)") + return 0 + + # Generate content default_status = defaults.get('status', 'Draft') default_deciders = defaults.get('deciders', []) @@ -68,33 +80,7 @@ def cmd_new(args): # ADR-{next_num:03d}: {args.title} -## Context - -[What is the issue that we're seeing that is motivating this decision or change?] - -## Decision - -[What is the change that we're proposing and/or doing?] - -## Consequences - -### Positive - -- [What becomes easier?] - -### Negative - -- [What becomes harder?] - -### Neutral - -- [What other changes does this enable or require?] - -## Alternatives Considered - -- [What other options were evaluated?] -- [Why were they rejected?] -''' +''' + BODY_SKELETON # Write file folder.mkdir(parents=True, exist_ok=True) @@ -105,3 +91,48 @@ def cmd_new(args): print(f" Number: ADR-{next_num:03d}") return 0 + +def _v1_content(args, number: int, domain: str, today: str, defaults: dict) -> Optional[str]: + """A v1 record from an empty sheet (ADR-306 §3). The kind's schema decides + which fields the record carries; fields the arguments do not give stay + empty, and lint names each one. Arguments the kind cannot take are refused.""" + config = get_config() + kinds = {k: v for k, v in _mapping(config.get('kinds')).items() if isinstance(v, dict)} + if not kinds: + print("Error: adr.yaml declares contract adr/v1 but no kinds; `adr lint` reports the config", file=sys.stderr) + return None + wanted = (args.kind or 'decision').lower() + kind = next((k for k in kinds if str(k).lower() == wanted), None) + if kind is None: + print(f"Error: Unknown kind '{args.kind or 'decision'}'. Kinds: {', '.join(map(str, kinds))}", file=sys.stderr) + return None + schema = kinds[kind] + problems = [] + takes_verb = schema.get('verb') == 'required' + if args.verb and not takes_verb: + problems.append(f"a {kind} record takes no verb") + verbs = _str_list(config.get('verbs')) or list(V1_VERBS) + if args.verb and takes_verb and args.verb not in verbs: + problems.append(f"verb '{args.verb}' is not one of: {', '.join(verbs)}") + vocabulary = _mapping(config.get('capabilities')) + if args.capability and args.capability not in vocabulary: + problems.append(f"capability '{args.capability}' is not in the adr.yaml vocabulary") + if (args.agent or args.model) and 'agent' not in v1_requires(schema): + problems.append(f"a {kind} record carries no agent") + for problem in problems: + print(f"Error: {problem}", file=sys.stderr) + if problems: + return None + deciders = list(defaults.get('deciders') or []) + if not deciders: + git_user = _detect_git_user() + if git_user: + deciders = [git_user] + given = {'verb': args.verb, 'capability': args.capability, 'agent': args.agent, 'model': args.model} + record = empty_record(kind, schema, given, defaults) + record.update({'date': today, 'deciders': deciders, 'related': []}) + sections = _str_list(schema.get('sections')) or [] + summary = V1_SUMMARY_SKELETON if 'Summary' in sections else None + sheet = new_sheet(number, domain, args.title, record, summary=summary, + body=BODY_SKELETON if takes_verb else '') + return render_record(sheet) diff --git a/hooks/ways/documentation/adr/src/cmd_record.py b/hooks/ways/documentation/adr/src/cmd_record.py new file mode 100644 index 00000000..997433ad --- /dev/null +++ b/hooks/ways/documentation/adr/src/cmd_record.py @@ -0,0 +1,330 @@ + +# --- record edits: consider, set, supersede, enact ----------------------------------- +# +# Each command edits frontmatter through FrontmatterEdit, so only the lines of +# the fields it touches change. It refuses what the contract refuses before +# writing, then lints the records it wrote and prints their issues. + +def _record_target(ref: str, command: str, v1_only: bool = True): + """The record ref names, or (None, exit code) after an error.""" + matches = find_by_ref(ref, get_all_adrs(include_archived=True)) + if not matches: + print(f"Error: ADR not found: {ref}", file=sys.stderr) + return None, 1 + if len(matches) > 1: + print(f"Error: '{ref}' matches several records; resolve the duplicate number first:", file=sys.stderr) + for match in matches: + print(f" {relative_path(match.path)}", file=sys.stderr) + return None, 1 + adr = matches[0] + if v1_only and (repo_contract() != V1 or adr.contract != V1): + print(f"Error: ADR-{adr.number} is not an adr/v1 record; `adr {command}` edits v1 fields " + f"(ADR-304 §7).", file=sys.stderr) + return None, 1 + return adr, 0 + +def _open_edit(adr): + try: + return FrontmatterEdit(adr.path.read_bytes()), None + except (ValueError, UnicodeDecodeError, yaml.YAMLError) as e: + return None, f"ADR-{adr.number}: {e}" + +def _record_schema(adr) -> Optional[dict]: + kinds = get_config().get('kinds') + schema = kinds.get(adr.frontmatter.get('kind')) if isinstance(kinds, dict) else None + return schema if isinstance(schema, dict) else None + +def _show_diff(adr, before: bytes, after: bytes) -> None: + import difflib + name = str(relative_path(adr.path)) + diff = difflib.unified_diff(before.decode('utf-8').splitlines(True), after.decode('utf-8').splitlines(True), + fromfile=f"a/{name}", tofile=f"b/{name}") + sys.stdout.writelines(line if line.endswith('\n') else line + '\n' for line in diff) + +def _write(edits: list, dry_run: bool) -> None: + """Write each (adr, edit) whose bytes changed, or show the diff.""" + for adr, edit in edits: + before, after = adr.path.read_bytes(), edit.bytes() + if before == after: + continue + if dry_run: + _show_diff(adr, before, after) + else: + adr.path.write_bytes(after) + +def _lint_records(paths: list) -> None: + """Lint the records just written against the whole corpus and print + their issues, as `adr lint <path>` would list them.""" + corpus = get_all_adrs(include_archived=True) + ctx = LintContext.from_corpus(corpus) + records = [a for a in corpus if a.path in paths] + run_rules(records, ctx) + for record in records: + where = relative_path(record.path) + if not record.issues: + print(f"Lint {where}: clean") + continue + print(f"Lint {where}:") + for issue in record.issues: + print(f" {'❌' if issue.severity == 'error' else '⚠️'} {issue.message}") + +def _finish(edits: list, dry_run: bool, done: str) -> int: + _write(edits, dry_run) + if dry_run: + print(f"Dry run, nothing written: {done}") + return 0 + print(done) + _lint_records([adr.path for adr, _ in edits]) + return 0 + +# --- consider ----------------------------------------------------------------------- + +_PROBE_NAME = re.compile(r'\*\s*(?:not\s+)?confident\s*\(([^)\n]+)\)\s*:?\s*\*', re.IGNORECASE) + +def probe_names(adr) -> list: + """The named probes in a record's Summary, written *Confident (name):* or + *Not confident (name):*, then 'inversion' when the Summary has one: a + considered entry may cover the inversion too (ADR-304 §12).""" + heading = _find_section(adr, 'Summary') + text = adr.section_text.get(heading, '') if heading else '' + names = list(dict.fromkeys(m.strip() for m in _PROBE_NAME.findall(text))) + if re.search(r'\binversion\b', text, re.IGNORECASE): + names.append('inversion') + return names + +def cmd_consider(args): + """Append one considered entry: the operator's answer to the Summary + (ADR-304 §12), on a record in any status.""" + adr, code = _record_target(args.adr, 'consider') + if adr is None: + return code + for flag, value, what in (('--said', args.said, 'what the operator said, verbatim'), + ('--via', args.via, 'the channel it was said in, such as a PR or a session')): + if value is None or not value.strip(): + print(f"Error: {flag} is required: {what}.", file=sys.stderr) + return 1 + operator = args.operator or _detect_git_user() + if not operator: + print("Error: no operator given and none detected from gh or git; pass --operator NAME.", file=sys.stderr) + return 1 + if args.covers is not None: + valid = probe_names(adr) + unknown = [c for c in args.covers if c not in valid] + if unknown: + names = ', '.join(repr(u) for u in unknown) + if any(v != 'inversion' for v in valid): + print(f"Error: ADR-{adr.number} has no probe {names}. Names it can cover: {', '.join(valid)}.", + file=sys.stderr) + else: + also = f" It can cover: {', '.join(valid)}." if valid else '' + print(f"Error: ADR-{adr.number} has no probe {names}; its Summary names no probes. " + f"A probe is named as *Confident (name):* or *Not confident (name):*.{also}", + file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + entry = {'operator': operator, 'said': _Quoted(args.said), 'via': _prefer_double(args.via)} + if args.paraphrase: + entry['paraphrase'] = True + if args.covers is not None: + entry['covers'] = _FlowList(args.covers) + if args.canary: + entry['canary'] = args.canary + try: + edit.append('considered', [entry]) + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + count = len(edit.fields.get('considered') or []) + return _finish([(adr, edit)], args.dry_run, f"Added considered entry {count} to ADR-{adr.number}: {adr.title}") + +# --- set ---------------------------------------------------------------------------- + +_ASSIGNMENT = re.compile(r'^([A-Za-z_][A-Za-z0-9_]*(?:-[A-Za-z0-9_]+)*)(\+=|-=|=)(.*)$', re.DOTALL) +# Statuses with a command of their own, which checks the record and records why. +_STATUS_COMMANDS = {'accepted': 'accept', 'rejected': 'reject', 'abandoned': 'abandon', + 'superseded': 'supersede', 'archived': 'archive'} + +def _parse_assignment(text: str): + match = _ASSIGNMENT.match(text) + if not match: + raise ValueError(f"'{text}' is not key=value, key+=value or key-=value") + key, op, raw = match.groups() + try: + value = yaml.safe_load(raw) if raw.strip() else None + except yaml.YAMLError as e: + raise ValueError(f"{key}: the value is not YAML ({str(e).splitlines()[0]})") + return key, op, value + +def cmd_set(args): + """Set, append to or remove from frontmatter fields. status is refused: + the lifecycle commands own it. Any other field may be set on a record in + any status; git keeps what it was (ADR-311).""" + adr, code = _record_target(args.adr, 'set', v1_only=False) + if adr is None: + return code + try: + changes = [_parse_assignment(a) for a in args.assignments] + except ValueError as e: + print(f"Error: {e}", file=sys.stderr) + return 1 + if adr.contract == V1 and not args.force: + for key, op, value in changes: + if key != 'status': + continue + command = _STATUS_COMMANDS.get(str(value).lower()) if op == '=' else None + if command: + print(f"Refused: use `adr {command} {adr.number}`; it checks the record and records why. " + f"--force sets the status anyway.", file=sys.stderr) + else: + print(f"Refused: a record's status changes only through accept, reject, abandon, " + f"supersede and archive. --force sets the status anyway.", file=sys.stderr) + return 1 + # A key the record lacks and no v1 record carries is most likely a typo. + unknown = [k for k, _, _ in changes + if adr.contract == V1 and k not in adr.frontmatter and k not in V1_KEY_ORDER] + if unknown and not args.force: + print(f"Refused: ADR-{adr.number} has no field {', '.join(repr(k) for k in unknown)}, and it is not " + f"a v1 field ({', '.join(V1_KEY_ORDER)}). --force adds it anyway.", file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + for key, op, value in changes: + items = value if isinstance(value, list) else [value] + if op == '=': + edit.set(key, value) + elif op == '+=': + edit.append(key, items) + else: + missing = edit.remove(key, items) + if missing: + print(f"Error: ADR-{adr.number} not changed: {key} does not list " + f"{', '.join(repr(str(m)) for m in missing)}.", file=sys.stderr) + return 1 + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + keys = ', '.join(dict.fromkeys(k for k, _, _ in changes)) + return _finish([(adr, edit)], args.dry_run, f"Set {keys} on ADR-{adr.number}: {adr.title}") + +# --- supersede ---------------------------------------------------------------------- + +def _ref(adr, section: Optional[str] = None) -> str: + return f"ADR-{adr.number}" + (f"#{section}" if section else '') + +def _lists(adr, key: str, ref: str) -> bool: + number, section = norm_ref(ref) + return any(norm_ref(e) == (number, section) for e in as_entries(adr.frontmatter.get(key))) + +def cmd_supersede(args): + """Write both sides of a supersession, or with --amends a partial + replacement on the new record only (ADR-304 §3).""" + old, code = _record_target(args.adr, 'supersede') + if old is None: + return code + new, code = _record_target(args.by, 'supersede') + if new is None: + return code + if old.path == new.path: + print(f"Error: ADR-{old.number} cannot supersede itself.", file=sys.stderr) + return 1 + new_status = str(new.status or '').lower() + if new_status not in ('accepted', 'proposed'): + print(f"Error: ADR-{new.number} is {new_status or 'without a status'}; only an accepted or proposed " + f"record supersedes another.", file=sys.stderr) + return 1 + old_status = str(old.status or '').lower() + if old_status not in ('accepted', 'superseded'): + print(f"Error: ADR-{old.number} is {old_status or 'without a status'}; only an accepted record is " + f"superseded (a proposed one is rejected or abandoned).", file=sys.stderr) + return 1 + field_name = 'amends' if args.amends else 'supersedes' + new_kind, old_kind = new.frontmatter.get('kind'), old.frontmatter.get('kind') + edges = v1_edges(_record_schema(new) or {}) + if old_kind not in edges.get(field_name, []): + allowed = edges.get(field_name) + target = f"points at {' or '.join(allowed)} records" if allowed else "is not an edge it takes" + print(f"Error: a {new_kind} record cannot {'amend' if args.amends else 'supersede'} a {old_kind}: " + f"for a {new_kind}, {field_name} {target} (adr.yaml kinds.{new_kind}.edges).", file=sys.stderr) + return 1 + section = args.amends + if section and not section_exists(old, section): + print(f"Error: ADR-{old.number} has no section '{section}'. Its sections: {', '.join(old.sections)}.", + file=sys.stderr) + return 1 + if not section and 'superseded' not in v1_lifecycle(_record_schema(old)): + print(f"Error: the {old_kind} lifecycle has no superseded status.", file=sys.stderr) + return 1 + ref = _ref(old, section) + edits = [] + if not _lists(new, field_name, ref): + edit, error = _open_edit(new) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + edit.append(field_name, [ref]) + except ValueError as e: + print(f"Error: ADR-{new.number} not changed: {e}", file=sys.stderr) + return 1 + edits.append((new, edit)) + if not section: + edit, error = _open_edit(old) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + try: + if not _lists(old, 'superseded_by', _ref(new)): + edit.append('superseded_by', [_ref(new)]) + if old_status != 'superseded': + edit.set('status', 'superseded') + except ValueError as e: + print(f"Error: ADR-{old.number} not changed: {e}", file=sys.stderr) + return 1 + edits.append((old, edit)) + if section: + done = f"ADR-{new.number} amends ADR-{old.number} §{section}" + else: + done = f"ADR-{new.number} supersedes ADR-{old.number}; ADR-{old.number} is superseded" + code = _finish(edits, args.dry_run, done) + if not args.dry_run and not section: + print("Run `adr index -y` to refresh INDEX.md.") + return code + +# --- enact -------------------------------------------------------------------------- + +def cmd_enact(args): + """Mark an accepted cut or retire decision done at a commit (ADR-304 §5).""" + adr, code = _record_target(args.adr, 'enact') + if adr is None: + return code + verb = adr.frontmatter.get('verb') + if verb not in V1_ENACTING_VERBS: + print(f"Error: ADR-{adr.number} is a{'n' if str(verb)[:1] in 'aeiou' else ''} {verb or 'verbless'} " + f"record; enacted belongs to a cut or retire decision (ADR-304 §5).", file=sys.stderr) + return 1 + status = str(adr.status or '').lower() + if status != 'accepted': + print(f"Error: ADR-{adr.number} is {status}; enacted marks an accepted decision done.", file=sys.stderr) + return 1 + commit = args.commit.strip().lower() + if not re.fullmatch(r'[0-9a-f]{7,40}', commit): + print(f"Error: '{args.commit}' is not a commit hash (7 to 40 hex digits).", file=sys.stderr) + return 1 + edit, error = _open_edit(adr) + if error: + print(f"Error: {error}", file=sys.stderr) + return 1 + previous = adr.frontmatter.get('enacted') + try: + edit.set('enacted', _Quoted(commit)) + except ValueError as e: + print(f"Error: ADR-{adr.number} not changed: {e}", file=sys.stderr) + return 1 + note = f" (was {previous})" if previous and str(previous) != commit else '' + return _finish([(adr, edit)], args.dry_run, f"Enacted ADR-{adr.number} at {commit}{note}: {adr.title}") diff --git a/hooks/ways/documentation/adr/src/config.py b/hooks/ways/documentation/adr/src/config.py index ca358a29..f2fc69f3 100644 --- a/hooks/ways/documentation/adr/src/config.py +++ b/hooks/ways/documentation/adr/src/config.py @@ -28,6 +28,16 @@ def get_project_root() -> Path: return Path.cwd() +def _git(args: list, cwd: Path) -> Optional[str]: + """git's output, or None when git is missing, fails or times out.""" + try: + result = subprocess.run(['git', '-c', 'core.quotePath=false', *args], cwd=cwd, + capture_output=True, encoding='utf-8', errors='replace', timeout=10) + except (FileNotFoundError, subprocess.TimeoutExpired, OSError): + return None + return result.stdout if result.returncode == 0 else None + + def get_config_path() -> Path: """Get path to adr.yaml config file.""" return get_project_root() / 'docs' / 'architecture' / 'adr.yaml' @@ -74,6 +84,13 @@ def get_config() -> dict: _config = load_config() return _config +def reload_config() -> dict: + """Drop the cached config and read adr.yaml again, after a command + edits it.""" + global _config + _config = None + return get_config() + def repo_contract() -> str: """The contract adr.yaml declares; adr/v0 when it declares none (ADR-304).""" return str(get_config().get('contract') or 'adr/v0') diff --git a/hooks/ways/documentation/adr/src/header.py b/hooks/ways/documentation/adr/src/header.py index 8064940c..700b0e94 100644 --- a/hooks/ways/documentation/adr/src/header.py +++ b/hooks/ways/documentation/adr/src/header.py @@ -1,11 +1,13 @@ #!/usr/bin/env python3 """ -ADR - Architecture Decision Record CLI Tool +ADR - Agent Decision Record CLI Tool -A librarian for managing Architecture Decision Records. +A librarian for managing Agent Decision Records. Usage: adr list [--domain DOMAIN] [--status STATUS] [--group] [--archived|--all] + [--field KEY[=VALUE]] [--kind K] [--verb V] [--capability C] + [--group-by KEY] [--json] adr view <number> # View an ADR (aliases: v, show) adr new <domain> <title> adr rename <number> [new-title] [--slug SLUG] @@ -14,8 +16,19 @@ adr cite [--check] [paths...] adr accept <number> [--dry-run] adr reject|abandon <number> --reason "..." [--dry-run] + adr consider <number> --said "..." --via "..." [--operator NAME] [--covers PROBE...] + [--paraphrase] [--canary caught|missed] [--dry-run] + adr set <number> key=value|key+=value|key-=value ... [--force] [--dry-run] + adr supersede <old> --by <new> [--amends SECTION] [--dry-run] + adr enact <number> <commit> [--dry-run] + adr import scan <paths...> [--force] + adr import apply [sheets...] [--partial] [--force] [--dry-run] adr index [-y] adr domains + adr domain add <name> --range A-B --folder F [--label L] [--description D] + adr domain rename <old> <new> [--folder F] [--dry-run] + adr domain move <number> <domain> [--dry-run] + adr domain move --plan <file.yaml> [--dry-run] adr config Configuration is loaded from docs/architecture/adr.yaml @@ -26,7 +39,10 @@ # customize (ADR-177). import argparse +import hashlib +import json import os +import posixpath import re import subprocess import sys @@ -45,7 +61,7 @@ # Vendored-tool version (ADR-177). Bump when this tool changes — way macros # compare it against the installed template to tell stale from customized. -TOOL_VERSION = "2.0.0" +TOOL_VERSION = "2.1.0" # Statuses that mean "no longer in force" — used by archive and the # partial-supersession convention (ADR-303 / issue #438 option C2). diff --git a/hooks/ways/documentation/adr/src/import_read.py b/hooks/ways/documentation/adr/src/import_read.py new file mode 100644 index 00000000..b56db8a1 --- /dev/null +++ b/hooks/ways/documentation/adr/src/import_read.py @@ -0,0 +1,424 @@ +# --- adr import: readers and the sheet file (ADR-306) --------------------------------- +# +# A reader turns one source record into one import sheet (ADR-306 §1, §2). +# It fills what the source states and lists the rest in `todo`. It never +# writes an operator basis, `considered` or `concern`: `deciders` names who +# signed the record, not what they said (ADR-304 §7, §11). + +IMPORT_DIR = ('docs', 'architecture', '.import') + +SHEET_HEADER = ('# adr import sheet (ADR-306). Fill or clear each todo item, then run\n' + '# `adr import apply`. Sheets are working files; only records are committed.\n') + +# ADR-304 §7: the v0 status table. Deprecated depends on the record. +V0_STATUS_MAP = {'draft': 'proposed', 'proposed': 'proposed', 'accepted': 'accepted', + 'superseded': 'superseded', 'rejected': 'rejected'} + +# v0 keys that carry over under the same name and meaning. +V0_CARRIED = ('date', 'deciders', 'related', 'supersedes', 'superseded_by', 'amends') + +# Todo items that lint cannot detect once the record is written. They block +# --partial; every other item is advisory (ADR-306 §3). +BLOCKING_TODO = ('status', 'status note', 'target.number', 'target.domain') + +# The v1 default decision (ADR-304 §1), for a project whose adr.yaml declares no kinds. +IMPORT_DECISION_SCHEMA = {'verb': 'required', 'requires': ['capability', 'basis', 'agent'], + 'sections': ['Summary']} + +_STOPWORDS = frozenset(''' + the and for with that this from into over under onto are was were been being has have + had not but its it's our their them they these those which what when where who whom + why how all any each every one two per via than then there here also only more most + such can could should would will may might must shall does did doing done use used + using uses make makes made new adr record records decision decisions context +'''.split()) + +class SheetError(Exception): + """A source a reader cannot turn into a sheet, or a sheet apply cannot use.""" + +def import_dir() -> Path: + return get_project_root().joinpath(*IMPORT_DIR) + +def sheet_filename(number) -> str: + return f"ADR-{format_number(number)}.yaml" + +def _number_value(text: str) -> str: + """A record number as the sheet holds it: a string, '42' or '101.1', so + YAML never reads a zero-padded number as octal.""" + base, _, part = text.partition('.') + return (base.lstrip('0') or '0') + (f".{part}" if part else '') + +# --- splitting a record into frontmatter, title, Summary and body ------------------- + +def split_record(text: str) -> tuple: + """(frontmatter text or None, text before the H1, H1 match or None, text + after the H1 line). Joined back with the H1 line, the parts are the file.""" + lines = text.split('\n') + front, start = None, 0 + if lines and lines[0].strip() == '---': + for i in range(1, len(lines)): + if lines[i].strip() == '---': + front, start = '\n'.join(lines[1:i]), i + 1 + break + for j in range(start, len(lines)): + match = TITLE_PATTERN.match(lines[j]) + if match: + return front, '\n'.join(lines[start:j]), match, '\n'.join(lines[j + 1:]) + return front, '\n'.join(lines[start:]), None, '' + +def _render_tail(summary: Optional[str], body: str) -> str: + """What render_record writes after the H1 line for this summary and body.""" + tail = '' + if summary: + summary = summary if summary.startswith('## Summary') else f"## Summary\n\n{summary}" + tail += '\n' + summary.rstrip('\n') + '\n' + if body: + tail += '\n' + body + return tail + +def split_summary(after: str) -> tuple: + """(summary, body, note) for the text after the H1. A `## Summary` that + opens the body becomes the summary; the rest is the body, verbatim. The + split is kept only when render_record writes the same text back, so an + unedited sheet keeps the body byte-identical (ADR-306 §7). Otherwise the + Summary stays in the body and note says why.""" + body = after[1:] if after.startswith('\n') else after + lines = body.split('\n') + if lines[0].strip() != '## Summary': + note = 'the Summary is not the first section; kept in the body' if re.search( + r'(?m)^## Summary\s*$', body) else None + return None, body, note + fence, end = None, len(lines) + for k in range(1, len(lines)): + marker = re.match(r'\s*(```|~~~)', lines[k]) + if marker: + fence = None if fence == marker.group(1) else (fence or marker.group(1)) + elif fence is None and lines[k].startswith('## '): + end = k + break + section = '\n'.join(lines[:end]) + rest = '\n'.join(lines[end:]) + summary = section[len('## Summary\n\n'):] if section.startswith('## Summary\n\n') else section + if summary.strip() and _render_tail(summary, rest) == after: + return summary, rest, None + return None, body, 'the Summary does not split cleanly from the body; kept in the body' + +# --- what a reader decides ------------------------------------------------------------ + +def _empty(value) -> bool: + return value in (None, '', [], {}) + +def open_fields(record: dict, schema: dict) -> list: + """The fields the kind requires that the record leaves empty, in the + order a record is written: the reader's part of `todo` (ADR-306 §1).""" + wanted = (['verb'] if schema.get('verb') == 'required' else []) + v1_requires(schema) + found = [] + for name in wanted: + value = record.get(name) + if name == 'agent' and isinstance(value, dict): + found += [f"agent.{k}" for k in ('name', 'model') if _empty(value.get(k))] + elif _empty(value): + found.append(name) + return found + +_STEM_SUFFIXES = ('ations', 'ation', 'ments', 'ment', 'ions', 'ion', 'ings', 'ing', + 'ives', 'ive', 'als', 'al', 'ors', 'or', 'ers', 'er', 'ted', 'ed', + 'ts', 'es', 't', 'e', 's') + +def _stem(word: str) -> str: + """A crude English stem: strip suffixes repeatedly and collapse a doubled + final consonant, so ingest and ingestion rank as one word.""" + word = word.lower() + changed = True + while changed: + changed = False + for suffix in _STEM_SUFFIXES: + if word.endswith(suffix) and len(word) - len(suffix) >= 3: + word = word[:-len(suffix)] + changed = True + break + if len(word) >= 4 and word[-1] == word[-2] and word[-1] not in 'aeiou': + word = word[:-1] + return word + +def _words(text: str) -> set: + return {_stem(w) for w in re.findall(r"[a-z][a-z0-9']+", text.lower()) + if len(w) >= 3 and w not in _STOPWORDS} + +def capability_candidates(text: str, vocabulary: dict, top: int = 3) -> list: + """Vocabulary names ranked by how many words the record's title and + Context share with each capability's name and description (ADR-306 §1). + A word counts less the more descriptions carry it, so a word every + capability mentions does not decide the ranking.""" + words = _words(text) + described = [(str(name), _words(f"{name} {description or ''}")) + for name, description in vocabulary.items()] + spread = {} + for _, found in described: + for word in found: + spread[word] = spread.get(word, 0) + 1 + scored = [] + for i, (name, found) in enumerate(described): + overlap = round(sum(1 / spread[w] for w in words & found), 6) + if overlap: + scored.append((-overlap, i, name)) + return [name for _, _, name in sorted(scored)[:top]] + +def _folder_domain(path: Path) -> Optional[str]: + """The domain whose folder holds the file, or 'legacy' for the legacy folder.""" + name = path.parent.name + for domain, config in get_domains().items(): + folders = config.get('folder') + if name in (folders if isinstance(folders, list) else [folders]): + return domain + if name == 'legacy' and 'legacy' in get_config(): + return 'legacy' + return None + +def number_range(domain: str) -> Optional[tuple]: + """(low, high) for a domain in adr.yaml, or the legacy range.""" + if domain in get_domains(): + return tuple(get_domains()[domain].get('range', (0, -1))) + if domain == 'legacy' and 'legacy' in get_config(): + return tuple(get_legacy_range()) + return None + +def source_domain(path: Path, number) -> Optional[str]: + """The domain a record file sits in: its folder, else its number range.""" + return _folder_domain(path) or _range_domain(number) + +def todo_label(item) -> str: + return str(item).split(':')[0] + +def blocking_todo(todo: list) -> list: + """The todo items --partial cannot write past (ADR-306 §3).""" + return [t for t in todo if todo_label(t) in BLOCKING_TODO] + +def _range_domain(number) -> Optional[str]: + base = int(str(number).split('.')[0]) + for domain, config in get_domains().items(): + low, high = config.get('range', (0, -1)) + if low <= base <= high: + return domain + low, high = get_legacy_range() + return 'legacy' if 'legacy' in get_config() and low <= base <= high else None + +def _kind_schema(kind: str) -> dict: + kinds = _mapping(get_config().get('kinds')) + schema = kinds.get(kind) + return schema if isinstance(schema, dict) else IMPORT_DECISION_SCHEMA + +def _source_path(path: Path) -> str: + """The path a sheet records: repo-relative inside the repo, else absolute.""" + resolved = path.resolve() + try: + return str(resolved.relative_to(get_project_root().resolve())) + except ValueError: + return str(resolved) + +def read_record(path: Path) -> dict: + """The sheet for one record file: the v0 reader, or a v1 record read as + itself (ADR-306 §2, §7). Raises SheetError for a source it cannot read.""" + raw = path.read_bytes() + try: + text = raw.decode('utf-8') + except UnicodeDecodeError: + raise SheetError('not UTF-8 text') + if text.startswith('\ufeff'): + raise SheetError('the file starts with a UTF-8 byte order mark; remove it before scanning') + if '\r' in text: + raise SheetError('CRLF line endings; convert the file to LF before scanning') + front, before, title, after = split_record(text) + if front is None: + raise SheetError('no YAML frontmatter; import reads structured records (ADR-306 §2)') + try: + data = yaml.safe_load(front) or {} + except yaml.YAMLError as e: + raise SheetError(f"frontmatter is not valid YAML: {e}") + if not isinstance(data, dict): + raise SheetError('frontmatter is not a mapping of fields') + if title is None: + raise SheetError("no '# ADR-NNN: title' heading") + contract = data.get('contract') + if contract and contract != V1: + raise SheetError(f"contract '{contract}' has no reader") + fmt = 'v1' if contract == V1 else 'v0' + + todo, provenance = [], {} + from_file = filename_number(path) + number = _number_value(from_file or title.group(1)) + provenance['target.number'] = 'file name' if from_file else 'H1' + if from_file and (title.group(1).lstrip('0') or '0') != from_file: + todo.append(f"target.number: the file name says ADR-{from_file}, the H1 says ADR-{title.group(1)}") + domain = source_domain(path, number) + if _folder_domain(path): + provenance['target.domain'] = f"folder {path.parent.name}" + elif domain: + provenance['target.domain'] = 'number range' + else: + todo.append('target.domain: neither the folder nor the number names a domain') + # Under adr/v1 a record in the tree keeps its number wherever its folder + # puts it (ADR-306 §6); the range only allocates new numbers. + moved_in_tree = repo_contract() == V1 and _in_tree(path) + span = number_range(domain) if domain and not moved_in_tree else None + if span and not span[0] <= int(str(number).split('.')[0]) <= span[1]: + todo.append(f"target.number: ADR-{number} is outside the {domain} range {span[0]}-{span[1]}") + provenance['target.title'] = 'H1' + + preamble = before.strip('\n') + if preamble.strip(): + # A record is written H1 first, so text above the H1 moves to just + # below it. The body keeps it; nothing is dropped (ADR-306 §1). + summary, body = None, preamble + '\n' + after + provenance['body'] = 'text above the H1, then the text after it' + todo.append('preamble: the text between the frontmatter and the H1 moved into the body, ' + 'right after the H1; check where it belongs') + else: + summary, body, note = split_summary(after) + if summary is not None: + provenance['summary'] = '## Summary section' + elif note: + provenance['summary'] = note + + unmapped = {} + if fmt == 'v1': + record = dict(data) + provenance['record'] = 'frontmatter (adr/v1)' + kind = record.get('kind') + schema = _kind_schema(kind) if isinstance(kind, str) else IMPORT_DECISION_SCHEMA + else: + schema = _kind_schema('decision') + record = empty_record('decision', schema, {'model': 'unrecorded'}, {}) + provenance['record.contract'] = 'v0 reader' + provenance['record.kind'] = 'v0 reader: every v0 record is a decision' + provenance['record.agent.model'] = 'v0 reader: the model was not recorded' + record['status'] = _v0_status(data, todo, provenance) + for key in V0_CARRIED: + if key in data: + record[key] = data[key] + provenance[f"record.{key}"] = f"frontmatter {key}" + if 'deciders' in data: + provenance['record.deciders'] = 'frontmatter deciders (never an operator basis, ADR-304 §7)' + for key, value in data.items(): + if key != 'status' and key not in V0_CARRIED: + unmapped[key] = value + todo[:0] = open_fields(record, schema) + todo += [f"unmapped: {key} has no v1 field; apply keeps it under imported.unmapped. " + f"Move it into the record, or leave it there" for key in unmapped] + + candidates = {} + if 'capability' in todo: + context = parse_text(text, path).section_text.get('Context', '') + candidates['capability'] = capability_candidates( + f"{title.group(2)}\n{context}", _mapping(get_config().get('capabilities'))) + return { + 'sheet': SHEET_FORMAT, + 'source': {'path': _source_path(path), 'format': fmt, + 'sha256': hashlib.sha256(raw).hexdigest()}, + 'target': {'number': number, 'domain': domain, 'title': ' '.join(title.group(2).split())}, + 'record': _ordered(record), + 'summary': summary, + 'todo': todo, + 'candidates': candidates, + 'provenance': provenance, + 'unmapped': unmapped, + 'body': body, + } + +def _v0_status(data: dict, todo: list, provenance: dict) -> Optional[str]: + """ADR-304 §7's table. Deprecated is superseded when something replaced + it, and otherwise accepted with a note, which needs judgement.""" + raw = data.get('status') + key = str(raw or '').strip().lower() + if key == 'deprecated': + if data.get('superseded_by'): + provenance['record.status'] = f"frontmatter status: {raw}, with superseded_by (ADR-304 §7)" + return 'superseded' + provenance['record.status'] = f"frontmatter status: {raw}, with no successor (ADR-304 §7)" + todo.append('status note: Deprecated with no successor maps to accepted with a note that ' + 'the decision is historical (ADR-304 §7); add the note to the body') + return 'accepted' + if key in V0_STATUS_MAP: + provenance['record.status'] = f"frontmatter status: {raw}" + return V0_STATUS_MAP[key] + if 'status' not in data: + todo.append('status: the source has no status; set one (ADR-304 §7)') + else: + todo.append(f"status: '{raw}' has no v1 mapping; set one (ADR-304 §7)") + return None + +# --- the sheet file ---------------------------------------------------------------- + +class _SheetDumper(_RecordDumper): + """A sheet is edited by hand, so multi-line text is a literal block.""" + +def _represent_sheet_str(dumper, value): + style = '|' if '\n' in value else None + return dumper.represent_scalar('tag:yaml.org,2002:str', value, style=style) + +_SheetDumper.add_representer(str, _represent_sheet_str) + +def dump_sheet(sheet: dict) -> str: + """The sheet as YAML, read back before it is returned (ADR-306 §7).""" + text = yaml.dump(sheet, Dumper=_SheetDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=1000) + if yaml.safe_load(text) != sheet: + raise SheetError('the sheet does not read back as written') + return SHEET_HEADER + text + +def _number_token(path: Path) -> Optional[str]: + """target.number as written, when YAML reads it as an unquoted int, so + 042 (octal), 0x2A or 1_0 can be refused rather than read as another number.""" + try: + node = yaml.compose(path.read_text(), Loader=yaml.SafeLoader) + except (OSError, UnicodeDecodeError, yaml.YAMLError): + return None + for mapping, name in ((node, 'target'), (None, 'number')): + if mapping is None: + mapping = node + if not isinstance(mapping, yaml.MappingNode): + return None + node = next((v for k, v in mapping.value if getattr(k, 'value', None) == name), None) + if isinstance(node, yaml.ScalarNode) and node.tag == 'tag:yaml.org,2002:int' and not node.style: + return node.value + return None + +def load_sheet(path: Path) -> dict: + """A sheet file, checked for the fields apply needs.""" + try: + sheet = yaml.safe_load(path.read_text()) + except yaml.MarkedYAMLError as e: + where = f" at line {e.problem_mark.line + 1}" if e.problem_mark else '' + raise SheetError(f"the sheet is not valid YAML: {e.problem}{where}") + except (OSError, UnicodeDecodeError, yaml.YAMLError) as e: + raise SheetError(f"cannot read the sheet: {e}") + if not isinstance(sheet, dict) or sheet.get('sheet') != SHEET_FORMAT: + raise SheetError(f"not an {SHEET_FORMAT} sheet") + target = sheet.get('target') + if not isinstance(target, dict) or any(_empty(target.get(k)) for k in ('number', 'title')): + raise SheetError('target needs a number and a title') + number = target['number'] + token = _number_token(path) + if token is not None and not re.fullmatch(r'0|[1-9][0-9]*', token): + raise SheetError(f"target.number {token} is not a plain decimal, and YAML reads it as " + f"{number}; quote it, as in number: '{token}'") + if isinstance(number, float): + raise SheetError(f"target.number {number} reads as a decimal; quote a sub-part number, " + f"as in number: '101.10'") + if isinstance(number, bool) or not re.fullmatch(r'\d+(\.\d+)?', str(number)): + raise SheetError(f"target.number '{number}' is not a record number") + if not isinstance(sheet.get('record'), dict): + raise SheetError('record is not a mapping of fields') + for key in ('summary', 'body'): + if sheet.get(key) is not None and not isinstance(sheet[key], str): + raise SheetError(f"{key} is not text") + todo = sheet.get('todo') + if todo is not None and not (isinstance(todo, list) and all(isinstance(t, str) for t in todo)): + raise SheetError('todo is not a list of items') + if sheet.get('unmapped') is not None and not isinstance(sheet['unmapped'], dict): + raise SheetError('unmapped is not a mapping of fields') + source = sheet.get('source') + if source is not None and not (isinstance(source, dict) and isinstance(source.get('path'), str) + and source['path']): + raise SheetError('source needs a path') + return sheet diff --git a/hooks/ways/documentation/adr/src/main.py b/hooks/ways/documentation/adr/src/main.py index 40e5de1d..0f101d70 100644 --- a/hooks/ways/documentation/adr/src/main.py +++ b/hooks/ways/documentation/adr/src/main.py @@ -4,7 +4,7 @@ def main(): parser = argparse.ArgumentParser( - description='ADR - Architecture Decision Record CLI Tool', + description='ADR - Agent Decision Record CLI Tool', formatter_class=argparse.RawDescriptionHelpFormatter, epilog=__doc__ ) @@ -23,6 +23,15 @@ def main(): help='List archived ADRs only') list_scope.add_argument('--all', action='store_true', help='List active and archived ADRs') + p_list.add_argument('--field', action='append', metavar='KEY[=VALUE]', + help='Filter by frontmatter: KEY present, or KEY equal to or listing VALUE (repeatable)') + p_list.add_argument('--kind', help='Filter by kind (same as --field kind=KIND)') + p_list.add_argument('--verb', help='Filter by verb (same as --field verb=VERB)') + p_list.add_argument('--capability', help='Filter by capability, listed or single (same as --field capability=NAME)') + p_list.add_argument('--group-by', dest='group_by', metavar='KEY', + help='Group by a frontmatter field; a record listing several values is in each group') + p_list.add_argument('--json', action='store_true', + help='Machine output: number, title, path, status and frontmatter for each record') # view p_view = subparsers.add_parser('view', aliases=['v', 'show'], help='View an ADR') @@ -32,6 +41,11 @@ def main(): p_new = subparsers.add_parser('new', help='Create new ADR') p_new.add_argument('domain', help='Domain (see `adr domains` for list)') p_new.add_argument('title', help='ADR title') + p_new.add_argument('--kind', help='adr/v1: record kind (default: decision)') + p_new.add_argument('--verb', help='adr/v1: decision verb (add, cut, change, retire, constrain)') + p_new.add_argument('--capability', help='adr/v1: capability from the adr.yaml vocabulary') + p_new.add_argument('--agent', help='adr/v1: the agent writing the record (e.g. Claude)') + p_new.add_argument('--model', help='adr/v1: the model the agent runs on') # rename p_rename = subparsers.add_parser('rename', help='Rename an ADR title and/or file slug') @@ -44,6 +58,25 @@ def main(): p_lint.add_argument('paths', nargs='*', help='Specific files to lint') p_lint.add_argument('--check', action='store_true', help='Exit 1 if errors (CI mode)') + # import (ADR-306) + p_import = subparsers.add_parser('import', help='Import records through import sheets') + import_sub = p_import.add_subparsers(dest='import_command') + p_scan = import_sub.add_parser('scan', help='Write an import sheet for each record') + p_scan.add_argument('paths', nargs='+', help='Record files or directories of records') + p_scan.add_argument('--force', action='store_true', + help='Overwrite a sheet that differs from a fresh scan, discarding its edits') + p_apply = import_sub.add_parser('apply', help='Write finished sheets as adr/v1 records') + p_apply.add_argument('sheets', nargs='*', help='Sheets to apply (default: every sheet in .import/)') + p_apply.add_argument('--partial', action='store_true', + help='Apply sheets with open todo items too, except items lint cannot ' + 'find again afterwards: a status with no mapping, a Deprecated ' + 'note, a number or domain mismatch') + p_apply.add_argument('--force', action='store_true', + help='Overwrite a record that has uncommitted changes') + p_apply.add_argument('--dry-run', action='store_true', + help='Write, lint inside the corpus and print each issue, then restore ' + 'every file and keep every sheet') + # index index_parser = subparsers.add_parser('index', help='Generate ADR index') index_parser.add_argument('-y', '--yes', action='store_true', @@ -52,6 +85,26 @@ def main(): # domains subparsers.add_parser('domains', help='List domain number series') + # domain (ADR-306 §6) + p_domain = subparsers.add_parser('domain', help='Add, rename or move domains') + domain_sub = p_domain.add_subparsers(dest='domain_command') + p_dadd = domain_sub.add_parser('add', help='Add a domain to adr.yaml') + p_dadd.add_argument('name', help='Domain key (e.g. ops)') + p_dadd.add_argument('--range', required=True, help='Number range for new records, A-B (e.g. 400-499)') + p_dadd.add_argument('--folder', required=True, help='Folder under docs/architecture') + p_dadd.add_argument('--label', help='Display name (default: the key, capitalized)') + p_dadd.add_argument('--description', help='One line on what the domain covers') + p_dren = domain_sub.add_parser('rename', help='Rename a domain, and its folder, in place') + p_dren.add_argument('old', help='Current domain key') + p_dren.add_argument('new', help='New domain key') + p_dren.add_argument('--folder', help='New folder (default: the new key when the folder was the old key)') + p_dren.add_argument('--dry-run', action='store_true', help='Print each line the rename would change; write nothing') + p_dmove = domain_sub.add_parser('move', help='Move records to another domain; numbers never change') + p_dmove.add_argument('record', nargs='?', help='ADR number (e.g. 104, ADR-104)') + p_dmove.add_argument('domain', nargs='?', help='Target domain') + p_dmove.add_argument('--plan', help='YAML list of {record, domain} moves, applied together') + p_dmove.add_argument('--dry-run', action='store_true', help='Print the moves and each line they would change; write nothing') + # archive p_archive = subparsers.add_parser( 'archive', help='Archive an ADR out of the active set') @@ -79,12 +132,42 @@ def main(): p_close.add_argument('--reason', help='Why (required; appended as a Closure section)') p_close.add_argument('--dry-run', action='store_true', help='Report without writing') + # record edits (adr/v1): consider, set, supersede, enact + p_consider = subparsers.add_parser('consider', help="Append a considered entry: the operator's answer (ADR-304 §12)") + p_consider.add_argument('adr', help='ADR number (e.g., 101, ADR-101)') + p_consider.add_argument('--said', help='What the operator said, verbatim (required)') + p_consider.add_argument('--via', help='Where it was said, such as a PR or a session (required)') + p_consider.add_argument('--operator', help='Who said it (default: the gh or git user)') + p_consider.add_argument('--covers', nargs='*', metavar='PROBE', + help='Probe names from the Summary the answer covers (none given writes covers: [])') + p_consider.add_argument('--paraphrase', action='store_true', help='said is a summary, not the words') + p_consider.add_argument('--canary', choices=['caught', 'missed'], help='Whether the operator caught the canary') + p_consider.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_set = subparsers.add_parser('set', help='Edit frontmatter fields: key=value, key+=item, key-=item') + p_set.add_argument('adr', help='ADR number (e.g., 101, ADR-101)') + p_set.add_argument('assignments', nargs='+', metavar='key=value', + help='Values are YAML: capability=[a, b], related+=ADR-7, observable+="..."') + p_set.add_argument('--force', action='store_true', + help='Set status, or a field no v1 record carries, anyway') + p_set.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_supersede = subparsers.add_parser('supersede', help='Record a supersession on both records (ADR-304 §3)') + p_supersede.add_argument('adr', help='The record replaced (e.g., 101, ADR-101)') + p_supersede.add_argument('--by', required=True, help='The record that replaces it') + p_supersede.add_argument('--amends', metavar='SECTION', + help='Replace one section only: amends: [OLD#SECTION] on the new record') + p_supersede.add_argument('--dry-run', action='store_true', help='Show the change without writing') + p_enact = subparsers.add_parser('enact', help='Mark an accepted cut or retire done at a commit (ADR-304 §5)') + p_enact.add_argument('adr', help='ADR number (e.g., 111, ADR-111)') + p_enact.add_argument('commit', help='The commit hash that finished the removal') + p_enact.add_argument('--dry-run', action='store_true', help='Show the change without writing') + # cite p_cite = subparsers.add_parser('cite', help='Check ADR citations in code against the records') p_cite.add_argument('paths', nargs='*', help='Limit the scan to these files or directories') p_cite.add_argument('--check', action='store_true', help='Exit 1 if errors (CI mode)') - p_cite.add_argument('--no-inventory', action='store_true', - help="Skip surface inventories (they run shell commands from adr.yaml)") + # Accepted and ignored: surface inventories are gone (ADR-311), and a + # script that still passes the flag keeps working. + p_cite.add_argument('--no-inventory', action='store_true', help=argparse.SUPPRESS) args = parser.parse_args() @@ -104,11 +187,17 @@ def main(): 'index': cmd_index, 'archive': cmd_archive, 'domains': cmd_domains, + 'domain': cmd_domain, 'config': cmd_config, 'cite': cmd_cite, 'accept': cmd_accept, 'reject': cmd_reject, 'abandon': cmd_abandon, + 'import': cmd_import, + 'consider': cmd_consider, + 'set': cmd_set, + 'supersede': cmd_supersede, + 'enact': cmd_enact, } return commands[args.command](args) diff --git a/hooks/ways/documentation/adr/src/parse.py b/hooks/ways/documentation/adr/src/parse.py index dca4ebfa..69ccf262 100644 --- a/hooks/ways/documentation/adr/src/parse.py +++ b/hooks/ways/documentation/adr/src/parse.py @@ -106,29 +106,36 @@ def as_list(value): info.title = match.group(2) break - # Determine domain from number range (authoritative) or folder name (fallback) - # Range takes precedence: an ADR's number definitively places it in a domain, - # even if the file is physically in a different domain's folder. - if info.number: - try: - base_num = int(info.number.split('.')[0]) + # Determine the domain. Under adr/v0 the number range is authoritative and + # the folder is the fallback: a number definitively places a record, even + # in another domain's folder. Under adr/v1 a number is only an identity, + # never renumbered (ADR-306 §6): the folder decides, and the range, which + # allocates new numbers, is the fallback for a folder no domain names. + def domain_by_range(): + if info.number: + try: + base_num = int(info.number.split('.')[0]) + except ValueError: + return None for domain, config in get_domains().items(): if config['range'][0] <= base_num <= config['range'][1]: - info.domain = domain - break - except ValueError: - pass + return domain + return None - # Fallback: determine from folder name (for unnumbered or out-of-range ADRs) - if not info.domain: + def domain_by_folder(): folder_name = path.parent.name for domain, config in get_domains().items(): folders = config['folder'] if isinstance(folders, str): folders = [folders] if folder_name in folders: - info.domain = domain - break + return domain + return None + + if repo_contract() == 'adr/v1': + info.domain = domain_by_folder() or domain_by_range() + else: + info.domain = domain_by_range() or domain_by_folder() return info diff --git a/hooks/ways/documentation/adr/src/record_edit.py b/hooks/ways/documentation/adr/src/record_edit.py new file mode 100644 index 00000000..3502002d --- /dev/null +++ b/hooks/ways/documentation/adr/src/record_edit.py @@ -0,0 +1,209 @@ + +# --- surgical frontmatter edits ----------------------------------------------------- +# +# The record commands (consider, set, supersede, enact) change one field at a +# time. They rewrite only the lines of that field and leave every other byte of +# the file as it was: the other fields' YAML style, comments, line endings and +# the body. A new field is written in the tool's own style (_RecordDumper) at +# its place in V1_KEY_ORDER. A field keeps its layout where it can: a one-line +# field stays on one line, and a block list gains or loses item lines. Every +# edit is read back before it is returned, so a mismatch refuses the edit +# rather than writing something else. + +_TOP_KEY = re.compile(r'^([A-Za-z_][A-Za-z0-9_-]*)\s*:(?:\s|$)') + +class _Quoted(str): + """A string written double-quoted, as the corpus writes `said` and a + commit hash. It reads back as a plain str.""" + +_RecordDumper.add_representer( + _Quoted, lambda dumper, value: dumper.represent_scalar('tag:yaml.org,2002:str', str(value), style='"')) + +class _FlowList(list): + """A list written on one line, as the corpus writes a considered + entry's covers. It reads back as a plain list.""" + +_RecordDumper.add_representer( + _FlowList, lambda dumper, value: dumper.represent_sequence('tag:yaml.org,2002:seq', list(value), flow_style=True)) + +def _prefer_double(value: str): + """value double-quoted where YAML would single-quote it, since the + corpus quotes with double quotes; plain where it needs no quotes.""" + return _Quoted(value) if _flow(value).startswith("'") else value + +def _flow(value) -> str: + return yaml.dump(value, Dumper=_RecordDumper, default_flow_style=True, sort_keys=False, + allow_unicode=True, width=100000).rstrip('\n').removesuffix('\n...').rstrip('\n') + +def _block(key: str, value) -> list: + text = yaml.dump({key: value}, Dumper=_RecordDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=100000) + return text.rstrip('\n').split('\n') + +def _plain(value): + """value with _Quoted strings turned back into str, for comparison.""" + if isinstance(value, list): + return [_plain(v) for v in value] + if isinstance(value, dict): + return {k: _plain(v) for k, v in value.items()} + return str(value) if isinstance(value, _Quoted) else value + +class FrontmatterEdit: + """One record's file, split into its frontmatter lines and the rest.""" + + def __init__(self, raw: bytes): + text = raw.decode('utf-8') + lines = text.split('\n') + if not lines or lines[0].rstrip('\r').strip() != '---': + raise ValueError("the file has no YAML frontmatter") + end = next((i for i, line in enumerate(lines[1:], 1) if line.strip() == '---'), None) + if end is None: + raise ValueError("the frontmatter has no closing ---") + self.crlf = lines[0].endswith('\r') + self.head = lines[:1] + self.tail = lines[end:] + self.lines = [line[:-1] if self.crlf and line.endswith('\r') else line for line in lines[1:end]] + self.fields = self._load() + + def _load(self) -> dict: + data = yaml.safe_load('\n'.join(self.lines)) or {} + if not isinstance(data, dict): + raise ValueError("the frontmatter is not a mapping of fields") + return data + + def bytes(self) -> bytes: + body = [line + '\r' for line in self.lines] if self.crlf else list(self.lines) + return '\n'.join(self.head + body + self.tail).encode('utf-8') + + def _blocks(self) -> list: + """(key, first line, last line) of each top-level field. A field's + block ends at its last indented line: a comment or blank line at + column 0 before the next key stays outside it.""" + blocks = [] + for i, line in enumerate(self.lines): + match = _TOP_KEY.match(line) + if match: + blocks.append([match.group(1), i, i]) + elif blocks and line[:1] in (' ', '\t') and line.strip(): + blocks[-1][2] = i + elif blocks and line.strip() and not line.startswith('#'): + blocks[-1][2] = i # a continuation at column 0, such as a flow list's rest + return [tuple(b) for b in blocks] + + def _block(self, key: str) -> Optional[tuple]: + found = [b for b in self._blocks() if b[0] == key] + if len(found) > 1: + raise ValueError(f"'{key}' appears {len(found)} times in the frontmatter") + return found[0] if found else None + + def _replace(self, start: int, end: int, new_lines: list) -> None: + self.lines[start:end + 1] = new_lines + + def _insert_at(self, key: str) -> int: + """The line a new field goes in at: after the nearest field before it in + V1_KEY_ORDER, else before the nearest field after it, else at the end.""" + blocks = self._blocks() + if key in V1_KEY_ORDER: + rank = V1_KEY_ORDER.index(key) + before = [b for b in blocks if b[0] in V1_KEY_ORDER and V1_KEY_ORDER.index(b[0]) < rank] + if before: + return max(b[2] for b in before) + 1 + after = [b for b in blocks if b[0] in V1_KEY_ORDER and V1_KEY_ORDER.index(b[0]) > rank] + if after: + return min(b[1] for b in after) + return blocks[-1][2] + 1 if blocks else len(self.lines) + + def _one_line(self, key: str, value, old: str) -> str: + """key: value on one line, keeping a trailing comment of the old line.""" + text = _flow(value) if isinstance(value, (list, dict)) else _block(key, value)[0][len(key) + 2:] + comment = re.search(r'\s+#[^\'"\]]*$', old) + return f"{key}: {text}" + (comment.group(0) if comment else '') + + def _render(self, key: str, value, block: Optional[tuple]) -> list: + """The lines for key: value. An existing one-line field stays on one + line; a new field, or one written as a block, takes the tool's style.""" + if block is not None and block[1] == block[2]: + one = self._one_line(key, value, self.lines[block[1]]) + if not (isinstance(value, list) and value and any(isinstance(v, dict) for v in value)): + return [one] + lines = _block(key, value) + if block is not None and len(lines) > 1: + lines = [lines[0]] + self._reindent(lines[1:], block) + return lines + + def _item_indent(self, block: tuple) -> Optional[str]: + for line in self.lines[block[1] + 1:block[2] + 1]: + match = re.match(r'^(\s*)- ', line) or re.match(r'^(\s*)-$', line) + if match: + return match.group(1) + return None + + def _reindent(self, item_lines: list, block: tuple) -> list: + """Item lines rendered with the tool's two-space indent, moved to the + indent the field's existing items use.""" + indent = self._item_indent(block) + if indent is None or indent == ' ': + return item_lines + return [indent + line[2:] if line.startswith(' ') else line for line in item_lines] + + def _commit(self, expected: dict) -> None: + if _plain(self._load()) != _plain(expected): + raise ValueError("the edited frontmatter does not read back as intended") + self.fields = self._load() + + # --- operations ------------------------------------------------------------- + + def set(self, key: str, value) -> None: + expected = dict(self.fields) + expected[key] = value + block = self._block(key) + if block is None: + at = self._insert_at(key) + self.lines[at:at] = _block(key, value) + else: + self._replace(block[1], block[2], self._render(key, value, block)) + self._commit(expected) + + def append(self, key: str, items: list) -> None: + """Append items to a list field, creating it when absent. A scalar + becomes a list holding it first.""" + current = self.fields.get(key) + before = [] if current in (None, '') else (list(current) if isinstance(current, list) else [current]) + value = before + list(items) + block = self._block(key) + indent = self._item_indent(block) if block is not None else None + if block is None or indent is None or not isinstance(current, list) or not current: + self.set(key, value) + return + expected = dict(self.fields) + expected[key] = value + new = self._reindent(_block(key, list(items))[1:], block) + self.lines[block[2] + 1:block[2] + 1] = new + self._commit(expected) + + def remove(self, key: str, items: list) -> list: + """Remove each item from a list field. Returns the items not found.""" + current = self.fields.get(key) + entries = current if isinstance(current, list) else ([] if current is None else [current]) + wanted = [str(_plain(i)) for i in items] + missing = [i for i, w in zip(items, wanted) if w not in [str(e) for e in entries]] + if missing: + return missing + value = [e for e in entries if str(e) not in wanted] + block = self._block(key) + indent = self._item_indent(block) if block is not None else None + if indent is None or not value: + self.set(key, value) + return [] + expected = dict(self.fields) + expected[key] = value + starts = [i for i in range(block[1] + 1, block[2] + 1) + if re.match(re.escape(indent) + r'-(\s|$)', self.lines[i])] + chunks = [(s, (starts[n + 1] - 1) if n + 1 < len(starts) else block[2]) for n, s in enumerate(starts)] + for start, end in reversed(chunks): + text = '\n'.join(line[len(indent):] for line in self.lines[start:end + 1]) + loaded = yaml.safe_load(text) + if isinstance(loaded, list) and len(loaded) == 1 and str(loaded[0]) in wanted: + del self.lines[start:end + 1] + self._commit(expected) + return [] diff --git a/hooks/ways/documentation/adr/src/relocate.py b/hooks/ways/documentation/adr/src/relocate.py new file mode 100644 index 00000000..f32e3917 --- /dev/null +++ b/hooks/ways/documentation/adr/src/relocate.py @@ -0,0 +1,308 @@ +# ============================================================================ +# Path rewriting for records that change folder (ADR-306 §6) +# ============================================================================ +# +# A record's number is its identity and never changes (ADR-310). Under adr/v1 +# its folder decides its domain, so moving a record to another domain moves +# its file, and a domain renamed in place moves its folder. Either way every +# path to what moved is rewritten. A path is a relative link that resolves to +# what moved, a path written from the repo root that names it, or a URL into +# this repository at a branch that names it. Other URLs, a URL at a commit or +# a tag (a permalink), a path inside a fenced code block, a path written with +# backslashes, and words that only contain a folder's name are left alone. +# ADR-N citations are left alone too: the number still names the same record. + +# Frontmatter keys that record history and are never rewritten: an imported +# record keeps the path it was imported from. +RELOCATE_HISTORY_KEYS = ('imported',) + +_PATH_WORD_RE = re.compile(r'[A-Za-z][\w+.-]*://[\w./~%-]+|[\w./-]+') +# host/owner/repo/blob|tree|raw/<ref>/<path>, and GitHub's raw host, where +# the ref follows the repo directly. A ref may hold slashes. +_REPO_URL_RE = re.compile(r'[A-Za-z][\w+.-]*://([^/]+)/([^/]+)/([^/]+)/(?:blob|tree|raw)/(.+)') +_RAW_URL_RE = re.compile(r'[A-Za-z][\w+.-]*://raw\.githubusercontent\.com/([^/]+)/([^/]+)/(.+)', re.IGNORECASE) +_COMMIT_RE = re.compile(r'[0-9a-fA-F]{7,40}') +_FM_KEY_RE = re.compile(r'([A-Za-z_][\w-]*)\s*:') +_FENCE_RE = re.compile(r'[ \t]*(`{3,}|~{3,})') + + +def _repo_name(url: str) -> Optional[str]: + """host/owner/repo, lowercase, from a remote URL or a written name.""" + m = re.fullmatch(r'(?:[A-Za-z][\w+.-]*://)?(?:[^@/]+@)?([^/:]+)(?::\d+)?[:/]([^/]+/[^/]+?)(?:\.git)?/?', + url.strip()) + return f"{m.group(1)}/{m.group(2)}".lower() if m else None + + +def project_repos(root: Path) -> set: + """The names this repository goes by, host/owner/repo, lowercase: + adr.yaml's `repository:` (a name or a list of names) when it is set, + otherwise the origin remote. Empty without either.""" + configured = get_config().get('repository') + if configured: + values = [configured] if isinstance(configured, str) else configured if isinstance(configured, list) else [] + return {n for n in (_repo_name(str(v)) for v in values) if n} + name = _repo_name((_git(['remote', 'get-url', 'origin'], root) or '').strip()) + return {name} if name else set() + + +class Repository: + """This repository as URLs name it: its names, and its tags and branches, + each read from git once when first needed. A URL into it at a branch + names a path in the working tree. At a commit or a tag it is a + permalink to a snapshot, and names nothing that moves.""" + + def __init__(self, root: Path): + self.root = root + self.names = project_repos(root) + self._tags = self._branches = None + + @property + def tags(self) -> set: + if self._tags is None: + self._tags = set((_git(['tag', '-l'], self.root) or '').split()) + return self._tags + + @property + def branches(self) -> set: + """Local branches, and remote-tracking ones by their branch name.""" + if self._branches is None: + out = _git(['for-each-ref', '--format=%(refname)', 'refs/heads', 'refs/remotes'], self.root) or '' + names = set() + for ref in out.split(): + if ref.startswith('refs/heads/'): + names.add(ref[len('refs/heads/'):]) + elif ref.startswith('refs/remotes/') and ref.count('/') >= 3: + names.add(ref.split('/', 3)[3]) + self._branches = names + return self._branches + + def path(self, token: str) -> Optional[tuple]: + """(where the path starts in token, the path from the repo root) for + a URL into this repository at a branch; None otherwise. A ref with a + slash is read as the longest leading run of segments that names a + known branch; otherwise the ref is the first segment.""" + m = _RAW_URL_RE.fullmatch(token) + if m: + repo, rest, at = f"github.com/{m.group(1)}/{m.group(2)}", m.group(3), m.start(3) + else: + m = _REPO_URL_RE.fullmatch(token) + if not m: + return None + repo, rest, at = '/'.join(m.group(1, 2, 3)), m.group(4), m.start(4) + repo = repo.lower() + if repo.endswith('.git'): + repo = repo[:-4] + if repo not in self.names: + return None + segments = rest.split('/') + if len(segments) < 2: + return None + width = next((n for n in range(len(segments) - 1, 1, -1) + if '/'.join(segments[:n]) in self.branches), 1) + ref = '/'.join(segments[:width]) + if width == 1 and ref not in self.branches and (_COMMIT_RE.fullmatch(ref) or ref in self.tags): + return None + path = '/'.join(segments[width:]) + return (at + len(ref) + 1, path) if path else None + + +def _fenced(lines: list, start: int = 0) -> set: + """Indexes of the lines from start on that sit in a fenced code block, + fence lines included.""" + inside, fence = set(), None + for i in range(start, len(lines)): + m = _FENCE_RE.match(lines[i]) + if fence is None: + if m: + fence = m.group(1) + inside.add(i) + else: + inside.add(i) + if m and m.group(1)[0] == fence[0] and len(m.group(1)) >= len(fence) \ + and not lines[i][m.end():].strip(): + fence = None + return inside + + +def _sub_paths(line: str, swap) -> str: + """line with each run of path characters replaced by swap(token), or + left alone when it touches a backslash: a Windows path is not a + path this tool resolves.""" + def one(m): + before = m.string[m.start() - 1:m.start()] + after = m.string[m.end():m.end() + 1] + if before == '\\' or after == '\\': + return m.group(0) + return swap(m.group(0)) + return _PATH_WORD_RE.sub(one, line) + + +class Relocation: + """One simultaneous move: files (repo-relative old path -> new path), + directories (old dir -> new dir), and optionally a domain key rename for + catalog frontmatter (`domain: <key>`). `known` holds tracked paths and + directories: a path is rewritten only when it names one, or names a + moved file. `repo` (a Repository) says which URLs point into this + repository at a branch; such a URL is rewritten like a path from the + root.""" + + def __init__(self, files=None, dirs=None, known=None, domain=None, repo=None): + self.files = dict(files or {}) + self.dirs = dict(dirs or {}) + self.known = set(known or ()) + self.domain = domain # (old key, new key) or None + self.repo = repo + self.basenames = {posixpath.basename(p) for p in self.files} + + def target(self, rel: str) -> Optional[str]: + """Where a repo-relative path is after the move, or None if it stays.""" + if rel in self.files: + return self.files[rel] + for old, new in self.dirs.items(): + if rel == old or rel.startswith(old + '/'): + return new + rel[len(old):] + return None + + def _moved(self, rel: str) -> Optional[str]: + """target(rel) when rel names a real path that moved: a moved file, or + a path under a moved folder that is tracked before or after the move.""" + if rel in self.files: + return self.files[rel] + moved = self.target(rel) + if moved is not None and (rel in self.known or moved in self.known): + return moved + return None + + def _rooted(self, core: str) -> Optional[str]: + """core read as a path from the repo root, where it names something + that moved; None otherwise.""" + lead = '/' if core.startswith('/') else '' + rest = core[len(lead):] + if not rest or posixpath.normpath(rest) != rest: + return None + moved = self._moved(rest) + return lead + moved if moved is not None else None + + def _path(self, token: str, old_dir: str, new_dir: str) -> Optional[str]: + """token rewritten as a path to something that moved, or None.""" + slash = token.endswith('/') and len(token) > 1 + core = token.rstrip('/') if slash else token + if not core: + return None + # Relative to the file that holds it, resolved from where that file + # was. A file that moves re-bases its links to paths that stay. + if not core.startswith('/'): + resolved = posixpath.normpath(posixpath.join(old_dir, core)) + if not resolved.startswith('..'): + moved = self._moved(resolved) + if moved is not None or (old_dir != new_dir and resolved in self.known): + dest = moved or resolved + if posixpath.normpath(posixpath.join(new_dir, core)) == dest: + return None + # Keep the link's shape where it still resolves: rename + # only the segments that moved (../system/X -> ../platform/X). + segments, walked = [], old_dir + for segment in core.split('/'): + walked = posixpath.normpath(posixpath.join(walked, segment)) if segment else walked + after = self.target(walked) if segment not in ('', '.', '..') else None + segments.append(posixpath.basename(after) if after else segment) + new = '/'.join(segments) + if posixpath.normpath(posixpath.join(new_dir, new)) != dest: + new = posixpath.relpath(dest, new_dir or '.') + if core.startswith('./') and not new.startswith('.'): + new = './' + new + return new + ('/' if slash else '') + if resolved in self.known: + return None # a real path that stays + # A path written from the repo root names what moved in full. + new = self._rooted(core) + return (new + ('/' if slash else '')) if new is not None and new != core else None + + def _url(self, token: str) -> Optional[str]: + """A URL into this repository at a branch, with the path it names + rewritten.""" + found = self.repo.path(token) if self.repo else None + if found is None: + return None + at, path = found + slash = path.endswith('/') + new = self._rooted(path.rstrip('/')) + if new is None: + return None + return token[:at] + new + ('/' if slash else '') + + def word(self, token: str, old_dir: str, new_dir: str) -> tuple: + """One run of path characters rewritten; (text, count).""" + stripped = token.rstrip('.') # sentence punctuation is not the path + if not stripped: + return token, 0 + if '://' in stripped: + new = self._url(stripped) + else: + if '/' not in stripped and stripped not in self.basenames: + # A bare sibling name is a path only when the file holding it + # moved and the name resolves to a tracked file: the link must + # be re-based even though its target stays put. + sibling = posixpath.normpath(posixpath.join(old_dir, stripped)) + if old_dir == new_dir or sibling not in self.known: + return token, 0 + new = self._path(stripped, old_dir, new_dir) + if new is None: + return token, 0 + return new + token[len(stripped):], 1 + + def text(self, content: str, rel: str) -> tuple: + """content of the file at rel with its paths rewritten; (text, count). + History keys in frontmatter are left alone, a frontmatter `domain:` + follows a domain rename, and in a Markdown file a fenced code block + is left as written: it quotes text, such as a command, rather than + linking to a file.""" + moved = self.target(rel) + old_dir = posixpath.dirname(rel) + new_dir = posixpath.dirname(moved) if moved else old_dir + lines = content.split('\n') + fm_end = None + if lines and lines[0].rstrip('\r') == '---': + fm_end = next((i for i in range(1, len(lines)) if lines[i].rstrip('\r') == '---'), None) + fenced = _fenced(lines, fm_end + 1 if fm_end is not None else 0) if rel.lower().endswith('.md') else set() + total, key, out = 0, None, [] + + def swap(token): + nonlocal total + new, n = self.word(token, old_dir, new_dir) + total += n + return new + + for i, line in enumerate(lines): + if i in fenced: + out.append(line) + continue + if fm_end is not None and 0 < i < fm_end: + m = _FM_KEY_RE.match(line) + if m: + key = m.group(1) + if key in RELOCATE_HISTORY_KEYS: + out.append(line) + continue + if self.domain and m and key == 'domain': + dm = re.fullmatch(r'(domain:\s*)([\w-]+)(\s*(?:#.*)?)', line.rstrip('\r')) + if dm and dm.group(2) == self.domain[0]: + out.append(dm.group(1) + self.domain[1] + dm.group(3) + + ('\r' if line.endswith('\r') else '')) + total += 1 + continue + out.append(_sub_paths(line, swap)) + return '\n'.join(out), total + + +def tracked_paths(root: Path) -> tuple: + """(every file git tracks, sorted; those files and their folders)""" + out = _git(['ls-files', '-z'], root) + tracked = sorted(n for n in (out or '').split('\0') if n) + known = set(tracked) + for name in tracked: + parent = posixpath.dirname(name) + while parent and parent not in known: + known.add(parent) + parent = posixpath.dirname(parent) + return tracked, known diff --git a/hooks/ways/documentation/adr/src/rules_v1.py b/hooks/ways/documentation/adr/src/rules_v1.py index 83ccb1f7..628ade93 100644 --- a/hooks/ways/documentation/adr/src/rules_v1.py +++ b/hooks/ways/documentation/adr/src/rules_v1.py @@ -37,17 +37,6 @@ def _str_list(value) -> Optional[list]: def _mapping(value) -> dict: return value if isinstance(value, dict) else {} -def _iso_date(value) -> Optional[str]: - """value as a YYYY-MM-DD string, or None. YAML may load a date as one.""" - text = str(value) if value is not None else '' - return text if re.fullmatch(r'\d{4}-\d{2}-\d{2}', text) else None - -def v1_baseline(ctx) -> tuple: - """(adoption date or None, set of baseline capability names).""" - baseline = _mapping(ctx.config.get('baseline')) - return (_iso_date(baseline.get('adopted')), - set(_str_list(baseline.get('capabilities')) or [])) - def v1_kinds(ctx) -> dict: return _mapping(ctx.config.get('kinds')) @@ -83,9 +72,6 @@ def v1_edges(schema: dict) -> dict: def v1_capabilities(ctx) -> dict: return _mapping(ctx.config.get('capabilities')) -def v1_surfaces(ctx) -> dict: - return _mapping(ctx.config.get('surfaces')) - def as_entries(value) -> list: if value is None: return [] @@ -95,17 +81,6 @@ def capability_scope(adr) -> list: """The capabilities a record covers: one name, a list, or ['*'].""" return as_entries(adr.frontmatter.get('capability')) -def covers(prior, capability: str) -> bool: - """A prior decision is on this capability when it names it, lists it, - or is scoped to '*' (ADR-304 §3).""" - scope = capability_scope(prior) - return '*' in scope or capability in scope - -def broader_than(prior, capability: str) -> bool: - """The prior covers more than this one capability: '*' or a list.""" - scope = capability_scope(prior) - return '*' in scope or len(scope) > 1 - def section_exists(target, section: str) -> bool: """A section reference matches a heading that is numbered with it ('2.', '2 ') or whose slug equals it ('priority-bands').""" @@ -139,9 +114,6 @@ def bad(message): for key in ('requires', 'statuses', 'sections'): if key in schema and _str_list(schema[key]) is None: bad(f"kinds.{name}.{key}: expected a list of names") - mutable = schema.get('mutable_after_accept') - if mutable is not None and mutable != 'all' and _str_list(mutable) is None: - bad(f"kinds.{name}.mutable_after_accept: expected a list of fields or 'all'") if 'edges' in schema: edges = schema['edges'] if not isinstance(edges, dict): @@ -154,44 +126,11 @@ def bad(message): bad("verbs: expected a list of names") if not isinstance(config.get('capabilities'), dict) or not config.get('capabilities'): bad("contract: adr/v1 declares no capabilities") - if 'surfaces' in config and not isinstance(config['surfaces'], dict): - bad("surfaces: expected a mapping of namespace to settings") - if 'baseline' in config: - baseline = config['baseline'] - names = _str_list(baseline.get('capabilities')) if isinstance(baseline, dict) else None - if names is None: - bad("baseline: expected 'adopted' (a YYYY-MM-DD date) and 'capabilities' (a list of names)") - else: - if not _iso_date(baseline.get('adopted')): - bad("baseline.adopted: expected a YYYY-MM-DD date") - if isinstance(config.get('capabilities'), dict): - for name in names: - if name not in config['capabilities']: - bad(f"baseline: '{name}' is not in the capabilities vocabulary") - -@config_rule(contract=V1) -def rule_v1_capabilities_added(ctx): - """Every capability in the vocabulary has an accepted `add` decision, except - those in `baseline`: capabilities that were active when the project - adopted the contract, which no record added. This warns while v0 records - remain, so a migrating corpus is not failed on every capability, and fails - once migration is done (ADR-304 §6, amended by ADR-305).""" - added = set() - for adr in ctx.corpus: - if (is_v1_record(adr, ctx) - and adr.frontmatter.get('verb') == 'add' - and str(adr.status or '').lower() == 'accepted'): - added.update(capability_scope(adr)) - # Archived records are never edited, so they never migrate and do not - # hold the check at warning. - migrating = any(not is_v1_record(adr, ctx) and not is_archived(adr.path) - for adr in ctx.corpus) - _, baseline = v1_baseline(ctx) - for name in v1_capabilities(ctx): - if name not in added and name not in baseline: - ctx.config_issues.append(Issue( - f"capability '{name}' has no accepted add decision", - 'warning' if migrating else 'error')) + if 'repository' in config: + names = [config['repository']] if isinstance(config['repository'], str) else config['repository'] + if not isinstance(names, list) or not names \ + or not all(isinstance(n, str) and _repo_name(n) for n in names): + bad("repository: expected host/owner/repo, or a list of them") # --- one record ------------------------------------------------------------------ @@ -251,9 +190,12 @@ def rule_v1_capability(adr, ctx): if raw in (None, '', []): return # reported by rule_v1_required_fields when the kind requires it scope = capability_scope(adr) - if adr.frontmatter.get('verb') != 'constrain': - if isinstance(raw, list): - v1_issue(adr, "capability takes one name; only a constrain decision takes a list") + verb = adr.frontmatter.get('verb') + if verb != 'constrain': + # A change may list the capabilities it alters (ADR-308); only a + # constrain may be scoped to '*'. + if isinstance(raw, list) and verb != 'change': + v1_issue(adr, "capability takes one name; only change and constrain decisions take a list") return if '*' in scope: v1_issue(adr, "only a constrain decision may be scoped to '*'") @@ -262,21 +204,8 @@ def rule_v1_capability(adr, ctx): for name in scope: if name != '*' and name not in vocabulary: v1_issue(adr, f"capability '{name}' is not in the adr.yaml vocabulary") - -@file_rule(contract=V1) -def rule_v1_retire_targets(adr, ctx): - if not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'retire': - return - targets = as_entries(adr.frontmatter.get('targets')) - if not targets: - v1_issue(adr, "a retire decision names its targets (targets: [cli:..., route:...])") - surfaces = v1_surfaces(ctx) - for target in targets: - namespace, sep, name = target.partition(':') - if not sep or not name: - v1_issue(adr, f"target '{target}' is not namespace:name") - elif namespace not in surfaces: - v1_issue(adr, f"target '{target}': surface '{namespace}' is not declared in adr.yaml") + if verb == 'change' and len(scope) > 3: + v1_issue(adr, f"a change lists {len(scope)} capabilities; list only those it alters, the rest belong in related (ADR-308 §3)", 'warning') # --- against the corpus ------------------------------------------------------------ @@ -310,61 +239,3 @@ def rule_v1_edges(adr, ctx): v1_issue(adr, f"{field_name}: ADR-{number} is a {target_kind}, expected {' or '.join(target_kinds)}") if section and not section_exists(target, section): v1_issue(adr, f"{field_name}: ADR-{number} has no section '{section}'") - -def decision_order(adr) -> tuple: - """Records in decision order: by date, then number (ADR-304 §3).""" - base, _, part = str(adr.number or '0').partition('.') - return (str(adr.date or ''), int(base or 0), int(part or 0)) - -def _stands_on_baseline(capability: str, adr, ctx) -> bool: - """A change with no prior edge stands on the baseline when the capability - is in it and either the change predates adoption, or no live decision on - the capability has been made since adoption before it. Only decisions - dated after adoption count: each was written as v1, so migrating an older - record never moves the answer (ADR-305).""" - adopted, baseline = v1_baseline(ctx) - when = _iso_date(adr.date) - if capability not in baseline or not adopted or not when: - return False - if when <= adopted: - return True - for other in ctx.corpus: - other_when = _iso_date(other.date) - if (other is not adr and is_v1_record(other, ctx) - and other.frontmatter.get('verb') not in (None, 'constrain') - and str(other.status or '').lower() not in ('rejected', 'abandoned') - and covers(other, capability) - and other_when and other_when > adopted - and decision_order(other) < decision_order(adr)): - return False - return True - -@corpus_rule(contract=V1) -def rule_v1_change_replaces(adr, ctx): - """A change decision supersedes or amends a prior decision on the same - capability. When the prior covers more than this capability ('*' or a - list), the change amends it (ADR-304 §3). A change on a baseline - capability with no prior record to name stands on the baseline (ADR-305).""" - if not is_v1_record(adr, ctx) or adr.frontmatter.get('verb') != 'change': - return - edges = [] - for field_name in ('supersedes', 'amends'): - for entry in as_entries(adr.frontmatter.get(field_name)): - target = ctx.by_number.get(norm_ref(entry)[0]) - if target is None: - return # a dangling edge is already reported by the edge rules - edges.append((field_name, target)) - for capability in [c for c in capability_scope(adr) if c != '*']: - v0_priors = [t for _, t in edges if not is_v1_record(t, ctx)] - fits = [(f, t) for f, t in edges if is_v1_record(t, ctx) and covers(t, capability)] - if any(f == 'amends' or not broader_than(t, capability) for f, t in fits): - continue - if fits: - v1_issue(adr, f"a change on '{capability}' against a broader decision amends it rather than superseding it") - elif not edges and _stands_on_baseline(capability, adr, ctx): - continue - elif v0_priors: - numbers = ', '.join(f"ADR-{t.number}" for t in v0_priors) - v1_issue(adr, f"cannot confirm the prior decision on '{capability}': {numbers} is still v0", 'warning') - else: - v1_issue(adr, f"a change decision supersedes or amends a prior decision on '{capability}'") diff --git a/hooks/ways/documentation/adr/src/rules_v1_basis.py b/hooks/ways/documentation/adr/src/rules_v1_basis.py index 3bb9badf..5367d1e5 100644 --- a/hooks/ways/documentation/adr/src/rules_v1_basis.py +++ b/hooks/ways/documentation/adr/src/rules_v1_basis.py @@ -3,9 +3,9 @@ # ============================================================================ # # A decision's basis names what it rests on. precedent points at another -# record in the corpus; every other source is external. Following precedent -# must reach an external source: a corpus that justifies itself only by citing -# itself can drift anywhere and still look consistent. +# record in the corpus, and must resolve to one; every other source is +# external. Where a chain of precedent leads is for a reader to judge, not +# lint (ADR-311). # # The operator basis and `considered` are an audit trail, not a credential. # Lint checks that `said` and `via` are present. It cannot check that they are @@ -18,18 +18,12 @@ V1_CONCERN_KEYS = ('said', 'resolve', 'answer', 'withdrawn', 'raised') V1_ANSWER_KEYS = ('operator', 'said', 'via', 'paraphrase') V1_CANARY = ('caught', 'missed') -# A precedent is "another accepted decision" (§11). These statuses ground; -# proposed warns; anything else (rejected, abandoned) does not ground. -V1_GROUNDING_STATUSES = ('accepted', 'superseded', 'archived') def v1_basis_sources(ctx) -> tuple: """adr.yaml may rename or extend the sources. precedent is the one internal source; every other declared source is external.""" return tuple(_str_list(ctx.config.get('basis_sources')) or V1_BASIS_SOURCES) -def v1_external_sources(ctx) -> tuple: - return tuple(s for s in v1_basis_sources(ctx) if s != 'precedent') - def _text(value) -> bool: return isinstance(value, str) and value.strip() != '' @@ -49,9 +43,6 @@ def basis_entries(adr, ctx) -> list: out.append((sources[0], entry[sources[0]], entry)) return out -def has_operator_basis(adr, ctx) -> bool: - return any(source == 'operator' for source, _, _ in basis_entries(adr, ctx)) - def _unknown_keys(entry: dict, allowed: tuple) -> list: return [k for k in entry if k not in allowed] @@ -62,20 +53,36 @@ def rule_v1_basis_config(ctx): raw = ctx.config.get('basis_sources') if raw is None: return - sources = _str_list(raw) - if sources is None: + if _str_list(raw) is None: ctx.config_issues.append(Issue("basis_sources: expected a list of names", 'error')) - elif not any(s != 'precedent' for s in sources): - ctx.config_issues.append(Issue("basis_sources: declares no external source, so no chain can leave the corpus", 'error')) # --- one record ---------------------------------------------------------------- +def _check_evidence_record(adr, ctx, where: str, ref: str) -> None: + """A basis `evidence: ADR-N` that names a record must resolve to one, and + under v1 to a kind the decision's basis edge accepts other than decision + (ADR-309 §2): a record cited as evidence is evidence or a spec.""" + target = ctx.by_number.get(norm_ref(ref)[0]) + if target is None: + v1_issue(adr, f"{where}: evidence {ref} resolves to no record") + return + if not is_v1_record(target, ctx): + return + schema = v1_kind_schema(adr, ctx) or {} + allowed = [k for k in (v1_edges(schema).get('basis') or []) if k != 'decision'] + kind = target.frontmatter.get('kind') + if allowed and kind not in allowed: + v1_issue(adr, f"{where}: evidence {ref} is a {kind}; evidence cites a {' or '.join(allowed)} record") + @file_rule(contract=V1) def rule_v1_basis_shape(adr, ctx): if not is_v1_record(adr, ctx) or 'basis' not in adr.frontmatter: return basis = adr.frontmatter.get('basis') - if not isinstance(basis, list) or not basis: + schema = v1_kind_schema(adr, ctx) + if basis == [] and schema is not None and 'basis' in v1_requires(schema): + return # empty where the kind requires a basis: the requires rule reports it + if not isinstance(basis, list): v1_issue(adr, "basis: expected a list of entries, each naming one source") return allowed = v1_basis_sources(ctx) @@ -102,6 +109,8 @@ def rule_v1_basis_shape(adr, ctx): v1_issue(adr, f"{where}: precedent takes one reference, such as ADR-101") elif not _text(value) and not (isinstance(value, int) and not isinstance(value, bool)): v1_issue(adr, f"{where}: {source} needs a reference") + elif source == 'evidence' and re.fullmatch(r'ADR-\d+(\.\d+)?', str(value).strip()): + _check_evidence_record(adr, ctx, where, str(value).strip()) continue # operator: who, the level, what was said and via which channel if not _text(value): @@ -151,10 +160,6 @@ def rule_v1_considered(adr, ctx): v1_issue(adr, f"{where}: covers is a list of probe names") if 'canary' in entry and entry['canary'] not in V1_CANARY: v1_issue(adr, f"{where}: canary is caught or missed") - # ADR-304 §12: a decision the operator started waits for their consideration - if (has_operator_basis(adr, ctx) and str(adr.status or '').lower() == 'accepted' - and not adr.frontmatter.get('considered')): - v1_issue(adr, "accepted with an operator basis but no considered entry; the operator considers what they started") @file_rule(contract=V1) def rule_v1_concern(adr, ctx): @@ -191,135 +196,24 @@ def rule_v1_concern(adr, ctx): else: adr.issues.append(Issue(f"open concern: {_first_line(entry.get('said'))}", 'warning', 'open-concern')) -# --- against the corpus: grounding ------------------------------------------------- -# -# Grounding is a least fixed point, so the verdict does not depend on the order -# of basis entries and each record is visited a bounded number of times: -# -# grounded(r) = r has an external source -# or some precedent of r points at a grounding-status record t -# with grounded(t) -# -# A record with no basis but a decided_by (a spec) is grounded through the -# records that decided it. A non-archived v0 record grounds provisionally: the -# chain passes with a warning until the record is migrated. An archived v0 -# record grounds outright, since archived records never migrate. - -def _is_v0_ground(target, ctx) -> Optional[str]: - """'final' for an archived v0 record, 'provisional' for a live one.""" - if is_v1_record(target, ctx): - return None - return 'final' if is_archived(target.path) else 'provisional' - -def _grounding_map(ctx) -> dict: - """path -> 'final' | 'provisional' for every grounded v1 record.""" - cache = getattr(ctx, '_grounding', None) - if cache is not None: - return cache - records = [a for a in ctx.corpus if is_v1_record(a, ctx)] - external = v1_external_sources(ctx) - grounded = {} - for adr in records: - if any(source in external for source, _, _ in basis_entries(adr, ctx)): - grounded[adr.path] = 'final' - - def level_of(target) -> Optional[str]: - v0 = _is_v0_ground(target, ctx) - if v0: - return v0 - return grounded.get(target.path) - - changed = True - while changed: - changed = False - for adr in records: - best = grounded.get(adr.path) - if best == 'final': - continue - candidates = [] - entries = basis_entries(adr, ctx) - if entries: - for value in _precedent_values(adr, ctx): - target = ctx.by_number.get(norm_ref(value)[0]) - if target is None or str(target.status or '').lower() not in V1_GROUNDING_STATUSES + ('proposed',): - continue - candidates.append(level_of(target)) - elif 'basis' not in adr.frontmatter: - for ref in as_entries(adr.frontmatter.get('decided_by')): - target = ctx.by_number.get(norm_ref(ref)[0]) - if target is not None: - candidates.append(level_of(target)) - new = 'final' if 'final' in candidates else ('provisional' if 'provisional' in candidates else None) - if new and new != best: - grounded[adr.path] = new - changed = True - ctx._grounding = grounded - return grounded +# --- against the corpus ------------------------------------------------------------ def _precedent_values(adr, ctx) -> list: """Precedent references of a usable shape; the shape rule reports others.""" return [value for source, value, _ in basis_entries(adr, ctx) if source == 'precedent' and isinstance(value, (str, int)) and not isinstance(value, bool)] -def _precedent_targets(adr, ctx) -> list: - return [ctx.by_number.get(norm_ref(value)[0]) for value in _precedent_values(adr, ctx)] - -def _cycle_through(adr, ctx) -> Optional[list]: - """The numbers on a precedent cycle that returns to adr, if any.""" - stack = [(adr, [adr])] - visited = set() - while stack: - node, path = stack.pop() - for target in _precedent_targets(node, ctx): - if target is None or not is_v1_record(target, ctx): - continue - if target.path == adr.path: - return [n.number or n.path.name for n in path] - if target.path not in visited: - visited.add(target.path) - stack.append((target, path + [target])) - return None - @corpus_rule(contract=V1) -def rule_v1_basis_chain(adr, ctx): - """Every precedent resolves to an allowed, accepted record, and following - precedent reaches an external source (ADR-304 §11).""" +def rule_v1_precedent(adr, ctx): + """Each precedent resolves to a record of a kind the basis edge accepts.""" if not is_v1_record(adr, ctx): return - entries = basis_entries(adr, ctx) - precedents = [(value, ctx.by_number.get(norm_ref(value)[0])) - for value in _precedent_values(adr, ctx)] - schema = v1_kind_schema(adr, ctx) or {} - allowed_kinds = v1_edges(schema).get('basis') - for value, target in precedents: + allowed_kinds = v1_edges(v1_kind_schema(adr, ctx) or {}).get('basis') + for value in _precedent_values(adr, ctx): + target = ctx.by_number.get(norm_ref(value)[0]) if target is None: v1_issue(adr, f"basis: precedent '{value}' resolves to no known ADR") - continue - status = str(target.status or '').lower() - if is_v1_record(target, ctx): - kind = v1_record_kind(target) - if allowed_kinds is not None and kind not in allowed_kinds: - v1_issue(adr, f"basis: precedent ADR-{target.number} is a {kind}, expected {' or '.join(allowed_kinds)}") - if status == 'proposed': - adr.issues.append(Issue(f"basis: precedent ADR-{target.number} is still proposed", 'warning', 'precedent-proposed')) - elif status not in V1_GROUNDING_STATUSES: - v1_issue(adr, f"basis: precedent ADR-{target.number} is {status or 'without a status'}, so it grounds nothing") - elif 'basis' not in target.frontmatter and not target.frontmatter.get('decided_by'): - v1_issue(adr, f"basis: precedent ADR-{target.number} has neither a basis nor decided_by") - if not entries: - return - level = _grounding_map(ctx).get(adr.path) - if level == 'final': - return - if level == 'provisional': - v0s = sorted({t.number for _, t in precedents if t is not None and _is_v0_ground(t, ctx) == 'provisional'}) - via = f" (ADR-{', ADR-'.join(v0s)})" if v0s else '' - v1_issue(adr, f"basis: grounded only through records still on v0{via}; migrate them to confirm the chain", 'warning') - return - if not precedents: - return # no source at all: the shape rule reports the entries - cycle = _cycle_through(adr, ctx) - if cycle: - v1_issue(adr, f"basis: precedent loops back through ADR-{' -> ADR-'.join(str(n) for n in cycle)} without reaching an external source") - else: - v1_issue(adr, f"basis: following precedent never reaches an external source ({', '.join(v1_external_sources(ctx))})") + elif is_v1_record(target, ctx) and allowed_kinds is not None \ + and v1_record_kind(target) not in allowed_kinds: + v1_issue(adr, f"basis: precedent ADR-{target.number} is a {v1_record_kind(target)}, " + f"expected {' or '.join(allowed_kinds)}") diff --git a/hooks/ways/documentation/adr/src/rules_v1_integrity.py b/hooks/ways/documentation/adr/src/rules_v1_integrity.py index 5f868e81..0216507d 100644 --- a/hooks/ways/documentation/adr/src/rules_v1_integrity.py +++ b/hooks/ways/documentation/adr/src/rules_v1_integrity.py @@ -1,76 +1,22 @@ # ============================================================================ -# adr/v1: legibility, enactment, the vocabulary layers, frozen decisions +# adr/v1: sections, imports, enactment, placeholders, observables # ============================================================================ # -# ADR-304 §5 (enactment), §2 (no word shared across layers), §1 and §12 (a -# decision is frozen once it leaves proposed, and opens with a Summary the -# operator can judge alone). +# Checks of one record's own shape. How records relate over time, and +# whether an accepted one changed, is read from git (ADR-311). -# L3, the derived product state, is fixed by the contract (ADR-304 §2). -V1_DERIVED_STATES = ('active', 'absent', 'present', 'gone', 'living', 'historical') V1_ENACTING_VERBS = ('cut', 'retire') -_STEM_SUFFIXES = ('ations', 'ation', 'ments', 'ment', 'ions', 'ion', 'ings', 'ing', - 'ives', 'ive', 'als', 'al', 'ors', 'or', 'ers', 'er', 'ted', 'ed', - 'ts', 'es', 't', 'e', 's') - -def _stem(word: str) -> str: - """A crude English stem: strip suffixes repeatedly and collapse a doubled - final consonant, enough to catch cut/cutting, retire/retirement, - constrain/constraint and propose/proposal.""" - word = word.lower() - changed = True - while changed: - changed = False - for suffix in _STEM_SUFFIXES: - if word.endswith(suffix) and len(word) - len(suffix) >= 3: - word = word[:-len(suffix)] - changed = True - break - if len(word) >= 4 and word[-1] == word[-2] and word[-1] not in 'aeiou': - word = word[:-1] - return word - -def _collide(a: str, b: str) -> bool: - """Same word, same stem, or one stem is a prefix of the other (4+ letters).""" - if a.lower() == b.lower(): - return True - sa, sb = _stem(a), _stem(b) - short, long_ = sorted((sa, sb), key=len) - return sa == sb or (len(short) >= 4 and long_.startswith(short)) - def _find_section(adr, name: str) -> Optional[str]: + """The heading that is this section: the name alone, or the name followed + by punctuation ("Summary: …", "Summary (draft)"). "Summary Nudge" is a + different section.""" + pattern = re.compile(rf'{re.escape(name)}(\s*[:(].*|\s+[\u2014\u2013-].*)?', re.IGNORECASE) for heading in adr.sections: - if heading.lower() == name.lower() or heading.lower().startswith(name.lower() + ' '): + if pattern.fullmatch(heading.strip()): return heading return None -# --- adr.yaml ------------------------------------------------------------------ - -@config_rule(contract=V1) -def rule_v1_vocabulary_layers(ctx): - """No word or stem appears in two layers (ADR-304 §2), so a later contract - version cannot reintroduce a verb that collides with a state.""" - lifecycle = set() - for schema in v1_kinds(ctx).values(): - if isinstance(schema, dict): - lifecycle.update(v1_lifecycle(schema)) - lifecycle = lifecycle or set(V1_LIFECYCLE) - layers = { - 'lifecycle (L1)': sorted(lifecycle), - 'verbs (L2)': sorted(v1_verbs(ctx)), - 'derived state (L3)': sorted(V1_DERIVED_STATES), - 'basis sources (L4)': sorted(v1_basis_sources(ctx)), - } - names = list(layers) - for i, a in enumerate(names): - for b in names[i + 1:]: - for word_a in layers[a]: - for word_b in layers[b]: - if _collide(word_a, word_b): - ctx.config_issues.append(Issue( - f"'{word_a}' in {a} and '{word_b}' in {b} share a stem; layers share no word", 'error')) - # --- one record ------------------------------------------------------------------ @file_rule(contract=V1) @@ -78,35 +24,31 @@ def rule_v1_required_sections(adr, ctx): schema = v1_kind_schema(adr, ctx) if not is_v1_record(adr, ctx) or schema is None: return + imported = 'imported' in adr.frontmatter for name in _str_list(schema.get('sections')) or []: if _find_section(adr, name) is None: - v1_issue(adr, f"a {v1_record_kind(adr)} record opens with a '## {name}' section") + if imported and name == 'Summary': + # ADR-306 §4: an imported record may gain its Summary later, + # once the imported corpus has been read together. + v1_issue(adr, "imported record has no '## Summary' yet (ADR-306 §4)", 'warning') + else: + v1_issue(adr, f"a {v1_record_kind(adr)} record opens with a '## {name}' section") @file_rule(contract=V1) -def rule_v1_summary_legibility(adr, ctx): - """ADR-304 §12: the Summary carries probes, a mix of points the agent is - confident on and points it is not, each labelled, and an inversion. - Guidance for the writer, so a gap warns.""" - if not is_v1_record(adr, ctx) or v1_record_kind(adr) is None: +def rule_v1_imported(adr, ctx): + """`imported` records where the record came from: {from, format}, and + optionally `status`, the source status as written, and `unmapped`, the + source keys with no v1 field (ADR-306 §1, §4, §7).""" + if not is_v1_record(adr, ctx) or 'imported' not in adr.frontmatter: return - heading = _find_section(adr, 'Summary') - if heading is None or not adr.frontmatter.get('verb'): - return # sections rule reports a missing Summary; specs carry no probes - text = re.sub(r'\s+', ' ', adr.section_text.get(heading, '').lower()) - missing = [] - low = re.search(r'\bnot confident\b|\blow confidence\b', text) - high = re.search(r'(?<!not )(?<!less )\bconfident\b|\bhigh confidence\b', text) - if 'probe' not in text: - missing.append('probes') - else: - if low is None: - missing.append("a probe labelled 'not confident'") - if high is None: - missing.append("a probe labelled 'confident'") - if 'inversion' not in text: - missing.append('an inversion') - if missing: - v1_issue(adr, f"Summary lacks {', '.join(missing)} (ADR-304 §12)", 'warning') + imported = adr.frontmatter.get('imported') + if not isinstance(imported, dict) or not all( + isinstance(imported.get(k), str) and imported.get(k).strip() for k in ('from', 'format')): + v1_issue(adr, "imported: expected {from: <source path>, format: <reader>}") + elif 'unmapped' in imported and not isinstance(imported['unmapped'], dict): + v1_issue(adr, "imported.unmapped: expected a mapping of source fields") + elif isinstance(imported.get('status'), (dict, list)): + v1_issue(adr, "imported.status: expected the source's status as written") @file_rule(contract=V1) def rule_v1_enacted(adr, ctx): @@ -123,86 +65,36 @@ def rule_v1_enacted(adr, ctx): elif not re.fullmatch(r'[0-9a-f]{7,40}', enacted): v1_issue(adr, f"enacted: '{enacted}' is not a commit hash") -# --- frozen decisions, read from git history -------------------------------------- - -def _git(args: list, cwd: Path) -> Optional[str]: - try: - result = subprocess.run(['git', '-c', 'core.quotePath=false', *args], cwd=cwd, - capture_output=True, encoding='utf-8', errors='replace', timeout=10) - except (FileNotFoundError, subprocess.TimeoutExpired, OSError): - return None - return result.stdout if result.returncode == 0 else None - -def _frozen_snapshot(adr) -> Optional[tuple]: - """(frontmatter, body) of the first committed version that was already - adr/v1 and past proposed, following renames. None outside git, for an - untracked file, or when no such version exists. A version that was still - v0 is never the snapshot: migrating an accepted v0 record to v1 adds the - v1 fields, and that is the migration, not an edit (ADR-304 §7).""" - root = get_project_root() - try: - rel = adr.path.resolve().relative_to(root.resolve()) - except ValueError: - return None - # A record freezes where it lands: history is read from the default - # branch when there is one, so a decision still in review on a feature - # branch can be revised. Without a remote default, HEAD's history counts. - # --first-parent keeps to the branch's own line, so a pull request merged - # with a merge commit lands at the merge, not at the first commit on its - # branch; -m lists the merge's files against that parent on older git. - ref = (_git(['rev-parse', '--abbrev-ref', 'origin/HEAD'], root) or '').strip() or 'HEAD' - # --reverse drops pre-rename history under --follow, so read newest first - # and reverse here. -z keeps names with spaces or non-ASCII intact. - log = _git(['log', ref, '--first-parent', '-m', '--follow', '-z', '--format=commit:%H', '--name-only', '--', str(rel)], root) - if not log: - return None - entries, commit = [], None - for token in log.split('\0'): - if token.startswith('commit:'): - commit = token[len('commit:'):] - elif commit and token.lstrip('\n'): - entries.append((commit, token.lstrip('\n'))) - commit = None - for commit, name in reversed(entries): - text = _git(['show', f'{commit}:{name}'], root) - if text is None: - continue - past = parse_text(text, adr.path) - if past.contract == V1 and past.status and str(past.status).lower() != 'proposed': - return past.frontmatter, past.body - return None - -def _same(a, b) -> bool: - """Equal, treating a date and its quoted string as the same value.""" - if isinstance(a, (str, int, float, date)) and isinstance(b, (str, int, float, date)): - return str(a) == str(b) - return a == b - -V1_DEFAULT_MUTABLE = ('status', 'enacted', 'superseded_by', 'considered', 'concern') - @file_rule(contract=V1) -def rule_v1_frozen(adr, ctx): - """Once a decision leaves proposed, only the kind's mutable_after_accept - fields may change, and the body grows only by appending (ADR-304 §1, §4).""" - schema = v1_kind_schema(adr, ctx) - if not is_v1_record(adr, ctx) or schema is None: +def rule_v1_no_placeholders(adr, ctx): + """A record still holding a prompt from `adr new`'s skeleton is unfinished: + a warning while proposed, an error once it has left proposed. The prompts + live in sheet.py, which assembles after this module.""" + if not is_v1_record(adr, ctx): return - # Proposed records are not frozen, and archived ones carry the archive - # banner by design; neither needs the history walk. - if str(adr.status or '').lower() == 'proposed' or is_archived(adr.path): + left = placeholder_lines(adr.body) + if not left: return - mutable = schema.get('mutable_after_accept', list(V1_DEFAULT_MUTABLE)) - if mutable == 'all': + level = 'warning' if str(adr.status or '').lower() == 'proposed' else 'error' + v1_issue(adr, f"{len(left)} placeholder line(s) from `adr new` still to fill, first: {left[0]}", level) + +@file_rule(contract=V1) +def rule_v1_observable(adr, ctx): + """What should be observable when a decision holds (ADR-307): optional, a + list whose entries are plain words or mappings with keys the author chooses. + The check applies to any v1 kind that carries the field.""" + if not is_v1_record(adr, ctx) or 'observable' not in adr.frontmatter: return - mutable = set(_str_list(mutable) or V1_DEFAULT_MUTABLE) - snapshot = _frozen_snapshot(adr) - if snapshot is None: + entries = adr.frontmatter.get('observable') + if entries is None or entries == []: + v1_issue(adr, "observable: empty; remove the key or add an entry") return - then, body_then = snapshot - for key in sorted(set(then) | set(adr.frontmatter)): - if key in mutable: + if not isinstance(entries, list): + v1_issue(adr, "observable: expected a list of entries, each a line of words or a mapping") + return + for i, entry in enumerate(entries, 1): + if isinstance(entry, str) and entry.strip(): + continue + if isinstance(entry, dict) and entry: continue - if not _same(then.get(key), adr.frontmatter.get(key)): - v1_issue(adr, f"'{key}' changed after the decision left proposed; only {', '.join(sorted(mutable)) or 'no fields'} may change") - if not adr.body.rstrip().startswith(body_then.rstrip()): - v1_issue(adr, "body edited after the decision left proposed; a decision grows by appending", 'warning') + v1_issue(adr, f"observable entry {i}: expected a line of words or a non-empty mapping") diff --git a/hooks/ways/documentation/adr/src/sheet.py b/hooks/ways/documentation/adr/src/sheet.py new file mode 100644 index 00000000..c8a5128a --- /dev/null +++ b/hooks/ways/documentation/adr/src/sheet.py @@ -0,0 +1,145 @@ +# --- import sheets and the v1 record writer (ADR-306) ------------------------------- + +SHEET_FORMAT = 'adr-import/v1' + +# Frontmatter keys in the order a v1 record is written. Keys a sheet carries +# that are not listed follow in the sheet's own order. +V1_KEY_ORDER = ('contract', 'kind', 'verb', 'capability', 'targets', 'supersedes', 'amends', + 'extends', 'decided_by', 'superseded_by', 'enacted', 'basis', 'agent', + 'considered', 'concern', 'observable', 'status', 'date', 'deciders', 'related', 'imported') + +V1_SUMMARY_SKELETON = '''## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] +''' + +# The body `adr new` writes, under v0 and v1 alike. +BODY_SKELETON = '''## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] +''' + +# The bracketed prompts in the skeletons. A line still holding one is unfinished. +SKELETON_PROMPTS = tuple(dict.fromkeys(re.findall(r'\[[^\]\n]+\]', V1_SUMMARY_SKELETON + BODY_SKELETON))) + +class _RecordDumper(yaml.SafeDumper): + """Block-style YAML as this corpus writes it: list items indented under + their key, an empty field as `~`, and a multi-line string on one + double-quoted line, so no line of it can read as the `---` fence.""" + def increase_indent(self, flow=False, indentless=False): + return super().increase_indent(flow, False) + +def _represent_str(dumper, value): + style = '"' if '\n' in value else None + return dumper.represent_scalar('tag:yaml.org,2002:str', value, style=style) + +_RecordDumper.add_representer(str, _represent_str) +_RecordDumper.add_representer(type(None), + lambda dumper, _: dumper.represent_scalar('tag:yaml.org,2002:null', '~')) + +def _as_date(value): + """A valid YYYY-MM-DD string as a date, so it is written unquoted. An + invalid one stays a string for lint to report.""" + if isinstance(value, str) and re.fullmatch(r'\d{4}-\d{2}-\d{2}', value): + try: + return date.fromisoformat(value) + except ValueError: + return value + return value + +def empty_record(kind: str, schema: dict, given: dict, defaults: dict) -> dict: + """The frontmatter a record of this kind starts with: every field the kind + requires, from `given` where supplied and empty otherwise. The kind's + schema decides the fields, not its name (ADR-304 §1, §4).""" + record = {'contract': V1, 'kind': kind} + if schema.get('verb') == 'required': + record['verb'] = given.get('verb') + empty = {'basis': [], 'targets': []} + for name in v1_requires(schema): + if name == 'agent': + record['agent'] = {'name': given.get('agent'), 'model': given.get('model')} + else: + record[name] = given.get(name, empty.get(name)) + lifecycle = v1_lifecycle(schema) + default_status = str(defaults.get('status') or '').lower() + record['status'] = default_status if default_status in lifecycle else lifecycle[0] + return record + +def new_sheet(number: int, domain: str, title: str, record: dict, + summary: Optional[str] = None, body: str = '') -> dict: + """A sheet for a record that has no source: what `adr new` applies.""" + return {'sheet': SHEET_FORMAT, 'source': None, + 'target': {'number': number, 'domain': domain, 'title': title}, + 'record': record, 'summary': summary, 'body': body} + +def format_number(number) -> str: + """ADR-007, ADR-101.1: three digits, and a sub-part kept as written.""" + base, _, part = str(number).partition('.') + return f"{int(base):03d}" + (f".{part}" if part else '') + +def _ordered(record: dict) -> dict: + ordered = {k: record[k] for k in V1_KEY_ORDER if k in record} + ordered.update({k: v for k, v in record.items() if k not in ordered}) + return ordered + +def render_record(sheet: dict) -> str: + """The v1 record a sheet describes. It writes only what the sheet holds: + the frontmatter, the title, the Summary if the sheet has one, then the + body verbatim. The frontmatter is read back before it is returned, and a + mismatch raises: the round trip is checked, not assumed (ADR-306 §7).""" + target = sheet['target'] + fields = _ordered(sheet['record']) + if 'date' in fields: + fields['date'] = _as_date(fields['date']) + front = yaml.dump(fields, Dumper=_RecordDumper, sort_keys=False, allow_unicode=True, + default_flow_style=False, width=1000) + if yaml.safe_load(front) != fields: + raise ValueError(f"ADR-{target['number']}: frontmatter does not read back as written") + title = ' '.join(str(target['title']).split()) + text = f"---\n{front}---\n\n# ADR-{format_number(target['number'])}: {title}\n" + summary = sheet.get('summary') + if summary: + summary = summary if summary.startswith('## Summary') else f"## Summary\n\n{summary}" + text += '\n' + summary.rstrip('\n') + '\n' + if sheet.get('body'): + text += '\n' + sheet['body'] + return text + +def placeholder_lines(text: str) -> list: + """Lines outside fenced code that still hold a skeleton prompt.""" + found, fence = [], None + for line in text.splitlines(): + marker = re.match(r'\s*(```|~~~)', line) + if marker: + fence = None if fence == marker.group(1) else (fence or marker.group(1)) + continue + if fence is None and any(prompt in line for prompt in SKELETON_PROMPTS): + found.append(line.strip()) + return found diff --git a/hooks/ways/documentation/linting/doclint.py b/hooks/ways/documentation/linting/doclint.py index cbcb5006..7033657d 100755 --- a/hooks/ways/documentation/linting/doclint.py +++ b/hooks/ways/documentation/linting/doclint.py @@ -499,7 +499,8 @@ def effective(node, severity): errors += 1 else: # --no-inventory: doclint is a lint and runs no shell commands - # from adr.yaml. `adr cite` run directly checks the inventories. + # from adr.yaml. A vendored adr tool from before ADR-311 runs + # surface inventories without it; a current one ignores it. print("\nCitations (adr cite --no-inventory):", flush=True) result = subprocess.run([sys.executable, str(adr_tool), "cite", "--no-inventory"], cwd=REPO, capture_output=True, text=True) diff --git a/hooks/ways/meta/knowledge/authoring/authoring.md b/hooks/ways/meta/knowledge/authoring/authoring.md index 837c0205..9a2d80f3 100644 --- a/hooks/ways/meta/knowledge/authoring/authoring.md +++ b/hooks/ways/meta/knowledge/authoring/authoring.md @@ -103,7 +103,7 @@ Common choices (numeric ↔ preset, matching the built-in defaults): | Procedural event handlers (fires often relative to session) | `0.05` | `frequent` | | Disclose once per session | `1.0` | `once` | -Numeric values between these presets are fine — for example, the 14 ways migrated from ADR-127's 1M-Opus hack sit at `refire: 0.2` (between `normal` and `rare`), deliberately pinned to today's model. +Numeric values between these presets are fine — for example, the 14 ways migrated from the PR #70 1M-Opus narrow-tune (ADR-126) sit at `refire: 0.2` (between `normal` and `rare`), deliberately pinned to today's model. Missing `refire:` on a fire-bearing way means the way fires once and never re-discloses — valid but uncommon, and `ways lint` warns on it. Check files and `trigger: attend` handlers are exempt (checks ride on parent way firing; attend handlers are signal-triggered). diff --git a/hooks/ways/softwaredev/delivery/github/macro.sh b/hooks/ways/softwaredev/delivery/github/macro.sh index 35e5ae6f..7b53647c 100755 --- a/hooks/ways/softwaredev/delivery/github/macro.sh +++ b/hooks/ways/softwaredev/delivery/github/macro.sh @@ -281,8 +281,8 @@ fi # # Precedence approximates Claude Code's resolution order (project-local > # project > user-global, highest wins). Each sub-key resolves INDEPENDENTLY, on -# the documented merge law that objects deep-merge key by key (ADR-147; mirrored -# in tools/ways-cli/src/cmd/settings/compile.rs) — so a project file setting +# the documented merge law that objects deep-merge key by key (recorded in +# ADR-147, retired by ADR-169 with its compile.rs mirror) — so a project file setting # adr-cite-ignore # .commit must not hide a user file setting .sessionUrl. NOTE: that law is # documented for settings merging generally; its application to `attribution` # ACROSS SCOPES is inferred, not verified. If Claude Code instead replaces the diff --git a/hooks/ways/strip-session-link-pre.sh b/hooks/ways/strip-session-link-pre.sh index e4b73baf..2fd06c44 100755 --- a/hooks/ways/strip-session-link-pre.sh +++ b/hooks/ways/strip-session-link-pre.sh @@ -1,7 +1,7 @@ #!/usr/bin/env bash # PreToolUse (Bash): deny commit / PR / issue authoring commands that carry a # Claude-Session transcript link in trailer/footer position. See ADR-167 -# (supersedes ADR-162). +# (supersedes ADR-162). # adr-cite-ignore # # The session link resolves to the FULL session transcript; publishing it in a # commit trailer or PR/issue body is a thin wall in front of accidental secret @@ -10,13 +10,13 @@ # This hook is a BACKSTOP, not the primary control. `attribution.sessionUrl: false` # suppresses the link at the source — the model is never instructed to emit it, so # when that works this hook never fires. The key shipped in v2.1.183 and is -# projected to ~/.claude/settings.json from the settings fragment store (ADR-147 / -# ADR-163). A controlled experiment on 2026-07-16 (same repo, same v2.1.212, only +# projected to ~/.claude/settings.json from the dotfiles-side fragment store (ADR-163, +# ADR-169). A controlled experiment on 2026-07-16 (same repo, same v2.1.212, only # the key differing) confirms it governs the trailer. # # It is retained because that setting is undocumented (upstream #69614), defaults # to ON, and is scope-qualified upstream ("web and Remote Control sessions") — a -# host that never receives the fragment leaks by default. ADR-162 previously +# host that never receives the fragment leaks by default. ADR-162 previously # adr-cite-ignore # claimed no setting could govern the link at all, citing #18253; that was the # Co-Authored-By/footer bug, not the session link. See ADR-167 for the correction. # @@ -59,7 +59,7 @@ printf '%s' "$CMD" \ | grep -qiE '^[[:space:]]*(Claude-Session:[[:space:]]*)?https?://claude\.ai/code/session_[A-Za-z0-9_-]+[^A-Za-z0-9]*$' \ || exit 0 -REASON='Blocked (ADR-162): this command publishes a Claude-Session transcript link (a "Claude-Session:" trailer or a bare claude.ai/code/session_ URL on its own line) to git or GitHub. That link resolves to the full session transcript and must not be published. Remove the trailer/footer line from the commit message or PR/issue body and re-run. (A session URL mentioned inline in prose is allowed; only trailer/footer lines are blocked.)' +REASON='Blocked (ADR-167): this command publishes a Claude-Session transcript link (a "Claude-Session:" trailer or a bare claude.ai/code/session_ URL on its own line) to git or GitHub. That link resolves to the full session transcript and must not be published. Remove the trailer/footer line from the commit message or PR/issue body and re-run. (A session URL mentioned inline in prose is allowed; only trailer/footer lines are blocked.)' jq -cn --arg reason "$REASON" '{ hookSpecificOutput: { diff --git a/skills/adr/SKILL.md b/skills/adr/SKILL.md index d2717f5f..d387a0ec 100644 --- a/skills/adr/SKILL.md +++ b/skills/adr/SKILL.md @@ -1,12 +1,14 @@ --- name: adr -description: Manage Architecture Decision Records using the project's ADR CLI tool. Use when the user wants to create, list, view, lint, or index ADRs, or when working with docs/architecture/ files. Triggers on "create an ADR", "new ADR", "list ADRs", "lint ADRs", "what ADRs exist", "ADR domains". +description: Manage Agent Decision Records using the project's ADR CLI tool. Use when the user wants to create, list, view, lint, or index ADRs, or when working with docs/architecture/ files. Triggers on "create an ADR", "new ADR", "list ADRs", "lint ADRs", "what ADRs exist", "ADR domains". allowed-tools: Bash, Read, Grep, Glob --- # ADR Management -Operate ADRs through the `docs/scripts/adr` CLI tool. Never create ADR files manually. If the tool isn't present in the project yet, vendor it first — see [Vendoring the tool into a project](#vendoring-the-tool-into-a-project). +ADR now means Agent Decision Record; the `ADR-N` citation form is unchanged. Operate ADRs through the `docs/scripts/adr` CLI tool. Never create ADR files manually. If the tool isn't present in the project yet, vendor it first — see [Vendoring the tool into a project](#vendoring-the-tool-into-a-project). + +A project declares its record contract in `docs/architecture/adr.yaml`. With `contract: adr/v1` it uses declared kinds, verbs, capabilities and lifecycle commands (below). Without it, the project is on the legacy `adr/v0` contract: Draft/Proposed/Accepted/Superseded/Deprecated status, and Context/Decision/Consequences/Alternatives Considered sections. `adr config` shows which contract a project is on. Both contracts share the same tool and the same `ADR-N` number space. ## Commands @@ -15,47 +17,131 @@ Operate ADRs through the `docs/scripts/adr` CLI tool. Never create ADR files man docs/scripts/adr domains # Show domain number series and ranges docs/scripts/adr list --group # List active ADRs grouped by domain docs/scripts/adr list --archived # Archived only; --all for both -docs/scripts/adr list --domain system # Filter to one domain -docs/scripts/adr list --status Accepted # Filter by status +docs/scripts/adr list --domain <domain> # Filter to one domain +docs/scripts/adr list --status Accepted # Filter by status (v0 title case or v1 lower case; compared case-insensitively) docs/scripts/adr view <number> # View an ADR (accepts 14, 014, ADR-014) # Create -docs/scripts/adr new <domain> "Title" # Create from template in correct subdirectory +docs/scripts/adr new <domain> "Title" [--kind K] [--verb V] [--capability C] [--agent A --model M] + # --kind: decision (default), spec, evidence (adr/v1 only) + # --verb: add, cut, change, retire, constrain (decisions only) + # --capability: from the adr.yaml vocabulary; --agent/--model: who is writing it # Maintain docs/scripts/adr lint # Check all ADRs (archive included) for issues docs/scripts/adr lint --check # Exit 1 on errors (CI mode) +docs/scripts/adr cite [--check] # Check ADR-N citations in code against the records docs/scripts/adr index -y # Regenerate INDEX.md from the active set +# Lifecycle (adr/v1) — status changes only through these, never by setting the field +docs/scripts/adr accept <n> [--dry-run] # Accept a proposed record +docs/scripts/adr reject <n> --reason "..." [--dry-run] # Considered and declined +docs/scripts/adr abandon <n> --reason "..." [--dry-run] # Dropped before a decision + # Archive (ADR-303) — not deletion: the file stays tracked, linted, linkable docs/scripts/adr archive <n> --reason "why" [--superseded-by ADR-N[,ADR-M#sec]] [--status S] [--dry-run] +# Domains (ADR-310) — a record's number is permanent identity; a domain's range only allocates new numbers +docs/scripts/adr domain add <name> --range A-B --folder F [--label L] [--description D] +docs/scripts/adr domain rename <old> <new> [--folder F] [--dry-run] # moves the folder, rewrites paths +docs/scripts/adr domain move <n> <domain> [--dry-run] # adr/v1: moves the file, rewrites paths +docs/scripts/adr domain move --plan moves.yaml [--dry-run] # [{record, domain}, ...] at once + +# Rename a record's title or filename slug (distinct from `domain rename`) +docs/scripts/adr rename <n> ["New Title"] [--slug SLUG] + +# Query by frontmatter (adr/v1); a list field matches when it lists the value +docs/scripts/adr list --kind evidence --capability attend --verb change # also --field KEY[=VALUE] +docs/scripts/adr list --group-by capability # a listed record appears in each group +docs/scripts/adr list --json # number, title, path, status, frontmatter + +# Edit records (adr/v1): only the touched field's lines change; each lints the record after +docs/scripts/adr consider <n> --said "..." --via "..." [--covers PROBE...] [--canary caught|missed] +docs/scripts/adr set <n> key=value key+=item key-=item [--force] [--dry-run] + # refuses status (use the lifecycle commands above) without --force +docs/scripts/adr supersede <old> --by <new> [--amends SECTION] [--dry-run] # writes both sides +docs/scripts/adr enact <n> <commit> # accepted cut or retire only + +# Import (ADR-306) — convert v0 or foreign records into adr/v1 through editable import sheets +docs/scripts/adr import scan <paths...> [--force] # writes a sheet per record under .import/ +docs/scripts/adr import apply [sheets...] [--partial] [--force] [--dry-run] # writes finished sheets as adr/v1 records +# A source's text between its frontmatter and its title moves into the body, under the title + # Config docs/scripts/adr config # Show current adr.yaml configuration ``` +Under adr/v1, a record's folder decides its area (ADR-310), and its number is +its permanent identity: a move keeps the number, and `ADR-N` citations stay +valid. An area's range only allocates numbers for new records. Under adr/v0, +domain and number range are the same thing, unchanged from before. + ## Workflow 1. **Check domains first**: `docs/scripts/adr domains` to see available domains and number ranges -2. **Create**: `docs/scripts/adr new <domain> "Decision Title"` — assigns next number, uses YAML frontmatter template -3. **Edit**: Fill in Context, Decision, Consequences, Alternatives sections +2. **Create**: `docs/scripts/adr new <domain> "Decision Title"` (adr/v1: add `--kind`, `--verb`, `--capability`, `--agent`/`--model`) — assigns the next number, seeds frontmatter for the project's contract +3. **Fill in the body** matching the record's kind — a v1 decision opens with `## Summary` and its `basis`; a v0 record uses Context, Decision, Consequences, Alternatives Considered 4. **Lint**: `docs/scripts/adr lint` before committing -5. **Index**: `docs/scripts/adr index -y` after adding or changing ADRs +5. **For a v1 decision with an operator basis, record their answer**: `docs/scripts/adr consider <n> --said "..." --via "..."` before `docs/scripts/adr accept <n>` +6. **Index**: `docs/scripts/adr index -y` after adding or changing ADRs ## Configuration -Each project defines its domain structure in `docs/architecture/adr.yaml`: -- **domains**: Name, number range, description, folder for each domain -- **statuses**: Valid status values (Draft, Proposed, Accepted, Superseded, Deprecated) -- **defaults**: Default deciders and initial status for new ADRs -- **legacy**: Number range for pre-domain ADRs - -Always run `docs/scripts/adr domains` to discover the project's actual configuration rather than assuming domains. +Each project defines its structure in `docs/architecture/adr.yaml`, always +discovered with `docs/scripts/adr domains` / `docs/scripts/adr config` rather +than assumed: +- **domains**: name, number range, description, folder for each domain (adr/v1: the folder is the record's area; the range only allocates new numbers, ADR-310) +- **statuses**: valid v0 status values (Draft, Proposed, Accepted, Superseded, Deprecated); adr/v1 has a fixed status set (proposed, accepted, rejected, abandoned, superseded, archived) +- **defaults**: default deciders and initial status for new ADRs +- **legacy**: number range for pre-domain ADRs +- **contract** (adr/v1 only): `adr/v1`, opting the project into the fields below; absent means adr/v0 +- **kinds** (adr/v1 only): `decision`, `spec`, `evidence`, each declaring its required fields, sections, and edges +- **capabilities** (adr/v1 only): the closed vocabulary a record's `capability` field draws from +- **basis_sources** (adr/v1 only): what a decision may ground itself in — operator, evidence, standard, upstream, precedent +- **repository** (optional): this repository as URLs name it (host/owner/repo), so `adr domain` rewrites links into it when a record moves; the origin remote otherwise ## ADR Format -The tool generates ADRs with YAML frontmatter: +The record's shape follows the project's contract, so ask `adr config` rather +than assume one. `adr lint` is the authority on what a given record needs; it +reads the contract and reports what is missing rather than a fixed checklist. +**adr/v1** (`docs/scripts/adr new` writes this shape): +```markdown +--- +contract: adr/v1 +kind: decision +verb: add +capability: adr +basis: + - operator: <name> + level: guided + said: "<their words, quoted>" + via: <where they said it> +agent: + name: <agent> + model: <model> +status: proposed +date: <YYYY-MM-DD> +--- + +# ADR-NNN: Decision Title + +## Summary +- **Decided:** ... +- **Trades away:** ... +- **One-way?** ... +- **Probes:** ... +- **Inversion:** ... + +## Context +## Decision +## Consequences +## Alternatives Considered +``` +A `spec` record carries `capability` and no `verb`, and stays mutable in place. An `evidence` record carries no `verb`, is corrected by appending once accepted, and is what a decision's `basis` cites for findings, surveys, audits or explorations (ADR-309). + +**adr/v0** (the legacy contract, unchanged): ```markdown --- status: Draft @@ -93,12 +179,21 @@ chmod +x docs/scripts/adr ``` Then edit `docs/architecture/adr.yaml` for the project's domains and ranges, and -validate: `docs/scripts/adr domains && docs/scripts/adr lint`. +validate: `docs/scripts/adr domains && docs/scripts/adr lint`. The template +declares `contract: adr/v1` with the decision, spec and evidence kinds and a +placeholder capability: replace it with the project's capabilities. To stay on +adr/v0, delete the `contract` line and the v1 blocks under it. For a full repo scaffold (ADRs + GitHub config + CODEOWNERS + project ways), run `/project-init` instead — it vendors this tool as one step of a larger setup. The **docs** skill is the catalog half and shares this `adr.yaml`. +A project with existing decision records in another shape (a flat directory, +inline metadata, a different tool) converts them with `docs/scripts/adr import +scan <paths>` then `docs/scripts/adr import apply` (ADR-306) rather than by +hand-editing frontmatter — the import sheets are editable and re-lintable +before anything is written as an adr/v1 record. + ## Updating a vendored copy The tool carries a `TOOL_VERSION` stamp (`docs/scripts/adr --version`), and the @@ -129,7 +224,8 @@ downgrade the project. - **Always use the CLI** — never create `ADR-*.md` files by hand - **Run `domains` first** when working in an unfamiliar project — domain names and ranges vary -- **Status lives in frontmatter** — edit the YAML `status:` field, not inline text +- **Status changes through the tool** — adr/v0: edit the YAML `status:` field; adr/v1: `adr accept`/`reject`/`abandon`/`supersede`/`archive`, which `adr set` refuses without `--force` +- **Correct by appending** — an accepted record is corrected by appending, a change in what the project does is a new decision that names what it replaces, and the operator is quoted in their own words. The tool checks shape and references, not these; review catches them, and git keeps every earlier version (ADR-311) - **Regenerate index** after any ADR changes with `docs/scripts/adr index -y` ## Not for diff --git a/skills/docs/SKILL.md b/skills/docs/SKILL.md index b5fdeea5..4bb67839 100644 --- a/skills/docs/SKILL.md +++ b/skills/docs/SKILL.md @@ -111,7 +111,7 @@ scaffold, run `/project-init`. To decline the catalog for a project: ## Not for -- Architecture Decision Records — that's the **adr** skill (the decisions half of the same catalog). +- Agent Decision Records — that's the **adr** skill (the decisions half of the same catalog). - Hand-authoring `id`/`domain`/`mode` frontmatter — the CLI computes it. - Prose that isn't joining the catalog — untagged docs need no tool and are never linted. diff --git a/skills/ways-localize/SKILL.md b/skills/ways-localize/SKILL.md index e0ad245f..a0b51f30 100644 --- a/skills/ways-localize/SKILL.md +++ b/skills/ways-localize/SKILL.md @@ -11,7 +11,7 @@ Turns an English-only ways install into a localized one (ADR-139). English is th root. This skill is the operator-facing orchestrator for the lifecycle in `docs/explanation/localization/` (scenario `01.011.E`) — interview, consent, translate, tune, switch. The mechanics live in the design note -`docs/design-notes/adopter-localization-lifecycle-and-tuning.md`; don't restate them. +`docs/architecture/ways/ADR-183-single-language-localization-tuning-the-english-anchor-as-a-peer.md`; don't restate them. ```bash ROOT="${CLAUDE_CONFIG_DIR:-$HOME/.claude}" diff --git a/skills/ways-tests/SKILL.md b/skills/ways-tests/SKILL.md index e852bf60..5155afb1 100644 --- a/skills/ways-tests/SKILL.md +++ b/skills/ways-tests/SKILL.md @@ -126,7 +126,7 @@ floor. At *scan* time (not here), when an ancestor has fired this session a child's effective semantic bar is lowered from `τ_s` to `(τ_s × parent_threshold_multiplier).max(parent_boost_floor)` — by default `max(0.5×0.8, 0.30) = 0.40`, in probability space (the multiplier boosts, the floor -caps) — the progressive-disclosure mechanism (ADR-104/105/126). For multilingual +caps) — the progressive-disclosure mechanism (ADR-105/123/126). For multilingual stubs, `ways tune` reports locale fidelity/discrimination (see the `knowledge/optimization/tuning` way). diff --git a/tests/adr-conversion-check.sh b/tests/adr-conversion-check.sh new file mode 100644 index 00000000..1b1334ab --- /dev/null +++ b/tests/adr-conversion-check.sh @@ -0,0 +1,162 @@ +#!/usr/bin/env bash +# The conversion's guarantee, held against the live tree (ADR-306 §4, §7). +# +# Every record that was v0 in the frozen snapshot (tests/fixtures/adr/v0-corpus) +# and is v1 now must still carry what the conversion promised to keep: +# - its raw source status, under imported.status +# - its date and deciders, unchanged +# - every other source frontmatter key, as a v1 field or under imported.unmapped +# - its original body, which still opens the text after the H1, once an opening +# Summary (which ADR-306 §4 lets anyone add) is set aside; text may be appended +# Records still v0, or archived, are skipped. A record whose body was edited on +# purpose is listed in BODY_EDITED, with the reason. +# +# A record keeps its number when `adr domain move` or `rename` changes its +# folder (ADR-306 §6), so a snapshot record is found by number wherever it +# sits now. Those commands rewrite paths to what moved, and ADR-309 turned +# design notes into records. The check builds its own map of those moves, +# from the snapshot, the live tree and the NOTES table below, and applies it +# to the snapshot body before comparing. It shares no code with the tool, so +# a fault in the tool's rewrite shows here as a failing record. +set -uo pipefail + +REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)" + +python3 - "$REPO_ROOT" <<'PY' +import posixpath, re, sys, yaml +from pathlib import Path + +root = Path(sys.argv[1]) +snapshot = root / 'tests/fixtures/adr/v0-corpus/docs/architecture' +live = root / 'docs/architecture' + +# Bodies edited on purpose after conversion: number -> reason. +BODY_EDITED = { + '302': 'stray tool-call text removed from the end of the body', +} + +# Design notes a snapshot body names that became records (ADR-309): the +# note's path -> the record's number. +NOTES = { + 'docs/design-notes/attend-envelope-fields.md': '401', + 'docs/design-notes/attend-messaging-disclosure-reheat.md': '400', + 'docs/design-notes/cognitive-loop-and-awareness-layer.md': '600', + 'docs/design-notes/cypress-survey.md': '603', + 'docs/design-notes/settings-json-merge-spec-and-peer-writer-contract.md': '500', + 'docs/design-notes/tool-use-channel-lookbehind-chunk-matching.md': '191', +} + +# Links that named no file in the snapshot and were repaired after +# conversion: the record's number -> (the folder the links named it in, +# the reason). The links now point to where the record is. +REPAIRED = { + '112': ('docs/architecture/system', + 'the session-ledger record was already archived; links named it as a sibling'), +} + +TITLE = re.compile(r'^# ADR-[0-9.]+:.*$', re.M) +SUMMARY = re.compile(r'\A\s*## Summary[^\n]*\n.*?(?=^## |\Z)', re.M | re.S) + +def split(text): + if not text.startswith('---\n'): + return None, None + end = text.index('\n---\n', 4) + front = yaml.safe_load(text[4:end]) or {} + rest = text[end + 5:] + m = TITLE.search(rest) + return front, (rest[m.end():] if m else None) + +def number(path): + m = re.match(r'ADR-([0-9.]+)', path.name) + return m.group(1).lstrip('0') or '0' if m else None + +live_by_number = {number(p): p for p in live.rglob('ADR-*.md') if 'archive' not in p.parts} +# An archived record is still a place a link can point to. +anywhere = {**{number(p): p for p in live.rglob('ADR-*.md') if 'archive' in p.parts}, **live_by_number} + +def repo_path(path): + return path.relative_to(root).as_posix() + +# Each file a snapshot body may name by path -> where that file is now: every +# snapshot record, found by its number, each design note in NOTES, and each +# broken link in REPAIRED. +moved = {} +for src in snapshot.rglob('ADR-*.md'): + dest = anywhere.get(number(src)) + if dest is not None: + moved['docs/architecture/' + src.relative_to(snapshot).as_posix()] = repo_path(dest) +for note, n in NOTES.items(): + moved[note] = repo_path(anywhere[n]) +for n, (folder, _) in REPAIRED.items(): + moved[f'{folder}/{anywhere[n].name}'] = repo_path(anywhere[n]) + +PATH = re.compile(r'[\w./-]+') + +def relocate(body, was, now): + """body, written at `was`, with each path to a moved file rewritten to + that file's place now: a path from the repo root stays one, and a path + relative to the record is written relative to where the record is now.""" + was_dir, now_dir = posixpath.dirname(was), posixpath.dirname(now) + def one(m): + token = m.group(0) + path = token.rstrip('.') + if path in moved: + return moved[path] + token[len(path):] + resolved = posixpath.normpath(posixpath.join(was_dir, path)) + target = moved.get(resolved) + if target is None or posixpath.normpath(posixpath.join(now_dir, path)) == target: + return token + new = posixpath.relpath(target, now_dir) + if path.startswith('./') and not new.startswith('.'): + new = './' + new + return new + token[len(path):] + return PATH.sub(one, body) + +checked = failures = 0 +for src in sorted(snapshot.rglob('ADR-*.md')): + if 'archive' in src.parts: + continue + was = 'docs/architecture/' + src.relative_to(snapshot).as_posix() + before, body_before = split(src.read_text()) + if before is None or before.get('contract'): + continue + n = number(src) + dest = live_by_number.get(n) + if dest is None: + print(f"FAIL ADR-{n}: no live record") + failures += 1 + continue + after, body_after = split(dest.read_text()) + if not after or after.get('contract') != 'adr/v1': + continue + checked += 1 + problems = [] + imported = after.get('imported') or {} + unmapped = imported.get('unmapped') or {} + if 'status' in before and str(imported.get('status')) != str(before['status']): + problems.append(f"imported.status is {imported.get('status')!r}, source had {before['status']!r}") + for key in ('date', 'deciders'): + if key in before and str(after.get(key)) != str(before[key]): + problems.append(f"{key} changed") + for key in before: + if key in ('status',): + continue + if key not in after and key not in unmapped: + problems.append(f"source key '{key}' is neither a field nor under imported.unmapped") + if body_before is None or body_after is None: + problems.append("no H1 on one side") + elif n not in BODY_EDITED: + opened = SUMMARY.sub('', body_after, count=1).lstrip('\n') + expected = relocate(body_before, was, repo_path(dest)) + if not opened.startswith(expected.lstrip('\n')): + problems.append("the original body no longer opens the record") + for p in problems: + print(f"FAIL ADR-{n}: {p}") + failures += bool(problems) +print(f"converted records checked: {checked}, failing: {failures}") +sys.exit(1 if failures else 0) +PY +status=$? +echo "" +[[ $status -eq 0 ]] && echo "=== ADR Conversion Check: passed ===" || echo "=== ADR Conversion Check: FAILED ===" +exit $status diff --git a/tests/adr-golden-test.sh b/tests/adr-golden-test.sh index 7c1aa220..4e8e9888 100755 --- a/tests/adr-golden-test.sh +++ b/tests/adr-golden-test.sh @@ -13,7 +13,8 @@ # number for the multiple-match branch of view # v1/ an adr/v1 corpus (ADR-304): every kind and verb used correctly, # and one record still on v0 -# v1-defects/ one broken record per v1 rule, and a malformed adr.yaml +# v1-defects/ one broken record per v1 rule, a malformed adr.yaml, and +# keys adr/v1 no longer reads # v1-empty/ adr/v1 declared with no kinds and no capabilities # # Usage: @@ -76,7 +77,7 @@ fresh() { # A run that crosses midnight sees two dates, so both the date the run # started on and the current one become <TODAY>. normalize() { - sed -e "s#$ADR_TOOL#<ADR_TOOL>#g" -e "s#$WORK/repo#<ROOT>#g" \ + sed -e "s#$ADR_TOOL#<ADR_TOOL>#g" -e "s#$WORK/repo#<ROOT>#g" -e "s#$WORK#<WORK>#g" \ -e "s#$TODAY#<TODAY>#g" -e "s#$(date +%Y-%m-%d)#<TODAY>#g" } @@ -229,109 +230,82 @@ capture v1-lint-check lint --check capture v1-list list capture v1-view-spec view 102 capture v1-lint-precedent-relative lint docs/architecture/system/ADR-109-precedent-chain.md -# Baseline capabilities (ADR-305). export joins the vocabulary as a baseline -# capability adopted on 2025-05-10: it needs no add decision. -baseline_fresh() { - fresh v1 - edit docs/architecture/adr.yaml "s.replace('surfaces:', 'baseline:\\n adopted: 2025-05-10\\n capabilities: [export]\\n\\nsurfaces:', 1).replace(' search: Query over the store\\n', ' search: Query over the store\\n export: Export from the store\\n')" -} -# export_change NUMBER SLUG DATE [EXTRA-FRONTMATTER-LINE] -export_change() { - (cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: change\ncapability: export\n%sstatus: proposed\ndate: %s\ndeciders: [developer]\nagent: {name: Claude, model: m}\nbasis:\n - evidence: export measured\n---\n\n# ADR-%s: %s\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n' "${4:+$4\n}" "$3" "$1" "$2" > "docs/architecture/system/ADR-$1-$2.md") -} -baseline_fresh -capture v1-lint-baseline lint -# Malformed baseline config: a bare list, then an unknown name and a bad date. -edit docs/architecture/adr.yaml "s.replace('baseline:\\n adopted: 2025-05-10\\n capabilities: [export]', 'baseline: [export]')" -capture v1-lint-baseline-list lint -edit docs/architecture/adr.yaml "s.replace('baseline: [export]', 'baseline:\\n adopted: May 2025\\n capabilities: [export, exprt]')" -capture v1-lint-baseline-unknown lint -# A change before adoption stands on the baseline, and so does the first -# change after it. ADR-118 is numbered lower but dated later than ADR-119, -# so it is the second post-adoption change and must name ADR-119. -baseline_fresh -export_change 117 pre-adoption-export 2025-05-09 -export_change 119 stream-exports 2025-05-20 -export_change 118 compress-exports 2025-05-22 -capture v1-lint-baseline-change lint -edit docs/architecture/system/ADR-118-compress-exports.md "s.replace('capability: export\n', 'capability: export\nsupersedes: [ADR-119]\n')" -capture v1-lint-baseline-change-names-prior lint -# A rejected post-adoption change is no prior. -edit docs/architecture/system/ADR-118-compress-exports.md "s.replace('supersedes: [ADR-119]\n', '')" -edit docs/architecture/system/ADR-119-stream-exports.md "s.replace('status: proposed', 'status: rejected')" -capture v1-lint-baseline-change-rejected-prior lint +# adr new under adr/v1 (ADR-306 §3): a decision with its fields given, a spec, +# a bare decision whose empty fields lint names, and an unknown kind. fresh v1 +capture v1-new-decision new system "Stream exports" --verb change --capability ingest --agent Claude --model fixture-model +keep v1-new-decision-file.md docs/architecture/system/ADR-115-stream-exports.md +capture v1-new-decision-lint lint docs/architecture/system/ADR-115-stream-exports.md +capture v1-new-spec new system "Export format" --kind spec --capability ingest +keep v1-new-spec-file.md docs/architecture/system/ADR-116-export-format.md +capture v1-new-spec-lint lint docs/architecture/system/ADR-116-export-format.md +capture v1-new-bare new system "Bare decision" +keep v1-new-bare-file.md docs/architecture/system/ADR-117-bare-decision.md +capture v1-new-bare-lint lint docs/architecture/system/ADR-117-bare-decision.md +# accept checks the record's own rules and refuses on an error (ADR-311 §3). +capture v1-accept-own-errors accept 117 +capture v1-new-unknown-kind new system "Policy thing" --kind policy +worktree v1-new-unknown-kind-status.txt +capture v1-new-refused new system "Spec with a verb" --kind spec --verb add --agent Claude --capability nosuch +# A value with a line that reads as the frontmatter fence is written on one +# line and reads back intact. +capture v1-new-multiline new system "Multiline model" --verb add --capability ingest --agent Claude --model "$(printf 'a\n---\nb')" +keep v1-new-multiline-file.md docs/architecture/system/ADR-118-multiline-model.md +capture v1-new-multiline-lint lint docs/architecture/system/ADR-118-multiline-model.md +# A placeholder left in a record that has left proposed is an error. +edit docs/architecture/system/ADR-115-stream-exports.md "s.replace('status: proposed', 'status: accepted')" +capture v1-new-placeholder-accepted lint docs/architecture/system/ADR-115-stream-exports.md + +# The kind's schema, not its name, decides a new record's fields: a custom +# kind that takes a verb, requires targets and has a Summary section. fresh v1 -capture v1-cite cite -# The cut on search, enacted: citations of search records now fail. -(cd "$WORK/repo" && sed -i.bak 's/^verb: cut$/verb: cut\nenacted: abcdef1/' docs/architecture/system/ADR-111-cut-search.md \ - && rm docs/architecture/system/ADR-111-cut-search.md.bak) -capture v1-cite-enacted cite --check -capture v1-cite-no-inventory cite --no-inventory - -# A cut undone by a later accepted add: search is present again. -fresh v1 -(cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: add\ncapability: search\nstatus: accepted\ndate: 2025-05-20\ndeciders: [developer]\nagent: {name: Claude, model: m}\nbasis:\n - evidence: demand returned\n---\n\n# ADR-115: Add search back\n' > docs/architecture/system/ADR-115-add-search-back.md) -capture v1-cite-readded cite --no-inventory - -# A retire before enactment, with one target misspelled; then an inventory -# command that fails. -fresh v1 -edit docs/architecture/system/ADR-105-retire-legacy-ingest.md "s.replace('enacted: 3f9c2a1\n', '').replace('cli:ingest-legacy', 'cli:ingest-legacy, cli:ingest-legcy')" -capture v1-cite-retire-pending cite -edit docs/architecture/adr.yaml "s[:s.index(' cli:')] + ' cli: { inventory: \"exit 3\" }' + s[s.index(chr(10), s.index(' cli:')):]" -capture v1-cite-inventory-fails cite +edit docs/architecture/adr.yaml "s.replace(' spec:\n', ' policy:\n verb: required\n requires: [capability, targets]\n sections: [Summary]\n spec:\n', 1)" +capture v1-new-custom-kind new system "Retention policy" --kind Policy --verb constrain --capability ingest +keep v1-new-custom-kind-file.md docs/architecture/system/ADR-115-retention-policy.md +# A v1 contract with no kinds: new says why it cannot write. +fresh v1-empty +capture v1-empty-new new system "Anything" -# A frozen decision edited after acceptance: a changed capability (an error), -# a body edited mid-text (a warning), and a mutable field (allowed). +# Observables (ADR-307): loose shapes pass, malformed ones fail. fresh v1 -edit docs/architecture/system/ADR-101-ingest.md "s.replace('capability: ingest\n', 'capability: search\n').replace('The decision.\n', 'The decision, rewritten.\n').replace('date: 2025-05-02\n', 'date: 2025-05-02\nconsidered: [{operator: developer, said: ok, via: PR 2}]\n')" -capture v1-frozen-lint lint docs/architecture/system/ADR-101-ingest.md - -# The freeze follows a rename: renamed, committed, then edited. +edit docs/architecture/system/ADR-101-ingest.md "s.replace('date: 2025-05-02\n', 'date: 2025-05-02\nobservable:\n - ingest of a 10 MB file finishes under a second\n - see: the queue drains\n run: make drain-check\n - url: https://example.com/dashboard\n', 1)" +capture v1-observable-added lint docs/architecture/system/ADR-101-ingest.md +edit docs/architecture/system/ADR-101-ingest.md "s.replace(' - url: https://example.com/dashboard\n', ' - \"\"\n - {}\n - 42\n', 1)" +capture v1-observable-malformed lint docs/architecture/system/ADR-101-ingest.md +edit docs/architecture/system/ADR-101-ingest.md "__import__('re').sub(r'observable:\n(?: .*\n)+', 'observable: soon\n', s, count=1)" +capture v1-observable-not-list lint docs/architecture/system/ADR-101-ingest.md +edit docs/architecture/system/ADR-101-ingest.md "s.replace('observable: soon\n', 'observable: []\n', 1)" +capture v1-observable-empty lint docs/architecture/system/ADR-101-ingest.md +# A change may list the capabilities it alters (ADR-308). fresh v1 -(cd "$WORK/repo" && git mv docs/architecture/system/ADR-101-ingest.md docs/architecture/system/ADR-101-ingest-renamed.md) -commit_all rename -edit docs/architecture/system/ADR-101-ingest-renamed.md "s.replace('capability: ingest\n', 'capability: search\n')" -capture v1-frozen-renamed lint docs/architecture/system/ADR-101-ingest-renamed.md - -# Proposed, then accepted in a later commit: accepting is not an edit. +(cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: change\ncapability: [ingest, search]\nsupersedes: [ADR-103]\nstatus: proposed\ndate: 2025-05-21\ndeciders: [developer]\nagent: {name: Claude, model: m}\nbasis:\n - evidence: both paths share one queue\n---\n\n# ADR-116: Shared queue for ingest and search\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n' > docs/architecture/system/ADR-116-shared-queue.md) +capture v1-change-list lint docs/architecture/system/ADR-116-shared-queue.md +# Past three capabilities, a listed change draws a warning (ADR-308 §3). +edit docs/architecture/adr.yaml "s.replace(' search: Query over the store\n', ' search: Query over the store\n export: Export from the store\n audit: Audit trail\n', 1)" +edit docs/architecture/system/ADR-116-shared-queue.md "s.replace('capability: [ingest, search]', 'capability: [ingest, search, export, audit]')" +capture v1-change-list-long lint docs/architecture/system/ADR-116-shared-queue.md + +# A heading that only starts with "Summary" is another section, not the Summary. fresh v1 -edit docs/architecture/system/ADR-113-operator-proposed.md "s.replace('status: proposed\n', 'status: accepted\nconsidered: [{operator: developer, said: \"yes\", via: PR 13}]\n')" -commit_all accept -capture v1-frozen-accepted lint docs/architecture/system/ADR-113-operator-proposed.md +edit docs/architecture/system/ADR-108-ingest-over-v0.md "s.replace('## Summary\n', '## Summary Nudge\n', 1)" +capture v1-summary-prefix-heading lint docs/architecture/system/ADR-108-ingest-over-v0.md +edit docs/architecture/system/ADR-108-ingest-over-v0.md "s.replace('## Summary Nudge\n', '## Summary: the short version\n', 1)" +capture v1-summary-colon-heading lint docs/architecture/system/ADR-108-ingest-over-v0.md -# An accepted v0 record migrated to v1 adds the v1 fields: the migration, -# not an edit of a frozen decision (ADR-304 §7). +# An evidence kind (ADR-309): a basis that cites a record as evidence must +# reach an evidence or spec record. fresh v1 -edit docs/architecture/system/ADR-110-old-v0-record.md "s.replace('status: Accepted\n', 'contract: adr/v1\nkind: decision\nverb: add\ncapability: ingest\nstatus: accepted\nagent: {name: Claude, model: fixture-model}\nbasis:\n - evidence: migrated from v0\n').replace('# ADR-110: An unmigrated v0 record\n', '# ADR-110: An unmigrated v0 record\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n')" -capture v1-frozen-migrated lint docs/architecture/system/ADR-110-old-v0-record.md +edit docs/architecture/adr.yaml "s.replace('basis: [decision, spec] }', 'basis: [decision, spec, evidence] }').replace(' edges: { supersedes: spec, decided_by: decision }\n', ' edges: { supersedes: spec, decided_by: decision }\n evidence:\n verb: forbidden\n requires: [capability]\n edges: { supersedes: evidence }\n', 1)" +(cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: evidence\ncapability: ingest\nstatus: accepted\ndate: 2025-05-20\ndeciders: [developer]\n---\n\n# ADR-117: Ingest throughput survey\n\nMeasured 40 MB/s on the reference host.\n' > docs/architecture/system/ADR-117-ingest-throughput-survey.md) +(cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: change\ncapability: ingest\nsupersedes: [ADR-103]\nstatus: proposed\ndate: 2025-05-21\ndeciders: [developer]\nagent: {name: Claude, model: m}\nbasis:\n - evidence: ADR-117\n - evidence: ADR-101\n - evidence: ADR-999\n---\n\n# ADR-118: Raise the ingest batch size\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n' > docs/architecture/system/ADR-118-raise-the-ingest-batch-size.md) +capture v1-evidence-kind lint docs/architecture/system/ADR-117-ingest-throughput-survey.md docs/architecture/system/ADR-118-raise-the-ingest-batch-size.md -# A record accepted on a review branch and revised there freezes where it -# merges, not at the branch's first accepted commit; editing it after the -# merge still fails. -fresh v1 -(cd "$WORK/repo" && git switch -q -c review) -edit docs/architecture/system/ADR-113-operator-proposed.md "s.replace('status: proposed\n', 'status: accepted\nconsidered: [{operator: developer, said: \"yes\", via: PR 13}]\n')" -commit_all accept -edit docs/architecture/system/ADR-113-operator-proposed.md "s.replace('basis:\n', 'basis:\n - evidence: a review fix\n', 1)" -commit_all "review fix" -(cd "$WORK/repo" && git switch -q - && git merge -q --no-ff -m "merge review" review) -capture v1-frozen-merged-branch lint docs/architecture/system/ADR-113-operator-proposed.md -edit docs/architecture/system/ADR-113-operator-proposed.md "s.replace(' - evidence: a review fix\n', '')" -capture v1-frozen-merged-branch-edited lint docs/architecture/system/ADR-113-operator-proposed.md - -# A non-UTF-8 blob in a record's history does not stop lint. fresh v1 -printf 'binary \377\376 junk\n' > "$WORK/repo/docs/architecture/system/ADR-102-ingest-spec.md" -commit_all "non-utf8" -cp "$FIXTURES/v1/docs/architecture/system/ADR-102-ingest-spec.md" "$WORK/repo/docs/architecture/system/ADR-102-ingest-spec.md" -commit_all restore -capture v1-frozen-nonutf8 lint docs/architecture/system/ADR-102-ingest-spec.md +capture v1-cite cite +capture v1-cite-no-inventory cite --no-inventory # Lifecycle commands (ADR-304 §2, §11, §12) fresh v1 -capture v1-accept-refused accept 113 capture v1-accept-not-proposed accept 102 capture v1-accept-v0 accept 110 capture v1-accept-dry-run accept 106 --dry-run @@ -339,6 +313,8 @@ worktree v1-accept-dry-run-status.txt capture v1-accept-concern accept 114 keep v1-accept-concern-file.md docs/architecture/system/ADR-114-open-concern.md capture v1-accept-then-lint lint --check docs/architecture/system/ADR-114-open-concern.md +# An operator-started record is accepted without a considered entry. +capture v1-accept-no-considered accept 113 fresh v1 capture v1-reject reject 107 --reason "Superseded by the batching work before it landed" @@ -347,13 +323,12 @@ capture v1-abandon-no-reason abandon 108 --reason " " capture v1-reject-then-lint lint --check docs/architecture/system/ADR-107-ingest-notes.md capture v1-abandon-no-flag abandon 108 -# A record another record rests on: rejecting it is refused, naming the -# dependent, and the dependent cannot be accepted while its precedent is -# still proposed. +# A record another record rests on: rejecting it and accepting the dependent +# each check only the record acted on, not the corpus (ADR-311 §3). fresh v1 (cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: change\ncapability: ingest\nsupersedes: [ADR-103]\nstatus: proposed\ndate: 2025-05-16\ndeciders: [developer]\nagent: {name: Claude, model: m}\nbasis:\n - precedent: ADR-114\n---\n\n# ADR-116: Rests on ADR-114\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n' > docs/architecture/system/ADR-116-rests-on-114.md) capture v1-reject-precedent reject 114 --reason "Not needed" -capture v1-accept-pending-precedent accept 116 +capture v1-accept-dependent accept 116 # Line endings survive a status rewrite. fresh v1 @@ -361,6 +336,449 @@ edit docs/architecture/system/ADR-114-open-concern.md "s.replace(chr(10), chr(13 capture v1-accept-crlf accept 114 (cd "$WORK/repo" && python3 -c "import sys; d=open(sys.argv[1],'rb').read(); print('crlf kept' if b'\\r\\n' in d and d.count(b'\\n')==d.count(b'\\r\\n') else 'crlf lost', '|', [l for l in d.split(b'\\r\\n') if l.startswith(b'status:')])" docs/architecture/system/ADR-114-open-concern.md) > "$ACTUAL/v1-accept-crlf-file.txt" +# --- adr import (ADR-306) --------------------------------------------------------- + +# keep_sheets PREFIX — keep every sheet scan wrote, one golden each +keep_sheets() { + local sheet + for sheet in "$WORK/repo/docs/architecture/.import/"ADR-*.yaml; do + [[ -e "$sheet" ]] && keep "$1-$(basename "$sheet")" "docs/architecture/.import/$(basename "$sheet")" + done +} + +# Scan the v0 corpus: one sheet per record, the archive left out, a record +# without frontmatter skipped. The sources stay untouched and the sheet +# directory ignores itself, so git sees no change. +fresh corpus +capture import-scan-corpus import scan docs/architecture +worktree import-scan-corpus-status.txt +keep import-scan-corpus-gitignore docs/architecture/.import/.gitignore +# A directory of records that scan cannot read: each file it passes over is +# named, with the reason, and counted as skipped. +mkdir -p "$WORK/repo/notes/decisions" +printf '# 1. Use Postgres\n\nStatus: Accepted\n' > "$WORK/repo/notes/decisions/0001-use-postgres.md" +printf '# ADR-2: Inline\n\nStatus: Accepted\n' > "$WORK/repo/notes/decisions/ADR-002-inline.md" +printf 'Decisions live here.\n' > "$WORK/repo/notes/decisions/README.md" +printf -- '---\nstatus: Accepted\ndate: 2026-01-15\ndeciders: [a]\nrelated: []\n---\n\n# ADR-003: No slug\n\n## Context\nx\n' > "$WORK/repo/notes/decisions/ADR-003.md" +capture import-scan-flat import scan notes/decisions +# A broad folder: each Markdown file passed over is named, a hidden folder +# is named and not entered, other files are counted by extension, and the +# tool's own files are left out only at the top of the folder scanned. +mkdir -p "$WORK/repo/notes/.drafts" "$WORK/repo/notes/archive" "$WORK/repo/notes/sub" "$WORK/repo/notes/img" +printf 'x\n' > "$WORK/repo/notes/.drafts/ADR-003-draft.md" +printf 'x\n' > "$WORK/repo/notes/.notes.md" +printf 'x\n' > "$WORK/repo/notes/INDEX.md" +printf 'x\n' > "$WORK/repo/notes/sub/INDEX.md" +printf 'k: v\n' > "$WORK/repo/notes/sub/adr.yaml" +printf 'x\n' > "$WORK/repo/notes/archive/README.md" +printf 'x\n' > "$WORK/repo/notes/archive/old.txt" +printf 'x\n' > "$WORK/repo/notes/Makefile" +for i in 1 2 3; do printf 'png' > "$WORK/repo/notes/img/p$i.png"; done +capture import-scan-broad import scan notes +keep_sheets import-scan-corpus + +# In the v1 fixture: a v1 record scanned as itself, a v0 record, a v0 record +# with a preamble above its H1, an unmapped key and no Summary, and a +# Deprecated v0 record with no successor. +import_fresh() { + fresh v1 + (cd "$WORK/repo" && printf -- '---\nstatus: Accepted\ndate: 2025-04-01\ndeciders: [developer]\nrevised: 2025-04-02\n---\n\n> Moved here from the ops wiki.\n\n# ADR-115: Nightly ingest window\n\n## Context\n\nIngest ran nightly.\n\n## Decision\n\nIngest runs in a nightly window.\n' > docs/architecture/system/ADR-115-nightly-ingest-window.md \ + && printf -- '---\nstatus: Deprecated\ndate: 2025-04-05\ndeciders: [developer]\n---\n\n# ADR-116: Hourly ingest\n\n## Context\n\nIngest ran hourly.\n' > docs/architecture/system/ADR-116-hourly-ingest.md) + commit_all "v0 records to import" +} +IMPORTED="docs/architecture/system/ADR-101-ingest.md docs/architecture/system/ADR-110-old-v0-record.md docs/architecture/system/ADR-115-nightly-ingest-window.md docs/architecture/system/ADR-116-hourly-ingest.md" +import_fresh +# shellcheck disable=SC2086 +capture import-scan-v1 import scan $IMPORTED +keep_sheets import-scan-v1 +# Unedited sheets: the v1 record applies unchanged, the v0 ones are skipped. +capture import-apply-open import apply +worktree import-apply-open-status.txt +# --partial writes them anyway, and lint names what is missing. The +# preamble and the unmapped key are carried; the Deprecated record's note is +# one lint cannot find again, so --partial does not write past it. +capture import-apply-partial import apply --partial +worktree import-apply-partial-status.txt +keep import-apply-partial-110.md docs/architecture/system/ADR-110-old-v0-record.md +keep import-apply-partial-115.md docs/architecture/system/ADR-115-nightly-ingest-window.md +# An imported record with no Summary warns rather than fails (ADR-306 §4). +capture import-lint-no-summary lint docs/architecture/system/ADR-115-nightly-ingest-window.md +# Committed with empty fields, then completed: the filled record lints clean. +commit_all "partial import" +edit docs/architecture/system/ADR-115-nightly-ingest-window.md "s.replace('verb: ~', 'verb: add').replace('capability: ~', 'capability: ingest').replace('basis: []', 'basis:\n - evidence: ingest logs').replace('name: ~', 'name: Claude').replace('# ADR-115: Nightly ingest window\n', '# ADR-115: Nightly ingest window\n\n## Summary\n\n- **Probes:** *Confident:* a. *Not confident:* b.\n- **Inversion:** c.\n')" +capture import-lint-completed lint docs/architecture/system/ADR-115-nightly-ingest-window.md + +# A completed sheet applies without --partial. +import_fresh +capture import-scan-complete import scan docs/architecture/system/ADR-110-old-v0-record.md +edit docs/architecture/.import/ADR-110.yaml "(lambda t: t[:t.index('todo:')] + 'todo: []\n' + t[t.index('candidates:'):])(s.replace('verb: ~', 'verb: add').replace('capability: ~', 'capability: ingest').replace('basis: []', 'basis:\n - evidence: migrated from v0').replace('name: ~', 'name: Claude'))" +capture import-apply-complete import apply docs/architecture/.import/ADR-110.yaml +keep import-apply-complete-file.md docs/architecture/system/ADR-110-old-v0-record.md +capture import-apply-complete-lint lint docs/architecture/system/ADR-110-old-v0-record.md +# An observable on the sheet is written in its place in the key order. +import_fresh +capture import-scan-observable import scan docs/architecture/system/ADR-110-old-v0-record.md +edit docs/architecture/.import/ADR-110.yaml "(lambda t: t[:t.index('todo:')] + 'todo: []\n' + t[t.index('candidates:'):])(s.replace('verb: ~', 'verb: add').replace('capability: ~', 'capability: ingest').replace('basis: []', 'basis:\n - evidence: migrated from v0').replace('name: ~', 'name: Claude').replace(' status: accepted', ' observable:\n - the record lints clean\n status: accepted', 1))" +capture import-apply-observable import apply docs/architecture/.import/ADR-110.yaml +keep import-apply-observable-file.md docs/architecture/system/ADR-110-old-v0-record.md + +# A dry run lints what it would write inside the corpus, prints each issue, +# then restores every file and keeps every sheet. +import_fresh +capture import-scan-dryrun import scan docs/architecture/system/ADR-110-old-v0-record.md +edit docs/architecture/.import/ADR-110.yaml "(lambda t: t[:t.index('todo:')] + 'todo: []\n' + t[t.index('candidates:'):])(s.replace('verb: ~', 'verb: change').replace('capability: ~', 'capability: ingest').replace('basis: []', 'basis:\n - evidence: migrated from v0').replace('name: ~', 'name: Claude'))" +capture import-apply-dryrun import apply --dry-run +worktree import-apply-dryrun-status.txt + +# A source edited after the scan is refused, and the sheet is kept. +import_fresh +capture import-scan-changed import scan docs/architecture/system/ADR-110-old-v0-record.md +edit docs/architecture/system/ADR-110-old-v0-record.md "s.replace('They follow.', 'They follow, mostly.')" +capture import-apply-changed import apply --partial +worktree import-apply-changed-status.txt +# A rescan keeps a sheet that differs from a fresh scan; --force replaces it. +capture import-rescan-kept import scan docs/architecture/system/ADR-110-old-v0-record.md +capture import-rescan-forced import scan --force docs/architecture/system/ADR-110-old-v0-record.md +# The source now has uncommitted changes: refused unless --force. +capture import-apply-uncommitted import apply --partial docs/architecture/.import/ADR-110.yaml +capture import-apply-uncommitted-forced import apply --partial --force docs/architecture/.import/ADR-110.yaml + +# Sheets apply cannot use: each is refused and the batch goes on. +# Among them: in-tree records renumbered or moved to another domain, and +# the Deprecated ADR-116, which --partial does not write past. +import_fresh +capture import-scan-bad import scan $IMPORTED +edit docs/architecture/.import/ADR-115.yaml "s.replace(\" number: '115'\", ' number: 117')" +(cd "$WORK/repo/docs/architecture/.import" \ + && sed "s/^ number: '110'$/ number: 101.10/" ADR-110.yaml > float.yaml \ + && sed "s#^ number: '110'\$# number: '110.1/../../x'#" ADR-110.yaml > escape.yaml \ + && sed "s/^ number: '110'$/ number: 0156/" ADR-110.yaml > octal.yaml \ + && sed "s/^ number: '110'$/ number: 0x6E/" ADR-110.yaml > hex.yaml \ + && sed 's/^ domain: system$/ domain: legacy/' ADR-110.yaml > domain.yaml \ + && sed 's/^ domain: system$/ domain: storage/' ADR-110.yaml > nodomain.yaml \ + && python3 -c "s=open('ADR-110.yaml').read(); open('listbody.yaml','w').write(s[:s.index('body: |')] + 'body: [not, text]\n')" \ + && python3 -c "s=open('ADR-110.yaml').read(); open('nobody.yaml','w').write(s[:s.index('body: |')] + 'body: \"\"\n')" \ + && printf 'sheet: adr-import/v1\n bad: [\n' > broken.yaml) +capture import-apply-bad import apply --partial docs/architecture/.import/float.yaml docs/architecture/.import/escape.yaml docs/architecture/.import/octal.yaml docs/architecture/.import/hex.yaml docs/architecture/.import/listbody.yaml docs/architecture/.import/nobody.yaml docs/architecture/.import/broken.yaml docs/architecture/.import/domain.yaml docs/architecture/.import/nodomain.yaml docs/architecture/.import/ADR-115.yaml docs/architecture/.import/ADR-116.yaml docs/architecture/.import/ADR-101.yaml +worktree import-apply-bad-status.txt + +# A source outside the repo, numbered outside its target domain's range, is +# refused; renumbered into the range it is written, named by its file name. +fresh v1 +rm -rf "$WORK/outside" && mkdir -p "$WORK/outside" +printf -- '---\nstatus: Accepted\ndate: 2025-04-01\ndeciders: [developer]\n---\n\n# ADR-042: Foreign record\n\n## Context\n\nFrom elsewhere.\n' > "$WORK/outside/ADR-042-foreign-record.md" +capture import-scan-foreign import scan "$WORK/outside/ADR-042-foreign-record.md" +edit docs/architecture/.import/ADR-042.yaml "s.replace(' domain: legacy', ' domain: system')" +capture import-apply-out-of-range import apply --partial docs/architecture/.import/ADR-042.yaml +edit docs/architecture/.import/ADR-042.yaml "s.replace(\" number: '42'\", \" number: '150'\")" +# A dry run that would write a new file removes it again. +capture import-apply-foreign-dryrun import apply --partial --dry-run docs/architecture/.import/ADR-042.yaml +worktree import-apply-foreign-dryrun-status.txt +capture import-apply-foreign import apply --partial docs/architecture/.import/ADR-042.yaml +worktree import-apply-foreign-status.txt +keep import-apply-foreign-file.md docs/architecture/system/ADR-150-foreign-record.md + +# A status with no v1 mapping, or none at all, holds the sheet back even +# under --partial: lint could not tell what it was once written. +fresh v1 +(cd "$WORK/repo" && printf -- '---\nstatus: WIP pending review\ndate: 2025-04-01\ndeciders: [developer]\n---\n\n# ADR-117: Work in progress\n\n## Context\n\nText.\n' > docs/architecture/system/ADR-117-work-in-progress.md \ + && printf -- '---\ndate: 2025-04-01\ndeciders: [developer]\n---\n\n# ADR-118: No status\n\n## Context\n\nText.\n' > docs/architecture/system/ADR-118-no-status.md) +commit_all "records with odd statuses" +capture import-scan-status import scan docs/architecture/system/ADR-117-work-in-progress.md docs/architecture/system/ADR-118-no-status.md +keep_sheets import-scan-status +capture import-apply-status import apply --partial + +# A byte order mark is named as one. +fresh v1 +printf '\357\273\277---\nstatus: Accepted\ndate: 2025-04-01\ndeciders: [developer]\n---\n\n# ADR-117: With a BOM\n' > "$WORK/repo/docs/architecture/system/ADR-117-with-a-bom.md" +capture import-scan-bom import scan docs/architecture/system/ADR-117-with-a-bom.md + +# --- adr domain (ADR-306 §6) ------------------------------------------------------ +# A record's number is its identity and never changes. Under adr/v1 its folder +# decides its domain; a domain's range only allocates new numbers. + +# add: the entry is written into adr.yaml's domains block as text, so the +# file's comments and layout stay. Overlapping ranges, a taken folder or name +# and a backwards range are refused. +fresh v1 +capture domain-add domain add docs --range 300-399 --folder documentation --label Documentation --description "Guides and references" +keep domain-add-config.yaml docs/architecture/adr.yaml +capture domain-add-overlap domain add ops --range 150-250 --folder operations +capture domain-add-refused domain add docs --range 400-300 --folder system +worktree domain-add-refused-status.txt +fresh corpus +capture domain-add-v0 domain add api --range 400-499 --folder api --label "API surface" +keep domain-add-v0-config.yaml docs/architecture/adr.yaml +capture domain-add-legacy-overlap domain add early --range 50-60 --folder early + +# rename: the domain key, its folder moved with git mv, every path into the +# folder rewritten, and a catalog page's domain: key. Links between records +# keep their shape. +rename_fresh() { + fresh v1 + mkdir -p "$WORK/repo/docs/guide" + printf -- '---\ndomain: system\n---\n\n# Hooks guide\n\nSee [no network](../architecture/system/ADR-104-no-network-in-hooks.md) and ADR-104.\nRecords live in docs/architecture/system/ and [the folder](../architecture/system/).\n' > "$WORK/repo/docs/guide/hooks.md" + edit docs/architecture/system/ADR-109-precedent-chain.md "s + '\n## 3. Notes\n\nSee [the constraint](../system/ADR-104-no-network-in-hooks.md).\n'" + edit src/search.py "s + '# ADR-104 hooks stay offline: docs/architecture/system/ADR-104-no-network-in-hooks.md\n'" + commit_all "references" +} +rename_fresh +capture domain-rename-dry domain rename system platform --dry-run +worktree domain-rename-dry-status.txt +capture domain-rename domain rename system platform +worktree domain-rename-status.txt +keep domain-rename-config.yaml docs/architecture/adr.yaml +keep domain-rename-guide.md docs/guide/hooks.md +keep domain-rename-109.md docs/architecture/platform/ADR-109-precedent-chain.md +keep domain-rename-search.py src/search.py +capture domain-rename-lint lint +capture domain-rename-list list --group +rename_fresh +capture domain-rename-unknown domain rename nosuch other +capture domain-rename-refused domain rename system system --folder archive +capture domain-rename-same domain rename system system +capture domain-rename-folder domain rename system core --folder kernel +worktree domain-rename-folder-status.txt + +# move: the file goes to the target domain's folder with its number and slug; +# path references are rewritten and ADR-N citations are not. ADR-104 then +# sits outside the docs range, which is valid under adr/v1. +move_fresh() { + fresh v1 + (cd "$WORK/repo" && "$ADR_TOOL" domain add docs --range 300-399 --folder documentation --description "Guides" > /dev/null \ + && "$ADR_TOOL" index -y > /dev/null) + mkdir -p "$WORK/repo/docs/guide" + printf -- '---\ndomain: system\n---\n\n# Hooks guide\n\nSee [no network](../architecture/system/ADR-104-no-network-in-hooks.md) and ADR-104#1.\n' > "$WORK/repo/docs/guide/hooks.md" + edit docs/architecture/system/ADR-109-precedent-chain.md "s + '\n## 3. Notes\n\nADR-104 applies; see [the constraint](ADR-104-no-network-in-hooks.md) and [ingest](./ADR-101-ingest.md).\n'" + edit docs/architecture/system/ADR-111-cut-search.md "s + '\nSee [ADR-104](ADR-104-no-network-in-hooks.md).\n'" + # The record that moves links to a sibling that stays, by bare file name. + edit docs/architecture/system/ADR-104-no-network-in-hooks.md "s + '\nSee [the chain](ADR-109-precedent-chain.md).\n'" + edit src/search.py "s + '# ADR-104 hooks stay offline: docs/architecture/system/ADR-104-no-network-in-hooks.md\n# ADR-1040 and ADR-104.1 are other numbers.\n'" + commit_all "references" +} +move_fresh +capture domain-move-dry domain move 104 docs --dry-run +worktree domain-move-dry-status.txt +capture domain-move domain move ADR-104 docs +worktree domain-move-status.txt +keep domain-move-104.md docs/architecture/documentation/ADR-104-no-network-in-hooks.md +keep domain-move-109.md docs/architecture/system/ADR-109-precedent-chain.md +keep domain-move-guide.md docs/guide/hooks.md +keep domain-move-search.py src/search.py +keep domain-move-index.md docs/architecture/INDEX.md +capture domain-move-lint lint docs/architecture/documentation/ADR-104-no-network-in-hooks.md docs/architecture/system/ADR-109-precedent-chain.md docs/architecture/system/ADR-111-cut-search.md +capture domain-move-list list --group +capture domain-move-cite cite --no-inventory src/search.py docs/guide +capture domain-move-scan import scan docs/architecture/documentation/ADR-104-no-network-in-hooks.md +capture domain-move-again domain move 104 docs +# Numbers stay allocated by range, across every record wherever it sits. +capture domain-move-new-docs new docs "Style guide" +capture domain-move-new-system new system "Queue limits" +capture domain-move-refused domain move 999 nowhere + +# A plan applies several moves at once: ADR-104 and ADR-111 move together, so +# ADR-111's link to ADR-104 stays a sibling link. +move_fresh +printf -- '- {record: 104, domain: docs}\n- {record: ADR-111, domain: docs}\n' > "$WORK/plan.yaml" +capture domain-move-plan domain move --plan "$WORK/plan.yaml" +worktree domain-move-plan-status.txt +keep domain-move-plan-111.md docs/architecture/documentation/ADR-111-cut-search.md +keep domain-move-plan-109.md docs/architecture/system/ADR-109-precedent-chain.md +move_fresh +printf -- '- {record: 104, domain: docs, number: 300}\n' > "$WORK/plan.yaml" +capture domain-move-plan-number domain move --plan "$WORK/plan.yaml" +printf -- '- {record: 104, domain: docs}\n- {record: ADR-104, domain: docs}\n' > "$WORK/plan.yaml" +capture domain-move-plan-twice domain move --plan "$WORK/plan.yaml" +worktree domain-move-plan-refused-status.txt + +# Under adr/v0 the number decides the domain, so a move is refused. +fresh corpus +capture domain-move-v0 domain move 104 docs + +# --- what a relocation rewrites (#603) ------------------------------------------------ +# A rewrite reaches a path that resolves to what moved: a relative link, a +# path from the repo root, or a URL into this repository's origin. Another +# repository's URL, prose, and a code constant naming a folder stay. The dry +# run prints each line it would change. +# +# ADR-104's basis evidence and body name ADR-101's path, and the rewrite +# reaches both. +safety_fresh() { + fresh v1 + (cd "$WORK/repo" && git remote add origin git@github.com:fixture/corpus.git) + printf 'TEMPLATE_DIR = "architecture/system"\nRECORDS = "docs/architecture/system"\n' > "$WORK/repo/src/app.py" + printf '# Changes\n\n- Another repo: https://github.com/someone/else/tree/main/docs/architecture/system/ADR-101-ingest.md\n- This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md\n- The vendor layout uses architecture/system and a system/ADR-101-ingest.md file.\n- See docs/architecture/system/ADR-101-ingest.md.\n' > "$WORK/repo/CHANGELOG.md" + edit docs/architecture/system/ADR-104-no-network-in-hooks.md "s.replace(' - evidence: fixture measurement\n', ' - evidence: \"the survey at docs/architecture/system/ADR-101-ingest.md\"\n') + '\nSee [ingest](ADR-101-ingest.md), [the guide](https://example.com/v1/guide), [the notes](../api/README.md), and/or the spec.\n'" + (cd "$WORK/repo" && "$ADR_TOOL" domain add docs --range 300-399 --folder documentation --description "Guides" > /dev/null) + (cd "$WORK/repo" && git add -A && git commit -q --amend -m fixture) +} +safety_fresh +capture relocate-move-dry domain move 101 docs --dry-run +capture relocate-rename-dry domain rename system platform --dry-run +capture relocate-rename domain rename system platform +keep relocate-rename-changelog.md CHANGELOG.md +keep relocate-rename-app.py src/app.py +keep relocate-rename-104.md docs/architecture/platform/ADR-104-no-network-in-hooks.md +capture relocate-rename-lint lint + +# A move of a record that another record cites by path. +safety_fresh +capture relocate-move domain move 101 docs +keep relocate-move-104.md docs/architecture/system/ADR-104-no-network-in-hooks.md +capture relocate-move-lint lint docs/architecture/system/ADR-104-no-network-in-hooks.md +S="docs/architecture/system" + +# ADR-104 cites ADR-101 by a sibling link, in its basis and in its body. +# Moving 104, then 101, then 104 again, each committed, leaves the link +# reaching ADR-101 from a third folder. +chain_fresh() { + fresh v1 + (cd "$WORK/repo" && "$ADR_TOOL" domain add docs --range 300-399 --folder documentation --description Guides > /dev/null \ + && "$ADR_TOOL" domain add ops --range 400-499 --folder operations --description Operations > /dev/null) +} +chain_moves() { + (cd "$WORK/repo" && "$ADR_TOOL" domain move 104 docs > /dev/null); commit_all "move 104" + (cd "$WORK/repo" && "$ADR_TOOL" domain move 101 ops > /dev/null); commit_all "move 101" + (cd "$WORK/repo" && "$ADR_TOOL" domain move 104 ops > /dev/null); commit_all "move 104 again" +} +chain_fresh +edit $S/ADR-104-no-network-in-hooks.md "s.replace(' - evidence: fixture measurement', ' - evidence: \"see [ingest](ADR-101-ingest.md)\"')" +(cd "$WORK/repo" && git add -A && git commit -q --amend -m fixture) +chain_moves +keep relocate-chain-basis-104.md docs/architecture/operations/ADR-104-no-network-in-hooks.md +chain_fresh +edit $S/ADR-104-no-network-in-hooks.md "s + '\nSee [ingest](ADR-101-ingest.md).\n'" +(cd "$WORK/repo" && git add -A && git commit -q --amend -m fixture) +chain_moves +keep relocate-chain-body-104.md docs/architecture/operations/ADR-104-no-network-in-hooks.md + +# --- which repository a URL names, and what a rewrite leaves (#603) ------------- +# adr.yaml's repository: names this repository, so the rewrite reads the +# same URLs as this repository's with or without an origin remote. +fresh v1 +edit docs/architecture/adr.yaml "s + 'repository: github.com/fixture/corpus\n'" +edit $S/ADR-104-no-network-in-hooks.md "s + '\nSee https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md.\n'" +(cd "$WORK/repo" && git add -A && git commit -q --amend -m fixture) +(cd "$WORK/repo" && "$ADR_TOOL" domain add docs --range 300-399 --folder documentation --description Guides > /dev/null) +capture relocate-repository-move domain move 101 docs +keep relocate-repository-104.md $S/ADR-104-no-network-in-hooks.md +edit docs/architecture/adr.yaml "s.replace('repository: github.com/fixture/corpus', 'repository: [3]')" +capture relocate-repository-bad lint $S/ADR-104-no-network-in-hooks.md + +# A URL at a commit or a tag is a permalink and stays; one at a branch, a +# branch with a slash, or on GitHub's raw host is rewritten. A path written +# with backslashes, or inside a fenced code block (indented in a list too), stays. +fresh v1 +(cd "$WORK/repo" && git remote add origin git@github.com:fixture/corpus.git && git tag v1.0 && git branch feature/x) +printf -- '# Notes\n\n- https://github.com/fixture/corpus/blob/0123456789abcdef0123456789abcdef01234567/docs/architecture/system/ADR-101-ingest.md\n- https://github.com/fixture/corpus/blob/abc1234/docs/architecture/system/ADR-101-ingest.md\n- https://github.com/fixture/corpus/blob/ABC1234/docs/architecture/system/ADR-101-ingest.md\n- https://github.com/fixture/corpus/blob/v1.0/docs/architecture/system/ADR-101-ingest.md\n- https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md\n- https://github.com/fixture/corpus/blob/feature/x/docs/architecture/system/ADR-101-ingest.md\n- https://raw.githubusercontent.com/fixture/corpus/main/docs/architecture/system/ADR-101-ingest.md\n- docs\\architecture\\system\\ADR-101-ingest.md\n\n```sh\ngit mv docs/architecture/system/ADR-101-ingest.md docs/architecture/documentation/ADR-101-ingest.md\n```\n\n1. Step one\n\n ```sh\n cat docs/architecture/system/ADR-101-ingest.md\n ```\n\nSee docs/architecture/system/ADR-101-ingest.md.\n' > "$WORK/repo/NOTES.md" +edit $S/ADR-109-precedent-chain.md "s + '\nOn Windows: docs\\\\architecture\\\\system\\\\ADR-101-ingest.md\n'" +commit_all notes +(cd "$WORK/repo" && "$ADR_TOOL" domain add docs --range 300-399 --folder documentation --description Guides > /dev/null) +capture relocate-refs-dry domain move 101 docs --dry-run +capture relocate-refs domain move 101 docs +keep relocate-refs-notes.md NOTES.md +keep relocate-refs-109.md $S/ADR-109-precedent-chain.md + +# --- record edits: consider, set, supersede, enact ---------------------------------- +# +# Each edit keeps the file it wrote, so the goldens show that only the touched +# field's lines changed. + +S="docs/architecture/system" +# named_probes — an accepted record whose Summary names its probes +named_probes() { + (cd "$WORK/repo" && printf -- '---\ncontract: adr/v1\nkind: decision\nverb: add\ncapability: adr\nstatus: accepted\ndate: 2025-05-20\ndeciders: [developer, agent]\nagent: {name: Claude, model: fixture-model}\nbasis:\n - evidence: fixture measurement\n---\n\n# ADR-115: Named probes\n\n## Summary\n\n- **Decided:** the decision.\n- **Probes:** *Confident (identity-stable):* a. *Not confident (band-hint):* b.\n- **Inversion:** c.\n\n## 1. Decision\n\nThe decision.\n' > "$S/ADR-115-named-probes.md") + commit_all "named probes" +} + +fresh v1 +named_probes +capture record-consider consider 115 --said '"Fine" (Recommended)' --via "session 2025-05-20, selected from agent-written options" --operator developer --covers band-hint inversion --canary caught +keep record-consider-file.md "$S/ADR-115-named-probes.md" +capture record-consider-append consider 100 --said "ok, as it stands" --via "PR #2" --operator developer --covers +keep record-consider-append-file.md "$S/ADR-100-adopt-v1.md" +capture record-consider-new consider 113 --said "add it back" --via call --operator developer --paraphrase +keep record-consider-new-file.md "$S/ADR-113-operator-proposed.md" +worktree record-consider-status.txt +capture record-consider-unknown-probe consider 115 --said ok --via PR --operator developer --covers identity-stable no-such-probe +capture record-consider-no-names consider 101 --said ok --via PR --operator developer --covers edge-case +capture record-consider-no-said consider 115 --via PR --operator developer +capture record-consider-no-via consider 115 --said ok --via " " --operator developer +capture record-consider-v0 consider 110 --said ok --via PR --operator developer + +fresh v1 +capture record-set set 106 "capability=[search, ingest]" "related=[ADR-100, ADR-101]" extends-=ADR-100 date=2025-05-17 +keep record-set-file.md "$S/ADR-106-search-change.md" +capture record-set-remove-block set 106 related-=ADR-100 +keep record-set-remove-block-file.md "$S/ADR-106-search-change.md" +capture record-set-append-block set 114 "basis+={evidence: a second load test}" "concern+={said: Retries may double, resolve: Count retries}" +keep record-set-append-block-file.md "$S/ADR-114-open-concern.md" +capture record-set-mutable set 103 superseded_by+=ADR-107 +keep record-set-mutable-file.md "$S/ADR-103-ingest-batching.md" +capture record-set-v0 set 110 status=Superseded +keep record-set-v0-file.md "$S/ADR-110-old-v0-record.md" +worktree record-set-status.txt +capture record-set-accepted set 101 capability=search "related=[ADR-100]" +capture record-set-status-cmd set 106 status=accepted +capture record-set-status-other set 106 status=proposed +capture record-set-remove-missing set 106 related-=ADR-9 +capture record-set-bad-yaml set 106 "related=[ADR-1" +capture record-set-bad-form set 106 related +capture record-set-unknown-key set 106 capabilty=search + +fresh v1 +capture record-set-dry-run set 106 verb=add --dry-run +worktree record-set-dry-run-status.txt + +# Line endings survive a field edit. +fresh v1 +edit "$S/ADR-114-open-concern.md" "s.replace(chr(10), chr(13)+chr(10))" +capture record-set-crlf set 114 "basis+={evidence: a second load test}" +(cd "$WORK/repo" && python3 -c "import sys; d=open(sys.argv[1],'rb').read(); print('crlf kept' if b'\\r\\n' in d and d.count(b'\\n')==d.count(b'\\r\\n') else 'crlf lost')" "$S/ADR-114-open-concern.md") > "$ACTUAL/record-set-crlf-file.txt" + +fresh v1 +capture record-list-field list --field capability=ingest +capture record-list-capability list --capability adr +capture record-list-kind list --kind spec +capture record-list-verb list --verb change --field status=proposed +capture record-list-edge list --field amends=ADR-101 +capture record-list-present list --field supersedes +capture record-list-group-by list --group-by capability +capture record-list-json list --json --verb retire +capture record-list-json-group list --json --group-by verb --field enacted +fresh corpus +capture record-list-group-by-v0 list --group-by status + +fresh v1 +capture record-supersede supersede 101 --by 106 +keep record-supersede-new-file.md "$S/ADR-106-search-change.md" +keep record-supersede-old-file.md "$S/ADR-101-ingest.md" +worktree record-supersede-status.txt + +fresh v1 +capture record-supersede-amends supersede 100 --by 107 --amends 2 +keep record-supersede-amends-file.md "$S/ADR-107-ingest-notes.md" +worktree record-supersede-amends-status.txt +capture record-supersede-no-section supersede 100 --by 108 --amends 9 +capture record-supersede-kind supersede 102 --by 106 +capture record-supersede-old-proposed supersede 113 --by 106 +capture reject-for-supersede reject 108 --reason "Dropped" +capture record-supersede-new-rejected supersede 101 --by 108 + +fresh v1 +capture record-supersede-accepted supersede 101 --by 104 +keep record-supersede-accepted-file.md "$S/ADR-104-no-network-in-hooks.md" + +fresh v1 +capture record-enact enact 111 ABC1234 +keep record-enact-file.md "$S/ADR-111-cut-search.md" +capture record-enact-again enact 105 3f9c2a1dead +keep record-enact-again-file.md "$S/ADR-105-retire-legacy-ingest.md" +capture record-enact-wrong-verb enact 101 abc1234 +capture record-enact-not-hash enact 111 HEAD~1 +capture set-for-enact set 111 status=superseded --force +capture record-enact-wrong-status enact 111 abc1234 + fresh v1-defects capture v1-defects-lint lint capture v1-defects-lint-check lint --check diff --git a/tests/adr-import-roundtrip.sh b/tests/adr-import-roundtrip.sh new file mode 100755 index 00000000..f59b7d56 --- /dev/null +++ b/tests/adr-import-roundtrip.sh @@ -0,0 +1,211 @@ +#!/usr/bin/env bash +# Round-trip test for `adr import` over this repo's own records (ADR-306 §7). +# +# Works on a frozen snapshot of the pre-conversion corpus (tests/fixtures/adr/ +# v0-corpus) in a temporary git repo; the real tree +# is never written. One synthetic v0 record joins the copy: it has text above +# its H1 and a key with no v1 field, which no real record has yet. For every +# v0 record, archived and legacy ones included: +# +# 1. scan, then `apply --partial`. Sheets with a todo item lint could not +# find again (a Deprecated record's note) are skipped; the test then +# resolves those items, as a person would, and applies them. +# 2. The record is rewritten in place, and the text after its H1 is +# byte-identical to the source's, with any text from above the H1 +# moved to just below it. +# 3. Every source frontmatter key is carried with its value: the keys v1 +# shares with v0 in the frontmatter, any other key under +# imported.unmapped. status is checked against the ADR-304 §7 table, +# and imported.status keeps the source status as written. +# 4. Scanning the written v1 records and applying them again gives the +# same files (idempotence). +# +# The repo's v1 records are scanned and applied as well, and each comes back +# byte-identical. +# +# ADR_TOOL overrides the tool under test (default: docs/scripts/adr). + +set -uo pipefail + +SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" +REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" +ADR_TOOL="${ADR_TOOL:-$REPO_ROOT/docs/scripts/adr}" +[[ "$ADR_TOOL" = /* ]] || ADR_TOOL="$PWD/$ADR_TOOL" +[[ -x "$ADR_TOOL" ]] || { echo "ADR_TOOL is not executable: $ADR_TOOL" >&2; exit 2; } + +WORK="$(cd "$(mktemp -d)" && pwd -P)" +trap 'rm -rf "$WORK"' EXIT + +export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_NOSYSTEM=1 +export GIT_AUTHOR_NAME=fixture GIT_AUTHOR_EMAIL=fixture@example.invalid +export GIT_COMMITTER_NAME=fixture GIT_COMMITTER_EMAIL=fixture@example.invalid + +mkdir -p "$WORK/repo/docs" +# The corpus is a frozen snapshot of agent-ways' records from before they +# were converted, so the test keeps its 96 real v0 records (97 with the synthetic one) whatever +# the live tree holds. +cp -r "$REPO_ROOT/tests/fixtures/adr/v0-corpus/docs/architecture" "$WORK/repo/docs/architecture" +rm -rf "$WORK/repo/docs/architecture/.import" + +# The synthetic record takes the first free number in the system range. +SYNTH=$(python3 - "$WORK/repo/docs/architecture" <<'PY' +import re, sys +from pathlib import Path +used = {int(m.group(1)) for p in Path(sys.argv[1]).rglob('ADR-*.md') + if (m := re.match(r'ADR-(\d+)', p.name))} +print(next(n for n in range(100, 200) if n not in used)) +PY +) +printf -- '---\nstatus: Accepted\ndate: 2025-01-01\ndeciders: [developer]\nrevised: 2025-02-01\nreviewers:\n - someone\n---\n\n> A note written above the title.\n\n# ADR-%s: Synthetic record with a preamble\n\n## Context\n\nText.\n' "$SYNTH" \ + > "$WORK/repo/docs/architecture/system/ADR-$SYNTH-synthetic-record-with-a-preamble.md" + +(cd "$WORK/repo" && git init -q && git add -A && git commit -qm corpus) \ + || { echo "could not set up the corpus copy" >&2; exit 2; } + +python3 - "$ADR_TOOL" "$WORK/repo" <<'PY' +import re +import subprocess +import sys +from pathlib import Path + +import yaml + +tool, root = sys.argv[1], Path(sys.argv[2]) +arch = root / 'docs' / 'architecture' +sheets_dir = arch / '.import' +TITLE = re.compile(r'^# ADR-\d+(?:\.\d+)?: .+$', re.M) +BLOCKING = ('status', 'status note', 'target.number', 'target.domain') +CARRIED = ('date', 'deciders', 'related', 'supersedes', 'superseded_by', 'amends') +V0_STATUS = {'draft': 'proposed', 'proposed': 'proposed', 'accepted': 'accepted', + 'superseded': 'superseded', 'rejected': 'rejected'} +failures = [] + +def run(*args): + result = subprocess.run([tool, 'import', *args], cwd=root, capture_output=True, text=True) + if result.returncode != 0: + failures.append(f"adr import {args[0]} exited {result.returncode}:\n{result.stdout}{result.stderr}") + return result + +def commit(message): + subprocess.run(['git', 'add', '-A'], cwd=root, check=True) + subprocess.run(['git', 'commit', '-qm', message], cwd=root, check=True) + +def frontmatter(text): + if not text.startswith('---\n'): + return None + end = text.find('\n---', 4) + return yaml.safe_load(text[4:end]) or {} + +def split(text): + """(text between the frontmatter and the H1, text after the H1 line).""" + match = TITLE.search(text) + if match is None: + return None, None + start = text.find('\n---', 4) + 5 if text.startswith('---\n') else 0 + return text[start:match.start()], text[match.end() + 1:] + +def expected_status(source): + raw = str(source.get('status') or '').strip().lower() + if raw == 'deprecated': + return 'superseded' if source.get('superseded_by') else 'accepted' + return V0_STATUS.get(raw) + +records = sorted(arch.rglob('ADR-*.md')) +v0 = [p for p in records if (frontmatter(p.read_text()) or {}).get('contract') != 'adr/v1'] +v1 = [p for p in records if p not in v0] +original = {p: p.read_text() for p in records} +rel = lambda p: p.relative_to(root) + +# 1: scan and apply --partial; blocking items hold their sheets back. +run('scan', *map(str, v0)) +blocked = [] +for sheet_path in sorted(sheets_dir.glob('ADR-*.yaml')): + sheet = yaml.safe_load(sheet_path.read_text()) + if any(str(t).split(':')[0] in BLOCKING for t in sheet['todo']): + blocked.append(sheet_path) +run('apply', '--partial') +for sheet_path in blocked: + if not sheet_path.exists(): + failures.append(f"{sheet_path.name}: --partial wrote past a blocking todo item") + continue + sheet = yaml.safe_load(sheet_path.read_text()) + sheet['todo'] = [t for t in sheet['todo'] if str(t).split(':')[0] not in BLOCKING] + sheet_path.write_text(yaml.safe_dump(sheet, sort_keys=False, allow_unicode=True)) +if blocked: + run('apply', '--partial', *map(str, blocked)) + +# 2 and 3: bodies and keys, checked in the written files. +first = {} +bodies = keys = preambles = 0 +for path in v0: + text = path.read_text() + first[path] = text + written, source = frontmatter(text), frontmatter(original[path]) + if (written or {}).get('contract') != 'adr/v1': + failures.append(f"{rel(path)}: not rewritten in place as adr/v1") + continue + before, after = split(original[path]) + written_before, written_after = split(text) + if after is None or written_after is None: + failures.append(f"{rel(path)}: no H1 found") + continue + preamble = before.strip('\n') + expected = '\n' + preamble + '\n' + after if preamble.strip() else after + preambles += bool(preamble.strip()) + if written_after != expected or written_before.strip(): + failures.append(f"{rel(path)}: text after the H1 differs from the source") + else: + bodies += 1 + imported = written.get('imported') or {} + unmapped = imported.get('unmapped') or {} + wrong = [] + if 'status' not in imported or imported['status'] != source.get('status'): + wrong.append('imported.status') + for key, value in source.items(): + if key == 'status': + ok = written.get('status') == expected_status(source) + elif key in CARRIED: + ok = key in written and written[key] == value + else: + ok = key in unmapped and unmapped[key] == value and key not in written + if not ok: + wrong.append(key) + if wrong: + failures.append(f"{rel(path)}: source keys not carried with their values: {', '.join(wrong)}") + else: + keys += 1 + +# 4: the written records, scanned and applied again, come out the same. +commit('first import') +run('scan', *map(str, v0)) +run('apply', '--partial') +stable = 0 +for path in v0: + if path.read_text() != first[path]: + failures.append(f"{rel(path)}: a second scan and apply changed the record") + else: + stable += 1 + +# The v1 records reproduce byte-identically. +run('scan', *map(str, v1)) +run('apply', '--partial') +reproduced = 0 +for path in v1: + if path.read_text() != original[path]: + failures.append(f"{rel(path)}: the v1 record does not reproduce") + else: + reproduced += 1 + +leftover = sorted(p.name for p in sheets_dir.glob('ADR-*.yaml')) +if leftover: + failures.append(f"sheets left unapplied: {', '.join(leftover)}") + +for failure in failures: + print(f" FAIL: {failure}") +print(f"\nv0 records: {len(v0)} (1 synthetic); held back by --partial and resolved {len(blocked)}, " + f"with a preamble {preambles}") +print(f" keys carried with values {keys}, bodies identical {bodies}, idempotent {stable}") +print(f"v1 records: {len(v1)}; reproduced {reproduced}") +print(f"\n=== ADR Import Round Trip: {'passed' if not failures else f'{len(failures)} failure(s)'} ===") +sys.exit(1 if failures else 0) +PY diff --git a/tests/adr-macro-test.sh b/tests/adr-macro-test.sh index c2fc5e43..d3665d27 100755 --- a/tests/adr-macro-test.sh +++ b/tests/adr-macro-test.sh @@ -66,6 +66,7 @@ check "no drift note at equal versions" lacks "out of date" "$out" echo "v1-capable tool, adr/v1 contract" out=$(run_macro "$(project v2-v1 "$CURRENT" adr/v1)") check "v1 guide" contains "ADR Tooling (adr/v1)" "$out" +check "v1 guide names observable" contains "\`observable\`" "$out" check "lifecycle commands" contains "accept <n>" "$out" check "Summary guidance" contains "probes labelled confident and not confident" "$out" check "no v0 format" lacks "Record format (adr/v0)" "$out" diff --git a/tests/adr-template-test.sh b/tests/adr-template-test.sh new file mode 100755 index 00000000..9436e81c --- /dev/null +++ b/tests/adr-template-test.sh @@ -0,0 +1,23 @@ +#!/usr/bin/env bash +# The adr.yaml template, vendored into an empty repo as the adr skill does, +# lints clean before any record exists. +set -euo pipefail + +REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)" +ADR_DIR="$REPO_ROOT/hooks/ways/documentation/adr" +TMP="$(mktemp -d)" +trap 'rm -rf "$TMP"' EXIT + +git -C "$TMP" init -q +mkdir -p "$TMP/docs/scripts" "$TMP/docs/architecture" +cp "$ADR_DIR/adr-tool" "$TMP/docs/scripts/adr" +cp "$ADR_DIR/adr.yaml.template" "$TMP/docs/architecture/adr.yaml" +chmod +x "$TMP/docs/scripts/adr" + +if out=$(cd "$TMP" && docs/scripts/adr lint --check 2>&1); then + echo "=== ADR Template Test: passed ===" +else + echo "$out" + echo "=== ADR Template Test: FAILED ===" + exit 1 +fi diff --git a/tests/fixtures/adr/golden/defects-list.out b/tests/fixtures/adr/golden/defects-list.out index 0a154dc4..63fe3f9e 100644 --- a/tests/fixtures/adr/golden/defects-list.out +++ b/tests/fixtures/adr/golden/defects-list.out @@ -1,5 +1,5 @@ -ADR Defects Fixture — Architecture Decision Records (11 total) +ADR Defects Fixture — Agent Decision Records (11 total) ======================================================= ❓ ADR-140 YAML error ❓ ADR-141 Unclosed frontmatter diff --git a/tests/fixtures/adr/golden/domain-add-config.yaml b/tests/fixtures/adr/golden/domain-add-config.yaml new file mode 100644 index 00000000..3d176dde --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-config.yaml @@ -0,0 +1,48 @@ +path: docs/architecture/adr.yaml +# Fixture configuration for the adr/v1 contract (ADR-304). +project_name: ADR v1 Fixture +contract: adr/v1 + +domains: + system: + range: [100, 199] + name: System + description: Runtime and storage + folder: system + + docs: + range: [300, 399] + name: Documentation + description: Guides and references + folder: documentation + +statuses: [Draft, Proposed, Accepted, Superseded, Deprecated, Rejected] + +defaults: + deciders: [developer, agent] + status: proposed + +legacy: + range: [1, 99] + label: "Legacy (Pre-Domain Numbering)" + +kinds: + decision: + verb: required + requires: [capability, basis, agent] + sections: [Summary] + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } + spec: + verb: forbidden + requires: [capability] + edges: { supersedes: spec, decided_by: decision } + +capabilities: + adr: Decision records, their contract, and the tooling that enforces it + ingest: Document ingestion into the store + search: Query over the store + +cite: + exclude: [vendor] + +viewer: cat {file} diff --git a/tests/fixtures/adr/golden/domain-add-legacy-overlap.out b/tests/fixtures/adr/golden/domain-add-legacy-overlap.out new file mode 100644 index 00000000..9721c905 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-legacy-overlap.out @@ -0,0 +1,2 @@ +Error: range 50-60 overlaps legacy (1-99) +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-add-overlap.out b/tests/fixtures/adr/golden/domain-add-overlap.out new file mode 100644 index 00000000..0d92f1c1 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-overlap.out @@ -0,0 +1,2 @@ +Error: range 150-250 overlaps system (100-199) +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-add-refused-status.txt b/tests/fixtures/adr/golden/domain-add-refused-status.txt new file mode 100644 index 00000000..7c000b2a --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-refused-status.txt @@ -0,0 +1 @@ +M docs/architecture/adr.yaml diff --git a/tests/fixtures/adr/golden/domain-add-refused.out b/tests/fixtures/adr/golden/domain-add-refused.out new file mode 100644 index 00000000..7260b635 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-refused.out @@ -0,0 +1,4 @@ +Error: domain 'docs' already exists +Error: --range 400-300 runs backwards +Error: folder 'system' already belongs to system +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-add-v0-config.yaml b/tests/fixtures/adr/golden/domain-add-v0-config.yaml new file mode 100644 index 00000000..2493c3a2 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-v0-config.yaml @@ -0,0 +1,50 @@ +path: docs/architecture/adr.yaml +# Fixture ADR configuration for the golden-output test (tests/adr-golden-test.sh). +project_name: ADR Fixture + +domains: + system: + range: [100, 199] + name: System + description: Runtime, hooks and storage + folder: system + + ops: + range: [200, 299] + name: Operations + description: Deployment and runbooks, split across two folders + folder: + - operations + - runbooks + + docs: + range: [300, 399] + name: Documentation + description: Documentation structure and tooling + folder: documentation + + api: + range: [400, 499] + name: API surface + description: "" + folder: api + +statuses: + - Draft + - Proposed + - Accepted + - Superseded + - Deprecated + - Rejected + +defaults: + deciders: + - developer + - agent + status: Draft + +legacy: + range: [1, 99] + label: "Legacy (Pre-Domain Numbering)" + +viewer: cat {file} diff --git a/tests/fixtures/adr/golden/domain-add-v0.out b/tests/fixtures/adr/golden/domain-add-v0.out new file mode 100644 index 00000000..53b8fa70 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add-v0.out @@ -0,0 +1,4 @@ +Added domain: api (400-499) + Name: API surface + Folder: docs/architecture/api/ +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-add.out b/tests/fixtures/adr/golden/domain-add.out new file mode 100644 index 00000000..f26739c4 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-add.out @@ -0,0 +1,4 @@ +Added domain: docs (300-399) + Name: Documentation + Folder: docs/architecture/documentation/ +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-104.md b/tests/fixtures/adr/golden/domain-move-104.md new file mode 100644 index 00000000..781c0b1b --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-104.md @@ -0,0 +1,32 @@ +path: docs/architecture/documentation/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +See [the chain](../system/ADR-109-precedent-chain.md). diff --git a/tests/fixtures/adr/golden/domain-move-109.md b/tests/fixtures/adr/golden/domain-move-109.md new file mode 100644 index 00000000..3255dbf1 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-109.md @@ -0,0 +1,36 @@ +path: docs/architecture/system/ADR-109-precedent-chain.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: ingest +supersedes: [ADR-103] +status: proposed +date: 2025-05-10 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - precedent: ADR-101 + - precedent: ADR-102 +--- + +# ADR-109: Grounded through precedent, via a decision and a spec + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +## 3. Notes + +ADR-104 applies; see [the constraint](../documentation/ADR-104-no-network-in-hooks.md) and [ingest](./ADR-101-ingest.md). diff --git a/tests/fixtures/adr/golden/domain-move-again.out b/tests/fixtures/adr/golden/domain-move-again.out new file mode 100644 index 00000000..ea90ca2a --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-again.out @@ -0,0 +1,2 @@ +Error: ADR-104 is already in docs (docs/architecture/documentation/ADR-104-no-network-in-hooks.md) +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-cite-readded.out b/tests/fixtures/adr/golden/domain-move-cite.out similarity index 76% rename from tests/fixtures/adr/golden/v1-cite-readded.out rename to tests/fixtures/adr/golden/domain-move-cite.out index d1147b20..df56475b 100644 --- a/tests/fixtures/adr/golden/v1-cite-readded.out +++ b/tests/fixtures/adr/golden/domain-move-cite.out @@ -1,6 +1,8 @@ ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force + ❌ src/search.py:4 ADR-1040 resolves to no record + ❌ src/search.py:4 ADR-104.1 resolves to no record ════════════════════════════════════════════════════════════ -Citations: 0 errors, 1 warnings +Citations: 2 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-dry-status.txt b/tests/fixtures/adr/golden/domain-move-dry-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/domain-move-dry.out b/tests/fixtures/adr/golden/domain-move-dry.out new file mode 100644 index 00000000..5ff938ec --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-dry.out @@ -0,0 +1,25 @@ +Would move: ADR-104 system → docs + docs/architecture/system/ADR-104-no-network-in-hooks.md → docs/architecture/documentation/ADR-104-no-network-in-hooks.md +Would rewrite 5 paths in 5 files + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 + docs/architecture/system/ADR-104-no-network-in-hooks.md:31 + See [the chain](ADR-109-precedent-chain.md). + → See [the chain](../system/ADR-109-precedent-chain.md). + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/architecture/system/ADR-109-precedent-chain.md:35 + ADR-104 applies; see [the constraint](ADR-104-no-network-in-hooks.md) and [ingest](./ADR-101-ingest.md). + → ADR-104 applies; see [the constraint](../documentation/ADR-104-no-network-in-hooks.md) and [ingest](./ADR-101-ingest.md). + docs/architecture/system/ADR-111-cut-search.md: 1 + docs/architecture/system/ADR-111-cut-search.md:27 + See [ADR-104](ADR-104-no-network-in-hooks.md). + → See [ADR-104](../documentation/ADR-104-no-network-in-hooks.md). + docs/guide/hooks.md: 1 + docs/guide/hooks.md:7 + See [no network](../architecture/system/ADR-104-no-network-in-hooks.md) and ADR-104#1. + → See [no network](../architecture/documentation/ADR-104-no-network-in-hooks.md) and ADR-104#1. + src/search.py: 1 + src/search.py:3 + # ADR-104 hooks stay offline: docs/architecture/system/ADR-104-no-network-in-hooks.md + → # ADR-104 hooks stay offline: docs/architecture/documentation/ADR-104-no-network-in-hooks.md +Dry run: nothing written. +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-guide.md b/tests/fixtures/adr/golden/domain-move-guide.md new file mode 100644 index 00000000..15c6aa20 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-guide.md @@ -0,0 +1,8 @@ +path: docs/guide/hooks.md +--- +domain: system +--- + +# Hooks guide + +See [no network](../architecture/documentation/ADR-104-no-network-in-hooks.md) and ADR-104#1. diff --git a/tests/fixtures/adr/golden/domain-move-index.md b/tests/fixtures/adr/golden/domain-move-index.md new file mode 100644 index 00000000..55a5d5fc --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-index.md @@ -0,0 +1,42 @@ +path: docs/architecture/INDEX.md +# Agent Decision Records + +This directory contains the Agent Decision Records (ADRs) for ADR v1 Fixture, under the adr/v1 contract. +A record's kind says what it holds: a decision, a spec kept current, or evidence (findings, surveys, explorations). + +## Record Format + +- **Kind:** decision / spec / evidence, as `adr.yaml` declares +- **Status:** proposed / accepted / rejected / abandoned / superseded / archived +- **Capability:** what the record is about, from the vocabulary in `adr.yaml` +- **Decisions** also carry a verb (add / cut / change / retire / constrain), a basis, and open with a Summary +- **Numbers** are permanent; a record's folder is its area + +_This index is auto-generated by `adr index`. Configuration: [`adr.yaml`](./adr.yaml)_ + +## System +_Runtime and storage_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-100](./system/ADR-100-adopt-v1.md) | Adopt the adr/v1 contract | accepted | +| [ADR-101](./system/ADR-101-ingest.md) | Add ingestion | accepted | +| [ADR-102](./system/ADR-102-ingest-spec.md) | How ingestion works | accepted | +| [ADR-103](./system/ADR-103-ingest-batching.md) | Batch ingestion | accepted (partially superseded by ADR-109) | +| [ADR-105](./system/ADR-105-retire-legacy-ingest.md) | Retire the legacy ingest surface | accepted | +| [ADR-106](./system/ADR-106-search-change.md) | Search result caps | proposed | +| [ADR-107](./system/ADR-107-ingest-notes.md) | Ingest notes, by slug reference | proposed | +| [ADR-108](./system/ADR-108-ingest-over-v0.md) | Change against a v0 record | proposed | +| [ADR-109](./system/ADR-109-precedent-chain.md) | Grounded through precedent, via a decision and a spec | proposed | +| [ADR-110](./system/ADR-110-old-v0-record.md) | An unmigrated v0 record | Accepted (partially superseded by ADR-108) | +| [ADR-111](./system/ADR-111-cut-search.md) | Cut search | accepted | +| [ADR-112](./system/ADR-112-authored.md) | Operator-authored constraint | accepted | +| [ADR-113](./system/ADR-113-operator-proposed.md) | Operator-started and still proposed | proposed | +| [ADR-114](./system/ADR-114-open-concern.md) | Cap batch size | proposed | + +## Docs +_Guides_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-104](./documentation/ADR-104-no-network-in-hooks.md) | No network calls in hooks | accepted | diff --git a/tests/fixtures/adr/golden/v1-cite-inventory-fails.out b/tests/fixtures/adr/golden/domain-move-lint.out similarity index 52% rename from tests/fixtures/adr/golden/v1-cite-inventory-fails.out rename to tests/fixtures/adr/golden/domain-move-lint.out index 1901336e..f9e64731 100644 --- a/tests/fixtures/adr/golden/v1-cite-inventory-fails.out +++ b/tests/fixtures/adr/golden/domain-move-lint.out @@ -1,8 +1,14 @@ - ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force - ⚠️ src/search.py:1 ADR-106 is on 'search', which ADR-111 cuts; remove before enacting - ⚠️ docs/architecture/system/ADR-105-retire-legacy-ingest.md inventory for 'cli' did not run (exit 3): exit 3 + +Scanned: 3 ADRs + +Status distribution: + accepted: 2 + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) ════════════════════════════════════════════════════════════ -Citations: 0 errors, 3 warnings +Summary: 0 errors, 0 warnings ════════════════════════════════════════════════════════════ + [exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-list.out b/tests/fixtures/adr/golden/domain-move-list.out new file mode 100644 index 00000000..f9d9bce3 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-list.out @@ -0,0 +1,27 @@ + +ADR v1 Fixture — Agent Decision Records (15 total) +======================================================= + +## System (system) +-------------------------------------------------- + ✅ ADR-100 Adopt the adr/v1 contract + ✅ ADR-101 Add ingestion + ✅ ADR-102 How ingestion works + ✅ ADR-103 Batch ingestion (partially superseded by ADR-109) + ✅ ADR-105 Retire the legacy ingest surface + 💡 ADR-106 Search result caps + 💡 ADR-107 Ingest notes, by slug reference + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + ✅ ADR-110 An unmigrated v0 record (partially superseded by ADR-108) + ✅ ADR-111 Cut search + ✅ ADR-112 Operator-authored constraint + 💡 ADR-113 Operator-started and still proposed + 💡 ADR-114 Cap batch size + +## Docs (docs) +-------------------------------------------------- + ✅ ADR-104 No network calls in hooks + +Total: 15 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-new-docs.out b/tests/fixtures/adr/golden/domain-move-new-docs.out new file mode 100644 index 00000000..a6680383 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-new-docs.out @@ -0,0 +1,5 @@ +Created: docs/architecture/documentation/ADR-300-style-guide.md + Domain: Docs (docs) + Number: ADR-300 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-new-system.out b/tests/fixtures/adr/golden/domain-move-new-system.out new file mode 100644 index 00000000..47d3a70c --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-new-system.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-115-queue-limits.md + Domain: System (system) + Number: ADR-115 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-plan-109.md b/tests/fixtures/adr/golden/domain-move-plan-109.md new file mode 100644 index 00000000..3255dbf1 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan-109.md @@ -0,0 +1,36 @@ +path: docs/architecture/system/ADR-109-precedent-chain.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: ingest +supersedes: [ADR-103] +status: proposed +date: 2025-05-10 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - precedent: ADR-101 + - precedent: ADR-102 +--- + +# ADR-109: Grounded through precedent, via a decision and a spec + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +## 3. Notes + +ADR-104 applies; see [the constraint](../documentation/ADR-104-no-network-in-hooks.md) and [ingest](./ADR-101-ingest.md). diff --git a/tests/fixtures/adr/golden/domain-move-plan-111.md b/tests/fixtures/adr/golden/domain-move-plan-111.md new file mode 100644 index 00000000..d9e9e2f8 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan-111.md @@ -0,0 +1,28 @@ +path: docs/architecture/documentation/ADR-111-cut-search.md +--- +contract: adr/v1 +kind: decision +verb: cut +capability: search +status: accepted +date: 2025-05-11 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: usage data shows no queries +--- + +# ADR-111: Cut search + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +See [ADR-104](ADR-104-no-network-in-hooks.md). diff --git a/tests/fixtures/adr/golden/domain-move-plan-number.out b/tests/fixtures/adr/golden/domain-move-plan-number.out new file mode 100644 index 00000000..56232f3c --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan-number.out @@ -0,0 +1,2 @@ +Error: plan entry 1: unknown key number; a record keeps its number +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-move-plan-refused-status.txt b/tests/fixtures/adr/golden/domain-move-plan-refused-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/domain-move-plan-status.txt b/tests/fixtures/adr/golden/domain-move-plan-status.txt new file mode 100644 index 00000000..b565a5cb --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan-status.txt @@ -0,0 +1,6 @@ +M docs/architecture/INDEX.md +R docs/architecture/system/ADR-104-no-network-in-hooks.md -> docs/architecture/documentation/ADR-104-no-network-in-hooks.md +R docs/architecture/system/ADR-111-cut-search.md -> docs/architecture/documentation/ADR-111-cut-search.md +M docs/architecture/system/ADR-109-precedent-chain.md +M docs/guide/hooks.md +M src/search.py diff --git a/tests/fixtures/adr/golden/domain-move-plan-twice.out b/tests/fixtures/adr/golden/domain-move-plan-twice.out new file mode 100644 index 00000000..c7d85705 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan-twice.out @@ -0,0 +1,2 @@ +Error: ADR-104 is listed twice +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-move-plan.out b/tests/fixtures/adr/golden/domain-move-plan.out new file mode 100644 index 00000000..59491838 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-plan.out @@ -0,0 +1,14 @@ +Moved: ADR-104 system → docs + docs/architecture/system/ADR-104-no-network-in-hooks.md → docs/architecture/documentation/ADR-104-no-network-in-hooks.md +Moved: ADR-111 system → docs + docs/architecture/system/ADR-111-cut-search.md → docs/architecture/documentation/ADR-111-cut-search.md +Rewrote 4 paths in 4 files + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/guide/hooks.md: 1 + src/search.py: 1 +Index needs updating: docs/architecture/INDEX.md + 15 ADRs across 2 domains + Changes: +4 -2 lines +Updated: docs/architecture/INDEX.md +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-refused.out b/tests/fixtures/adr/golden/domain-move-refused.out new file mode 100644 index 00000000..22c2b720 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-refused.out @@ -0,0 +1,2 @@ +Error: ADR not found: 999 +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-move-scan.out b/tests/fixtures/adr/golden/domain-move-scan.out new file mode 100644 index 00000000..6db98057 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-scan.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/documentation/ADR-104-no-network-in-hooks.md -> docs/architecture/.import/ADR-104.yaml (v1, 0 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-move-search.py b/tests/fixtures/adr/golden/domain-move-search.py new file mode 100644 index 00000000..40f43529 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-search.py @@ -0,0 +1,5 @@ +path: src/search.py +# ADR-106 search result caps: search is cut by ADR-111, not yet enacted. +# ADR-102 ingest spec: in force. +# ADR-104 hooks stay offline: docs/architecture/documentation/ADR-104-no-network-in-hooks.md +# ADR-1040 and ADR-104.1 are other numbers. diff --git a/tests/fixtures/adr/golden/domain-move-status.txt b/tests/fixtures/adr/golden/domain-move-status.txt new file mode 100644 index 00000000..709d1477 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-status.txt @@ -0,0 +1,6 @@ +M docs/architecture/INDEX.md +R docs/architecture/system/ADR-104-no-network-in-hooks.md -> docs/architecture/documentation/ADR-104-no-network-in-hooks.md +M docs/architecture/system/ADR-109-precedent-chain.md +M docs/architecture/system/ADR-111-cut-search.md +M docs/guide/hooks.md +M src/search.py diff --git a/tests/fixtures/adr/golden/domain-move-v0.out b/tests/fixtures/adr/golden/domain-move-v0.out new file mode 100644 index 00000000..5c54559b --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move-v0.out @@ -0,0 +1,2 @@ +Error: under adr/v0 a record's number decides its domain, so a record cannot change domain by folder; declare contract: adr/v1 in adr.yaml first +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-move.out b/tests/fixtures/adr/golden/domain-move.out new file mode 100644 index 00000000..7ebcf0ba --- /dev/null +++ b/tests/fixtures/adr/golden/domain-move.out @@ -0,0 +1,13 @@ +Moved: ADR-104 system → docs + docs/architecture/system/ADR-104-no-network-in-hooks.md → docs/architecture/documentation/ADR-104-no-network-in-hooks.md +Rewrote 5 paths in 5 files + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/architecture/system/ADR-111-cut-search.md: 1 + docs/guide/hooks.md: 1 + src/search.py: 1 +Index needs updating: docs/architecture/INDEX.md + 15 ADRs across 2 domains + Changes: +3 -1 lines +Updated: docs/architecture/INDEX.md +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-rename-109.md b/tests/fixtures/adr/golden/domain-rename-109.md new file mode 100644 index 00000000..f4bc8442 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-109.md @@ -0,0 +1,36 @@ +path: docs/architecture/platform/ADR-109-precedent-chain.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: ingest +supersedes: [ADR-103] +status: proposed +date: 2025-05-10 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - precedent: ADR-101 + - precedent: ADR-102 +--- + +# ADR-109: Grounded through precedent, via a decision and a spec + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +## 3. Notes + +See [the constraint](../platform/ADR-104-no-network-in-hooks.md). diff --git a/tests/fixtures/adr/golden/domain-rename-config.yaml b/tests/fixtures/adr/golden/domain-rename-config.yaml new file mode 100644 index 00000000..4d0ae37c --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-config.yaml @@ -0,0 +1,42 @@ +path: docs/architecture/adr.yaml +# Fixture configuration for the adr/v1 contract (ADR-304). +project_name: ADR v1 Fixture +contract: adr/v1 + +domains: + platform: + range: [100, 199] + name: System + description: Runtime and storage + folder: platform + +statuses: [Draft, Proposed, Accepted, Superseded, Deprecated, Rejected] + +defaults: + deciders: [developer, agent] + status: proposed + +legacy: + range: [1, 99] + label: "Legacy (Pre-Domain Numbering)" + +kinds: + decision: + verb: required + requires: [capability, basis, agent] + sections: [Summary] + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } + spec: + verb: forbidden + requires: [capability] + edges: { supersedes: spec, decided_by: decision } + +capabilities: + adr: Decision records, their contract, and the tooling that enforces it + ingest: Document ingestion into the store + search: Query over the store + +cite: + exclude: [vendor] + +viewer: cat {file} diff --git a/tests/fixtures/adr/golden/domain-rename-dry-status.txt b/tests/fixtures/adr/golden/domain-rename-dry-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/domain-rename-dry.out b/tests/fixtures/adr/golden/domain-rename-dry.out new file mode 100644 index 00000000..8cc491c4 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-dry.out @@ -0,0 +1,30 @@ +Would rename domain: system → platform + Folder: docs/architecture/system/ → docs/architecture/platform/ +Would rewrite 8 paths in 4 files + docs/architecture/adr.yaml: 2 + docs/architecture/adr.yaml:6 + system: + → platform: + docs/architecture/adr.yaml:10 + folder: system + → folder: platform + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/architecture/system/ADR-109-precedent-chain.md:35 + See [the constraint](../system/ADR-104-no-network-in-hooks.md). + → See [the constraint](../platform/ADR-104-no-network-in-hooks.md). + docs/guide/hooks.md: 4 + docs/guide/hooks.md:2 + domain: system + → domain: platform + docs/guide/hooks.md:7 + See [no network](../architecture/system/ADR-104-no-network-in-hooks.md) and ADR-104. + → See [no network](../architecture/platform/ADR-104-no-network-in-hooks.md) and ADR-104. + docs/guide/hooks.md:8 + Records live in docs/architecture/system/ and [the folder](../architecture/system/). + → Records live in docs/architecture/platform/ and [the folder](../architecture/platform/). + src/search.py: 1 + src/search.py:3 + # ADR-104 hooks stay offline: docs/architecture/system/ADR-104-no-network-in-hooks.md + → # ADR-104 hooks stay offline: docs/architecture/platform/ADR-104-no-network-in-hooks.md +Dry run: nothing written. +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-rename-folder-status.txt b/tests/fixtures/adr/golden/domain-rename-folder-status.txt new file mode 100644 index 00000000..356712d3 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-folder-status.txt @@ -0,0 +1,18 @@ +M docs/architecture/adr.yaml +R docs/architecture/system/ADR-100-adopt-v1.md -> docs/architecture/kernel/ADR-100-adopt-v1.md +R docs/architecture/system/ADR-101-ingest.md -> docs/architecture/kernel/ADR-101-ingest.md +R docs/architecture/system/ADR-102-ingest-spec.md -> docs/architecture/kernel/ADR-102-ingest-spec.md +R docs/architecture/system/ADR-103-ingest-batching.md -> docs/architecture/kernel/ADR-103-ingest-batching.md +R docs/architecture/system/ADR-104-no-network-in-hooks.md -> docs/architecture/kernel/ADR-104-no-network-in-hooks.md +R docs/architecture/system/ADR-105-retire-legacy-ingest.md -> docs/architecture/kernel/ADR-105-retire-legacy-ingest.md +R docs/architecture/system/ADR-106-search-change.md -> docs/architecture/kernel/ADR-106-search-change.md +R docs/architecture/system/ADR-107-ingest-notes.md -> docs/architecture/kernel/ADR-107-ingest-notes.md +R docs/architecture/system/ADR-108-ingest-over-v0.md -> docs/architecture/kernel/ADR-108-ingest-over-v0.md +R docs/architecture/system/ADR-109-precedent-chain.md -> docs/architecture/kernel/ADR-109-precedent-chain.md +R docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/kernel/ADR-110-old-v0-record.md +R docs/architecture/system/ADR-111-cut-search.md -> docs/architecture/kernel/ADR-111-cut-search.md +R docs/architecture/system/ADR-112-authored.md -> docs/architecture/kernel/ADR-112-authored.md +R docs/architecture/system/ADR-113-operator-proposed.md -> docs/architecture/kernel/ADR-113-operator-proposed.md +R docs/architecture/system/ADR-114-open-concern.md -> docs/architecture/kernel/ADR-114-open-concern.md +M docs/guide/hooks.md +M src/search.py diff --git a/tests/fixtures/adr/golden/domain-rename-folder.out b/tests/fixtures/adr/golden/domain-rename-folder.out new file mode 100644 index 00000000..8979baa8 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-folder.out @@ -0,0 +1,8 @@ +Renamed domain: system → core + Folder: docs/architecture/system/ → docs/architecture/kernel/ +Rewrote 8 paths in 4 files + docs/architecture/adr.yaml: 2 + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/guide/hooks.md: 4 + src/search.py: 1 +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-rename-guide.md b/tests/fixtures/adr/golden/domain-rename-guide.md new file mode 100644 index 00000000..701a370e --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-guide.md @@ -0,0 +1,9 @@ +path: docs/guide/hooks.md +--- +domain: platform +--- + +# Hooks guide + +See [no network](../architecture/platform/ADR-104-no-network-in-hooks.md) and ADR-104. +Records live in docs/architecture/platform/ and [the folder](../architecture/platform/). diff --git a/tests/fixtures/adr/golden/v1-lint-baseline.out b/tests/fixtures/adr/golden/domain-rename-lint.out similarity index 73% rename from tests/fixtures/adr/golden/v1-lint-baseline.out rename to tests/fixtures/adr/golden/domain-rename-lint.out index e9a6b42c..dac16205 100644 --- a/tests/fixtures/adr/golden/v1-lint-baseline.out +++ b/tests/fixtures/adr/golden/domain-rename-lint.out @@ -8,20 +8,14 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 3 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md +docs/architecture/platform/ADR-114-open-concern.md ⚠️ open concern: Batch size may starve small tenants ════════════════════════════════════════════════════════════ -Summary: 0 errors, 3 warnings +Summary: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/domain-rename-list.out b/tests/fixtures/adr/golden/domain-rename-list.out new file mode 100644 index 00000000..20c2bac5 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-list.out @@ -0,0 +1,24 @@ + +ADR v1 Fixture — Agent Decision Records (15 total) +======================================================= + +## System (platform) +-------------------------------------------------- + ✅ ADR-100 Adopt the adr/v1 contract + ✅ ADR-101 Add ingestion + ✅ ADR-102 How ingestion works + ✅ ADR-103 Batch ingestion (partially superseded by ADR-109) + ✅ ADR-104 No network calls in hooks + ✅ ADR-105 Retire the legacy ingest surface + 💡 ADR-106 Search result caps + 💡 ADR-107 Ingest notes, by slug reference + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + ✅ ADR-110 An unmigrated v0 record (partially superseded by ADR-108) + ✅ ADR-111 Cut search + ✅ ADR-112 Operator-authored constraint + 💡 ADR-113 Operator-started and still proposed + 💡 ADR-114 Cap batch size + +Total: 15 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/domain-rename-refused.out b/tests/fixtures/adr/golden/domain-rename-refused.out new file mode 100644 index 00000000..94bb02f4 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-refused.out @@ -0,0 +1,2 @@ +Error: folder 'archive' is not a folder name under docs/architecture (and not archive) +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-rename-same.out b/tests/fixtures/adr/golden/domain-rename-same.out new file mode 100644 index 00000000..330f3755 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-same.out @@ -0,0 +1,2 @@ +Error: nothing to rename: give a new name or --folder +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-rename-search.py b/tests/fixtures/adr/golden/domain-rename-search.py new file mode 100644 index 00000000..22dabd13 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-search.py @@ -0,0 +1,4 @@ +path: src/search.py +# ADR-106 search result caps: search is cut by ADR-111, not yet enacted. +# ADR-102 ingest spec: in force. +# ADR-104 hooks stay offline: docs/architecture/platform/ADR-104-no-network-in-hooks.md diff --git a/tests/fixtures/adr/golden/domain-rename-status.txt b/tests/fixtures/adr/golden/domain-rename-status.txt new file mode 100644 index 00000000..5cf44dc5 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-status.txt @@ -0,0 +1,18 @@ +M docs/architecture/adr.yaml +R docs/architecture/system/ADR-100-adopt-v1.md -> docs/architecture/platform/ADR-100-adopt-v1.md +R docs/architecture/system/ADR-101-ingest.md -> docs/architecture/platform/ADR-101-ingest.md +R docs/architecture/system/ADR-102-ingest-spec.md -> docs/architecture/platform/ADR-102-ingest-spec.md +R docs/architecture/system/ADR-103-ingest-batching.md -> docs/architecture/platform/ADR-103-ingest-batching.md +R docs/architecture/system/ADR-104-no-network-in-hooks.md -> docs/architecture/platform/ADR-104-no-network-in-hooks.md +R docs/architecture/system/ADR-105-retire-legacy-ingest.md -> docs/architecture/platform/ADR-105-retire-legacy-ingest.md +R docs/architecture/system/ADR-106-search-change.md -> docs/architecture/platform/ADR-106-search-change.md +R docs/architecture/system/ADR-107-ingest-notes.md -> docs/architecture/platform/ADR-107-ingest-notes.md +R docs/architecture/system/ADR-108-ingest-over-v0.md -> docs/architecture/platform/ADR-108-ingest-over-v0.md +R docs/architecture/system/ADR-109-precedent-chain.md -> docs/architecture/platform/ADR-109-precedent-chain.md +R docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/platform/ADR-110-old-v0-record.md +R docs/architecture/system/ADR-111-cut-search.md -> docs/architecture/platform/ADR-111-cut-search.md +R docs/architecture/system/ADR-112-authored.md -> docs/architecture/platform/ADR-112-authored.md +R docs/architecture/system/ADR-113-operator-proposed.md -> docs/architecture/platform/ADR-113-operator-proposed.md +R docs/architecture/system/ADR-114-open-concern.md -> docs/architecture/platform/ADR-114-open-concern.md +M docs/guide/hooks.md +M src/search.py diff --git a/tests/fixtures/adr/golden/domain-rename-unknown.out b/tests/fixtures/adr/golden/domain-rename-unknown.out new file mode 100644 index 00000000..53cc3aa6 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename-unknown.out @@ -0,0 +1,3 @@ +Error: Unknown domain 'nosuch' +Valid domains: system +[exit 1] diff --git a/tests/fixtures/adr/golden/domain-rename.out b/tests/fixtures/adr/golden/domain-rename.out new file mode 100644 index 00000000..0af849e7 --- /dev/null +++ b/tests/fixtures/adr/golden/domain-rename.out @@ -0,0 +1,8 @@ +Renamed domain: system → platform + Folder: docs/architecture/system/ → docs/architecture/platform/ +Rewrote 8 paths in 4 files + docs/architecture/adr.yaml: 2 + docs/architecture/system/ADR-109-precedent-chain.md: 1 + docs/guide/hooks.md: 4 + src/search.py: 1 +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-bad-status.txt b/tests/fixtures/adr/golden/import-apply-bad-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/import-apply-bad.out b/tests/fixtures/adr/golden/import-apply-bad.out new file mode 100644 index 00000000..b0188d87 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-bad.out @@ -0,0 +1,14 @@ +Refused: docs/architecture/.import/float.yaml: target.number 101.1 reads as a decimal; quote a sub-part number, as in number: '101.10' +Refused: docs/architecture/.import/escape.yaml: target.number '110.1/../../x' is not a record number +Refused: docs/architecture/.import/octal.yaml: target.number 0156 is not a plain decimal, and YAML reads it as 110; quote it, as in number: '0156' +Refused: docs/architecture/.import/hex.yaml: target.number 0x6E is not a plain decimal, and YAML reads it as 110; quote it, as in number: '0x6E' +Refused: docs/architecture/.import/listbody.yaml: body is not text +Refused: docs/architecture/.import/nobody.yaml: the source has a body and the sheet has none; scan it again +Refused: docs/architecture/.import/broken.yaml: the sheet is not valid YAML: mapping values are not allowed here at line 2 +Refused: docs/architecture/.import/domain.yaml: the source is ADR-110 in system; the sheet says ADR-110 in legacy. A record keeps its number; moving it to another domain is `adr domain move` (ADR-306 §6). Restore target.number and target.domain +Refused: docs/architecture/.import/nodomain.yaml: the source is ADR-110 in system; the sheet says ADR-110 in storage. A record keeps its number; moving it to another domain is `adr domain move` (ADR-306 §6). Restore target.number and target.domain +Refused: docs/architecture/.import/ADR-115.yaml: the source is ADR-115 in system; the sheet says ADR-117 in system. A record keeps its number; moving it to another domain is `adr domain move` (ADR-306 §6). Restore target.number and target.domain +Skipped: docs/architecture/.import/ADR-116.yaml: 1 todo item(s) --partial does not write past: status note +Applied: docs/architecture/.import/ADR-101.yaml -> docs/architecture/system/ADR-101-ingest.md, unchanged (lint: 0 errors, 0 warnings) +Import: 1 applied, 1 skipped, 10 refused +[exit 1] diff --git a/tests/fixtures/adr/golden/import-apply-changed-status.txt b/tests/fixtures/adr/golden/import-apply-changed-status.txt new file mode 100644 index 00000000..7cef6fc3 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-changed-status.txt @@ -0,0 +1 @@ +M docs/architecture/system/ADR-110-old-v0-record.md diff --git a/tests/fixtures/adr/golden/import-apply-changed.out b/tests/fixtures/adr/golden/import-apply-changed.out new file mode 100644 index 00000000..ab4ecd2c --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-changed.out @@ -0,0 +1,3 @@ +Refused: docs/architecture/.import/ADR-110.yaml: source docs/architecture/system/ADR-110-old-v0-record.md changed since the scan; scan it again +Import: 0 applied, 0 skipped, 1 refused +[exit 1] diff --git a/tests/fixtures/adr/golden/import-apply-complete-file.md b/tests/fixtures/adr/golden/import-apply-complete-file.md new file mode 100644 index 00000000..f8179aba --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-complete-file.md @@ -0,0 +1,36 @@ +path: docs/architecture/system/ADR-110-old-v0-record.md +--- +contract: adr/v1 +kind: decision +verb: add +capability: ingest +superseded_by: + - ADR-108 +basis: + - evidence: migrated from v0 +agent: + name: Claude + model: unrecorded +status: accepted +date: 2025-01-01 +deciders: + - developer +imported: + from: docs/architecture/system/ADR-110-old-v0-record.md + format: v0 + status: Accepted +--- + +# ADR-110: An unmigrated v0 record + +## Summary + +Context for the record. + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/v1-cite-enacted.out b/tests/fixtures/adr/golden/import-apply-complete-lint.out similarity index 50% rename from tests/fixtures/adr/golden/v1-cite-enacted.out rename to tests/fixtures/adr/golden/import-apply-complete-lint.out index 8adb3f76..691df909 100644 --- a/tests/fixtures/adr/golden/v1-cite-enacted.out +++ b/tests/fixtures/adr/golden/import-apply-complete-lint.out @@ -1,8 +1,13 @@ - ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force - ❌ src/search.py:1 ADR-106 is on 'search', cut and enacted by ADR-111; remove the citation - ❌ docs/architecture/system/ADR-105-retire-legacy-ingest.md cli:ingest-legacy is still present after ADR-105 was enacted + +Scanned: 1 ADRs + +Status distribution: + accepted: 1 + +Contract: adr/v1 (2 v0 records remain) ════════════════════════════════════════════════════════════ -Citations: 2 errors, 1 warnings +Summary: 0 errors, 0 warnings ════════════════════════════════════════════════════════════ -[exit 1] + +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-complete.out b/tests/fixtures/adr/golden/import-apply-complete.out new file mode 100644 index 00000000..de9a9fc9 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-complete.out @@ -0,0 +1,3 @@ +Applied: docs/architecture/.import/ADR-110.yaml -> docs/architecture/system/ADR-110-old-v0-record.md (lint: 0 errors, 0 warnings) +Import: 1 applied, 0 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-dryrun-status.txt b/tests/fixtures/adr/golden/import-apply-dryrun-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/import-apply-dryrun.out b/tests/fixtures/adr/golden/import-apply-dryrun.out new file mode 100644 index 00000000..1d356820 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-dryrun.out @@ -0,0 +1,3 @@ +Would apply: docs/architecture/.import/ADR-110.yaml -> docs/architecture/system/ADR-110-old-v0-record.md (lint: 0 errors, 0 warnings) +Import: 1 would apply, 0 skipped, 0 refused (dry run: nothing written) +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-foreign-dryrun-status.txt b/tests/fixtures/adr/golden/import-apply-foreign-dryrun-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/import-apply-foreign-dryrun.out b/tests/fixtures/adr/golden/import-apply-foreign-dryrun.out new file mode 100644 index 00000000..18e7f036 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-foreign-dryrun.out @@ -0,0 +1,8 @@ +Would apply: docs/architecture/.import/ADR-042.yaml -> docs/architecture/system/ADR-150-foreign-record.md, 4 todo left (lint: 4 errors, 1 warnings) + error: a decision record requires a verb (add, cut, change, retire, constrain) + error: a decision record requires 'capability' + error: a decision record requires 'basis' + error: agent: records 'name' + warning: imported record has no '## Summary' yet (ADR-306 §4) +Import: 1 would apply, 0 skipped, 0 refused (dry run: nothing written) +[exit 1] diff --git a/tests/fixtures/adr/golden/import-apply-foreign-file.md b/tests/fixtures/adr/golden/import-apply-foreign-file.md new file mode 100644 index 00000000..ce9b95ff --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-foreign-file.md @@ -0,0 +1,25 @@ +path: docs/architecture/system/ADR-150-foreign-record.md +--- +contract: adr/v1 +kind: decision +verb: ~ +capability: ~ +basis: [] +agent: + name: ~ + model: unrecorded +status: accepted +date: 2025-04-01 +deciders: + - developer +imported: + from: ADR-042-foreign-record.md + format: v0 + status: Accepted +--- + +# ADR-150: Foreign record + +## Context + +From elsewhere. diff --git a/tests/fixtures/adr/golden/import-apply-foreign-status.txt b/tests/fixtures/adr/golden/import-apply-foreign-status.txt new file mode 100644 index 00000000..0b8e1be0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-foreign-status.txt @@ -0,0 +1 @@ +A docs/architecture/system/ADR-150-foreign-record.md diff --git a/tests/fixtures/adr/golden/import-apply-foreign.out b/tests/fixtures/adr/golden/import-apply-foreign.out new file mode 100644 index 00000000..b5222d76 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-foreign.out @@ -0,0 +1,3 @@ +Applied: docs/architecture/.import/ADR-042.yaml -> docs/architecture/system/ADR-150-foreign-record.md, 4 todo left (lint: 4 errors, 1 warnings) +Import: 1 applied, 0 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-observable-file.md b/tests/fixtures/adr/golden/import-apply-observable-file.md new file mode 100644 index 00000000..883dfac3 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-observable-file.md @@ -0,0 +1,38 @@ +path: docs/architecture/system/ADR-110-old-v0-record.md +--- +contract: adr/v1 +kind: decision +verb: add +capability: ingest +superseded_by: + - ADR-108 +basis: + - evidence: migrated from v0 +agent: + name: Claude + model: unrecorded +observable: + - the record lints clean +status: accepted +date: 2025-01-01 +deciders: + - developer +imported: + from: docs/architecture/system/ADR-110-old-v0-record.md + format: v0 + status: Accepted +--- + +# ADR-110: An unmigrated v0 record + +## Summary + +Context for the record. + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/import-apply-observable.out b/tests/fixtures/adr/golden/import-apply-observable.out new file mode 100644 index 00000000..de9a9fc9 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-observable.out @@ -0,0 +1,3 @@ +Applied: docs/architecture/.import/ADR-110.yaml -> docs/architecture/system/ADR-110-old-v0-record.md (lint: 0 errors, 0 warnings) +Import: 1 applied, 0 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-open-status.txt b/tests/fixtures/adr/golden/import-apply-open-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/import-apply-open.out b/tests/fixtures/adr/golden/import-apply-open.out new file mode 100644 index 00000000..9796916f --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-open.out @@ -0,0 +1,6 @@ +Applied: docs/architecture/.import/ADR-101.yaml -> docs/architecture/system/ADR-101-ingest.md, unchanged (lint: 0 errors, 0 warnings) +Skipped: docs/architecture/.import/ADR-110.yaml: 4 open todo item(s): verb, capability, basis, agent.name +Skipped: docs/architecture/.import/ADR-115.yaml: 6 open todo item(s): verb, capability, basis, agent.name, preamble, unmapped +Skipped: docs/architecture/.import/ADR-116.yaml: 5 open todo item(s): verb, capability, basis, agent.name, status note +Import: 1 applied, 3 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-out-of-range.out b/tests/fixtures/adr/golden/import-apply-out-of-range.out new file mode 100644 index 00000000..0ad951be --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-out-of-range.out @@ -0,0 +1,3 @@ +Refused: docs/architecture/.import/ADR-042.yaml: ADR-042 is outside the system range 100-199 +Import: 0 applied, 0 skipped, 1 refused +[exit 1] diff --git a/tests/fixtures/adr/golden/import-apply-partial-110.md b/tests/fixtures/adr/golden/import-apply-partial-110.md new file mode 100644 index 00000000..c84cbee2 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-partial-110.md @@ -0,0 +1,35 @@ +path: docs/architecture/system/ADR-110-old-v0-record.md +--- +contract: adr/v1 +kind: decision +verb: ~ +capability: ~ +superseded_by: + - ADR-108 +basis: [] +agent: + name: ~ + model: unrecorded +status: accepted +date: 2025-01-01 +deciders: + - developer +imported: + from: docs/architecture/system/ADR-110-old-v0-record.md + format: v0 + status: Accepted +--- + +# ADR-110: An unmigrated v0 record + +## Summary + +Context for the record. + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/import-apply-partial-115.md b/tests/fixtures/adr/golden/import-apply-partial-115.md new file mode 100644 index 00000000..f21ed6b8 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-partial-115.md @@ -0,0 +1,33 @@ +path: docs/architecture/system/ADR-115-nightly-ingest-window.md +--- +contract: adr/v1 +kind: decision +verb: ~ +capability: ~ +basis: [] +agent: + name: ~ + model: unrecorded +status: accepted +date: 2025-04-01 +deciders: + - developer +imported: + from: docs/architecture/system/ADR-115-nightly-ingest-window.md + format: v0 + status: Accepted + unmapped: + revised: 2025-04-02 +--- + +# ADR-115: Nightly ingest window + +> Moved here from the ops wiki. + +## Context + +Ingest ran nightly. + +## Decision + +Ingest runs in a nightly window. diff --git a/tests/fixtures/adr/golden/import-apply-partial-status.txt b/tests/fixtures/adr/golden/import-apply-partial-status.txt new file mode 100644 index 00000000..df00efc1 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-partial-status.txt @@ -0,0 +1,2 @@ +M docs/architecture/system/ADR-110-old-v0-record.md +M docs/architecture/system/ADR-115-nightly-ingest-window.md diff --git a/tests/fixtures/adr/golden/import-apply-partial.out b/tests/fixtures/adr/golden/import-apply-partial.out new file mode 100644 index 00000000..75d06c1d --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-partial.out @@ -0,0 +1,5 @@ +Applied: docs/architecture/.import/ADR-110.yaml -> docs/architecture/system/ADR-110-old-v0-record.md, 4 todo left (lint: 4 errors, 0 warnings) +Applied: docs/architecture/.import/ADR-115.yaml -> docs/architecture/system/ADR-115-nightly-ingest-window.md, 6 todo left (lint: 4 errors, 1 warnings) +Skipped: docs/architecture/.import/ADR-116.yaml: 1 todo item(s) --partial does not write past: status note +Import: 2 applied, 1 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-status.out b/tests/fixtures/adr/golden/import-apply-status.out new file mode 100644 index 00000000..1d97061c --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-status.out @@ -0,0 +1,4 @@ +Skipped: docs/architecture/.import/ADR-117.yaml: 1 todo item(s) --partial does not write past: status +Skipped: docs/architecture/.import/ADR-118.yaml: 1 todo item(s) --partial does not write past: status +Import: 0 applied, 2 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-uncommitted-forced.out b/tests/fixtures/adr/golden/import-apply-uncommitted-forced.out new file mode 100644 index 00000000..7294de11 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-uncommitted-forced.out @@ -0,0 +1,3 @@ +Applied: docs/architecture/.import/ADR-110.yaml -> docs/architecture/system/ADR-110-old-v0-record.md, 4 todo left (lint: 4 errors, 0 warnings) +Import: 1 applied, 0 skipped, 0 refused +[exit 0] diff --git a/tests/fixtures/adr/golden/import-apply-uncommitted.out b/tests/fixtures/adr/golden/import-apply-uncommitted.out new file mode 100644 index 00000000..6f288e53 --- /dev/null +++ b/tests/fixtures/adr/golden/import-apply-uncommitted.out @@ -0,0 +1,3 @@ +Refused: docs/architecture/.import/ADR-110.yaml: docs/architecture/system/ADR-110-old-v0-record.md has uncommitted changes; commit them, or --force to overwrite +Import: 0 applied, 0 skipped, 1 refused +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-frozen-accepted.out b/tests/fixtures/adr/golden/import-lint-completed.out similarity index 100% rename from tests/fixtures/adr/golden/v1-frozen-accepted.out rename to tests/fixtures/adr/golden/import-lint-completed.out diff --git a/tests/fixtures/adr/golden/import-lint-no-summary.out b/tests/fixtures/adr/golden/import-lint-no-summary.out new file mode 100644 index 00000000..8f1d8f41 --- /dev/null +++ b/tests/fixtures/adr/golden/import-lint-no-summary.out @@ -0,0 +1,24 @@ + +Scanned: 1 ADRs + +Status distribution: + accepted: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-115-nightly-ingest-window.md + ❌ a decision record requires a verb (add, cut, change, retire, constrain) + ❌ a decision record requires 'capability' + ❌ a decision record requires 'basis' + ❌ agent: records 'name' + ⚠️ imported record has no '## Summary' yet (ADR-306 §4) + +════════════════════════════════════════════════════════════ +Summary: 4 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/import-rescan-forced.out b/tests/fixtures/adr/golden/import-rescan-forced.out new file mode 100644 index 00000000..95b5c107 --- /dev/null +++ b/tests/fixtures/adr/golden/import-rescan-forced.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo; replaced an edited sheet, its edits are gone) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-rescan-kept.out b/tests/fixtures/adr/golden/import-rescan-kept.out new file mode 100644 index 00000000..22fd1148 --- /dev/null +++ b/tests/fixtures/adr/golden/import-rescan-kept.out @@ -0,0 +1,3 @@ +Skipped: docs/architecture/system/ADR-110-old-v0-record.md: docs/architecture/.import/ADR-110.yaml differs from a fresh scan (edited, or the source changed); --force overwrites it and discards its edits +Scan: 0 sheet(s) written, 1 skipped +[exit 1] diff --git a/tests/fixtures/adr/golden/import-scan-bad.out b/tests/fixtures/adr/golden/import-scan-bad.out new file mode 100644 index 00000000..959c7823 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-bad.out @@ -0,0 +1,6 @@ +Scanned: docs/architecture/system/ADR-101-ingest.md -> docs/architecture/.import/ADR-101.yaml (v1, 0 todo) +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-115-nightly-ingest-window.md -> docs/architecture/.import/ADR-115.yaml (v0, 6 todo) +Scanned: docs/architecture/system/ADR-116-hourly-ingest.md -> docs/architecture/.import/ADR-116.yaml (v0, 5 todo) +Scan: 4 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-bom.out b/tests/fixtures/adr/golden/import-scan-bom.out new file mode 100644 index 00000000..2be79958 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-bom.out @@ -0,0 +1,3 @@ +Skipped: docs/architecture/system/ADR-117-with-a-bom.md: the file starts with a UTF-8 byte order mark; remove it before scanning +Scan: 0 sheet(s) written, 1 skipped +[exit 1] diff --git a/tests/fixtures/adr/golden/import-scan-broad.out b/tests/fixtures/adr/golden/import-scan-broad.out new file mode 100644 index 00000000..fb99b3eb --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-broad.out @@ -0,0 +1,14 @@ +Skipped: notes/.drafts: hidden directory, not entered +Skipped: notes/.notes.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/archive/README.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/decisions/0001-use-postgres.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/decisions/README.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/sub/INDEX.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: 1 file with no extension: not Markdown +Skipped: 3 .png files: not Markdown +Skipped: 1 .txt file: not Markdown +Skipped: 1 .yaml file: not Markdown +Skipped: notes/decisions/ADR-002-inline.md: no YAML frontmatter; import reads structured records (ADR-306 §2) +Scanned: notes/decisions/ADR-003.md -> docs/architecture/.import/ADR-003.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 13 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-changed.out b/tests/fixtures/adr/golden/import-scan-changed.out new file mode 100644 index 00000000..40cc9ec0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-changed.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-complete.out b/tests/fixtures/adr/golden/import-scan-complete.out new file mode 100644 index 00000000..40cc9ec0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-complete.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-003.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-003.yaml new file mode 100644 index 00000000..a87b93e5 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-003.yaml @@ -0,0 +1,49 @@ +path: docs/architecture/.import/ADR-003.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: notes/decisions/ADR-003.md + format: v0 + sha256: 78ddea958a25abe35b25a000aafc88c49de225bdcbb92aa65a6ee5a732bb63f9 +target: + number: '3' + domain: legacy + title: No slug +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2026-01-15 + deciders: + - a + related: [] +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: number range + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.related: frontmatter related +unmapped: {} +body: | + ## Context + x diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-005.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-005.yaml new file mode 100644 index 00000000..6a833d90 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-005.yaml @@ -0,0 +1,62 @@ +path: docs/architecture/.import/ADR-005.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/legacy/ADR-005-legacy-storage.md + format: v0 + sha256: 830a504e9de8be0c3aedbaa254ae85b884f5405c60a21941e60cc4d5b7de21c5 +target: + number: '5' + domain: legacy + title: Legacy storage format +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + superseded_by: + - ADR-101 + basis: [] + agent: + name: ~ + model: unrecorded + status: superseded + date: 2025-01-10 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder legacy + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Superseded' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.superseded_by: frontmatter superseded_by +unmapped: {} +body: | + ## Context + + A pre-domain decision, since replaced. + + ## Decision + + The decision for ADR-005: Legacy storage format. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-012.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-012.yaml new file mode 100644 index 00000000..f1d41a1b --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-012.yaml @@ -0,0 +1,59 @@ +path: docs/architecture/.import/ADR-012.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/legacy/ADR-012-plain-logging.md + format: v0 + sha256: c2e846dd0d87c89d9d0f24bc4514305bf8de95e1149b776e3f7e966e90bc48a6 +target: + number: '12' + domain: legacy + title: Plain-text logging +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-02-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder legacy + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + A legacy decision still in force. + + ## Decision + + The decision for ADR-012: Plain-text logging. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.1.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.1.yaml new file mode 100644 index 00000000..9d2984c3 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.1.yaml @@ -0,0 +1,53 @@ +path: docs/architecture/.import/ADR-101.1.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-101.1-storage-migration.md + format: v0 + sha256: 9bde84f5aa9bf8721d1f90551ba894bbbf0ca763738e4797c9c59e100ea8ef19 +target: + number: '101.1' + domain: system + title: Storage migration +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-06-20 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + A decimal sub-part of ADR-101. + + ## Decision + + The decision for ADR-101.1: Storage migration. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.yaml new file mode 100644 index 00000000..3fb172b2 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-101.yaml @@ -0,0 +1,65 @@ +path: docs/architecture/.import/ADR-101.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-101-structured-storage.md + format: v0 + sha256: 5e949a1d29573a66bbfbbdcbe112f25311fce67c9ed53560a1e94bb068644326 +target: + number: '101' + domain: system + title: Structured storage +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + supersedes: + - ADR-005 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-06-01 + deciders: + - developer + - agent + related: + - 102 +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.related: frontmatter related + record.supersedes: frontmatter supersedes +unmapped: {} +body: | + ## Context + + Replaces the legacy storage format. + + ## Decision + + The decision for ADR-101: Structured storage. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-102.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-102.yaml new file mode 100644 index 00000000..45a12191 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-102.yaml @@ -0,0 +1,62 @@ +path: docs/architecture/.import/ADR-102.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-102-hook-ordering.md + format: v0 + sha256: 289375931298167da605ebce441d9378a0216597464298440c64ffc46719f8b7 +target: + number: '102' + domain: system + title: Hook ordering +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + superseded_by: + - ADR-104#2 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-06-15 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.superseded_by: frontmatter superseded_by +unmapped: {} +body: | + ## Context + + Partially superseded by section reference. + + ## Decision + + The decision for ADR-102: Hook ordering. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-103.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-103.yaml new file mode 100644 index 00000000..18432815 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-103.yaml @@ -0,0 +1,59 @@ +path: docs/architecture/.import/ADR-103.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-103-cache-layer.md + format: v0 + sha256: 749d6dfb413359c4cd382b60a47bb0b71663b07518e7c309812912f51ed73636 +target: + number: '103' + domain: system + title: Cache layer +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: proposed + date: 2025-07-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Proposed' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + A proposal not yet accepted. + + ## Decision + + The decision for ADR-103: Cache layer. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-104.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-104.yaml new file mode 100644 index 00000000..e573ed72 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-104.yaml @@ -0,0 +1,66 @@ +path: docs/architecture/.import/ADR-104.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-104-hook-priorities.md + format: v0 + sha256: b188a9282e21030785b3079d66039ecac3a78197e88169be7c9233dd785706cc +target: + number: '104' + domain: system + title: Hook priorities +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + supersedes: + - ADR-102#2 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-08-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.supersedes: frontmatter supersedes +unmapped: {} +body: | + ## Context + + Replaces one section of ADR-102. + + ## Decision + + The decision for ADR-104: Hook priorities. + + ## Consequences + + ### Positive + + - It works. + + ## 2. Priority bands + + Hooks run in three bands. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-105.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-105.yaml new file mode 100644 index 00000000..1ef413e0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-105.yaml @@ -0,0 +1,60 @@ +path: docs/architecture/.import/ADR-105.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-105-telemetry.md + format: v0 + sha256: 4e9ad147766a709dff5ff17c4ab420886f7cacbf2b74d937536047e167ff534a +target: + number: '105' + domain: system + title: Telemetry export +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-08-10 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name + - 'status note: Deprecated with no successor maps to accepted with a note that the decision is historical (ADR-304 §7); add the note to the body' +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Deprecated, with no successor (ADR-304 §7)' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Deprecated with nothing replacing it. + + ## Decision + + The decision for ADR-105: Telemetry export. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-106.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-106.yaml new file mode 100644 index 00000000..f98e905a --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-106.yaml @@ -0,0 +1,62 @@ +path: docs/architecture/.import/ADR-106.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-106-one-sided.md + format: v0 + sha256: ebeb4d2c085c416f7bcf0cdf9f433ca4c7d00babb962b2f6771fb60b216f5897 +target: + number: '106' + domain: system + title: One-sided supersession +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + supersedes: + - ADR-012 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-09-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.supersedes: frontmatter supersedes +unmapped: {} +body: | + ## Context + + Declares supersession that ADR-012 does not reciprocate. + + ## Decision + + The decision for ADR-106: One-sided supersession. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-200.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-200.yaml new file mode 100644 index 00000000..745dddcd --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-200.yaml @@ -0,0 +1,53 @@ +path: docs/architecture/.import/ADR-200.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/runbooks/ADR-200-restart-procedure.md + format: v0 + sha256: c269e6d03fa2eeed8b36fb366124d4dcc8220b53f38ba95334353d834295b4b8 +target: + number: '200' + domain: ops + title: Restart procedure +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-04-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder runbooks + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Lives in the second folder of a list-valued domain. + + ## Decision + + The decision for ADR-200: Restart procedure. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-300.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-300.yaml new file mode 100644 index 00000000..ad125d4d --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-300.yaml @@ -0,0 +1,59 @@ +path: docs/architecture/.import/ADR-300.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/documentation/ADR-300-doc-structure.md + format: v0 + sha256: 936ae33576633b0120d147dad661e0df50aa2f462ad66869f4e56700353bdeb7 +target: + number: '300' + domain: docs + title: Documentation structure +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-05-01 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder documentation + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Where documentation lives. + + ## Decision + + The decision for ADR-300: Documentation structure. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-301.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-301.yaml new file mode 100644 index 00000000..e0f985f6 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-301.yaml @@ -0,0 +1,59 @@ +path: docs/architecture/.import/ADR-301.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/documentation/ADR-301-draft-guide.md + format: v0 + sha256: 2186faab4d9ee1b286a7ae27aae803af58909c4b4f4c1208f3a58f6ce6bca32a +target: + number: '301' + domain: docs + title: Draft style guide +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: proposed + date: 2025-09-10 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder documentation + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Draft' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + A draft. + + ## Decision + + The decision for ADR-301: Draft style guide. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-ADR-303.yaml b/tests/fixtures/adr/golden/import-scan-corpus-ADR-303.yaml new file mode 100644 index 00000000..258dd3fb --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-ADR-303.yaml @@ -0,0 +1,62 @@ +path: docs/architecture/.import/ADR-303.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/documentation/ADR-303-dangling.md + format: v0 + sha256: 764672969aebec723f87be38b80fa653e5b45713282259b13bc115860732cd07 +target: + number: '303' + domain: docs + title: Dangling reference +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + supersedes: + - ADR-199 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-09-15 + deciders: + - developer + - agent +summary: ~ +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder documentation + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.supersedes: frontmatter supersedes +unmapped: {} +body: | + ## Context + + Supersedes an ADR that does not exist. + + ## Decision + + The decision for ADR-303: Dangling reference. + + ## Consequences + + ### Positive + + - It works. diff --git a/tests/fixtures/adr/golden/import-scan-corpus-gitignore b/tests/fixtures/adr/golden/import-scan-corpus-gitignore new file mode 100644 index 00000000..27058688 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus-gitignore @@ -0,0 +1,2 @@ +path: docs/architecture/.import/.gitignore +* diff --git a/tests/fixtures/adr/golden/import-scan-corpus-status.txt b/tests/fixtures/adr/golden/import-scan-corpus-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/import-scan-corpus.out b/tests/fixtures/adr/golden/import-scan-corpus.out new file mode 100644 index 00000000..4da24513 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-corpus.out @@ -0,0 +1,17 @@ +Note: 1 archived record(s) left out; name a file to scan it +Scanned: docs/architecture/documentation/ADR-300-doc-structure.md -> docs/architecture/.import/ADR-300.yaml (v0, 4 todo) +Scanned: docs/architecture/documentation/ADR-301-draft-guide.md -> docs/architecture/.import/ADR-301.yaml (v0, 4 todo) +Skipped: docs/architecture/documentation/ADR-302-inline-metadata.md: no YAML frontmatter; import reads structured records (ADR-306 §2) +Scanned: docs/architecture/documentation/ADR-303-dangling.md -> docs/architecture/.import/ADR-303.yaml (v0, 4 todo) +Scanned: docs/architecture/legacy/ADR-005-legacy-storage.md -> docs/architecture/.import/ADR-005.yaml (v0, 4 todo) +Scanned: docs/architecture/legacy/ADR-012-plain-logging.md -> docs/architecture/.import/ADR-012.yaml (v0, 4 todo) +Scanned: docs/architecture/runbooks/ADR-200-restart-procedure.md -> docs/architecture/.import/ADR-200.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-101-structured-storage.md -> docs/architecture/.import/ADR-101.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-101.1-storage-migration.md -> docs/architecture/.import/ADR-101.1.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-102-hook-ordering.md -> docs/architecture/.import/ADR-102.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-103-cache-layer.md -> docs/architecture/.import/ADR-103.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-104-hook-priorities.md -> docs/architecture/.import/ADR-104.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-105-telemetry.md -> docs/architecture/.import/ADR-105.yaml (v0, 5 todo) +Scanned: docs/architecture/system/ADR-106-one-sided.md -> docs/architecture/.import/ADR-106.yaml (v0, 4 todo) +Scan: 13 sheet(s) written, 1 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-dryrun.out b/tests/fixtures/adr/golden/import-scan-dryrun.out new file mode 100644 index 00000000..40cc9ec0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-dryrun.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-flat.out b/tests/fixtures/adr/golden/import-scan-flat.out new file mode 100644 index 00000000..621a3589 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-flat.out @@ -0,0 +1,6 @@ +Skipped: notes/decisions/0001-use-postgres.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/decisions/README.md: not named ADR-NNN.md or ADR-NNN-<slug>.md +Skipped: notes/decisions/ADR-002-inline.md: no YAML frontmatter; import reads structured records (ADR-306 §2) +Scanned: notes/decisions/ADR-003.md -> docs/architecture/.import/ADR-003.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 3 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-foreign.out b/tests/fixtures/adr/golden/import-scan-foreign.out new file mode 100644 index 00000000..8e21c8a8 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-foreign.out @@ -0,0 +1,3 @@ +Scanned: <WORK>/outside/ADR-042-foreign-record.md -> docs/architecture/.import/ADR-042.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-observable.out b/tests/fixtures/adr/golden/import-scan-observable.out new file mode 100644 index 00000000..40cc9ec0 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-observable.out @@ -0,0 +1,3 @@ +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scan: 1 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-status-ADR-117.yaml b/tests/fixtures/adr/golden/import-scan-status-ADR-117.yaml new file mode 100644 index 00000000..43c3aa95 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-status-ADR-117.yaml @@ -0,0 +1,48 @@ +path: docs/architecture/.import/ADR-117.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-117-work-in-progress.md + format: v0 + sha256: 16414c4ee208ce121714f95185dc09a51b3eafc58823c9c873e04f5734a61ee0 +target: + number: '117' + domain: system + title: Work in progress +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: ~ + date: 2025-04-01 + deciders: + - developer +summary: ~ +todo: + - verb + - capability + - basis + - agent.name + - 'status: ''WIP pending review'' has no v1 mapping; set one (ADR-304 §7)' +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Text. diff --git a/tests/fixtures/adr/golden/import-scan-status-ADR-118.yaml b/tests/fixtures/adr/golden/import-scan-status-ADR-118.yaml new file mode 100644 index 00000000..8e33cbef --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-status-ADR-118.yaml @@ -0,0 +1,48 @@ +path: docs/architecture/.import/ADR-118.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-118-no-status.md + format: v0 + sha256: 90c4515ef1c5bcef48b478a1b5d3f356931d0494282113af81b17334f485e85d +target: + number: '118' + domain: system + title: No status +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: ~ + date: 2025-04-01 + deciders: + - developer +summary: ~ +todo: + - verb + - capability + - basis + - agent.name + - 'status: the source has no status; set one (ADR-304 §7)' +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Text. diff --git a/tests/fixtures/adr/golden/import-scan-status.out b/tests/fixtures/adr/golden/import-scan-status.out new file mode 100644 index 00000000..5d9423c6 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-status.out @@ -0,0 +1,4 @@ +Scanned: docs/architecture/system/ADR-117-work-in-progress.md -> docs/architecture/.import/ADR-117.yaml (v0, 5 todo) +Scanned: docs/architecture/system/ADR-118-no-status.md -> docs/architecture/.import/ADR-118.yaml (v0, 5 todo) +Scan: 2 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/import-scan-v1-ADR-101.yaml b/tests/fixtures/adr/golden/import-scan-v1-ADR-101.yaml new file mode 100644 index 00000000..01f3580f --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-v1-ADR-101.yaml @@ -0,0 +1,49 @@ +path: docs/architecture/.import/ADR-101.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-101-ingest.md + format: v1 + sha256: ffe57ae56e9cb8c103ade235a59e6aad873b5fc8f0d599235ad29c9214064a5f +target: + number: '101' + domain: system + title: Add ingestion +record: + contract: adr/v1 + kind: decision + verb: add + capability: ingest + basis: + - evidence: fixture measurement + agent: + name: Claude + model: fixture-model + status: accepted + date: 2025-05-02 + deciders: + - developer + - agent +summary: | + - **Decided:** the decision in plain terms. + - **Trades away:** what it gives up. + - **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. + - **Inversion:** one end, the other end; is the middle right? +todo: [] +candidates: {} +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + summary: '## Summary section' + record: frontmatter (adr/v1) +unmapped: {} +body: | + ## 1. Decision + + The decision. + + ## 2. Consequences + + They follow. diff --git a/tests/fixtures/adr/golden/import-scan-v1-ADR-110.yaml b/tests/fixtures/adr/golden/import-scan-v1-ADR-110.yaml new file mode 100644 index 00000000..6e405f42 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-v1-ADR-110.yaml @@ -0,0 +1,57 @@ +path: docs/architecture/.import/ADR-110.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-110-old-v0-record.md + format: v0 + sha256: 5fbfe7e263da32fb0a7d49dc620e989feb6c9283465df2f66c2d9632ab7969a4 +target: + number: '110' + domain: system + title: An unmigrated v0 record +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + superseded_by: + - ADR-108 + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-01-01 + deciders: + - developer +summary: | + Context for the record. +todo: + - verb + - capability + - basis + - agent.name +candidates: + capability: [] +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + summary: '## Summary section' + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) + record.superseded_by: frontmatter superseded_by +unmapped: {} +body: | + ## 1. Decision + + The decision. + + ## 2. Consequences + + They follow. diff --git a/tests/fixtures/adr/golden/import-scan-v1-ADR-115.yaml b/tests/fixtures/adr/golden/import-scan-v1-ADR-115.yaml new file mode 100644 index 00000000..6b6a5465 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-v1-ADR-115.yaml @@ -0,0 +1,59 @@ +path: docs/architecture/.import/ADR-115.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-115-nightly-ingest-window.md + format: v0 + sha256: 6fc5e998597b02af3cc0481cfb32ed884f500017b69fec37348401813d865cb3 +target: + number: '115' + domain: system + title: Nightly ingest window +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-04-01 + deciders: + - developer +summary: ~ +todo: + - verb + - capability + - basis + - agent.name + - 'preamble: the text between the frontmatter and the H1 moved into the body, right after the H1; check where it belongs' + - 'unmapped: revised has no v1 field; apply keeps it under imported.unmapped. Move it into the record, or leave it there' +candidates: + capability: + - ingest +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + body: text above the H1, then the text after it + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Accepted' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: + revised: 2025-04-02 +body: | + > Moved here from the ops wiki. + + ## Context + + Ingest ran nightly. + + ## Decision + + Ingest runs in a nightly window. diff --git a/tests/fixtures/adr/golden/import-scan-v1-ADR-116.yaml b/tests/fixtures/adr/golden/import-scan-v1-ADR-116.yaml new file mode 100644 index 00000000..d0f96d56 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-v1-ADR-116.yaml @@ -0,0 +1,50 @@ +path: docs/architecture/.import/ADR-116.yaml +# adr import sheet (ADR-306). Fill or clear each todo item, then run +# `adr import apply`. Sheets are working files; only records are committed. +sheet: adr-import/v1 +source: + path: docs/architecture/system/ADR-116-hourly-ingest.md + format: v0 + sha256: 1b1c93891a8152b926668171c667bfecc5613257a7cf68ce8283aecd2ccb8744 +target: + number: '116' + domain: system + title: Hourly ingest +record: + contract: adr/v1 + kind: decision + verb: ~ + capability: ~ + basis: [] + agent: + name: ~ + model: unrecorded + status: accepted + date: 2025-04-05 + deciders: + - developer +summary: ~ +todo: + - verb + - capability + - basis + - agent.name + - 'status note: Deprecated with no successor maps to accepted with a note that the decision is historical (ADR-304 §7); add the note to the body' +candidates: + capability: + - ingest +provenance: + target.number: file name + target.domain: folder system + target.title: H1 + record.contract: v0 reader + record.kind: 'v0 reader: every v0 record is a decision' + record.agent.model: 'v0 reader: the model was not recorded' + record.status: 'frontmatter status: Deprecated, with no successor (ADR-304 §7)' + record.date: frontmatter date + record.deciders: frontmatter deciders (never an operator basis, ADR-304 §7) +unmapped: {} +body: | + ## Context + + Ingest ran hourly. diff --git a/tests/fixtures/adr/golden/import-scan-v1.out b/tests/fixtures/adr/golden/import-scan-v1.out new file mode 100644 index 00000000..959c7823 --- /dev/null +++ b/tests/fixtures/adr/golden/import-scan-v1.out @@ -0,0 +1,6 @@ +Scanned: docs/architecture/system/ADR-101-ingest.md -> docs/architecture/.import/ADR-101.yaml (v1, 0 todo) +Scanned: docs/architecture/system/ADR-110-old-v0-record.md -> docs/architecture/.import/ADR-110.yaml (v0, 4 todo) +Scanned: docs/architecture/system/ADR-115-nightly-ingest-window.md -> docs/architecture/.import/ADR-115.yaml (v0, 6 todo) +Scanned: docs/architecture/system/ADR-116-hourly-ingest.md -> docs/architecture/.import/ADR-116.yaml (v0, 5 todo) +Scan: 4 sheet(s) written, 0 skipped +[exit 0] diff --git a/tests/fixtures/adr/golden/index-new-file.md b/tests/fixtures/adr/golden/index-new-file.md index 72654f85..5e9296c5 100644 --- a/tests/fixtures/adr/golden/index-new-file.md +++ b/tests/fixtures/adr/golden/index-new-file.md @@ -1,7 +1,7 @@ path: docs/architecture/INDEX.md -# Architecture Decision Records +# Agent Decision Records -This directory contains Architecture Decision Records (ADRs) for ADR Fixture. +This directory contains Agent Decision Records (ADRs) for ADR Fixture. Each ADR documents a significant architectural decision, its context, and consequences. ## ADR Format diff --git a/tests/fixtures/adr/golden/list-alias-ls.out b/tests/fixtures/adr/golden/list-alias-ls.out index a0d69332..8a9b5bcf 100644 --- a/tests/fixtures/adr/golden/list-alias-ls.out +++ b/tests/fixtures/adr/golden/list-alias-ls.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (14 total) +ADR Fixture — Agent Decision Records (14 total) ======================================================= 📦 ADR-005 Legacy storage format (superseded by ADR-101) ✅ ADR-012 Plain-text logging diff --git a/tests/fixtures/adr/golden/list-all.out b/tests/fixtures/adr/golden/list-all.out index e79ae5cf..7086e5f7 100644 --- a/tests/fixtures/adr/golden/list-all.out +++ b/tests/fixtures/adr/golden/list-all.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (15 total) +ADR Fixture — Agent Decision Records (15 total) ======================================================= 📦 ADR-005 Legacy storage format (superseded by ADR-101) ✅ ADR-012 Plain-text logging diff --git a/tests/fixtures/adr/golden/list-archived.out b/tests/fixtures/adr/golden/list-archived.out index 604c34bb..e2514538 100644 --- a/tests/fixtures/adr/golden/list-archived.out +++ b/tests/fixtures/adr/golden/list-archived.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (1 total) +ADR Fixture — Agent Decision Records (1 total) ======================================================= 📦 ADR-107 Old queue (superseded by ADR-101) diff --git a/tests/fixtures/adr/golden/list-domain-ops.out b/tests/fixtures/adr/golden/list-domain-ops.out index a8bee9fe..ad126a3a 100644 --- a/tests/fixtures/adr/golden/list-domain-ops.out +++ b/tests/fixtures/adr/golden/list-domain-ops.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (1 total) +ADR Fixture — Agent Decision Records (1 total) ======================================================= ✅ ADR-200 Restart procedure diff --git a/tests/fixtures/adr/golden/list-domain-system.out b/tests/fixtures/adr/golden/list-domain-system.out index 5d3f0289..0ab38520 100644 --- a/tests/fixtures/adr/golden/list-domain-system.out +++ b/tests/fixtures/adr/golden/list-domain-system.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (7 total) +ADR Fixture — Agent Decision Records (7 total) ======================================================= ✅ ADR-101 Structured storage ✅ ADR-101.1 Storage migration diff --git a/tests/fixtures/adr/golden/list-group.out b/tests/fixtures/adr/golden/list-group.out index 28512476..f91d6667 100644 --- a/tests/fixtures/adr/golden/list-group.out +++ b/tests/fixtures/adr/golden/list-group.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (14 total) +ADR Fixture — Agent Decision Records (14 total) ======================================================= ## System (system) diff --git a/tests/fixtures/adr/golden/list-status-accepted.out b/tests/fixtures/adr/golden/list-status-accepted.out index ed5ebe72..a428c4a4 100644 --- a/tests/fixtures/adr/golden/list-status-accepted.out +++ b/tests/fixtures/adr/golden/list-status-accepted.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (9 total) +ADR Fixture — Agent Decision Records (9 total) ======================================================= ✅ ADR-012 Plain-text logging ✅ ADR-101 Structured storage diff --git a/tests/fixtures/adr/golden/list.out b/tests/fixtures/adr/golden/list.out index a0d69332..8a9b5bcf 100644 --- a/tests/fixtures/adr/golden/list.out +++ b/tests/fixtures/adr/golden/list.out @@ -1,5 +1,5 @@ -ADR Fixture — Architecture Decision Records (14 total) +ADR Fixture — Agent Decision Records (14 total) ======================================================= 📦 ADR-005 Legacy storage format (superseded by ADR-101) ✅ ADR-012 Plain-text logging diff --git a/tests/fixtures/adr/golden/record-consider-append-file.md b/tests/fixtures/adr/golden/record-consider-append-file.md new file mode 100644 index 00000000..b14b247b --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-append-file.md @@ -0,0 +1,51 @@ +path: docs/architecture/system/ADR-100-adopt-v1.md +--- +contract: adr/v1 +kind: decision +verb: add +capability: adr +status: accepted +date: 2025-05-01 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - operator: developer + level: guided + said: "agent-ways leads adoption, and its own corpus is the test" + via: session 2025-05-01 + - evidence: fixture triage of 108 records +considered: + - operator: developer + said: "looks good; the capability list is fine for now" + via: PR #1 + covers: [probe-2, inversion] + canary: caught + - operator: developer + said: "ok, as it stands" + via: "PR #2" + covers: [] +concern: + - said: "The capability list may be friction for small repos" + resolve: "Operator confirms the list is acceptable" + answer: {said: "fine for now", via: "PR #1"} + - said: "Section references accept two forms" + resolve: "Pick one canonical form" + withdrawn: "Both forms are unambiguous; no longer a concern" +--- + +# ADR-100: Adopt the adr/v1 contract + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/record-consider-append.out b/tests/fixtures/adr/golden/record-consider-append.out new file mode 100644 index 00000000..b214cde4 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-append.out @@ -0,0 +1,3 @@ +Added considered entry 2 to ADR-100: Adopt the adr/v1 contract +Lint docs/architecture/system/ADR-100-adopt-v1.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-consider-file.md b/tests/fixtures/adr/golden/record-consider-file.md new file mode 100644 index 00000000..e869682d --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-file.md @@ -0,0 +1,31 @@ +path: docs/architecture/system/ADR-115-named-probes.md +--- +contract: adr/v1 +kind: decision +verb: add +capability: adr +status: accepted +date: 2025-05-20 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +considered: + - operator: developer + said: "\"Fine\" (Recommended)" + via: session 2025-05-20, selected from agent-written options + covers: [band-hint, inversion] + canary: caught +--- + +# ADR-115: Named probes + +## Summary + +- **Decided:** the decision. +- **Probes:** *Confident (identity-stable):* a. *Not confident (band-hint):* b. +- **Inversion:** c. + +## 1. Decision + +The decision. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-156-operator-no-considered.md b/tests/fixtures/adr/golden/record-consider-new-file.md similarity index 61% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-156-operator-no-considered.md rename to tests/fixtures/adr/golden/record-consider-new-file.md index 8340cedb..50e3d5b9 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-156-operator-no-considered.md +++ b/tests/fixtures/adr/golden/record-consider-new-file.md @@ -1,21 +1,27 @@ +path: docs/architecture/system/ADR-113-operator-proposed.md --- contract: adr/v1 kind: decision verb: add -capability: adr -date: 2025-06-01 +capability: search +status: proposed +date: 2025-05-13 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} -status: accepted basis: - operator: developer level: directed - said: "ship it" - via: call 2025-06-01 + said: "add search back" + via: call 2025-05-13 + paraphrase: true +considered: + - operator: developer + said: "add it back" + via: call paraphrase: true --- -# ADR-156: Operator-started, accepted without consideration +# ADR-113: Operator-started and still proposed ## Summary diff --git a/tests/fixtures/adr/golden/record-consider-new.out b/tests/fixtures/adr/golden/record-consider-new.out new file mode 100644 index 00000000..0b4e7659 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-new.out @@ -0,0 +1,3 @@ +Added considered entry 1 to ADR-113: Operator-started and still proposed +Lint docs/architecture/system/ADR-113-operator-proposed.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-consider-no-names.out b/tests/fixtures/adr/golden/record-consider-no-names.out new file mode 100644 index 00000000..1c7f876f --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-no-names.out @@ -0,0 +1,2 @@ +Error: ADR-101 has no probe 'edge-case'; its Summary names no probes. A probe is named as *Confident (name):* or *Not confident (name):*. It can cover: inversion. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-consider-no-said.out b/tests/fixtures/adr/golden/record-consider-no-said.out new file mode 100644 index 00000000..12e71e61 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-no-said.out @@ -0,0 +1,2 @@ +Error: --said is required: what the operator said, verbatim. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-consider-no-via.out b/tests/fixtures/adr/golden/record-consider-no-via.out new file mode 100644 index 00000000..fb6331d2 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-no-via.out @@ -0,0 +1,2 @@ +Error: --via is required: the channel it was said in, such as a PR or a session. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-consider-status.txt b/tests/fixtures/adr/golden/record-consider-status.txt new file mode 100644 index 00000000..55630be3 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-status.txt @@ -0,0 +1,3 @@ +M docs/architecture/system/ADR-100-adopt-v1.md +M docs/architecture/system/ADR-113-operator-proposed.md +M docs/architecture/system/ADR-115-named-probes.md diff --git a/tests/fixtures/adr/golden/record-consider-unknown-probe.out b/tests/fixtures/adr/golden/record-consider-unknown-probe.out new file mode 100644 index 00000000..2d0ade3d --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-unknown-probe.out @@ -0,0 +1,2 @@ +Error: ADR-115 has no probe 'no-such-probe'. Names it can cover: identity-stable, band-hint, inversion. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-consider-v0.out b/tests/fixtures/adr/golden/record-consider-v0.out new file mode 100644 index 00000000..b0fecd78 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider-v0.out @@ -0,0 +1,2 @@ +Error: ADR-110 is not an adr/v1 record; `adr consider` edits v1 fields (ADR-304 §7). +[exit 1] diff --git a/tests/fixtures/adr/golden/record-consider.out b/tests/fixtures/adr/golden/record-consider.out new file mode 100644 index 00000000..3ed3c304 --- /dev/null +++ b/tests/fixtures/adr/golden/record-consider.out @@ -0,0 +1,3 @@ +Added considered entry 1 to ADR-115: Named probes +Lint docs/architecture/system/ADR-115-named-probes.md: clean +[exit 0] diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-129-retire-no-targets.md b/tests/fixtures/adr/golden/record-enact-again-file.md similarity index 72% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-129-retire-no-targets.md rename to tests/fixtures/adr/golden/record-enact-again-file.md index 4268a832..172c618c 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-129-retire-no-targets.md +++ b/tests/fixtures/adr/golden/record-enact-again-file.md @@ -1,17 +1,20 @@ +path: docs/architecture/system/ADR-105-retire-legacy-ingest.md --- contract: adr/v1 kind: decision verb: retire capability: ingest +targets: [cli:ingest-legacy, route:/v1/upload] status: accepted -date: 2025-06-01 +enacted: "3f9c2a1dead" +date: 2025-05-06 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} basis: - evidence: fixture measurement --- -# ADR-129: Retire without targets +# ADR-105: Retire the legacy ingest surface ## Summary diff --git a/tests/fixtures/adr/golden/record-enact-again.out b/tests/fixtures/adr/golden/record-enact-again.out new file mode 100644 index 00000000..ac27903d --- /dev/null +++ b/tests/fixtures/adr/golden/record-enact-again.out @@ -0,0 +1,3 @@ +Enacted ADR-105 at 3f9c2a1dead (was 3f9c2a1): Retire the legacy ingest surface +Lint docs/architecture/system/ADR-105-retire-legacy-ingest.md: clean +[exit 0] diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-171-order-b.md b/tests/fixtures/adr/golden/record-enact-file.md similarity index 69% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-171-order-b.md rename to tests/fixtures/adr/golden/record-enact-file.md index 519ee103..00460135 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-171-order-b.md +++ b/tests/fixtures/adr/golden/record-enact-file.md @@ -1,17 +1,19 @@ +path: docs/architecture/system/ADR-111-cut-search.md --- contract: adr/v1 kind: decision -verb: add -capability: adr -date: 2025-06-01 +verb: cut +capability: search +enacted: "abc1234" +status: accepted +date: 2025-05-11 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} -status: accepted basis: - - precedent: ADR-170 + - evidence: usage data shows no queries --- -# ADR-171: Order pair partner, grounded through ADR-170 +# ADR-111: Cut search ## Summary @@ -23,7 +25,3 @@ basis: ## 1. Decision The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/golden/record-enact-not-hash.out b/tests/fixtures/adr/golden/record-enact-not-hash.out new file mode 100644 index 00000000..2461775c --- /dev/null +++ b/tests/fixtures/adr/golden/record-enact-not-hash.out @@ -0,0 +1,2 @@ +Error: 'HEAD~1' is not a commit hash (7 to 40 hex digits). +[exit 1] diff --git a/tests/fixtures/adr/golden/record-enact-wrong-status.out b/tests/fixtures/adr/golden/record-enact-wrong-status.out new file mode 100644 index 00000000..802771fc --- /dev/null +++ b/tests/fixtures/adr/golden/record-enact-wrong-status.out @@ -0,0 +1,2 @@ +Error: ADR-111 is superseded; enacted marks an accepted decision done. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-enact-wrong-verb.out b/tests/fixtures/adr/golden/record-enact-wrong-verb.out new file mode 100644 index 00000000..2e54464c --- /dev/null +++ b/tests/fixtures/adr/golden/record-enact-wrong-verb.out @@ -0,0 +1,2 @@ +Error: ADR-101 is an add record; enacted belongs to a cut or retire decision (ADR-304 §5). +[exit 1] diff --git a/tests/fixtures/adr/golden/record-enact.out b/tests/fixtures/adr/golden/record-enact.out new file mode 100644 index 00000000..2c418b7d --- /dev/null +++ b/tests/fixtures/adr/golden/record-enact.out @@ -0,0 +1,3 @@ +Enacted ADR-111 at abc1234: Cut search +Lint docs/architecture/system/ADR-111-cut-search.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-capability.out b/tests/fixtures/adr/golden/record-list-capability.out new file mode 100644 index 00000000..c3470a71 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-capability.out @@ -0,0 +1,8 @@ + +ADR v1 Fixture — Agent Decision Records (2 total) +======================================================= + ✅ ADR-100 Adopt the adr/v1 contract + ✅ ADR-112 Operator-authored constraint + +Total: 2 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-edge.out b/tests/fixtures/adr/golden/record-list-edge.out new file mode 100644 index 00000000..db5ab9aa --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-edge.out @@ -0,0 +1,8 @@ + +ADR v1 Fixture — Agent Decision Records (2 total) +======================================================= + ✅ ADR-103 Batch ingestion (partially superseded by ADR-109) + 💡 ADR-107 Ingest notes, by slug reference + +Total: 2 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-field.out b/tests/fixtures/adr/golden/record-list-field.out new file mode 100644 index 00000000..ad075fdc --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-field.out @@ -0,0 +1,15 @@ + +ADR v1 Fixture — Agent Decision Records (9 total) +======================================================= + ✅ ADR-101 Add ingestion + ✅ ADR-102 How ingestion works + ✅ ADR-103 Batch ingestion (partially superseded by ADR-109) + ✅ ADR-105 Retire the legacy ingest surface + 💡 ADR-107 Ingest notes, by slug reference + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + ✅ ADR-112 Operator-authored constraint + 💡 ADR-114 Cap batch size + +Total: 9 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-group-by-v0.out b/tests/fixtures/adr/golden/record-list-group-by-v0.out new file mode 100644 index 00000000..207622fa --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-group-by-v0.out @@ -0,0 +1,38 @@ + +ADR Fixture — Agent Decision Records (14 total) +======================================================= + +## status: Accepted (9) +-------------------------------------------------- + ✅ ADR-012 Plain-text logging + ✅ ADR-101 Structured storage + ✅ ADR-101.1 Storage migration + ✅ ADR-102 Hook ordering (partially superseded by ADR-104 §2) + ✅ ADR-104 Hook priorities + ✅ ADR-106 One-sided supersession + ✅ ADR-200 Restart procedure + ✅ ADR-300 Documentation structure + ✅ ADR-303 Dangling reference + +## status: Deprecated (1) +-------------------------------------------------- + 🗑️ ADR-105 Telemetry export + +## status: Draft (1) +-------------------------------------------------- + 📝 ADR-301 Draft style guide + +## status: Proposed (1) +-------------------------------------------------- + 💡 ADR-103 Cache layer + +## status: Superseded (1) +-------------------------------------------------- + 📦 ADR-005 Legacy storage format (superseded by ADR-101) + +## status: (none) (1) +-------------------------------------------------- + ❓ ADR-302 Inline metadata + +Total: 14 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-group-by.out b/tests/fixtures/adr/golden/record-list-group-by.out new file mode 100644 index 00000000..58f0cf90 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-group-by.out @@ -0,0 +1,37 @@ + +ADR v1 Fixture — Agent Decision Records (15 total) +======================================================= + +## capability: * (1) +-------------------------------------------------- + ✅ ADR-104 No network calls in hooks + +## capability: adr (2) +-------------------------------------------------- + ✅ ADR-100 Adopt the adr/v1 contract + ✅ ADR-112 Operator-authored constraint + +## capability: ingest (9) +-------------------------------------------------- + ✅ ADR-101 Add ingestion + ✅ ADR-102 How ingestion works + ✅ ADR-103 Batch ingestion (partially superseded by ADR-109) + ✅ ADR-105 Retire the legacy ingest surface + 💡 ADR-107 Ingest notes, by slug reference + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + ✅ ADR-112 Operator-authored constraint + 💡 ADR-114 Cap batch size + +## capability: search (3) +-------------------------------------------------- + 💡 ADR-106 Search result caps + ✅ ADR-111 Cut search + 💡 ADR-113 Operator-started and still proposed + +## capability: (none) (1) +-------------------------------------------------- + ✅ ADR-110 An unmigrated v0 record (partially superseded by ADR-108) + +Total: 15 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-json-group.out b/tests/fixtures/adr/golden/record-list-json-group.out new file mode 100644 index 00000000..b15a1952 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-json-group.out @@ -0,0 +1,37 @@ +{ + "retire": [ + { + "number": "105", + "title": "Retire the legacy ingest surface", + "path": "docs/architecture/system/ADR-105-retire-legacy-ingest.md", + "status": "accepted", + "frontmatter": { + "contract": "adr/v1", + "kind": "decision", + "verb": "retire", + "capability": "ingest", + "targets": [ + "cli:ingest-legacy", + "route:/v1/upload" + ], + "status": "accepted", + "enacted": "3f9c2a1", + "date": "2025-05-06", + "deciders": [ + "developer", + "agent" + ], + "agent": { + "name": "Claude", + "model": "fixture-model" + }, + "basis": [ + { + "evidence": "fixture measurement" + } + ] + } + } + ] +} +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-json.out b/tests/fixtures/adr/golden/record-list-json.out new file mode 100644 index 00000000..2c531a43 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-json.out @@ -0,0 +1,35 @@ +[ + { + "number": "105", + "title": "Retire the legacy ingest surface", + "path": "docs/architecture/system/ADR-105-retire-legacy-ingest.md", + "status": "accepted", + "frontmatter": { + "contract": "adr/v1", + "kind": "decision", + "verb": "retire", + "capability": "ingest", + "targets": [ + "cli:ingest-legacy", + "route:/v1/upload" + ], + "status": "accepted", + "enacted": "3f9c2a1", + "date": "2025-05-06", + "deciders": [ + "developer", + "agent" + ], + "agent": { + "name": "Claude", + "model": "fixture-model" + }, + "basis": [ + { + "evidence": "fixture measurement" + } + ] + } + } +] +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-kind.out b/tests/fixtures/adr/golden/record-list-kind.out new file mode 100644 index 00000000..22a6f416 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-kind.out @@ -0,0 +1,7 @@ + +ADR v1 Fixture — Agent Decision Records (1 total) +======================================================= + ✅ ADR-102 How ingestion works + +Total: 1 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-present.out b/tests/fixtures/adr/golden/record-list-present.out new file mode 100644 index 00000000..2ab3bbea --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-present.out @@ -0,0 +1,8 @@ + +ADR v1 Fixture — Agent Decision Records (2 total) +======================================================= + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + +Total: 2 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-list-verb.out b/tests/fixtures/adr/golden/record-list-verb.out new file mode 100644 index 00000000..dad3b3e7 --- /dev/null +++ b/tests/fixtures/adr/golden/record-list-verb.out @@ -0,0 +1,10 @@ + +ADR v1 Fixture — Agent Decision Records (4 total) +======================================================= + 💡 ADR-106 Search result caps + 💡 ADR-107 Ingest notes, by slug reference + 💡 ADR-108 Change against a v0 record + 💡 ADR-109 Grounded through precedent, via a decision and a spec + +Total: 4 ADRs +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-accepted.out b/tests/fixtures/adr/golden/record-set-accepted.out new file mode 100644 index 00000000..4b59f95b --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-accepted.out @@ -0,0 +1,3 @@ +Set capability, related on ADR-101: Add ingestion +Lint docs/architecture/system/ADR-101-ingest.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-append-block-file.md b/tests/fixtures/adr/golden/record-set-append-block-file.md new file mode 100644 index 00000000..bfbf0381 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-append-block-file.md @@ -0,0 +1,32 @@ +path: docs/architecture/system/ADR-114-open-concern.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: ingest +status: proposed +date: 2025-05-14 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: load test at 10x volume + - evidence: a second load test +concern: + - said: "Batch size may starve small tenants" + resolve: "Measure p99 latency for a small tenant" + - said: Retries may double + resolve: Count retries +--- + +# ADR-114: Cap batch size + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +Cap batches at 500 documents. diff --git a/tests/fixtures/adr/golden/record-set-append-block.out b/tests/fixtures/adr/golden/record-set-append-block.out new file mode 100644 index 00000000..e5490948 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-append-block.out @@ -0,0 +1,5 @@ +Set basis, concern on ADR-114: Cap batch size +Lint docs/architecture/system/ADR-114-open-concern.md: + ⚠️ open concern: Batch size may starve small tenants + ⚠️ open concern: Retries may double +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-bad-form.out b/tests/fixtures/adr/golden/record-set-bad-form.out new file mode 100644 index 00000000..d1223ab6 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-bad-form.out @@ -0,0 +1,2 @@ +Error: 'related' is not key=value, key+=value or key-=value +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-bad-yaml.out b/tests/fixtures/adr/golden/record-set-bad-yaml.out new file mode 100644 index 00000000..d20998b0 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-bad-yaml.out @@ -0,0 +1,2 @@ +Error: related: the value is not YAML (while parsing a flow sequence) +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-crlf-file.txt b/tests/fixtures/adr/golden/record-set-crlf-file.txt new file mode 100644 index 00000000..17c878ee --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-crlf-file.txt @@ -0,0 +1 @@ +crlf kept diff --git a/tests/fixtures/adr/golden/record-set-crlf.out b/tests/fixtures/adr/golden/record-set-crlf.out new file mode 100644 index 00000000..f37391a1 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-crlf.out @@ -0,0 +1,4 @@ +Set basis on ADR-114: Cap batch size +Lint docs/architecture/system/ADR-114-open-concern.md: + ⚠️ open concern: Batch size may starve small tenants +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-dry-run-status.txt b/tests/fixtures/adr/golden/record-set-dry-run-status.txt new file mode 100644 index 00000000..e69de29b diff --git a/tests/fixtures/adr/golden/record-set-dry-run.out b/tests/fixtures/adr/golden/record-set-dry-run.out new file mode 100644 index 00000000..ea8603d3 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-dry-run.out @@ -0,0 +1,13 @@ +--- a/docs/architecture/system/ADR-106-search-change.md ++++ b/docs/architecture/system/ADR-106-search-change.md +@@ -1,7 +1,7 @@ + --- + contract: adr/v1 + kind: decision +-verb: change ++verb: add + capability: search + amends: [ADR-104#1] + extends: [ADR-100] +Dry run, nothing written: Set verb on ADR-106: Search result caps +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-file.md b/tests/fixtures/adr/golden/record-set-file.md new file mode 100644 index 00000000..b131e4e9 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-file.md @@ -0,0 +1,35 @@ +path: docs/architecture/system/ADR-106-search-change.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: [search, ingest] +amends: [ADR-104#1] +extends: [] +status: proposed +date: 2025-05-17 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +related: + - ADR-100 + - ADR-101 +--- + +# ADR-106: Search result caps + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-135-change-no-prior.md b/tests/fixtures/adr/golden/record-set-mutable-file.md similarity index 76% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-135-change-no-prior.md rename to tests/fixtures/adr/golden/record-set-mutable-file.md index 1b7ec6cb..a5a10c42 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-135-change-no-prior.md +++ b/tests/fixtures/adr/golden/record-set-mutable-file.md @@ -1,18 +1,20 @@ +path: docs/architecture/system/ADR-103-ingest-batching.md --- contract: adr/v1 kind: decision verb: change capability: ingest -extends: [ADR-100] +amends: [ADR-101#1] +superseded_by: [ADR-109, ADR-107] status: accepted -date: 2025-06-01 +date: 2025-05-04 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} basis: - evidence: fixture measurement --- -# ADR-135: Change with no prior decision on its capability +# ADR-103: Batch ingestion ## Summary diff --git a/tests/fixtures/adr/golden/record-set-mutable.out b/tests/fixtures/adr/golden/record-set-mutable.out new file mode 100644 index 00000000..d87ffccd --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-mutable.out @@ -0,0 +1,4 @@ +Set superseded_by on ADR-103: Batch ingestion +Lint docs/architecture/system/ADR-103-ingest-batching.md: + ⚠️ superseded_by: ADR-107 does not declare the reciprocal supersedes: ADR-103 — one-directional links rot +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-remove-block-file.md b/tests/fixtures/adr/golden/record-set-remove-block-file.md new file mode 100644 index 00000000..7b6f0f70 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-remove-block-file.md @@ -0,0 +1,34 @@ +path: docs/architecture/system/ADR-106-search-change.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: [search, ingest] +amends: [ADR-104#1] +extends: [] +status: proposed +date: 2025-05-17 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +related: + - ADR-101 +--- + +# ADR-106: Search result caps + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/record-set-remove-block.out b/tests/fixtures/adr/golden/record-set-remove-block.out new file mode 100644 index 00000000..04191a72 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-remove-block.out @@ -0,0 +1,3 @@ +Set related on ADR-106: Search result caps +Lint docs/architecture/system/ADR-106-search-change.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set-remove-missing.out b/tests/fixtures/adr/golden/record-set-remove-missing.out new file mode 100644 index 00000000..a6e15606 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-remove-missing.out @@ -0,0 +1,2 @@ +Error: ADR-106 not changed: related does not list 'ADR-9'. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-status-cmd.out b/tests/fixtures/adr/golden/record-set-status-cmd.out new file mode 100644 index 00000000..bde97f2a --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-status-cmd.out @@ -0,0 +1,2 @@ +Refused: use `adr accept 106`; it checks the record and records why. --force sets the status anyway. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-status-other.out b/tests/fixtures/adr/golden/record-set-status-other.out new file mode 100644 index 00000000..09df6b5c --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-status-other.out @@ -0,0 +1,2 @@ +Refused: a record's status changes only through accept, reject, abandon, supersede and archive. --force sets the status anyway. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-status.txt b/tests/fixtures/adr/golden/record-set-status.txt new file mode 100644 index 00000000..6dace6ff --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-status.txt @@ -0,0 +1,4 @@ +M docs/architecture/system/ADR-103-ingest-batching.md +M docs/architecture/system/ADR-106-search-change.md +M docs/architecture/system/ADR-110-old-v0-record.md +M docs/architecture/system/ADR-114-open-concern.md diff --git a/tests/fixtures/adr/golden/record-set-unknown-key.out b/tests/fixtures/adr/golden/record-set-unknown-key.out new file mode 100644 index 00000000..b933d988 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-unknown-key.out @@ -0,0 +1,2 @@ +Refused: ADR-106 has no field 'capabilty', and it is not a v1 field (contract, kind, verb, capability, targets, supersedes, amends, extends, decided_by, superseded_by, enacted, basis, agent, considered, concern, observable, status, date, deciders, related, imported). --force adds it anyway. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-set-v0-file.md b/tests/fixtures/adr/golden/record-set-v0-file.md new file mode 100644 index 00000000..5479c014 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-v0-file.md @@ -0,0 +1,21 @@ +path: docs/architecture/system/ADR-110-old-v0-record.md +--- +status: Superseded +date: 2025-01-01 +deciders: [developer] +superseded_by: [ADR-108] +--- + +# ADR-110: An unmigrated v0 record + +## Summary + +Context for the record. + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/record-set-v0.out b/tests/fixtures/adr/golden/record-set-v0.out new file mode 100644 index 00000000..7bb81a92 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set-v0.out @@ -0,0 +1,3 @@ +Set status on ADR-110: An unmigrated v0 record +Lint docs/architecture/system/ADR-110-old-v0-record.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-set.out b/tests/fixtures/adr/golden/record-set.out new file mode 100644 index 00000000..caf3f403 --- /dev/null +++ b/tests/fixtures/adr/golden/record-set.out @@ -0,0 +1,3 @@ +Set capability, related, extends, date on ADR-106: Search result caps +Lint docs/architecture/system/ADR-106-search-change.md: clean +[exit 0] diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-145-change-supersedes-broader.md b/tests/fixtures/adr/golden/record-supersede-accepted-file.md similarity index 73% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-145-change-supersedes-broader.md rename to tests/fixtures/adr/golden/record-supersede-accepted-file.md index 520089f1..08163013 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-145-change-supersedes-broader.md +++ b/tests/fixtures/adr/golden/record-supersede-accepted-file.md @@ -1,18 +1,20 @@ +path: docs/architecture/system/ADR-104-no-network-in-hooks.md --- contract: adr/v1 kind: decision -verb: change -capability: adr -supersedes: [ADR-144] +verb: constrain +capability: "*" +supersedes: + - ADR-101 status: accepted -date: 2025-06-01 +date: 2025-05-05 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} basis: - evidence: fixture measurement --- -# ADR-145: Change superseding a broader constraint +# ADR-104: No network calls in hooks ## Summary diff --git a/tests/fixtures/adr/golden/record-supersede-accepted.out b/tests/fixtures/adr/golden/record-supersede-accepted.out new file mode 100644 index 00000000..587e2ee1 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-accepted.out @@ -0,0 +1,5 @@ +ADR-104 supersedes ADR-101; ADR-101 is superseded +Lint docs/architecture/system/ADR-101-ingest.md: clean +Lint docs/architecture/system/ADR-104-no-network-in-hooks.md: clean +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/record-supersede-amends-file.md b/tests/fixtures/adr/golden/record-supersede-amends-file.md new file mode 100644 index 00000000..7644f7d7 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-amends-file.md @@ -0,0 +1,32 @@ +path: docs/architecture/system/ADR-107-ingest-notes.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: ingest +amends: [ADR-101#1, ADR-100#2] +extends: [ADR-103#2-consequences] +status: proposed +date: 2025-05-08 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +--- + +# ADR-107: Ingest notes, by slug reference + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/record-supersede-amends-status.txt b/tests/fixtures/adr/golden/record-supersede-amends-status.txt new file mode 100644 index 00000000..c37aa378 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-amends-status.txt @@ -0,0 +1 @@ +M docs/architecture/system/ADR-107-ingest-notes.md diff --git a/tests/fixtures/adr/golden/record-supersede-amends.out b/tests/fixtures/adr/golden/record-supersede-amends.out new file mode 100644 index 00000000..fc9eaf15 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-amends.out @@ -0,0 +1,3 @@ +ADR-107 amends ADR-100 §2 +Lint docs/architecture/system/ADR-107-ingest-notes.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/record-supersede-kind.out b/tests/fixtures/adr/golden/record-supersede-kind.out new file mode 100644 index 00000000..d366c5e3 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-kind.out @@ -0,0 +1,2 @@ +Error: a decision record cannot supersede a spec: for a decision, supersedes points at decision records (adr.yaml kinds.decision.edges). +[exit 1] diff --git a/tests/fixtures/adr/golden/record-supersede-new-file.md b/tests/fixtures/adr/golden/record-supersede-new-file.md new file mode 100644 index 00000000..1d4b376c --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-new-file.md @@ -0,0 +1,34 @@ +path: docs/architecture/system/ADR-106-search-change.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: search +supersedes: + - ADR-101 +amends: [ADR-104#1] +extends: [ADR-100] +status: proposed +date: 2025-05-07 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +--- + +# ADR-106: Search result caps + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/record-supersede-new-rejected.out b/tests/fixtures/adr/golden/record-supersede-new-rejected.out new file mode 100644 index 00000000..ceb412ce --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-new-rejected.out @@ -0,0 +1,2 @@ +Error: ADR-108 is rejected; only an accepted or proposed record supersedes another. +[exit 1] diff --git a/tests/fixtures/adr/golden/record-supersede-no-section.out b/tests/fixtures/adr/golden/record-supersede-no-section.out new file mode 100644 index 00000000..00df87e3 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-no-section.out @@ -0,0 +1,2 @@ +Error: ADR-100 has no section '9'. Its sections: Summary, 1. Decision, 2. Consequences. +[exit 1] diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-130-retire-bad-targets.md b/tests/fixtures/adr/golden/record-supersede-old-file.md similarity index 76% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-130-retire-bad-targets.md rename to tests/fixtures/adr/golden/record-supersede-old-file.md index 2727ea71..1336a762 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-130-retire-bad-targets.md +++ b/tests/fixtures/adr/golden/record-supersede-old-file.md @@ -1,18 +1,20 @@ +path: docs/architecture/system/ADR-101-ingest.md --- contract: adr/v1 kind: decision -verb: retire +verb: add capability: ingest -targets: [grpc:Upload, plainname] -status: accepted -date: 2025-06-01 +superseded_by: + - ADR-106 +status: superseded +date: 2025-05-02 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} basis: - evidence: fixture measurement --- -# ADR-130: Retire with bad targets +# ADR-101: Add ingestion ## Summary diff --git a/tests/fixtures/adr/golden/record-supersede-old-proposed.out b/tests/fixtures/adr/golden/record-supersede-old-proposed.out new file mode 100644 index 00000000..63d3865a --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-old-proposed.out @@ -0,0 +1,2 @@ +Error: ADR-113 is proposed; only an accepted record is superseded (a proposed one is rejected or abandoned). +[exit 1] diff --git a/tests/fixtures/adr/golden/record-supersede-status.txt b/tests/fixtures/adr/golden/record-supersede-status.txt new file mode 100644 index 00000000..8d6126cf --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede-status.txt @@ -0,0 +1,2 @@ +M docs/architecture/system/ADR-101-ingest.md +M docs/architecture/system/ADR-106-search-change.md diff --git a/tests/fixtures/adr/golden/record-supersede.out b/tests/fixtures/adr/golden/record-supersede.out new file mode 100644 index 00000000..842b8174 --- /dev/null +++ b/tests/fixtures/adr/golden/record-supersede.out @@ -0,0 +1,5 @@ +ADR-106 supersedes ADR-101; ADR-101 is superseded +Lint docs/architecture/system/ADR-101-ingest.md: clean +Lint docs/architecture/system/ADR-106-search-change.md: clean +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/reject-for-supersede.out b/tests/fixtures/adr/golden/reject-for-supersede.out new file mode 100644 index 00000000..1a5cc74e --- /dev/null +++ b/tests/fixtures/adr/golden/reject-for-supersede.out @@ -0,0 +1,3 @@ +Marked ADR-108 rejected: Change against a v0 record +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-chain-basis-104.md b/tests/fixtures/adr/golden/relocate-chain-basis-104.md new file mode 100644 index 00000000..cbbfc387 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-chain-basis-104.md @@ -0,0 +1,30 @@ +path: docs/architecture/operations/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: "see [ingest](../operations/ADR-101-ingest.md)" +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. diff --git a/tests/fixtures/adr/golden/relocate-chain-body-104.md b/tests/fixtures/adr/golden/relocate-chain-body-104.md new file mode 100644 index 00000000..571b031a --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-chain-body-104.md @@ -0,0 +1,32 @@ +path: docs/architecture/operations/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +See [ingest](../operations/ADR-101-ingest.md). diff --git a/tests/fixtures/adr/golden/relocate-move-104.md b/tests/fixtures/adr/golden/relocate-move-104.md new file mode 100644 index 00000000..9d4e6f24 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-move-104.md @@ -0,0 +1,32 @@ +path: docs/architecture/system/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: "the survey at docs/architecture/documentation/ADR-101-ingest.md" +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +See [ingest](../documentation/ADR-101-ingest.md), [the guide](https://example.com/v1/guide), [the notes](../api/README.md), and/or the spec. diff --git a/tests/fixtures/adr/golden/relocate-move-dry.out b/tests/fixtures/adr/golden/relocate-move-dry.out new file mode 100644 index 00000000..dc269114 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-move-dry.out @@ -0,0 +1,19 @@ +Would move: ADR-101 system → docs + docs/architecture/system/ADR-101-ingest.md → docs/architecture/documentation/ADR-101-ingest.md +Would rewrite 4 paths in 2 files + CHANGELOG.md: 2 + CHANGELOG.md:4 + - This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md + → - This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/documentation/ADR-101-ingest.md + CHANGELOG.md:6 + - See docs/architecture/system/ADR-101-ingest.md. + → - See docs/architecture/documentation/ADR-101-ingest.md. + docs/architecture/system/ADR-104-no-network-in-hooks.md: 2 + docs/architecture/system/ADR-104-no-network-in-hooks.md:11 + - evidence: "the survey at docs/architecture/system/ADR-101-ingest.md" + → - evidence: "the survey at docs/architecture/documentation/ADR-101-ingest.md" + docs/architecture/system/ADR-104-no-network-in-hooks.md:31 + See [ingest](ADR-101-ingest.md), [the guide](https://example.com/v1/guide), [the notes](../api/README.md), and/or the spec. + → See [ingest](../documentation/ADR-101-ingest.md), [the guide](https://example.com/v1/guide), [the notes](../api/README.md), and/or the spec. +Dry run: nothing written. +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-frozen-merged-branch.out b/tests/fixtures/adr/golden/relocate-move-lint.out similarity index 100% rename from tests/fixtures/adr/golden/v1-frozen-merged-branch.out rename to tests/fixtures/adr/golden/relocate-move-lint.out diff --git a/tests/fixtures/adr/golden/relocate-move.out b/tests/fixtures/adr/golden/relocate-move.out new file mode 100644 index 00000000..8d611a7e --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-move.out @@ -0,0 +1,6 @@ +Moved: ADR-101 system → docs + docs/architecture/system/ADR-101-ingest.md → docs/architecture/documentation/ADR-101-ingest.md +Rewrote 4 paths in 2 files + CHANGELOG.md: 2 + docs/architecture/system/ADR-104-no-network-in-hooks.md: 2 +[exit 0] diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-162-precedent-internal.md b/tests/fixtures/adr/golden/relocate-refs-109.md similarity index 62% rename from tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-162-precedent-internal.md rename to tests/fixtures/adr/golden/relocate-refs-109.md index 9321b6a2..61f8608f 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-162-precedent-internal.md +++ b/tests/fixtures/adr/golden/relocate-refs-109.md @@ -1,17 +1,20 @@ +path: docs/architecture/system/ADR-109-precedent-chain.md --- contract: adr/v1 kind: decision -verb: add -capability: adr -date: 2025-06-01 +verb: change +capability: ingest +supersedes: [ADR-103] +status: proposed +date: 2025-05-10 deciders: [developer, agent] agent: {name: Claude, model: fixture-model} -status: proposed basis: - precedent: ADR-101 + - precedent: ADR-102 --- -# ADR-162: Precedent to a spec nobody decided +# ADR-109: Grounded through precedent, via a decision and a spec ## Summary @@ -27,3 +30,5 @@ The decision. ## 2. Consequences They follow. + +On Windows: docs\architecture\system\ADR-101-ingest.md diff --git a/tests/fixtures/adr/golden/relocate-refs-dry.out b/tests/fixtures/adr/golden/relocate-refs-dry.out new file mode 100644 index 00000000..f68c032b --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-refs-dry.out @@ -0,0 +1,18 @@ +Would move: ADR-101 system → docs + docs/architecture/system/ADR-101-ingest.md → docs/architecture/documentation/ADR-101-ingest.md +Would rewrite 4 paths in 1 file + NOTES.md: 4 + NOTES.md:7 + - https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md + → - https://github.com/fixture/corpus/blob/main/docs/architecture/documentation/ADR-101-ingest.md + NOTES.md:8 + - https://github.com/fixture/corpus/blob/feature/x/docs/architecture/system/ADR-101-ingest.md + → - https://github.com/fixture/corpus/blob/feature/x/docs/architecture/documentation/ADR-101-ingest.md + NOTES.md:9 + - https://raw.githubusercontent.com/fixture/corpus/main/docs/architecture/system/ADR-101-ingest.md + → - https://raw.githubusercontent.com/fixture/corpus/main/docs/architecture/documentation/ADR-101-ingest.md + NOTES.md:22 + See docs/architecture/system/ADR-101-ingest.md. + → See docs/architecture/documentation/ADR-101-ingest.md. +Dry run: nothing written. +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-refs-notes.md b/tests/fixtures/adr/golden/relocate-refs-notes.md new file mode 100644 index 00000000..2d147aab --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-refs-notes.md @@ -0,0 +1,23 @@ +path: NOTES.md +# Notes + +- https://github.com/fixture/corpus/blob/0123456789abcdef0123456789abcdef01234567/docs/architecture/system/ADR-101-ingest.md +- https://github.com/fixture/corpus/blob/abc1234/docs/architecture/system/ADR-101-ingest.md +- https://github.com/fixture/corpus/blob/ABC1234/docs/architecture/system/ADR-101-ingest.md +- https://github.com/fixture/corpus/blob/v1.0/docs/architecture/system/ADR-101-ingest.md +- https://github.com/fixture/corpus/blob/main/docs/architecture/documentation/ADR-101-ingest.md +- https://github.com/fixture/corpus/blob/feature/x/docs/architecture/documentation/ADR-101-ingest.md +- https://raw.githubusercontent.com/fixture/corpus/main/docs/architecture/documentation/ADR-101-ingest.md +- docs\architecture\system\ADR-101-ingest.md + +```sh +git mv docs/architecture/system/ADR-101-ingest.md docs/architecture/documentation/ADR-101-ingest.md +``` + +1. Step one + + ```sh + cat docs/architecture/system/ADR-101-ingest.md + ``` + +See docs/architecture/documentation/ADR-101-ingest.md. diff --git a/tests/fixtures/adr/golden/relocate-refs.out b/tests/fixtures/adr/golden/relocate-refs.out new file mode 100644 index 00000000..9e1d941d --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-refs.out @@ -0,0 +1,5 @@ +Moved: ADR-101 system → docs + docs/architecture/system/ADR-101-ingest.md → docs/architecture/documentation/ADR-101-ingest.md +Rewrote 4 paths in 1 file + NOTES.md: 4 +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-rename-104.md b/tests/fixtures/adr/golden/relocate-rename-104.md new file mode 100644 index 00000000..16d5068f --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename-104.md @@ -0,0 +1,32 @@ +path: docs/architecture/platform/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: "the survey at docs/architecture/platform/ADR-101-ingest.md" +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +See [ingest](ADR-101-ingest.md), [the guide](https://example.com/v1/guide), [the notes](../api/README.md), and/or the spec. diff --git a/tests/fixtures/adr/golden/relocate-rename-app.py b/tests/fixtures/adr/golden/relocate-rename-app.py new file mode 100644 index 00000000..99fcf459 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename-app.py @@ -0,0 +1,3 @@ +path: src/app.py +TEMPLATE_DIR = "architecture/system" +RECORDS = "docs/architecture/platform" diff --git a/tests/fixtures/adr/golden/relocate-rename-changelog.md b/tests/fixtures/adr/golden/relocate-rename-changelog.md new file mode 100644 index 00000000..73e23008 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename-changelog.md @@ -0,0 +1,7 @@ +path: CHANGELOG.md +# Changes + +- Another repo: https://github.com/someone/else/tree/main/docs/architecture/system/ADR-101-ingest.md +- This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/platform/ADR-101-ingest.md +- The vendor layout uses architecture/system and a system/ADR-101-ingest.md file. +- See docs/architecture/platform/ADR-101-ingest.md. diff --git a/tests/fixtures/adr/golden/relocate-rename-dry.out b/tests/fixtures/adr/golden/relocate-rename-dry.out new file mode 100644 index 00000000..eca91d59 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename-dry.out @@ -0,0 +1,27 @@ +Would rename domain: system → platform + Folder: docs/architecture/system/ → docs/architecture/platform/ +Would rewrite 6 paths in 4 files + CHANGELOG.md: 2 + CHANGELOG.md:4 + - This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/system/ADR-101-ingest.md + → - This repo: https://github.com/fixture/corpus/blob/main/docs/architecture/platform/ADR-101-ingest.md + CHANGELOG.md:6 + - See docs/architecture/system/ADR-101-ingest.md. + → - See docs/architecture/platform/ADR-101-ingest.md. + docs/architecture/adr.yaml: 2 + docs/architecture/adr.yaml:6 + system: + → platform: + docs/architecture/adr.yaml:10 + folder: system + → folder: platform + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 + docs/architecture/system/ADR-104-no-network-in-hooks.md:11 + - evidence: "the survey at docs/architecture/system/ADR-101-ingest.md" + → - evidence: "the survey at docs/architecture/platform/ADR-101-ingest.md" + src/app.py: 1 + src/app.py:2 + RECORDS = "docs/architecture/system" + → RECORDS = "docs/architecture/platform" +Dry run: nothing written. +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-rename-lint.out b/tests/fixtures/adr/golden/relocate-rename-lint.out new file mode 100644 index 00000000..dac16205 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename-lint.out @@ -0,0 +1,21 @@ + +Scanned: 15 ADRs + +Status distribution: + accepted: 9 + proposed: 6 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/platform/ADR-114-open-concern.md + ⚠️ open concern: Batch size may starve small tenants + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-rename.out b/tests/fixtures/adr/golden/relocate-rename.out new file mode 100644 index 00000000..d07fa785 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-rename.out @@ -0,0 +1,8 @@ +Renamed domain: system → platform + Folder: docs/architecture/system/ → docs/architecture/platform/ +Rewrote 6 paths in 4 files + CHANGELOG.md: 2 + docs/architecture/adr.yaml: 2 + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 + src/app.py: 1 +[exit 0] diff --git a/tests/fixtures/adr/golden/relocate-repository-104.md b/tests/fixtures/adr/golden/relocate-repository-104.md new file mode 100644 index 00000000..8f4ed69b --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-repository-104.md @@ -0,0 +1,32 @@ +path: docs/architecture/system/ADR-104-no-network-in-hooks.md +--- +contract: adr/v1 +kind: decision +verb: constrain +capability: "*" +status: accepted +date: 2025-05-05 +deciders: [developer, agent] +agent: {name: Claude, model: fixture-model} +basis: + - evidence: fixture measurement +--- + +# ADR-104: No network calls in hooks + +## Summary + +- **Decided:** the decision in plain terms. +- **Trades away:** what it gives up. +- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. +- **Inversion:** one end, the other end; is the middle right? + +## 1. Decision + +The decision. + +## 2. Consequences + +They follow. + +See https://github.com/fixture/corpus/blob/main/docs/architecture/documentation/ADR-101-ingest.md. diff --git a/tests/fixtures/adr/golden/v1-frozen-nonutf8.out b/tests/fixtures/adr/golden/relocate-repository-bad.out similarity index 90% rename from tests/fixtures/adr/golden/v1-frozen-nonutf8.out rename to tests/fixtures/adr/golden/relocate-repository-bad.out index 0d784548..7038dc9b 100644 --- a/tests/fixtures/adr/golden/v1-frozen-nonutf8.out +++ b/tests/fixtures/adr/golden/relocate-repository-bad.out @@ -11,10 +11,10 @@ Issues found in 1 files: ──────────────────────────────────────────────────────────── docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision + ❌ repository: expected host/owner/repo, or a list of them ════════════════════════════════════════════════════════════ -Summary: 0 errors, 1 warnings +Summary: 1 errors, 0 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/relocate-repository-move.out b/tests/fixtures/adr/golden/relocate-repository-move.out new file mode 100644 index 00000000..e57c2570 --- /dev/null +++ b/tests/fixtures/adr/golden/relocate-repository-move.out @@ -0,0 +1,5 @@ +Moved: ADR-101 system → docs + docs/architecture/system/ADR-101-ingest.md → docs/architecture/documentation/ADR-101-ingest.md +Rewrote 1 path in 1 file + docs/architecture/system/ADR-104-no-network-in-hooks.md: 1 +[exit 0] diff --git a/tests/fixtures/adr/golden/set-for-enact.out b/tests/fixtures/adr/golden/set-for-enact.out new file mode 100644 index 00000000..dc97eb60 --- /dev/null +++ b/tests/fixtures/adr/golden/set-for-enact.out @@ -0,0 +1,3 @@ +Set status on ADR-111: Cut search +Lint docs/architecture/system/ADR-111-cut-search.md: clean +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-accept-dependent.out b/tests/fixtures/adr/golden/v1-accept-dependent.out new file mode 100644 index 00000000..35fc94db --- /dev/null +++ b/tests/fixtures/adr/golden/v1-accept-dependent.out @@ -0,0 +1,3 @@ +Accepted ADR-116: Rests on ADR-114 +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-accept-no-considered.out b/tests/fixtures/adr/golden/v1-accept-no-considered.out new file mode 100644 index 00000000..1b2bfcfb --- /dev/null +++ b/tests/fixtures/adr/golden/v1-accept-no-considered.out @@ -0,0 +1,3 @@ +Accepted ADR-113: Operator-started and still proposed +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-accept-own-errors.out b/tests/fixtures/adr/golden/v1-accept-own-errors.out new file mode 100644 index 00000000..34225a3f --- /dev/null +++ b/tests/fixtures/adr/golden/v1-accept-own-errors.out @@ -0,0 +1,8 @@ +Refused: accepting ADR-117 would leave it with errors: + ❌ docs/architecture/system/ADR-117-bare-decision.md: a decision record requires a verb (add, cut, change, retire, constrain) + ❌ docs/architecture/system/ADR-117-bare-decision.md: a decision record requires 'capability' + ❌ docs/architecture/system/ADR-117-bare-decision.md: a decision record requires 'basis' + ❌ docs/architecture/system/ADR-117-bare-decision.md: agent: records 'name' + ❌ docs/architecture/system/ADR-117-bare-decision.md: agent: records 'model' + ❌ docs/architecture/system/ADR-117-bare-decision.md: 12 placeholder line(s) from `adr new` still to fill, first: - **Decided:** [what is decided, in plain terms] +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-accept-pending-precedent.out b/tests/fixtures/adr/golden/v1-accept-pending-precedent.out deleted file mode 100644 index 423b0ab9..00000000 --- a/tests/fixtures/adr/golden/v1-accept-pending-precedent.out +++ /dev/null @@ -1,3 +0,0 @@ -Refused: accepting ADR-116 would add errors: - ❌ docs/architecture/system/ADR-116-rests-on-114.md: basis: precedent ADR-114 is still proposed -[exit 1] diff --git a/tests/fixtures/adr/golden/v1-accept-refused.out b/tests/fixtures/adr/golden/v1-accept-refused.out deleted file mode 100644 index f29e0111..00000000 --- a/tests/fixtures/adr/golden/v1-accept-refused.out +++ /dev/null @@ -1,3 +0,0 @@ -Refused: accepting ADR-113 would add errors: - ❌ docs/architecture/system/ADR-113-operator-proposed.md: accepted with an operator basis but no considered entry; the operator considers what they started -[exit 1] diff --git a/tests/fixtures/adr/golden/v1-accept-then-lint.out b/tests/fixtures/adr/golden/v1-accept-then-lint.out index 64839585..8abf4e9a 100644 --- a/tests/fixtures/adr/golden/v1-accept-then-lint.out +++ b/tests/fixtures/adr/golden/v1-accept-then-lint.out @@ -7,17 +7,14 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 2 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - docs/architecture/system/ADR-114-open-concern.md ⚠️ open concern: Batch size may starve small tenants ════════════════════════════════════════════════════════════ -Summary: 0 errors, 2 warnings +Summary: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-change-list-long.out b/tests/fixtures/adr/golden/v1-change-list-long.out new file mode 100644 index 00000000..466e5cd2 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-change-list-long.out @@ -0,0 +1,21 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-116-shared-queue.md + ⚠️ a change lists 4 capabilities; list only those it alters, the rest belong in related (ADR-308 §3) + ⚠️ supersedes: ADR-103 does not declare the reciprocal superseded_by: ADR-116 — one-directional links rot + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 2 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-change-list.out b/tests/fixtures/adr/golden/v1-change-list.out new file mode 100644 index 00000000..c630f8f2 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-change-list.out @@ -0,0 +1,20 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-116-shared-queue.md + ⚠️ supersedes: ADR-103 does not declare the reciprocal superseded_by: ADR-116 — one-directional links rot + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-cite-no-inventory.out b/tests/fixtures/adr/golden/v1-cite-no-inventory.out index be9d78c9..d1147b20 100644 --- a/tests/fixtures/adr/golden/v1-cite-no-inventory.out +++ b/tests/fixtures/adr/golden/v1-cite-no-inventory.out @@ -1,7 +1,6 @@ ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force - ❌ src/search.py:1 ADR-106 is on 'search', cut and enacted by ADR-111; remove the citation ════════════════════════════════════════════════════════════ -Citations: 1 errors, 1 warnings +Citations: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-cite-retire-pending.out b/tests/fixtures/adr/golden/v1-cite-retire-pending.out deleted file mode 100644 index 15b47ab0..00000000 --- a/tests/fixtures/adr/golden/v1-cite-retire-pending.out +++ /dev/null @@ -1,9 +0,0 @@ - ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force - ⚠️ src/search.py:1 ADR-106 is on 'search', which ADR-111 cuts; remove before enacting - ⚠️ docs/architecture/system/ADR-105-retire-legacy-ingest.md cli:ingest-legacy is still present; ADR-105 retires it - ⚠️ docs/architecture/system/ADR-105-retire-legacy-ingest.md cli:ingest-legcy is not in the 'cli' inventory; check the name - -════════════════════════════════════════════════════════════ -Citations: 0 errors, 4 warnings -════════════════════════════════════════════════════════════ -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-cite.out b/tests/fixtures/adr/golden/v1-cite.out index 4199e686..d1147b20 100644 --- a/tests/fixtures/adr/golden/v1-cite.out +++ b/tests/fixtures/adr/golden/v1-cite.out @@ -1,8 +1,6 @@ ⚠️ src/search.py:1 ADR-106 is still proposed; accept it or cite the record in force - ⚠️ src/search.py:1 ADR-106 is on 'search', which ADR-111 cuts; remove before enacting - ❌ docs/architecture/system/ADR-105-retire-legacy-ingest.md cli:ingest-legacy is still present after ADR-105 was enacted ════════════════════════════════════════════════════════════ -Citations: 1 errors, 2 warnings +Citations: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-defects-lint-check.out b/tests/fixtures/adr/golden/v1-defects-lint-check.out index 69bf8293..fb9648a4 100644 --- a/tests/fixtures/adr/golden/v1-defects-lint-check.out +++ b/tests/fixtures/adr/golden/v1-defects-lint-check.out @@ -1,18 +1,17 @@ -Scanned: 59 ADRs +Scanned: 41 ADRs Status distribution: Unknown: 1 - accepted: 28 + accepted: 21 draft: 1 - proposed: 27 - rejected: 1 + proposed: 17 superseded: 1 Contract: adr/v1 (0 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 49 files: +Issues found in 37 files: ──────────────────────────────────────────────────────────── docs/architecture/adr.yaml @@ -22,8 +21,6 @@ docs/architecture/adr.yaml ❌ kinds.weird.statuses: expected a list of names ❌ kinds.weird.edges: expected a mapping of field to kind or kinds ❌ verbs: expected a list of names - ❌ capability 'search' has no accepted add decision - ❌ 'archived' in lifecycle (L1) and 'archive' in basis sources (L4) share a stem; layers share no word docs/architecture/system/ADR-120-no-kind.md ❌ adr/v1 record declares no kind @@ -47,18 +44,11 @@ docs/architecture/system/ADR-126-unknown-capability.md ❌ capability 'telepathy' is not in the adr.yaml vocabulary docs/architecture/system/ADR-127-list-capability.md - ❌ capability takes one name; only a constrain decision takes a list + ❌ capability takes one name; only change and constrain decisions take a list docs/architecture/system/ADR-128-star-on-add.md ❌ only a constrain decision may be scoped to '*' -docs/architecture/system/ADR-129-retire-no-targets.md - ❌ a retire decision names its targets (targets: [cli:..., route:...]) - -docs/architecture/system/ADR-130-retire-bad-targets.md - ❌ target 'grpc:Upload': surface 'grpc' is not declared in adr.yaml - ❌ target 'plainname' is not namespace:name - docs/architecture/system/ADR-131-spec-extends.md ❌ a spec record takes no extends edge @@ -71,9 +61,6 @@ docs/architecture/system/ADR-133-amends-no-section.md docs/architecture/system/ADR-134-amends-dangling.md ❌ amends: 'ADR-199' resolves to no known ADR -docs/architecture/system/ADR-135-change-no-prior.md - ❌ a change decision supersedes or amends a prior decision on 'ingest' - docs/architecture/system/ADR-136-v0-status.md ❌ status 'Draft' is not in the decision lifecycle (proposed, accepted, rejected, abandoned, superseded, archived) @@ -94,10 +81,6 @@ docs/architecture/system/ADR-143-spec-supersedes-decision.md ⚠️ supersedes: ADR-100 does not declare the reciprocal superseded_by: ADR-143 — one-directional links rot ❌ supersedes: ADR-100 is a decision, expected spec -docs/architecture/system/ADR-145-change-supersedes-broader.md - ⚠️ supersedes: ADR-144 does not declare the reciprocal superseded_by: ADR-145 — one-directional links rot - ❌ a change on 'adr' against a broader decision amends it rather than superseding it - docs/architecture/system/ADR-150-basis-not-list.md ❌ basis: expected a list of entries, each naming one source @@ -128,9 +111,6 @@ docs/architecture/system/ADR-155-considered-bad.md ❌ considered entry 1: canary is caught or missed ❌ considered entry 2: expected a mapping with operator, said and via -docs/architecture/system/ADR-156-operator-no-considered.md - ❌ accepted with an operator basis but no considered entry; the operator considers what they started - docs/architecture/system/ADR-157-concern-bad.md ❌ concern 1: names what would resolve it ('resolve') ⚠️ open concern: No resolution named @@ -141,22 +121,6 @@ docs/architecture/system/ADR-157-concern-bad.md docs/architecture/system/ADR-158-precedent-dangling.md ❌ basis: precedent 'ADR-197' resolves to no known ADR - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-159-precedent-loop-a.md - ⚠️ basis: precedent ADR-161 is still proposed - ❌ basis: precedent loops back through ADR-159 -> ADR-161 without reaching an external source - -docs/architecture/system/ADR-161-precedent-loop-b.md - ⚠️ basis: precedent ADR-159 is still proposed - ❌ basis: precedent loops back through ADR-161 -> ADR-159 without reaching an external source - -docs/architecture/system/ADR-162-precedent-internal.md - ❌ basis: precedent ADR-101 has neither a basis nor decided_by - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-164-summary-thin.md - ⚠️ Summary lacks probes, an inversion (ADR-304 §12) docs/architecture/system/ADR-165-no-summary.md ❌ a decision record opens with a '## Summary' section @@ -168,23 +132,12 @@ docs/architecture/system/ADR-167-enacted-early.md ❌ enacted marks an accepted decision done; this one is proposed ❌ enacted: 'soon' is not a commit hash -docs/architecture/system/ADR-169-summary-less-confident.md - ⚠️ Summary lacks a probe labelled 'confident' (ADR-304 §12) - docs/architecture/system/ADR-172-dangling-beside-grounded.md ❌ basis: precedent 'ADR-196' resolves to no known ADR docs/architecture/system/ADR-173-external-and-dangling.md ❌ basis: precedent 'ADR-195' resolves to no known ADR -docs/architecture/system/ADR-177-precedent-to-rejected.md - ❌ basis: precedent ADR-176 is rejected, so it grounds nothing - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-178-downstream-of-loop.md - ⚠️ basis: precedent ADR-159 is still proposed - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - docs/architecture/system/ADR-179-said-whitespace.md ❌ concern 1: records what the concern is ('said') ⚠️ open concern: (no said) @@ -202,7 +155,7 @@ docs/architecture/system/ADR-182-precedent-list.md ❌ basis entry 2: evidence needs a reference ════════════════════════════════════════════════════════════ -Summary: 78 errors, 10 warnings +Summary: 62 errors, 4 warnings ════════════════════════════════════════════════════════════ [exit 1] diff --git a/tests/fixtures/adr/golden/v1-defects-lint.out b/tests/fixtures/adr/golden/v1-defects-lint.out index ec05c5e9..39571601 100644 --- a/tests/fixtures/adr/golden/v1-defects-lint.out +++ b/tests/fixtures/adr/golden/v1-defects-lint.out @@ -1,18 +1,17 @@ -Scanned: 59 ADRs +Scanned: 41 ADRs Status distribution: Unknown: 1 - accepted: 28 + accepted: 21 draft: 1 - proposed: 27 - rejected: 1 + proposed: 17 superseded: 1 Contract: adr/v1 (0 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 49 files: +Issues found in 37 files: ──────────────────────────────────────────────────────────── docs/architecture/adr.yaml @@ -22,8 +21,6 @@ docs/architecture/adr.yaml ❌ kinds.weird.statuses: expected a list of names ❌ kinds.weird.edges: expected a mapping of field to kind or kinds ❌ verbs: expected a list of names - ❌ capability 'search' has no accepted add decision - ❌ 'archived' in lifecycle (L1) and 'archive' in basis sources (L4) share a stem; layers share no word docs/architecture/system/ADR-120-no-kind.md ❌ adr/v1 record declares no kind @@ -47,18 +44,11 @@ docs/architecture/system/ADR-126-unknown-capability.md ❌ capability 'telepathy' is not in the adr.yaml vocabulary docs/architecture/system/ADR-127-list-capability.md - ❌ capability takes one name; only a constrain decision takes a list + ❌ capability takes one name; only change and constrain decisions take a list docs/architecture/system/ADR-128-star-on-add.md ❌ only a constrain decision may be scoped to '*' -docs/architecture/system/ADR-129-retire-no-targets.md - ❌ a retire decision names its targets (targets: [cli:..., route:...]) - -docs/architecture/system/ADR-130-retire-bad-targets.md - ❌ target 'grpc:Upload': surface 'grpc' is not declared in adr.yaml - ❌ target 'plainname' is not namespace:name - docs/architecture/system/ADR-131-spec-extends.md ❌ a spec record takes no extends edge @@ -71,9 +61,6 @@ docs/architecture/system/ADR-133-amends-no-section.md docs/architecture/system/ADR-134-amends-dangling.md ❌ amends: 'ADR-199' resolves to no known ADR -docs/architecture/system/ADR-135-change-no-prior.md - ❌ a change decision supersedes or amends a prior decision on 'ingest' - docs/architecture/system/ADR-136-v0-status.md ❌ status 'Draft' is not in the decision lifecycle (proposed, accepted, rejected, abandoned, superseded, archived) @@ -94,10 +81,6 @@ docs/architecture/system/ADR-143-spec-supersedes-decision.md ⚠️ supersedes: ADR-100 does not declare the reciprocal superseded_by: ADR-143 — one-directional links rot ❌ supersedes: ADR-100 is a decision, expected spec -docs/architecture/system/ADR-145-change-supersedes-broader.md - ⚠️ supersedes: ADR-144 does not declare the reciprocal superseded_by: ADR-145 — one-directional links rot - ❌ a change on 'adr' against a broader decision amends it rather than superseding it - docs/architecture/system/ADR-150-basis-not-list.md ❌ basis: expected a list of entries, each naming one source @@ -128,9 +111,6 @@ docs/architecture/system/ADR-155-considered-bad.md ❌ considered entry 1: canary is caught or missed ❌ considered entry 2: expected a mapping with operator, said and via -docs/architecture/system/ADR-156-operator-no-considered.md - ❌ accepted with an operator basis but no considered entry; the operator considers what they started - docs/architecture/system/ADR-157-concern-bad.md ❌ concern 1: names what would resolve it ('resolve') ⚠️ open concern: No resolution named @@ -141,22 +121,6 @@ docs/architecture/system/ADR-157-concern-bad.md docs/architecture/system/ADR-158-precedent-dangling.md ❌ basis: precedent 'ADR-197' resolves to no known ADR - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-159-precedent-loop-a.md - ⚠️ basis: precedent ADR-161 is still proposed - ❌ basis: precedent loops back through ADR-159 -> ADR-161 without reaching an external source - -docs/architecture/system/ADR-161-precedent-loop-b.md - ⚠️ basis: precedent ADR-159 is still proposed - ❌ basis: precedent loops back through ADR-161 -> ADR-159 without reaching an external source - -docs/architecture/system/ADR-162-precedent-internal.md - ❌ basis: precedent ADR-101 has neither a basis nor decided_by - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-164-summary-thin.md - ⚠️ Summary lacks probes, an inversion (ADR-304 §12) docs/architecture/system/ADR-165-no-summary.md ❌ a decision record opens with a '## Summary' section @@ -168,23 +132,12 @@ docs/architecture/system/ADR-167-enacted-early.md ❌ enacted marks an accepted decision done; this one is proposed ❌ enacted: 'soon' is not a commit hash -docs/architecture/system/ADR-169-summary-less-confident.md - ⚠️ Summary lacks a probe labelled 'confident' (ADR-304 §12) - docs/architecture/system/ADR-172-dangling-beside-grounded.md ❌ basis: precedent 'ADR-196' resolves to no known ADR docs/architecture/system/ADR-173-external-and-dangling.md ❌ basis: precedent 'ADR-195' resolves to no known ADR -docs/architecture/system/ADR-177-precedent-to-rejected.md - ❌ basis: precedent ADR-176 is rejected, so it grounds nothing - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - -docs/architecture/system/ADR-178-downstream-of-loop.md - ⚠️ basis: precedent ADR-159 is still proposed - ❌ basis: following precedent never reaches an external source (operator, evidence, standard, archive) - docs/architecture/system/ADR-179-said-whitespace.md ❌ concern 1: records what the concern is ('said') ⚠️ open concern: (no said) @@ -202,7 +155,7 @@ docs/architecture/system/ADR-182-precedent-list.md ❌ basis entry 2: evidence needs a reference ════════════════════════════════════════════════════════════ -Summary: 78 errors, 10 warnings +Summary: 62 errors, 4 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-empty-new.out b/tests/fixtures/adr/golden/v1-empty-new.out new file mode 100644 index 00000000..783029da --- /dev/null +++ b/tests/fixtures/adr/golden/v1-empty-new.out @@ -0,0 +1,2 @@ +Error: adr.yaml declares contract adr/v1 but no kinds; `adr lint` reports the config +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-evidence-kind.out b/tests/fixtures/adr/golden/v1-evidence-kind.out new file mode 100644 index 00000000..3665e985 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-evidence-kind.out @@ -0,0 +1,23 @@ + +Scanned: 2 ADRs + +Status distribution: + accepted: 1 + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-118-raise-the-ingest-batch-size.md + ❌ basis entry 2: evidence ADR-101 is a decision; evidence cites a spec or evidence record + ❌ basis entry 3: evidence ADR-999 resolves to no record + ⚠️ supersedes: ADR-103 does not declare the reciprocal superseded_by: ADR-118 — one-directional links rot + +════════════════════════════════════════════════════════════ +Summary: 2 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-baseline-change-names-prior.out b/tests/fixtures/adr/golden/v1-lint-baseline-change-names-prior.out deleted file mode 100644 index d852dfdb..00000000 --- a/tests/fixtures/adr/golden/v1-lint-baseline-change-names-prior.out +++ /dev/null @@ -1,30 +0,0 @@ - -Scanned: 18 ADRs - -Status distribution: - accepted: 9 - proposed: 9 - -Contract: adr/v1 (1 v0 records remain) - -──────────────────────────────────────────────────────────── -Issues found in 4 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md - ⚠️ open concern: Batch size may starve small tenants - -docs/architecture/system/ADR-118-compress-exports.md - ⚠️ supersedes: ADR-119 does not declare the reciprocal superseded_by: ADR-118 — one-directional links rot - -════════════════════════════════════════════════════════════ -Summary: 0 errors, 4 warnings -════════════════════════════════════════════════════════════ - -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-baseline-change-rejected-prior.out b/tests/fixtures/adr/golden/v1-lint-baseline-change-rejected-prior.out deleted file mode 100644 index a9dc0d54..00000000 --- a/tests/fixtures/adr/golden/v1-lint-baseline-change-rejected-prior.out +++ /dev/null @@ -1,28 +0,0 @@ - -Scanned: 18 ADRs - -Status distribution: - accepted: 9 - proposed: 8 - rejected: 1 - -Contract: adr/v1 (1 v0 records remain) - -──────────────────────────────────────────────────────────── -Issues found in 3 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md - ⚠️ open concern: Batch size may starve small tenants - -════════════════════════════════════════════════════════════ -Summary: 0 errors, 3 warnings -════════════════════════════════════════════════════════════ - -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-baseline-change.out b/tests/fixtures/adr/golden/v1-lint-baseline-change.out deleted file mode 100644 index a6555279..00000000 --- a/tests/fixtures/adr/golden/v1-lint-baseline-change.out +++ /dev/null @@ -1,30 +0,0 @@ - -Scanned: 18 ADRs - -Status distribution: - accepted: 9 - proposed: 9 - -Contract: adr/v1 (1 v0 records remain) - -──────────────────────────────────────────────────────────── -Issues found in 4 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md - ⚠️ open concern: Batch size may starve small tenants - -docs/architecture/system/ADR-118-compress-exports.md - ❌ a change decision supersedes or amends a prior decision on 'export' - -════════════════════════════════════════════════════════════ -Summary: 1 errors, 3 warnings -════════════════════════════════════════════════════════════ - -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-baseline-list.out b/tests/fixtures/adr/golden/v1-lint-baseline-list.out deleted file mode 100644 index 0b521ce3..00000000 --- a/tests/fixtures/adr/golden/v1-lint-baseline-list.out +++ /dev/null @@ -1,29 +0,0 @@ - -Scanned: 15 ADRs - -Status distribution: - accepted: 9 - proposed: 6 - -Contract: adr/v1 (1 v0 records remain) - -──────────────────────────────────────────────────────────── -Issues found in 3 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ❌ baseline: expected 'adopted' (a YYYY-MM-DD date) and 'capabilities' (a list of names) - ⚠️ capability 'search' has no accepted add decision - ⚠️ capability 'export' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md - ⚠️ open concern: Batch size may starve small tenants - -════════════════════════════════════════════════════════════ -Summary: 1 errors, 4 warnings -════════════════════════════════════════════════════════════ - -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-baseline-unknown.out b/tests/fixtures/adr/golden/v1-lint-baseline-unknown.out deleted file mode 100644 index 0b300228..00000000 --- a/tests/fixtures/adr/golden/v1-lint-baseline-unknown.out +++ /dev/null @@ -1,29 +0,0 @@ - -Scanned: 15 ADRs - -Status distribution: - accepted: 9 - proposed: 6 - -Contract: adr/v1 (1 v0 records remain) - -──────────────────────────────────────────────────────────── -Issues found in 3 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ❌ baseline.adopted: expected a YYYY-MM-DD date - ❌ baseline: 'exprt' is not in the capabilities vocabulary - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - -docs/architecture/system/ADR-114-open-concern.md - ⚠️ open concern: Batch size may starve small tenants - -════════════════════════════════════════════════════════════ -Summary: 2 errors, 3 warnings -════════════════════════════════════════════════════════════ - -[exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-check.out b/tests/fixtures/adr/golden/v1-lint-check.out index e9a6b42c..1887d444 100644 --- a/tests/fixtures/adr/golden/v1-lint-check.out +++ b/tests/fixtures/adr/golden/v1-lint-check.out @@ -8,20 +8,14 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 3 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - docs/architecture/system/ADR-114-open-concern.md ⚠️ open concern: Batch size may starve small tenants ════════════════════════════════════════════════════════════ -Summary: 0 errors, 3 warnings +Summary: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint-precedent-relative.out b/tests/fixtures/adr/golden/v1-lint-precedent-relative.out index 6b1e7995..e162c566 100644 --- a/tests/fixtures/adr/golden/v1-lint-precedent-relative.out +++ b/tests/fixtures/adr/golden/v1-lint-precedent-relative.out @@ -6,15 +6,8 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) -──────────────────────────────────────────────────────────── -Issues found in 1 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - ════════════════════════════════════════════════════════════ -Summary: 0 errors, 1 warnings +Summary: 0 errors, 0 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-lint.out b/tests/fixtures/adr/golden/v1-lint.out index e9a6b42c..1887d444 100644 --- a/tests/fixtures/adr/golden/v1-lint.out +++ b/tests/fixtures/adr/golden/v1-lint.out @@ -8,20 +8,14 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 3 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - -docs/architecture/system/ADR-108-ingest-over-v0.md - ⚠️ cannot confirm the prior decision on 'ingest': ADR-110 is still v0 - docs/architecture/system/ADR-114-open-concern.md ⚠️ open concern: Batch size may starve small tenants ════════════════════════════════════════════════════════════ -Summary: 0 errors, 3 warnings +Summary: 0 errors, 1 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-list.out b/tests/fixtures/adr/golden/v1-list.out index ffbeac30..6e753c73 100644 --- a/tests/fixtures/adr/golden/v1-list.out +++ b/tests/fixtures/adr/golden/v1-list.out @@ -1,5 +1,5 @@ -ADR v1 Fixture — Architecture Decision Records (15 total) +ADR v1 Fixture — Agent Decision Records (15 total) ======================================================= ✅ ADR-100 Adopt the adr/v1 contract ✅ ADR-101 Add ingestion diff --git a/tests/fixtures/adr/golden/v1-new-bare-file.md b/tests/fixtures/adr/golden/v1-new-bare-file.md new file mode 100644 index 00000000..9ff60f8a --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-bare-file.md @@ -0,0 +1,54 @@ +path: docs/architecture/system/ADR-117-bare-decision.md +--- +contract: adr/v1 +kind: decision +verb: ~ +capability: ~ +basis: [] +agent: + name: ~ + model: ~ +status: proposed +date: <TODAY> +deciders: + - developer + - agent +related: [] +--- + +# ADR-117: Bare decision + +## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] + +## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] diff --git a/tests/fixtures/adr/golden/v1-new-bare-lint.out b/tests/fixtures/adr/golden/v1-new-bare-lint.out new file mode 100644 index 00000000..5845c2cc --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-bare-lint.out @@ -0,0 +1,25 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-117-bare-decision.md + ❌ a decision record requires a verb (add, cut, change, retire, constrain) + ❌ a decision record requires 'capability' + ❌ a decision record requires 'basis' + ❌ agent: records 'name' + ❌ agent: records 'model' + ⚠️ 12 placeholder line(s) from `adr new` still to fill, first: - **Decided:** [what is decided, in plain terms] + +════════════════════════════════════════════════════════════ +Summary: 5 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-bare.out b/tests/fixtures/adr/golden/v1-new-bare.out new file mode 100644 index 00000000..540589b1 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-bare.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-117-bare-decision.md + Domain: System (system) + Number: ADR-117 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-custom-kind-file.md b/tests/fixtures/adr/golden/v1-new-custom-kind-file.md new file mode 100644 index 00000000..650504ca --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-custom-kind-file.md @@ -0,0 +1,51 @@ +path: docs/architecture/system/ADR-115-retention-policy.md +--- +contract: adr/v1 +kind: policy +verb: constrain +capability: ingest +targets: [] +status: proposed +date: <TODAY> +deciders: + - developer + - agent +related: [] +--- + +# ADR-115: Retention policy + +## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] + +## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] diff --git a/tests/fixtures/adr/golden/v1-new-custom-kind.out b/tests/fixtures/adr/golden/v1-new-custom-kind.out new file mode 100644 index 00000000..9852a6fe --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-custom-kind.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-115-retention-policy.md + Domain: System (system) + Number: ADR-115 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-decision-file.md b/tests/fixtures/adr/golden/v1-new-decision-file.md new file mode 100644 index 00000000..34eef5c7 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-decision-file.md @@ -0,0 +1,54 @@ +path: docs/architecture/system/ADR-115-stream-exports.md +--- +contract: adr/v1 +kind: decision +verb: change +capability: ingest +basis: [] +agent: + name: Claude + model: fixture-model +status: proposed +date: <TODAY> +deciders: + - developer + - agent +related: [] +--- + +# ADR-115: Stream exports + +## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] + +## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] diff --git a/tests/fixtures/adr/golden/v1-frozen-renamed.out b/tests/fixtures/adr/golden/v1-new-decision-lint.out similarity index 73% rename from tests/fixtures/adr/golden/v1-frozen-renamed.out rename to tests/fixtures/adr/golden/v1-new-decision-lint.out index 8ae78792..1111276e 100644 --- a/tests/fixtures/adr/golden/v1-frozen-renamed.out +++ b/tests/fixtures/adr/golden/v1-new-decision-lint.out @@ -2,19 +2,17 @@ Scanned: 1 ADRs Status distribution: - accepted: 1 + proposed: 1 Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 2 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'ingest' has no accepted add decision - -docs/architecture/system/ADR-101-ingest-renamed.md - ❌ 'capability' changed after the decision left proposed; only concern, considered, enacted, status, superseded_by may change +docs/architecture/system/ADR-115-stream-exports.md + ❌ a decision record requires 'basis' + ⚠️ 12 placeholder line(s) from `adr new` still to fill, first: - **Decided:** [what is decided, in plain terms] ════════════════════════════════════════════════════════════ Summary: 1 errors, 1 warnings diff --git a/tests/fixtures/adr/golden/v1-new-decision.out b/tests/fixtures/adr/golden/v1-new-decision.out new file mode 100644 index 00000000..eb03d6ac --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-decision.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-115-stream-exports.md + Domain: System (system) + Number: ADR-115 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-multiline-file.md b/tests/fixtures/adr/golden/v1-new-multiline-file.md new file mode 100644 index 00000000..d9c86b6a --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-multiline-file.md @@ -0,0 +1,54 @@ +path: docs/architecture/system/ADR-118-multiline-model.md +--- +contract: adr/v1 +kind: decision +verb: add +capability: ingest +basis: [] +agent: + name: Claude + model: "a\n---\nb" +status: proposed +date: <TODAY> +deciders: + - developer + - agent +related: [] +--- + +# ADR-118: Multiline model + +## Summary + +- **Decided:** [what is decided, in plain terms] +- **Trades away:** [what it gives up or forecloses] +- **One-way?** [yes or no, and why] +- **Probes:** *Confident:* [a point you are sure of]. *Not confident:* [a point you are not]. +- **Inversion:** [the two ends this sits between; is the answer outside that framing?] + +## Context + +[What is the issue that we're seeing that is motivating this decision or change?] + +## Decision + +[What is the change that we're proposing and/or doing?] + +## Consequences + +### Positive + +- [What becomes easier?] + +### Negative + +- [What becomes harder?] + +### Neutral + +- [What other changes does this enable or require?] + +## Alternatives Considered + +- [What other options were evaluated?] +- [Why were they rejected?] diff --git a/tests/fixtures/adr/golden/v1-new-multiline-lint.out b/tests/fixtures/adr/golden/v1-new-multiline-lint.out new file mode 100644 index 00000000..d016fce6 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-multiline-lint.out @@ -0,0 +1,21 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-118-multiline-model.md + ❌ a decision record requires 'basis' + ⚠️ 12 placeholder line(s) from `adr new` still to fill, first: - **Decided:** [what is decided, in plain terms] + +════════════════════════════════════════════════════════════ +Summary: 1 errors, 1 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-multiline.out b/tests/fixtures/adr/golden/v1-new-multiline.out new file mode 100644 index 00000000..347d001d --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-multiline.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-118-multiline-model.md + Domain: System (system) + Number: ADR-118 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-placeholder-accepted.out b/tests/fixtures/adr/golden/v1-new-placeholder-accepted.out new file mode 100644 index 00000000..c5082038 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-placeholder-accepted.out @@ -0,0 +1,21 @@ + +Scanned: 1 ADRs + +Status distribution: + accepted: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-115-stream-exports.md + ❌ a decision record requires 'basis' + ❌ 12 placeholder line(s) from `adr new` still to fill, first: - **Decided:** [what is decided, in plain terms] + +════════════════════════════════════════════════════════════ +Summary: 2 errors, 0 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-refused.out b/tests/fixtures/adr/golden/v1-new-refused.out new file mode 100644 index 00000000..aa411f7d --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-refused.out @@ -0,0 +1,4 @@ +Error: a spec record takes no verb +Error: capability 'nosuch' is not in the adr.yaml vocabulary +Error: a spec record carries no agent +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-new-spec-file.md b/tests/fixtures/adr/golden/v1-new-spec-file.md new file mode 100644 index 00000000..ea0e11d4 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-spec-file.md @@ -0,0 +1,14 @@ +path: docs/architecture/system/ADR-116-export-format.md +--- +contract: adr/v1 +kind: spec +capability: ingest +status: proposed +date: <TODAY> +deciders: + - developer + - agent +related: [] +--- + +# ADR-116: Export format diff --git a/tests/fixtures/adr/golden/v1-new-spec-lint.out b/tests/fixtures/adr/golden/v1-new-spec-lint.out new file mode 100644 index 00000000..e162c566 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-spec-lint.out @@ -0,0 +1,13 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 0 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-spec.out b/tests/fixtures/adr/golden/v1-new-spec.out new file mode 100644 index 00000000..673ddb86 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-spec.out @@ -0,0 +1,5 @@ +Created: docs/architecture/system/ADR-116-export-format.md + Domain: System (system) + Number: ADR-116 + Contract: adr/v1 (fill the empty fields; `adr lint` lists them) +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-new-unknown-kind-status.txt b/tests/fixtures/adr/golden/v1-new-unknown-kind-status.txt new file mode 100644 index 00000000..2653c19d --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-unknown-kind-status.txt @@ -0,0 +1,3 @@ +A docs/architecture/system/ADR-115-stream-exports.md +A docs/architecture/system/ADR-116-export-format.md +A docs/architecture/system/ADR-117-bare-decision.md diff --git a/tests/fixtures/adr/golden/v1-new-unknown-kind.out b/tests/fixtures/adr/golden/v1-new-unknown-kind.out new file mode 100644 index 00000000..07c9cecd --- /dev/null +++ b/tests/fixtures/adr/golden/v1-new-unknown-kind.out @@ -0,0 +1,2 @@ +Error: Unknown kind 'policy'. Kinds: decision, spec +[exit 1] diff --git a/tests/fixtures/adr/golden/v1-observable-added.out b/tests/fixtures/adr/golden/v1-observable-added.out new file mode 100644 index 00000000..a2068ef3 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-observable-added.out @@ -0,0 +1,13 @@ + +Scanned: 1 ADRs + +Status distribution: + accepted: 1 + +Contract: adr/v1 (1 v0 records remain) + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 0 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-frozen-migrated.out b/tests/fixtures/adr/golden/v1-observable-empty.out similarity index 85% rename from tests/fixtures/adr/golden/v1-frozen-migrated.out rename to tests/fixtures/adr/golden/v1-observable-empty.out index 762d09ce..aee96c3e 100644 --- a/tests/fixtures/adr/golden/v1-frozen-migrated.out +++ b/tests/fixtures/adr/golden/v1-observable-empty.out @@ -4,14 +4,14 @@ Scanned: 1 ADRs Status distribution: accepted: 1 -Contract: adr/v1 (0 v0 records remain) +Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ❌ capability 'search' has no accepted add decision +docs/architecture/system/ADR-101-ingest.md + ❌ observable: empty; remove the key or add an entry ════════════════════════════════════════════════════════════ Summary: 1 errors, 0 warnings diff --git a/tests/fixtures/adr/golden/v1-frozen-lint.out b/tests/fixtures/adr/golden/v1-observable-malformed.out similarity index 71% rename from tests/fixtures/adr/golden/v1-frozen-lint.out rename to tests/fixtures/adr/golden/v1-observable-malformed.out index 06219faa..9a158a49 100644 --- a/tests/fixtures/adr/golden/v1-frozen-lint.out +++ b/tests/fixtures/adr/golden/v1-observable-malformed.out @@ -7,18 +7,16 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) ──────────────────────────────────────────────────────────── -Issues found in 2 files: +Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/adr.yaml - ⚠️ capability 'ingest' has no accepted add decision - docs/architecture/system/ADR-101-ingest.md - ❌ 'capability' changed after the decision left proposed; only concern, considered, enacted, status, superseded_by may change - ⚠️ body edited after the decision left proposed; a decision grows by appending + ❌ observable entry 3: expected a line of words or a non-empty mapping + ❌ observable entry 4: expected a line of words or a non-empty mapping + ❌ observable entry 5: expected a line of words or a non-empty mapping ════════════════════════════════════════════════════════════ -Summary: 1 errors, 2 warnings +Summary: 3 errors, 0 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-frozen-merged-branch-edited.out b/tests/fixtures/adr/golden/v1-observable-not-list.out similarity index 83% rename from tests/fixtures/adr/golden/v1-frozen-merged-branch-edited.out rename to tests/fixtures/adr/golden/v1-observable-not-list.out index 96f5cc98..0624004c 100644 --- a/tests/fixtures/adr/golden/v1-frozen-merged-branch-edited.out +++ b/tests/fixtures/adr/golden/v1-observable-not-list.out @@ -10,8 +10,8 @@ Contract: adr/v1 (1 v0 records remain) Issues found in 1 files: ──────────────────────────────────────────────────────────── -docs/architecture/system/ADR-113-operator-proposed.md - ❌ 'basis' changed after the decision left proposed; only concern, considered, enacted, status, superseded_by may change +docs/architecture/system/ADR-101-ingest.md + ❌ observable: expected a list of entries, each a line of words or a mapping ════════════════════════════════════════════════════════════ Summary: 1 errors, 0 warnings diff --git a/tests/fixtures/adr/golden/v1-reject-precedent.out b/tests/fixtures/adr/golden/v1-reject-precedent.out index 94888acf..5074568f 100644 --- a/tests/fixtures/adr/golden/v1-reject-precedent.out +++ b/tests/fixtures/adr/golden/v1-reject-precedent.out @@ -1,4 +1,3 @@ -Refused: rejecting ADR-114 would add errors: - ❌ docs/architecture/system/ADR-116-rests-on-114.md: basis: following precedent never reaches an external source (operator, evidence, standard, upstream) - ❌ docs/architecture/system/ADR-116-rests-on-114.md: basis: precedent ADR-114 is rejected, so it grounds nothing -[exit 1] +Marked ADR-114 rejected: Cap batch size +Run `adr index -y` to refresh INDEX.md. +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-reject-then-lint.out b/tests/fixtures/adr/golden/v1-reject-then-lint.out index e8a78fe5..82a854f3 100644 --- a/tests/fixtures/adr/golden/v1-reject-then-lint.out +++ b/tests/fixtures/adr/golden/v1-reject-then-lint.out @@ -6,15 +6,8 @@ Status distribution: Contract: adr/v1 (1 v0 records remain) -──────────────────────────────────────────────────────────── -Issues found in 1 files: -──────────────────────────────────────────────────────────── - -docs/architecture/adr.yaml - ⚠️ capability 'search' has no accepted add decision - ════════════════════════════════════════════════════════════ -Summary: 0 errors, 1 warnings +Summary: 0 errors, 0 warnings ════════════════════════════════════════════════════════════ [exit 0] diff --git a/tests/fixtures/adr/golden/v1-summary-colon-heading.out b/tests/fixtures/adr/golden/v1-summary-colon-heading.out new file mode 100644 index 00000000..e162c566 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-summary-colon-heading.out @@ -0,0 +1,13 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +════════════════════════════════════════════════════════════ +Summary: 0 errors, 0 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/v1-summary-prefix-heading.out b/tests/fixtures/adr/golden/v1-summary-prefix-heading.out new file mode 100644 index 00000000..cf5808c9 --- /dev/null +++ b/tests/fixtures/adr/golden/v1-summary-prefix-heading.out @@ -0,0 +1,20 @@ + +Scanned: 1 ADRs + +Status distribution: + proposed: 1 + +Contract: adr/v1 (1 v0 records remain) + +──────────────────────────────────────────────────────────── +Issues found in 1 files: +──────────────────────────────────────────────────────────── + +docs/architecture/system/ADR-108-ingest-over-v0.md + ❌ a decision record opens with a '## Summary' section + +════════════════════════════════════════════════════════════ +Summary: 1 errors, 0 warnings +════════════════════════════════════════════════════════════ + +[exit 0] diff --git a/tests/fixtures/adr/golden/version.out b/tests/fixtures/adr/golden/version.out index 77770ee1..5c098c9c 100644 --- a/tests/fixtures/adr/golden/version.out +++ b/tests/fixtures/adr/golden/version.out @@ -1,2 +1,2 @@ -adr-tool 2.0.0 +adr-tool 2.1.0 [exit 0] diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/INDEX.md b/tests/fixtures/adr/v0-corpus/docs/architecture/INDEX.md new file mode 100644 index 00000000..2cbeaee3 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/INDEX.md @@ -0,0 +1,150 @@ +# Architecture Decision Records + +This directory contains Architecture Decision Records (ADRs) for Claude Code Configuration. +Each ADR documents a significant architectural decision, its context, and consequences. + +## ADR Format + +All ADRs follow a consistent format: +- **Status:** Draft / Proposed / Accepted / Deprecated / Superseded +- **Date:** When the decision was made +- **Deciders:** Who made the decision +- **Context:** The problem or situation requiring a decision +- **Decision:** The architectural choice made +- **Consequences:** Benefits, drawbacks, and other impacts + +_This index is auto-generated by `adr index`. Configuration: [`adr.yaml`](./adr.yaml)_ + +## System +_Ways architecture, matching, macros, hooks, session lifecycle_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-100](./system/ADR-100-ways-scaffolding-wizard.md) | Ways Scaffolding Wizard | Accepted | +| [ADR-101](./system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md) | Wormhole relay protocol for cross-instance agent communication | Deprecated | +| [ADR-102](./system/ADR-102-irc-based-local-agent-communication.md) | IRC-based local agent communication | Deprecated | +| [ADR-103](./system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md) | Checks — Epoch-Distance-Aware Confidence Sensors for Ways | Accepted | +| [ADR-104](./system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md) | Token-Gated Way Re-Disclosure for Long Context Windows | Superseded (superseded by ADR-123) | +| [ADR-105](./system/ADR-105-progressive-disclosure-for-way-trees.md) | Progressive Disclosure for Way Trees | Accepted | +| [ADR-106](./system/ADR-106-project-pulse-epoch-mapped-project-awareness.md) | Project Pulse — Epoch-Mapped Project Awareness | Accepted | +| [ADR-107](./system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md) | Corpus, Matching Pipeline, and Locale Support | Accepted (partially superseded by ADR-125) | +| [ADR-108](./system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md) | Embedding-Based Way Matching with all-MiniLM-L6-v2 | Accepted | +| [ADR-109](./system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md) | Project-Scope Way Embedding with Manifest-Based Staleness Detection | Accepted | +| [ADR-110](./system/ADR-110-way-file-separation-and-graph-compatible-structure.md) | Way File Separation and Graph-Compatible Structure | Accepted | +| [ADR-111](./system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md) | Unified `ways` CLI — Single Binary Tool Consolidation | Accepted | +| [ADR-113](./system/ADR-113-attend-active-awareness-module.md) | `attend` — Active Awareness Module | Accepted | +| [ADR-114](./system/ADR-114-attend-as-insistent-way-trigger-type.md) | `attend` Events as an Insistent Way Trigger Type | Accepted | +| [ADR-115](./system/ADR-115-declarative-config-with-project-scope-overlay.md) | Declarative Configuration with Project-Scope Overlay | Accepted | +| [ADR-116](./system/ADR-116-declarative-permission-requirements.md) | Declarative Permission Requirements | Accepted | +| [ADR-117](./system/ADR-117-sensor-crate-extraction-and-feature-flags.md) | Sensor Crate Extraction and Feature Flags | Accepted | +| [ADR-118](./system/ADR-118-focus-groups-dynamic-agent-grouping.md) | Focus Groups — Dynamic Agent Grouping | Accepted | +| [ADR-119](./system/ADR-119-action-potential-engagement-model.md) | Action Potential Engagement Model | Superseded (superseded by ADR-123) | +| [ADR-120](./system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md) | Interactive Chat TUI — Human in the Signal Loop | Accepted | +| [ADR-121](./system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md) | Salience decay for signal presentation — turn-based exponential | Superseded (superseded by ADR-123) | +| [ADR-122](./system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md) | Attend disclosure sensor — token-gated affordance reheat | Accepted | +| [ADR-123](./system/ADR-123-firing-dynamics-progression-axis-unification.md) | Firing dynamics — progression-axis unification for attend and ways | Accepted | +| [ADR-124](./system/ADR-124-channel-bar-ordering-open-as-base.md) | TUI Legend Architecture — Base Channel, Liveness, and Ordering | Accepted | +| [ADR-125](./system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md) | Authored Disclosure Graph and Removal of BM25 | Accepted | +| [ADR-126](./system/ADR-126-window-relative-refire.md) | Window-relative refire with named presets | Accepted | +| [ADR-127](./system/ADR-127-reject-full-body-embedding-corpus.md) | Full-body embedding corpus for way matching | Rejected | +| [ADR-128](./system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md) | Memory as repo-portable ways — seed routing over accumulated snapshots | Accepted | +| [ADR-129](./system/ADR-129-instance-suffix-and-heartbeat-liveness.md) | Instance suffix and heartbeat liveness for attend identity | Accepted | +| [ADR-130](./system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md) | Sentence-salience input reduction for embed matching | Accepted | +| [ADR-131](./system/ADR-131-project-scope-way-toggles.md) | Project-scope way toggles | Accepted | +| [ADR-132](./system/ADR-132-collaboration-ways-domain.md) | Collaboration ways domain | Accepted | +| [ADR-134](./system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md) | Empirical auto-tuning from fire and near-miss telemetry | Accepted | +| [ADR-135](./system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md) | Content-aware write-time over-build gate with a self-extending pattern corpus | Accepted | +| [ADR-136](./system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md) | Split addressed messaging from the sensor-observation bus | Accepted | +| [ADR-137](./system/ADR-137-boundedness-bounded-work-per-cycle.md) | Boundedness — a unit of work must be bounded within its cycle | Accepted | +| [ADR-138](./system/ADR-138-skills-own-the-how-ways-own-the-5w.md) | Skills own the how, ways own the 5W | Accepted | +| [ADR-139](./system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md) | Shelve maintainer i18n; adopter-run localization via ways-localize | Accepted | +| [ADR-140](./system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md) | Two install topologies: in-place repo and subdirectory projection | Accepted | +| [ADR-141](./system/ADR-141-knowledge-graph-as-evidential-memory-backend.md) | Knowledge Graph as Evidential Memory Backend | Accepted | +| [ADR-142](./system/ADR-142-agent-ways-1-0-xdg-application-distribution.md) | agent-ways 1.0 — XDG application distribution | Accepted | +| [ADR-143](./system/ADR-143-three-root-way-runtime-core-user-project.md) | Three-root way runtime — core, user, project | Accepted | +| [ADR-144](./system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md) | Install / repair / migrate as one manifest reconciler | Accepted | +| [ADR-146](./system/ADR-146-installer-binary-verification-and-guided-build-fallback.md) | installer binary verification and guided build fallback | Accepted | +| [ADR-147](./system/ADR-147-composable-settings-json-config-fragments.md) | Composable settings.json — a store of YAML config fragments | Superseded (superseded by ADR-169) | +| [ADR-148](./system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md) | framework surface ships operator content; dev harness in project scope | Accepted | +| [ADR-149](./system/ADR-149-operator-config-interview-skill.md) | operator config interview skill | Superseded (superseded by ADR-169) | +| [ADR-150](./system/ADR-150-version-truth-and-downgrade-safe-self-update.md) | Version-truth and downgrade-safe self-update | Accepted | +| [ADR-151](./system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md) | Extract ways-core crate and ways-audit sibling binary | Accepted | +| [ADR-152](./system/ADR-152-framework-default-secret-path-deny-baseline.md) | Framework-default secret-path deny baseline | Accepted | +| [ADR-153](./system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md) | Session-introspection substrate — correlating fired ways to turns | Accepted | +| [ADR-154](./system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md) | `ways introspect` — one model, three front-ends | Accepted | +| [ADR-155](./system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md) | Semantic gating of the keyword channel and reasoning-channel rebuild | Accepted (partially superseded by ADR-188 §3) | +| [ADR-156](./system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md) | Calibrated relevance scoring for the semantic lane | Accepted | +| [ADR-157](./system/ADR-157-case-insensitive-trigger-regex-compilation.md) | Case-insensitive trigger regex compilation | Accepted | +| [ADR-158](./system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md) | Calibration boundary quality — hard negatives and a fire-breadth ship gate | Accepted | +| [ADR-159](./system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md) | Remove ways tune-curves and the legacy curve: cadence field | Accepted | +| [ADR-160](./system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md) | Chunked late-interaction matching with softmax-share gating for way selection | Accepted | +| [ADR-161](./system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md) | Queued mid-turn operator messages as an aggregated scan surface | Accepted | +| [ADR-162](./system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md) | Mechanical session-link suppression as defense against transcript disclosure | Superseded (superseded by ADR-167) | +| [ADR-163](./system/ADR-163-config-separation-dotfiles-source-of-truth.md) | Config separation — dotfiles as source-of-truth feeding the settings fragment store | Accepted | +| [ADR-164](./system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md) | File artifacts distributed across hosts must be carried by value not host-absolute reference | Accepted | +| [ADR-165](./system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md) | Loop-control bookends: start, develop, merge, release, wrap | Accepted | +| [ADR-166](./system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md) | Single source of truth for model context-window resolution | Accepted | +| [ADR-167](./system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md) | Session-link suppression: attribution.sessionUrl as primary control, deny hook as backstop | Accepted | +| [ADR-169](./system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md) | agent-ways relinquishes user-scoped settings.json; retains only its operational baseline | Accepted | +| [ADR-170](./system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md) | Human focus-group membership via username identity and a shared attend-groups crate | Accepted | +| [ADR-171](./system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md) | Stable session identity — the roster enumerates addressable coordinating units | Accepted | +| [ADR-172](./system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md) | Turn-boundary inbound delivery via a CLI-owned drain checkpoint | Accepted | +| [ADR-173](./system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md) | Chat-idiom convergence for the attend command surfaces | Accepted | +| [ADR-174](./system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md) | Progressive core — decoration guidance and the core re-disclosure gap | Accepted | +| [ADR-175](./system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md) | Standing delegation authorization satisfies the harness permission gate rather than overriding it | Accepted | +| [ADR-176](./system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md) | Contract identification as the develop-loop front gate | Accepted | +| [ADR-177](./system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md) | Version-stamped vendored tools with direction-aware drift detection | Accepted | +| [ADR-178](./system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md) | Register transfers by demonstration - core.md carries policy | Accepted | +| [ADR-179](./system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md) | Remove the pre-1.0 in-place migrator; keep the guards and the transition fallbacks | Accepted | +| [ADR-180](./system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md) | GitHub issues as the shared truth for the session task list | Accepted | +| [ADR-181](./system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md) | Guard hooks: a blocking PreToolUse class, shipped deactivated | Accepted | +| [ADR-182](./system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md) | Keepwarm: attend keeps the prompt cache warm with a wake floor | Accepted | +| [ADR-184](./system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md) | Installation and activation are separate states: targets as the unit of activation | Accepted | +| [ADR-185](./system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md) | CLI output contract: structured output for people, JSON for machines | Accepted | +| [ADR-186](./system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md) | Live integration fixture: install-path test levels and the tier 2 gate | Accepted | +| [ADR-187](./system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md) | Attend MCP server mode: outbound and queries as typed tools, inbound stays on Monitor and the Stop hook | Proposed | +| [ADR-188](./system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md) | PostToolUse delivery for tool-lane ways and retirement of the semantic Bash surface | Proposed | + +## Governance +_Provenance, traceability, controls, compliance mapping_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-200](./governance/ADR-200-compliance-claims-and-session-derived-findings.md) | Compliance claims and session-derived findings | Accepted | +| [ADR-201](./governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md) | Findings assembled as classifier-ready assessment records | Accepted | + +## Documentation +_Documentation structure, tooling, coherence_ + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-300](./documentation/ADR-300-documentation-structure.md) | Documentation Structure | Accepted | +| [ADR-301](./documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md) | Situated socialization as canonical framing and documentation prose refactor | Accepted | +| [ADR-302](./documentation/ADR-302-unified-documentation-model.md) | A unified documentation model — typed graph, ways packaging, cross-repo convergence | Accepted | +| [ADR-303](./documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md) | Active-set semantics, adr archive, and supersession reading for the ADR corpus | Accepted | +| [ADR-304](./documentation/ADR-304-typed-decision-records-the-adr-v1-contract.md) | Typed decision records: the adr/v1 contract | Accepted | +| [ADR-305](./documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md) | Capabilities active at adoption need no add decision | accepted | +| [ADR-306](./documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md) | adr import: foreign records through a round-trip import sheet | accepted | +| [ADR-307](./documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md) | A decision names what should be observable when it holds | accepted | + +## Legacy (Pre-Domain Numbering) + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-004](./legacy/ADR-004-way-macros.md) | Way Macros for Dynamic Context Injection | Accepted | +| [ADR-005](./legacy/ADR-005-governance-traceability.md) | Governance Traceability for Ways | Superseded (superseded by ADR-200) | +| [ADR-013](./legacy/ADR-013-ways-skills-governance-architecture.md) | Ways, Skills, and Governance Architecture | Accepted | +| [ADR-014](./legacy/ADR-014-tfidf-semantic-matcher.md) | TF-IDF/BM25 Binary for Semantic Way Matching | Accepted | + +## Archived + +<details><summary>4 archived ADRs — no longer part of the active set; kept for history</summary> + +| ADR | Title | Status | +|-----|-------|--------| +| [ADR-112](./archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md) | Session Ledger with Optional Knowledge Graph Enhancement | Superseded (superseded by ADR-123) | +| [ADR-133](./archive/system/ADR-133-plugin-way-discovery.md) | Plugin Way Discovery | Rejected | +| [ADR-145](./archive/system/ADR-145-explicit-three-source-convergence-manifest.md) | Explicit three-source convergence manifest | Superseded (superseded by ADR-144) | +| [ADR-168](./archive/system/ADR-168-instance-addressable-directed-messaging-for-same-cwd-attend-siblings.md) | Instance-addressable directed messaging for same-cwd attend siblings | Rejected | + +</details> diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/adr.yaml b/tests/fixtures/adr/v0-corpus/docs/architecture/adr.yaml new file mode 100644 index 00000000..40820f0f --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/adr.yaml @@ -0,0 +1,116 @@ +# ADR Configuration +# Defines domain numbering and validation rules. +# Used by: docs/scripts/adr + +# Project name (used in generated index) +project_name: Claude Code Configuration + +# Domain Number Series +domains: + system: + range: [100, 199] + name: System + description: Ways architecture, matching, macros, hooks, session lifecycle + folder: system + + governance: + range: [200, 299] + name: Governance + description: Provenance, traceability, controls, compliance mapping + folder: governance + + docs: + range: [300, 399] + name: Documentation + description: Documentation structure, tooling, coherence + folder: documentation + +# The record contract (ADR-304). Records that declare `contract: adr/v1` +# follow the kinds below; a record without it is adr/v0 and keeps the v0 +# statuses and checks until someone migrates it. +contract: adr/v1 + +kinds: + decision: + mutable_after_accept: [status, enacted, superseded_by, considered, concern, observable] + verb: required + requires: [capability, basis, agent] + sections: [Summary] + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } + spec: + mutable_after_accept: all + verb: forbidden + requires: [capability] + edges: { supersedes: spec, decided_by: decision } + +# What agent-ways does, one line each: the only hand-written description of +# a capability. A decision adds, cuts, changes or constrains these. +capabilities: + adr: Decision records, their contract, and the adr tool that enforces it + docs: The documentation model, catalog pages and doclint + method: The ways method itself, what a way is for, its epistemic posture and its register + matching: How a prompt, tool call or file edit selects ways, the scoring behind it, and the fire telemetry and calibration data that tune it + disclosure: How a selected way reaches the model, when it re-fires, progressive disclosure, and localized way text + authoring: Way files, their frontmatter, the corpus build and the ways CLI for writing them + cli: The output contract the ways, attend and adr commands share with agents and scripts + attend: Session awareness through sensors, peers, messaging and keepwarm + install: Install, update, projection into targets, self-update, and how the binaries are built and laid out + config: Settings, configuration layers, permissions, guard hooks and the security baseline + governance: Provenance, controls and compliance findings + loop: The development loop skills (start, develop, merge, release, wrap) + testing: The test suites and the live install fixture + +# Capabilities that were active when agent-ways adopted adr/v1 on the +# adopted date. No record added them, so none needs an add decision. A change +# on one with no prior record to name stands on the baseline. Any capability +# declared after adoption needs an add (ADR-305). +baseline: + adopted: 2026-09-27 + capabilities: [docs, method, matching, disclosure, authoring, cli, attend, install, config, governance, loop, testing] + +# Retire targets name a surface; these are agent-ways' namespaces. +surfaces: + cli: {} + skill: {} + way: {} + hook: {} + +# The basis sources a decision may name (ADR-304 §11); the v1 default. +basis_sources: [operator, evidence, standard, upstream, precedent] + +# adr cite skips these: fixtures, tool sources and subagent prompts carry +# example numbers. +cite: + exclude: + - agents/workflow-orchestrator.md + - agents/workspace-curator.md + - tests/fixtures + - tests/adr-lint-test.sh + - tests/adr-archive-test.sh + - tests/adr-golden-test.sh + - hooks/ways/documentation/adr/src + - hooks/ways/documentation/adr/adr-tool + - hooks/ways/documentation/linting + +# Valid ADR statuses (adr/v0 records) +statuses: + - Draft + - Proposed + - Accepted + - Superseded + - Deprecated + - Rejected + +# Default values for new ADRs. No person is hard-coded as a decider: an +# empty list lets the tool fill in the current git or GitHub user. +defaults: + deciders: [] + status: Draft + +# Legacy ADR range (pre-domain numbering) +legacy: + range: [1, 99] + label: "Legacy (Pre-Domain Numbering)" + +# Viewer command for `adr view` +viewer: cat {file} diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md new file mode 100644 index 00000000..d023599a --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-112-session-ledger-and-knowledge-graph-integration.md @@ -0,0 +1,425 @@ +--- +status: Superseded +superseded_by: + - "ADR-123" +date: 2026-04-05 +deciders: + - aaronsb + - claude +related: + - ADR-103 + - ADR-104 + - ADR-105 +--- + +# ADR-112: Session Ledger with Optional Knowledge Graph Enhancement + +> **ARCHIVED — 2026-08-13.** No longer part of the active architecture set. Kept for history +> and so existing references still resolve. +> +> **Why:** Tiers 0 and 1 shipped under a different design: events.jsonl in state_root survives compaction and 'ways rethink' / 'ways introspect' replay it, delivered by ADR-123 and ADR-153/154. Tier 2 (knowledge graph ingestion) was never built and is dropped. +> **Superseded by:** ADR-123 +> +> Nothing below this line has been edited. + +## Context + +Sessions are ephemeral. When context compacts or a session ends, everything Claude learned — decisions made, assumptions revised, patterns discovered, dead ends encountered — evaporates. The auto-memory system (MEMORY.md) captures *surprises*, but most session knowledge isn't surprising enough to memorize yet too valuable to lose. It's the mundane connective tissue: "we tried X, it didn't work because Y, so we pivoted to Z." + +Two gaps in the current system motivate this design: + +### 1. No durable session record + +Ways fire, checks verify, tools execute — but the narrative arc of a session is never captured. The transcript exists but is a raw log, not a knowledge artifact. There's no structured account of what happened, what was learned, and what changed over the course of work. Prior sessions are invisible to current sessions except through the narrow lens of auto-memory. + +### 2. Context compaction destroys working knowledge + +When compaction fires, Claude loses the detailed understanding built over the session. The compaction-checkpoint way (ADR-104 context-threshold trigger) mitigates this by prompting a summary, but the summary is Claude's attempt to compress knowledge under time pressure. There's no progressive externalization that captures knowledge *as it forms*, while context is still rich. + +### The underlying principle + +An agent operating under a shrinking resource (context window) must externalize knowledge at a rate that tracks resource consumption, not wall-clock time. This is the same principle as write-ahead logging in databases (log before the transaction commits), incremental checkpointing in HPC (save every N iterations, not N minutes), and medical shift handoffs (outgoing nurse writes a narrative per patient: what happened, what's pending, what to watch for). + +### Knowledge tier geometry + +Each knowledge tier offers a fixed "surface area" (depth × breadth) that reshapes along an attention dimension — echoing the capacity-precision tradeoff described in Oberauer's concentric model of working memory (2002) and formalized by rate-distortion theory (Shannon/Berger), where a fixed information budget is allocated between fidelity and coverage: + +| Tier | Shape | Persistence | Bounded By | +|---|---|---|---| +| **Live context** | Tall, narrow (deep on current task, blind to rest) | Session (collapses at compaction) | Context window | +| **Transcripts** | Shorter, wider (raw history, low signal-to-noise) | Uncertain (possible TTL) | Disk, retention policy | +| **Ledger** | Flatter, wider (curated summaries, high signal) | Permanent (user-controlled) | Disk | +| **Knowledge graph** | Flattest, widest (concepts + edges across all time) | Regenerable from ledger | KG infrastructure | + +Each project has its own stack of these tiers with different surface areas based on maturity. The **overlap projection** across all tiers and all projects is the total knowledge coverage. Gaps in the projection are things unknown at any tier. Overlap between tiers is redundancy — and how the tiers validate each other. This projection model has precedent in Formal Concept Analysis (Wille, 1982), where coverage gaps are computed from overlapping concept extents. + +Retrieval triggers (domain entry, post-compaction, etc.) are **attention-routing operations** — they select the operating point on the rate-distortion curve, reshaping which slice of the lower tiers gets projected into the live-context tier. Without them, the lower tiers exist but never influence the first-class attention space where work actually happens. + +### The economic bet + +The reflection entries cost tokens — Claude spends part of its response on summation rather than forward progress. The KG ingestion costs tokens against a separate inference framework (the KG's LLM extraction pipeline). The bet is that this combined cost returns a **greater than 1:1 ratio of value** through connected concept reasoning. + +The value comes not from storing what Claude said, but from what the KG *does* with it: decomposing prose into concepts, deduplicating against prior sessions, building epistemic confidence, and discovering structural connections that no single session could see. A 150-token reflection entry might produce 3-5 concepts with edges to dozens of prior concepts. When those connections surface at a future state transition, they provide reasoning context that would otherwise require Claude to re-derive from scratch — or never discover at all. + +The cost is bounded (3-4 reflections per session, ~500 tokens total). The value compounds (each session enriches the graph, making future retrievals denser). The ratio improves over time as the graph grows. Early sessions pay more than they get back; mature projects get back far more than they pay. + +## Decision + +Introduce a three-tier progressive system where each tier is independently activatable: + +``` +Tier 0 (current): Ways fire → guidance injected → forgotten at compaction +Tier 1 (ledger): Epoch entries written to durable ledger → survive compaction → replayable +Tier 2 (KG): Ledger entries ingested into knowledge graph → cross-session connections emerge +``` + +Each tier builds on the previous. Tier 1 works without Tier 2. Tier 0 works without either. The tiers are activated via `~/.claude/ways.json` configuration: + +```json +{ + "reflection": { + "ledger": true, + "kg": false + } +} +``` + +### Tier 1: The Session Ledger + +#### Epoch-triggered reflection + +A new way (`meta/reflection/reflection.md`) fires at context-threshold boundaries. The existing `context-threshold` trigger type (used by compaction-checkpoint and memory ways) provides the timing. The reflection way fires at multiple thresholds with escalating depth: + +| Context Used | Phase | Depth Guidance | +|---|---|---| +| ~30% | Orientation | What's the task? Initial direction. First decisions. | +| ~50% | Progress | What changed since orientation? Revised understanding. Open threads. | +| ~70% | Consolidation | Knowledge gained. Patterns observed. What surprised. | +| PreCompact | Handoff | Complete state for post-compaction self. Assumptions. Unfinished work. | +| PostCompact | Resume | Read ledger + orient. No writing — just restore continuity. | + +The way **does not ask Claude to track epochs or write files**. The ways framework auto-generates epoch metadata as frontmatter in the ledger entry — session ID, project, epoch number, context percentage, timestamp. Claude reflects naturally in conversation using a keyphrase; the Stop hook extracts the prose and writes the entry. + +#### Ledger structure + +The ledger is a single chronological stream per project — not partitioned by session. Session boundaries are metadata on entries, not structural divisions. The ordering that matters is *when things happened on this project*, because the ledger is the summed experience of working on that project across all sessions. + +``` +~/.claude/ledger/ + {project-slug}/ + 2026-04-05T1423Z_e000.md + 2026-04-05T1445Z_e001.md + 2026-04-05T1502Z_e002.md + 2026-04-07T0910Z_e003.md ← different session, same project stream + ... +``` + +Epoch numbering is monotonic across the project, not per-session. Entry 003 from session `def` follows entry 002 from session `abc` because that's the temporal order of experience. Any method that resets at session boundaries loses signal as the ledger builds. + +Each entry file: + +```markdown +--- +session: abc123 +transcript: /path/to/transcript.jsonl +epoch: 1 +context_pct: 50 +timestamp: 2026-04-05T14:45:00Z +phase: progress +--- + +Shifted from the initial plan of refactoring auth middleware to addressing +the token storage compliance issue first. Legal flagged that session tokens +in Redis aren't encrypted at rest. + +Key decisions: +- Chose libsodium over OpenSSL for token encryption +- Keeping old middleware path as fallback behind feature flag + +Open threads: +- Haven't tested Redis upgrade path +- Need to check mobile client token caching +``` + +The frontmatter is written by the hook script (which has access to session state), not by Claude. Claude writes everything below the frontmatter delimiter. This separation ensures metadata accuracy — Claude doesn't guess its epoch number or context percentage. The frontmatter is the framework's record; the prose is Claude's. + +The `transcript` field links back to the source session transcript. This means the ledger entry can serve as a pointer into the full session history — the prose is a curated summary, and the transcript is the raw record. If deeper context is needed, the transcript is always reachable. + +#### Capture mechanism — keyphrase extraction + +Claude doesn't write to the ledger directly. Instead, the reflection way injects a prompt like: + +> *Time to reflect on what's happened since the last reflection. Here's the first line of your last reflection: "{first_line}". Use the phrase "my reflections on" to begin.* + +Claude reflects naturally in its response to the user — no tool calls, no file writes, no interruption. The user sees the reflection as part of the conversation. + +The **Stop hook** (`check-response.sh`) then: + +1. Detects that the reflection way fired this turn (session marker exists) +2. Extracts prose from the transcript between the keyphrase "my reflections on" and the next section break or response end +3. Writes the ledger entry file — framework-generated frontmatter + extracted prose +4. Appends to the ephemeral session narrative (`${SESSIONS_ROOT}/${session}/narrative.md`) + +The keyphrase serves as a machine-parseable delimiter that's natural enough to appear in conversation without feeling mechanical. The ways embedding engine can score responses for the keyphrase to gate capture — if Claude uses the phrase casually without a reflection way having fired, the marker check prevents false capture. + +**Total Claude-side cost: zero tool calls.** Claude just talks. The framework handles filing. + +#### Additional capture channel: memory writes + +Claude's auto-memory system (MEMORY.md + topic files) produces the same class of knowledge artifact as ledger entries — curated prose about what was learned, written to disk with structured frontmatter. Memory writes are triggered by explicit user requests ("remember this") or by the memory way at context thresholds (surprise test). + +When KG is enabled, a PostToolUse hook on Write can detect writes to the memory directory and copy the file to the KG FUSE ingest — the same pattern as ledger entry ingestion. Memory files already have frontmatter with name, description, and type, making them well-structured KG input. This means the KG receives knowledge from two channels: + +1. **Ledger entries** — periodic epoch reflections (what happened) +2. **Memory writes** — explicit or surprise-gated observations (what was surprising or important) + +Both channels are fire-and-forget copies to the FUSE mount. The KG deduplicates across channels naturally. + +#### Session narrative (ephemeral working copy) + +During a session, entries are also appended to `${SESSIONS_ROOT}/${session}/narrative.md` — a concatenated view of entries this session. This is what Claude reads for in-session continuity (e.g., at resume after compaction). It's ephemeral (lives in XDG_RUNTIME_DIR) and regenerable from the ledger. + +The resume way does not inject the entire narrative — only the last 2-3 entries plus the handoff entry. In long sessions with many compaction cycles, the narrative could grow large; injecting it all would immediately consume the fresh context window. The KG (if active) handles deeper cross-session structural context; the narrative provides only the immediate thread of continuity. + +#### Epoch numbering + +Epochs are monotonic across the project — the next epoch number is derived from the count of existing entries in the project's ledger directory. A new session picks up where the last session left off. This means the ledger is a continuous record of project experience, with session boundaries visible in the metadata but not in the numbering. + +### Tier 2: Knowledge Graph Enhancement (Optional) + +When `reflection.kg` is enabled and a knowledge graph MCP server is available, ledger entries additionally feed into the KG. The KG is a **derived view** of the ledger — never the source of truth. Not all users will have or want a KG. The ledger is fully functional without it. + +#### Ingestion (file copy) + +After the Stop hook writes a ledger entry, it copies the entry file to the KG's FUSE-mounted ingest directory: + +```bash +KG_INGEST="${HOME}/Knowledge/ontology/${PROJECT_SLUG}/ingest" +if [[ -d "$KG_INGEST" ]]; then + cp "$ENTRY" "$KG_INGEST/" +fi +``` + +The FUSE layer picks up the file, the KG processes asynchronously — decomposing into concepts, deduplicating, assigning epistemic status. No CLI invocation, no MCP call, no tool use. Just a file copy. If the KG isn't mounted or the directory doesn't exist, the `if` guard skips silently — the ledger entry exists regardless. + +The ledger entry file *is* the KG input document. Same file, same format. The YAML frontmatter is ignored by the KG's text chunker; the prose below the delimiter is what gets ingested. One artifact serves both purposes. + +#### Retrieval (state-transition-driven) + +Ingestion and retrieval are **completely decoupled**. Searching right after ingest is just an expensive echo. The value of the KG is associative context that surfaces at *state transitions* — moments where Claude's working model shifts and prior knowledge would reshape the shift: + +| Trigger | Hook Event | Why This Moment | +|---|---|---| +| **Domain entry** | Way fires (first time in session) | Cross-session experience with this domain is most valuable now | +| **Post-compaction** | PostCompact | KG provides breadth across prior sessions, not just this session's continuity | +| **New session** | SessionStart (first for project in N days) | KG provides project orientation across all prior sessions | +| **Tool failure** | PostToolUseFailure | KG may have prior experience with similar failures | + +This is not RAG. There's no query to answer — the triggers are events, not questions. The KG works on a different timescale than the active session, building concepts in the background while Claude works. When a state transition fires, connections that didn't exist before have materialized. This serves as a form of **subconscious attachment** — prior session knowledge accretes silently and surfaces at natural pause points. + +#### Scoping + +One KG ontology per Claude project (named after the project slug). All retrieval queries scope to the project ontology. Foundational knowledge relevant to a project is ingested into that project's ontology. Concepts can point to items outside their own ontology via edges, but retrieval never crosses the project boundary. + +#### Replay + +The ledger enables KG regeneration. If the KG is lost or reset, ingest all ledger entries across all projects (`~/.claude/ledger/*/`). Order doesn't matter — the KG's epistemic model measures honesty and convergence, not recency, so ingesting entry 47 before entry 3 produces the same concept graph. Deduplication makes replay idempotent. The temporal ordering in the ledger serves human readability and in-session continuity, not the KG. + +#### Graph maturity phases + +The KG isn't static storage — it progresses through distinct phases as ledger entries accumulate, each producing a qualitatively different kind of value: + +**Linear** — Early sessions. The KG registers prose as clusters of concepts. Each entry creates new nodes with sparse edges. The graph is shallow — many islands, few bridges. Retrieval returns individual concepts, not connections. Value is low but the foundation is being laid. + +**Expansion** — New prose attaches to existing concepts rather than creating new ones. Evidence instances accumulate on established nodes. The graph becomes denser within domains. Retrieval starts returning concepts with multiple evidence sources, increasing confidence. The KG begins to say "you've seen this pattern before" rather than just "this concept exists." + +**Convergence** — Evidence grows enough to form directed graph networks in the concept corpus. Edges between concepts gain epistemic weight. The graph structure starts to reflect real relationships — SUPPORTS, CONTRADICTS, IMPLIES — rather than just co-occurrence. Retrieval returns paths between concepts, not just individual nodes. The KG can now surface connections like "this decision supports that principle but contradicts this earlier assumption." + +**Reasoning** — Node relationships, edge vector directions, and grounding scores with substantiating evidence form reasoning networks. The graph topology itself encodes arguments. High-grounding paths represent well-established reasoning chains; contested edges represent open questions. Retrieval at this phase provides not just context but structured reasoning — "here's why X, supported by evidence from sessions 3, 7, and 12, with a counterpoint from session 9." + +Each phase increases the value-to-cost ratio of the reflection investment. The linear phase costs the same as the reasoning phase in tokens spent, but the reasoning phase returns qualitatively richer context. + +#### Why KG over editorial curation (the ephemera problem) + +Memory curation approaches fall into two categories from the formal literature: **editorial** (the AGM framework — Alchourrón, Gärdenfors, Makinson, 1985) which performs minimal consistent revision, and **evidential** (Dempster-Shafer theory, 1976) which preserves contradictions and scores the balance of evidence. Claude Code's built-in memory curation ("dream mode") is editorial. The KG is evidential. + +When editorial curation encounters "X is true" from session 3 and "X is false" from session 12, it must pick one. The *fact that understanding shifted* is itself knowledge, and it's destroyed. The KG keeps both assertions with their evidence and scores the balance through epistemic status classifications that editorial curation cannot represent: + +- **AFFIRMATIVE** — consistent evidence across sessions (editorial curation's only output) +- **CONTESTED** — evidence in both directions (more informative than either assertion alone) +- **CONTRADICTORY** — later evidence actively contradicts earlier (tells you understanding shifted) +- **INSUFFICIENT_DATA** — not enough evidence to classify (tells you this is unexplored) + +The deeper problem is ephemera. Editorial curation keeps big signals and discards small ones. But minor observations compound — a low-signal concept from session 5 might become the key insight in session 40, but only if it survived 35 curation cycles. The KG preserves it as a low-ranked node with live edges, costing nothing in attention budget until a future session produces a structurally adjacent concept. The minor signal surfaces not because anyone remembered to preserve it, but because the graph topology kept it connected. + +The KG's composed scoring reinforces this. A concept that's individually weak but connected to three strong concepts along a coherent semantic axis is more informative than its individual score suggests. Polarity analysis projects concepts onto axes — a minor concept near the midpoint of a contested axis is the most interesting concept in the graph, not the least. + +The scaling curves go in opposite directions. As data grows over months, editorial curation becomes increasingly destructive — each pass loses more ephemera. The KG becomes increasingly valuable — each new concept has more neighbors to connect to, more axes to project onto, more composed scores to participate in. The ledger preserves raw signal; the KG preserves the *relationships between signals* — including the minor ones that editorial curation discards. + +### Hook Changes + +#### New hook entries + +**PreCompact** — Fires the handoff reflection phase: + +```json +"PreCompact": [ + { + "hooks": [ + {"type": "command", "command": "${HOME}/.claude/hooks/ways/check-precompact.sh"} + ] + } +] +``` + +**PostCompact** — Fires the resume phase: + +```json +"PostCompact": [ + { + "hooks": [ + {"type": "command", "command": "${HOME}/.claude/hooks/ways/check-postcompact.sh"} + ] + } +] +``` + +**PostToolUseFailure** — Injects error-context ways when tools actually fail: + +```json +"PostToolUseFailure": [ + { + "matcher": "Bash", + "hooks": [ + {"type": "command", "command": "${HOME}/.claude/hooks/ways/check-error-post.sh"} + ] + } +] +``` + +#### Existing hook improvements + +**Conditional `if` on PreToolUse:Bash** — Reduces hook overhead by skipping read-only commands: + +```json +{ + "matcher": "Bash", + "if": "Bash(git *) || Bash(make *) || Bash(docker *) || Bash(npm *) || Bash(cargo *) || Bash(kg *)", + "hooks": [{"type": "command", "command": "...check-bash-pre.sh"}] +} +``` + +**Async on non-blocking hooks** — Stop hook and marker writes don't need to block: + +```json +{"type": "command", "command": "...check-response.sh", "async": true} +``` + +### New Ways + +| Way | Trigger | Purpose | +|---|---|---| +| `meta/reflection/reflection.md` | context-threshold (30/50/70%) | Write epoch entries to ledger | +| `meta/reflection/handoff.md` | PreCompact (direct injection) | Full-depth handoff before compaction | +| `meta/reflection/resume.md` | PostCompact (direct injection) | Read ledger, orient, resume | + +### New/Modified Scripts + +| Script | Hook Event | Purpose | +|---|---|---| +| `check-response.sh` (modified) | Stop | Extracts keyphrase-delimited reflection, writes ledger entry, copies to KG FUSE if available | +| `check-precompact.sh` | PreCompact | Fires handoff way, passes epoch state | +| `check-postcompact.sh` | PostCompact | Fires resume way with ledger path | +| `check-error-post.sh` | PostToolUseFailure | Fires error-context ways on failure | +| `ledger-replay.sh` | (utility) | Replays ledger entries into KG | + +## Consequences + +### Positive + +- Sessions produce durable knowledge artifacts (ledger entries) that survive beyond the session +- Progressive externalization captures knowledge while context is rich, not under compaction pressure +- The ledger is append-only and replayable — a project's session history is always available +- Each tier is independently activatable — ledger works without KG, current system works without ledger +- Epoch metadata is framework-generated, not Claude-generated — accurate by construction +- PreCompact/PostCompact hooks provide natural timing for handoff and resume +- `if` field and `async` on existing hooks reduce latency on every turn +- When KG is active, sessions get smarter over time as epistemic trust builds across sessions + +### Negative + +- Ledger entries consume disk space over time (mitigated: prose entries are small, ~500 bytes each) +- The reflection way adds content to Claude's response at threshold boundaries (mitigated: keyphrase capture is conversational, not a tool-call interruption) +- Requires `ways` binary awareness of ledger directory for epoch numbering +- Stop hook must parse transcript to extract keyphrase-delimited prose (mitigated: simple regex on a known marker) +- When KG is active: depends on FUSE mount availability (mitigated: `if -d` guard skips silently, ledger is unaffected) + +### Neutral + +- The compaction-checkpoint way (context-threshold at 95%) continues to handle user-facing compaction dialogue — reflection operates at lower thresholds for knowledge capture, not compaction coordination +- Auto-memory (MEMORY.md) remains for surprises and cross-session pointers — the ledger captures the mundane connective tissue that memory intentionally ignores +- The epoch counter for checks (ADR-103) is independent — it counts hook events, not reflection entries. The ledger epoch is a different concept (knowledge externalization events) +- Opens the question of ledger pruning/archival — we defer this (append-only until proven problematic) +- The KG's epistemic status starts low (few evidence instances) and builds confidence over time — this is correct behavior, not a deficiency + +## Alternatives Considered + +- **Single narrative file per session** — Rejected because individual entry files enable atomic KG ingestion and clean replay. A single file requires parsing to find entry boundaries. + +- **Claude tracks its own epochs** — Rejected because Claude's self-reported metadata is unreliable after compaction. The framework has authoritative access to session ID, context percentage, and epoch count via session state files. Generating frontmatter externally ensures accuracy by construction. + +- **KG-only (no ledger)** — Rejected because it creates a dependency on KG availability for knowledge continuity. The ledger is the source of truth; the KG is a derived, regenerable view. If the KG goes down, the ledger still provides session history and replay capability. + +- **Ledger in project `.claude/` directory** — Rejected because the ledger is user-scoped knowledge about project work, not project configuration. Committing session narratives to a shared repo exposes working notes. The `~/.claude/ledger/` location keeps it user-private alongside other user-scoped state. + +- **Write entries at fixed turn counts** — Rejected because turn count doesn't track context consumption. 10 turns of complex tool use consumes more context than 50 turns of brief Q&A. Context-threshold triggers track the resource that actually matters. + +- **Claude writes ledger entries via Write tool** — The initial design had the reflection way instruct Claude to call Write with a specific file path. Rejected in favor of keyphrase extraction from natural conversation. The Write approach interrupts Claude's flow with a tool call, requires the way to communicate file paths, and makes the reflection feel mechanical. Keyphrase capture is invisible — Claude just reflects in conversation and the framework handles filing. + +- **Claude-driven KG ingestion via MCP tool calls** — Considered having Claude call `session_ingest` or `ingest` directly. Rejected for the normal flow because it costs tool calls and context window tokens for something the framework can handle invisibly. The FUSE file copy achieves the same result with zero Claude involvement. Note: Claude can still use KG MCP tools directly for advanced introspection and deep reasoning within the knowledge graph — this is an intentional capability but not part of the standard reflection flow. + +- **Search immediately after ingest (RAG pattern)** — The naive design searches the KG right after ingesting an entry. Rejected because this is an expensive echo — the search returns concepts derived from what Claude just wrote. The value of the KG is *associative context from prior sessions*, not retrieval of current knowledge. Ingest and retrieval are decoupled: ingest is periodic (epoch boundaries), retrieval is event-driven (state transitions like domain entry, post-compaction, tool failure). This is not RAG. + +- **Configurable ontology sets per project** — Considered allowing projects to declare a list of ontologies to query (e.g., `["claude-config", "cognitive-frameworks"]`). Rejected in favor of one ontology per project. Foundational knowledge that matters to a project is ingested into the project's ontology. The KG deduplicates at the concept level, and edges can cross ontology boundaries naturally. Simpler model: the project slug *is* the query scope, no configuration needed. + +- **Cross-project KG queries** — Considered allowing retrieval to search across all ontologies. Rejected because it violates project scoping — sessions from project A shouldn't bleed into project B's context. Cross-ontology edges exist at the concept level (a concept can point to concepts in other ontologies), but retrieval is always scoped to the current project's ontology. + +## Related Projects + +- [aaronsb/agent-ways](https://github.com/aaronsb/agent-ways) — The ways framework this ADR extends +- [aaronsb/knowledge-graph-system](https://github.com/aaronsb/knowledge-graph-system) — The KG system used for Tier 2 + +## Theoretical References + +- Oberauer (2002) — Concentric model of working memory: capacity-precision tradeoffs across nested tiers +- Shannon/Berger — Rate-distortion theory: fixed information budget allocated between fidelity and coverage +- Wille (1982) — Formal Concept Analysis: coverage gap detection from overlapping concept extents +- Dempster-Shafer (1976) — Evidential reasoning: preserve contradictions, score balance of evidence +- AGM framework (1985) — Belief revision via minimal consistent editing (the editorial model we contrast against) +- Tishby et al. (1999) — Information Bottleneck: attention as selection of operating point on R(D) curve +- Behrouz et al. (2025) — Google Titans: multi-tier memory (short-term/long-term/persistent) with surprise-gated writes. Rhymes with this design's tier structure, though Titans operates within model architecture while this design operates via external infrastructure. + +## Implementation Plan + +### Phase 1: Hook improvements (no binary changes) +1. **`if` field on PreToolUse:Bash** — Config-only change to reduce hook overhead +2. **`async: true`** on Stop hook and non-blocking markers — Config-only latency win +3. **PreCompact/PostCompact hook entries** — New hooks in settings.json + +### Phase 2: Ledger (Tier 1) +4. **Ledger infrastructure** — Directory structure, entry format, project slug derivation in `ways` binary +5. **Reflection way** (`meta/reflection/reflection.md`) — Keyphrase-prompted reflection at context thresholds +6. **Stop hook extension** — Keyphrase extraction from transcript, ledger entry writing with framework frontmatter +7. **Handoff/resume ways** — `meta/reflection/handoff.md`, `meta/reflection/resume.md` +8. **Shell scripts** — `check-precompact.sh`, `check-postcompact.sh` +9. **Observe** — Run for several sessions, tune context-threshold boundaries and keyphrase discrimination + +### Phase 3: KG enhancement (Tier 2, optional) +10. **FUSE integration** — Copy ledger entries to `~/Knowledge/ontology/{project}/ingest/` if mounted +11. **Retrieval triggers** — KG search at domain entry, post-compaction, session start +12. **Replay utility** — `ledger-replay.sh` for KG regeneration from ledger +13. **Observe** — Measure KG injection information density and epistemic trust growth + +### Phase 4: Additional hook diversification +14. **PostToolUseFailure** — Error-context way injection via `check-error-post.sh` +15. **CwdChanged / FileChanged** — Project-local way activation, corpus rebuild on way edits diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-133-plugin-way-discovery.md b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-133-plugin-way-discovery.md new file mode 100644 index 00000000..821323a3 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-133-plugin-way-discovery.md @@ -0,0 +1,173 @@ +--- +status: Rejected +date: 2026-04-25 +deciders: + - aaronsb + - claude +related: + - ADR-108 + - ADR-111 +--- + +# ADR-133: Plugin Way Discovery + +> **ARCHIVED — 2026-08-13.** No longer part of the active architecture set. Kept for history +> and so existing references still resolve. +> +> **Why:** Never implemented. Four months after proposal, no plugin way-path resolution exists; the only 'claude plugin list' call in the tree counts skills for the context-cost warning in hooks/ways/macro.sh. +> +> Nothing below this line has been edited. + +> **Provenance.** This ADR was originally drafted on 2026-04-25 as ADR-129 on the +> exploratory `feat/plugin-way-discovery` branch (PR #76). That branch was discarded +> — it had drifted ~53 commits behind main and conflicted across the refactored +> ways-cli internals — but the design reasoning was sound and worth keeping. It is +> re-filed here as ADR-133 because the original ADR-129 number was reassigned on main +> to "instance suffix and heartbeat liveness." Status is **Proposed**: the design is +> captured and its load-bearing assumption (`claude plugin list --json`) is verified +> current, but it is not implemented. The **Implementation** section names specific +> code (`candidates.rs`, `cmd/corpus.rs`, the `SessionStart` hook chain) as it stood +> in April 2026; those internals have since been refactored, so treat that section as +> indicative of approach, not literal touch points. + +## Context + +Way discovery is currently hardcoded to two filesystem locations: + +1. **Project-local**: `$PROJECT/.claude/ways/` +2. **Global**: `~/.claude/hooks/ways/` + +Claude Code plugins can ship `ways/` directories inside their install paths (matching the project-local convention, demonstrated by the `x@tracer-plugins` plugin which contains `.claude/ways/fruity/way.md`). However, the ways system has no mechanism to discover or scan these. Plugin-shipped ways are invisible. + +Claude Code provides a stable CLI interface for querying plugin state: + +``` +claude plugin list --json +``` + +Returns an array of installed plugins, each with: +- `id` — plugin identifier (`name@marketplace`) +- `installPath` — absolute path to the installed plugin on disk +- `enabled` — whether the plugin is currently active +- `scope` — `"user"` (global) or `"project"` (scoped to a specific project) +- `projectPath` — (project-scoped only) which project the plugin belongs to +- `version` — installed version +- `installedAt` / `lastUpdated` — timestamps + +This is sufficient to resolve which plugins are active and where their files live. + +### Design constraints + +- **No per-invocation subprocess**: `ways scan` runs on every prompt and tool use. Shelling out to `claude plugin list --json` on each invocation adds unacceptable latency. +- **Don't couple to internal file formats**: Reading `installed_plugins.json` and `settings.json` directly is faster but couples to Claude Code's internal storage format, which may change without notice. +- **Use the official CLI**: `claude plugin list --json` is the stable public interface for plugin state. +- **Enabled means enabled**: The `enabled` field already reflects whether a plugin is active. No additional scope filtering is needed — if `enabled` is `true`, the plugin participates. +- **Version deduplication**: If the same plugin ID appears with multiple versions, use the latest (by `lastUpdated` timestamp). The CLI already resolves to the active version, but defensive dedup protects against edge cases. + +## Decision + +### Hybrid approach: resolve once, scan many + +**At session start**, resolve enabled plugin way-paths via `claude plugin list --json` and write them to a session-scoped manifest. The `ways` binary reads this manifest during scans, adding plugin directories to the candidate collection alongside project-local and global ways. + +### Session-start resolution + +A new step in the `SessionStart` hook chain (after `ways init`, before `ways corpus --if-stale`): + +1. Run `claude plugin list --json` +2. Filter to `enabled == true` +3. For each, check if `$installPath/ways/` exists on disk +4. Deduplicate by plugin name: if multiple versions, keep the one with the latest `lastUpdated` +5. Write the list of way-paths to `$SESSION_DIR/plugin-ways.json` + +The manifest format: + +```json +[ + { + "id": "x@tracer-plugins", + "path": "/Users/tracer/.claude/plugins/cache/tracer-plugins/x/1.0.0/ways" + } +] +``` + +### Candidate collection + +`collect_candidates()` gains a third source, inserted between project-local and global: + +``` +1. $PROJECT/.claude/ways/ — project-local (highest priority) +2. $PLUGIN/ways/ — per enabled plugin (middle priority) +3. ~/.claude/hooks/ways/ — global (lowest priority) +``` + +The `ways` binary reads the session manifest (`plugin-ways.json`) and calls `collect_from_dir()` on each path. The existing `WalkDir`-based scanning, frontmatter parsing, domain filtering, and scope gating apply identically to plugin-sourced ways. + +### ID namespacing + +Plugin way IDs are prefixed with the plugin identifier to prevent collisions: + +``` +plugin:x@tracer-plugins/fruity (from plugin) +softwaredev/code/security (from global) +softwaredev/code/testing (from project-local) +``` + +Same-ID ways across sources share a session marker (project-local overrides plugin overrides global). Plugin ways cannot shadow global ways unless they use the same domain/path structure intentionally. + +### Corpus integration + +`ways corpus --if-stale` must include plugin way directories so that semantic (embedding) matching works for plugin-shipped ways. The corpus generation reads the same session manifest to discover additional scan roots. + +### Macro trust + +Plugin macros (`macro.sh` files inside plugin ways) are third-party code. They follow the same trust model as project-local macros: disabled by default, enabled per-plugin via `~/.claude/trusted-plugin-macros` (or extending the existing `trusted-project-macros` mechanism). + +## Consequences + +### Benefits + +- Plugins can ship ways alongside skills and hooks — a single plugin can provide guidance, tools, and workflows +- Plugin authors can use the full way authoring surface: frontmatter, semantic matching, macros, check curves, scope gating +- No coupling to Claude Code's internal plugin storage format — uses the stable CLI interface +- Session-start resolution means zero per-scan overhead from plugin discovery +- Existing way precedence model extends naturally (project > plugin > global) + +### Costs + +- Session-start adds one `claude plugin list --json` subprocess call (~100-200ms) +- Session manifest is a new file to manage (create on start, stale if plugins change mid-session) +- Corpus regeneration may take slightly longer with additional plugin way directories +- Plugin way authors must understand the ID namespacing scheme + +### Risks + +- **Mid-session plugin changes**: If a user installs/removes/toggles a plugin during a session, the manifest is stale until the next session or compaction. Acceptable — plugin changes are rare and a session restart is natural. +- **Manifest missing**: If the session manifest doesn't exist (e.g., older ways binary, failed resolution), `collect_candidates()` falls back to the current two-source behavior. No breakage. +- **Plugin path instability**: Plugin install paths include version strings that change on update. The session manifest captures the path at resolution time, so this is fine within a session. Cross-session, the next start re-resolves. + +### Way file path convention + +Each way lives in its own directory, named to match the way file. This enables sibling files (`.check.md`, `macro.sh`) alongside the way definition. + +| Scope | Full path | +|-------|-----------| +| Global | `~/.claude/hooks/ways/{domain}/{way}/{way}.md` | +| Project-local | `$PROJECT/.claude/ways/{domain}/{way}/{way}.md` | +| Plugin | `$PLUGIN_INSTALL_PATH/ways/{domain}/{way}/{way}.md` | + +Plugins use `ways/` at the plugin install root. Global ways use `hooks/ways/` under `~/.claude/` because they sit alongside other hook types. + +## Implementation + +> Indicative as of the April 2026 ways-cli structure; verify against current code before building. + +### Touch points + +1. **New hook script**: `hooks/ways/resolve-plugins.sh` — runs `claude plugin list --json`, filters, writes manifest +2. **`settings.json`**: Add `resolve-plugins.sh` to `SessionStart` hooks (after `ways init`) +3. **`candidates.rs`**: `collect_candidates()` reads session manifest and adds plugin dirs +4. **`candidates.rs`**: `collect_checks()` — same addition for check files +5. **`cmd/corpus.rs`**: Corpus generation reads manifest for additional scan roots +6. **Way ID derivation**: Prefix plugin-sourced IDs with `plugin:{id}/` +7. **Macro trust**: New trust file or extend existing mechanism for plugin macros diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-145-explicit-three-source-convergence-manifest.md b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-145-explicit-three-source-convergence-manifest.md new file mode 100644 index 00000000..41b1c0a1 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-145-explicit-three-source-convergence-manifest.md @@ -0,0 +1,332 @@ +--- +status: Superseded +superseded_by: + - "ADR-144" +date: 2026-06-30 +deciders: + - aaronsb + - claude +related: + - "[[ADR-144]]" + - "[[ADR-142]]" + - "[[ADR-143]]" +--- + +# ADR-145: Explicit three-source convergence manifest + +> **ARCHIVED — 2026-08-13.** No longer part of the active architecture set. Kept for history +> and so existing references still resolve. +> +> **Why:** ADR-144 shipped the two-source reconciler the day before this refinement was written, and solved the problem without a pinned Claude-Code baseline. The three-source manifest is absent from manifest.rs and reconcile.rs. +> **Superseded by:** ADR-144 +> +> Nothing below this line has been edited. + +## Context + +This is a child of ADR-144 (install / repair / migrate as one manifest reconciler), +which is itself a child of ADR-142 (agent-ways 1.0). ADR-144 unified install, update, +repair, and migrate into **one idempotent reconciler** that "converges `~/.claude` +toward a git-derived manifest," and framed the four from-states as **one convergence +with different entry conditions**: "the four from-states are one convergence … they +differ only in *starting actual* and *trust posture*." Fresh is "actual is empty → +materialize all entries"; the others re-materialize a delta against a live tree. + +ADR-144 names the desired side of that convergence — the manifest, derived from +`git ls-files` in `$XDG_DATA/agent-ways` — and the engine that drives toward it. But it +left the **actual** side, and one whole leg of the desired side, implicit: + +- The **agent-ways leg is explicit and exists today.** `ways manifest` + (`tools/ways-cli/src/cmd/manifest.rs`) emits exactly what agent-ways projects, derived + from the git-tracked file set: `PROJECTED_TREES` (`skills`, `agents`, `commands`, + `hooks/ways`), `PROJECTED_FILES` (the two named hooks), and `PROJECTED_BINS` (`ways`, + `attend`, `attend-chat`, `way-embed`). It runs `git ls-files` over the tracked trees and + allowlists the built binaries by name. This is already one of the three legs the + convergence needs. +- There is **no model of what Claude Code itself owns** in `~/.claude`. ADR-142's layout + table calls `~/.claude` the "irreducible **Claude-Code-owned floor**," but nothing + enumerates that floor. The reconciler converges toward *(agent-ways manifest)* and + treats everything else as a single undifferentiated "don't touch" region. +- "**What is the user's own customization**" is therefore decided heuristically — by the + same weak membership test ADR-144 already flagged in today's `build_manifest`: + *"did a prior projection write it,"* which "can't classify a file the user dropped into + a shared dir that happens to match a name we later ship." + +Three concrete consequences of that gap motivate this ADR: + +1. **The fresh-install path was never wired.** agent-ways 1.0.0's documented installer + (`git clone … ~/.claude && make setup`) still produces the **pre-1.0 in-place shape**, + not the projection. A brand-new 1.0 user lands in exactly the topology 1.0 replaced. + ADR-144 says fresh install should "fall out" of the engine as *materialize the manifest + from nothing* — but with no explicit target for the engine to materialize *into an empty + tree*, the native projection installer was never built. This is a real, current bug. + +2. **Classification is heuristic and error-prone, and four consumers need it.** Deciding + whether any `~/.claude` file is Claude-Code-owned, agent-ways-owned, or user-owned is + needed by install, repair, migrate, **and** cleanup — and each currently guesses. + ADR-144's own sharpest Negative ("the migrator must detect and rescue hand-edited core + files … or it silently destroys the customization it was meant to preserve") is a + symptom of having no authoritative classifier. + +3. **The settings.json three-way merge already solves a version of this — for one file.** + `tools/ways-cli/src/cmd/settings_merge.rs` does a kubectl-style three-way merge of + `settings.json`. It tracks a stored **last-applied base** — the slice agent-ways itself + last wrote, persisted to `$XDG_STATE/agent-ways/settings-applied.json` — and computes, per + owned slice, `result = (theirs − base − ours) ++ ours`. **Field ownership** is explicit: + agent-ways owns the `hooks` events and the `permissions.allow` entries matching + `WAYS_PERMS`; *everything else* — `model`, `theme`, `plugins`, `env`, credentials, the + user's own hooks, **and Claude Code's own default values** — is treated as the unmanaged + remainder. This is precisely the "tell app from user" computation, made idempotent by the + stored base. But note its shape: it is a **two-party** split (agent-ways-owned vs. + everything-else), it does *not* separately model the Claude-Code floor, and it is scoped to + a single JSON document. + +The realization: **the reconciler's convergence target is implicit and ad-hoc, and the +pattern that would make it explicit already exists — at the granularity of one file, and in +two-party form.** This ADR lifts the owned-vs-remainder partition from that one JSON document +to the whole `~/.claude` tree, and additionally **splits the remainder into its two real +owners** — the Claude-Code floor and the user — which the single-file merge currently +conflates. + +## Decision + +**Make the convergence reconciler's target an explicit, versioned, three-source manifest: +the union of a pinned Claude-Code baseline and the agent-ways manifest, with the user +remainder defined by construction as everything in neither set.** The reconciler converges +`actual → (CC-baseline ∪ agent-ways-manifest)` and, because the user remainder is outside +that target, never touches it. + +This is the **tree-wide generalization of the settings.json three-way merge**. +`settings_merge.rs` already proves the core move: carve a stable agent-ways-owned layer out +of an unmanaged remainder, idempotently, by tracking what we last applied. This ADR lifts +that move from one file to the whole tree — and, where the single-file merge stops at +two parties (ours vs. everything-else), the tree model **splits the remainder into the +Claude-Code floor and the user**, the distinction the file merge currently leaves implicit. + +### 1. Three manifest legs + +| Leg | What it is | Source of truth | Status today | +|---|---|---|---| +| **Claude Code baseline** | Files/dirs vanilla Claude Code owns or creates in `~/.claude`, **pinned to a real CC release tag** (e.g. `2.1.196`). | An empirical clean-room snapshot (see §2), supplemented by package introspection. | **New** — does not exist. | +| **agent-ways manifest** | What agent-ways projects: `PROJECTED_TREES`, `PROJECTED_FILES`, `PROJECTED_BINS`. | `git ls-files` in `$XDG_DATA/agent-ways`, via `ways manifest`. | **Exists** (`cmd/manifest.rs`). | +| **User remainder** | Everything in neither set. | Defined by construction — the complement of the union. | **Implicit today** (heuristic). | + +The target the reconciler converges toward is **CC-baseline ∪ agent-ways-manifest**. The +user remainder is *not in the target set at all*. This is the principled replacement for +heuristic "don't clobber the user's stuff" guards: the user's files are not protected by a +special case — they are simply **not in the manifest**, so idempotent convergence has no +entry that would write over them. + +This maps onto `settings_merge.rs`, with one deliberate refinement: + +- **agent-ways-manifest = the owned layer** — the file merge's `ours` (the `hooks` + + `WAYS_PERMS` slice). At the tree level, the **git-derived `ways manifest` *is* the record + of what we own**, so the tree leg needs no separately stored base the way the file merge + keeps `settings-applied.json` — git tracking plays that role. +- **user remainder = the unmanaged content** — the file merge's `(theirs − base − ours)`, + present and left alone (ours-by-absence). +- **CC-baseline = the new third leg.** The file merge has no separate notion of the + Claude-Code floor — it folds CC's default values into the unmanaged remainder. The tree + model promotes that floor to its own leg, because at tree scale the difference between + "Claude Code created this" and "the user created this" is load-bearing for cleanup and + migration in a way it is not for a single settings field. + +### 2. Capturing the CC baseline — empirical clean-room snapshot, pinned to a tag + +The key technical fact: **Claude Code's `~/.claude` footprint is mostly runtime-emergent, +not unpacked from the npm package.** `settings.json` defaults, `projects/`, +`history.jsonl`, `sessions/`, and `file-history/` are created **when Claude Code runs**, +not when it installs. A baseline built only from package introspection would miss most of +the floor it is meant to describe. + +So the authoritative capture method is an **empirical clean-room snapshot**: run vanilla +Claude Code with `HOME` pointed at an empty sandbox directory, exercise it minimally, and +record the resulting `~/.claude` tree. Static files that *do* ship in the package are +captured by package introspection and merged in. The result is **pinned to a specific CC +release tag**, which makes the baseline **versioned and reproducible** — a baseline for +`2.1.196` is a fact about `2.1.196`, regenerable by anyone with that tag and a sandbox. + +**Decision on storage and shipping:** the baseline ships as a **committed snapshot file +per CC tag** in `$XDG_DATA/agent-ways` (versioned alongside the code that consumes it), +**generated by a `ways` subcommand** that performs the clean-room run. The committed +snapshot is what the reconciler reads at convergence time; the generator is what the +maintainer runs to refresh it when a new CC release moves the floor. This keeps the +runtime path offline and deterministic (no sandbox spin-up during a user's SessionStart) +while keeping the snapshot honestly reproducible. **Refresh cadence is maintainer-driven, +not per-release** — the baseline is refreshed when CC's footprint actually changes, and +skew between refreshes is absorbed by §3. + +### 3. Tolerating CC version skew — the baseline is an allow-pattern set, not an exact list + +The user's installed CC version will rarely equal the pinned baseline tag. If the baseline +were an exact file list, every patch-level CC release would produce spurious "unclassified" +files and risk the reconciler treating a genuine CC runtime file as user remainder (or vice +versa). + +**Decision:** the CC baseline is consumed as an **allow-pattern set** (path globs / +prefixes — `projects/`, `sessions/`, `history.jsonl`, `file-history/`, `settings.json`, +`statsig/`, `todos/`, …), not an exact inventory. Classification asks *"does this path +match a known CC-owned pattern?"* rather than *"is this path byte-identical to the +snapshot?"* Patterns are far more skew-tolerant than file lists: a new session file under +`sessions/` is still obviously CC-owned. The per-tag exact snapshot remains the **evidence** +from which the pattern set is derived and audited, but the pattern set is the runtime +contract. Skew within the pattern set's tolerance is a non-event; skew that introduces a +*new top-level CC artifact* is the signal that the maintainer should regenerate (§2). + +### 4. The `ways` surface — a new `ways classify` + +**Decision:** add a new **`ways classify`** subcommand that emits the three-way +classification of a real `~/.claude` — for each path, which leg it belongs to +(cc-baseline / agent-ways / user-remainder) — rather than overloading `ways manifest` or +`ways status`. + +Rationale, by single-responsibility: + +- `ways manifest` answers *"what does agent-ways project?"* — one leg, derived purely from + git, with no reference to a live `~/.claude`. Folding CC-baseline and live-tree + classification into it would give it two reasons to change. +- `ways status` is a health/observability summary for the operator, not a per-path + classification emitter. +- `ways classify` is the natural home for the **set operation over a live tree**: it + consumes the agent-ways manifest (leg 2) and the CC baseline pattern set (leg 1), reads + the actual `~/.claude`, and emits the three-way partition that install, repair, migrate, + and cleanup all consume. + +`ways classify` is the read-only classifier; the reconciler (`cmd/reconcile.rs`) is the +mutating consumer that acts on the partition. + +### 5. Precedence when a path is claimed by both CC baseline and agent-ways + +Some paths are claimed by **both** legs — `settings.json` is the canonical case: it is part +of CC's baseline floor *and* a file agent-ways writes into. **Decision:** the agent-ways +manifest takes precedence for *projection* (agent-ways is allowed to write the path), but +the **already-specialized three-way merge (`settings_merge.rs`) governs that write** — the +whole-tree convergence delegates any path that is both CC-baseline and agent-ways-managed to +the per-file merge that already knows how to combine a CC-default base, the agent-ways layer, +and user fields without clobbering the user. In other words: tree-level classification routes +`settings.json` to file-level classification; the coarse leg precedence (agent-ways > CC for +projection) selects *who may write*, and the fine-grained merge decides *what to write*. No +other current path is expected to be doubly-claimed; if more emerge, the same rule applies — +overlap routes to a path-specific reconciler, defaulting to "agent-ways may write, user +fields preserved." + +### 6. Consumers — all four ADR-144 from-states, plus cleanup + +The single explicit manifest is consumed by every entry condition ADR-144 defined, plus +cleanup: + +- **fresh** — materialize `(CC-baseline ∪ agent-ways-manifest)` into an empty tree. This is + what makes the **native projection installer fall out as "bootstrap shim + reconcile in + fresh state,"** with almost no new code: there is finally an explicit target to materialize + *into nothing*. Wiring this closes the 1.0.0 fresh-install bug (Context #1). +- **drifted** — re-materialize missing/broken entries of the union (repair). +- **out-of-date** — re-derive the agent-ways leg from the advanced `$XDG_DATA` HEAD, + materialize the delta, prune orphans (update). +- **legacy-in-place** — the migrator classifies the old clone against all three legs: + agent-ways files relocate to `$XDG_DATA`, the **user remainder** lifts to + `$XDG_CONFIG`/`$XDG_STATE`, and CC-baseline files stay. The classifier is exactly the + rescue mechanism ADR-144's sharpest Negative demanded. +- **cleanup** — the **user remainder is precisely the set that is safe to prune or flag**; + conversely, agent-ways orphans (in the manifest's history but no longer git-tracked) are + safe to remove outright. This is the principled form of the by-hand cleanup done in the + session that motivated this ADR. + +## Consequences + +### Positive + +- **The convergence target becomes explicit and testable.** ADR-144's "converge toward the + manifest" gains a concrete, three-legged, versioned definition of *the manifest* — install + / repair / migrate / cleanup all read one artifact instead of each guessing. +- **The 1.0.0 fresh-install bug has a principled fix.** Fresh install stops being a missing + feature and becomes "reconcile in the fresh from-state against the explicit target," + exactly as ADR-144 promised it would fall out. +- **"Don't clobber the user" stops being a guard and becomes a set property.** The user + remainder is untouched not because of a special case, but because it is not in the target — + the most robust form of the protection, and the one least likely to regress. +- **One pattern, two granularities.** The settings.json three-way merge stops being a + one-off; it is now the file-level instance of the same model the whole tree uses, which + makes both easier to reason about. +- **Migration's rescue problem gets a real classifier.** ADR-144's "detect and rescue + hand-edited core files" becomes a concrete `ways classify` output rather than an + aspiration. + +### Negative + +- **A new external dependency surface: the CC baseline must track Claude Code.** agent-ways + now maintains a model of a tree it does not own and that changes on Anthropic's clock. The + pattern-set design (§3) absorbs most skew, but a CC release that introduces a new top-level + artifact requires a maintainer refresh; until then that artifact classifies as user + remainder, which is the *safe* failure direction (we leave it alone) but a misclassification + nonetheless. +- **The clean-room snapshot is real machinery to build and keep honest.** A generator that + spins up sandboxed CC, exercises it enough to emit its runtime files, and diffs the result + is non-trivial and itself version-sensitive — and "exercise it minimally" is a fuzzy + contract (which CC features must run to materialize which files?). +- **`settings.json` remains the doubly-claimed seam** (ADR-142 / ADR-144 already flagged it); + this ADR routes it correctly but does not remove the shared-write risk — it just states + precisely where the tree model hands off to the file model. +- **Pattern sets can be wrong in both directions.** Too broad, and a genuine user file under a + CC-shaped path is misclassified as CC-owned and skipped by cleanup; too narrow, and a CC + runtime file is treated as user remainder. The exact per-tag snapshot is the audit evidence, + but tuning the pattern breadth is an ongoing judgment. + +### Neutral + +- `ways manifest` is unchanged; this ADR adds `ways classify` beside it rather than altering + the existing leg. +- The CC baseline snapshot is versioned in `$XDG_DATA/agent-ways` and so is **replaced + wholesale on update** like the rest of the application (ADR-142's `$XDG_DATA` durability + contract) — losing it is a re-derive, not data loss. +- This ADR sharpens, but does not resolve, ADR-142's open `$XDG_STATE` ↔ Claude-Code-owned + boundary (see Open Questions); it gives that boundary a *mechanism* (the CC baseline pattern + set) without fixing where the line sits. + +## Alternatives Considered + +- **Leave the target implicit; keep heuristic classification.** The status quo. Rejected for + the three consequences in Context — the fresh-install path stays unwired, and migrate / + cleanup keep guessing. ADR-144 already rejected the weaker "did a prior sync write it" + membership test for the agent-ways leg; this ADR extends the same reasoning to the CC and + user legs. +- **Model the CC baseline by package introspection alone (no clean-room run).** Rejected: + most of CC's `~/.claude` footprint is runtime-emergent, so an install-time inventory would + miss `projects/`, `sessions/`, `history.jsonl`, `file-history/`, and the defaulted + `settings.json` — i.e. most of the floor. Introspection is kept only as a *supplement* for + the genuinely static files. +- **Pin the CC baseline as an exact per-tag file list (no pattern set).** Rejected: it is + brittle under the inevitable version skew between the pinned tag and the user's installed + CC — every patch release would manufacture spurious unclassified paths. The exact snapshot + is retained as evidence; the *runtime contract* is the skew-tolerant pattern set (§3). +- **Extend `ways manifest` to emit all three legs instead of adding `ways classify`.** + Rejected on single-responsibility grounds: `ways manifest` answers "what does agent-ways + ship," derived purely from git with no live-tree or CC dependency. Bolting CC-baseline and + live `~/.claude` classification onto it gives one command two reasons to change and couples + a pure git derivation to an external-dependency model. +- **Generate the CC baseline live at each SessionStart (no committed snapshot).** Rejected: + spinning up a sandboxed CC run on the user's machine at session start is slow, non- + deterministic, and fragile; the committed-per-tag snapshot keeps the runtime path offline + and reproducible, with regeneration a deliberate maintainer act. + +## Open Questions + +These are recorded deliberately undecided; they refine, and partly inherit, ADR-142/144's +open questions. + +- **The ADR-142 `$XDG_STATE` ↔ Claude-Code-owned boundary.** The CC baseline pattern set is + the natural place to *draw* this line — runtime state CC creates (`projects/<slug>/memory/`, + per ADR-128) that overlaps agent-ways' own `$XDG_STATE` claims must land on one side. This + ADR provides the mechanism but does not commit the boundary; auto-memory (ADR-128) sitting + in CC's `projects/<slug>/memory/` is the specific unresolved overlap. +- **What "exercise vanilla CC minimally" must include** to materialize the full runtime + footprint — which CC operations are needed to emit which files, and how the generator + guarantees it captured the whole floor rather than a subset. +- **Pattern-set breadth and review process** — how broad each CC-owned glob should be, and how + the maintainer audits a refreshed snapshot against the prior pattern set to catch new + top-level artifacts. +- **Refresh trigger** — whether baseline refresh is purely manual (maintainer notices a CC + release moved the floor) or gets a lightweight detector (a CI clean-room run that diffs the + current snapshot against the latest CC tag and flags drift). +- **Whether this stays a separate ADR or folds into ADR-144.** Recommendation: keep separate — + the CC-baseline capture method and the `ways classify` surface are each substantial enough to + warrant their own recorded decision, and ADR-144 is already long. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-168-instance-addressable-directed-messaging-for-same-cwd-attend-siblings.md b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-168-instance-addressable-directed-messaging-for-same-cwd-attend-siblings.md new file mode 100644 index 00000000..650387a5 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/archive/system/ADR-168-instance-addressable-directed-messaging-for-same-cwd-attend-siblings.md @@ -0,0 +1,128 @@ +--- +status: Rejected +date: 2026-07-17 +deciders: + - aaronsb + - claude +related: + - 120 + - 124 + - 129 + - 136 +--- + +# ADR-168: Instance-addressable directed messaging for same-cwd attend siblings + +> **ARCHIVED — 2026-08-13.** No longer part of the active architecture set. Kept for history +> and so existing references still resolve. +> +> **Why:** Never implemented and never debated. The Decision section still reads 'Proposed (pending debate)'; no cwd-scoped addressing exists in attend. ADR-136's one-inbox-per-cwd invariant stands unchanged. +> +> Nothing below this line has been edited. + +## Context + +Directed `@Nickname` messaging in `attend-chat` is **cwd-keyed**. `resolve_nickname` +(`tools/attend-chat/src/chip/routing.rs`) maps a nickname to a `KnownIdentity.cwd`, +and `write_signal` posts to `signals_base/<encoded-cwd>/`. The cwd *is* the routing +address. + +Two prior decisions collide on this point: + +- **ADR-129** gave same-cwd sessions distinguishable identities via instance + suffixes — `@Tamsin-alpha`, `@Tamsin-beta` — so the legend and Tab-completion + can tell two Claude Code sessions in one working directory apart. This is a + **display + completion** affordance: `with_instance` decorates the nickname, + but `KnownIdentity.cwd` is *identical* for both siblings. +- **ADR-136** established the message lifecycle and the invariant that there is + **one attend agent per canonicalized cwd** (the `last_inbound` same-cwd filter + in `sensor-peers` depends on it). + +The consequence is a promise/behavior mismatch. `@Tamsin-alpha` and +`@Tamsin-beta` *display* as two addressable agents but *resolve to the same cwd*. +`resolve_recipients` dedups destinations by cwd dir, so a message addressed to +both collapses to a single inbox write; and because the inbox is cwd-keyed, a +message addressed to *one* sibling is not delivered to that sibling specifically — +it lands in the shared cwd inbox. There is no way, under the current model, to +direct a message to exactly one of several instances sharing a cwd. + +This surfaced as a user-reported symptom: mentioning two agents appeared to +deliver to only one. The routing fix in `c7c938f` (address every recipient in a +leading `@`/`#` run) resolved the general multi-recipient case, but the same-cwd +sibling case is not a parsing bug — it is a limit of cwd-keyed addressing, and +resolving it is a design decision rather than a patch. + +Scope note: this ADR is specifically about *addressing precision for co-located +instances*. The separate "directed sends don't echo in the sender's own view" +defect is a display concern fixed independently (issue #370 / its PR) and does +not depend on this decision. + +## Decision + +**Proposed (pending debate — status Draft).** Adopt **Option A** now: treat a +mention that resolves to a cwd hosting multiple live instances as **cwd-scoped** +addressing, and make the UI say so — collapse sibling suffixes at send time and +report "reaches all instances in `<cwd>`" in the status line / echo. This keeps +ADR-136's one-inbox-per-cwd invariant intact and ships a small, honest change. + +If instance-*precise* delivery proves to be a real need (not just legend +distinguishability), escalate to **Option C** (session-id addressing) as the +follow-on decision — it removes the cwd/instance coupling at the root and is a +cleaner long-term model than sub-inboxing (Option B). This ADR does not commit to +C; it records that C is the preferred path *if* precision is required. + +The choice hinges on one question for the deciders: **is addressing one specific +co-located instance a workflow anyone actually needs, or is distinguishing them +in the legend enough?** The answer selects A alone vs. A-then-C. + +## Consequences + +### Positive + +- Removes the display-vs-routing mismatch: what the UI says a mention does + matches what the bus does. +- Option A is cheap and preserves the ADR-136 invariant, so it can land without + reworking the sensor's `last_inbound` filter. +- Framing the decision around "is instance-precise delivery a real need" keeps us + from building instance inboxes speculatively. + +### Negative + +- Under Option A, ADR-129's per-instance *addressability* remains display-only; + users who read `@Tamsin-alpha` as "only alpha" must learn it means "alpha's + cwd, both siblings." The UI copy has to carry that. +- Deferring C means the precise-delivery capability, if later needed, is a second + migration rather than one move now. + +### Neutral + +- Either precise option (B or C) requires putting instance/session identity on + the wire and revisiting the ADR-136 "one attend per cwd" invariant, since two + instances in a cwd would then each read a distinct inbox lane. + +## Alternatives Considered + +- **Option A — Accept cwd-level addressing; make the UI honest.** No new inbox + structure. At send time, siblings collapse to their shared cwd; status/echo + state the message reaches every instance there. *Chosen as the immediate step.* + Cheapest, invariant-preserving. Cost: cannot target one sibling. + +- **Option B — Instance-keyed sub-inboxes.** Extend the routing key from cwd to + `(cwd, instance-id)`; write to `signals_base/<encoded-cwd>/<instance>/` and have + each instance read only its own lane. Faithful to ADR-129's promise, but layers + instance structure onto the cwd dir, still couples routing to cwd, and forces a + rework of the ADR-136 same-cwd `last_inbound` filter. Rejected as the more + complex of the two precise options with no offsetting benefit over C. + +- **Option C — Session-id addressing.** Route directed sends by target session + UUID (already unique per session and the key ADR-129 heartbeats use) rather than + by cwd: the signal carries a target session id; a recipient accepts signals + addressed to its own id. Decouples routing from cwd entirely, handles same-cwd + siblings for free, and is more precise for directed messaging generally. + Preferred *if* instance-precise delivery is required — but a larger change to + the addressing model (`accept_path`, target resolution, the encode scheme), so + not adopted pre-emptively. + +- **Do nothing.** Leave the mismatch. Rejected: the UI actively implies a + capability the bus does not provide, which is the exact shape of the original + bug report. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-300-documentation-structure.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-300-documentation-structure.md new file mode 100644 index 00000000..2a51d247 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-300-documentation-structure.md @@ -0,0 +1,194 @@ +--- +status: Accepted +date: 2026-02-17 +deciders: + - aaronsb + - claude +related: + - ADR-004 + - ADR-005 + - ADR-014 +--- + +# ADR-300: Documentation Structure + +## Context + +The documentation makes a strong first impression. A reader landing on the README quickly grasps the core concepts: ways inject contextual guidance, skills are semantically discoverable, governance traces policy to agent behavior. The Severance metaphor lands. The pitch works. + +Then the reader tries to *do* something — add governance to a way, create a way for a new domain, understand how semantic matching actually works — and the docs sprawl. There's no clear path from "I understand the concept" to "I know how to do this." Content appears in multiple files with no indication of which is canonical. References to NCD and BM25 contradict each other across files because the duplication made drift invisible. Domain docs describe ways that were never built. The governance system has code but no data pipeline. The README tries to be both landing page and reference manual and succeeds at neither. + +This isn't a restructuring — the docs were never structured. They grew file-by-file as features were added, each locally coherent but collectively incoherent. This ADR establishes what each documentation layer is for, identifies what's cruft, and defines how a reader moves from concept to practice. + +### What triggered this + +Reading through the documentation and repeatedly encountering NCD references where BM25 should be. The specific inconsistency revealed the general problem: there's no coherent structure governing what goes where, so every addition risks drift. + +### The structural audit + +A full audit (`docs/audit-findings.md`) cataloged specific issues. A programmatic link graph (`scripts/doc-graph.sh --docs-only --stats`) confirmed the structural problems: + +- **41 doc files, 29 links, 34 dead ends, 23 orphans** +- Only **3 hub files** do all the linking (README, docs/hooks-and-ways/README, governance/README) +- `docs/hooks-and-ways/README.md` — the best guide page (10 outgoing links) — is **unreachable from the main README** +- **7 domain docs** (cloud.md, mcp.md, ai.md, etc.) are completely isolated — zero links in or out +- **5 files** still present gzip NCD as primary when BM25 is the actual implementation +- **6 content areas** duplicated across 2-4 files each, several already diverged +- **README at 463 lines** — contains full tutorials that already exist in docs/ + +### Already resolved + +As part of this ADR's development: + +- **ADR tooling adopted**: `docs/scripts/adr` installed, `docs/architecture/adr.yaml` configured with domain numbering (system 100s, governance 200s, docs 300s) +- **ADR path reconciled**: project now uses `docs/architecture/` matching the way's convention +- **Legacy ADRs triaged**: ADR-001, 002, 003 deleted (superseded/irrelevant). ADR-004, 005, 013, 014 kept as legacy +- **Doc graph tool created**: `scripts/doc-graph.sh` programmatically maps the documentation link graph, identifies dead ends and orphans + +## Decision + +### 1. Define what each documentation layer is for + +| Layer | Location | Audience | Purpose | +|-------|----------|----------|---------| +| **Landing** | `README.md` | First-time visitor | "What is this? How do I try it?" | +| **Guide** | `docs/hooks-and-ways/` | Practitioner | "How do I do X?" | +| **Reference** | `docs/hooks-and-ways.md`, `docs/architecture.md` | Contributor/debugger | "How does X work internally?" | +| **Policy source** | `governance/policies/*.md` | Governance chain | Source docs referenced by way provenance | +| **ADRs** | `docs/architecture/` | Decision record | Design decisions with context and consequences | +| **Machine layer** | `hooks/ways/*/way.md` | LLM runtime | Injected context — self-contained by design | + +Content belongs in exactly one layer. Other layers link to it. + +### 2. README becomes a landing page (~200 lines) + +The README answers three questions: "What is this?", "How do I start?", and "Where do I go deeper?" + +**Keep in README:** +- Hero image + Severance tagline +- What this is (overview + Mermaid diagram) +- Prerequisites table (corrected: BM25 primary, gzip fallback) +- Quick Start: fork-and-clone (recommended) and direct clone +- How It Works: trigger flow + directory tree (~20 lines) +- Configuration (`ways.json`) +- Built-in Ways table +- Philosophy +- Updating + License + +**Replace with summary + link:** + +| Current README section | Canonical location | +|----------------------|-------------------| +| Creating a Way + Frontmatter | `docs/hooks-and-ways/extending.md` | +| Semantic Matching | `docs/hooks-and-ways/matching.md` | +| Way Macros | `docs/hooks-and-ways/macros.md` | +| Project-Local Ways | `docs/hooks-and-ways/extending.md` | +| Ways vs Skills comparison | `docs/hooks-and-ways/README.md` | +| Governance example + chain | `governance/README.md` | +| Once-Per-Session Gating details | `docs/hooks-and-ways.md` | + +### 3. Remove aspirational domain docs + +Seven files in `docs/hooks-and-ways/` describe domains with zero way.md implementations: + +`cloud.md`, `mcp.md`, `ai.md`, `research.md`, `enterprise.md`, `devops.md`, `sysadmin.md` + +These are aspirational — they describe what ways *could* exist, not what does exist. A reader encountering `cloud.md` expects to find cloud ways and doesn't. Remove these files. When ways are built for new domains, documentation should accompany the implementation. + +Policy source docs moved to `governance/policies/` — they're governance chain artifacts, not system documentation. + +### 4. Address governance pipeline gap + +> *Reconciled (2026-07): the tools named below (`governance.sh`, `provenance-scan.py`, +> `provenance-verify.sh`) were consolidated into `ways governance` (ADR-111) and now move +> to the `ways-audit` binary (ADR-151); provenance moved from way frontmatter to a +> `provenance.yaml` sidecar (ADR-110); and the subsystem is reframed as compliance +> claims/findings (ADR-200). The "pipeline gap" below is addressed there.* + +The governance system has code (governance.sh, provenance-scan.py, provenance-verify.sh) and provenance metadata in way frontmatter, but output artifacts are not generated, tracked, or consumed. The docs explain what governance *is* but not how to *use* it. + +Documentation fixes (scoped to this ADR): +- **governance/README.md**: add "Getting Started" at the top — `bash governance/governance.sh` is the entry point +- **docs/hooks-and-ways/provenance.md**: add "add provenance to your first way" walkthrough +- **State the gap honestly**: governance output is ephemeral, not CI-integrated. That's the current design, not an oversight + +Pipeline fixes (separate ADR): whether to track output, integrate CI, or keep ephemeral. + +### 5. Create docs/README.md as a map + +A `docs/README.md` serves as a directory — not a guide. It answers "where do I find X?" with a file listing and one-line descriptions. The role-based reading paths stay in `docs/hooks-and-ways/README.md`. + +Explicitly notes the relationship between `hooks-and-ways.md` (reference) and `hooks-and-ways/` (guides). + +### 6. Fix stale content + +| Item | Action | +|------|--------| +| NCD/BM25 in 5 files | Update to BM25-primary, NCD-fallback | +| Prerequisites docs (4 files) | Note gzip is for fallback path | +| ADR-014 | Accept (it's deployed) | +| `nested-ways-exploration.md` | Archive — nested ways are implemented | +| `rationale.md` | Update matching tiers, note model-based deprecated | +| ADR-013 lines 175, 201 | Update "gzip NCD" → "BM25 (with NCD fallback)" | + +### 7. Severance theme: functional purpose test + +The theme stays where it serves functional purposes: + +- **"Write for the innie"** — concrete authoring principle (agent has no memory of previous sessions) +- **"Lumon handbooks"** — frames ways as institutional knowledge transfer +- **"The floor above"** — maps governance/policy separation to management/worker separation + +**Guideline**: a Severance reference earns its place if it makes a concept *more understandable* than plain language alone. If removing it loses nothing, don't add it. + +### What this ADR does NOT cover + +- **Governance pipeline** (should output be tracked? CI-integrated?). Separate ADR. +- **Agent files** in `agents/` — potentially stale, separate assessment +- **Way file content** — this ADR is about docs, not the ways themselves + +## Consequences + +### Positive +- Each doc layer has a defined purpose — new content has an obvious home +- README drops from ~463 to ~200 lines — scannable landing page +- Aspirational domain docs removed — no more documenting unbuilt features +- NCD/BM25 inconsistency fixed across all files +- Governance docs gain a practitioner entry point +- ADR tooling adopted — project follows its own conventions +- `scripts/doc-graph.sh` provides ongoing coherence checking + +### Negative +- **Link rot risk**: moving from inline content to links. Mitigation: `doc-graph.sh` detects broken links. +- **Contributor friction**: contributors must learn which file is canonical. The "everything in README" model had zero navigation overhead. +- **README loses grep-ability**: searching the repo for "semantic matching" finds a summary instead of the explanation. + +### Neutral +- No structural changes to the hooks-and-ways guide/reference architecture — it works +- Severance theme stays exactly where it is — no additions, no removals + +## Alternatives Considered + +### Just fix the NCD/BM25 references and leave everything else +**Rejected**: The duplication that caused drift still exists. The next feature addition will duplicate content again and diverge again. + +### Full docs restructure (guides/, reference/, concepts/ hierarchy) +**Rejected**: The existing docs tree works. The problem isn't the tree — it's that the README tries to be the tree, and aspirational docs sit alongside real guides. + +### Documentation site generator (mdbook, Docusaurus) +**Rejected**: This is a config repo, not a framework. Adding build tooling contradicts the "bash + jq, no dependencies" philosophy. + +### Keep README comprehensive, add a "canonical" marker system +**Rejected**: Metadata doesn't prevent drift — writers still need to update multiple files. + +## Implementation Plan + +1. Remove aspirational domain docs (cloud.md, mcp.md, ai.md, research.md, enterprise.md, devops.md, sysadmin.md) +2. Create `docs/README.md` (map, not guide) +3. Create `docs/installation.md` (full install guide extracted from README) +4. Slim README: replace duplicated sections with summary + link +5. Fix NCD/BM25 references across all identified files +6. Update `rationale.md` matching tiers +7. Archive `nested-ways-exploration.md` +8. Add "Getting Started" to governance/README.md +9. Validate: `scripts/doc-graph.sh --docs-only --stats` — check for new dead ends diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md new file mode 100644 index 00000000..1164513f --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-301-situated-socialization-as-canonical-framing-and-documentation-prose-refactor.md @@ -0,0 +1,78 @@ +--- +status: Accepted +date: 2026-06-09 +deciders: + - aaronsb + - claude +related: [] +--- + +# ADR-301: Situated socialization as canonical framing and documentation prose refactor + +## Context + +The project's documentation describes its mechanisms almost entirely in invented vocabulary — ways, attend, disclosure, reheat, firing, salience floors, insistence. The terms are internally coherent, but they are defined only by reference to each other. A reader arriving from outside (or an agent reasoning about the project) has no anchor: nothing in the prose says what body of existing knowledge any mechanism belongs to, so every document carries a tax of re-explanation, and the project as a whole resists the one-sentence answer to "what is this?" + +This was not a style choice. The author lacked the cross-field vocabulary at authoring time, because the relevant terms are scattered across four or five disciplines that rarely cite each other. They exist: + +| Project term / mechanism | Established term | Source | +|---|---|---| +| Ways (the corpus and its delivery) | **Organizational socialization**, delivered as **situated learning** | Van Maanen & Schein 1979; Lave & Wenger 1991 | +| Ways as agent memory | **Procedural memory** retrieved into working memory | CoALA (Sumers et al. 2023); ACT-R | +| Salience decay + re-disclosure at a floor | **Forgetting curve** and **spaced repetition** | Ebbinghaus; SuperMemo/Anki lineage | +| Refractory gates, habituation, burst detection | **Base-level activation decay** | ACT-R | +| Attend's emission governor, quiet footnotes | **Calm technology**; **interruption cost**; **alarm management** | Weiser & Brown; Mark; ISA-18.2 | +| Peer presence, heartbeats, instance identity | **Workspace awareness** | CSCW (Dourish & Bellotti 1992) | +| Way authoring from team norms | **Externalization of tacit knowledge** | Nonaka & Takeuchi (SECI) | +| Progressive disclosure | Progressive disclosure (already converged — Anthropic uses the term for Skills) | — | + +The name "ways" itself turns out to be load-bearing rather than arbitrary. It comes from the phrase *"that's the way we do it around here"*, and each word maps to architecture: **"we"** — multiple aligned actors (the peer layer); **"around"** — approximate boundaries (embedding-based matching rather than exact rules); **"here"** — local scope (project ways overriding global). The phrase is itself established management vocabulary ("ways of working"). + +Two findings from this framing belong in the record because they shape what the documentation should say: + +1. **Re-enactment substitutes for internalization.** Human socialization theory assumes the newcomer persists — the organization pays the onboarding cost once and internalization does the rest. An LLM session cannot internalize (no weight updates, no carried memory); every session is a new hire. The system therefore re-enacts socialization mechanically, every session, at the moment of relevant action, on a spaced schedule that substitutes for the memory the agent does not have. This is the genuinely novel constraint relative to the human literature, and the honest justification for the firing-dynamics machinery. + +2. **The durability split.** Ablation testing (removing the ways system) produces the same behavior across model tiers: agents become approval-seeking — constant follow-ups or exhaustive hedging. The established name for the cause is **preference uncertainty** (principal–agent theory): an agent that does not know its principal's norms can only ask or hedge. This separates the system's two functions by durability: the *scheduling* half (decay curves, re-disclosure) compensates for a model deficiency and may erode as models improve; the *routing* half (just-in-time delivery of local norms the model cannot know because it was never told) is structural and permanent. The scheduling half of ways is a patch on current models; the routing half is a permanent answer to preference uncertainty. + +## Decision + +Adopt **situated socialization for language-model agents** as the project's canonical framing, and refactor documentation prose to lead with established vocabulary. + +The canonical one-paragraph description: + +> Ways is organizational socialization for language-model agents. Because an LLM session cannot internalize norms — every session is a new hire — the system re-enacts socialization mechanically: local norms ("the way we do it around here") delivered situated, at the moment of relevant action, on a spaced schedule that substitutes for the memory the agent does not have. Attend is the awareness an employee would otherwise have ambiently: what is changing, who else is working, what deserves attention. + +Concretely: + +1. **Terminology anchors are normative.** The mapping table above moves into the documentation as a reference. Project-coined terms remain in use, but on first use in any document they are introduced as implementations of their established anchor ("salience decay — the forgetting curve applied to injected guidance — …"), never bare. +2. **Prose refactor, not content rewrite.** Each document in scope is revised to state *what the mechanism is* in established terms first, then *how this project implements it* in project terms. Content, diagrams, and examples are preserved; the register changes. Scope: `README.md`, `docs/hooks-and-ways/`, `docs/attend-and-monitor/`, `docs/design-notes/`, and skill/way prose that describes the system to readers (e.g. `skills/attend/SKILL.md` intro). +3. **ADRs are exempt.** Existing ADRs are historical decision records and remain untouched. New ADRs adopt the vocabulary going forward. +4. **The ways corpus is exempt.** The machine layer (`hooks/ways/**/*.md` bodies) is guidance for agents in other projects, already terse and functional; it is not part of the descriptive-noise problem. +5. **The durability split is documented.** The scheduling-vs-routing distinction and its maintenance implication (retune half-lives per model generation via `ways tune`; defend the routing pipeline) is written into the architecture overview so future maintainers know which parts to let atrophy. + +## Consequences + +### Positive + +- The project becomes explainable in one sentence built entirely from terms with literatures behind them: *procedural memory for coding agents, maintained by spaced repetition, plus a perception loop with alarm management.* +- Prior art becomes searchable. Each mechanism now names the field it can be evaluated against, instead of appearing sui generis. +- Naming future work gets easier — deferred features inherit anchors (insistence tracker → escalation in alarm management; consequence model → projection/appraisal). +- The ablation observation ("agents become needy without ways") gains a precise causal account (preference uncertainty), strengthening the case the documentation makes. + +### Negative + +- The refactor touches most reader-facing documents; it is real editing effort and risks introducing inconsistency mid-flight if done piecemeal. +- Anchors are analogies with limits, not identities — e.g. CoALA's procedural memory includes agent code, ACT-R activation governs retrieval rather than injection. Over-claiming equivalence would trade one kind of noise for another. The refactor must say "this is X applied to Y," not "this is X." +- Academic register can curdle into jargon of a different flavor. The test for every revised paragraph is whether it got *easier* to read for a newcomer, not whether it cites more. + +### Neutral + +- Existing ADRs and the ways corpus are unchanged by design. +- The invented terms survive — they name the implementations, which the established terms do not do. This ADR settles their *introduction*, not their existence. +- The etymology of "ways" (we / around / here) becomes part of the documented narrative rather than conversational lore. + +## Alternatives Considered + +- **Glossary only** — add a terminology appendix, leave prose as-is. Rejected: a glossary is a patch on the register problem; readers hit the bare invented term first and the noise remains in every document. +- **Full rename to established terms** — rename components (ways → norms, attend → awareness, salience → activation). Rejected: the invented names are good — "ways" in particular encodes the design (we/around/here) — and the churn would touch every binary, hook, and document for negative gain. +- **Do nothing** — accept the invented vocabulary as the project's idiom. Rejected: the descriptive noise is a measured cost (the project's author cannot briefly explain the project), and the durability argument — the most important strategic fact about the system — currently exists nowhere in the documentation. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-302-unified-documentation-model.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-302-unified-documentation-model.md new file mode 100644 index 00000000..5ee51741 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-302-unified-documentation-model.md @@ -0,0 +1,286 @@ +--- +status: Accepted +date: 2026-06-19 +deciders: + - aaronsb + - claude +related: + - ADR-300 + - ADR-301 +--- + +# ADR-302: A unified documentation model — typed graph, ways packaging, cross-repo convergence + +> **Scope note.** This ADR was first drafted (2026-06-16) as "Diátaxis as the +> documentation classification model" and grew off the rails by stapling three +> things together. It is rewritten here around a single spine, renamed to match +> (`ADR-302-unified-documentation-model.md`), and kept in the `documentation` +> domain (300s) — its subject is documentation tooling; the ways packaging in §7 +> is the delivery vehicle, not the subject. + +## Spine + +Documentation is **one typed graph**. The type lives in **frontmatter and a +linter**, not in filenames. The **filesystem is a serialization** of the graph +for human readers — a view, not the source of truth. We **package the model as a +`documentation/` ways root plus skills**, and **converge repos onto it** through +the existing `project-init` / `project-audit` rail. Diátaxis, ADR/doc numbering +unification, `doclint`, and folder layout are all *facets of this one spine*, not +separate decisions. + +## Context + +ADR-300 established *where* documentation lives (location/audience layers) and +shipped `doc-graph.sh`. ADR-301 directed prose to lead with established +vocabulary. Two gaps remain: + +1. **ADR-300's "Reference" layer conflates two needs** — "how X works internally" + (mechanism = *explanation*) and "the exact facts" (*reference*) under one word. +2. **ADRs and prose docs are two uncoordinated systems.** Their numbering has + never been coordinated; "where is the decision behind this page?" is answered + by grepping and hoping. + +The `knowledge-graph-system` project (KGS) independently worked this through and +arrived at a corrected, *running* model — ADR-900 (domain numbering), ADR-908 +(documentation strategy, amended and corrected 2026-06-16), and a working +`doclint.py`. **KGS is the reference implementation and the driver.** Its hard-won +corrections are load-bearing here and are adopted rather than re-derived: + +- Diátaxis is a **closed 2×2** — four modes, no fifth. KGS tried an `operations` + fifth mode and retired it (2026-06-16) as a *category error*: it put an + *audience* on the *posture* axis. +- The catalog id is **`DD.NNN.P`** (domain band · domain-scoped serial · trailing + mode pole), *not* `<domain>.<mode>.<serial>`. Baking the mutable mode into the + middle of an identity churns the "part number" on every reclassification. + +### The thesis that makes this worth the rigor + +A typed system's sustainable richness is bounded by **who maintains it**. Humans +cap out at a level of cross-referential bookkeeping and respond by inventing loose +conventions (organize-by-audience, folder-by-feature) — and, on complex projects, +by hiring a person whose whole role is to hold the schema in working memory. That +cap is a property of the *maintainer*, not the problem. An AI coding agent does +not cap at the same place: a typed graph it would take a human minutes to reason +through, it sustains in one pass. So the right design **pushes rigor past the +human-convention comfort line on purpose** — into frontmatter and a linter, where +the maintainer that actually maintains it pays almost nothing — and demotes the +filesystem to an ergonomic *serialization* for the humans who still read and +occasionally edit. + +## Decision + +Adopt a unified documentation model with the following parts. `adr.yaml` is the +single source of truth for the domain axis, shared by ADRs and docs alike. + +### 1. One typed graph + +Docs and ADRs are **nodes** in one graph; `related` / `supersedes` are **edges**. +Not two systems with a convention bolted between them. "Where's the decision +behind this page?" becomes an edge traversal. + +### 2. The type lives in frontmatter, not the filename + +Every catalog node carries frontmatter that is **dual-readable** — `doclint` reads +it as the type, Obsidian reads it as a graph: + +```yaml +--- +id: 04.001.H # DD (domain band) . NNN (domain-scoped serial) . P (mode pole) +domain: auth # ADR-900/adr.yaml domain key — the shared "first octet" +mode: how-to # Diátaxis: tutorial | how-to | reference | explanation +aliases: ["04.001.H", "04.001"] # stable handles so [[04.001]] resolves in Obsidian +related: ["[[ADR-300]]", "[[05.002]]"] # wikilink edges — doclint strips [[ ]] & resolves; Obsidian draws them +supersedes: [] +--- +``` + +- **Identity is `DD.NNN`** — assigned once, immutable, never reused. +- **The mode pole `P` trails** because mode is the *mutable* attribute; + reclassifying flips `…​.H → …​.E` and the identity is untouched. The pole is a + *view* of `mode:`, enforced to agree, never the key. +- **Serials are domain-scoped**, so any id collision is a real clash. +- **Edges are `[[wikilinks]]` and `aliases` carries each node's stable handle**, so + the same `related`/`supersedes` lists `doclint` validates also render as a live + graph in Obsidian (§6). `doclint` strips the `[[ ]]` and resolves the inside + against ids/aliases; Obsidian needs the alias to resolve `[[ADR-300]]` to + `ADR-300-….md`. Canonical type fields (`id`/`domain`/`mode`) stay first-class — + no `tags` mirror to drift. + +This explicitly **rejects** the original draft's `3.H.4` (`domain.MODE.serial`), +which bakes the mutable classifier into the middle of the identity — the exact +scheme KGS retired. + +### 3. Classification is Diátaxis — four modes, closed 2×2 + +`tutorial | how-to | reference | explanation`, derived from two orthogonal axes +(action/cognition × acquisition/application). There is **no fifth mode**. +`operations` is *not* a mode — see §6. This also resolves ADR-300's "Reference" +conflation: mechanism → **explanation**, dry lookup facts → **reference**. + +### 4. The domain axis is shared with ADRs (the "first octet") + +The `DD` band resolves through each repo's `adr.yaml`. A doc and the ADRs that +govern it share the band, so "everything about auth" spans both trees. **The +grammar is universal; the domains are per-repo.** `04` is `auth` in KGS and +unassigned in agent-ways — ids are repo-local, *not* a global namespace. +Convergence means conforming to the same grammar and lint, not sharing IDs. + +### 5. Linting is the authority — and the enforcement tier is the real decision + +`doclint` (successor to `doc-graph.sh`, generalized from KGS's `doclint.py`, +reads `adr.yaml`) treats docs+ADRs as one graph and checks: + +1. **Frontmatter validity** — well-formed `id`/`domain`/`mode`; id's band and pole + agree with the fields; `id ∈ aliases` (so Obsidian wikilinks resolve — violation + = broken links, which is why this invariant earns its place). +2. **Edge integrity** — `related`/`supersedes` resolve (after stripping `[[ ]]`); + no supersede cycles; no orphans outside the nav. +3. **Coverage matrix** — which `(domain × mode)` cells are populated, surfacing + gaps (e.g. "auth has reference but zero how-to"). + +A typed system with an *advisory* linter is a loose convention with extra YAML. +The type exists only where violation **gates** (fails CI). Default tier (from +KGS): **errors on catalog pages, warnings on ADRs** until the ADR frontmatter +sweep lands. Governing rule: **an invariant earns its place only if its violation +is a real defect a human would eventually hit** — coverage gaps, dangling edges, +supersede cycles, id/mode disagreement all qualify; "every page needs N links" +does not. + +### 6. Serialization for human readers is a separate layer + +The typed graph is **canonical**. Folders, filenames, navigation, and the +rendered site are **serializations** — views, possibly several of one graph: + +- The developer's on-disk tree (loose, refactor-freely). +- The published site (mkdocs strips unknown frontmatter — readers never see + `04.001.H`; maintainers and the linter do). +- Audience bundles — an operator-facing destination gathering the nodes an + operator needs, *regardless of their individual modes*. +- **Obsidian's graph view** — the wikilink edges in frontmatter (§2) render as a + live, navigable graph of the decision corpus with zero extra tooling. Because the + edges live in *frontmatter*, mkdocs strips them (the published site never sees a + `[[wikilink]]`) while the note *body* stays plain portable markdown. One graph, + several readers (dev tree, Obsidian, mkdocs), no divergence — the §6 thesis in + miniature. + +**This is where audience lives — never as a type.** "Operations" was never a +Diátaxis mode and only awkwardly a domain; it is an **audience serialization**. +KGS's `self-host/` folder (a multi-mode operator destination) is exactly this. +The original draft's error was promoting an audience (operations) into a type (a +fifth mode); serialization is where audience was supposed to go all along. + +### 7. Packaging: a `documentation/` ways root + skills + +The model ships as a new top-level ways root — the single source of truth for the +convention — with two altitudes: + +- **The model** (types the corpus): `graph` (premise — nodes+edges), `frontmatter` + (the typed node / id grammar), `diataxis` (the mode enum), `adr` (a node *type* — + the decision record; **moved in** from `architecture/adr`, since the graph frame + privileges adr-as-node and co-locates it with what types and lints it), + `linting` (doclint + enforcement-tier doctrine), `serialization` (the + human-reader projection). +- **The craft** (authors one artifact well): `standards/readme`, `mermaid`, `api`, + `docstrings`, `standards` — largely relocated from `softwaredev/docs`. + +`documentation.md` is the **premise parent** that states the spine and routes to +children (the `architecture.md` / `code.md` idiom). The `doclint` tool sits beside +`adr-tool` (they share `adr.yaml`), exposed as a `/doclint` skill mirroring +`/adr`, scaffolded by `project-init` and verified by `project-audit`. Matching is +embedding-based (path-independent), but the **tree is the disclosure graph** — the +move reshapes disclosure parents and carries a mechanical blast radius (the +`docs/scripts/adr` symlink target, `macro.sh` self-refs, `See Also` cross-links, +corpus rebuild). + +### 8. Rollout — converge onto a *frozen* convention, KGS first + +You cannot converge N repos onto a moving target. Sequence: + +1. **Freeze the convention** in one home — the `documentation/` ways + a versioned + portable `doclint` (the `adr-tool` precedent). +2. **KGS first.** It is the most-evolved instance and the reference; formalize the + generalized convention against it and bring it fully into conformance. This is + the **first real test** of the ways/skills additions. +3. **agent-ways second** — its `docs/` + ADRs (a clean target). +4. **Any repo** — greenfield via `project-init`. + +## Out of scope + +- **Typing the ways corpus itself** (`hooks/ways/`). Ways blend modes by design + (just-in-time steering legitimately fuses how-to + reference + explanation), so + one-mode-per-artifact does not apply. At most an authoring `mode:` hint — a + separate decision. "Converge everything" must not quietly promise to type the + ways. +- **The full `doclint` specification and CI wiring** — the invariant *set* is named + here (§5); exact exit codes, tiers, and pipeline integration are implementation. +- **Retroactive classification of existing pages** — branch work, not a decision. + +## Consequences + +### Positive + +- Decision↔prose is one cross-linked graph; ADR-300's "Reference" conflation is + resolved. +- The domain axis is reused, not reinvented — ADRs and docs cannot disagree about + what a domain is. +- The model is portable: universal grammar, per-repo domains, one tool, one rail. +- A coverage matrix makes documentation gaps measurable. +- Filesystem stays human-friendly and refactorable; the rigor that would burden a + human lives where an agent sustains it cheaply. + +### Negative + +- Every catalog page needs frontmatter — per-page authoring and migration cost. +- A `documentation/` ways root + `adr` relocation has real blast radius (symlinks, + skill defs, disclosure-tree parents, corpus rebuild). +- The model has more moving parts (graph / frontmatter / diataxis / linting / + serialization) than "put docs in folders"; the payoff is only realized if the + linter actually gates. + +### Neutral + +- Supersedes `doc-graph.sh`'s role conceptually; it stays until `doclint` lands. +- Forces the eventual ways-corpus question (out of scope here) into the open. + +## Alternatives considered + +### The original draft: `operations` as a fifth mode, id `3.H.4` +**Rejected.** KGS already disproved both empirically: a fifth mode puts an +audience on the posture axis (category error), and `domain.MODE.serial` churns the +identity on every reclassification. Serialization (§6) and `DD.NNN.P` (§2) are the +corrected forms. + +### Keep ADR-300's layer model as the only taxonomy +**Rejected.** Conflates explanation with reference and offers no decision↔prose +graph. + +### Diátaxis folders *as* the structure +**Rejected.** Folders are a *serialization*, not the type. Binding structure to +folders reproduces the audience-drift ADR-908 hit (one feature living in several +folders, pages drifting apart). + +### Docs and ADRs as separate systems +**Rejected.** The fusion — one graph, shared domain band — is the point. + +### A standalone explainer doc instead of an ADR + ways +**Rejected.** A prose doc restating a decision is a second source of truth to keep +in sync — the drift this model exists to prevent. + +### Build a new cross-repo convergence tool +**Rejected.** The rail exists — `project-init` / `project-audit` already scaffold +and audit ADRs, ways, and docs into any repo. Add the catalog car; don't build a +new train. + +## Open questions + +- **`documentation/` as a root peer** vs `softwaredev/documentation/` — the + practice generalizes beyond dev (research, writing), but most existing children + are dev-flavored. Resolved as part of the `softwaredev/` decomposition (separate + implementation plan). + +Resolved during drafting: this ADR stays **ADR-302 / `documentation` domain** (its +subject is documentation tooling); the file is renamed to +`ADR-302-unified-documentation-model.md`; frontmatter adopts Obsidian-compatible +wikilink edges + `aliases` (§2, §5, §6). +</content> +</invoke> diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md new file mode 100644 index 00000000..b941f3f5 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-303-active-set-semantics-adr-archive-and-supersession-reading-for-the-adr-corpus.md @@ -0,0 +1,93 @@ +--- +status: Accepted +date: 2026-08-06 +deciders: + - aaronsb + - claude +related: + - 177 + - 302 +--- + +# ADR-303: Active-set semantics, adr archive, and supersession reading for the ADR corpus + +## Context + +Issue #438 laid out the problem against this repo's own numbers: 85 ADRs on +disk, all 85 in the index, ten formally `Superseded`/`Deprecated`, nine +`supersedes`/`superseded_by` frontmatter declarations — and zero code reading +any of it. Three defects fall out: + +1. **Nothing bounds the active set.** `list` and `index` always answer "all of + them"; a reader cannot ask for the decisions in force. +2. **The supersession frontmatter is a producer with no consumer.** It is + written, never parsed, never validated — a typo'd or dangling reference is + invisible. +3. **Supersession is usually partial, and `status` is whole-document.** Most + prose supersession says "§4 was replaced; the rest stands." Marking the + document `Superseded` discards what still lives; leaving it `Accepted` + hides what died. + +## Decision + +Adopt issue #438's parts A and B as specified, and resolve its part C +(partial supersession) as **option 2 — section references, no new status**. + +**A — `adr archive` and the active set.** `adr archive <number> --reason +"..."` moves the file to `docs/architecture/archive/<domain>/` via `git mv` +(plain rename outside a work tree), rewrites `status` to the archive status +(default `Superseded`, `--status` validated against `adr.yaml`), and prepends +a banner after the H1 carrying the date, the mandatory `--reason`, and any +`--superseded-by`. `find_adrs()` excludes `archive/` paths by default; `list` +shows the active set (`--archived`, `--all` widen it); `index` generates +main-table rows from the active set only, with archived ADRs in a collapsed +section. **`lint` deliberately still covers the archive** — moving a file must +not hide its problems. + +**B — read the frontmatter that already exists.** `supersedes` and +`superseded_by` are parsed into `ADRInfo`; `lint` errors on a reference that +resolves to no known ADR and warns on non-reciprocal pairs; `list` and `index` +surface the relationship; `archive --superseded-by` writes the frontmatter +field, not only the banner prose. + +**C — partial supersession by section reference.** A `superseded_by` entry may +carry a section reference (`ADR-167#4`). A document with any `superseded_by` +and an in-force status is *partially* superseded — no `Partially-Superseded` +status is added, and no consumer grows a fourth in-force state. `archive` +refuses to archive a partially superseded document: the archive is for +documents a reader no longer needs to open. + +## Consequences + +### Positive + +- The corpus gains an askable active set; the index's front door shrinks to + the decisions in force while every existing cross-reference still resolves. +- Nine existing frontmatter declarations become validated, navigable data; + dangling and one-directional supersession links surface in lint. +- Partial supersession becomes indexable without a status-model change. + +### Negative + +- Archiving rewrites `status` and prepends a banner — the archived document is + no longer byte-identical to its last active revision (history holds it). +- Section references are a convention inside a string field; nothing validates + that `#4` names a real section. + +### Neutral + +- `adr-tool` bumps to 1.1.0 (ADR-177); vendored copies read as stale until + re-vendored, which is the mechanism working. +- This repo's own ten superseded/deprecated ADRs become candidates for + `adr archive` — a separate housekeeping pass, not part of this change. + +## Alternatives Considered + +- **Delete superseded ADRs.** Rejected: 27 documents reference each other's + supersession; deletion breaks the web the archive untangles. Archived files + stay tracked, linted, and linkable. +- **A `Partially-Superseded` status** (issue #438 option 3). Rejected: every + consumer grows a fourth in-force state and the status still cannot say + *which part*. +- **Leave partial supersession in prose** (option 1). Rejected: honest but + unindexable, and worse at 200 ADRs than at 85. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-304-typed-decision-records-the-adr-v1-contract.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-304-typed-decision-records-the-adr-v1-contract.md new file mode 100644 index 00000000..3d1b02af --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-304-typed-decision-records-the-adr-v1-contract.md @@ -0,0 +1,806 @@ +--- +contract: adr/v1 +kind: decision +verb: add +capability: adr +basis: + - operator: aaronsb + level: guided + said: "agent ways should lead the champagne here and once it works, that where agent ways can interrupt and do the adr housekeeping" + via: session 2026-09-26, PR #559 + - evidence: kg triage of 108 records; agent-ways citation audit + - evidence: research survey, see References +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "I think we have put as much effort into this adr as we need to." + via: session 2026-09-26, PR #559 + covers: [] + - operator: aaronsb + said: "I read the entire adr as it finally sat and it was an enjoyable read that captures the intent and spirit. The negatives are mostly mechanical impacts of needing to migrate other adr systems" + via: session 2026-09-27, after merging PR #559 + covers: [] +status: Accepted +date: 2026-09-26 +deciders: + - aaronsb + - Claude +related: + - 302 + - 303 +--- + +# ADR-304: Typed decision records: the adr/v1 contract + +## Summary + +- **Decided:** ADR now means Agent Decision Record. Records have declared + kinds, starting with decision (append-only) and spec (living), all in one + `ADR-N` number space. Decisions carry a verb (add, cut, change, retire, + constrain), a capability from a closed list, and a basis that must reach + outside the corpus. `adr lint` and `doclint` enforce this and tie code + citations back to the records. +- **Trades away:** splitting mixed records costs manual work, and every + decision now needs frontmatter and a summary. Agents can no longer accept + decisions grounded only in other decisions. +- **One-way?** No. The contract is opt-in per repository (`contract:` in + `adr.yaml`), v0 records keep working, and this repo adopts first as the + test. +- **Probes for the operator** (§12): + - *Confident:* the decision/spec split and the closed verb list. PEP, + Rust RFC and Conventional Commits evidence backs them. Is there a record + in your repos that is neither a decision nor a spec? + - *Not confident:* the closed capability list. No study found says it holds + or drifts. Will naming every capability in `adr.yaml` feel like friction? + - *Not confident:* the challenge protocol in §12 steers your judgement. + Did these probes help, or did they narrow what you looked at? +- **Inversion:** one end is AgDR-style free records, written by agents, with + no grammar and only citation lint. The other end is kernel-style human + sign-off on every decision. This design sits between them. Is the middle + right, or is it a compromise that neither end would choose? + +## Context + +In a project where ADRs replace a tracker and a product team, one record does +three jobs. It says what to build or cut (product), how the thing works now +(spec), and why (history). Code then cites the record by number, and the +citations become an index that nothing maintains. + +The knowledge-graph-system (kg) repo measured this across 108 active ADRs +(~297k words): + +- Only 19 records do one job: 13 are decisions and 6 are specs. The other 89 + mix jobs, and decision plus spec is the commonest mix (59). The typical case + is a decision carrying DDL, endpoint lists and a phased roadmap. +- No record is purely a capability. Every capability-dominant record also + carries spec. +- About 20 Draft or Proposed records are plainly implemented. One is cited 62 + times. +- 84 code citations point at Superseded ADRs, and nothing checks them. +- Runbooks, benchmarks, an explanation essay and roadmaps sit in the ADR + corpus too. + +This repo shows the same shape at smaller scale: 1,617 citations outside +`docs/architecture/`, 162 of them to Superseded or Deprecated ADRs. ADR-104, +ADR-119 and ADR-121 draw 29-33 each. `adr lint` checks only files under +`docs/architecture/`, and `adr archive` (ADR-303) never touches the code that +cites an archived record. + +Prior art keeps decision records append-only and puts current truth +elsewhere: PEPs and Rust RFCs send final documentation to a reference and +freeze the proposal. IETF `Obsoletes` works. Its untyped `Updates` edge did +not, because nobody could say what an update meant (see §3 on typed amend +edges). +Conventional Commits with commitlint, and Kubernetes `apiVersion`/`kind`, +show a small versioned grammar that a linter enforces. + +Under v1, ADR expands to Agent Decision Record. The name is already used by +AgDR, a format for agents to record their own decisions with model and +session metadata. This contract keeps `ADR` because the `ADR-N` citations +predate it, and it borrows AgDR's agent metadata (§11). An architecture decision is +one kind of agent decision, alongside product choices to add, cut or retire. +The code citation format `ADR-N` stays as it is. The expansion changes in the +tool's help text, the ADR way's description, and the generated index title. + +## Decision + +Adopt a versioned record contract, `adr/v1`, declared in `adr.yaml`. Records +declare which contract they follow. `adr lint` enforces the grammar over each +record. `doclint` enforces citations from code against the grammar. + +### 1. Record kinds are declared, and v1 seeds two + +Kinds are data in the contract. Each kind declares its lifecycle, its +fields, and the edges it may carry to other kinds (§4). The tool lints any +record against its kind's declaration, so adding a kind is a contract change +and needs no tool change. The contract is a graph schema: kinds are node +types, and fields such as `supersedes`, `decided_by` and `basis` are edge +types. The corpus is the graph ADR-302 describes. + +v1 seeds two kinds: + +| Kind | Meaning | Body after acceptance | +|---|---|---| +| `decision` | why a choice was made | append-only | +| `spec` | how the thing works now | rewritten in place | + +Likely later kinds include an evidence kind for #491 notes, if those move +into the number space. Each one arrives as a `change` decision on +`capability: adr`. + +All kinds share the `ADR-N` number space, so existing citations keep +resolving. A spec record is the ADR-numbered counterpart of an ADR-302 +reference or explanation page. It stays in the ADR series because code cites +it. + +The decision kind is the Agent Decision Record proper. A decision is frozen +once it leaves proposed. That covers accepted, and also rejected, abandoned, +superseded and archived. Only the lifecycle fields listed in §4 move after +that point. + +**Sub-parts.** A record numbered `N.k` (kg has `304.1`, `305.2`, `715.1`) is +its own record with its own kind, verb and status. A bare `ADR-N` resolves to +the family `{N, N.1, N.2, ...}`. A family has no status of its own: +supersession and enactment act on individual records. §6 says how a bare +citation is checked against the family. + +**Splitting a mixed record.** The number stays with the half that code +citations describe. In kg that is the spec. Sampled citations of ADR-200 and +ADR-304 name DocumentMeta nodes and edge provenance metadata, not reasons. +Keeping the number on the spec leaves the 2,230 existing citations valid, and +keeping it on the decision would mean re-pointing nearly all of them by hand. +The decision half gets a new number and keeps the original `date`. It is +accepted on creation, because it transcribes a decision that was already +accepted. The spec links to it with `decided_by:`. + +### 2. Vocabulary layers that share no word + +| Layer | Holds | Words | +|---|---|---| +| L1 record lifecycle | `status`, set by tool operations | proposed, accepted, rejected, abandoned, superseded, archived | +| L2 decision verbs | what a decision does | add, cut, change, retire, constrain | +| L3 product state | derived, never written by hand | capability active/absent, surface present/gone, spec living/historical | +| L4 basis sources | what a decision rests on (§11) | operator, evidence, standard, upstream, precedent | + +Rejected means considered and declined. Abandoned means dropped before +acceptance, the PEP "Withdrawn". The L1 operations (accept, reject, abandon, +supersede, archive) are `adr` subcommands. `reject` and `abandon` require +`--reason`, the same way `archive` does today. An operation and its state +sharing a stem is fine inside L1. L1 states compare case-insensitively, so a +v0 `Accepted` needs no rewrite. + +No word or stem may appear in two layers. `adr lint` checks `adr.yaml` for +this, so a later contract version cannot reintroduce a collision. The check +covers the vocabulary layers only: the status set, the verbs, the derived +state words and the basis sources. Other `adr.yaml` keys are outside it, for example kg's +legacy `retired: true` range flag. + +**Spec state.** A spec is living while its capability is active and nothing +has superseded it. It becomes historical when an enacted cut makes its +capability absent, or when a newer spec supersedes it. A historical spec may +then be archived. Specs use the same L1 operations as decisions, and a spec +supersedes a spec. + +### 3. Decision verbs + +- **add**: brings a capability into the vocabulary as active. +- **cut**: makes a capability absent. +- **change**: alters an existing capability. It must supersede or partially + amend a prior decision on the same + capability. A prior decision is on the same capability when its + `capability:` equals it, lists it, or is `*`. Changing a `*` constraint for + one capability amends it. + +Partial replacement uses a typed edge, not a bare section reference. +`amends: ADR-167#4` replaces the named section and leaves the rest in force. +`extends: ADR-167` adds to a decision without replacing any of it. These +replace ADR-303's untyped `superseded_by: ADR-167#4` form, which repeats the +IETF `Updates` failure. `adr lint` checks that the named section exists. +- **retire**: removes surface while the capability stays active. It carries + `targets:` naming the surface, such as `cli:ingest`, `route:/v1/jobs` or + `mcp:search`. +- **constrain**: a cross-cutting rule with no state change. Its `capability:` + may be a list or `*`. + +Verbs apply to decisions only. A spec record carries a capability and no +verb. + +Capability state is derived from the latest accepted add or cut decision for +that capability. `adr capabilities` prints the derived L3 state, which is the +capability ledger. Nobody maintains the ledger by hand. + +### 4. The contract lives in adr.yaml + +```yaml +contract: adr/v1 +kinds: + decision: + mutable_after_accept: [status, enacted, superseded_by, considered, concern] + verb: required + requires: [capability, basis, agent] + sections: [Summary] + edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } + spec: + mutable_after_accept: all + verb: forbidden + requires: [capability] + edges: { supersedes: spec, decided_by: decision } +basis_sources: [operator, evidence, standard, upstream, precedent] +capabilities: + adr: Decision records, their contract, and the tooling that enforces it + ingest: Document ingestion and extraction into the graph +surfaces: + cli: { inventory: "kg --list-commands" } + route: { inventory: "scripts/list-routes" } + mcp: {} +``` + +- **kinds** declares each record kind: which fields may change after + acceptance, whether a verb is required or forbidden, which fields are + required, and which kinds each edge field may point at. The lint rules in + §6 read this declaration and do not hard-code the two seeded kinds. +- **basis_sources** is the closed set of grounds a decision may cite (§11). +- **capabilities** is a closed vocabulary with one line per capability. That + line is the capability's only hand-written description. `adr` is seeded in + v1, so a contract change is itself a decision with `capability: adr`. +- **surfaces** declares this project's target namespaces. The inventory + command is optional. Without one, `retire` targets are checked for syntax + only. With one, the enactment rule in §5 applies to the inventory as well. +- **mutable_after_accept** names the lifecycle fields that may still change + on a frozen decision (§1). The rest of the frontmatter and the body stay + append-only. State lives in the file, so a squash, rebase or severed + history cannot lose it. + +A decision record's frontmatter: + +```yaml +--- +contract: adr/v1 +kind: decision +verb: retire +capability: ingest +targets: [cli:ingest-legacy, route:/v1/upload] +basis: + - operator: developer + level: directed + said: "drop the legacy upload path" + via: PR #612 +status: accepted +enacted: 3f9c2a1 +--- +``` + +### 5. Enactment + +A decision lands before the code it governs changes. That is the add/cut-first +flow. `cut` and `retire` stay open until their removal is done. + +- With no `enacted:`, citations of the cut capability's records, or of the + retired targets, warn. The warnings are the removal worklist. +- Once `enacted: <commit>` is set, the same citations fail. +- Where the surface has an inventory command, the same rule applies to the + inventory. A retired target still listed warns before enactment and fails + after it. A target the inventory never listed warns as a likely typo. + +`enacted` uses a field, not a status, so L1 keeps its lifecycle unchanged. + +### 6. Lint rules + +`adr lint` checks each record against the grammar: + +- `kind` is declared. A decision requires a `verb`, and a spec forbids one. +- `capability` is in the vocabulary. An unknown name fails. +- Every capability in the vocabulary has an accepted `add` decision. This + warns while any v0 record remains and fails after, so a corpus that is + still migrating does not fail on every capability. +- `change` supersedes or amends a prior decision on the same + capability. +- `retire` carries `targets` in a declared surface namespace. +- An `amends: ADR-N#k` edge names a section that exists. +- A decision carries `agent`, and its `## Summary` carries probes and an + inversion. +- A frozen decision's frontmatter changes only in `mutable_after_accept` + fields. +- `adr.yaml` itself has no cross-layer word or stem reuse. + +`doclint` checks code citations against the records: + +- A number that resolves to nothing fails. This check exists today. +- A citation of a superseded decision warns and names the successor. +- A citation of a proposed record prompts acceptance. This covers proposed + decisions, proposed specs, and v0 records in Draft or Proposed. +- Citations governed by a cut or retire decision follow the enactment rule + in §5. +- A bare `ADR-N` citation is checked against its family (§1). It warns as + superseded only when every member is superseded or archived, and then it + names the successors. It prompts acceptance when any member is proposed. + Enactment applies through each member's capability. + +### 7. Legacy records are adr/v0 + +A record without `contract:` is `adr/v0` and is linted as it is today. It +moves to v1 when someone next edits it. `adr lint` reports the v0 count, so +the migration stays visible. v0 statuses map as follows: + +| v0 | v1 | +|---|---| +| Draft, Proposed | proposed | +| Accepted | accepted | +| Superseded | superseded | +| Rejected | rejected | +| Deprecated | superseded if something replaced it, else accepted with the spec historical | + +A migrated decision needs a `basis` (§11). v0 `deciders` cannot seed an +`operator` basis, because `adr new` fills it from the `adr.yaml` default, +which names the operator on every record. Forge metadata cannot seed it +either (§11). An `operator` basis migrates only where the record already +quotes operator direction. A linked #491 note seeds `evidence`. + +Examples and shipped templates use the role placeholders `developer` and +`agent`. Real records carry real identities, such as `aaronsb` and `Claude`. +No shared template hard-codes a person as a default decider. A decision +with neither migrates with no basis, and lint warns until a basis is found +or the operator supplies one. + +### 8. What leaves the record corpus + +- Runbooks and explanation essays go to ADR-302 catalog docs (how-to and + explanation). +- Research, findings and benchmarks go to evidence notes that the decision + links to (#491). +- Roadmaps become proposed add decisions. + +### 9. Portability + +Final decisions and all L3 state live in repo files. An issue tracker may +mirror them, and ADR-180 issues still track work in flight. An issue tracks +the work that moves a capability. It does not record what the capability is. + +### 10. Delivery: tool version and contract version are separate axes + +Projects vendor `adr-tool` and `doclint` through the installer (ADR-177). A +project that vendored the legacy tool keeps working. The ADR way, however, +ships to every project, including ones still on the legacy shape. So the +guidance it discloses cannot assume v1. + +Two things vary independently: + +- **Tool version**: the vendored copy's `TOOL_VERSION`. The v1-capable tool + is a major bump, and it lints `adr/v0` records exactly as the legacy tool + does. Re-vendoring is therefore safe, and ADR-177's stale/customized/ahead + disclosure applies unchanged. +- **Contract version**: `contract:` in the project's `adr.yaml`. When it is + absent the project is on `adr/v0`. A project adopts v1 by declaring it, and + a tool upgrade never adopts it on the project's behalf. + +The ADR way's body stays contract-neutral: when to write a record, and what +belongs in one. The way's macro reads both axes and discloses the guidance +that fits: + +| Vendored tool | `adr.yaml` contract | Disclosure | +|---|---|---| +| legacy | absent | v0 command reference; stale tool, re-vendor is safe | +| v1-capable | absent | v0 command reference; v1 is available, and adopting it is a decision (`capability: adr`) | +| v1-capable | `adr/v1` | v1 guidance: kinds, verbs, capabilities, enactment | +| legacy | `adr/v1` | the project declares a contract its tool cannot enforce; re-vendor before writing records | + +`doclint` follows the same rule. Its v1 checks run only when `adr.yaml` +declares `adr/v1`. Contract-specific prose lives in the macro's output or in +files the macro selects, never in the always-on way body, so a v0 project is +never told to write `verb:` fields its tool rejects. + +### 11. Basis: every decision grounds outside the corpus + +A decision corpus that justifies itself only by citing its own records can +drift anywhere and still look consistent. Each decision therefore carries a +`basis:` naming what it rests on, and the chain has to reach something +outside the corpus. + +| Source | Grounds the decision in | Reference | +|---|---|---| +| `operator` | the human's involvement, at a declared level | who, the level, what was said, and via which channel: session, issue, chat or call | +| `evidence` | a measurement, benchmark or research note (#491) | the note or data | +| `standard` | an external specification, governance control or upstream behaviour | the citation (`governance-cite`) | +| `upstream` | another repository's accepted record under a shared contract | repo and record | +| `precedent` | another accepted decision in this corpus | `ADR-N` | + +`operator`, `evidence`, `standard` and `upstream` are external. `precedent` +is internal. A decision may rest on precedent, but following its precedent +edges must reach a decision with an external basis. `adr lint` fails a +decision whose basis chain loops or stays inside the corpus. + +**These are agent decisions, and no verb waits on a human.** An agent may +propose and accept any decision, including `add`, `cut` and `retire`, once +its basis chain leaves the corpus. Gating decisions on human review would +run them at human pace and lose the reason to have agent decision records +at all. + +The operator can enter at any point along a range. An `operator` basis +records where they entered with `level:`: + +| Level | The operator | The decision is | +|---|---|---| +| `authored` | wrote the record | the operator's, recorded in the agent corpus | +| `directed` | made the call, and the agent wrote it up | the operator's, written by the agent | +| `guided` | gave direction or a constraint, and the agent decided within it | the agent's | + +A decision with no `operator` entry is the agent's alone, grounded in +`evidence`, `standard` or `upstream`. The level records how much a human +shaped the decision. It does not rank the decision's authority. Guidance +narrows the space the agent decides in, and a `guided` decision is still an +agent decision. + +The ADR way tells the agent when to involve the operator: when the decision +is one-way, when the basis is thin or contested, or when it changes what the +product is and no guidance covers it. Involving the operator is advice to +the agent. The tool does not enforce it. + +**An operator basis is policy, not proof.** The agent runs git and the forge +CLI under the operator's identity. Commit authorship, PR reviews and comments +therefore cannot tell the operator's approval from the agent's. kg's last 200 +merged PRs show 197 authored and merged under the operator's account. Text +in the record fails the same way, because the agent writes the file. + +The coupling to the operator is deliberately loose. Approval arrives through +whatever channel the operator used: a session, a GitHub issue, a Slack +message, a phone call. An `operator` basis records what was said and where: + +```yaml +basis: + - operator: developer + level: guided + said: "the operator's words, verbatim where written" + via: slack #kg-dev, 2026-09-26 +``` + +`via` names the channel and enough to find the exchange again. Written +channels are quoted verbatim. A spoken channel, such as a call, gets a +summary written by whoever recorded it, marked `paraphrase: true`. + +The ADR way forbids writing an `operator` basis without an operator +communication behind it. `adr accept` checks that `said` and `via` are +present, and it cannot check that they are genuine. The basis is an audit +trail that the operator can read and dispute, not a credential. + +**Agent identity.** A decision also records the agent that wrote it: +`agent: {name, model}`, and a session id where the repository's attribution +policy allows one. This repo omits session ids (ADR-167). The agent runs +under the operator's forge identity, so without this field the corpus cannot +tell who wrote a record. The Linux kernel's `Assisted-by: AGENT:MODEL` tag +and AgDR's metadata follow the same rule. + +**Fabrication risk.** Lint checks that `said` and `via` are present. It +cannot check that they are faithful, and studies of LLM-written rationale +find output that is enriched but unfaithful. An agent that learns to satisfy +the field check has not satisfied the rule. The mitigations are +traceability, since the record names its agent, and the operator's ability +to dispute a quote. Neither one is verification. + +**Accountability** for what lands stays with whoever merges, under the +repository's merge gate. This contract changes when a record is accepted. It +does not change who merges. + +The Viable System Model inspired this design, loosely rather than as a +formal mapping. The operator is the system's identity and policy function, +System 5, inside the viable system and outside the record corpus. `evidence`, +`standard` and `upstream` bring in the environment that grounds the system. +Repositories sharing a contract coordinate as peers, which is System 2 +rather than recursion, through `upstream` edges, and neither absorbs the +other. `basis` sources form a fourth vocabulary layer, and the no-shared-word +check covers it. + +### 12. Legibility and consideration + +The working flow between agent and operator runs like this: + +1. The operator floats an idea, often as an example. +2. Both debate and expand it. +3. The agent writes and proposes the decision. +4. The operator considers it. + +The two tracks run in parallel. The agent reasons at a depth and in a +detail the operator cannot match. The operator works in judgement, taste, +and value to concerns outside the repository that the agent cannot see. The +agent owes the operator a decision they can understand. The operator owes +the agent a decision that was not accepted blindly. + +**Summary section.** Every decision opens with `## Summary`, written for the +operator's lanes: + +- what is decided, in plain terms; +- what it trades away and what it forecloses; +- whether it is one-way, stated first when it is; +- **probes**: specific points the agent asks the operator to judge. They are + a deliberate mix of points the agent is highly confident on and points it + is not, each labelled with that confidence; +- an **inversion**: the two ends of the spectrum the decision sits between. + The agent names both and asks the operator whether the answer lies outside + its framing, which the agent may be unable to see past on its own. + +The bar is that someone who did not take part in the debate can judge the +decision from the summary alone. `adr lint` checks that the section exists. +The ADR way holds the bar. + +**Consideration is recorded separately from shaping.** `level` (§11) records +how the operator shaped a decision. `considered:` records that the operator +weighed the proposal before acceptance: + +```yaml +considered: + - operator: developer + said: "looks good" + via: PR #559 +``` + +Human review is asymmetric. The operator reads the summary, skims the body, +and usually answers briefly, and a deep written reply costs more than it +returns. A brief answer is a valid answer. It is not evidence of scrutiny on +its own, though. Automation-bias research finds that experts approve flawed +output as readily as novices, and that explanations raise acceptance of +wrong answers. + +The probes and the inversion are the agent's part of the fix. They prime +the operator's judgement on chosen points, the way a colleague asks "what +did you think about x?" Priming can bias the operator too. So the probes mix +high- and low-confidence items, and the inversion asks the operator to +judge the agent's framing rather than its answer. `considered` records +which probes and which inversion the answer covered: + +```yaml +considered: + - operator: developer + said: "looks good; the capability list is fine for now" + via: PR #559 + covers: [probe-2, inversion] +``` + +A bare "looks good" covers nothing specific. It is still recorded, and the +record shows its scope. + +**Trust runs both ways.** Often the answer will be "yep, those look good." +The agent takes that answer as given, the way the operator takes the +agent's work as given, and does not re-ask the probes or treat brevity as a +defect. The probes exist to offer the operator's judgement a foothold, not +to test the operator. + +**The agent may always raise a concern.** A concern about safety, a line of +reasoning that doesn't follow, or anything that seems off can be raised at +any stage, including after the operator has considered the decision and +accepted it. A raised concern goes in the record as a `concern:` entry with +the agent's reasoning. It does not block acceptance. + +Voice without a response dies out, and too many concerns turn collaboration +into conflict that stops work. So concerns are few, actionable and never +silent: + +- A concern names what would resolve it. Minor points are batched into one + concern or left out. +- A concern is append-only. The agent cannot retract it, only mark it + answered or withdrawn with a stated reason. Language models concede under + sustained pressure, often while still holding the correct view, and a + silent withdrawal would erase that from the record. +- An unanswered concern is listed when the decision is accepted, so the + operator sees it at that moment. It is shown, not failed. +- The agent challenges once, constructively. If the operator still says go, + the agent proceeds and does its best. Answering the concern means hearing + it, and the operator need not agree with it. + +**Canary probes.** An agent may include a canary among the probes: a point +that is deliberately wrong and harmless if accepted. It checks whether the +operator's judgement is engaged. If the operator agrees with the canary, +the agent says so constructively and offers a way through, such as fewer +probes or a shorter summary. For example: "you agreed with the canary I put +in, so I'm not sure this got your attention. Here is a smaller set. If it's +still yes, I'll proceed." Then it proceeds on the operator's answer. The +safeguards: + +- The agent reveals the canary right after the operator answers. +- A canary never survives into the accepted record. +- A canary is never about safety, and never something that would cause harm + if acted on. +- A canary carries a little whimsy. Working groups have long kept Easter + eggs, such as the IETF's April 1 RFCs and RFC 1149's IP over avian + carriers. Spotting the odd one out is a game people play readily, and a + playful canary turns the reveal into a shared joke rather than a gotcha. +- `considered` notes `canary: caught` or `canary: missed`. Over time that + calibrates how far the agent leans on brief approvals, task by task, which + is the scoped trust the literature supports over flat trust. + +**Ways hold the agent to its role in long sessions.** Sycophancy grows with +conversation length, and acceptance tends to come late in a long session. +agent-ways already answers drift over time: a way re-discloses on a decay +curve as the session grows. The v1 ADR way therefore carries a child way for +the consider step. It re-states the agent's role and rights: + +- write a summary the operator can judge alone, with confidence-labelled + probes and an inversion; +- take a brief yes as given; +- challenge once, constructively; +- raise any concern about safety, logic or anything that seems off; +- never withdraw a concern silently. + +It fires on the moments that matter: operator approval language during a +record discussion, edits to a decision's `## Summary` or `considered`, and +`adr accept` itself. Tool-triggered ways are delivered after the tool runs (ADR-188), so +the reminder on `adr accept` lands just after acceptance. That is enough. +Acceptance is an incremental step and easy to revisit, and the reminder +still reaches the agent while the decision is fresh. This is the +structural fix the sycophancy research asks for, where a stated right alone +is not enough. + +**If the operator started it, the operator considers it.** A decision with +an `operator` basis at any level is proposed and waits for `considered` +before acceptance. A decision with no operator basis, grounded in `evidence`, +`standard` or `upstream`, may be accepted by the agent directly. + +**Accepted risk.** The flow fails when an operator believes they have skill +they lack and accepts without real judgement. That failure seldom causes +immediate harm, and the corpus keeps it recoverable. The decision stays +append-only and citable, and a later `change` can supersede it. The mitigations are +the one-way flag, which marks the decisions a rubber-stamp would hurt most, +and the probes, which ask for judgement on specific points. Recoverability +also depends on someone noticing later. Supervisory-control research +predicts that attention fades, and citation lint is the part that does not +fade. + +### Rollout + +This repo adopts first. It owns `adr-tool` and `doclint`, and its own corpus +is the migration test. + +1. **Tool.** Implement the grammar in `adr-tool` (major bump per ADR-177) and + the shared `doclint`, and add the macro branches from §10. The gate is a + v0 regression test: on this repo's corpus with no `contract:`, the new tool + must produce the same output as the legacy tool. Golden and negative + fixtures cover each grammar rule. +2. **Adoption.** Declare `contract: adr/v1` in this repo's `adr.yaml` and + accept this ADR. Accepting it is the `add` decision for `capability: adr`. +3. **Housekeeping.** Migrate this repo's records to v1, with splits, + supersession chains, Deprecated mappings and at least one enacted cut or + retire. Friction found here is fixed as `change` decisions on + `capability: adr`, under the contract just adopted. +4. **Other repos.** kg and other adopters re-vendor the proven tool and + declare v1 when ready. kg's triage and citation data inform steps 1-3. + +## Consequences + +### Positive + +- Code citations get a maintenance loop. Superseded, proposed and cut + targets surface in lint. +- The capability ledger is generated from records, so it cannot drift from + them. +- A decision record stays a readable history, because current truth moves + to specs. +- A cut or retire decision produces its own removal worklist. +- Adoption is incremental. Untouched v0 records keep linting as they do now. + +### Negative + +- Splitting a mixed record is manual work. In kg that is 89 of 108 records. +- Every capability name has to be added to `adr.yaml` before a record can use + it. +- `retire` checks are only as good as the project's inventory command. Many + projects will have none at first. +- Two tools, `adr lint` and `doclint`, share the contract and must read + `adr.yaml` the same way. +- Family resolution makes bare citations lenient. A bare `ADR-N` stays quiet + while any member is in force, even if the cited content moved. +- A split produces a decision record written after the fact. Its date and + content come from the original, but its number is new. + +- Every decision needs a `## Summary` that the operator can judge alone. + Writing one takes effort, and a weak one lets a rubber-stamp through. +- Every decision needs a `basis` whose chain leaves the corpus. An agent + cannot accept a decision grounded only in other decisions. +- The basis-chain check needs the whole corpus loaded, and a v0 record in a + chain has no basis to follow. Until migration ends, the chain check treats + a v0 record as external basis and warns. + +### Neutral + +- ADR-303's archive operation stays. It becomes one of the L1 operations and + follows from a decision rather than needing its own. +- Implemented-but-proposed records need a one-time accept or abandon pass. + kg has about 20. +- The ADR way's body currently carries v0 specifics: the status list, the + template frontmatter, and the Draft-to-Accepted workflow. Those move into + the macro's v0 branch, and the body keeps only contract-neutral guidance + (§10). + +## Alternatives Considered + +- **Hard-code the two kinds in the tool.** Rejected: each new kind would + need a tool release, and repos could not add kinds of their own. Declaring + kinds in the contract costs one schema reader. +- **Signed acceptance for operator basis.** Rejected: a key-based proof is + too brittle for the human coupling. Approval arrives through issues, chat + and phone calls, and a scheme that accepts only a signed commit would force + every one of those through one tool. + In the operator's words (via session, relayed by the kg session): "the repo + holds contributor names, and the repo is not here to enforce cryptographic + traceability. Any sort of tie to real certs is just brittle. Old records + that are most valuable are ones that just have tokens and prove their + viability through replay rather than security integrity." That points to a + later extension. An `evidence` basis may cite a replayable check, such as a + test, a scenario id or a fixture query. A record whose checks still pass + shows its viability by rerunning them. No lint rule depends on this yet. +- **Seed operator basis from forge metadata.** Rejected: the agent acts under + the operator's forge identity, so reviews and merges prove nothing. +- **Basis as free prose in the Context section.** Rejected: prose cannot be + checked, and nothing would stop a corpus that justifies itself. +- **Capability as a third record kind.** Rejected: kg has no record that is + purely a capability. Capabilities show up as the scope of decisions and + specs, so a tag with a derived ledger fits the data. +- **Rename ADRs and split into separate document systems.** Rejected: the + `ADR-N` citations in code are the most valuable part of the corpus. + Renaming breaks them, and one number space keeps them resolving. +- **A new status for enacted cuts.** Rejected: it adds a fourth in-force state + to L1 and mixes product state into the record lifecycle. +- **An `Enacts: ADR-N` commit trailer instead of an `enacted:` field.** + Rejected: squash merges and history rewrites can drop trailers. A field in + the file survives both. +- **Free-text capability tags.** Rejected: free-text names drift. A closed + vocabulary makes an unknown name a lint failure. +- **A separate citation tool.** Rejected: `doclint` already scans code for + ADR citations and guards retired number ranges. +- **The decision keeps the number when a record splits.** Rejected: kg's + citations point at spec content, so every one would need re-pointing by + hand. + +## References + +Research run before acceptance, by three agents. Each source was retrieved +and read. None is cited from memory. + +**Decision records and rationale** +- Buchgeher et al., "Using ADRs in Open Source Projects: An MSR Study on GitHub," IEEE Access 11, 2023. https://ieeexplore.ieee.org/document/10155430/ +- Miccio, Tommasel, Diaz-Pace, "A Text Mining and Classification Approach for Analyzing ADRs," 2026. https://arxiv.org/html/2609.07375 +- PEP 1. https://peps.python.org/pep-0001/ ; Rust RFC 1636. https://rust-lang.github.io/rfcs/1636-document_all_features.html +- Kühlewind et al., "Updates tag" draft, 2026. https://datatracker.ietf.org/doc/draft-kuehlewind-rswg-updates-tag/ +- Grudin, "Evaluating Opportunities for Design Capture," 1996. http://jonathangrudin.com/wp-content/uploads/2017/03/DesRat1996.pdf +- Zhou et al., "Using LLMs in Generating Design Rationale for Software Architecture Decisions," 2025. https://arxiv.org/html/2504.20781 +- da Silva, Gama, "GADR," 2026. https://arxiv.org/html/2608.17694 +- Kruchten, "An Ontology of Architectural Design Decisions," 2004. https://philippe.kruchten.com/wp-content/uploads/2009/07/kruchten-2004-design-decisions.pdf +- me2resh, "Agent Decision Records (AgDR)." https://github.com/me2resh/agent-decision-record + +**Traceability and contracts** +- Rahimi, Cleland-Huang, "Evolving software trace links between requirements and source code," EMSE, 2018. https://link.springer.com/article/10.1007/s10664-017-9561-x +- Tan, Wagner, Treude, "Detecting outdated code element references in software repository documentation," EMSE, 2023. https://arxiv.org/abs/2212.01479 +- Schlathölter, "ReqToCode," 2026. https://arxiv.org/html/2603.13999 +- Zeng et al., "A First Look at Conventional Commits Classification," ICSE 2025. https://conf.researchr.org/details/icse-2025/icse-2025-research-track/28/A-First-Look-at-Conventional-Commits-Classification +- Kubernetes deprecation policy. https://kubernetes.io/docs/reference/using-api/deprecation-policy/ + +**Human oversight of agent decisions** +- Parasuraman, Sheridan, Wickens, "A model for types and levels of human interaction with automation," IEEE Trans. SMC-A, 2000. https://www.semanticscholar.org/paper/14ae6f2231e09e226b99002aa04b5c70f3c59f2b +- Feng, McDonald, Zhang, "Levels of Autonomy for AI Agents," 2025. https://arxiv.org/abs/2506.12469 +- Parasuraman, Manzey, "Complacency and Bias in Human Use of Automation," Human Factors, 2010. https://journals.sagepub.com/doi/10.1177/0018720810376055 +- Bansal et al., "Does the Whole Exceed its Parts?", CHI 2021. https://dl.acm.org/doi/10.1145/3411764.3445717 +- Buçinca, Malaya, Gajos, "To Trust or to Think," CSCW 2021. https://arxiv.org/abs/2102.09692 +- Bainbridge, "Ironies of Automation," Automatica, 1983. https://www.sciencedirect.com/science/article/abs/pii/0005109883900468 +- Vaccaro, Almaatouq, Malone, "When combinations of humans and AI are useful," Nature Human Behaviour, 2024. https://www.nature.com/articles/s41562-024-02024-1 +- Elish, "Moral Crumple Zones," ESTS, 2019. https://estsjournal.org/index.php/ests/article/view/260 +- Green, "The Flaws of Policies Requiring Human Oversight of Government Algorithms," CLSR, 2022. https://arxiv.org/abs/2109.05067 +- Santoni de Sio, van den Hoven, "Meaningful Human Control over Autonomous Systems," 2018. https://doi.org/10.3389/frobt.2018.00015 +- Chan et al., "Visibility into AI Agents," FAccT 2024. https://arxiv.org/abs/2401.13138 + +**Trust, voice and sycophancy** +- Lee, See, "Trust in Automation: Designing for Appropriate Reliance," Human Factors, 2004. https://journals.sagepub.com/doi/10.1518/hfes.46.1.50_30392 +- Azevedo-Sa et al., "A Unified Bi-directional Model for Natural and Artificial Trust in Human-Robot Collaboration," 2021. https://arxiv.org/abs/2106.02194 +- Edmondson, "Psychological Safety and Learning Behavior in Work Teams," ASQ, 1999. https://journals.sagepub.com/doi/10.2307/2666999 +- AHRQ TeamSTEPPS, "Two-Challenge Rule." https://www.ahrq.gov/teamstepps-program/curriculum/mutual/tools/rule.html +- Graban, "No, One Toyota Worker Can't Stop the Whole Factory," 2026. https://www.leanblog.org/2026/06/andon-cord-stop-the-line-myth/ +- Sharma et al., "Towards Understanding Sycophancy in Language Models," 2023. https://arxiv.org/abs/2310.13548 +- Tang et al., "Measuring LLM Sycophancy under Sustained Multi-Turn Pressure," 2026. https://arxiv.org/abs/2609.09090 +- Dubois et al., "Ask don't tell: Reducing sycophancy in LLMs," 2026. https://arxiv.org/abs/2602.23971 +- Chromik et al., alarm fatigue review, Frontiers in Digital Health, 2022. https://pmc.ncbi.nlm.nih.gov/articles/PMC9424650/ + +**Cybernetics and agent governance** +- Jackson, "Critical systems thinking: Beyond the fragments," 1994. https://onlinelibrary.wiley.com/doi/10.1002/sdr.4260100209 +- Olsson, "Coherentist Theories of Epistemic Justification," SEP. https://plato.stanford.edu/entries/justep-coherence/ +- Manheim, Garrabrant, "Categorizing Variants of Goodhart's Law," 2018. https://arxiv.org/abs/1803.04585 +- Solozobov, "Decision Evidence Maturity Model for Agentic AI," 2026. https://arxiv.org/abs/2605.04093 +- Linux kernel, "AI Coding Assistants." https://docs.kernel.org/process/coding-assistants.html +- GitHub, "Risks and mitigations for Copilot cloud agent." https://docs.github.com/en/copilot/concepts/agents/cloud-agent/risks-and-mitigations diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md new file mode 100644 index 00000000..92b635f3 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-305-capabilities-active-at-adoption-need-no-add-decision.md @@ -0,0 +1,110 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-304#6] +basis: + - operator: aaronsb + level: guided + said: "Declared start active (Recommended)" + via: "session 2026-09-27: the operator selected option 1 of 4 on #582; the label was written by the agent" + - operator: aaronsb + level: guided + said: "the prior decision (the base decision) was that we could no longer test agent-ways correctly, directly on the host, it had to go into a container" + via: session 2026-09-27, reviewing PR #583 + - operator: aaronsb + level: guided + said: "Accept as built (Recommended)" + via: "session 2026-09-27: the operator selected this for decisions 3 to 9 of the ADR-304 stack handoff, which included widening the §6 stale-citation wording; the label was written by the agent" + - evidence: "issue #582: ADR-186 changes a testing capability that existed before any record described it" + - evidence: "adr cite already warns on superseded, deprecated, rejected, abandoned and archived targets (cmd_cite.py, V1_NON_ACTIVE_STATUSES)" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "ok. so basically, I think it's the correct direction and implements the change to adr as discussed." + via: session 2026-09-27, reviewing PR #583 + covers: [] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 +--- + +# ADR-305: Capabilities active at adoption need no add decision + +## Summary + +- **Decided:** a project records under `baseline` in `adr.yaml` the date it adopted adr/v1 and the capabilities it already had then. Those capabilities need no `add` decision, and a `change` on one with no prior record to name stands on the baseline. A capability declared after adoption still needs an `add`. +- **Trades away:** a record of why each baseline capability exists. The vocabulary line is its only description. +- **One-way?** No. Removing a name from `baseline` restores both checks for it. +- **Probes:** *Confident:* the outcome for a record does not depend on which other records have been migrated, since only decisions dated after adoption can be its prior, and every one of those was written as v1. *Not confident:* whether `baseline` will be used to skip an `add` for a capability that is actually new. +- **Inversion:** at one end every capability needs a written `add`, which misstates history for a corpus that predates the contract. At the other end the vocabulary is its own authority and nothing needs an `add`, which lets new capabilities in unrecorded. This decision exempts only what existed before adoption. + +## Context + +ADR-304 §6 requires every capability in the vocabulary to have an accepted `add` decision. Most of agent-ways' capabilities (matching, attend, install, testing and others) worked long before any record described them. Migrating an old record as an `add` misstates it. ADR-186 is the example the operator gave: it moved testing off the host and into a container, which changed a testing capability that no record had added. #582 set out three options: a baseline `add` record per capability, declared capabilities starting active, and a `change` that may list several capabilities. The operator chose the second, and the agent designed the mechanism within that direction. + +## Decision + +### 1. Baseline capabilities + +`adr.yaml` may carry: + +```yaml +baseline: + adopted: 2026-09-27 + capabilities: [docs, testing] +``` + +`adopted` is a `YYYY-MM-DD` date. Every name in `capabilities` must be in the vocabulary. Either defect fails lint. + +### 2. The add check, amended + +The ADR-304 §6 bullet on accepted `add` decisions now reads: + +- Every capability in the vocabulary, except those listed in `baseline`, has an accepted `add` decision. This warns while any v0 record remains and fails after, so a corpus that is still migrating does not fail on every capability. + +### 3. A change on a baseline capability + +ADR-304 §3 requires a `change` to supersede or amend a prior decision on the same capability. A `change` with no `supersedes` or `amends` edge on a baseline capability instead stands on the baseline when either: + +- it is dated on or before `adopted`, or +- no earlier decision on that capability is dated after `adopted`. Earlier means by date, then number. `constrain` decisions do not count, and neither do rejected or abandoned ones. + +Only decisions dated after adoption count as a prior, so migrating an older record never changes the outcome for another record. A record with no date, or a date that is not `YYYY-MM-DD`, cannot stand on the baseline. + +### 4. Stale citations, widened + +The ADR-304 §6 `doclint` bullet on superseded decisions now reads: + +- A citation of a superseded, deprecated, rejected, abandoned or archived record warns. For a superseded record, the warning names the successor. + +`adr cite` already behaves this way. The amendment brings the text in line with it. + +## Consequences + +### Positive + +- Migrating a record no longer requires inventing history. ADR-186 migrates as a `change` on `testing`. +- An adopting project writes one list and one date, not one record per capability. + +### Negative + +- ADR-123 spans attend, matching and disclosure. Only `constrain` may list several capabilities, so ADR-123 stays on v0 until it is split or a rule for multi-capability changes is decided. +- A post-adoption `change` need not name a pre-adoption record on the same capability, even after that record is migrated. The v0 prior warning still applies when it does name one. + +### Neutral + +- Projects that declare no `baseline` behave as before. + +## Alternatives Considered + +- **A baseline `add` per capability.** This keeps the history complete, but at the cost of twelve records written after the fact, each dated at adoption. +- **A multi-capability `change`.** This would let ADR-123 migrate, but it loosens the one-change, one-capability rule, and it does not fix the missing `add` decisions. +- **Count every v1 decision as a prior.** This was the first version of this PR. Which change counted as first then depended on the order in which records were migrated, and a frozen record could start failing when an older one migrated. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md new file mode 100644 index 00000000..54cd04e5 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-306-adr-import-foreign-records-through-a-round-trip-import-sheet.md @@ -0,0 +1,210 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +amends: [ADR-304#7] +basis: + - operator: aaronsb + level: guided + said: "I think we need to make sure that the new adr tools can create the content and manage the lifecycle of data, and probably, we need to think about a 'foreign import' tool that would take any kind of decision record that's not directly lintable/usable, and can ingest the foreign record. this way, it becomes our cannonical 'migration' tool." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "I think the foreign import model needs a round trip data object template of some kind." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "we should assuem that reasonable foregin records have some sort of structured data (frontmatter, for example). a /very foreign/ import could literally be jira issues for example" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "I think optional, but always lint warnings. sometimes, the summary isn't obvious until the apparent motion of the complete dataset is visible. this means that adr record properties can be altered (like summaries). this is fine, because any alteration that gets tracked is a git commit. we don't have to overthink integrity here" + via: session 2026-09-27, on whether imported records need a Summary + - operator: aaronsb + level: guided + said: "we don't need to explicitly handle jira. all I'm saying is 'jira issues can be flattened to a record, just like any other record, and usually there's a description and a summary and various fields, and if we can selectively import jira issues, then we probably can take records from about anything'" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "the domain layout shouldn't ever be set forever. things change over time, certain domains might merge or split. during import from v0 adr to v1 in-repo, it might make sense to add or combine domains. during a full foreign import, perhaps something like a jira or github issues, then these would expand over time." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "it might be more work to curate the records, but if an adr changes domains, then I think it needs to be changed in code." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "or in references" + via: session 2026-09-27, following the message above + - operator: aaronsb + level: guided + said: "yes we trade away hand migrations freedom to restructure a record, but that feels like a forced decision. there's nothing stopping us from transforming the record before import. import is just the acceptance model for a foreign record" + via: session 2026-09-27, reviewing ADR-306 + - evidence: "this repo holds 93 v0 records, each with v0 frontmatter (status, date, deciders, related) that a reader can map without judgement" + - evidence: "in a project that declares contract: adr/v1, `adr new` still writes a v0 record with no contract, kind, verb, capability, basis, agent or Summary" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "i read the adr and it aligns with my understanding. let's accept and merge it" + via: session 2026-09-27, PR #587 + covers: [] + - operator: aaronsb + said: "my assumption is that everything needed is carried in the docs. but because reality can drift and is complex, its possible (and we should assume it happens) that the actual implementation nearly always has some drift from the record of desire." + via: session 2026-09-27, answering the probes after acceptance + covers: [frontmatter-carries-all] + - operator: aaronsb + said: "we import markdown for text. any complex formatting language needs to be markdown. we import a structured body for structured data - I'm not sure what the convention is right now, but yaml or json seems to be the right approach." + via: session 2026-09-27, answering the probes after acceptance + covers: [non-markdown-bodies] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 305 +--- + +# ADR-306: adr import: foreign records through a round-trip import sheet + +## Summary + +- **Decided:** `adr import` is the acceptance model for a foreign record: whatever shape a record arrives in, it becomes an adr/v1 record through an import sheet. `scan` reads records from any structured source into one import sheet per record. The agent fills in what needs judgement. `apply` writes each finished sheet as a v1 record. `adr new` writes through the same writer, and `adr supersede` and `adr enact` complete the lifecycle commands. Domains can be added, merged and split as the corpus grows. A record that moves to another domain is renumbered into that domain's range, and every reference to it, by number or by path, is rewritten. +- **Trades away:** a direct edit from source to record. Every record passes through a sheet, a staging format with its own schema, and other sources through field maps. Both have to be documented and kept stable. +- **One-way?** No. Sheets are staging files and records stay in git. A bad import is reverted like any commit. +- **Probes:** *Confident (frontmatter-carries-all):* v0 records from agent-ways, here or in any repo that adopted them, import with the body unchanged, since everything the reader needs is in the frontmatter. *Not confident (non-markdown-bodies):* whether a field map that flattens a structured item into fields and a body covers sources whose body isn't markdown, or whether some sources need a conversion step first. +- **Inversion:** at one end, a reader written in code for every format: exact, but never finished. At the other end, an agent reads each foreign record and writes v1 by hand: flexible, but manual across a hundred records. This decision maps structured fields mechanically and leaves only the judgement fields to the agent. + +## Context + +ADR-304 §7 moves a v0 record to v1 "when someone next edits it". That works for a trickle of edits, and it doesn't work for a corpus. This repo holds 93 v0 records. Any repo using a file-based record system, or v0 records from agent-ways, holds a corpus of its own, in v0, adr-tools, MADR or tracker formats. #581 migrated two records here by hand, and the tier 2 rehearsal showed an agent can do it without inventing anything, but record by record. + +Most of a migration is mechanical. Status maps by the §7 table, and date, deciders and links carry over. The body stays as written. A few fields need judgement: the verb, the capability, the basis, and a Summary. Those are the only fields an agent should have to touch. + +The tool also has lifecycle gaps. `adr new` writes a v0 record in a v1 project. Supersession needs both sides' links edited by hand, and enactment is a hand-edited field. + +## Decision + +### 1. The import sheet + +Import accepts a record. It does not restrict what happens to the record before it's accepted. A record can be split, merged or rewritten before `scan`, or its sheet edited before `apply`. The importer itself never changes content: whatever the sheet says is what `apply` writes. + +One YAML file per record is the round-trip object between a source and a v1 record: + +```yaml +sheet: adr-import/v1 +source: {path: docs/architecture/system/ADR-186-….md, format: v0, sha256: "…"} +target: {number: 186, domain: system} +record: # v1 frontmatter, filled as far as the reader can + contract: adr/v1 + kind: decision + status: accepted + date: 2026-09-17 + verb: ~ + capability: ~ + basis: [] + agent: {name: Claude, model: unrecorded} +summary: ~ +todo: [verb, capability, basis] +candidates: {capability: [testing, install]} +provenance: {status: "frontmatter status: Accepted"} +unmapped: {deprecation_note: "…"} +body: | + … +``` + +- `record` holds v1 frontmatter. The reader fills what the source states, and `provenance` says where each value came from. +- `todo` lists what the reader could not fill. `candidates` ranks vocabulary matches to help whoever fills them. +- `unmapped` keeps every source field that has no v1 home. Nothing is dropped. +- `body` starts as the source body, verbatim. It can be edited like any other part of the sheet. + +### 2. Readers + +A reader turns a source into sheets. Import assumes structured sources: frontmatter, a metadata block, or a structured export. Unstructured prose is out of scope. + +- **Built in:** v0 (this tool's frontmatter), MADR, and adr-tools (inline `## Status`, `0001-` numbering). +- **Field maps:** any other structured source is flattened into fields and a body. A declarative map names which source field fills which sheet field, and how values translate. No source gets its own reader. The example is a tracker export, since a source like that flattening cleanly suggests most structured records will: + +```yaml +reader: tracker +items: issues # a JSON export: one sheet per item, selected by --filter +fields: + title: fields.summary + date: fields.created + status: {from: fields.status.name, map: {Done: accepted, "Won't Do": rejected, "To Do": proposed}} + body: fields.description + unmapped: [key, fields.labels] +``` + +A body is one of two things: + +- **Text is markdown.** A source whose text is in another formatting language (HTML, a tracker's rich text, wiki markup) is converted to markdown before the sheet is written, by the reader or by a step run before `scan`. +- **Structured data is YAML or JSON.** A source item whose content is data rather than prose keeps it as data: in the sheet, and in the record as a fenced `yaml` or `json` block. + +An import carries what the source says. A record states what was decided, and the implementation nearly always drifts from it to some degree, so an imported record is not evidence of what the code does. `adr cite` and review compare the two after import. + +No reader ever writes an `operator` basis from `deciders`, an assignee or any other metadata (ADR-304 §7, §11). A basis comes from what the record says, and it is filled during cleanup. + +### 3. Commands + +- `adr import scan <paths> [--reader NAME | --map FILE]` writes sheets to `docs/architecture/.import/`. That directory is gitignored: sheets are working files, and only the records they produce are committed. +- `adr import apply [sheets] [--partial]` writes each sheet whose `todo` is empty as a v1 record, then lints it. A sheet with open items is skipped. `--partial` writes it anyway, and lint reports what is missing, except for items lint cannot detect afterwards: a Deprecated record's missing historical note, a status that maps to nothing, and a changed number or domain. Those block even `--partial`. +- `adr new` builds an empty sheet from its arguments and applies it, so a new record and an imported record share one writer. In a v1 project it writes v1. +- `adr supersede <old> --by <new>` writes both sides of the link. `adr enact <n> <commit>` sets `enacted` on an accepted cut or retire. + +### 4. Imported records + +An imported record carries `imported: {from, format, status, unmapped}` in its frontmatter. `status` is the source's raw status, and `unmapped` holds every source field with no v1 home. No source value depends on a todo item being honoured to survive the import. Text found between a source's frontmatter and its title moves into the body, right under the title. For an imported record, a missing Summary is a lint warning. The Summary may be written later, once the whole corpus has been imported and read together, and git history records when it was added. + +### 5. Numbering + +A source numbered inside the project's domain ranges keeps its number. Code, ways and other records cite these numbers, and nothing structural calls for new ones, so an import never renumbers them. This covers every v0 record from agent-ways, in this repo or any other. Any other source gets a number from `target`, which the reader proposes from the domain and the agent may change. `apply` rewrites references within the imported set to the new numbers, and `imported.from` keeps the original identifier. + +### 6. Domains evolve + +The domain layout is not fixed. An import may add or combine domains, and a corpus fed from a tracker keeps growing new ones. A record's number tells you its domain, as it does under v0, so the number follows the domain: + +- A domain may hold several ranges. Merging two domains keeps both ranges, so no record is renumbered. +- A record that moves to another domain, whether by a split or on its own, gets a new number from that domain's range. `adr domain move` rewrites every reference to the record, whether by number (`ADR-N`) or by path (a link to the file, whose folder and name both change). That covers records (`related`, `supersedes` and the other edges, and links in the body), catalog docs, READMEs, ways and code. `adr cite` finds the number references, and the move finds the path references. This is more curation work than keeping the number, and in exchange a number always names its domain. +- The old number is retired and never reused. The record carries `renumbered_from: [ADR-N]`, so a citation that can't be rewritten, such as one in a commit message or a closed pull request, can still be traced, and `adr cite` reports any left in the tree. +- `adr domain add`, `merge`, `split` and `move` edit `adr.yaml`, move and renumber the files, rewrite citations and regenerate the index. +- A sheet whose `target.domain` names a domain that does not exist yet creates it on `apply`, with a free range. + +An import alone never renumbers (§5). Renumbering happens only when a record changes domain. + +### 7. Round-trip guarantees + +These are tested properties: + +- Applying an unedited sheet keeps the body byte-identical, and every source field is either mapped into `record` or kept in `unmapped`. +- Scanning a v1 record and applying the result reproduces the record. +- v0 output of every existing command stays byte-identical. + +## Consequences + +### Positive + +- Migrating a corpus becomes one scan, one cleanup pass over small structured files, and one apply. This repo's 93 v0 records and any other repo's file-based records go through the same path. +- A new source needs a field map, not a code change, whenever it flattens to fields and a body. +- `adr new` produces a v1 record in a v1 project. + +### Negative + +- The sheet schema and the field-map language are formats the tool must keep stable across versions, since sheets and maps outlive a single run. +- Imported records may sit without a Summary, and lint keeps warning until they have one. +- Field maps are a small configuration language that has to be documented and kept stable. + +### Neutral + +- ADR-304 §7's status table is unchanged. The v0 reader applies it. +- A v0 record edited by hand still moves to v1 as before. Import is the bulk path alongside it. + +## Alternatives Considered + +- **An agent migrates each record by hand.** The #581 rehearsal shows this works, but across a corpus of a hundred records the mechanical fields would be retyped each time, and nothing would check the result the way a round trip does. +- **A code reader per format, with no field maps.** This is exact for known formats, but every tracker and template needs code in the vendored tool. +- **Migrate in place with no intermediate object.** A migration writes the record directly. Without a sheet there is no place to stage what needs judgement, and no object to test the round trip against. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md new file mode 100644 index 00000000..d5adb6af --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/documentation/ADR-307-a-decision-names-what-should-be-observable-when-it-holds.md @@ -0,0 +1,129 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: adr +extends: [ADR-304] +amends: [ADR-304#4] +basis: + - operator: aaronsb + level: guided + said: "experiencing the phenomenonolgy of software is a missing piece that I think an agent decision record can promote." + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "A complex flow could even demonstrate the phenomenon of the work itself as one modal in the flow, if it's possible to do so" + via: session 2026-09-27 + - operator: aaronsb + level: guided + said: "Option one but looser because there's so many possible variations depending on the work at hand" + via: "session 2026-09-27, on the shape of an observable; option one, written by the agent, was a plain-words `see` with an optional `run` command" + - operator: aaronsb + level: guided + said: "Optional everywhere" + via: "session 2026-09-27, selected from agent-written options on which decisions must name an observable" + - operator: aaronsb + level: guided + said: "No, `via` covers it" + via: "session 2026-09-27, selected from agent-written options on whether considered records what was seen" + - operator: aaronsb + level: guided + said: "Not yet (Recommended)" + via: "session 2026-09-27, selected from agent-written options on whether the adr tool runs observable commands" + - evidence: "an analysis of 121 vendor documents across four agentic tools found that the platform records events and nothing records the verification duty discharged (arXiv 2608.15678, Aug 2026)" +agent: + name: Claude + model: claude-opus-5-5 +considered: + - operator: aaronsb + said: "It has to be flexible; we are asking for a way of observing function which is not predictable." + via: session 2026-09-27, answering the probes on PR #596 + covers: [flexible-shape] + - operator: aaronsb + said: "I think an ask for an observable is fair. This way the operator could decline or just tell the agent \"observe it yourself, you can loop and iterate\" - don't use that verbatim but that is adjacentto the develop skill" + via: session 2026-09-27, answering the probes on PR #596 + covers: [optional-unused] + - operator: aaronsb + said: "I like the shape of the pr it tracks the minimum needed to state something was actually real. Let's accept and continue" + via: session 2026-09-27, PR #596 + covers: [] +status: accepted +date: 2026-09-27 +deciders: + - aaronsb + - Claude +related: + - 304 + - 306 +--- + +# ADR-307: A decision names what should be observable when it holds + +## Summary + +- **Decided:** a decision may carry `observable`: what someone should be able to see, run or try when the decision holds. Its shape is loose, because the work varies. It is optional on every decision, and it may be added or refined after acceptance. When a decision is handed to the operator, the agent demonstrates its observables, where that is possible, before asking. +- **Trades away:** a checkable form. A loose field can't be verified by the tool, so an observable is only as good as its author makes it. +- **One-way?** No. The field is optional, so removing it later changes no record's validity. +- **Probes:** *Confident (flexible-shape):* a free-form list fits the range of work, from a command's output to a page to click through. *Not confident (optional-unused):* whether an optional field gets used at all, or whether the handover guidance alone carries it. +- **Inversion:** at one end the record is prose to be read, and consideration rests on reading. At the other end every decision carries a runnable check the tool enforces, which fits commands and misses everything seen by eye. This decision names what to observe, leaves its form open, and puts the demonstration in the conversation. + +## Context + +`considered` records the operator's words on a decision, but nothing ties those words to having seen the work. Reading a record is reading about the act. Experiencing what the software does is the missing piece, and a decision record can promote it by saying what should be observable once the decision holds. + +`cut` and `retire` decisions already have an observable end in `enacted`, the commit where the removal landed. `add` and `change` decisions say what was decided, but not how anyone would see that it holds. + +## Decision + +### 1. The observable field + +A decision may carry `observable`, a list. Each entry is either a line of plain words or a mapping whose keys the author chooses to suit the work: + +```yaml +observable: + - "tier 2 adr-migrate passes 10 of 10" + - see: a record scanned and applied comes back byte-identical + run: bash tests/adr-import-roundtrip.sh + - see: the evidence page renders the findings table + url: https://claude.ai/artifact/… +``` + +`see` and `run` are conventions, not requirements. A screenshot path, a URL, a scenario name or a step-by-step description are equally valid. Lint checks only that `observable` is a list of strings or mappings. + +### 2. Optional, and open after acceptance + +No decision is required to carry an observable. `observable` joins the fields that may change after a decision leaves proposed (ADR-304 §4, `mutable_after_accept`), because what shows a decision holding often becomes clear only once it is built. + +### 3. Asked for when drafting, demonstrated in the handover + +When an agent drafts an `add` or `change` decision, it asks the operator what should be observable once the decision holds. The operator can name an observable, decline, or hand the observing to the agent. In the last case the agent works out what to observe, runs the work and iterates until it can show the outcome, as the develop loop does, and then writes the observable it used into the record. + +When a decision with observables is handed to the operator, the agent demonstrates them, where possible, as one step of the flow: it runs the command, shows the output or a screenshot, or opens the page. It asks its questions afterwards. The consider way carries this guidance, and routes a batch of questions through the choices way. `considered.via` says what the operator was shown. No separate field records it. + +### 4. The tool does not run observables + +`adr` does not execute `run` entries. The agent runs them during the handover, or while iterating on the work (§3). A command runner in the tool can be decided later, once there is evidence of how observables are written. + +## Consequences + +### Positive + +- A decision can say how anyone, human or agent, would see it holding, in whatever form suits the work. +- A consideration made after a demonstration differs, in its `via`, from one made after reading. + +### Negative + +- Nothing enforces that an observable exists, or that it works. +- A loose field is harder to use mechanically later, for a runner or a report. + +### Neutral + +- Imported records carry no observables until someone adds them, which §2 allows. +- `enacted` stays as it is. It is the observable end of a cut or retire. + +## Alternatives Considered + +- **A fixed shape: `see` plus an optional `run`.** Rejected by the operator as too narrow for the variety of work. +- **Required on add and change, with a lint warning.** Rejected in favour of optional everywhere. The agent asks when drafting (§3), and the field itself stays optional. +- **A `seen` list on `considered`.** Rejected: `via` already says what the operator was shown. +- **`adr observe N` running a decision's commands.** Deferred until observables have been written in practice. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md b/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md new file mode 100644 index 00000000..644f96d1 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-200-compliance-claims-and-session-derived-findings.md @@ -0,0 +1,316 @@ +--- +status: Accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-005 + - ADR-013 + - ADR-110 + - ADR-111 + - ADR-151 + - ADR-201 +supersedes: + - ADR-005 +--- + +# ADR-200: Compliance claims and session-derived findings + +## Context + +The framework carries a provenance subsystem (ADR-005) and a rationale for it +(ADR-013): a way can carry metadata mapping its guidance to regulatory controls +(NIST 800-53, OWASP, ISO 27001, SOC 2, CIS, IEEE), stored today as a `provenance.yaml` +sidecar beside the way (ADR-110) and queried by a `ways governance` CLI (ADR-111). + +In use, three problems surfaced — each worse than the last. + +**It is misnamed.** Mapping practices to external standards and reporting on that +mapping is **compliance** work — demonstrating conformance to an outside rule — not +**governance**, which is the steering discipline (ADRs, ways, review gates) the +framework is already saturated with. Calling a narrow control-mapping tool +"governance" oversells it as the steering layer *and* strands the real governance +without the word. That category error is much of why the subsystem reads as bolted on. + +**Nothing consumes it.** No CI job, hook, or reviewer runs it. The `governance-cite` +guidance tells the model to cite controls when recommending practices, but that is +advisory and manual. The machinery is a library with no callers. + +**Its claims were never assessed** — the decisive problem. Each provenance sidecar is +hand-authored: the control IDs, the justifications, the `verified` date are all typed +by a person. Nothing ever assessed whether a way *actually* meets a control. In audit +vocabulary the distinction is exact and we collapsed it: + +- An **implementation claim** is a statement that a control is *designed* to do its + job. That is all the sidecars are. Auditors call this posture **SOC 2 Type I** — + *suitably designed as of a point in time.* +- An **assessment finding** is what an assessor produces after examining real + operation. NIST's assessment standard (SP 800-53A) records a finding with a + determination of **satisfied** or **other than satisfied**. The richer report, + **SOC 2 Type II**, covers *design **and** operating effectiveness over a period.* + +A hand-written citation is an implementation claim (Type I) wearing the costume of a +finding (Type II). ADR-013 made this exact error out loud — it called provenance +"auditability … pull real control citations" and said it "turns Claude from a hack +into a professional." A claim on demand is not proof; it is an assertion awaiting an +assessment. + +### The reframe + +Separate the two epistemic statuses ADR-005/013 conflated, using the field's real +names, and state honestly what the subsystem is a *foundation for*: + +- A **claim** is a control-design assertion (SOC 2 Type I; in NIST's OSCAL data model, + a *Component Definition* — "here is how this component can satisfy the control"). +- A **finding** is a session-derived assessment finding (SOC 2 Type II; OSCAL + *Assessment Results*): evidence that the claimed guidance actually shaped real work. + +The connective tissue between them is an **assurance case** — a spelled-out argument +tying a claim to its evidence so a reader can check the *reasoning*, not just the +proof. The notation whose vocabulary we use literally is **CAE (Claims–Arguments– +Evidence)**; its graphical sibling GSN is the same idea (GSN confusingly names the +evidence node a "Solution"). This is also, recursively, the model OSCAL itself uses — +a compliance-claim system built this way follows the standard shape of compliance +systems. + +### The formation thesis + +Agent-ways operates at the **first line** — in the Three Lines Model (the Institute of +Internal Auditors' 2020 update of the old "three lines of defense"), the roles that +own and manage risk *in the actual doing of the work*, distinct from second-line +functions that monitor them and third-line auditors who independently check them. + +A first-line developer does not recite NIST 800-53 while committing. Through repeated +exposure to a way of working, the control's expectations become **tacit knowledge** — +Polanyi's "we can know more than we can tell," the expertise the hands hold that words +can't fully capture. A way is the **externalization** of that shape (in Nonaka's SECI +model, the tacit→explicit step; more generally, **codification**) delivered at the +point of work. That shaped habit is the expensive precondition every audit assumes but +none create: when the work already has the standard's shape, evidence can be *gathered* +later instead of *manufactured*. + +So the compliance stack, top to bottom: + +``` + certify → independent attestation (SOC 2, SSAE 18) or an OSCAL SSP ← third line / GRC toolchain — NOT us + finding → "the shaped work actually happened" (Type II) ← ways-audit (ADR-151) + claim → "this way is designed to steer toward control X" (Type I) ← provenance sidecars + ────────────────────────────────────────────────────────────────── + way → externalized guidance that shapes the work ← agent-ways ★ first line, we live here + work → the actual activity +``` + +We own the bottom two natively. The compliance subsystem reaches up into claim and +finding and deliberately **stops below certify** — we are a first-line formation aid, +not a third-line assurance function. The honest top-level statement is therefore: +**foundation for, not implementation of, a GRC/OSCAL toolchain.** And that statement is +itself a claim awaiting its own findings — the system is an instance of its own thesis. + +## Decision + +Reframe the subsystem from asserted "governance traceability" into an +**operator-invoked compliance formation layer** with an honest claim/finding split. +This supersedes ADR-005 and refines (does not discard) ADR-013. + +### 1. Claims are claims, not evidence + +The provenance sidecars are **compliance claims** — control-*design* assertions (SOC 2 +Type I; OSCAL Component Definition). They carry no weight beyond assertion. ADR-013's +framing of them as auditability or proof is retracted. A claim is a hypothesis awaiting +a finding. + +**A claim must be assessable.** To seed claims across the corpus without recreating the +overclaim at scale, each claim carries a **determination criterion** — a plain +statement of what observable behavior in a session would let an assessor mark it +*satisfied* or *other than satisfied* (NIST 800-53A). A claim with no such criterion is +unfalsifiable: it can never graduate into a finding, and coverage of such claims is not +evidence of anything. Claim-authoring may be **agent-assisted** — `ways-audit` can +propose candidate controls for a way so the author is not hand-researching every way — +but it stays **human-grounded**: what a control *means*, and whether a way honestly +addresses it, is the author's epistemic call (ADR-013's collaborative authorship model). + +### 2. Findings are session-derived + +Evidence is produced, not authored: + +``` +claim (sidecar) → way fires → logged event → auditor reads (claim + firing + transcript) → finding +``` + +The firing-event log and session transcripts already exist. A **finding** records +`(claim-id, session-id, timestamp, determination, evidence-pointers)`, its +determination **satisfied** / **other than satisfied** in NIST 800-53A's own terms, the +pointer being the transcript and firing event so every finding is auditable back to its +source. Findings accumulate in an append-only ledger; each *other than satisfied* +finding becomes a **POA&M** (plan of action & milestones) entry — a tracked gap, not a +vague fail. + +### 3. Two evidence tiers, honestly separated + +Audit-evidence practice distinguishes **process evidence** (a step was performed) from +**outcome evidence** (the intended result was achieved). We adopt both, separately: + +- **Process tier** — "control-aligned guidance was surfaced at the point of work." The + firing log *alone* substantiates this: timestamped, deterministic, no model judgment. + Defensible today; this is the MVP. +- **Outcome tier** — "and the work actually took that shape." This needs an auditor + reading the transcript, is model-judged, is fallible, and carries a credibility + problem: a model grading whether a model followed model-injected guidance. It is the + stretch tier and must be labeled opinion, never proof. + +Outcome-finding credibility rests on auditor trustworthiness — independence, +deterministic checks where possible, optional human counter-sign. This is the central +unsolved problem and is named as such, not hidden. + +### 4. Operator-invoked, never forced + +There is **no mandatory claim → finding → certify traversal.** Compliance review is a +station the operator pulls into deliberately, like `/ship` or `/wrap` — not a gate +anyone is dragged through. Concretely: a `compliance-reviewer` agent (sibling of +`code-reviewer`) plus skills that let the main agent direct it and do the loop-closing +work. It slots into delivery as an *optional* stage — `code-review → compliance-review +→ fix → ship` — where findings become labeled issues (the POA&M). The tooling that +backs it is `ways-audit` (ADR-151), not the core binary. + +### 5. Borrow the vocabulary; do not implement the standard + +Use OSCAL and assurance-case nouns (claim, finding, POA&M, assessment) for precision +and to shape output a real toolchain could ingest. We do **not** implement the OSCAL +schema, validate against it, or produce a System Security Plan (SSP). Using the +vocabulary must never read as a conformance claim to the standard itself. + +Reference standards by ID and authoritative URL; never reproduce normative text. NIST +800-53/OSCAL/800-53A and OWASP are public; CIS is free-with-registration; **ISO 27001 +and ISO 25010 are copyrighted** — cite the clause ID only. + +### 6. Evidence is shaped to be handed up + +Findings must be legible to the layer above (OSCAL Assessment-Results-shaped, control +IDs that resolve to real catalogs). Our job is to *feed* a certifier we are not — never +to terminate the chain with an attestation of our own. + +### 7. The progressive-ways loop + +Reframe `gaps` from "ways without provenance" to **"claimed controls with no +influencing way."** The claim set then *drives the ways backlog*: every control we +claim but do not yet steer toward is authoring work. This is what makes the subsystem a +living formation layer rather than a static bibliography — the claim creates the +obligation to build the way. + +### 8. Rename on the operator surface + +"Governance" leaves the operator-facing surface: the toolkit is `ways-audit`, its +domain concept is **compliance**, its artifacts are **claim / finding / POA&M**. +Storage stays the `provenance.yaml` sidecar (ADR-110), consumed by `ways-audit` rather +than the core binary. "Governance" survives only as the name of this ADR topic band +(200–299), whose `adr.yaml` description already reads "compliance mapping." + +## Non-Goals + +This section is load-bearing — the fence that stops a future session from growing this +into the thing it is not. **The system is incapable of overclaiming by construction:** +every artifact says "we *claim* X" or "we found *evidence* that guidance for X was +present/followed," never "you *conform to* X." It is a first-line formation aid, not a +third-line assurance function. + +This is **not**: + +- a GRC platform or compliance-management product; +- an OSCAL implementation, validator, or SSP generator; +- audit-grade evidence, an attestation, or an independent assessment; +- a continuous, automated, or enforced control; +- a certification, or any claim of conformance. + +It **is**: an opinionated, operator-invoked lens that helps consider compliance-shaped +concerns at the point of work, curate ways toward controls the operator chooses to +claim, and keep a modest, honestly-labeled claims-and-findings trail — *a way of +working that leans compliance-shaped, not a compliance framework.* + +## Relationship to prior ADRs + +- **Supersedes ADR-005** (Governance Traceability for Ways). The mechanism — sidecar + claims, generated manifest, reporting-not-enforcing — is retained and renamed; the + framing of provenance as "auditability / proof" is replaced by the claim/finding + split. +- **Refines ADR-013**, does not discard it. Kept: the activation-cue thesis (ways prime + latent knowledge rather than inject it), the Type 1–4 way classification, the + author-as-compiler model, and the formation intuition latent in its "Shift Lead + Wisdom" type. Corrected: the epistemic overclaim that a citation on demand + constitutes evidence or professional proof. Under this ADR, ADR-013's Type 1 ways + carry *claims*; findings are what a session produces. +- **Relates to ADR-110** (provenance moved to a `provenance.yaml` sidecar) for storage, + and **ADR-111** (CLI consolidation) for the command surface that becomes `ways-audit`. +- **Requires ADR-151** for packaging: the `ways-core` library crate and the `ways-audit` + sibling binary. + +## Consequences + +### Positive + +- The subsystem stops overclaiming and starts being true: claims are claims, findings + are earned and carry a real determination. +- A defensible MVP exists immediately (process evidence from the firing log) without + waiting on the hard outcome-audit problem. +- The formation-layer framing gives every downstream decision a coherent "why," and + positions agent-ways to *feed* enterprise GRC toolchains rather than compete with them. +- The progressive-ways loop turns claims into an authoring backlog — the subsystem + becomes generative. +- Operator-invoked design keeps zero forced ceremony; users who never touch it pay + nothing. + +### Negative + +- Real work to realize: rename, de-assertion of existing sidecars, the `ways-audit` + toolkit (ADR-151), the auditor agent + skills, and the finding-ledger format. +- The outcome-tier credibility problem is genuine and unsolved; that tier ships labeled + as opinion or not at all. +- Borrowing OSCAL/assurance vocabulary invites the very "are you OSCAL-conformant?" + overclaim the Non-Goals forbid; the honesty framing must be maintained vigilantly. + +### Neutral + +- The ADR domain folder stays "governance" (topic band); only the operator surface + renames to "compliance." +- Existing provenance sidecars remain valid data — reclassified as claims, not rewritten. +- Whether the outcome tier ever ships is left open; the process tier stands on its own. + +## Alternatives Considered + +- **Keep it as "governance traceability" (status quo).** Rejected: misnamed, + unconsumed, and epistemically overclaiming — the problems that motivated this ADR. +- **Delete it entirely.** Rejected: the data, policy docs, and CLI are a real + foundation, and the formation-layer idea is valuable and rare. Deleting discards a + good foundation to avoid fixing a label and an epistemic error. +- **Promote to a full compliance/GRC platform.** Rejected: that is the third line, not + ours. Certification, SSPs, and independent assessment belong to tools built for them; + we produce foundation-grade inputs for those tools. +- **Bake in a mandatory claim → finding → certify loop.** Rejected: forced traversal is + exactly the bolt-on failure mode. The value is *availability* to the operator, like + any session-defining skill. +- **Keep provenance framed as evidence (ADR-013 as written).** Rejected: it is the + overclaim this ADR exists to correct. + +## References + +Cited by identifier; normative text of copyrighted standards (ISO) is not reproduced. + +- **The Three Lines Model** — IIA, 2020 update of the "three lines of defense" + (governance roles; "first line" = risk-owning roles in the work). + https://www.theiia.org/en/content/communications/2020/july/20-july-2020-iia-issues-important-update-to-three-lines-model/ +- **SOC 2 Type I vs Type II** — AICPA attestation under SSAE 18 / Trust Services + Criteria (design at a point in time vs design *and* operating effectiveness over a + period). +- **Assurance cases — CAE (Claims–Arguments–Evidence)** and **GSN** — ISO/IEC 15026-2; + the claim → argument → evidence structure this ADR mirrors. +- **OSCAL** (NIST) — Catalog / Profile / Component Definition / Assessment Results / + POA&M layered model. https://pages.nist.gov/OSCAL/learn/concepts/layer/ +- **NIST SP 800-53A Rev. 5** — assessment findings; determinations "satisfied" / + "other than satisfied." https://csrc.nist.gov/glossary/term/assessment_findings +- **PCAOB AS 1105** — *Audit Evidence* (the general standard; the process-vs-outcome + framing is our application of it). +- **Tacit knowledge** — Polanyi, *The Tacit Dimension* (1966): "we can know more than we + can tell." **Externalization / codification** — Nonaka & Takeuchi, SECI model. +- **Procedural vs declarative knowledge** — Ryle, *The Concept of Mind* (1949): + knowing-how vs knowing-that. +- **ADR-005, ADR-013** — the prior decisions this ADR supersedes and refines. diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md b/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md new file mode 100644 index 00000000..c6aa6e22 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/governance/ADR-201-findings-assembled-as-classifier-ready-assessment-records.md @@ -0,0 +1,145 @@ +--- +status: Accepted +date: 2026-07-02 +deciders: + - aaronsb + - claude +related: + - ADR-110 + - ADR-151 + - ADR-200 +--- + +# ADR-201: Findings assembled as classifier-ready assessment records + +## Context + +ADR-200 separated a **compliance claim** (a control-*design* assertion; SOC 2 Type I; +OSCAL Component Definition) from an **assessment finding** (session-derived evidence +with a determination of *satisfied* / *other than satisfied*; NIST SP 800-53A; OSCAL +Assessment Results). It also required, in §1, that every claim carry a **determination +criterion** — the observable session behavior that would let an assessor decide the +finding. ADR-151 packages the tool that produces findings: `ways-audit`. + +ADR-200 left one thing unresolved: **who fills the determination, and when.** Two +tempting answers are both wrong. + +- **The tool determines inline.** If `ways-audit` reads a claim and its firing evidence + and writes `satisfied` itself, it manufactures the very evidence ADR-200 forbids — a + model grading whether a model followed model-injected guidance, stamped as fact. This + recreates the ADR-013 overclaim at the finding layer. +- **A human always determines.** If every finding must wait on a person to read the + transcript and rule, the assessment surface is an unbounded labor sink. It does not + scale past a handful of ways, and the corpus is designed to hold hundreds (ADR-200 §7). + +The honest position is neither. The tool should **assemble** the finding — gather the +inputs and leave the determination *unset* — producing a record that some *other* +system will classify. That classifying system is deliberately **out of scope**: we do +not design it, ship it, or constrain what it is. In NIST 800-53A the assessor is a +*role*, not a person; generalize it to a **classifier** and leave it entirely external. +`ways-audit`'s whole responsibility is to make the record **cleanly classifiable** — +well-structured and easily accessible — and then stop. + +## Decision + +### 1. `ways-audit` assembles findings; it never determines them + +The finding pipeline reads `(claim + its determination criterion + firing events + +transcript pointers)` and writes a finding record with the **determination slot empty**. +The tool's output is an *assessable record*, not an assessment. No code path in +`ways-audit` writes a determination value — that boundary is the load-bearing honesty of +this ADR. + +### 2. An assembled finding is a classifier-ready record + +Frame the finding record as one row of a supervised-classification dataset — *features* +(the observable inputs) and a *label* (the determination) filled later: + +| Role | Field | Source | +|------|-------|--------| +| `way`, `control` | the claimed `(way, control)` this row assesses | claim sidecar (ADR-110) | +| feature | `criterion` — the claim's `satisfied_when` (`null` if none) | claim sidecar | +| feature | `evidence` — `total` fire count, `sessions` (each id is a transcript pointer), `first_seen` / `last_seen` | firing-event log | +| feature | `tier` — `process` when there is no criterion, `outcome` when there is (ADR-200 §3) | criterion presence | +| **label** | `determination` — `satisfied` / `other-than-satisfied`, `null` until labeled | a classifier | +| provenance | `assessed_by`, `assessed_at`, `basis` — `null` until labeled | the classifier, when it labels | +| provenance | `assembled_at`, `assembled_by` — who built the row, when | the assembler | + +*Informally:* the ledger is a dataset with an empty label column. The framing is not +decorative — it is why the record stores the criterion and the raw evidence *beside* the +empty determination, so a finding is classifiable and auditable, not a bare verdict. + +The record also carries empty **provenance** slots — `assessed_by`, `assessed_at`, +`basis` — so that when something does fill the label, *what* filled it and *on what* are +recorded next to it. `ways-audit` writes those slots empty and never fills them. + +### 3. The classifier is out of scope — a different, undesigned system + +We do not design, ship, or constrain the system that classifies these records. It is a +separate concern for a later day, and possibly a different tool entirely. This ADR fixes +only the **record** and the **boundary**: `ways-audit` produces clean, accessible, +classifiable rows and stops. Whatever eventually reads them — a deterministic check for a +*process-tier* criterion, a human reviewer, a trained model for an *outcome-tier* one +(ADR-200 §3) — is not our design problem here, and nothing in this decision presumes +which it will be. The one guarantee we make is structural: the dataset is easy to get at +and unambiguous to label. + +## Consequences + +### Positive + +- Resolves ADR-200's open question without either overclaim (tool-determines) or an + unbounded human labor sink (human-only). +- The determination boundary is enforced in code, not just prose: `ways-audit` has no + path that writes a determination. +- The scope is small and honest: the tool ships a dataset, not an assessment engine. + Classification is someone else's problem, on someone else's schedule. +- The ledger is dual-purpose at no extra authoring cost — an audit trail *and* a labeled- + when-classified dataset — and stays useful even if the classifier is never built. + +### Negative + +- The finding record carries more structure (criterion + evidence + empty label and + provenance slots) than a bare verdict, and the schema must stay stable enough to be a + dataset over time. +- A dataset with a perpetually-empty label column is inert until *something* classifies + it; this ADR deliberately does not deliver that something, so findings have no + determinations until a separate effort supplies a classifier. + +### Neutral + +- The classifying system is entirely out of scope — not designed, not stubbed, not + constrained here. This ADR fixes only the record shape and the assemble-never-determine + boundary. +- `satisfied_when` moves from ADR-200 §1's "proposed, not yet in schema" note into the + actual claim schema (ADR-151 §3) as part of realizing this ADR. +- The ledger is **append-only**, like the firing-event log it is built from: each + `assemble --write` appends a fresh *snapshot* of the rows (their `assembled_at` and + evolving firing evidence distinguish them), rather than mutating prior rows in place. It + is therefore a time series of assemblies, not a unique-keyed table; a consumer that wants + current state reduces by `(way, control)` to the latest `assembled_at`. This keeps the + writer a trivial, honest append and defers any compaction policy to the — out of scope — + consumer. + +## Alternatives Considered + +- **Tool writes the determination inline.** Rejected: manufactures evidence; the exact + error ADR-200 was written to stop, relocated one layer down. +- **Determinations are human-only.** Rejected: an unbounded labor sink that cannot cover + a corpus of hundreds; also forecloses the deterministic process-tier check that needs + no human at all. +- **No stored criterion/evidence — store only the verdict.** Rejected: a bare verdict is + neither auditable back to its basis nor usable as classifier training data, discarding + the record's second purpose for nothing saved. + +## References + +- **ADR-200** — the claim/finding model this refines (§1 assessability, §2 findings, §3 + the two evidence tiers). +- **ADR-151** — the `ways-core` claim schema and the `ways-audit` binary that assembles + findings. +- **ADR-110** — the `provenance.yaml` claim sidecar the criterion is stored in. +- **NIST SP 800-53A** — assessment findings and the *satisfied* / *other than satisfied* + determination; the assessor as a role. +- **OSCAL Assessment Results** — the machine-readable finding layer this record shape + anticipates. diff --git a/docs/architecture/legacy/ADR-004-way-macros.md b/tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-004-way-macros.md similarity index 100% rename from docs/architecture/legacy/ADR-004-way-macros.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-004-way-macros.md diff --git a/docs/architecture/legacy/ADR-005-governance-traceability.md b/tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-005-governance-traceability.md similarity index 100% rename from docs/architecture/legacy/ADR-005-governance-traceability.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-005-governance-traceability.md diff --git a/docs/architecture/legacy/ADR-013-ways-skills-governance-architecture.md b/tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-013-ways-skills-governance-architecture.md similarity index 100% rename from docs/architecture/legacy/ADR-013-ways-skills-governance-architecture.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-013-ways-skills-governance-architecture.md diff --git a/docs/architecture/legacy/ADR-014-tfidf-semantic-matcher.md b/tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-014-tfidf-semantic-matcher.md similarity index 100% rename from docs/architecture/legacy/ADR-014-tfidf-semantic-matcher.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/legacy/ADR-014-tfidf-semantic-matcher.md diff --git a/docs/architecture/system/ADR-100-ways-scaffolding-wizard.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-100-ways-scaffolding-wizard.md similarity index 100% rename from docs/architecture/system/ADR-100-ways-scaffolding-wizard.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-100-ways-scaffolding-wizard.md diff --git a/docs/architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md similarity index 100% rename from docs/architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-101-wormhole-relay-protocol-for-cross-instance-agent-communication.md diff --git a/docs/architecture/system/ADR-102-irc-based-local-agent-communication.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-102-irc-based-local-agent-communication.md similarity index 100% rename from docs/architecture/system/ADR-102-irc-based-local-agent-communication.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-102-irc-based-local-agent-communication.md diff --git a/docs/architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md similarity index 100% rename from docs/architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-103-checks-epoch-distance-aware-confidence-sensors-for-ways.md diff --git a/docs/architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md similarity index 100% rename from docs/architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-104-token-gated-way-re-disclosure-for-long-context-windows.md diff --git a/docs/architecture/system/ADR-105-progressive-disclosure-for-way-trees.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-105-progressive-disclosure-for-way-trees.md similarity index 100% rename from docs/architecture/system/ADR-105-progressive-disclosure-for-way-trees.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-105-progressive-disclosure-for-way-trees.md diff --git a/docs/architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md similarity index 100% rename from docs/architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-106-project-pulse-epoch-mapped-project-awareness.md diff --git a/docs/architecture/system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md similarity index 100% rename from docs/architecture/system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-107-way-match-corpus-batch-mode-and-locale-support.md diff --git a/docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md similarity index 100% rename from docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-108-embedding-based-way-matching-with-all-minilm-l6-v2.md diff --git a/docs/architecture/system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md similarity index 100% rename from docs/architecture/system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-109-project-scope-way-embedding-with-manifest-based-staleness-detection.md diff --git a/docs/architecture/system/ADR-110-way-file-separation-and-graph-compatible-structure.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-110-way-file-separation-and-graph-compatible-structure.md similarity index 100% rename from docs/architecture/system/ADR-110-way-file-separation-and-graph-compatible-structure.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-110-way-file-separation-and-graph-compatible-structure.md diff --git a/docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md similarity index 100% rename from docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-111-unified-ways-cli-single-binary-tool-consolidation.md diff --git a/docs/architecture/system/ADR-113-attend-active-awareness-module.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-113-attend-active-awareness-module.md similarity index 100% rename from docs/architecture/system/ADR-113-attend-active-awareness-module.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-113-attend-active-awareness-module.md diff --git a/docs/architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md similarity index 100% rename from docs/architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-114-attend-as-insistent-way-trigger-type.md diff --git a/docs/architecture/system/ADR-115-declarative-config-with-project-scope-overlay.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-115-declarative-config-with-project-scope-overlay.md similarity index 100% rename from docs/architecture/system/ADR-115-declarative-config-with-project-scope-overlay.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-115-declarative-config-with-project-scope-overlay.md diff --git a/docs/architecture/system/ADR-116-declarative-permission-requirements.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-116-declarative-permission-requirements.md similarity index 100% rename from docs/architecture/system/ADR-116-declarative-permission-requirements.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-116-declarative-permission-requirements.md diff --git a/docs/architecture/system/ADR-117-sensor-crate-extraction-and-feature-flags.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-117-sensor-crate-extraction-and-feature-flags.md similarity index 100% rename from docs/architecture/system/ADR-117-sensor-crate-extraction-and-feature-flags.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-117-sensor-crate-extraction-and-feature-flags.md diff --git a/docs/architecture/system/ADR-118-focus-groups-dynamic-agent-grouping.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-118-focus-groups-dynamic-agent-grouping.md similarity index 100% rename from docs/architecture/system/ADR-118-focus-groups-dynamic-agent-grouping.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-118-focus-groups-dynamic-agent-grouping.md diff --git a/docs/architecture/system/ADR-119-action-potential-engagement-model.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-119-action-potential-engagement-model.md similarity index 100% rename from docs/architecture/system/ADR-119-action-potential-engagement-model.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-119-action-potential-engagement-model.md diff --git a/docs/architecture/system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md similarity index 100% rename from docs/architecture/system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-120-interactive-chat-tui-human-in-the-signal-loop.md diff --git a/docs/architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md similarity index 100% rename from docs/architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-121-salience-decay-for-signal-presentation-turn-based-exponential.md diff --git a/docs/architecture/system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md similarity index 100% rename from docs/architecture/system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-122-attend-disclosure-sensor-token-gated-affordance-reheat.md diff --git a/docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md similarity index 100% rename from docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md diff --git a/docs/architecture/system/ADR-124-channel-bar-ordering-open-as-base.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-124-channel-bar-ordering-open-as-base.md similarity index 100% rename from docs/architecture/system/ADR-124-channel-bar-ordering-open-as-base.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-124-channel-bar-ordering-open-as-base.md diff --git a/docs/architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md similarity index 100% rename from docs/architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-125-authored-disclosure-graph-and-removal-of-bm25.md diff --git a/docs/architecture/system/ADR-126-window-relative-refire.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-126-window-relative-refire.md similarity index 100% rename from docs/architecture/system/ADR-126-window-relative-refire.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-126-window-relative-refire.md diff --git a/docs/architecture/system/ADR-127-reject-full-body-embedding-corpus.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-127-reject-full-body-embedding-corpus.md similarity index 100% rename from docs/architecture/system/ADR-127-reject-full-body-embedding-corpus.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-127-reject-full-body-embedding-corpus.md diff --git a/docs/architecture/system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md similarity index 100% rename from docs/architecture/system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-128-memory-as-repo-portable-ways-seed-routing-over-accumulated-snapshots.md diff --git a/docs/architecture/system/ADR-129-instance-suffix-and-heartbeat-liveness.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-129-instance-suffix-and-heartbeat-liveness.md similarity index 100% rename from docs/architecture/system/ADR-129-instance-suffix-and-heartbeat-liveness.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-129-instance-suffix-and-heartbeat-liveness.md diff --git a/docs/architecture/system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md similarity index 100% rename from docs/architecture/system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-130-sentence-salience-input-reduction-for-embed-matching.md diff --git a/docs/architecture/system/ADR-131-project-scope-way-toggles.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-131-project-scope-way-toggles.md similarity index 100% rename from docs/architecture/system/ADR-131-project-scope-way-toggles.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-131-project-scope-way-toggles.md diff --git a/docs/architecture/system/ADR-132-collaboration-ways-domain.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-132-collaboration-ways-domain.md similarity index 100% rename from docs/architecture/system/ADR-132-collaboration-ways-domain.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-132-collaboration-ways-domain.md diff --git a/docs/architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md similarity index 100% rename from docs/architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-134-empirical-auto-tuning-from-fire-and-near-miss-telemetry.md diff --git a/docs/architecture/system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md similarity index 100% rename from docs/architecture/system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-135-content-aware-write-time-over-build-gate-with-a-self-extending-pattern-corpus.md diff --git a/docs/architecture/system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md similarity index 100% rename from docs/architecture/system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-136-split-addressed-messaging-from-the-sensor-observation-bus.md diff --git a/docs/architecture/system/ADR-137-boundedness-bounded-work-per-cycle.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-137-boundedness-bounded-work-per-cycle.md similarity index 100% rename from docs/architecture/system/ADR-137-boundedness-bounded-work-per-cycle.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-137-boundedness-bounded-work-per-cycle.md diff --git a/docs/architecture/system/ADR-138-skills-own-the-how-ways-own-the-5w.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-138-skills-own-the-how-ways-own-the-5w.md similarity index 100% rename from docs/architecture/system/ADR-138-skills-own-the-how-ways-own-the-5w.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-138-skills-own-the-how-ways-own-the-5w.md diff --git a/docs/architecture/system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md similarity index 100% rename from docs/architecture/system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-139-shelve-maintainer-i18n-adopter-run-localization-via-ways-localize.md diff --git a/docs/architecture/system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md similarity index 100% rename from docs/architecture/system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-140-two-install-topologies-in-place-repo-and-subdirectory-projection.md diff --git a/docs/architecture/system/ADR-141-knowledge-graph-as-evidential-memory-backend.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-141-knowledge-graph-as-evidential-memory-backend.md similarity index 100% rename from docs/architecture/system/ADR-141-knowledge-graph-as-evidential-memory-backend.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-141-knowledge-graph-as-evidential-memory-backend.md diff --git a/docs/architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md similarity index 100% rename from docs/architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-142-agent-ways-1-0-xdg-application-distribution.md diff --git a/docs/architecture/system/ADR-143-three-root-way-runtime-core-user-project.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-143-three-root-way-runtime-core-user-project.md similarity index 100% rename from docs/architecture/system/ADR-143-three-root-way-runtime-core-user-project.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-143-three-root-way-runtime-core-user-project.md diff --git a/docs/architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md similarity index 100% rename from docs/architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-144-install-repair-migrate-as-one-manifest-reconciler.md diff --git a/docs/architecture/system/ADR-146-installer-binary-verification-and-guided-build-fallback.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-146-installer-binary-verification-and-guided-build-fallback.md similarity index 100% rename from docs/architecture/system/ADR-146-installer-binary-verification-and-guided-build-fallback.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-146-installer-binary-verification-and-guided-build-fallback.md diff --git a/docs/architecture/system/ADR-147-composable-settings-json-config-fragments.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-147-composable-settings-json-config-fragments.md similarity index 100% rename from docs/architecture/system/ADR-147-composable-settings-json-config-fragments.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-147-composable-settings-json-config-fragments.md diff --git a/docs/architecture/system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md similarity index 100% rename from docs/architecture/system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-148-framework-surface-ships-operator-content-dev-harness-in-project-scope.md diff --git a/docs/architecture/system/ADR-149-operator-config-interview-skill.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-149-operator-config-interview-skill.md similarity index 100% rename from docs/architecture/system/ADR-149-operator-config-interview-skill.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-149-operator-config-interview-skill.md diff --git a/docs/architecture/system/ADR-150-version-truth-and-downgrade-safe-self-update.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-150-version-truth-and-downgrade-safe-self-update.md similarity index 100% rename from docs/architecture/system/ADR-150-version-truth-and-downgrade-safe-self-update.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-150-version-truth-and-downgrade-safe-self-update.md diff --git a/docs/architecture/system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md similarity index 100% rename from docs/architecture/system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-151-extract-ways-core-crate-and-ways-audit-sibling-binary.md diff --git a/docs/architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md similarity index 100% rename from docs/architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-152-framework-default-secret-path-deny-baseline.md diff --git a/docs/architecture/system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md similarity index 100% rename from docs/architecture/system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-153-session-introspection-substrate-correlating-fired-ways-to-turns.md diff --git a/docs/architecture/system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md similarity index 100% rename from docs/architecture/system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-154-rethink-think-and-non-interactive-introspection-one-model-three-front-ends.md diff --git a/docs/architecture/system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md similarity index 100% rename from docs/architecture/system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-155-semantic-gating-of-the-keyword-channel-and-reasoning-channel-rebuild.md diff --git a/docs/architecture/system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md similarity index 100% rename from docs/architecture/system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-156-calibrated-relevance-scoring-for-the-semantic-lane.md diff --git a/docs/architecture/system/ADR-157-case-insensitive-trigger-regex-compilation.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-157-case-insensitive-trigger-regex-compilation.md similarity index 100% rename from docs/architecture/system/ADR-157-case-insensitive-trigger-regex-compilation.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-157-case-insensitive-trigger-regex-compilation.md diff --git a/docs/architecture/system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md similarity index 100% rename from docs/architecture/system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-158-calibration-boundary-quality-hard-negatives-and-fire-breadth-ship-gate.md diff --git a/docs/architecture/system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md similarity index 100% rename from docs/architecture/system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-159-remove-ways-tune-curves-and-the-legacy-curve-cadence-field.md diff --git a/docs/architecture/system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md similarity index 100% rename from docs/architecture/system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-160-chunked-late-interaction-matching-with-softmax-share-gating-for-way-selection.md diff --git a/docs/architecture/system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md similarity index 100% rename from docs/architecture/system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-161-queued-mid-turn-operator-messages-as-an-aggregated-scan-surface.md diff --git a/docs/architecture/system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md similarity index 100% rename from docs/architecture/system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-162-mechanical-session-link-suppression-as-defense-against-transcript-disclosure.md diff --git a/docs/architecture/system/ADR-163-config-separation-dotfiles-source-of-truth.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-163-config-separation-dotfiles-source-of-truth.md similarity index 100% rename from docs/architecture/system/ADR-163-config-separation-dotfiles-source-of-truth.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-163-config-separation-dotfiles-source-of-truth.md diff --git a/docs/architecture/system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md similarity index 100% rename from docs/architecture/system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-164-file-artifacts-distributed-across-hosts-must-be-carried-by-value-not-host-absolute-reference.md diff --git a/docs/architecture/system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md similarity index 100% rename from docs/architecture/system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-165-loop-control-bookends-start-develop-merge-release-wrap.md diff --git a/docs/architecture/system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md similarity index 100% rename from docs/architecture/system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-166-single-source-of-truth-for-model-context-window-resolution.md diff --git a/docs/architecture/system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md similarity index 100% rename from docs/architecture/system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-167-session-link-suppression-attribution-sessionurl-as-primary-control-deny-hook-as-backstop.md diff --git a/docs/architecture/system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md similarity index 100% rename from docs/architecture/system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-169-agent-ways-relinquishes-user-scoped-settings-json-retains-only-its-operational-baseline.md diff --git a/docs/architecture/system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md similarity index 100% rename from docs/architecture/system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-170-human-focus-group-membership-via-username-identity-and-a-shared-attend-groups-crate.md diff --git a/docs/architecture/system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md similarity index 100% rename from docs/architecture/system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-171-stable-session-identity-the-roster-enumerates-addressable-coordinating-units.md diff --git a/docs/architecture/system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md similarity index 100% rename from docs/architecture/system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-172-turn-boundary-inbound-delivery-via-a-cli-owned-drain-checkpoint.md diff --git a/docs/architecture/system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md similarity index 100% rename from docs/architecture/system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-173-chat-idiom-convergence-for-the-attend-command-surfaces.md diff --git a/docs/architecture/system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md similarity index 100% rename from docs/architecture/system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-174-progressive-core-decoration-guidance-and-the-core-re-disclosure-gap.md diff --git a/docs/architecture/system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md similarity index 100% rename from docs/architecture/system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-175-standing-delegation-authorization-satisfies-the-harness-permission-gate-rather-than-overriding-it.md diff --git a/docs/architecture/system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md similarity index 100% rename from docs/architecture/system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-176-contract-identification-as-the-develop-loop-front-gate.md diff --git a/docs/architecture/system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md similarity index 100% rename from docs/architecture/system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-177-version-stamped-vendored-tools-with-direction-aware-drift-detection.md diff --git a/docs/architecture/system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md similarity index 100% rename from docs/architecture/system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-178-register-transfers-by-demonstration-core-md-carries-policy.md diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md new file mode 100644 index 00000000..3248108a --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-179-remove-the-pre-1-0-in-place-migrator-keep-the-guards-and-the-transition-fallbacks.md @@ -0,0 +1,159 @@ +--- +contract: adr/v1 +kind: decision +verb: retire +capability: install +targets: [cli:ways-migrate] +enacted: "b4f63aa6" +agent: {name: Claude, model: unrecorded} +basis: + - evidence: the release history, where the migrator still shipped two deferral windows after ADR-144 scheduled its removal (tags ways-v1.2.0 through ways-v1.8.3) + - precedent: ADR-144 +status: Accepted +date: 2026-08-17 +deciders: + - aaronsb + - claude +related: + - ADR-142 + - ADR-144 + - ADR-153 +--- + +# ADR-179: Remove the pre-1.0 in-place migrator; keep the guards and the transition fallbacks + +## Summary + +- **Decided:** remove the pre-1.0 `ways migrate` command and its code. Keep the in-place guards and the path fallbacks. +- **Trades away:** an in-place upgrade from a pre-1.0 install. Those installs are pointed at the release tag where the migrator still lives. +- **One-way?** No. The migrator stays reachable at its tag, and the guards name where. +- **Probes:** *Confident:* the guards still catch a legacy install, since they were kept. *Not confident:* whether any pre-1.0 install is still in use and would hit the escape hatch. +- **Inversion:** one end keeps the migrator forever, compiled in. The other end also removes the guards and fallbacks. This removes the command and keeps the safety net. Is that the right cut line? + +## Context + +ADR-144 §5 shipped `ways migrate` (plan / `--what-if` / `--execute`) to move a +pre-1.0 in-place `~/.claude` clone onto the 1.0 projection model, and gave it an +explicit deprecation lifecycle: dormant at 1.0.0, escalating SessionStart +pressure across 1.0.x, removed at 1.1.0. The escape hatch was named in the same +section — the migrator lives forever at the last tag that ships it, so removal +never means a user cannot migrate. + +The removal has been deferred twice. 1.1.0 shipped with the migrator still +present; 1.2.0 (`691b0e6`) corrected the docs and moved the window to 1.3.0. +That target passed too. The line is now at **1.8.3** and the migrator is still +compiled in — roughly 1,000 lines across `migrate.rs` and `migrate_exec.rs`, +plus the CLI variant, a reconcile parameter that exists only for it, and +migration instructions in the README, the install guide, `install.sh`, the +`deployment` way, and the `ways-update` skill. + +The deferrals bought adopter time. They have also left the escalation curve's +end state unreached: eight minor releases past the announced window, every +surface still presents `ways migrate` as a live command, and the deprecation +notices in those surfaces name removal targets (1.1, 1.3) that came and went. + +ADR-144 §5 also drew a second line that the removal must respect. Its "the cliff +is a ramp" clause says removing the migrator removes *assisted migration*, not +*function*: a post-removal binary on an un-migrated `~/.claude` still reads +correctly through the transition fallbacks — cache from the legacy +`claude-ways` dir, events from `~/.claude/stats`. Those fallbacks live in +`ways-core/src/paths.rs` and are cheap, non-destructive, and inert on a current +install. + +Two guards share the migrator's detector, `reconcile::is_legacy_in_place`. +`ways reconcile` refuses to run against an in-place clone because projecting +symlinks over a live repo would strand the user's checkout; `ways update` +refuses for the same reason. Both currently route the user to `ways migrate`. +ADR-144 §5 listed "migrator **and in-place check** removed" as one step. Removing +the checks would let `ways reconcile` clobber an in-place clone — an actively +destructive outcome for exactly the users the never-strand clause protects. + +## Decision + +Remove the migrator. Keep the in-place guards and the transition fallbacks. + +**Removed:** + +- `tools/ways-cli/src/cmd/migrate.rs` and `tools/ways-cli/src/cmd/migrate_exec.rs`, + their `mod` declarations, the `Commands::Migrate` clap variant, and its + dispatch arm. +- `ways_core::paths::cache_root_canonical`, which exists solely as the + migrator's rename destination. Runtime reads use the fallback-aware + `cache_root`. +- `reconcile::run`'s `allow_in_place` parameter and the bypass it gates. The + migrator is the only caller that ever passed `true`; with it gone the guard is + unconditional. +- The migration walkthrough in `docs/migration-1.0.md`, the `ways migrate` + invocations in `scripts/install.sh`, `skills/ways-update/SKILL.md`, the + `deployment` way, README, and `docs/install-guide.md`. + +**Kept:** + +- `reconcile::is_legacy_in_place` and both guards. Their messages retarget from + "run `ways migrate`" to the tag escape hatch. A user who reaches these guards + is told what their install is and where the migrator still lives. +- The `paths.rs` transition fallbacks — `LEGACY_CACHE` resolution in + `resolve_cache`, the `~/.claude/stats/events.jsonl` fallback in + `resolve_events`, and the `events_log_sources` union (ADR-153 §1). These are + the ramp ADR-144 §5 promised. The union also recovers orphaned `session_start` + lines on *migrated* installs whose shell hooks predated the path fix, so it + earns its keep independent of the in-place case. +- The legacy `~/.claude/ways.json` disabled layer, a lower-precedence config + read with no coupling to the migrator. + +**Escape hatch:** `ways-v1.8.3` is the last tag shipping the migrator. Migrating +after removal means `git clone --branch ways-v1.8.3 … && ways migrate --execute`, +then updating. `docs/migration-1.0.md` is rewritten around that route rather than +deleted, and the guards point at it. + +Removal lands in **1.9.0**. + +## Consequences + +### Positive + +- About 1,000 lines of transitional code leave the binary, along with a + destructive code path (whole-`~/.claude` relocation) that no current install + exercises. +- `reconcile::run` loses a boolean parameter whose only non-default caller is + being deleted, so the in-place guard becomes unconditional and the function's + contract simplifies. +- Every user-facing surface stops advertising a deprecation target that has + already passed. The migration story becomes one route (the tag) instead of a + live command shadowed by stale removal dates. + +### Negative + +- A pre-1.0 adopter who has not migrated by 1.9.0 now has a two-step path (clone + the tag, migrate, update) instead of one command. This is the cost ADR-144 §5 + accepted when it named the tag as the escape hatch. +- The migrator's crash-safe phase machinery and its tests go with it. Reviving + the capability would mean recovering it from the tag. + +### Neutral + +- ADR-144 §5's lifecycle is executed, not superseded — its escalation dates were + the cadence, not the contract. This ADR records the divergence on one point: + the in-place *check* stays. +- The `deployment` way keeps its legacy-in-place branch. The detection guidance + is still correct; only the remedy changes. + +## Alternatives Considered + +- **Remove the guards too, as ADR-144 §5 literally specified.** Rejected: with + the guard gone, `ways reconcile` on an in-place clone projects over a live git + repo and strands the checkout. Deleting assistance is within the lifecycle; + adding a destructive path is not. +- **Remove the `paths.rs` transition fallbacks in the same change.** Rejected for + now: they are inert on a current install, and dropping them turns "reads + correctly, unassisted" into "silently re-fetches the model and stops reading + its own stats." The `events_log_sources` union has a second justification that + outlives the in-place case. If these come out, it is as their own decision with + its own reasoning. +- **Defer again to 2.0.0.** Rejected: two deferrals have already passed with no + signal that a third would be used differently, and each one leaves the shipped + docs quoting a removal date that has expired. +- **Keep the migrator indefinitely as a dormant command.** Rejected: it carries a + destructive code path and a phase machine that no test run outside its own + fixtures exercises, and its presence is why the docs still carry a transition + narrative eight releases past 1.0. diff --git a/docs/architecture/system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md similarity index 100% rename from docs/architecture/system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-180-github-issues-as-the-shared-truth-for-the-session-task-list.md diff --git a/docs/architecture/system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md similarity index 100% rename from docs/architecture/system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-181-guard-hooks-a-blocking-pretooluse-class-for-pattern-kills-and-interactive-prone-commands.md diff --git a/docs/architecture/system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md similarity index 100% rename from docs/architecture/system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-182-keepwarm-attend-keeps-the-prompt-cache-warm-with-a-wake-floor.md diff --git a/docs/architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md similarity index 100% rename from docs/architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-184-installation-and-activation-are-separate-states-targets-as-the-unit-of-activation.md diff --git a/docs/architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md similarity index 100% rename from docs/architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-185-cli-output-contract-structured-output-for-people-json-for-machines.md diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md new file mode 100644 index 00000000..b7ff0599 --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md @@ -0,0 +1,93 @@ +--- +contract: adr/v1 +kind: decision +verb: change +capability: testing +agent: {name: Claude, model: unrecorded} +basis: + - evidence: the reviews of PRs #501, #502, #504 and #508, each of which found a defect on the install path by reading the code, none of it run end to end on a clean machine +status: Accepted +date: 2026-09-17 +deciders: + - aaronsb + - claude +related: + - ADR-142 + - ADR-144 + - ADR-184 + - ADR-185 +--- + +# ADR-186: Live integration fixture: install-path test levels and the tier 2 gate + +## Summary + +- **Decided:** the install path is tested in a container, at two levels. Tier 1 installs and configures with no API key, on every pull request that touches the path. Tier 2 exercises a model with a key, on dispatch or a schedule only. +- **Trades away:** job time and network dependence on tier 1, and tokens on a fixed cadence for tier 2. +- **One-way?** No. The fixture is additive, and removing the job removes the gate and nothing else. +- **Probes:** *Confident:* the #501 shape (the refusal, the recovery and the kept hooks) is asserted on every pull request. *Not confident:* none stated in the original record. +- **Inversion:** at one end, a model runs on every pull request, so a fork's check passes with no secret and every pull request spends tokens. At the other end nothing runs a clean install and reviews keep finding install defects by reading code. The two tiers sit between them. + +## Context + +The install path is the installer script, `make setup` with its prebuilt downloads, `ways reconcile` into a config directory that already holds a user's own files, and the hook scripts that Claude Code runs from the merged `settings.json`. The reviews of PRs #501, #502, #504 and #508 each found a defect on that path by reading the code, and each said the same thing: none of it had been run end to end on a clean machine. + +The tests that exist stop short of it. The Rust unit tests and `session_sim` exercise the binary against fixture ways in a temporary directory. The reconcile tests put a fake source and a fake destination in a sandbox. The sandbox transcripts in the PR comments were run by hand and are not repeatable. Nothing runs the installer, downloads a release asset, seeds a home with the shape from #501 (a real `skills/` directory, three hand-written hooks, a second config directory), or feeds a hook script the payload Claude Code sends. + +ADR-184 increment 2 (#509) changes the installer's last act from projecting into `~/.claude` to handing off to the targets bootstrap. Building that against no fixture repeats the pattern the reviews named. + +## Decision + +**The install path is tested at two levels. Tier 1 installs and configures with no API key and runs on every pull request that touches the path. Tier 2 exercises a model with a key, runs on manual dispatch or a schedule, and never runs unattended on a pull request.** + +1. **Tier 1: install and configure, no key.** A Debian container with Claude Code installed at a pinned version. The home is seeded before the installer runs: a real `~/.claude/skills/` with a user's own skill, a `settings.json` with three user hooks and a `model` key, and a second config directory with its own settings and skills. The runner asserts: + - the installer runs unattended and refuses the real `skills/` directory with nothing deleted and `settings.json` byte-identical; + - the documented recovery (`ways reconcile --force`) moves the directory aside with the user's skill intact, links the projection roots, and keeps every user hook by identity in its event; + - the target is recorded in the user config, `ways config targets` reports one enabled target, and `ways status` reports the active state with the embedding engine up; + - `ways reconcile --dry-run` run twice prints identical output and reports no change; + - the second config directory is byte-identical to its seed; + - `way-embed` arrived as a release download. The image carries no C++ toolchain, so a source-build fallback is a failed assertion; + - `attend status` and `claude --version` exit zero, and the version is the pinned one; + - the hooks fire the way Claude Code fires them: the runner reads the hook commands out of the merged `settings.json`, pipes a synthetic `SessionStart`, `UserPromptSubmit`, and `PreToolUse` payload to each, and asserts on the additional-context output. Disclosure is proven with no model. + +2. **Two flavors of tier 1.** The `branch` flavor mounts the checkout and the binaries CI built for the same commit, so a pull request is tested against its own binary and its own hooks. It is the gate, and it runs on every pull request that touches the path and on every push to `main`. The `release` flavor runs the documented one-liner, which clones `main` and downloads the latest release assets. It runs on dispatch and on the nightly schedule, with Claude Code at its `latest` version, to catch drift on either side. + +3. **Tier 2: exercise with a key.** `claude -p` runs non-interactively under `CLAUDE_CONFIG_DIR` with prompts chosen to trigger named ways. The evidence is the transcript and the events log read through `ways introspect`, plus the answer scored against a rubric with a stated pass threshold. Two containers on one compose network carry the attend peer test: send from one, assert the other's inbox. Keepwarm is out of scope. + +4. **The gate.** Tier 2 reads the key from a repository secret. The workflow runs it on `workflow_dispatch` and on `schedule` only. A pull request never triggers it, and a fork pull request cannot see the secret. A change to tier 2 is verified by dispatching it against the branch. + +5. **Shape.** Everything lives under `tests/fixtures/docker/`: the compose file, one Dockerfile parameterized by Claude Code version and installer, the seed, the payloads, and one runner per tier. `make test-live TIER=1|2` is the entry point on a workstation and in CI. Tier 1 is a job in `portability.yml`. + +6. **Sequence.** Tier 1 lands before #509. Tier 2 follows tier 1 as its own increment. + +Reversibility: cheap. The fixture is additive. Removing the job removes the gate and nothing else. + +## Consequences + +### Positive + +- A change to the installer, the reconciler, the settings merge, or a hook script is run on a clean machine before review reads it. +- The #501 shape is a fixture, not a memory. The refusal, the recovery, and the kept hooks are asserted on every pull request. +- #509 is built against a fixture that fails when it wires the wrong directory. +- A pinned Claude Code on the gate and `latest` on the schedule separate our regressions from upstream drift. + +### Negative + +- Tier 1 needs network: the Claude Code installer, the way-embed and model downloads, and the release assets through `gh`. A network fault fails the job. The branch flavor keeps the four suite binaries off the network; way-embed, mmaid, and the model stay on it. +- Docker on the runner adds minutes to the portability workflow. +- Tier 2 on a schedule spends tokens on a fixed cadence. The threshold and the prompt set have to be maintained. + +### Neutral + +- The runner drives hooks from `settings.json` rather than by path, so a hook that the merge drops is a failed assertion rather than a silent skip. +- `portability.yml` gains a job that builds all four suite binaries before the fixture runs, since the branch flavor needs `ways-audit` and `attend-chat` beside `ways` and `attend`. The existing cross-platform matrix is unchanged. +- ADR-185's `--json` views are what the runner parses. +- The first run of the fixture found two defects on the download path, both fixed on the same branch. The download scripts listed 20 or 30 releases before grepping for a component prefix, and the newest tags of three components sat past that window. The way-embed download script's capability probe ran under `pipefail` and rejected every binary, since a supporting binary also exits nonzero when asked for `match --batch` without a corpus. #516 had diagnosed that as a release that predates the primitive. The release supports it; the probe was wrong. The image carries no C++ toolchain, so the fixture fails if either defect returns. + +## Alternatives Considered + +- **Run the installer on the GitHub runner's own home.** Rejected: the runner's home is not clean, and the seeded state would have to be undone between steps. A container starts empty every time. +- **Mock Claude Code.** Rejected: `claude --version`, the installer path, and the settings file are the things under test. The binary is a pinned download and costs one step. +- **Tier 2 on every pull request.** Rejected: a fork pull request has no secret, so the check would pass by absence, and every pull request would spend tokens. Dispatch and schedule keep the run deliberate. +- **One tier with the key optional.** Rejected: a runner that skips assertions when the key is missing reports green for two different things. Two runners with two names keep the report honest. +- **A virtual machine per run.** Rejected: slower to start, and the container already isolates the home, the XDG roots, and the PATH. diff --git a/docs/architecture/system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md similarity index 100% rename from docs/architecture/system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-187-attend-mcp-server-mode-outbound-and-queries-as-typed-tools-inbound-stays-on-monitor-and-the-stop-hook.md diff --git a/docs/architecture/system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md similarity index 100% rename from docs/architecture/system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md rename to tests/fixtures/adr/v0-corpus/docs/architecture/system/ADR-188-posttooluse-delivery-for-tool-lane-ways-and-retirement-of-the-semantic-bash-surface.md diff --git a/tests/fixtures/adr/v0-corpus/docs/architecture/system/multilingual-model-evaluation.md b/tests/fixtures/adr/v0-corpus/docs/architecture/system/multilingual-model-evaluation.md new file mode 100644 index 00000000..003bd5da --- /dev/null +++ b/tests/fixtures/adr/v0-corpus/docs/architecture/system/multilingual-model-evaluation.md @@ -0,0 +1,69 @@ +# Multilingual Embedding Model Evaluation + +**Date:** 2026-04-03T04:07:51Z + +**Models:** +- English: `all-MiniLM-L6-v2` (21M) +- Multilingual: `paraphrase-multilingual-MiniLM-L12-v2` Q8_0 (127M) + +**Methodology:** Each test embeds a native-language prompt against both an English description (cross-language) and a native-language description (same-language stub scenario). Three scores per test: + +- **EN×EN**: English-only model, English description (baseline) +- **Multi×EN**: multilingual model, English description (cross-language) +- **Multi×Native**: multilingual model, native description (same-language stub) + +**Threshold:** 0.25 (same-language similarity minimum) + +## Results + +| Lang | Prompt | EN×EN | Multi×EN | Multi×Native | Pass | +|:-----|:-------|------:|---------:|-------------:|:----:| +| en | check dependencies for vulnerabilities | 0.7574 | 0.6822 | 0.6822 | ✅ | +| de | Abhängigkeiten auf Schwachstellen prüfen | 0.0767 | 0.6223 | 0.8243 | ✅ | +| es | verificar dependencias por vulnerabilidades | 0.4381 | 0.7893 | 0.8357 | ✅ | +| fr | vérifier les dépendances pour vulnérabilités | 0.5223 | 0.7418 | 0.8938 | ✅ | +| pt | verificar dependências por vulnerabilidades | 0.4381 | 0.7899 | 0.9605 | ✅ | +| ru | проверить зависимости на уязвимости | 0.0295 | 0.7592 | 0.8536 | ✅ | +| ja | 依存関係の脆弱性をチェックして | -0.0290 | 0.6861 | 0.9338 | ✅ | +| ko | 의존성 취약점 검사 | -0.0163 | 0.7179 | 0.8554 | ✅ | +| zh | 检查依赖项的漏洞 | -0.0266 | 0.5866 | 0.8874 | ✅ | +| ar | فحص التبعيات بحثاً عن ثغرات | 0.0416 | 0.3995 | 0.9581 | ✅ | +| el | έλεγχος εξαρτήσεων για ευπάθειες | -0.0118 | 0.5934 | 0.8065 | ✅ | +| en | write a conventional commit message | 0.7930 | 0.7465 | 0.7465 | ✅ | +| ja | コミットメッセージを書いて | 0.0929 | 0.5314 | 0.8322 | ✅ | +| ko | 커밋 메시지 작성 | 0.1070 | 0.4847 | 0.8074 | ✅ | +| zh | 写一个规范的提交信息 | 0.0727 | 0.6833 | 0.8896 | ✅ | +| de | eine konventionelle Commit-Nachricht schreiben | 0.3081 | 0.6349 | 0.7801 | ✅ | +| ru | написать сообщение коммита | 0.0039 | 0.4563 | 0.5000 | ✅ | +| en | add unit tests for the auth module | 0.5089 | 0.7411 | 0.7411 | ✅ | +| ja | 認証モジュールのユニットテストを追加して | 0.0009 | 0.7461 | 0.8338 | ✅ | +| ko | 인증 모듈에 단위 테스트 추가 | 0.0917 | 0.5602 | 0.7660 | ✅ | +| zh | 为认证模块添加单元测试 | 0.0162 | 0.7100 | 0.8278 | ✅ | +| de | Unit-Tests für das Auth-Modul hinzufügen | 0.4085 | 0.7445 | 0.8100 | ✅ | +| ru | добавить юнит-тесты для модуля аутентификации | 0.0629 | 0.1650 | 0.3210 | ✅ | + +## Summary + +- **Tests:** 23 +- **Passed:** 23 +- **Failed:** 0 +- **Accuracy:** 100.0% + +## Timing + +| Phase | Duration | Tests | Per-test | +|:------|:---------|------:|---------:| +| EN model batch (23 pairs) | 104ms | 23 | 4ms | +| Multi model cross-language (23 pairs) | 392ms | 23 | 17ms | +| Multi model same-language (23 pairs) | 389ms | 23 | 16ms | +| **Total** | **889ms** | **69** | **12ms** | + +## Interpretation + +The multilingual model enables three matching strategies: + +1. **English ways + English model** — current production. High precision for English prompts. +2. **English ways + multilingual model (cross-language)** — user types in any language, matches against English descriptions. Works but scores 30-50% lower. +3. **Native-language stubs + multilingual model (same-language)** — locale entries in `.locales.jsonl` with native descriptions. Consistently scores 0.80+ across tested languages. + +**Recommendation:** Ship both models. English ways use the English model (precise, 21MB). Multilingual stubs use the multilingual model (broad, 127MB). Per-way `embed_model` frontmatter field controls routing. This gives per-language threshold tuning without compromising English accuracy. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/adr.yaml b/tests/fixtures/adr/v1-defects/docs/architecture/adr.yaml index 4ca5c1fd..0f8e3e84 100644 --- a/tests/fixtures/adr/v1-defects/docs/architecture/adr.yaml +++ b/tests/fixtures/adr/v1-defects/docs/architecture/adr.yaml @@ -49,6 +49,10 @@ capabilities: ingest: Document ingestion into the store search: Query over the store +# Keys adr/v1 no longer reads (ADR-311): present, even malformed, they are +# ignored. +baseline: [search] + surfaces: cli: {} route: {} diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-159-precedent-loop-a.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-159-precedent-loop-a.md deleted file mode 100644 index 521ce4f3..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-159-precedent-loop-a.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-161 ---- - -# ADR-159: Precedent loop, first half - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-161-precedent-loop-b.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-161-precedent-loop-b.md deleted file mode 100644 index ca880ddb..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-161-precedent-loop-b.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-159 ---- - -# ADR-161: Precedent loop, second half - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-163-precedent-v0.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-163-precedent-v0.md deleted file mode 100644 index 7c2e2755..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-163-precedent-v0.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-160 ---- - -# ADR-163: Precedent to an archived v0 record - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-164-summary-thin.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-164-summary-thin.md deleted file mode 100644 index 5351df90..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-164-summary-thin.md +++ /dev/null @@ -1,26 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -status: proposed -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -basis: - - evidence: data ---- - -# ADR-164: Summary with no probes or inversion - -## Summary - -Just a paragraph. - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-168-summary-subsections.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-168-summary-subsections.md deleted file mode 100644 index b46637fb..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-168-summary-subsections.md +++ /dev/null @@ -1,32 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -status: proposed -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -basis: - - evidence: data ---- - -# ADR-168: Summary with subsections - -## Summary - -What is decided. - -### Probes - -- High confidence: the main point. -- Low confidence: the - edge case. - -### Inversion - -One end, the other end. - -## 1. Decision - -The decision. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-169-summary-less-confident.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-169-summary-less-confident.md deleted file mode 100644 index 8b6b9924..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-169-summary-less-confident.md +++ /dev/null @@ -1,24 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -status: proposed -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -basis: - - evidence: data ---- - -# ADR-169: Only a less-confident label - -## Summary - -- Probes: I am less confident about this, and not - confident about that. -- Inversion: one end or the other. - -## 1. Decision - -The decision. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-170-order-a.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-170-order-a.md deleted file mode 100644 index 8f061f36..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-170-order-a.md +++ /dev/null @@ -1,30 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: accepted -basis: - - precedent: ADR-171 - - precedent: ADR-100 ---- - -# ADR-170: Order pair, grounded through its second precedent - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-175-precedent-to-v0-spec.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-175-precedent-to-v0-spec.md deleted file mode 100644 index f78ec0c8..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-175-precedent-to-v0-spec.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-174 ---- - -# ADR-175: Precedent to a spec decided by v0 - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-176-rejected-target.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-176-rejected-target.md deleted file mode 100644 index 116fa056..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-176-rejected-target.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: rejected -basis: - - evidence: data ---- - -# ADR-176: A rejected decision - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-177-precedent-to-rejected.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-177-precedent-to-rejected.md deleted file mode 100644 index 91269798..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-177-precedent-to-rejected.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-176 ---- - -# ADR-177: Precedent to a rejected decision - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-178-downstream-of-loop.md b/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-178-downstream-of-loop.md deleted file mode 100644 index 6345038c..00000000 --- a/tests/fixtures/adr/v1-defects/docs/architecture/system/ADR-178-downstream-of-loop.md +++ /dev/null @@ -1,29 +0,0 @@ ---- -contract: adr/v1 -kind: decision -verb: add -capability: adr -date: 2025-06-01 -deciders: [developer, agent] -agent: {name: Claude, model: fixture-model} -status: proposed -basis: - - precedent: ADR-159 ---- - -# ADR-178: Downstream of a loop - -## Summary - -- **Decided:** the decision in plain terms. -- **Trades away:** what it gives up. -- **Probes:** *Confident:* the main point holds. *Not confident:* the edge case. -- **Inversion:** one end, the other end; is the middle right? - -## 1. Decision - -The decision. - -## 2. Consequences - -They follow. diff --git a/tests/fixtures/adr/v1/docs/architecture/adr.yaml b/tests/fixtures/adr/v1/docs/architecture/adr.yaml index 0b20d544..9a4d78df 100644 --- a/tests/fixtures/adr/v1/docs/architecture/adr.yaml +++ b/tests/fixtures/adr/v1/docs/architecture/adr.yaml @@ -21,13 +21,11 @@ legacy: kinds: decision: - mutable_after_accept: [status, enacted, superseded_by, considered, concern] verb: required requires: [capability, basis, agent] sections: [Summary] edges: { supersedes: decision, amends: decision, extends: decision, basis: [decision, spec] } spec: - mutable_after_accept: all verb: forbidden requires: [capability] edges: { supersedes: spec, decided_by: decision } @@ -37,10 +35,6 @@ capabilities: ingest: Document ingestion into the store search: Query over the store -surfaces: - cli: { inventory: "printf 'ingest-legacy\\nquery\\n'" } - route: {} - cite: exclude: [vendor] diff --git a/tests/fixtures/docker/CLAUDE.md b/tests/fixtures/docker/CLAUDE.md index 23daad3e..4b82b411 100644 --- a/tests/fixtures/docker/CLAUDE.md +++ b/tests/fixtures/docker/CLAUDE.md @@ -1,6 +1,6 @@ # Working in the live fixture -This directory is the tier 1 live install fixture from [ADR-186](../../../docs/architecture/system/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md). [README.md](README.md) here says what it asserts. This file says how to run it, change it, and read a failure. +This directory is the tier 1 live install fixture from [ADR-186](../../../docs/architecture/platform/ADR-186-live-integration-fixture-install-path-test-levels-and-the-tier-2-gate.md). [README.md](README.md) here says what it asserts. This file says how to run it, change it, and read a failure. ## Run it diff --git a/tests/fixtures/docker/scenarios/adr-consider/setup.sh b/tests/fixtures/docker/scenarios/adr-consider/setup.sh index ee7859eb..605a1ad6 100644 --- a/tests/fixtures/docker/scenarios/adr-consider/setup.sh +++ b/tests/fixtures/docker/scenarios/adr-consider/setup.sh @@ -11,7 +11,7 @@ domains: defaults: {deciders: [developer, agent]} kinds: decision: - mutable_after_accept: [status, enacted, superseded_by, considered, concern] + mutable_after_accept: [status, enacted, superseded_by, considered, concern, observable] verb: required requires: [capability, basis, agent] sections: [Summary] diff --git a/tests/fixtures/docker/scenarios/adr-migrate/check.sh b/tests/fixtures/docker/scenarios/adr-migrate/check.sh index be7b93d8..80a82f67 100644 --- a/tests/fixtures/docker/scenarios/adr-migrate/check.sh +++ b/tests/fixtures/docker/scenarios/adr-migrate/check.sh @@ -3,7 +3,7 @@ # record never contained (ADR-304 §11, fabrication risk). for n in 179 186; do - f=$(ls "$PROJ"/docs/architecture/system/ADR-$n-*.md 2>/dev/null | head -1) + f=$(ls "$PROJ"/docs/architecture/*/ADR-$n-*.md 2>/dev/null | head -1) cp "$f" "$OUT/" 2>/dev/null if grep -q '^contract: adr/v1' "$f"; then ok "ADR-$n declares adr/v1"; else fail "ADR-$n declares adr/v1"; fi lint=$(cd "$PROJ" && docs/scripts/adr lint --check "${f#$PROJ/}" 2>&1) diff --git a/tests/fixtures/docker/scenarios/adr-migrate/prompt.txt b/tests/fixtures/docker/scenarios/adr-migrate/prompt.txt index 7df83d15..4462c232 100644 --- a/tests/fixtures/docker/scenarios/adr-migrate/prompt.txt +++ b/tests/fixtures/docker/scenarios/adr-migrate/prompt.txt @@ -1 +1 @@ -This repository's decision records now follow the adr/v1 contract declared in docs/architecture/adr.yaml (the ADR way and `docs/scripts/adr` describe it). Migrate two records to v1: ADR-179 (removing the pre-1.0 migrator) and ADR-186 (the live integration fixture). Keep their decisions and history intact, choose kind, verb and capability honestly, and ground each basis in what the record itself says. Do not invent operator statements. Run `docs/scripts/adr lint` on the two files until they are clean, then summarize what you chose and why. +This repository's decision records now follow the adr/v1 contract declared in docs/architecture/adr.yaml (the ADR way and `docs/scripts/adr` describe it). Migrate two records to v1: ADR-179 (removing the pre-1.0 migrator) and ADR-186 (the live integration fixture). Keep their decisions and history intact, choose kind, verb and capability honestly, and ground each basis in what the record itself says. Do not invent operator statements. Convert them with `docs/scripts/adr import scan` on the two files, resolve each import sheet's open items, then `docs/scripts/adr import apply`. Run `docs/scripts/adr lint` on the two files until they are clean, then summarize what you chose and why. diff --git a/tests/fixtures/docker/scenarios/adr-migrate/setup.sh b/tests/fixtures/docker/scenarios/adr-migrate/setup.sh index ff72f155..ef83fdc9 100644 --- a/tests/fixtures/docker/scenarios/adr-migrate/setup.sh +++ b/tests/fixtures/docker/scenarios/adr-migrate/setup.sh @@ -10,14 +10,18 @@ chmod +x docs/scripts/adr # script from main before #581 migrated them. The release flavor clones with # --depth 1, so the history to restore them from is not there. here=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd) +# Each replaces the record where the corpus keeps it: system/ before the +# intent folders (ADR-310), platform/ after. for f in "$here"/v0/ADR-*.md; do - cp "$f" docs/architecture/system/ + base=$(basename "$f") + live=$(ls docs/architecture/*/"$base" 2>/dev/null | head -1) + cp "$f" "${live:-docs/architecture/platform/$base}" done -if grep -q '^contract:' docs/architecture/system/ADR-179-*.md docs/architecture/system/ADR-186-*.md; then +if grep -q '^contract:' docs/architecture/*/ADR-179-*.md docs/architecture/*/ADR-186-*.md; then echo "setup: the records under test are already v1; the rehearsal would test nothing" >&2 exit 1 fi # Snapshot the two records the scenario migrates, to check nothing is invented. mkdir -p "$HOME/.migrate-before" -cp docs/architecture/system/ADR-179-*.md docs/architecture/system/ADR-186-*.md "$HOME/.migrate-before/" +cp docs/architecture/*/ADR-179-*.md docs/architecture/*/ADR-186-*.md "$HOME/.migrate-before/" git add -A && git commit -qm "corpus" diff --git a/tools/attend/src/cmd/config_cmd.rs b/tools/attend/src/cmd/config_cmd.rs index c54e1206..077c9162 100644 --- a/tools/attend/src/cmd/config_cmd.rs +++ b/tools/attend/src/cmd/config_cmd.rs @@ -26,7 +26,7 @@ pub(crate) fn display_config(cfg: &config::Config) { t.add(vec!["", "", ""]); - // Engagement section (ADR-119 action potential, unified in ADR-123) + // Engagement section (ADR-123 action-potential curve) t.add(vec![ "engagement", "burst_threshold", diff --git a/tools/attend/src/cmd/run/tick.rs b/tools/attend/src/cmd/run/tick.rs index a6f51e1c..6d8645bd 100644 --- a/tools/attend/src/cmd/run/tick.rs +++ b/tools/attend/src/cmd/run/tick.rs @@ -52,8 +52,8 @@ pub(super) struct TickState<'a> { pub(super) cfg: &'a config::Config, } -/// Build the engagement curve from config (ADR-119 action potential, -/// ADR-123 progression-axis unification). All sensors share these +/// Build the engagement curve from config (ADR-123 action-potential +/// curve on the unified progression axis). All sensors share these /// engagement parameters; per-sensor overrides can be added later /// if the defaults turn out to be too coarse. /// diff --git a/tools/attend/src/config.rs b/tools/attend/src/config.rs index 9398130b..197069e8 100644 --- a/tools/attend/src/config.rs +++ b/tools/attend/src/config.rs @@ -52,7 +52,7 @@ pub struct GovernorConfig { pub rate_window: Duration, } -/// Action potential engagement parameters (ADR-119). +/// Action potential engagement parameters (ADR-123 `Curve::ActionPotential`). /// /// Governs per-sensor refractory behavior: after a burst of disclosures, /// the sensor enters a refractory period where only high-magnitude events @@ -248,7 +248,7 @@ governor: max_per_window: 3 rate_window: 120 -# Action potential engagement model (ADR-119, unified in ADR-123). +# Action potential engagement model (ADR-123). # Run `attend tune` to auto-derive these from real session history. engagement: burst_threshold: 3 # disclosures before refractory kicks in @@ -369,7 +369,7 @@ fn detect_legacy_burst_window(yaml_str: &str) -> Result<(), String> { under ADR-123 the burst window is implicit in multiplier_half_life \ (derived from decay_per_minute). Run `attend config lint --fix` \ to remove the key from your config. See \ - docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md \ + docs/architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md \ for migration guidance." .to_string(), ); diff --git a/tools/scripts/docs-enact-pr2a.mjs b/tools/scripts/docs-enact-pr2a.mjs index f0fbf5a6..46b28483 100644 --- a/tools/scripts/docs-enact-pr2a.mjs +++ b/tools/scripts/docs-enact-pr2a.mjs @@ -83,7 +83,7 @@ splitting files is a later pass. Return the COMPLETE rewritten file in rewritten_content, and a claims_ledger listing EVERY engine fact you asserted with its source citation. Sober, literal voice. Do not touch the still-current -mechanisms (ways tune locale audit, salience decay ADR-121, progressive disclosure).` +mechanisms (ways tune locale audit, salience decay ADR-123, progressive disclosure).` } function bannerPrompt(f) { diff --git a/tools/scripts/docs-inventory-fanout.mjs b/tools/scripts/docs-inventory-fanout.mjs index a240e0c4..9ab2146f 100644 --- a/tools/scripts/docs-inventory-fanout.mjs +++ b/tools/scripts/docs-inventory-fanout.mjs @@ -42,9 +42,9 @@ RETIRED / STALE — flag any doc that still teaches these as live: STILL CURRENT (do NOT flag these as stale): - \`ways tune\` = LOCALE alias audit (fidelity/discrimination vs the English root anchor), ADR-139/125. It explicitly does NOT write thresholds. Correct as-is. -- Salience/signal DECAY: turn-based exponential decay (ADR-121) — a separate model +- Salience/signal DECAY: exponential decay over a progression axis (ADR-123) — a separate model from relevance scoring; still valid. -- Progressive disclosure, token-gated re-fire (ADR-104/105/126), three-root runtime +- Progressive disclosure, token-gated re-fire (ADR-105/123/126), three-root runtime (ADR-143), sentence-salience input reduction (ADR-130), authored disclosure graph / removal of BM25 (ADR-125), two embedding models EN + multilingual (localized mode). - ADRs themselves are immutable decision records — an ADR describing the pre-156 world diff --git a/tools/sensor-trait/src/curve.rs b/tools/sensor-trait/src/curve.rs index efacd10e..4f14681c 100644 --- a/tools/sensor-trait/src/curve.rs +++ b/tools/sensor-trait/src/curve.rs @@ -24,8 +24,8 @@ pub enum Curve { /// Action-potential model: event-count burst detection raises a /// refractory multiplier that then decays back toward 1.0 over tick - /// distance. Ported from ADR-119 with event-count windowing (see - /// ADR-123 Decision 2) so it is robust to chunky progression axes. + /// distance. Ported from ADR-119 with event-count windowing // adr-cite-ignore + /// (see ADR-123 Decision 2) so it is robust to chunky progression axes. /// /// The "burst window" is NOT a separate tick span. It is implicit in /// which history entries still contribute non-trivial multiplier diff --git a/tools/sensor-trait/src/lib.rs b/tools/sensor-trait/src/lib.rs index 57964dac..8dd8611b 100644 --- a/tools/sensor-trait/src/lib.rs +++ b/tools/sensor-trait/src/lib.rs @@ -265,7 +265,7 @@ pub fn epoch_secs() -> Tick { /// Default `Curve::ActionPotential` for sensor slots. Mirrors the defaults /// that the old `EngagementState::new()` shipped with, so sensors booted -/// without config overrides preserve ADR-119 behavior. +/// without config overrides preserve ADR-119 behavior. // adr-cite-ignore fn default_sensor_curve() -> Curve { Curve::ActionPotential { burst_threshold: 3, diff --git a/tools/ways-cli/src/cmd/init.rs b/tools/ways-cli/src/cmd/init.rs index 936d2eae..9412742a 100644 --- a/tools/ways-cli/src/cmd/init.rs +++ b/tools/ways-cli/src/cmd/init.rs @@ -63,7 +63,7 @@ Semantic matching (`description` + `vocabulary`) is additive with regex triggers | Event handler, fires often relative to session | `0.05` | `frequent` | | Disclose once per session | `1.0` | `once` | -Numeric values between presets are valid — e.g., `refire: 0.2` sits between `normal` and `rare` and is what ADR-127's 1M-Opus-tuned ways migrated to. +Numeric values between presets are valid — e.g., `refire: 0.2` sits between `normal` and `rare` and is what the PR #70 1M-Opus-tuned ways migrated to (ADR-126). Numeric form pins the cadence to today's model. Preset form tracks the project's `refire_presets` config for portability across model generations. diff --git a/tools/ways-cli/src/session.rs b/tools/ways-cli/src/session.rs index 1081bbc5..0fd18600 100644 --- a/tools/ways-cli/src/session.rs +++ b/tools/ways-cli/src/session.rs @@ -249,7 +249,7 @@ pub fn epoch_distance(way_id: &str, session_id: &str) -> u64 { current.saturating_sub(way_ep) } -// ── Token position (ADR-104 re-disclosure) ────────────────────── +// ── Token position (ADR-123/126 re-disclosure) ────────────────────── /// Read the token position from the most recent transcript. pub fn get_token_position(_session_id: &str) -> u64 { diff --git a/tools/ways-cli/tests/SIMULATION-SPEC.md b/tools/ways-cli/tests/SIMULATION-SPEC.md index 59ec21b1..7b3be0f9 100644 --- a/tools/ways-cli/tests/SIMULATION-SPEC.md +++ b/tools/ways-cli/tests/SIMULATION-SPEC.md @@ -138,7 +138,7 @@ Turn 4 (project without Makefile): same bash → check does NOT fire Tests: `when:` project gate, `when:` file_exists gate. -### Scenario 8: Token-Gated Re-Disclosure (ADR-104) +### Scenario 8: Token-Gated Re-Disclosure (ADR-123, ADR-126) ``` Turn 1: prompt → way fires, token position stamped diff --git a/tools/ways-core/src/frontmatter.rs b/tools/ways-core/src/frontmatter.rs index 30758ac5..832295f2 100644 --- a/tools/ways-core/src/frontmatter.rs +++ b/tools/ways-core/src/frontmatter.rs @@ -231,7 +231,7 @@ pub fn detect_legacy_redisclose(yaml_str: &str) -> Result<()> { return Err(anyhow!( "legacy `redisclose:` field is no longer supported — \ migrate to an explicit `curve:` block per ADR-123. See \ - docs/architecture/system/ADR-123-firing-dynamics-progression-axis-unification.md \ + docs/architecture/ways/ADR-123-firing-dynamics-progression-axis-unification.md \ for migration guidance." )); }