diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md
index 58faff50d..8b3dcf633 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md
@@ -1,30 +1,33 @@
-# Chunk Map: WS-QUAL-001 Backend Coverage Floor
+# Chunk Map: WS-QUAL-001 Behavior And Mutation Assurance
-## Historical completed work
+## Completed and superseded work
| Chunk | Durable outcome | State |
|---|---|---|
-| `WS-QUAL-001-PLAN` | Original 90-percent initiative plan | Merged PR #99; superseded by PLAN2 sequencing |
-| `WS-QUAL-001-01` | Original combined harness/baseline contract | Superseded by the 01A/01B split before implementation |
+| `WS-QUAL-001-PLAN` | Original coverage initiative plan | Merged PR #99; superseded |
| `WS-QUAL-001-01A` | Isolated least-privilege database runner | Merged PR #103 |
| `WS-QUAL-001-01B1A-R2` | Coverage configuration/evidence grammar | Merged PR #105 |
-| `WS-QUAL-001-01B1B-R10` | Conservative test-weakening semantic guard | Merged PR #108 |
+| `WS-QUAL-001-01B1B-R10` | Conservative test-weakening guard | Merged PR #108 |
+| `WS-QUAL-001-PLAN2` | Current-main coverage closure plan | Merged PR #260; succeeded by PLAN3 |
+| `WS-QUAL-001-02R` | Project/setup observable behavior coverage | Merged PR #265 |
+| `WS-QUAL-001-03R` | Checker observable behavior coverage | Merged PR #269; main at 90.316651% |
+| `WS-QUAL-001-04R` | Raise global floor from 78 to 90 | Superseded before implementation; 78 retained by human decision |
-All other 01B/01B1/01B1A/01B1B replacement attempts are stopped historical
-experiments. Do not resume them. `WS-QUAL-001-01B2` and the old 02-06 milestone
-ladder are superseded before implementation.
+All other old 01B/01B1 replacement attempts and the old 02-06 milestone ladder
+remain stopped historical experiments. Do not resume them.
## Current sequence
| Chunk | Purpose | Risk | State |
|---|---|---:|---|
-| `WS-QUAL-001-PLAN2` | Reconcile current hosted baseline, retire obsolete machinery, and define the small closure sequence | L1 | Merged PR #260 |
-| `WS-QUAL-001-02R` | Project/setup observable behavior coverage | L2 | Merged PR #265 |
-| `WS-QUAL-001-03R` | Checker observable behavior coverage | L2 | Implementation in progress |
-| `WS-QUAL-001-04R` | Change the exact global hosted CI floor from 78 to 90 after current-main proof | L1 | Proposed after measured >=90.25% proof |
-
-One chunk maps to one PR. A test chunk may close early when its behavioral scope
-is exhausted. The next contract refreshes from current `main`; stale missing-line
-inventories are never implementation authority. If 02R and 03R are
-insufficient, PLAN2 must be amended with one exact owner-specific successor;
-there is no mixed residual-coverage chunk.
+| `WS-QUAL-001-PLAN3` | Replace percentage-only closure with behavior/mutation assurance | L1 | Planning in progress |
+| `WS-QUAL-001-04M` | Pilot pinned changed-scope mutation evidence without a score gate | L1 | Proposed after PLAN3 merge and explicit instruction |
+| `WS-QUAL-001-05M` | Add calibrated blocking behavior-mutation policy | L1 | Proposed only after accepted 04M hosted evidence and explicit instruction |
+
+## Dependency rule
+
+`PLAN3 -> 04M -> human calibration checkpoint -> 05M`.
+
+Each chunk maps to one PR. `04M` may prove that the candidate engine or target
+strategy is unsuitable and stop without `05M`. Planning does not pre-authorize
+either implementation chunk.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md
index 7011b0fed..a8b787aa3 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md
@@ -1,40 +1,51 @@
-# Decisions: WS-QUAL-001 Backend Coverage Floor
+# Decisions: WS-QUAL-001 Behavior And Mutation Assurance
-## D1: Preserve the 90-percent final target
+## D1: Preserve the global 78-percent floor
-The user-set target remains 90 percent for the complete backend application.
+The user confirmed that 78 percent is the permitted complete-backend baseline.
+Named new or materially changed subsystem floors remain at 90 percent.
-## D2: Treat historical QUAL machinery as completed or stopped evidence
+## D2: Supersede the global-90 floor switch
-PRs #103, #105, and #108 remain durable. Stopped parser/semantic-analysis
-attempts and unimplemented 01B2 are not resumed.
+`WS-QUAL-001-04R` is superseded before implementation. Main already exceeds 90
+percent, and raising the floor would not prove assertion sensitivity.
-## D3: Use the hosted combined report as baseline truth
+## D3: Make observable behavior the quality claim
-The current baseline is the exact semantic-lane fan-in evidence, not a local
-developer-machine timing or partial test selection.
+Coverage remains a backstop. Behavior evidence identifies the production
+target, owning tests, and observable result, denial, persisted fact, mapped
+error, idempotent replay, or recovery outcome.
-## D4: Prefer behavior depth over infrastructure
+## D4: Pilot mutation testing before blocking
-The remaining coverage is added through meaningful tests at the cheapest valid
-layer. Real PostgreSQL, MinIO, and HTTP remain mandatory only for behavior that
-depends on those boundaries.
+One bounded non-blocking-score pilot must measure compatibility, result noise,
+and runtime. Infrastructure failure and invalid evidence still fail the pilot.
-## D5: Separate tests from the threshold switch
+## D5: Prefer mutmut provisionally
-The global floor changes only after a current exact head proves at least 90.25
-percent. The enforced floor remains 90 percent; the extra 0.25 is merge-race
-headroom.
+Current official documentation and project metadata make `mutmut` the leading
+candidate for pytest-aware, changed-scope execution on Python 3.11/3.12. The
+pilot may reject it if exact pinning, isolation, determinism, or runtime fails.
-## D6: Keep architecture and CI optimization separate
+## D6: Do not use a global mutation percentage
-Service decomposition, typed ports/UnitOfWork, mutation/property testing, type
-checking, and semantic-lane runtime optimization are worthwhile possible
-initiatives but are not QUAL coverage-closure work.
+Enforcement is based on complete outcomes for eligible changed targets. A
+surviving meaningful mutant is missing behavior proof. Typed equivalent or
+non-behavioral classifications may be designed only from pilot evidence.
-## D7: Never create a mixed residual-coverage bucket
+## D7: Include test-only behavior claims
-Project and checker test chunks retain one product owner each. If they do not
-reach the required headroom, planning adds one exact owner-specific successor
-from refreshed evidence rather than combining ART, AUTH, TASK, background-job,
-and adapter ownership to chase a percentage.
+A test-only PR that claims behavioral or coverage improvement must name bounded
+production targets and owning tests; otherwise it cannot bypass mutation
+assurance merely because application files did not change.
+
+## D8: Keep mutation work off the Backend critical path
+
+The pilot runs independently with hard command/job limits. A blocking rollout
+must preserve the existing complete Backend authority and practical PR latency.
+
+## D9: Require a second human checkpoint
+
+PLAN3 authorizes planning only. Pilot implementation requires its own explicit
+instruction, and blocking rollout requires another explicit human decision
+after exact hosted pilot evidence is reviewed.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md
index 48149bdd6..f7a715ba5 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md
@@ -1,113 +1,119 @@
-# Discovery: WS-QUAL-001 Current-Main Coverage Closure
-
-## Audited baseline
-
-### Current-main refresh for 03R
-
-Backend run `30921410531` on current-main merge commit `5b853d50` completed
-3,068 tests, covered 21,453 of 23,938 statements (89.619016 percent), recorded
-727.166 seconds total hosted wall time, and a 567.994-second slowest lane.
-Reaching 90.25 percent on this denominator requires 21,605 covered statements,
-a net gain of 152. The focused 03R test union covers 168 statements missing
-from this hosted report: checker service 107, runner 45, and compiler 16. That
-projects 21,621 / 23,938, or 90.320829 percent; hosted exact-head fan-in remains
-authoritative.
-
-Checker-owned gaps are sufficient and remain unchanged by ART: service 169,
-runner 45, compiler 26, router 12, repository 11, gate queue 2, and pre-review
-gate 1. Existing checker tests are integration-heavy. Direct fast coverage is
-still missing for observable policy-shape rejection, registry ordering and
-conflicts, routing priority, blocking-policy escalation, role-sensitive result
-redaction, and bounded gate recovery outcomes. These are the preferred 03R
-test seams; unrelated TASK or ART lifecycle paths remain out of scope.
-
-### PLAN2 historical baseline
-
-Hosted Backend run `30854931616` on the final PR #249 tested tree
-`19d48f7ea4bf20cb29f03cbba54f98683ce52661` produced:
-
-- 2,925 collected and completed tests;
-- 20,793 covered statements of 23,475;
-- 2,682 missed statements;
-- 88.575080 percent global statement coverage;
-- 640.284 seconds total backend wall time;
-- 468.506 seconds in the slowest semantic lane.
-
-At the current denominator, 90 percent permits at most 2,347 missed statements.
-The suite therefore needs 335 additional covered statements to reach 90
-percent and 394 to reach the required 90.25-percent pre-switch headroom.
-
-## Current CI behavior
-
-`.github/workflows/backend.yml` runs five semantic lanes, combines exactly five
-coverage files, runs the real API contract drill, blocks below 78 percent
-globally, and applies multiple 90-percent subsystem/per-file checks. Lane
-custody, PostgreSQL isolation, and coverage fan-in are already implemented.
-
-`backend/scripts/coverage_policy.py` and
-`backend/tests/test_coverage_contract.py` are merged historical integrity
-machinery. The current workflow does not invoke that policy script. PLAN2 does
-not wire it into CI or expand its static Python analysis.
-
-## Largest current gaps
-
-The latest hosted coverage JSON identifies these high-value gaps:
-
-| Module | Statements | Missing | Coverage |
-|---|---:|---:|---:|
-| `app/modules/projects/service.py` | 1,451 | 550 | 62.10% |
-| `app/modules/checkers/service.py` | 579 | 169 | 70.81% |
-| `app/modules/authorization/router.py` | 484 | 168 | 65.29% |
-| `app/modules/artifacts/service.py` | 959 | 138 | 85.61% |
-| `app/modules/tasks/service.py` | 682 | 108 | 84.16% |
-| `app/modules/projects/repository.py` | 285 | 96 | 66.32% |
-| `app/modules/artifacts/operator.py` | 204 | 80 | 60.78% |
-| `app/modules/projects/router.py` | 178 | 63 | 64.61% |
-| `app/modules/artifacts/guide_extraction_worker.py` | 237 | 65 | 72.57% |
-
-Smaller gaps exist in checker repository/router/runner/compiler, project setup
-queue and policy replay, authorization read/repository code, auth API/deps/
-schemas, artifact extraction/materialization, background-job modules, and actor
-services.
-
-## Existing ownership and test layers
-
-- Project behavior: `backend/tests/test_projects.py` and focused project files.
-- Task behavior: `backend/tests/test_tasks.py`.
-- Checker behavior: `backend/tests/test_checkers.py` and runner tests.
-- Artifact behavior: focused artifact, storage, guide, and recovery tests.
-- Authorization behavior: focused actor/authorization/API tests.
-- Test isolation: `backend/scripts/run_isolated_tests.py`.
-- Semantic execution: `backend/scripts/run_test_lanes.py`.
-
-The largest services depend directly on `AsyncSession`; this makes broad unit
-extraction an architectural concern outside QUAL. Tests may use small typed
-fakes or existing fixtures where behavior is observable, but QUAL must not
-refactor production services merely to raise coverage.
-
-## Risks discovered
-
-- Adding hundreds of database-heavy covered lines could worsen the current
- 10.7-minute hosted wall time.
-- Testing implementation branches without outcomes can manufacture percentage
- while adding little confidence.
-- Raising the floor in the same PR as broad tests makes failures harder to
- diagnose and encourages threshold bargaining.
-- Concurrent AUTH, ART, and REV work can increase the denominator; the final
- floor chunk must remeasure current `main` and retain headroom.
-
-## Conventions to preserve
-
-- Complete `backend/app` inventory and combined semantic-lane coverage.
-- Real PostgreSQL for constraints, locks, migrations, transactions, triggers,
- and concurrency.
-- Real MinIO for the protocol boundary.
-- Global 78-percent floor until the exact 90-percent switch merges.
-- Existing protected 90-percent subsystem gates.
-- Test-delta and CI-integrity review for every QUAL implementation PR.
-
-## Unknowns resolved per implementation chunk
-
-The exact missing lines and best observable tests must be refreshed from the
-then-current hosted coverage JSON. A contract may not promise a coverage gain
-from stale line numbers or require tests that merely execute code.
+# Discovery: WS-QUAL-001 Behavior And Mutation Assurance
+
+## Current hosted truth
+
+Main Backend run `30926337804` on merge `5f2baf90` completed 3,162 tests with
+21,620 / 23,938 statements covered (90.316651 percent), 620.264 seconds total
+hosted wall time, and a 464.471-second slowest lane. The complete suite is above
+90 percent, but `.github/workflows/backend.yml` intentionally blocks globally
+at 78 percent and applies more than ten named 90-percent subsystem/per-file
+checks.
+
+This means raising the global floor is neither necessary nor sufficient for the
+new human goal. The remaining gap is whether assertions detect behavioral
+changes.
+
+## Existing test-integrity controls
+
+| Control | Current implementation | What it proves | What it does not prove |
+|---|---|---|---|
+| Complete semantic lanes | `.github/workflows/backend.yml`, `backend/scripts/run_test_lanes.py` | Five canonical lanes collect and execute under isolated custody | Assertions are sensitive to faults |
+| Evidence validation | `backend/scripts/validate_test_lane_evidence.py`, `merge_test_lane_evidence.py` | No missing lane, skipped node, missing coverage, or invalid bundle | Tests kill plausible defects |
+| Global coverage | `coverage report --fail-under=78` | Complete app execution stays above the permitted baseline | Behavioral correctness |
+| Protected coverage | Named `--fail-under=90` checks | New/material subsystems retain deeper execution | Assertions reject wrong outcomes |
+| Weakening scan | `scripts/workstream_agent_gate.py` | Flags common skip/bypass/threshold suppression tokens | Semantic weakening expressed without those tokens |
+| Internal review | QA, test-delta, CI integrity and other routed reviewers | Human/agent reasoning examines behavior and scope | Deterministic executable fault sensitivity |
+
+`backend/tests/test_project_policy_mutations.py` tests project-policy mutation
+behavior; it is not a mutation-testing engine. No `mutmut`, Cosmic Ray, or
+equivalent package/configuration currently exists in backend dependencies or
+GitHub workflows.
+
+## Candidate engine evidence
+
+The current `mutmut` documentation says the tool supports pytest-aware test
+selection, function/module wildcards, incremental results, parallel execution,
+source-path restriction, and optional covered-line filtering. Its current
+project metadata supports Python 3.10 through 3.14, which includes Workstream's
+Python 3.11/3.12 range. It requires fork support, compatible with hosted Linux
+runners. Sources:
+
+-
+-
+
+Cosmic Ray is also viable and stores resumable mutation sessions, but its
+configuration centers on explicit module paths and test commands, and its
+official documentation notes that plugin options are not fully documented.
+That creates more wrapper/configuration ownership for the first pilot:
+
+-
+-
+
+Planning therefore selects `mutmut` only as the leading candidate. The pilot
+must prove an exact pinned release, async pytest compatibility, deterministic
+results, safe worktree isolation, and bounded hosted runtime before adoption.
+
+## Selection boundary
+
+Production changes can be derived from `origin/main...HEAD`. Test-only behavior
+changes have no changed production file, so a deterministic behavior-claim
+manifest is required to name the bounded production targets and owning tests.
+Without this second path, a coverage-only test PR could avoid mutation
+assurance entirely.
+
+The planned canonical boundary is
+`.ci/behavior-claims/.json` under a repository-owned schema. It is
+immutable PR content, not PR prose or a workflow input. Behavior claims name
+repository-relative targets, qualified callables, owning pytest nodes, and
+typed outcomes. Narrow non-behavioral test maintenance is classified through
+the same schema so “no production diff” cannot become an implicit bypass.
+
+Eligible pilot targets should begin with pure functions or direct service
+methods that have fast owning tests. Initial discovery candidates live in the
+project/checker policy, compiler, and runner layers already exercised by 02R
+and 03R. The implementation chunk must choose a much smaller representative
+set from current main and record why each target is eligible.
+
+Ineligible-by-default categories for the pilot:
+
+- migrations and generated/declarative files;
+- Pydantic/SQLAlchemy schemas whose mutations are primarily framework noise;
+- composition-only modules and adapter wiring;
+- external-effect adapters requiring network or real object storage per mutant;
+- modules whose only truthful proof requires the full PostgreSQL/HTTP suite;
+- unchanged modules not named by an explicit test-only behavior claim.
+
+## Evidence model
+
+A useful result must bind:
+
+- exact git tree/source digest;
+- mutation engine version and configuration digest;
+- target module/callable and owning test nodes;
+- generated, killed, survived, timeout, suspicious, excluded, and error counts;
+- stable mutant identifiers and classifications;
+- command timeout and elapsed time;
+- whether the result is pilot-only or blocking.
+
+A percentage without these facts is insufficient. Cache reuse is allowed only
+when the wrapper proves the cached inputs match the exact current inputs.
+
+## Unknowns the pilot must answer
+
+- Whether current mutmut works cleanly with Workstream's async pytest fixtures.
+- How precisely relevant tests are selected without broad incidental execution.
+- Which mutation operators create equivalent/noisy results in Workstream code.
+- Hosted runtime and p95 variability for representative changed targets.
+- Whether fresh execution is cheap enough or authenticated cache reuse is
+ needed.
+- Which narrow classification categories can be machine checked without
+ becoming an exclusion escape hatch.
+
+## Historical reconciliation
+
+PRs #103, #105, #108, #265, and #269 remain completed QUAL evidence. The old
+01B2/milestone ladder remains superseded. `WS-QUAL-001-04R`, which proposed
+raising the global floor to 90 percent, is superseded before implementation by
+the human decision to keep 78 percent and move to behavior/mutation assurance.
+Historical ENG-008 mutation planning is discovery input only; its retired
+signed-loop and machine-scope requirements are not current authority.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md
index 81a64b5b6..d4c8e64ac 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md
@@ -1,49 +1,108 @@
-# Intent: WS-QUAL-001 Backend Coverage Floor
+# Intent: WS-QUAL-001 Behavior And Mutation Assurance
-## Human goal
+## Problem being solved
-Make the complete backend test suite protect at least 90 percent statement
-coverage without weakening tests, hiding application files, or making CI slow
-through unnecessary PostgreSQL and HTTP duplication.
+Statement coverage shows that code executed, but it does not show that a test
+would detect a meaningful behavioral defect. Workstream now exceeds 90 percent
+global backend coverage while CI still correctly permits a 78-percent global
+floor. The unfinished QUAL problem is assertion sensitivity and behavior
+ownership, not another global percentage increase.
-## Why this matters
+## Why this work matters
-Coverage is a backstop for behavior proof, not the goal by itself. Workstream's
-authorization, artifact, project, task, checker, review, and contribution
-boundaries need meaningful failure and recovery tests while the repository
-remains practical for contributors.
+Humans and agents can add tests that execute lines without proving returned
+data, persisted state, denial, failure, recovery, audit, queue, or lifecycle
+behavior. A trusted repository must reject those weak tests without making
+every contributor run the entire backend once per mutant.
-## Current truth
+## Current behavior
-The latest complete hosted result after ART-03C ran 2,925 tests and covered
-20,793 of 23,475 application statements: 88.575080 percent. The
-global CI floor remains 78 percent, while named new or materially changed
-subsystems are already protected at 90 percent.
+- Main Backend run `30926337804` on merge `5f2baf90` completed 3,162 tests and
+ covered 21,620 / 23,938 statements (90.316651 percent).
+- The global blocking floor remains 78 percent.
+- Named new or materially changed subsystems retain blocking 90-percent floors.
+- Semantic lanes reject incomplete collection, skipped nodes, missing coverage,
+ and invalid evidence.
+- Agent Gates scan common test and CI weakening tokens.
+- No real mutation engine or mutation-result policy currently runs in CI.
-## Success state
+## Target behavior
-- The exact complete backend suite covers at least 90.00 percent globally
- across the complete importable `backend/app` inventory.
-- GitHub CI blocks below a global `--fail-under=90` floor.
-- New tests protect observable behavior, rejection, failure, or recovery.
-- Pure or adapter-contract tests are preferred when PostgreSQL and HTTP are not
- the behavior under test.
-- Existing semantic lanes, isolation, coverage combination, and protected
- 90-percent subsystem checks remain intact.
+- Preserve the 78-percent global floor and every protected 90-percent floor.
+- Require behavior claims to identify the production module and observable
+ outcome they protect.
+- Mutation-test only eligible changed production logic or explicitly claimed
+ production targets for test-only behavior PRs.
+- Treat surviving meaningful mutants as missing behavior proof, not as a reason
+ to increase statement coverage.
+- Bound runtime, isolate mutation evidence from ordinary coverage, and keep the
+ complete Backend suite authoritative.
+- Introduce blocking mutation policy only after a measured non-blocking pilot
+ proves deterministic selection, acceptable noise, and acceptable runtime.
-## Non-goals
+## Design chosen
-- No production behavior, schema, migration, API, authorization, or product
- lifecycle change.
-- No arbitrary sharding or infrastructure purchase.
-- No test deletion, weakened assertion, skip, xfail, coverage pragma, omit, or
- narrowed application inventory.
-- No revival of the historical signed-memory, base-evidence, semantic-parser,
- line-budget, or per-milestone ratchet process.
-- No promise that coverage alone proves correctness.
+Use a two-stage rollout. First, pilot one pinned mutation engine with
+deterministic target/test selection and complete non-blocking score evidence.
+Second, after human review of pilot evidence, introduce a separate fail-closed
+gate for eligible changed logic and explicit test-only behavior claims. The
+blocking policy is survivor-based with reviewed classifications, not a global
+mutation-score target.
-## Human decision already provided
+## Alternatives considered
-The user directed the orchestrator to restart QUAL only after current-main
-documentation reconciliation. This PLAN2 audit is authorized; implementation
-still begins with the first reviewed bounded successor.
+- Raise global coverage to 90 percent: rejected because coverage is already
+ above 90 and percentage alone does not prove assertion sensitivity.
+- Mutate the full backend on every PR: rejected because runtime would be
+ unbounded and would discourage contribution.
+- Require one killed mutant per test: rejected because it is easily gamed and
+ does not prove all eligible changed behavior.
+- Immediately block on an uncalibrated mutation percentage: rejected because
+ equivalent/noisy mutants and infrastructure behavior must be measured first.
+- Restore historical signed-loop or machine-scope machinery: rejected; those
+ systems were intentionally retired and are not prerequisites for quality.
+
+## Boundaries preserved
+
+- Coverage, real PostgreSQL, migration, trigger, lock, concurrency, MinIO, API,
+ and semantic-lane checks remain unchanged.
+- QUAL owns test-assurance policy, evidence, and CI integration only.
+- Production defects found by mutation testing move to the owning product
+ initiative; QUAL does not silently repair product behavior.
+- Mutation targets exclude migrations, generated/declarative code, schemas,
+ adapters requiring external effects, and modules without an explicitly
+ reviewed eligibility rule during the pilot.
+
+## Expected risks
+
+- Mutation runtime can multiply test time.
+- Equivalent or invalid mutants can create noisy false blockers.
+- Target or test selection can be gamed to omit behavior.
+- Test-only PRs need an explicit production target to avoid percentage padding.
+- A new pinned tool adds dependency and supply-chain maintenance.
+
+## What must not change
+
+- Global coverage floor stays at 78 percent.
+- Protected subsystem floors stay at 90 percent.
+- No skips, xfails, coverage exclusions, assertion deletion, or narrower test
+ inventory.
+- No mutation pragma or exclusion may be introduced casually to make CI pass.
+- No full-repository mutation run is added to the normal PR critical path.
+
+## How this will be proven
+
+- Unit tests prove diff-to-target eligibility, explicit test-only claims,
+ classification grammar, evidence completeness, and fail-closed behavior.
+- A hosted pilot records generated, killed, survived, timeout, suspicious,
+ excluded, and error outcomes with exact source/test identity and elapsed time.
+- Known behavior tests must kill seeded representative mutants.
+- Weak or vacuous fixture tests must leave representative mutants alive in the
+ policy regression suite.
+- Existing Backend and Agent Gates remain green and unchanged in authority.
+
+## Human decisions required
+
+The 78-percent global floor decision is complete. A separate human checkpoint
+is required after pilot evidence and before any mutation result becomes a
+blocking merge gate.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md
index b6aa7484e..7c1e439a9 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md
@@ -1,77 +1,122 @@
-# Plan: WS-QUAL-001 Current-Main Coverage Closure
+# Plan: WS-QUAL-001 Behavior And Mutation Assurance
## Approach
-Retire the old milestone ladder and close the remaining gap with two declared
-bounded test chunks followed by one floor switch. One additional owner-specific
-test chunk is permitted only when 02R and 03R exhaust their meaningful gaps
-without reaching the required headroom:
-
-1. Add fast project/setup behavior tests for observable service, repository,
- routing, queue, and replay gaps.
-2. If still needed, add fast checker behavior tests for observable service,
- repository, runner, compiler, and routing gaps.
-3. On current `main`, prove at least 90.25 percent globally and change only the
- canonical GitHub global floor from 78 to 90.
-
-Each test chunk starts by reading the current hosted coverage JSON and selecting
-behavioral gaps. It prefers pure functions, typed fakes, direct use-case calls,
-and adapter contracts. PostgreSQL, MinIO, or HTTP is used only when that
-boundary is itself the assertion.
-
-## Coverage target
-
-The last necessary test chunk must reach at least 90.25 percent before the CI
-switch. This is operational headroom, not a permanent higher policy floor. If
-concurrent main growth moves the measured result below 90.25 percent, the floor
-chunk stops and returns to one owner-specific test plan; it never lowers or
-rounds around the target.
-
-## Test-quality rule
-
-Every new test must name and assert one observable contract such as returned
-data, persisted state, emitted audit/outbox fact, queue decision, mapped error,
-authorization denial, idempotent replay, or recovery outcome. A test whose only
-effect is executing previously missed lines is invalid.
-
-No chunk may introduce skips, xfails, coverage pragmas, omit/include narrowing,
-deleted assertions, broad mocking of the behavior under test, or duplicated
-database/HTTP coverage already owned by another layer.
-
-## Boundaries
-
-- QUAL changes tests and, only in the final chunk, the global CI threshold and
- its lightweight invariant test.
-- A production defect discovered by a stronger test is reported and fixed in a
- separate owning initiative/chunk.
-- Production service decomposition, repository ports, UnitOfWork design, type
- checking, mutation testing, and property-test architecture require separate
- initiatives. They are not hidden inside coverage closure.
-- CI runtime optimization remains WS-CI-owned. QUAL records test-time impact and
- must avoid obvious regressions but does not redesign lane infrastructure.
+Retire the proposed global-90 floor switch and deliver behavior assurance in
+two independently reviewed implementation chunks.
+
+### Stage 1: changed-scope mutation pilot
+
+Add one exactly pinned mutation engine and a Workstream-owned policy wrapper.
+The wrapper derives a closed set of eligible production targets from the git
+delta or from an explicit test-only behavior claim. It selects the smallest
+owner test set, runs under a hard timeout, and emits machine-readable exact-head
+evidence.
+
+The pilot does not block on mutation score. It does block on infrastructure
+failure, malformed evidence, target escape, missing claimed tests, ordinary
+test failure, or any weakening of existing Backend checks. Pilot results must
+distinguish killed, survived, timeout, suspicious, excluded, and error mutants.
+
+`mutmut` is the leading pilot candidate because its current documentation
+supports pytest-aware test selection, function/module wildcards, incremental
+results, parallel execution, source/selection configuration, covered-line
+filtering, and Python 3.11/3.12. The implementation chunk must still prove a
+pinned release against Workstream's async pytest and isolated-service setup;
+planning does not pre-approve an unusable dependency.
+
+### Stage 2: blocking behavior-mutation gate
+
+Only after pilot review, add a separate required check for eligible changed
+production logic and test-only PRs that claim behavioral improvement. The gate
+uses the pilot's deterministic target and evidence grammar.
+
+There is no repository-wide mutation percentage. Every eligible survivor
+blocks unless it has a narrow, typed classification accepted by policy (for
+example, demonstrably equivalent or non-behavioral). Missing, stale, broad, or
+free-form exclusions fail closed. Timeout and tool errors do not count as
+killed mutants and cannot silently pass.
+
+## Behavior ownership
+
+A qualifying behavior claim identifies:
+
+- production module and callable or bounded target;
+- owning test nodes;
+- observable contract (return, persisted state, emitted fact, denial, mapped
+ error, idempotent replay, or recovery outcome);
+- relevant real boundary, if PostgreSQL, MinIO, HTTP, lock, trigger, or
+ concurrency is essential.
+
+The canonical input is a schema-v1 JSON file at
+`.ci/behavior-claims/.json`, validated by a repository-owned schema
+and policy parser. Chat, PR prose, labels, workflow inputs, and environment
+variables cannot widen targets. Behavior claims contain repository-relative
+production targets, qualified callables, owning pytest node IDs, and typed
+observable outcomes. Test-only non-behavioral maintenance uses a narrow typed
+classification defined by policy rather than free-form exemption text.
+
+Test-only changes that claim coverage or stronger behavior must provide this
+mapping. Documentation-only, fixture-only, generated-code, and non-behavioral
+maintenance changes are outside mutation selection but remain subject to
+ordinary tests and review.
+
+## Runtime and isolation strategy
+
+- Never mutate the full backend in ordinary PR CI.
+- Start with pure or direct-service logic whose owning tests avoid PostgreSQL
+ and HTTP unless those boundaries are the behavior being proved.
+- Run mutation work independently from the existing Backend critical path.
+- Pilot command limit: 12 minutes inside a 15-minute job limit.
+- A blocking rollout must demonstrate a practical hosted p95 and cannot extend
+ required PR latency by more than two minutes when run in parallel.
+- Mutation caches are acceleration only; evidence binds the exact source,
+ configuration, selected tests, tool version, and result set.
+
+## Dependency and evidence integrity
+
+- Pin the selected engine and its transitive dependency closure with hashes.
+- Do not add the mutation engine to production dependencies.
+- Install the engine only from `scripts/mutation-requirements.txt` with
+ `pip install --require-hashes`; `backend/pyproject.toml` may contain tool
+ configuration but cannot add the engine to ordinary dev extras.
+- Never apply mutants to the contributor worktree in CI.
+- Upload bounded result evidence without source secrets, environment values,
+ database contents, or artifact payloads.
+- The policy wrapper, not mutable PR prose, determines eligibility and validates
+ results.
+- Mutation CI runs only on an unprivileged `pull_request`/`push` boundary with
+ explicit read-only permissions, pinned Actions, checkout credentials
+ disabled, no secrets or writable token in the mutation subprocess, and
+ bounded non-restorable artifacts/caches.
## Alternatives rejected
-- Reviving `01B2` and the complex base-evidence ratchet: unnecessary now that
- exact lane custody and hosted coverage evidence exist.
-- One large cross-owner coverage PR: crosses project, checker, task, artifact, and
- authorization ownership and is difficult to review.
-- Raising the floor immediately: current measured coverage is below 90.
-- Excluding low-coverage services or files: makes the global percentage false.
-- More arbitrary shards: changes runtime distribution, not test architecture or
- coverage quality.
+- `WS-QUAL-001-04R` global floor switch: superseded before implementation.
+- Full-suite-per-mutant execution: too slow and poorly owned.
+- Score-only gating: hides which behavior remains unproved.
+- Non-blocking forever: measures quality without protecting it.
+- Immediate blocking rollout: lacks runtime and equivalent-mutant calibration.
+- Mutating only covered lines as the sole eligibility rule: can hide untested
+ changed behavior; covered-line filtering may optimize the pilot but cannot
+ define the full policy.
## Verification strategy
-Every implementation chunk runs focused tests, Ruff for changed tests, complete
-test-delta review, relevant stale-contract checks, and hosted Backend. The final
-floor chunk additionally proves the combined coverage JSON covers the complete
-application inventory at or above 90.25 percent and that every protected
-90-percent check remains blocking.
+Each implementation chunk runs focused policy tests, mutation-engine smoke
+tests, Ruff, Agent Gates, Markdown/stale scans, internal reviewer tracks, and
+hosted Backend. The pilot additionally proves at least one known strong test
+kills its representative mutants and at least one deliberately weak test leaves
+a representative mutant alive. The blocking chunk proves survivors, timeouts,
+errors, missing evidence, stale evidence, and target escape all stop the gate.
## Dependency order
-PLAN2 -> 02R -> optional 03R -> 04R. If those exact owner-scoped chunks do not
-provide enough headroom, stop and plan one additional owner-specific test chunk
-from the refreshed report. Do not create a percentage-driven residual bucket.
-The CI floor change always remains a separate final PR.
+`PLAN3 -> 04M pilot -> human calibration checkpoint -> 05M blocking gate`.
+`05M` cannot begin from planning alone; it requires accepted exact hosted pilot
+evidence and a new explicit human instruction.
+
+## Stop
+
+Planning does not install a mutation engine, change a workflow, or change a
+coverage threshold. Stop after the PLAN3 PR and human checkpoint.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md
index 6e5f2134b..ad6cb818f 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md
@@ -1,14 +1,22 @@
-# Risks: WS-QUAL-001 Current-Main Coverage Closure
+# Risks: WS-QUAL-001 Behavior And Mutation Assurance
| Risk | Consequence | Control |
|---|---|---|
-| Coverage-only tests | Higher percentage without stronger behavior proof | Require observable outcomes and QA/test-delta review |
-| More database/HTTP tests | Backend CI becomes slower | Prefer pure/use-case/adapter-contract tests unless the real boundary is essential |
-| Concurrent denominator growth | Candidate falls below 90 before floor merge | Remeasure current main and require >=90.25% headroom before 04R |
-| Threshold bundled with tests | Harder diagnosis and pressure to bargain | Keep the 90-percent switch in separate chunk 04R |
-| Production defect discovered | QUAL scope expands into product repair | Stop and hand defect to owning initiative |
-| Historical parser revival | Reintroduces complexity and maintenance burden | Mark 01B2 and old replacements superseded |
-| File exclusion or pragma | False global measurement | Preserve complete app inventory and existing stale/coverage guards |
-| Duplicate invariant tests | Slower suite and ambiguous ownership | Map each new test to its owning layer and review existing proof first |
+| Full-repository mutation | CI becomes unusably slow | Mutate only eligible changed or explicitly claimed targets under hard limits |
+| Equivalent/noisy mutants | Correct PRs are blocked without quality benefit | Non-blocking pilot, typed classifications, separate human checkpoint before enforcement |
+| Score gaming | Contributors kill one easy mutant or raise a percentage while behavior remains weak | Survivor-based exact evidence; no “one mutant” or global score success rule |
+| Target-selection escape | Important changed logic is silently omitted | Workstream-owned deterministic diff/claim parser; missing or broad evidence fails closed |
+| Test-selection escape | Mutants pass because relevant tests were not selected | Bind explicit owning nodes, baseline-run them first, validate selection in policy tests |
+| Test-only coverage padding | No production diff means no mutation work | Require bounded production targets for test-only behavior/coverage claims |
+| Timeout treated as success | Hanging mutants silently pass | Timeout is a distinct non-killed outcome and blocks in enforcement mode |
+| Cache poisoning/staleness | Results do not describe the current source | Bind tree, config, tool, targets, tests, and result digests; cache is acceleration only |
+| Mutation exclusions spread | Meaningful behavior is hidden | No source pragmas in pilot; classifications are narrow, typed, reviewed evidence |
+| Dependency compromise | CI executes an untrusted tool closure | Exact pin and hash-lock the development-only dependency closure |
+| Worktree mutation | Contributor source is left modified | Execute in disposable isolated workspace and verify tree custody |
+| Existing gates weakened | Mutation becomes a substitute for real tests | Preserve semantic lanes, full suite, API E2E, 78 global and protected 90 floors |
+| Runtime critical-path increase | Contribution slows despite scoped execution | Independent job, 12-minute command/15-minute job bounds, <=2-minute critical-path objective |
+| Production defect discovered | QUAL scope drifts into product repair | Stop and hand the defect to its owning initiative |
-No secret, credential, deployment, payment, or production-data access is needed.
+No secret, production credential, production data, payment, or deployment
+access is required. The CI dependency and executable-tool boundary requires
+security and CI-integrity review.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md
index ae60423b5..68432a496 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md
@@ -1,41 +1,31 @@
-# Status: WS-QUAL-001 Backend Coverage Floor
+# Status: WS-QUAL-001 Behavior And Mutation Assurance
## Current state
-`WS-QUAL-001-PLAN2` merged through PR #260. `WS-QUAL-001-02R` merged through
-PR #265 with all 2,996 exact-head tests passing and 21,004 / 23,455 statements
-covered (89.550203 percent).
+Coverage closure through `WS-QUAL-001-03R` is complete. PR #269 merged at
+`5f2baf90`; main Backend run `30926337804` completed all 3,162 tests and
+reported 21,620 / 23,938 statements (90.316651 percent), 620.264 seconds hosted
+wall time, and a 464.471-second slowest lane.
-The current-main baseline after ART PR #268 is Backend run `30921410531` on
-`5b853d50`: 3,068 tests completed, 21,453 / 23,938 statements covered
-(89.619016 percent), 727.166 seconds hosted wall time, and a 567.994-second
-slowest lane. The global CI floor remains 78 percent; named protected subsystem
-checks remain blocking at 90 percent.
-
-Historical QUAL work delivered the isolated database runner and test-integrity
-guards through PRs #103, #105, and #108. The many stopped semantic-analysis
-replacements remain historical evidence, not work to resume.
+The global blocking floor remains 78 percent by explicit human decision. Named
+new or materially changed subsystem checks remain blocking at 90 percent.
## Current gate
-`WS-QUAL-001-03R` is the current implementation chunk. It must add meaningful
-checker-owned behavior tests and gain at least 152 covered statements on the
-current-main denominator to reach the 90.25-percent headroom target. The
-focused test union measures
-168 unique previously missing checker statements and projects 21,621 / 23,938,
-or 90.320829 percent. Require Agent
-Gates, CodeRabbit, all Backend semantic lanes, final coverage fan-in, six
-internal reviewer tracks, and human review. An unexplained focused or hosted
-runtime increase above 10 percent stops merge readiness.
-
-The 94-case focused implementation selection passes in 37.76 seconds with
-narrow checker-module coverage. The isolated full `test_checkers.py` run
-reached the 1,200-second local ceiling after approximately 95 percent completion
-with 163 passing tests and no failure output. It is not a complete pass; hosted
-Backend remains mandatory.
+`WS-QUAL-001-PLAN3` is planning only. It replaces the unstarted 04R global-floor
+switch with a two-stage behavior/mutation assurance proposal:
+
+1. `04M` — bounded, pinned, changed-scope mutation pilot with complete evidence
+ and no blocking score.
+2. Human calibration checkpoint.
+3. `05M` — separately approved blocking survivor policy for eligible changed
+ logic and explicit test-only behavior claims.
+
+No mutation dependency, workflow, policy script, or blocking check has been
+implemented.
## Stop condition
-Planning does not change tests, application code, workflow code, or thresholds.
-Do not raise the global floor until hosted combined coverage is at least 90.25
-percent on the exact candidate head.
+Stop after PLAN3 planning review and PR. Do not start 04M automatically. Do not
+start 05M without accepted exact hosted pilot evidence and a new human
+instruction.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md
index ad3b62973..8a3d8d2c1 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md
@@ -1,13 +1,14 @@
# QUAL Chunk Records
-Only these PLAN2 records describe possible current work:
+Only these PLAN3 records describe possible current work:
-- `WS-QUAL-001-PLAN2-current-main-reconciliation.md`
-- `WS-QUAL-001-02R-project-setup-behavior-coverage.md`
-- `WS-QUAL-001-03R-checker-behavior-coverage.md`
-- `WS-QUAL-001-04R-global-90-floor.md`
+- `WS-QUAL-001-PLAN3-behavior-mutation-assurance.md`
+- `WS-QUAL-001-04M-changed-scope-mutation-pilot.md`
+- `WS-QUAL-001-05M-blocking-behavior-mutation-gate.md`
Every other file in this directory is historical evidence from the original
QUAL plan, a completed chunk, a stopped repair attempt, or a superseded
-contract. Historical records cannot be started or treated as current
+contract. In particular, `WS-QUAL-001-04R-global-90-floor.md` was superseded
+before implementation by the human decision to retain the 78-percent global
+floor. Historical records cannot be started or treated as current
implementation authority. `CHUNK_MAP.md` records their final disposition.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md
new file mode 100644
index 000000000..8665a5051
--- /dev/null
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md
@@ -0,0 +1,132 @@
+# Chunk Contract: WS-QUAL-001-04M — Changed-Scope Mutation Pilot
+
+## Parent initiative
+
+`WS-QUAL-001` — Behavior And Mutation Assurance
+
+## Goal
+
+Pilot one exactly pinned mutation engine on eligible changed or explicitly
+claimed production targets and publish complete exact-head result evidence
+without imposing an uncalibrated mutation-score gate.
+
+## Why this chunk exists
+
+Coverage proves execution, not assertion sensitivity. Workstream needs measured
+compatibility, mutant quality, selection integrity, and hosted runtime before a
+mutation outcome can block contributions.
+
+## Approved plan reference
+
+- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md`
+- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md`
+- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md`
+
+## Risk class
+
+L1 — CI executable/dependency/evidence policy.
+
+## SLA
+
+P2.
+
+## Allowed files
+
+```text
+backend/pyproject.toml
+backend/scripts/mutation_policy.py
+backend/tests/test_mutation_policy.py
+scripts/git_delta.py
+scripts/test_git_delta.py
+scripts/workstream_agent_gate.py
+scripts/behavior-claim.schema.json
+scripts/mutation-requirements.txt
+scripts/test_lightweight_agent_gates.py
+.ci/behavior-claims/WS-QUAL-001-04M.json
+.ci/behavior-claims/README.md
+.github/workflows/mutation-pilot.yml
+docs/operations_backend_testing.md
+.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/**
+```
+
+## Not allowed
+
+```text
+backend/app/** or backend/alembic/** changes
+global or blocking mutation percentage
+change to the global 78-percent or protected 90-percent coverage floors
+replacement, reduction, or bypass of Backend semantic lanes, API E2E, or fan-in
+full-backend mutation on ordinary PRs
+mutation pragmas or free-form exclusion lists
+production dependency changes
+pull_request_target, privileged PR-code execution, writable workflow token,
+checkout credentials, secrets in mutation execution, or unpinned Actions
+```
+
+## Acceptance criteria
+
+- [ ] One engine and transitive closure are exactly pinned and hash locked as
+ development/CI-only dependencies, installed exclusively from
+ `scripts/mutation-requirements.txt` with `--require-hashes`.
+- [ ] `backend/pyproject.toml` contains configuration only; the mutation engine
+ is absent from production dependencies and ordinary dev extras.
+ `scripts/mutation-requirements.txt` is the sole mutation-tool dependency
+ authority; `backend/uv.lock` remains unchanged and is not a second install
+ path.
+- [ ] Deterministic policy selects eligible changed targets or validates a
+ bounded test-only behavior claim with explicit owning test nodes.
+- [ ] Git-delta discovery extracts one shared `scripts/git_delta.py` primitive
+ reused by `scripts/workstream_agent_gate.py` and mutation policy, and
+ mutation evidence mirrors the
+ existing semantic-lane exact-tree/digest/fail-closed conventions rather
+ than creating a parallel custody dialect.
+- [ ] Schema-v1 `.ci/behavior-claims/.json` is the only test-only
+ claim input; PR prose, labels, workflow inputs, and environment variables
+ cannot widen production targets or owning test nodes.
+- [ ] `.ci/behavior-claims/README.md` and the backend testing operations guide
+ document the pilot format without presenting it as a blocking contributor
+ requirement before 05M.
+- [ ] Disposable execution cannot leave mutants in the checked-out source tree.
+- [ ] Evidence binds exact tree, tool/config, target/test identities, elapsed
+ time, and generated/killed/survived/timeout/suspicious/excluded/error
+ outcomes.
+- [ ] Known strong behavior tests kill representative mutants and a deliberately
+ weak fixture leaves a representative mutant alive.
+- [ ] Score is observational; infrastructure errors, malformed/stale evidence,
+ target escape, or baseline test failure remain blocking.
+- [ ] Mutation command is bounded to 12 minutes inside a 15-minute independent
+ job and records critical-path impact.
+- [ ] Workflow runs untrusted PR code only through `pull_request`/`push`, uses
+ explicit read-only permissions, pinned Actions, `persist-credentials:
+ false`, no secrets or writable token in mutation execution, and bounded
+ non-restorable artifacts/caches; invariant tests enforce each property.
+- [ ] Full Backend and existing coverage gates remain unchanged and green.
+
+## Verification commands
+
+```bash
+cd backend
+.venv/bin/python -m pytest -q tests/test_mutation_policy.py
+.venv/bin/ruff check scripts/mutation_policy.py tests/test_mutation_policy.py
+cd ..
+python3 scripts/check_markdown_links.py
+PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python3 -m pytest -q scripts/test_lightweight_agent_gates.py
+git diff --check
+```
+
+## Required reviewers
+
+Senior engineering, QA/test, security/auth, product/ops, architecture, CI
+integrity, docs, reuse/dedup, and test delta.
+
+## Human review focus
+
+- Is target/test selection deterministic and non-gamable?
+- Is result evidence complete enough to calibrate a later blocking policy?
+- Is hosted runtime practical without weakening Backend?
+
+## Stop conditions
+
+Stop if the engine cannot be pinned, mutates the contributor worktree, requires
+full-suite-per-mutant execution, produces unclassifiable noise, exceeds runtime
+bounds, or requires weakening any existing gate.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md
index cd9b52f5f..fdadd57f1 100644
--- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md
@@ -1,5 +1,9 @@
# Chunk Contract: WS-QUAL-001-04R — Global 90 Percent CI Floor
+> Superseded before implementation by `WS-QUAL-001-PLAN3`. The human decision
+> retains the 78-percent global floor and moves QUAL to behavior/mutation
+> assurance. This file is historical evidence and must not be started.
+
Parent initiative: `WS-QUAL-001`
Goal: after exact hosted proof at or above 90.25 percent, change the canonical
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md
new file mode 100644
index 000000000..432938d6a
--- /dev/null
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md
@@ -0,0 +1,93 @@
+# Chunk Contract: WS-QUAL-001-05M — Blocking Behavior-Mutation Gate
+
+## Parent initiative
+
+`WS-QUAL-001` — Behavior And Mutation Assurance
+
+## Goal
+
+After accepted pilot evidence and separate human approval, make complete
+mutation outcomes blocking for eligible changed production logic and explicit
+test-only behavior claims.
+
+## Why this chunk exists
+
+The pilot measures feasibility. This separate chunk converts only calibrated,
+deterministic evidence into contributor protection.
+
+## Approved plan reference
+
+- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md`
+- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md`
+- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md`
+- Required input: accepted exact hosted `WS-QUAL-001-04M` pilot evidence.
+
+## Risk class
+
+L1 — blocking CI/test policy.
+
+## SLA
+
+P2.
+
+## Allowed files
+
+Exact files must be refreshed from the merged 04M implementation before start.
+Expected ownership is limited to its mutation policy/tests, independent
+workflow, backend testing operations guide, Agent Gate invariant, and QUAL
+initiative evidence. Test-only inputs use only schema-v1
+`.ci/behavior-claims/.json` files validated by the merged policy.
+The refreshed allowed list must explicitly include `CONTRIBUTING.md`,
+`.ci/behavior-claims/README.md`, the policy-owned schema, and a copyable example
+so external contributors see the exact blocking workflow before it is enabled.
+
+## Not allowed
+
+```text
+start without accepted 04M evidence and explicit human instruction
+global mutation percentage
+change to 78-percent global or protected 90-percent coverage floors
+free-form exemptions, source mutation pragmas, silent timeout/error success
+full-repository mutation on ordinary PRs
+production behavior, migration, or dependency changes
+```
+
+## Acceptance criteria
+
+- [ ] Eligibility and evidence grammar are unchanged from accepted pilot proof
+ unless a separately reviewed correction is explicit.
+- [ ] Every eligible surviving mutant blocks by default.
+- [ ] Any allowed classification is narrow, typed, evidence-bound, and tested;
+ missing, stale, broad, or free-form classifications fail closed.
+- [ ] Timeout, suspicious, and error outcomes never count as killed.
+- [ ] Test-only behavior/coverage claims cannot bypass target mutation.
+- [ ] `CONTRIBUTING.md` and the canonical claim README/schema/example explain
+ when a claim is required, the permitted typed non-behavioral cases, local
+ verification, evidence interpretation, and repair of surviving mutants.
+- [ ] Non-eligible maintenance/docs/generated changes do not run irrelevant
+ mutants.
+- [ ] Hosted p95 and critical-path impact satisfy the accepted pilot bound.
+- [ ] Backend, Agent Gates, 78-percent global floor, and protected 90-percent
+ floors remain authoritative and green.
+
+## Verification commands
+
+Refresh exact commands from merged 04M; at minimum run mutation-policy unit and
+integration tests, strong/weak seeded behavior proof, Ruff, Agent Gates,
+Markdown/stale checks, full hosted Backend, and exact blocking-workflow proof.
+
+## Required reviewers
+
+Senior engineering, QA/test, security/auth, product/ops, architecture, CI
+integrity, docs, reuse/dedup, and test delta.
+
+## Human review focus
+
+- Does the gate block weak behavior proof without blocking unrelated work?
+- Can classifications or test selection be used as an escape hatch?
+- Does the gate remain practical for external contributors?
+
+## Stop conditions
+
+Stop on missing pilot evidence, unreviewed policy change, unacceptable hosted
+latency/noise, coverage/Backend weakening, or need for broad exemptions.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md
new file mode 100644
index 000000000..e08742ded
--- /dev/null
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md
@@ -0,0 +1,95 @@
+# Chunk Contract: WS-QUAL-001-PLAN3 — Behavior And Mutation Assurance Planning
+
+## Parent initiative
+
+`WS-QUAL-001` — Behavior And Mutation Assurance
+
+## Goal
+
+Replace the unstarted global-90 floor proposal with reviewed, bounded planning
+for behavior-owned changed-scope mutation assurance while retaining the global
+78-percent floor.
+
+## Why this chunk exists
+
+Main already exceeds 90-percent statement coverage, but no executable gate
+proves assertion sensitivity. Planning must define a safe pilot and a separate
+calibrated enforcement step before workflow or dependency changes.
+
+## Approved plan reference
+
+- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md`
+- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md`
+- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md`
+
+## Risk class
+
+L1 — CI/test policy planning.
+
+## SLA
+
+P2.
+
+## Allowed files
+
+```text
+.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/**
+```
+
+## Not allowed
+
+```text
+application or test implementation
+workflow, dependency, threshold, coverage, or mutation-engine changes
+revival of retired signed-loop, merge-intent, or machine-scope machinery
+automatic start of 04M or 05M
+```
+
+## Acceptance criteria
+
+- [ ] Current main coverage, runtime, and existing test-integrity gates are
+ recorded exactly.
+- [ ] The global 78-percent and protected 90-percent floors are preserved.
+- [ ] Planning distinguishes behavior evidence from statement coverage.
+- [ ] Pilot selection covers changed production targets and explicit test-only
+ behavior claims without full-repository mutation.
+- [ ] Runtime, dependency, evidence, classification, cache, and worktree-safety
+ risks have fail-closed controls.
+- [ ] Blocking rollout requires accepted pilot evidence and a new human gate.
+- [ ] Required internal planning reviewers pass with no open sessions.
+
+## Verification commands
+
+```bash
+python3 scripts/check_markdown_links.py
+python3 scripts/check_stale_workstream_wording.py
+python3 scripts/check_stale_authorization_docs.py
+python3 scripts/check_stale_artifact_contracts.py
+PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python3 -m pytest -q scripts/test_lightweight_agent_gates.py
+git diff --check
+```
+
+## Required reviewers
+
+- senior engineering
+- QA/test
+- security/auth
+- product/ops
+- architecture
+- CI integrity
+- docs
+- reuse/dedup
+- test delta
+
+## Human review focus
+
+- Confirm 78 percent remains the global floor.
+- Confirm mutation assurance measures behavior rather than another percentage.
+- Confirm the pilot is bounded and the blocking rollout has a separate human
+ checkpoint.
+
+## Stop conditions
+
+Stop if planning implies immediate blocking mutation, full-repository mutation,
+coverage/Backend weakening, an unbounded dependency/runtime commitment, or any
+implementation change.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md
new file mode 100644
index 000000000..498a2a36d
--- /dev/null
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md
@@ -0,0 +1,87 @@
+# WS-QUAL-001-PLAN3 Internal Review Evidence
+
+## Review scope
+
+Planning-only changes under
+`.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/**`. Review began
+against main merge `5f2baf90036839589cf3db5ad6949d4888da5e9e` and reconciled without conflict
+to current main `24f677cb` after CON PLAN5 PR #270 changed only CON/REV/general
+documentation. The QUAL delta is unchanged. No application, test, workflow,
+dependency, threshold, or mutation implementation changed.
+
+## Risk routing
+
+- Risk class: L1.
+- SLA: P2.
+- Work type: CI/test policy architecture and documentation.
+- Required reviewers: senior engineering, QA/test, security/auth, product/ops,
+ architecture, CI integrity, docs, reuse/dedup, and test delta.
+- Human gate: required for PLAN3 merge, required again before 04M, and required
+ after pilot evidence before 05M.
+- Budget posture: high scrutiny with bounded planning scope; no implementation
+ or runtime spend in PLAN3.
+- Why: future work executes third-party tooling against untrusted PR code and
+ may eventually become a blocking contributor gate.
+
+## Circuit breaker
+
+PASS with a planning-record size exception. The diff is larger than the
+preferred L1 implementation guideline because initiative planning must keep
+intent, discovery, plan, decisions, risks, status, chunk map, three contracts,
+and review evidence mutually consistent in one planning PR. It touches one
+major boundary (QUAL CI/test policy), changes no executable file, and has clear
+acceptance criteria and nine completed reviewer tracks. Splitting the records
+would publish internally contradictory partial policy without reducing the
+eventual 04M or 05M implementation boundary.
+
+## Reviewer results
+
+| Reviewer | Result | Final findings |
+|---|---:|---|
+| senior engineering | PASS | None |
+| QA/test | PASS | None |
+| security/auth | PASS | Prior CI privilege and dependency-custody findings resolved |
+| product/ops | PASS | None |
+| architecture | PASS | Shared git-delta boundary and canonical claim input conditions resolved |
+| CI integrity | PASS | None |
+| docs | PASS WITH LOW RISKS | Dependency authority and contributor onboarding conditions resolved |
+| reuse/dedup | PASS WITH LOW RISKS | Existing delta/evidence conventions are now explicit reuse requirements |
+| test delta | PASS | None |
+
+Open reviewer sessions: none.
+
+## Findings resolved
+
+- Added one shared `scripts/git_delta.py` boundary for Agent Gates and mutation
+ policy rather than parallel diff parsing.
+- Declared schema-v1 `.ci/behavior-claims/.json` as the only test-only
+ claim input; mutable PR prose, labels, workflow inputs, and environment
+ variables cannot widen scope.
+- Required mutation tooling to install only from a hash-locked sidecar
+ requirements file; `pyproject.toml` is configuration-only and `uv.lock` is
+ not a second install path.
+- Required unprivileged `pull_request`/`push`, explicit read-only permissions,
+ pinned Actions, disabled checkout credentials, no secrets/writable token in
+ the mutation subprocess, and bounded artifacts/caches.
+- Required canonical claim documentation during the pilot and explicit
+ `CONTRIBUTING.md` onboarding before any blocking rollout.
+
+## Deterministic evidence
+
+- Markdown link scan — passed for all changed Markdown files.
+- Stale Workstream wording scan — passed.
+- Stale authorization documentation scan — passed.
+- Stale artifact-contract scan — passed.
+- Lightweight Agent Gates — 10 passed.
+- `git diff --check` — passed.
+- Allowed scope — QUAL initiative planning tree only.
+- Official candidate-tool assumptions were checked against current mutmut and
+ Cosmic Ray primary documentation.
+
+## Remaining risks
+
+- `mutmut` is provisional; 04M must prove exact pinned compatibility with
+ Workstream's async pytest and isolation setup.
+- Equivalent/noisy mutants and hosted runtime remain unknown until 04M.
+- No mutation result may block merging until exact hosted pilot evidence is
+ accepted through a separate human checkpoint.
diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md
new file mode 100644
index 000000000..41e371317
--- /dev/null
+++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md
@@ -0,0 +1,112 @@
+# WS-QUAL-001-PLAN3 PR Trust Bundle
+
+## Chunk
+
+`WS-QUAL-001-PLAN3` — Behavior And Mutation Assurance Planning.
+
+## Goal
+
+Keep the global Backend floor at 78 percent and replace the unstarted
+percentage-only floor raise with a safe path for behavior-owned mutation
+assurance.
+
+## Human-approved intent
+
+Coverage remains a permitted baseline, not proof that tests verify behavior.
+The user confirmed 78 percent and directed planning for behavior and mutation
+quality before implementation.
+
+## What changed and why
+
+- Reframed QUAL from “raise global coverage” to “prove assertion sensitivity.”
+- Marked unstarted `WS-QUAL-001-04R` superseded.
+- Declared `04M`, a bounded non-blocking-score mutation pilot.
+- Declared `05M`, a separately approved blocking survivor policy only after
+ accepted hosted pilot evidence.
+- Added current-main facts, security/runtime risks, claim ownership, evidence
+ rules, contributor boundaries, and exact chunk contracts.
+
+## Design chosen
+
+Mutation targets are eligible changed production logic or explicit test-only
+behavior claims. Claims use schema-v1
+`.ci/behavior-claims/.json`; mutable PR prose cannot widen them.
+`04M` runs independently under hard limits and records complete outcomes.
+`05M` is not authorized until pilot evidence is accepted by a human.
+
+## Alternatives rejected
+
+- Raising global coverage from 78 to 90: execution percentage is not behavior
+ proof.
+- Full-backend mutation per PR: operationally impractical.
+- A global mutation score or “one mutant killed” rule: gameable and opaque.
+- Immediate blocking rollout: uncalibrated noise and runtime.
+- Retired signed-loop/machine-scope machinery: unnecessary for current simple
+ contribution flow.
+
+## Scope control
+
+Only QUAL initiative planning records change. PLAN3 installs no tool, changes
+no workflow or dependency, modifies no tests/application code, and changes no
+coverage threshold.
+
+## Product behavior
+
+None. Product review decisions remain `accept`, `needs_revision`, and `reject`;
+mutation outcomes are engineering evidence and never product decisions.
+
+## Acceptance criteria proof
+
+- Current main is recorded from Backend run `30926337804`: 3,162 completed
+ tests, 21,620 / 23,938 coverage (90.316651 percent), 620.264 seconds wall,
+ and 464.471 seconds slowest lane.
+- Global 78 and protected 90 floors are explicitly preserved.
+- Changed-production and test-only behavior paths are both planned.
+- Strong-vs-weak seeded mutant proof, complete outcome evidence, runtime bounds,
+ dependency custody, CI privilege controls, and a second human checkpoint are
+ explicit acceptance requirements.
+
+## Tests and checks run
+
+Markdown links, all three stale scans, ten lightweight Agent Gate regression
+tests, scope review, and whitespace validation pass.
+
+## Test delta and CI integrity
+
+No tests or CI files change. Future contracts forbid skips, xfails, assertion
+weakening, coverage exclusions, Backend replacement, and changes to the 78/90
+coverage policy.
+
+## Reviewer results
+
+All nine required tracks pass after resolving shared-helper, claim-boundary,
+CI-privilege, dependency-custody, and contributor-onboarding conditions. No
+reviewer session remains open.
+
+## External review
+
+Agent Gates, CodeRabbit, and human review are pending publication. Backend is
+not required by the planning diff unless GitHub policy schedules it; no Backend
+file changes.
+
+## Remaining risks
+
+Tool compatibility, equivalent-mutant noise, target/test selection quality,
+and hosted runtime remain deliberately assigned to 04M pilot evidence.
+
+## Follow-up work
+
+After PLAN3 merges, 04M may start only by explicit human instruction. 05M
+requires accepted 04M hosted evidence and a separate explicit decision.
+
+## Human review focus
+
+- Confirm the 78-percent global floor remains unchanged.
+- Confirm this measures behavior rather than another percentage.
+- Confirm the pilot is safe for untrusted contributor code and blocking remains
+ behind a second human checkpoint.
+
+## Human merge ownership
+
+GitHub checks and explicit human approval are required. PLAN3 authorizes no
+mutation implementation.