From 6bad521ea231d7f342e02d49b6c3d84a4ae5d8ee Mon Sep 17 00:00:00 2001 From: Abiorh001 Date: Tue, 4 Aug 2026 18:11:18 +0100 Subject: [PATCH] docs(qual): plan behavior mutation assurance --- .../CHUNK_MAP.md | 39 +-- .../DECISIONS.md | 63 +++-- .../DISCOVERY.md | 232 +++++++++--------- .../INTENT.md | 133 +++++++--- .../PLAN.md | 175 ++++++++----- .../RISKS.md | 28 ++- .../STATUS.md | 52 ++-- .../chunks/README.md | 13 +- ...AL-001-04M-changed-scope-mutation-pilot.md | 132 ++++++++++ .../chunks/WS-QUAL-001-04R-global-90-floor.md | 4 + ...001-05M-blocking-behavior-mutation-gate.md | 93 +++++++ ...L-001-PLAN3-behavior-mutation-assurance.md | 95 +++++++ ...QUAL-001-PLAN3-internal-review-evidence.md | 87 +++++++ .../WS-QUAL-001-PLAN3-pr-trust-bundle.md | 112 +++++++++ 14 files changed, 952 insertions(+), 306 deletions(-) create mode 100644 .agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md create mode 100644 .agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md create mode 100644 .agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md create mode 100644 .agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md create mode 100644 .agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md index 58faff50d..8b3dcf633 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md @@ -1,30 +1,33 @@ -# Chunk Map: WS-QUAL-001 Backend Coverage Floor +# Chunk Map: WS-QUAL-001 Behavior And Mutation Assurance -## Historical completed work +## Completed and superseded work | Chunk | Durable outcome | State | |---|---|---| -| `WS-QUAL-001-PLAN` | Original 90-percent initiative plan | Merged PR #99; superseded by PLAN2 sequencing | -| `WS-QUAL-001-01` | Original combined harness/baseline contract | Superseded by the 01A/01B split before implementation | +| `WS-QUAL-001-PLAN` | Original coverage initiative plan | Merged PR #99; superseded | | `WS-QUAL-001-01A` | Isolated least-privilege database runner | Merged PR #103 | | `WS-QUAL-001-01B1A-R2` | Coverage configuration/evidence grammar | Merged PR #105 | -| `WS-QUAL-001-01B1B-R10` | Conservative test-weakening semantic guard | Merged PR #108 | +| `WS-QUAL-001-01B1B-R10` | Conservative test-weakening guard | Merged PR #108 | +| `WS-QUAL-001-PLAN2` | Current-main coverage closure plan | Merged PR #260; succeeded by PLAN3 | +| `WS-QUAL-001-02R` | Project/setup observable behavior coverage | Merged PR #265 | +| `WS-QUAL-001-03R` | Checker observable behavior coverage | Merged PR #269; main at 90.316651% | +| `WS-QUAL-001-04R` | Raise global floor from 78 to 90 | Superseded before implementation; 78 retained by human decision | -All other 01B/01B1/01B1A/01B1B replacement attempts are stopped historical -experiments. Do not resume them. `WS-QUAL-001-01B2` and the old 02-06 milestone -ladder are superseded before implementation. +All other old 01B/01B1 replacement attempts and the old 02-06 milestone ladder +remain stopped historical experiments. Do not resume them. ## Current sequence | Chunk | Purpose | Risk | State | |---|---|---:|---| -| `WS-QUAL-001-PLAN2` | Reconcile current hosted baseline, retire obsolete machinery, and define the small closure sequence | L1 | Merged PR #260 | -| `WS-QUAL-001-02R` | Project/setup observable behavior coverage | L2 | Merged PR #265 | -| `WS-QUAL-001-03R` | Checker observable behavior coverage | L2 | Implementation in progress | -| `WS-QUAL-001-04R` | Change the exact global hosted CI floor from 78 to 90 after current-main proof | L1 | Proposed after measured >=90.25% proof | - -One chunk maps to one PR. A test chunk may close early when its behavioral scope -is exhausted. The next contract refreshes from current `main`; stale missing-line -inventories are never implementation authority. If 02R and 03R are -insufficient, PLAN2 must be amended with one exact owner-specific successor; -there is no mixed residual-coverage chunk. +| `WS-QUAL-001-PLAN3` | Replace percentage-only closure with behavior/mutation assurance | L1 | Planning in progress | +| `WS-QUAL-001-04M` | Pilot pinned changed-scope mutation evidence without a score gate | L1 | Proposed after PLAN3 merge and explicit instruction | +| `WS-QUAL-001-05M` | Add calibrated blocking behavior-mutation policy | L1 | Proposed only after accepted 04M hosted evidence and explicit instruction | + +## Dependency rule + +`PLAN3 -> 04M -> human calibration checkpoint -> 05M`. + +Each chunk maps to one PR. `04M` may prove that the candidate engine or target +strategy is unsuitable and stop without `05M`. Planning does not pre-authorize +either implementation chunk. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md index 7011b0fed..a8b787aa3 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DECISIONS.md @@ -1,40 +1,51 @@ -# Decisions: WS-QUAL-001 Backend Coverage Floor +# Decisions: WS-QUAL-001 Behavior And Mutation Assurance -## D1: Preserve the 90-percent final target +## D1: Preserve the global 78-percent floor -The user-set target remains 90 percent for the complete backend application. +The user confirmed that 78 percent is the permitted complete-backend baseline. +Named new or materially changed subsystem floors remain at 90 percent. -## D2: Treat historical QUAL machinery as completed or stopped evidence +## D2: Supersede the global-90 floor switch -PRs #103, #105, and #108 remain durable. Stopped parser/semantic-analysis -attempts and unimplemented 01B2 are not resumed. +`WS-QUAL-001-04R` is superseded before implementation. Main already exceeds 90 +percent, and raising the floor would not prove assertion sensitivity. -## D3: Use the hosted combined report as baseline truth +## D3: Make observable behavior the quality claim -The current baseline is the exact semantic-lane fan-in evidence, not a local -developer-machine timing or partial test selection. +Coverage remains a backstop. Behavior evidence identifies the production +target, owning tests, and observable result, denial, persisted fact, mapped +error, idempotent replay, or recovery outcome. -## D4: Prefer behavior depth over infrastructure +## D4: Pilot mutation testing before blocking -The remaining coverage is added through meaningful tests at the cheapest valid -layer. Real PostgreSQL, MinIO, and HTTP remain mandatory only for behavior that -depends on those boundaries. +One bounded non-blocking-score pilot must measure compatibility, result noise, +and runtime. Infrastructure failure and invalid evidence still fail the pilot. -## D5: Separate tests from the threshold switch +## D5: Prefer mutmut provisionally -The global floor changes only after a current exact head proves at least 90.25 -percent. The enforced floor remains 90 percent; the extra 0.25 is merge-race -headroom. +Current official documentation and project metadata make `mutmut` the leading +candidate for pytest-aware, changed-scope execution on Python 3.11/3.12. The +pilot may reject it if exact pinning, isolation, determinism, or runtime fails. -## D6: Keep architecture and CI optimization separate +## D6: Do not use a global mutation percentage -Service decomposition, typed ports/UnitOfWork, mutation/property testing, type -checking, and semantic-lane runtime optimization are worthwhile possible -initiatives but are not QUAL coverage-closure work. +Enforcement is based on complete outcomes for eligible changed targets. A +surviving meaningful mutant is missing behavior proof. Typed equivalent or +non-behavioral classifications may be designed only from pilot evidence. -## D7: Never create a mixed residual-coverage bucket +## D7: Include test-only behavior claims -Project and checker test chunks retain one product owner each. If they do not -reach the required headroom, planning adds one exact owner-specific successor -from refreshed evidence rather than combining ART, AUTH, TASK, background-job, -and adapter ownership to chase a percentage. +A test-only PR that claims behavioral or coverage improvement must name bounded +production targets and owning tests; otherwise it cannot bypass mutation +assurance merely because application files did not change. + +## D8: Keep mutation work off the Backend critical path + +The pilot runs independently with hard command/job limits. A blocking rollout +must preserve the existing complete Backend authority and practical PR latency. + +## D9: Require a second human checkpoint + +PLAN3 authorizes planning only. Pilot implementation requires its own explicit +instruction, and blocking rollout requires another explicit human decision +after exact hosted pilot evidence is reviewed. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md index 48149bdd6..f7a715ba5 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/DISCOVERY.md @@ -1,113 +1,119 @@ -# Discovery: WS-QUAL-001 Current-Main Coverage Closure - -## Audited baseline - -### Current-main refresh for 03R - -Backend run `30921410531` on current-main merge commit `5b853d50` completed -3,068 tests, covered 21,453 of 23,938 statements (89.619016 percent), recorded -727.166 seconds total hosted wall time, and a 567.994-second slowest lane. -Reaching 90.25 percent on this denominator requires 21,605 covered statements, -a net gain of 152. The focused 03R test union covers 168 statements missing -from this hosted report: checker service 107, runner 45, and compiler 16. That -projects 21,621 / 23,938, or 90.320829 percent; hosted exact-head fan-in remains -authoritative. - -Checker-owned gaps are sufficient and remain unchanged by ART: service 169, -runner 45, compiler 26, router 12, repository 11, gate queue 2, and pre-review -gate 1. Existing checker tests are integration-heavy. Direct fast coverage is -still missing for observable policy-shape rejection, registry ordering and -conflicts, routing priority, blocking-policy escalation, role-sensitive result -redaction, and bounded gate recovery outcomes. These are the preferred 03R -test seams; unrelated TASK or ART lifecycle paths remain out of scope. - -### PLAN2 historical baseline - -Hosted Backend run `30854931616` on the final PR #249 tested tree -`19d48f7ea4bf20cb29f03cbba54f98683ce52661` produced: - -- 2,925 collected and completed tests; -- 20,793 covered statements of 23,475; -- 2,682 missed statements; -- 88.575080 percent global statement coverage; -- 640.284 seconds total backend wall time; -- 468.506 seconds in the slowest semantic lane. - -At the current denominator, 90 percent permits at most 2,347 missed statements. -The suite therefore needs 335 additional covered statements to reach 90 -percent and 394 to reach the required 90.25-percent pre-switch headroom. - -## Current CI behavior - -`.github/workflows/backend.yml` runs five semantic lanes, combines exactly five -coverage files, runs the real API contract drill, blocks below 78 percent -globally, and applies multiple 90-percent subsystem/per-file checks. Lane -custody, PostgreSQL isolation, and coverage fan-in are already implemented. - -`backend/scripts/coverage_policy.py` and -`backend/tests/test_coverage_contract.py` are merged historical integrity -machinery. The current workflow does not invoke that policy script. PLAN2 does -not wire it into CI or expand its static Python analysis. - -## Largest current gaps - -The latest hosted coverage JSON identifies these high-value gaps: - -| Module | Statements | Missing | Coverage | -|---|---:|---:|---:| -| `app/modules/projects/service.py` | 1,451 | 550 | 62.10% | -| `app/modules/checkers/service.py` | 579 | 169 | 70.81% | -| `app/modules/authorization/router.py` | 484 | 168 | 65.29% | -| `app/modules/artifacts/service.py` | 959 | 138 | 85.61% | -| `app/modules/tasks/service.py` | 682 | 108 | 84.16% | -| `app/modules/projects/repository.py` | 285 | 96 | 66.32% | -| `app/modules/artifacts/operator.py` | 204 | 80 | 60.78% | -| `app/modules/projects/router.py` | 178 | 63 | 64.61% | -| `app/modules/artifacts/guide_extraction_worker.py` | 237 | 65 | 72.57% | - -Smaller gaps exist in checker repository/router/runner/compiler, project setup -queue and policy replay, authorization read/repository code, auth API/deps/ -schemas, artifact extraction/materialization, background-job modules, and actor -services. - -## Existing ownership and test layers - -- Project behavior: `backend/tests/test_projects.py` and focused project files. -- Task behavior: `backend/tests/test_tasks.py`. -- Checker behavior: `backend/tests/test_checkers.py` and runner tests. -- Artifact behavior: focused artifact, storage, guide, and recovery tests. -- Authorization behavior: focused actor/authorization/API tests. -- Test isolation: `backend/scripts/run_isolated_tests.py`. -- Semantic execution: `backend/scripts/run_test_lanes.py`. - -The largest services depend directly on `AsyncSession`; this makes broad unit -extraction an architectural concern outside QUAL. Tests may use small typed -fakes or existing fixtures where behavior is observable, but QUAL must not -refactor production services merely to raise coverage. - -## Risks discovered - -- Adding hundreds of database-heavy covered lines could worsen the current - 10.7-minute hosted wall time. -- Testing implementation branches without outcomes can manufacture percentage - while adding little confidence. -- Raising the floor in the same PR as broad tests makes failures harder to - diagnose and encourages threshold bargaining. -- Concurrent AUTH, ART, and REV work can increase the denominator; the final - floor chunk must remeasure current `main` and retain headroom. - -## Conventions to preserve - -- Complete `backend/app` inventory and combined semantic-lane coverage. -- Real PostgreSQL for constraints, locks, migrations, transactions, triggers, - and concurrency. -- Real MinIO for the protocol boundary. -- Global 78-percent floor until the exact 90-percent switch merges. -- Existing protected 90-percent subsystem gates. -- Test-delta and CI-integrity review for every QUAL implementation PR. - -## Unknowns resolved per implementation chunk - -The exact missing lines and best observable tests must be refreshed from the -then-current hosted coverage JSON. A contract may not promise a coverage gain -from stale line numbers or require tests that merely execute code. +# Discovery: WS-QUAL-001 Behavior And Mutation Assurance + +## Current hosted truth + +Main Backend run `30926337804` on merge `5f2baf90` completed 3,162 tests with +21,620 / 23,938 statements covered (90.316651 percent), 620.264 seconds total +hosted wall time, and a 464.471-second slowest lane. The complete suite is above +90 percent, but `.github/workflows/backend.yml` intentionally blocks globally +at 78 percent and applies more than ten named 90-percent subsystem/per-file +checks. + +This means raising the global floor is neither necessary nor sufficient for the +new human goal. The remaining gap is whether assertions detect behavioral +changes. + +## Existing test-integrity controls + +| Control | Current implementation | What it proves | What it does not prove | +|---|---|---|---| +| Complete semantic lanes | `.github/workflows/backend.yml`, `backend/scripts/run_test_lanes.py` | Five canonical lanes collect and execute under isolated custody | Assertions are sensitive to faults | +| Evidence validation | `backend/scripts/validate_test_lane_evidence.py`, `merge_test_lane_evidence.py` | No missing lane, skipped node, missing coverage, or invalid bundle | Tests kill plausible defects | +| Global coverage | `coverage report --fail-under=78` | Complete app execution stays above the permitted baseline | Behavioral correctness | +| Protected coverage | Named `--fail-under=90` checks | New/material subsystems retain deeper execution | Assertions reject wrong outcomes | +| Weakening scan | `scripts/workstream_agent_gate.py` | Flags common skip/bypass/threshold suppression tokens | Semantic weakening expressed without those tokens | +| Internal review | QA, test-delta, CI integrity and other routed reviewers | Human/agent reasoning examines behavior and scope | Deterministic executable fault sensitivity | + +`backend/tests/test_project_policy_mutations.py` tests project-policy mutation +behavior; it is not a mutation-testing engine. No `mutmut`, Cosmic Ray, or +equivalent package/configuration currently exists in backend dependencies or +GitHub workflows. + +## Candidate engine evidence + +The current `mutmut` documentation says the tool supports pytest-aware test +selection, function/module wildcards, incremental results, parallel execution, +source-path restriction, and optional covered-line filtering. Its current +project metadata supports Python 3.10 through 3.14, which includes Workstream's +Python 3.11/3.12 range. It requires fork support, compatible with hosted Linux +runners. Sources: + +- +- + +Cosmic Ray is also viable and stores resumable mutation sessions, but its +configuration centers on explicit module paths and test commands, and its +official documentation notes that plugin options are not fully documented. +That creates more wrapper/configuration ownership for the first pilot: + +- +- + +Planning therefore selects `mutmut` only as the leading candidate. The pilot +must prove an exact pinned release, async pytest compatibility, deterministic +results, safe worktree isolation, and bounded hosted runtime before adoption. + +## Selection boundary + +Production changes can be derived from `origin/main...HEAD`. Test-only behavior +changes have no changed production file, so a deterministic behavior-claim +manifest is required to name the bounded production targets and owning tests. +Without this second path, a coverage-only test PR could avoid mutation +assurance entirely. + +The planned canonical boundary is +`.ci/behavior-claims/.json` under a repository-owned schema. It is +immutable PR content, not PR prose or a workflow input. Behavior claims name +repository-relative targets, qualified callables, owning pytest nodes, and +typed outcomes. Narrow non-behavioral test maintenance is classified through +the same schema so “no production diff” cannot become an implicit bypass. + +Eligible pilot targets should begin with pure functions or direct service +methods that have fast owning tests. Initial discovery candidates live in the +project/checker policy, compiler, and runner layers already exercised by 02R +and 03R. The implementation chunk must choose a much smaller representative +set from current main and record why each target is eligible. + +Ineligible-by-default categories for the pilot: + +- migrations and generated/declarative files; +- Pydantic/SQLAlchemy schemas whose mutations are primarily framework noise; +- composition-only modules and adapter wiring; +- external-effect adapters requiring network or real object storage per mutant; +- modules whose only truthful proof requires the full PostgreSQL/HTTP suite; +- unchanged modules not named by an explicit test-only behavior claim. + +## Evidence model + +A useful result must bind: + +- exact git tree/source digest; +- mutation engine version and configuration digest; +- target module/callable and owning test nodes; +- generated, killed, survived, timeout, suspicious, excluded, and error counts; +- stable mutant identifiers and classifications; +- command timeout and elapsed time; +- whether the result is pilot-only or blocking. + +A percentage without these facts is insufficient. Cache reuse is allowed only +when the wrapper proves the cached inputs match the exact current inputs. + +## Unknowns the pilot must answer + +- Whether current mutmut works cleanly with Workstream's async pytest fixtures. +- How precisely relevant tests are selected without broad incidental execution. +- Which mutation operators create equivalent/noisy results in Workstream code. +- Hosted runtime and p95 variability for representative changed targets. +- Whether fresh execution is cheap enough or authenticated cache reuse is + needed. +- Which narrow classification categories can be machine checked without + becoming an exclusion escape hatch. + +## Historical reconciliation + +PRs #103, #105, #108, #265, and #269 remain completed QUAL evidence. The old +01B2/milestone ladder remains superseded. `WS-QUAL-001-04R`, which proposed +raising the global floor to 90 percent, is superseded before implementation by +the human decision to keep 78 percent and move to behavior/mutation assurance. +Historical ENG-008 mutation planning is discovery input only; its retired +signed-loop and machine-scope requirements are not current authority. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md index 81a64b5b6..d4c8e64ac 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md @@ -1,49 +1,108 @@ -# Intent: WS-QUAL-001 Backend Coverage Floor +# Intent: WS-QUAL-001 Behavior And Mutation Assurance -## Human goal +## Problem being solved -Make the complete backend test suite protect at least 90 percent statement -coverage without weakening tests, hiding application files, or making CI slow -through unnecessary PostgreSQL and HTTP duplication. +Statement coverage shows that code executed, but it does not show that a test +would detect a meaningful behavioral defect. Workstream now exceeds 90 percent +global backend coverage while CI still correctly permits a 78-percent global +floor. The unfinished QUAL problem is assertion sensitivity and behavior +ownership, not another global percentage increase. -## Why this matters +## Why this work matters -Coverage is a backstop for behavior proof, not the goal by itself. Workstream's -authorization, artifact, project, task, checker, review, and contribution -boundaries need meaningful failure and recovery tests while the repository -remains practical for contributors. +Humans and agents can add tests that execute lines without proving returned +data, persisted state, denial, failure, recovery, audit, queue, or lifecycle +behavior. A trusted repository must reject those weak tests without making +every contributor run the entire backend once per mutant. -## Current truth +## Current behavior -The latest complete hosted result after ART-03C ran 2,925 tests and covered -20,793 of 23,475 application statements: 88.575080 percent. The -global CI floor remains 78 percent, while named new or materially changed -subsystems are already protected at 90 percent. +- Main Backend run `30926337804` on merge `5f2baf90` completed 3,162 tests and + covered 21,620 / 23,938 statements (90.316651 percent). +- The global blocking floor remains 78 percent. +- Named new or materially changed subsystems retain blocking 90-percent floors. +- Semantic lanes reject incomplete collection, skipped nodes, missing coverage, + and invalid evidence. +- Agent Gates scan common test and CI weakening tokens. +- No real mutation engine or mutation-result policy currently runs in CI. -## Success state +## Target behavior -- The exact complete backend suite covers at least 90.00 percent globally - across the complete importable `backend/app` inventory. -- GitHub CI blocks below a global `--fail-under=90` floor. -- New tests protect observable behavior, rejection, failure, or recovery. -- Pure or adapter-contract tests are preferred when PostgreSQL and HTTP are not - the behavior under test. -- Existing semantic lanes, isolation, coverage combination, and protected - 90-percent subsystem checks remain intact. +- Preserve the 78-percent global floor and every protected 90-percent floor. +- Require behavior claims to identify the production module and observable + outcome they protect. +- Mutation-test only eligible changed production logic or explicitly claimed + production targets for test-only behavior PRs. +- Treat surviving meaningful mutants as missing behavior proof, not as a reason + to increase statement coverage. +- Bound runtime, isolate mutation evidence from ordinary coverage, and keep the + complete Backend suite authoritative. +- Introduce blocking mutation policy only after a measured non-blocking pilot + proves deterministic selection, acceptable noise, and acceptable runtime. -## Non-goals +## Design chosen -- No production behavior, schema, migration, API, authorization, or product - lifecycle change. -- No arbitrary sharding or infrastructure purchase. -- No test deletion, weakened assertion, skip, xfail, coverage pragma, omit, or - narrowed application inventory. -- No revival of the historical signed-memory, base-evidence, semantic-parser, - line-budget, or per-milestone ratchet process. -- No promise that coverage alone proves correctness. +Use a two-stage rollout. First, pilot one pinned mutation engine with +deterministic target/test selection and complete non-blocking score evidence. +Second, after human review of pilot evidence, introduce a separate fail-closed +gate for eligible changed logic and explicit test-only behavior claims. The +blocking policy is survivor-based with reviewed classifications, not a global +mutation-score target. -## Human decision already provided +## Alternatives considered -The user directed the orchestrator to restart QUAL only after current-main -documentation reconciliation. This PLAN2 audit is authorized; implementation -still begins with the first reviewed bounded successor. +- Raise global coverage to 90 percent: rejected because coverage is already + above 90 and percentage alone does not prove assertion sensitivity. +- Mutate the full backend on every PR: rejected because runtime would be + unbounded and would discourage contribution. +- Require one killed mutant per test: rejected because it is easily gamed and + does not prove all eligible changed behavior. +- Immediately block on an uncalibrated mutation percentage: rejected because + equivalent/noisy mutants and infrastructure behavior must be measured first. +- Restore historical signed-loop or machine-scope machinery: rejected; those + systems were intentionally retired and are not prerequisites for quality. + +## Boundaries preserved + +- Coverage, real PostgreSQL, migration, trigger, lock, concurrency, MinIO, API, + and semantic-lane checks remain unchanged. +- QUAL owns test-assurance policy, evidence, and CI integration only. +- Production defects found by mutation testing move to the owning product + initiative; QUAL does not silently repair product behavior. +- Mutation targets exclude migrations, generated/declarative code, schemas, + adapters requiring external effects, and modules without an explicitly + reviewed eligibility rule during the pilot. + +## Expected risks + +- Mutation runtime can multiply test time. +- Equivalent or invalid mutants can create noisy false blockers. +- Target or test selection can be gamed to omit behavior. +- Test-only PRs need an explicit production target to avoid percentage padding. +- A new pinned tool adds dependency and supply-chain maintenance. + +## What must not change + +- Global coverage floor stays at 78 percent. +- Protected subsystem floors stay at 90 percent. +- No skips, xfails, coverage exclusions, assertion deletion, or narrower test + inventory. +- No mutation pragma or exclusion may be introduced casually to make CI pass. +- No full-repository mutation run is added to the normal PR critical path. + +## How this will be proven + +- Unit tests prove diff-to-target eligibility, explicit test-only claims, + classification grammar, evidence completeness, and fail-closed behavior. +- A hosted pilot records generated, killed, survived, timeout, suspicious, + excluded, and error outcomes with exact source/test identity and elapsed time. +- Known behavior tests must kill seeded representative mutants. +- Weak or vacuous fixture tests must leave representative mutants alive in the + policy regression suite. +- Existing Backend and Agent Gates remain green and unchanged in authority. + +## Human decisions required + +The 78-percent global floor decision is complete. A separate human checkpoint +is required after pilot evidence and before any mutation result becomes a +blocking merge gate. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md index b6aa7484e..7c1e439a9 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md @@ -1,77 +1,122 @@ -# Plan: WS-QUAL-001 Current-Main Coverage Closure +# Plan: WS-QUAL-001 Behavior And Mutation Assurance ## Approach -Retire the old milestone ladder and close the remaining gap with two declared -bounded test chunks followed by one floor switch. One additional owner-specific -test chunk is permitted only when 02R and 03R exhaust their meaningful gaps -without reaching the required headroom: - -1. Add fast project/setup behavior tests for observable service, repository, - routing, queue, and replay gaps. -2. If still needed, add fast checker behavior tests for observable service, - repository, runner, compiler, and routing gaps. -3. On current `main`, prove at least 90.25 percent globally and change only the - canonical GitHub global floor from 78 to 90. - -Each test chunk starts by reading the current hosted coverage JSON and selecting -behavioral gaps. It prefers pure functions, typed fakes, direct use-case calls, -and adapter contracts. PostgreSQL, MinIO, or HTTP is used only when that -boundary is itself the assertion. - -## Coverage target - -The last necessary test chunk must reach at least 90.25 percent before the CI -switch. This is operational headroom, not a permanent higher policy floor. If -concurrent main growth moves the measured result below 90.25 percent, the floor -chunk stops and returns to one owner-specific test plan; it never lowers or -rounds around the target. - -## Test-quality rule - -Every new test must name and assert one observable contract such as returned -data, persisted state, emitted audit/outbox fact, queue decision, mapped error, -authorization denial, idempotent replay, or recovery outcome. A test whose only -effect is executing previously missed lines is invalid. - -No chunk may introduce skips, xfails, coverage pragmas, omit/include narrowing, -deleted assertions, broad mocking of the behavior under test, or duplicated -database/HTTP coverage already owned by another layer. - -## Boundaries - -- QUAL changes tests and, only in the final chunk, the global CI threshold and - its lightweight invariant test. -- A production defect discovered by a stronger test is reported and fixed in a - separate owning initiative/chunk. -- Production service decomposition, repository ports, UnitOfWork design, type - checking, mutation testing, and property-test architecture require separate - initiatives. They are not hidden inside coverage closure. -- CI runtime optimization remains WS-CI-owned. QUAL records test-time impact and - must avoid obvious regressions but does not redesign lane infrastructure. +Retire the proposed global-90 floor switch and deliver behavior assurance in +two independently reviewed implementation chunks. + +### Stage 1: changed-scope mutation pilot + +Add one exactly pinned mutation engine and a Workstream-owned policy wrapper. +The wrapper derives a closed set of eligible production targets from the git +delta or from an explicit test-only behavior claim. It selects the smallest +owner test set, runs under a hard timeout, and emits machine-readable exact-head +evidence. + +The pilot does not block on mutation score. It does block on infrastructure +failure, malformed evidence, target escape, missing claimed tests, ordinary +test failure, or any weakening of existing Backend checks. Pilot results must +distinguish killed, survived, timeout, suspicious, excluded, and error mutants. + +`mutmut` is the leading pilot candidate because its current documentation +supports pytest-aware test selection, function/module wildcards, incremental +results, parallel execution, source/selection configuration, covered-line +filtering, and Python 3.11/3.12. The implementation chunk must still prove a +pinned release against Workstream's async pytest and isolated-service setup; +planning does not pre-approve an unusable dependency. + +### Stage 2: blocking behavior-mutation gate + +Only after pilot review, add a separate required check for eligible changed +production logic and test-only PRs that claim behavioral improvement. The gate +uses the pilot's deterministic target and evidence grammar. + +There is no repository-wide mutation percentage. Every eligible survivor +blocks unless it has a narrow, typed classification accepted by policy (for +example, demonstrably equivalent or non-behavioral). Missing, stale, broad, or +free-form exclusions fail closed. Timeout and tool errors do not count as +killed mutants and cannot silently pass. + +## Behavior ownership + +A qualifying behavior claim identifies: + +- production module and callable or bounded target; +- owning test nodes; +- observable contract (return, persisted state, emitted fact, denial, mapped + error, idempotent replay, or recovery outcome); +- relevant real boundary, if PostgreSQL, MinIO, HTTP, lock, trigger, or + concurrency is essential. + +The canonical input is a schema-v1 JSON file at +`.ci/behavior-claims/.json`, validated by a repository-owned schema +and policy parser. Chat, PR prose, labels, workflow inputs, and environment +variables cannot widen targets. Behavior claims contain repository-relative +production targets, qualified callables, owning pytest node IDs, and typed +observable outcomes. Test-only non-behavioral maintenance uses a narrow typed +classification defined by policy rather than free-form exemption text. + +Test-only changes that claim coverage or stronger behavior must provide this +mapping. Documentation-only, fixture-only, generated-code, and non-behavioral +maintenance changes are outside mutation selection but remain subject to +ordinary tests and review. + +## Runtime and isolation strategy + +- Never mutate the full backend in ordinary PR CI. +- Start with pure or direct-service logic whose owning tests avoid PostgreSQL + and HTTP unless those boundaries are the behavior being proved. +- Run mutation work independently from the existing Backend critical path. +- Pilot command limit: 12 minutes inside a 15-minute job limit. +- A blocking rollout must demonstrate a practical hosted p95 and cannot extend + required PR latency by more than two minutes when run in parallel. +- Mutation caches are acceleration only; evidence binds the exact source, + configuration, selected tests, tool version, and result set. + +## Dependency and evidence integrity + +- Pin the selected engine and its transitive dependency closure with hashes. +- Do not add the mutation engine to production dependencies. +- Install the engine only from `scripts/mutation-requirements.txt` with + `pip install --require-hashes`; `backend/pyproject.toml` may contain tool + configuration but cannot add the engine to ordinary dev extras. +- Never apply mutants to the contributor worktree in CI. +- Upload bounded result evidence without source secrets, environment values, + database contents, or artifact payloads. +- The policy wrapper, not mutable PR prose, determines eligibility and validates + results. +- Mutation CI runs only on an unprivileged `pull_request`/`push` boundary with + explicit read-only permissions, pinned Actions, checkout credentials + disabled, no secrets or writable token in the mutation subprocess, and + bounded non-restorable artifacts/caches. ## Alternatives rejected -- Reviving `01B2` and the complex base-evidence ratchet: unnecessary now that - exact lane custody and hosted coverage evidence exist. -- One large cross-owner coverage PR: crosses project, checker, task, artifact, and - authorization ownership and is difficult to review. -- Raising the floor immediately: current measured coverage is below 90. -- Excluding low-coverage services or files: makes the global percentage false. -- More arbitrary shards: changes runtime distribution, not test architecture or - coverage quality. +- `WS-QUAL-001-04R` global floor switch: superseded before implementation. +- Full-suite-per-mutant execution: too slow and poorly owned. +- Score-only gating: hides which behavior remains unproved. +- Non-blocking forever: measures quality without protecting it. +- Immediate blocking rollout: lacks runtime and equivalent-mutant calibration. +- Mutating only covered lines as the sole eligibility rule: can hide untested + changed behavior; covered-line filtering may optimize the pilot but cannot + define the full policy. ## Verification strategy -Every implementation chunk runs focused tests, Ruff for changed tests, complete -test-delta review, relevant stale-contract checks, and hosted Backend. The final -floor chunk additionally proves the combined coverage JSON covers the complete -application inventory at or above 90.25 percent and that every protected -90-percent check remains blocking. +Each implementation chunk runs focused policy tests, mutation-engine smoke +tests, Ruff, Agent Gates, Markdown/stale scans, internal reviewer tracks, and +hosted Backend. The pilot additionally proves at least one known strong test +kills its representative mutants and at least one deliberately weak test leaves +a representative mutant alive. The blocking chunk proves survivors, timeouts, +errors, missing evidence, stale evidence, and target escape all stop the gate. ## Dependency order -PLAN2 -> 02R -> optional 03R -> 04R. If those exact owner-scoped chunks do not -provide enough headroom, stop and plan one additional owner-specific test chunk -from the refreshed report. Do not create a percentage-driven residual bucket. -The CI floor change always remains a separate final PR. +`PLAN3 -> 04M pilot -> human calibration checkpoint -> 05M blocking gate`. +`05M` cannot begin from planning alone; it requires accepted exact hosted pilot +evidence and a new explicit human instruction. + +## Stop + +Planning does not install a mutation engine, change a workflow, or change a +coverage threshold. Stop after the PLAN3 PR and human checkpoint. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md index 6e5f2134b..ad6cb818f 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/RISKS.md @@ -1,14 +1,22 @@ -# Risks: WS-QUAL-001 Current-Main Coverage Closure +# Risks: WS-QUAL-001 Behavior And Mutation Assurance | Risk | Consequence | Control | |---|---|---| -| Coverage-only tests | Higher percentage without stronger behavior proof | Require observable outcomes and QA/test-delta review | -| More database/HTTP tests | Backend CI becomes slower | Prefer pure/use-case/adapter-contract tests unless the real boundary is essential | -| Concurrent denominator growth | Candidate falls below 90 before floor merge | Remeasure current main and require >=90.25% headroom before 04R | -| Threshold bundled with tests | Harder diagnosis and pressure to bargain | Keep the 90-percent switch in separate chunk 04R | -| Production defect discovered | QUAL scope expands into product repair | Stop and hand defect to owning initiative | -| Historical parser revival | Reintroduces complexity and maintenance burden | Mark 01B2 and old replacements superseded | -| File exclusion or pragma | False global measurement | Preserve complete app inventory and existing stale/coverage guards | -| Duplicate invariant tests | Slower suite and ambiguous ownership | Map each new test to its owning layer and review existing proof first | +| Full-repository mutation | CI becomes unusably slow | Mutate only eligible changed or explicitly claimed targets under hard limits | +| Equivalent/noisy mutants | Correct PRs are blocked without quality benefit | Non-blocking pilot, typed classifications, separate human checkpoint before enforcement | +| Score gaming | Contributors kill one easy mutant or raise a percentage while behavior remains weak | Survivor-based exact evidence; no “one mutant” or global score success rule | +| Target-selection escape | Important changed logic is silently omitted | Workstream-owned deterministic diff/claim parser; missing or broad evidence fails closed | +| Test-selection escape | Mutants pass because relevant tests were not selected | Bind explicit owning nodes, baseline-run them first, validate selection in policy tests | +| Test-only coverage padding | No production diff means no mutation work | Require bounded production targets for test-only behavior/coverage claims | +| Timeout treated as success | Hanging mutants silently pass | Timeout is a distinct non-killed outcome and blocks in enforcement mode | +| Cache poisoning/staleness | Results do not describe the current source | Bind tree, config, tool, targets, tests, and result digests; cache is acceleration only | +| Mutation exclusions spread | Meaningful behavior is hidden | No source pragmas in pilot; classifications are narrow, typed, reviewed evidence | +| Dependency compromise | CI executes an untrusted tool closure | Exact pin and hash-lock the development-only dependency closure | +| Worktree mutation | Contributor source is left modified | Execute in disposable isolated workspace and verify tree custody | +| Existing gates weakened | Mutation becomes a substitute for real tests | Preserve semantic lanes, full suite, API E2E, 78 global and protected 90 floors | +| Runtime critical-path increase | Contribution slows despite scoped execution | Independent job, 12-minute command/15-minute job bounds, <=2-minute critical-path objective | +| Production defect discovered | QUAL scope drifts into product repair | Stop and hand the defect to its owning initiative | -No secret, credential, deployment, payment, or production-data access is needed. +No secret, production credential, production data, payment, or deployment +access is required. The CI dependency and executable-tool boundary requires +security and CI-integrity review. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md index ae60423b5..68432a496 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/STATUS.md @@ -1,41 +1,31 @@ -# Status: WS-QUAL-001 Backend Coverage Floor +# Status: WS-QUAL-001 Behavior And Mutation Assurance ## Current state -`WS-QUAL-001-PLAN2` merged through PR #260. `WS-QUAL-001-02R` merged through -PR #265 with all 2,996 exact-head tests passing and 21,004 / 23,455 statements -covered (89.550203 percent). +Coverage closure through `WS-QUAL-001-03R` is complete. PR #269 merged at +`5f2baf90`; main Backend run `30926337804` completed all 3,162 tests and +reported 21,620 / 23,938 statements (90.316651 percent), 620.264 seconds hosted +wall time, and a 464.471-second slowest lane. -The current-main baseline after ART PR #268 is Backend run `30921410531` on -`5b853d50`: 3,068 tests completed, 21,453 / 23,938 statements covered -(89.619016 percent), 727.166 seconds hosted wall time, and a 567.994-second -slowest lane. The global CI floor remains 78 percent; named protected subsystem -checks remain blocking at 90 percent. - -Historical QUAL work delivered the isolated database runner and test-integrity -guards through PRs #103, #105, and #108. The many stopped semantic-analysis -replacements remain historical evidence, not work to resume. +The global blocking floor remains 78 percent by explicit human decision. Named +new or materially changed subsystem checks remain blocking at 90 percent. ## Current gate -`WS-QUAL-001-03R` is the current implementation chunk. It must add meaningful -checker-owned behavior tests and gain at least 152 covered statements on the -current-main denominator to reach the 90.25-percent headroom target. The -focused test union measures -168 unique previously missing checker statements and projects 21,621 / 23,938, -or 90.320829 percent. Require Agent -Gates, CodeRabbit, all Backend semantic lanes, final coverage fan-in, six -internal reviewer tracks, and human review. An unexplained focused or hosted -runtime increase above 10 percent stops merge readiness. - -The 94-case focused implementation selection passes in 37.76 seconds with -narrow checker-module coverage. The isolated full `test_checkers.py` run -reached the 1,200-second local ceiling after approximately 95 percent completion -with 163 passing tests and no failure output. It is not a complete pass; hosted -Backend remains mandatory. +`WS-QUAL-001-PLAN3` is planning only. It replaces the unstarted 04R global-floor +switch with a two-stage behavior/mutation assurance proposal: + +1. `04M` — bounded, pinned, changed-scope mutation pilot with complete evidence + and no blocking score. +2. Human calibration checkpoint. +3. `05M` — separately approved blocking survivor policy for eligible changed + logic and explicit test-only behavior claims. + +No mutation dependency, workflow, policy script, or blocking check has been +implemented. ## Stop condition -Planning does not change tests, application code, workflow code, or thresholds. -Do not raise the global floor until hosted combined coverage is at least 90.25 -percent on the exact candidate head. +Stop after PLAN3 planning review and PR. Do not start 04M automatically. Do not +start 05M without accepted exact hosted pilot evidence and a new human +instruction. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md index ad3b62973..8a3d8d2c1 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/README.md @@ -1,13 +1,14 @@ # QUAL Chunk Records -Only these PLAN2 records describe possible current work: +Only these PLAN3 records describe possible current work: -- `WS-QUAL-001-PLAN2-current-main-reconciliation.md` -- `WS-QUAL-001-02R-project-setup-behavior-coverage.md` -- `WS-QUAL-001-03R-checker-behavior-coverage.md` -- `WS-QUAL-001-04R-global-90-floor.md` +- `WS-QUAL-001-PLAN3-behavior-mutation-assurance.md` +- `WS-QUAL-001-04M-changed-scope-mutation-pilot.md` +- `WS-QUAL-001-05M-blocking-behavior-mutation-gate.md` Every other file in this directory is historical evidence from the original QUAL plan, a completed chunk, a stopped repair attempt, or a superseded -contract. Historical records cannot be started or treated as current +contract. In particular, `WS-QUAL-001-04R-global-90-floor.md` was superseded +before implementation by the human decision to retain the 78-percent global +floor. Historical records cannot be started or treated as current implementation authority. `CHUNK_MAP.md` records their final disposition. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md new file mode 100644 index 000000000..8665a5051 --- /dev/null +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04M-changed-scope-mutation-pilot.md @@ -0,0 +1,132 @@ +# Chunk Contract: WS-QUAL-001-04M — Changed-Scope Mutation Pilot + +## Parent initiative + +`WS-QUAL-001` — Behavior And Mutation Assurance + +## Goal + +Pilot one exactly pinned mutation engine on eligible changed or explicitly +claimed production targets and publish complete exact-head result evidence +without imposing an uncalibrated mutation-score gate. + +## Why this chunk exists + +Coverage proves execution, not assertion sensitivity. Workstream needs measured +compatibility, mutant quality, selection integrity, and hosted runtime before a +mutation outcome can block contributions. + +## Approved plan reference + +- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md` +- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md` +- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md` + +## Risk class + +L1 — CI executable/dependency/evidence policy. + +## SLA + +P2. + +## Allowed files + +```text +backend/pyproject.toml +backend/scripts/mutation_policy.py +backend/tests/test_mutation_policy.py +scripts/git_delta.py +scripts/test_git_delta.py +scripts/workstream_agent_gate.py +scripts/behavior-claim.schema.json +scripts/mutation-requirements.txt +scripts/test_lightweight_agent_gates.py +.ci/behavior-claims/WS-QUAL-001-04M.json +.ci/behavior-claims/README.md +.github/workflows/mutation-pilot.yml +docs/operations_backend_testing.md +.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/** +``` + +## Not allowed + +```text +backend/app/** or backend/alembic/** changes +global or blocking mutation percentage +change to the global 78-percent or protected 90-percent coverage floors +replacement, reduction, or bypass of Backend semantic lanes, API E2E, or fan-in +full-backend mutation on ordinary PRs +mutation pragmas or free-form exclusion lists +production dependency changes +pull_request_target, privileged PR-code execution, writable workflow token, +checkout credentials, secrets in mutation execution, or unpinned Actions +``` + +## Acceptance criteria + +- [ ] One engine and transitive closure are exactly pinned and hash locked as + development/CI-only dependencies, installed exclusively from + `scripts/mutation-requirements.txt` with `--require-hashes`. +- [ ] `backend/pyproject.toml` contains configuration only; the mutation engine + is absent from production dependencies and ordinary dev extras. + `scripts/mutation-requirements.txt` is the sole mutation-tool dependency + authority; `backend/uv.lock` remains unchanged and is not a second install + path. +- [ ] Deterministic policy selects eligible changed targets or validates a + bounded test-only behavior claim with explicit owning test nodes. +- [ ] Git-delta discovery extracts one shared `scripts/git_delta.py` primitive + reused by `scripts/workstream_agent_gate.py` and mutation policy, and + mutation evidence mirrors the + existing semantic-lane exact-tree/digest/fail-closed conventions rather + than creating a parallel custody dialect. +- [ ] Schema-v1 `.ci/behavior-claims/.json` is the only test-only + claim input; PR prose, labels, workflow inputs, and environment variables + cannot widen production targets or owning test nodes. +- [ ] `.ci/behavior-claims/README.md` and the backend testing operations guide + document the pilot format without presenting it as a blocking contributor + requirement before 05M. +- [ ] Disposable execution cannot leave mutants in the checked-out source tree. +- [ ] Evidence binds exact tree, tool/config, target/test identities, elapsed + time, and generated/killed/survived/timeout/suspicious/excluded/error + outcomes. +- [ ] Known strong behavior tests kill representative mutants and a deliberately + weak fixture leaves a representative mutant alive. +- [ ] Score is observational; infrastructure errors, malformed/stale evidence, + target escape, or baseline test failure remain blocking. +- [ ] Mutation command is bounded to 12 minutes inside a 15-minute independent + job and records critical-path impact. +- [ ] Workflow runs untrusted PR code only through `pull_request`/`push`, uses + explicit read-only permissions, pinned Actions, `persist-credentials: + false`, no secrets or writable token in mutation execution, and bounded + non-restorable artifacts/caches; invariant tests enforce each property. +- [ ] Full Backend and existing coverage gates remain unchanged and green. + +## Verification commands + +```bash +cd backend +.venv/bin/python -m pytest -q tests/test_mutation_policy.py +.venv/bin/ruff check scripts/mutation_policy.py tests/test_mutation_policy.py +cd .. +python3 scripts/check_markdown_links.py +PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python3 -m pytest -q scripts/test_lightweight_agent_gates.py +git diff --check +``` + +## Required reviewers + +Senior engineering, QA/test, security/auth, product/ops, architecture, CI +integrity, docs, reuse/dedup, and test delta. + +## Human review focus + +- Is target/test selection deterministic and non-gamable? +- Is result evidence complete enough to calibrate a later blocking policy? +- Is hosted runtime practical without weakening Backend? + +## Stop conditions + +Stop if the engine cannot be pinned, mutates the contributor worktree, requires +full-suite-per-mutant execution, produces unclassifiable noise, exceeds runtime +bounds, or requires weakening any existing gate. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md index cd9b52f5f..fdadd57f1 100644 --- a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-04R-global-90-floor.md @@ -1,5 +1,9 @@ # Chunk Contract: WS-QUAL-001-04R — Global 90 Percent CI Floor +> Superseded before implementation by `WS-QUAL-001-PLAN3`. The human decision +> retains the 78-percent global floor and moves QUAL to behavior/mutation +> assurance. This file is historical evidence and must not be started. + Parent initiative: `WS-QUAL-001` Goal: after exact hosted proof at or above 90.25 percent, change the canonical diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md new file mode 100644 index 000000000..432938d6a --- /dev/null +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-05M-blocking-behavior-mutation-gate.md @@ -0,0 +1,93 @@ +# Chunk Contract: WS-QUAL-001-05M — Blocking Behavior-Mutation Gate + +## Parent initiative + +`WS-QUAL-001` — Behavior And Mutation Assurance + +## Goal + +After accepted pilot evidence and separate human approval, make complete +mutation outcomes blocking for eligible changed production logic and explicit +test-only behavior claims. + +## Why this chunk exists + +The pilot measures feasibility. This separate chunk converts only calibrated, +deterministic evidence into contributor protection. + +## Approved plan reference + +- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md` +- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md` +- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md` +- Required input: accepted exact hosted `WS-QUAL-001-04M` pilot evidence. + +## Risk class + +L1 — blocking CI/test policy. + +## SLA + +P2. + +## Allowed files + +Exact files must be refreshed from the merged 04M implementation before start. +Expected ownership is limited to its mutation policy/tests, independent +workflow, backend testing operations guide, Agent Gate invariant, and QUAL +initiative evidence. Test-only inputs use only schema-v1 +`.ci/behavior-claims/.json` files validated by the merged policy. +The refreshed allowed list must explicitly include `CONTRIBUTING.md`, +`.ci/behavior-claims/README.md`, the policy-owned schema, and a copyable example +so external contributors see the exact blocking workflow before it is enabled. + +## Not allowed + +```text +start without accepted 04M evidence and explicit human instruction +global mutation percentage +change to 78-percent global or protected 90-percent coverage floors +free-form exemptions, source mutation pragmas, silent timeout/error success +full-repository mutation on ordinary PRs +production behavior, migration, or dependency changes +``` + +## Acceptance criteria + +- [ ] Eligibility and evidence grammar are unchanged from accepted pilot proof + unless a separately reviewed correction is explicit. +- [ ] Every eligible surviving mutant blocks by default. +- [ ] Any allowed classification is narrow, typed, evidence-bound, and tested; + missing, stale, broad, or free-form classifications fail closed. +- [ ] Timeout, suspicious, and error outcomes never count as killed. +- [ ] Test-only behavior/coverage claims cannot bypass target mutation. +- [ ] `CONTRIBUTING.md` and the canonical claim README/schema/example explain + when a claim is required, the permitted typed non-behavioral cases, local + verification, evidence interpretation, and repair of surviving mutants. +- [ ] Non-eligible maintenance/docs/generated changes do not run irrelevant + mutants. +- [ ] Hosted p95 and critical-path impact satisfy the accepted pilot bound. +- [ ] Backend, Agent Gates, 78-percent global floor, and protected 90-percent + floors remain authoritative and green. + +## Verification commands + +Refresh exact commands from merged 04M; at minimum run mutation-policy unit and +integration tests, strong/weak seeded behavior proof, Ruff, Agent Gates, +Markdown/stale checks, full hosted Backend, and exact blocking-workflow proof. + +## Required reviewers + +Senior engineering, QA/test, security/auth, product/ops, architecture, CI +integrity, docs, reuse/dedup, and test delta. + +## Human review focus + +- Does the gate block weak behavior proof without blocking unrelated work? +- Can classifications or test selection be used as an escape hatch? +- Does the gate remain practical for external contributors? + +## Stop conditions + +Stop on missing pilot evidence, unreviewed policy change, unacceptable hosted +latency/noise, coverage/Backend weakening, or need for broad exemptions. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md new file mode 100644 index 000000000..e08742ded --- /dev/null +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/chunks/WS-QUAL-001-PLAN3-behavior-mutation-assurance.md @@ -0,0 +1,95 @@ +# Chunk Contract: WS-QUAL-001-PLAN3 — Behavior And Mutation Assurance Planning + +## Parent initiative + +`WS-QUAL-001` — Behavior And Mutation Assurance + +## Goal + +Replace the unstarted global-90 floor proposal with reviewed, bounded planning +for behavior-owned changed-scope mutation assurance while retaining the global +78-percent floor. + +## Why this chunk exists + +Main already exceeds 90-percent statement coverage, but no executable gate +proves assertion sensitivity. Planning must define a safe pilot and a separate +calibrated enforcement step before workflow or dependency changes. + +## Approved plan reference + +- INTENT: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/INTENT.md` +- PLAN: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/PLAN.md` +- CHUNK_MAP: `.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/CHUNK_MAP.md` + +## Risk class + +L1 — CI/test policy planning. + +## SLA + +P2. + +## Allowed files + +```text +.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/** +``` + +## Not allowed + +```text +application or test implementation +workflow, dependency, threshold, coverage, or mutation-engine changes +revival of retired signed-loop, merge-intent, or machine-scope machinery +automatic start of 04M or 05M +``` + +## Acceptance criteria + +- [ ] Current main coverage, runtime, and existing test-integrity gates are + recorded exactly. +- [ ] The global 78-percent and protected 90-percent floors are preserved. +- [ ] Planning distinguishes behavior evidence from statement coverage. +- [ ] Pilot selection covers changed production targets and explicit test-only + behavior claims without full-repository mutation. +- [ ] Runtime, dependency, evidence, classification, cache, and worktree-safety + risks have fail-closed controls. +- [ ] Blocking rollout requires accepted pilot evidence and a new human gate. +- [ ] Required internal planning reviewers pass with no open sessions. + +## Verification commands + +```bash +python3 scripts/check_markdown_links.py +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_stale_authorization_docs.py +python3 scripts/check_stale_artifact_contracts.py +PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python3 -m pytest -q scripts/test_lightweight_agent_gates.py +git diff --check +``` + +## Required reviewers + +- senior engineering +- QA/test +- security/auth +- product/ops +- architecture +- CI integrity +- docs +- reuse/dedup +- test delta + +## Human review focus + +- Confirm 78 percent remains the global floor. +- Confirm mutation assurance measures behavior rather than another percentage. +- Confirm the pilot is bounded and the blocking rollout has a separate human + checkpoint. + +## Stop conditions + +Stop if planning implies immediate blocking mutation, full-repository mutation, +coverage/Backend weakening, an unbounded dependency/runtime commitment, or any +implementation change. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md new file mode 100644 index 000000000..498a2a36d --- /dev/null +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-internal-review-evidence.md @@ -0,0 +1,87 @@ +# WS-QUAL-001-PLAN3 Internal Review Evidence + +## Review scope + +Planning-only changes under +`.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/**`. Review began +against main merge `5f2baf90036839589cf3db5ad6949d4888da5e9e` and reconciled without conflict +to current main `24f677cb` after CON PLAN5 PR #270 changed only CON/REV/general +documentation. The QUAL delta is unchanged. No application, test, workflow, +dependency, threshold, or mutation implementation changed. + +## Risk routing + +- Risk class: L1. +- SLA: P2. +- Work type: CI/test policy architecture and documentation. +- Required reviewers: senior engineering, QA/test, security/auth, product/ops, + architecture, CI integrity, docs, reuse/dedup, and test delta. +- Human gate: required for PLAN3 merge, required again before 04M, and required + after pilot evidence before 05M. +- Budget posture: high scrutiny with bounded planning scope; no implementation + or runtime spend in PLAN3. +- Why: future work executes third-party tooling against untrusted PR code and + may eventually become a blocking contributor gate. + +## Circuit breaker + +PASS with a planning-record size exception. The diff is larger than the +preferred L1 implementation guideline because initiative planning must keep +intent, discovery, plan, decisions, risks, status, chunk map, three contracts, +and review evidence mutually consistent in one planning PR. It touches one +major boundary (QUAL CI/test policy), changes no executable file, and has clear +acceptance criteria and nine completed reviewer tracks. Splitting the records +would publish internally contradictory partial policy without reducing the +eventual 04M or 05M implementation boundary. + +## Reviewer results + +| Reviewer | Result | Final findings | +|---|---:|---| +| senior engineering | PASS | None | +| QA/test | PASS | None | +| security/auth | PASS | Prior CI privilege and dependency-custody findings resolved | +| product/ops | PASS | None | +| architecture | PASS | Shared git-delta boundary and canonical claim input conditions resolved | +| CI integrity | PASS | None | +| docs | PASS WITH LOW RISKS | Dependency authority and contributor onboarding conditions resolved | +| reuse/dedup | PASS WITH LOW RISKS | Existing delta/evidence conventions are now explicit reuse requirements | +| test delta | PASS | None | + +Open reviewer sessions: none. + +## Findings resolved + +- Added one shared `scripts/git_delta.py` boundary for Agent Gates and mutation + policy rather than parallel diff parsing. +- Declared schema-v1 `.ci/behavior-claims/.json` as the only test-only + claim input; mutable PR prose, labels, workflow inputs, and environment + variables cannot widen scope. +- Required mutation tooling to install only from a hash-locked sidecar + requirements file; `pyproject.toml` is configuration-only and `uv.lock` is + not a second install path. +- Required unprivileged `pull_request`/`push`, explicit read-only permissions, + pinned Actions, disabled checkout credentials, no secrets/writable token in + the mutation subprocess, and bounded artifacts/caches. +- Required canonical claim documentation during the pilot and explicit + `CONTRIBUTING.md` onboarding before any blocking rollout. + +## Deterministic evidence + +- Markdown link scan — passed for all changed Markdown files. +- Stale Workstream wording scan — passed. +- Stale authorization documentation scan — passed. +- Stale artifact-contract scan — passed. +- Lightweight Agent Gates — 10 passed. +- `git diff --check` — passed. +- Allowed scope — QUAL initiative planning tree only. +- Official candidate-tool assumptions were checked against current mutmut and + Cosmic Ray primary documentation. + +## Remaining risks + +- `mutmut` is provisional; 04M must prove exact pinned compatibility with + Workstream's async pytest and isolation setup. +- Equivalent/noisy mutants and hosted runtime remain unknown until 04M. +- No mutation result may block merging until exact hosted pilot evidence is + accepted through a separate human checkpoint. diff --git a/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md new file mode 100644 index 000000000..41e371317 --- /dev/null +++ b/.agent-loop/initiatives/WS-QUAL-001-backend-coverage-floor/reviews/WS-QUAL-001-PLAN3-pr-trust-bundle.md @@ -0,0 +1,112 @@ +# WS-QUAL-001-PLAN3 PR Trust Bundle + +## Chunk + +`WS-QUAL-001-PLAN3` — Behavior And Mutation Assurance Planning. + +## Goal + +Keep the global Backend floor at 78 percent and replace the unstarted +percentage-only floor raise with a safe path for behavior-owned mutation +assurance. + +## Human-approved intent + +Coverage remains a permitted baseline, not proof that tests verify behavior. +The user confirmed 78 percent and directed planning for behavior and mutation +quality before implementation. + +## What changed and why + +- Reframed QUAL from “raise global coverage” to “prove assertion sensitivity.” +- Marked unstarted `WS-QUAL-001-04R` superseded. +- Declared `04M`, a bounded non-blocking-score mutation pilot. +- Declared `05M`, a separately approved blocking survivor policy only after + accepted hosted pilot evidence. +- Added current-main facts, security/runtime risks, claim ownership, evidence + rules, contributor boundaries, and exact chunk contracts. + +## Design chosen + +Mutation targets are eligible changed production logic or explicit test-only +behavior claims. Claims use schema-v1 +`.ci/behavior-claims/.json`; mutable PR prose cannot widen them. +`04M` runs independently under hard limits and records complete outcomes. +`05M` is not authorized until pilot evidence is accepted by a human. + +## Alternatives rejected + +- Raising global coverage from 78 to 90: execution percentage is not behavior + proof. +- Full-backend mutation per PR: operationally impractical. +- A global mutation score or “one mutant killed” rule: gameable and opaque. +- Immediate blocking rollout: uncalibrated noise and runtime. +- Retired signed-loop/machine-scope machinery: unnecessary for current simple + contribution flow. + +## Scope control + +Only QUAL initiative planning records change. PLAN3 installs no tool, changes +no workflow or dependency, modifies no tests/application code, and changes no +coverage threshold. + +## Product behavior + +None. Product review decisions remain `accept`, `needs_revision`, and `reject`; +mutation outcomes are engineering evidence and never product decisions. + +## Acceptance criteria proof + +- Current main is recorded from Backend run `30926337804`: 3,162 completed + tests, 21,620 / 23,938 coverage (90.316651 percent), 620.264 seconds wall, + and 464.471 seconds slowest lane. +- Global 78 and protected 90 floors are explicitly preserved. +- Changed-production and test-only behavior paths are both planned. +- Strong-vs-weak seeded mutant proof, complete outcome evidence, runtime bounds, + dependency custody, CI privilege controls, and a second human checkpoint are + explicit acceptance requirements. + +## Tests and checks run + +Markdown links, all three stale scans, ten lightweight Agent Gate regression +tests, scope review, and whitespace validation pass. + +## Test delta and CI integrity + +No tests or CI files change. Future contracts forbid skips, xfails, assertion +weakening, coverage exclusions, Backend replacement, and changes to the 78/90 +coverage policy. + +## Reviewer results + +All nine required tracks pass after resolving shared-helper, claim-boundary, +CI-privilege, dependency-custody, and contributor-onboarding conditions. No +reviewer session remains open. + +## External review + +Agent Gates, CodeRabbit, and human review are pending publication. Backend is +not required by the planning diff unless GitHub policy schedules it; no Backend +file changes. + +## Remaining risks + +Tool compatibility, equivalent-mutant noise, target/test selection quality, +and hosted runtime remain deliberately assigned to 04M pilot evidence. + +## Follow-up work + +After PLAN3 merges, 04M may start only by explicit human instruction. 05M +requires accepted 04M hosted evidence and a separate explicit decision. + +## Human review focus + +- Confirm the 78-percent global floor remains unchanged. +- Confirm this measures behavior rather than another percentage. +- Confirm the pilot is safe for untrusted contributor code and blocking remains + behind a second human checkpoint. + +## Human merge ownership + +GitHub checks and explicit human approval are required. PLAN3 authorizes no +mutation implementation.