Problem
The first real runs of aidd-qa:01-acceptance-qa exposed gaps that still force the operator to improvise. The current contract does not fully define how to select a browser entry, prove test-only storage before starting it, configure a mobile or permission-sensitive session, assess qualitative or short-lived outcomes, or preserve the user-facing flow when QA is delegated.
Some follow-ups are already complete and are not part of this issue: the runner supports a visually verified tail cut for command latency, and #925 removed the one-time release-as: 1.0.0 pin.
Outcome
As a reviewer, I want Acceptance QA to prepare, execute, and report real browser journeys from explicit evidence, so desktop and mobile criteria produce safe, reproducible results without host-specific improvisation or overstated proof.
Scope
Preparation and safety
- Resolve entry, authentication, fixtures, reset, viewport, and browser permissions from the
Browser QA project memory first, then a related browser test or fixture. Separate resolution from starting the application.
- Treat storage as proven test-only only when project memory explicitly says so, an isolated test fixture/configuration establishes it, or the user confirms it. Environment names such as
test or staging are not proof by themselves.
- When reset can delete data, prove test-only storage before starting the entry or taking a browser snapshot. If proof is absent, ask once and do not start.
- Validate that the selected entry serves the expected browser UI. If it fails or starts only a backend, try at most one browser entry explicitly documented by the same evidence chain; otherwise report
blocked with the shortest decisive error. Never invent a start command.
- Define a destructive action as one that can alter data outside the isolated scenario fixture or cause an irreversible external side effect. A verified teardown limited to test-only data created by the scenario is not destructive.
- Resolve viewport in this order: acceptance criteria,
Browser QA memory, then the existing 1280×720 fallback. When mobile behavior is required but no dimensions are declared, ask once instead of guessing.
- Resolve required permission and consent state before recording. Set it explicitly when criteria or project memory declare it; otherwise retain the browser default and record that fact in the prepared run.
Observable and timed criteria
- Turn qualitative wording such as “jumps” or “flashes” into an explicit browser-observable condition before recording. Ask once when more than one materially different interpretation exists; reject the scenario when no observable condition can be established.
- When a criterion names a timing boundary, define probes on both sides, record their chosen margins, and report measured values. Ask once when the margin could change the product meaning; never invent a product tolerance silently.
- Assess short-lived transitions inside the single
run-code execution with a bounded browser observation and include the measured result in Actual. Frame sampling validates the video artifact but is not sufficient behavioral proof for a transition shorter than the sampling interval.
Flow, discovery, and host portability
- Resolve an existing feature evidence folder through its
backlog-link.json relation when present; otherwise use the source's project folder, then the dated fallback folder.
- Make
prepare-run read the Playwright CLI reference before it uses a browser snapshot or emits runner commands.
- Reconcile scope visibility with the no-narration rule: return scope details only when a user decision is required; otherwise continue silently and report the final result and paths.
- When running at the user-facing root, offer to open
happy-path.webm. When delegated, return the review offer and evidence path to the parent so the parent can ask the user.
- Document that
page is the only guaranteed run-code binding. Browser globals such as URL must be used inside the page context, for example through page.evaluate, rather than assumed in the runner sandbox.
- Apply the no-pipes rule to the runner command itself. A host-native wrapper is allowed only when it preserves the exact invocation, stdout, stderr, and exit status and does not hide or coerce failures.
Report semantics
- Add a
Limitations section to qa.md.
- Keep the existing run statuses. A run may remain
pass with a recorded limitation only when every scoped criterion has sufficient evidence and the limitation does not weaken the required proof. A limitation that prevents sufficient proof produces blocked.
- Keep
Findings for fail or blocked; do not turn a limitation into a product finding unless it affects a criterion.
Acceptance criteria
Out of scope
- API or CLI interfaces.
- Patching the application under test.
- Adding QA dependencies to the application manifest.
- Adding a new run status.
- Revisiting the completed release-pin cleanup or the existing verified tail-cut mechanism.
References
Problem
The first real runs of
aidd-qa:01-acceptance-qaexposed gaps that still force the operator to improvise. The current contract does not fully define how to select a browser entry, prove test-only storage before starting it, configure a mobile or permission-sensitive session, assess qualitative or short-lived outcomes, or preserve the user-facing flow when QA is delegated.Some follow-ups are already complete and are not part of this issue: the runner supports a visually verified tail cut for command latency, and #925 removed the one-time
release-as: 1.0.0pin.Outcome
As a reviewer, I want Acceptance QA to prepare, execute, and report real browser journeys from explicit evidence, so desktop and mobile criteria produce safe, reproducible results without host-specific improvisation or overstated proof.
Scope
Preparation and safety
Browser QAproject memory first, then a related browser test or fixture. Separate resolution from starting the application.testorstagingare not proof by themselves.blockedwith the shortest decisive error. Never invent a start command.Browser QAmemory, then the existing1280×720fallback. When mobile behavior is required but no dimensions are declared, ask once instead of guessing.Observable and timed criteria
run-codeexecution with a bounded browser observation and include the measured result inActual. Frame sampling validates the video artifact but is not sufficient behavioral proof for a transition shorter than the sampling interval.Flow, discovery, and host portability
backlog-link.jsonrelation when present; otherwise use the source's project folder, then the dated fallback folder.prepare-runread the Playwright CLI reference before it uses a browser snapshot or emits runner commands.happy-path.webm. When delegated, return the review offer and evidence path to the parent so the parent can ask the user.pageis the only guaranteedrun-codebinding. Browser globals such asURLmust be used inside the page context, for example throughpage.evaluate, rather than assumed in the runner sandbox.Report semantics
Limitationssection toqa.md.passwith a recorded limitation only when every scoped criterion has sufficient evidence and the limitation does not weaken the required proof. A limitation that prevents sufficient proof producesblocked.Findingsforfailorblocked; do not turn a limitation into a product finding unless it affects a criterion.Acceptance criteria
blockedresult.run-code; the report does not cite 4 fps frame inspection as sufficient behavioral proof.backlog-link.jsonwhen that relation exists.prepare-runloads the runner reference before any snapshot or runner invocation.qa.mdrecords limitations without adding a new status or allowing insufficient evidence to pass.Out of scope
References
plugins/aidd-qa/skills/01-acceptance-qa/