Skip to content

feat(aidd-qa): make acceptance evidence reliable across real browser journeys #918

Description

@blafourcade

Problem

The first real runs of aidd-qa:01-acceptance-qa exposed gaps that still force the operator to improvise. The current contract does not fully define how to select a browser entry, prove test-only storage before starting it, configure a mobile or permission-sensitive session, assess qualitative or short-lived outcomes, or preserve the user-facing flow when QA is delegated.

Some follow-ups are already complete and are not part of this issue: the runner supports a visually verified tail cut for command latency, and #925 removed the one-time release-as: 1.0.0 pin.

Outcome

As a reviewer, I want Acceptance QA to prepare, execute, and report real browser journeys from explicit evidence, so desktop and mobile criteria produce safe, reproducible results without host-specific improvisation or overstated proof.

Scope

Preparation and safety

  • Resolve entry, authentication, fixtures, reset, viewport, and browser permissions from the Browser QA project memory first, then a related browser test or fixture. Separate resolution from starting the application.
  • Treat storage as proven test-only only when project memory explicitly says so, an isolated test fixture/configuration establishes it, or the user confirms it. Environment names such as test or staging are not proof by themselves.
  • When reset can delete data, prove test-only storage before starting the entry or taking a browser snapshot. If proof is absent, ask once and do not start.
  • Validate that the selected entry serves the expected browser UI. If it fails or starts only a backend, try at most one browser entry explicitly documented by the same evidence chain; otherwise report blocked with the shortest decisive error. Never invent a start command.
  • Define a destructive action as one that can alter data outside the isolated scenario fixture or cause an irreversible external side effect. A verified teardown limited to test-only data created by the scenario is not destructive.
  • Resolve viewport in this order: acceptance criteria, Browser QA memory, then the existing 1280×720 fallback. When mobile behavior is required but no dimensions are declared, ask once instead of guessing.
  • Resolve required permission and consent state before recording. Set it explicitly when criteria or project memory declare it; otherwise retain the browser default and record that fact in the prepared run.

Observable and timed criteria

  • Turn qualitative wording such as “jumps” or “flashes” into an explicit browser-observable condition before recording. Ask once when more than one materially different interpretation exists; reject the scenario when no observable condition can be established.
  • When a criterion names a timing boundary, define probes on both sides, record their chosen margins, and report measured values. Ask once when the margin could change the product meaning; never invent a product tolerance silently.
  • Assess short-lived transitions inside the single run-code execution with a bounded browser observation and include the measured result in Actual. Frame sampling validates the video artifact but is not sufficient behavioral proof for a transition shorter than the sampling interval.

Flow, discovery, and host portability

  • Resolve an existing feature evidence folder through its backlog-link.json relation when present; otherwise use the source's project folder, then the dated fallback folder.
  • Make prepare-run read the Playwright CLI reference before it uses a browser snapshot or emits runner commands.
  • Reconcile scope visibility with the no-narration rule: return scope details only when a user decision is required; otherwise continue silently and report the final result and paths.
  • When running at the user-facing root, offer to open happy-path.webm. When delegated, return the review offer and evidence path to the parent so the parent can ask the user.
  • Document that page is the only guaranteed run-code binding. Browser globals such as URL must be used inside the page context, for example through page.evaluate, rather than assumed in the runner sandbox.
  • Apply the no-pipes rule to the runner command itself. A host-native wrapper is allowed only when it preserves the exact invocation, stdout, stderr, and exit status and does not hide or coerce failures.

Report semantics

  • Add a Limitations section to qa.md.
  • Keep the existing run statuses. A run may remain pass with a recorded limitation only when every scoped criterion has sufficient evidence and the limitation does not weaken the required proof. A limitation that prevents sufficient proof produces blocked.
  • Keep Findings for fail or blocked; do not turn a limitation into a product finding unless it affects a criterion.

Acceptance criteria

  • Test-only proof and application start are ordered so no destructive reset target is mounted or inspected through a started app before proof.
  • A documented entry that crashes or serves no browser UI yields one documented fallback attempt or a precise blocked result.
  • Mobile viewport and browser permission state come from explicit evidence and appear in the prepared run; ambiguous mobile dimensions produce one question.
  • Qualitative criteria have one explicit observable condition before recording, or are rejected with a reason.
  • A timed boundary records probes on both sides and the measured values used for the result.
  • A sub-sampling transition is assessed in run-code; the report does not cite 4 fps frame inspection as sufficient behavioral proof.
  • Scenario-owned teardown is distinguished from destructive external mutation and remains subject to test-only storage proof.
  • Existing feature evidence is found through backlog-link.json when that relation exists.
  • prepare-run loads the runner reference before any snapshot or runner invocation.
  • Scope reporting obeys the no-narration contract unless a decision is required.
  • A delegated run returns the video-review offer to its parent instead of attempting to question the user directly.
  • Runner examples use only documented bindings and keep browser globals inside page context.
  • Host wrappers are accepted only when command output and exit semantics remain intact.
  • qa.md records limitations without adding a new status or allowing insufficient evidence to pass.
  • A rerun on a mobile-first application completes without inventing entry, viewport, permission, timing, or evidence semantics.

Out of scope

  • API or CLI interfaces.
  • Patching the application under test.
  • Adding QA dependencies to the application manifest.
  • Adding a new run status.
  • Revisiting the completed release-pin cleanup or the existing verified tail-cut mechanism.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Fields

    Priority

    Medium

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions