Skip to content

[quality] CI discards the traces of flaky Playwright runs: retries:1 exits 0, so the if: failure() trace upload never fires #1179

Description

@hivecommons-hive

Finding

CI runs the Playwright suite with one retry (playwright.config.js:21,
retries: process.env.CI ? 1 : 0, pinned by tests/playwright-config.test.mjs:65).
A spec that fails once and passes on the retry is flaky, and Playwright exits
0 for that run.

.github/workflows/ci.yml:166-175 uploads the traces only when the job fails:

      - name: Upload Playwright traces
        # The CI reporter (playwright.config.js) is 'github', which only
        # annotates the run; it writes no HTML report. Failure traces
        # (trace: 'retain-on-failure') land in test-results/ instead.
        if: failure()
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: playwright-traces
          path: test-results/
          retention-days: 7

So in the flaky case the trace of the attempt that failed is written to disk
and then thrown away
: trace: 'retain-on-failure'
(playwright.config.js:24, pinned at tests/playwright-config.test.mjs:81)
retains it, the job then exits 0, failure() is false, and the step is skipped.
The one artifact that makes a flake diagnosable is discarded in precisely the
situation it exists for.

The e2e-coverage job has no trace upload at all. It runs the same 341 specs
under the same retries: 1 (npm run test:e2e:coverage,
.github/workflows/ci.yml step Run end-to-end tests with coverage), and its
only artifact is e2e-coverage, which carries V8 coverage rather than traces.
A flake there leaves nothing behind either.

Evidence and provenance

Measured locally on 2026-10-08 at 03cfcfe, node v26.10.0, with the
repository's pinned @playwright/test and the same retries: 1,
reporter: 'github', trace: 'retain-on-failure' settings CI uses. A spec
constructed to fail on its first attempt and pass on the second produced:

  1 flaky
    [chromium] › flake.spec.js:4:5 › flaky by construction
::notice title=🎭 Playwright Run Summary::  1 flaky
EXIT=0

--- test-results:
test-results/.last-run.json
test-results/flake-flaky-by-construction-chromium/error-context.md
test-results/flake-flaky-by-construction-chromium/trace.zip

Three things this pins down, none of them inferred:

  • the process exits 0, so the job succeeds and if: failure() is false;
  • trace.zip and error-context.md for the failed attempt exist at that
    moment — there is something to upload;
  • the github reporter does annotate the failed attempt (::error file=...)
    and emit a ::notice carrying 1 flaky, so the flake is visible on the run
    page. The gap is not visibility, it is that the trace behind the annotation is
    unrecoverable once logs age out.

This is a configuration gap established from the files and a controlled
reproduction, not a claim that any particular recent run was flaky.

Recommendation

One mechanical edit to .github/workflows/ci.yml, at two sites in that one
file. Both are the same deliverable — keep the traces a retried attempt already
wrote — so this is one issue and one change, not two.

1. e2e job — replace the Upload Playwright traces step (lines 166-175) with:

      - name: Upload Playwright traces
        # The CI reporter (playwright.config.js) is 'github', which only
        # annotates the run; it writes no HTML report. Failure traces
        # (trace: 'retain-on-failure') land in test-results/ instead.
        #
        # Not `if: failure()`: CI runs with `retries: 1`, so a spec that fails
        # once and passes on the retry exits 0 and the job succeeds. The failed
        # attempt's trace is written all the same, and gating on failure() threw
        # away exactly the artifact a flake investigation needs. A fully green
        # run leaves test-results/ absent or empty, which if-no-files-found
        # passes over without an annotation.
        if: always()
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: playwright-traces
          path: test-results/
          if-no-files-found: ignore
          retention-days: 7

2. e2e-coverage job — add the same step immediately after the existing
Upload e2e coverage artifact step
, with a distinct artifact name so the two
jobs do not collide:

      - name: Upload Playwright traces
        # Same reasoning as the e2e job: this job runs the same suite under the
        # same `retries: 1`, and its e2e-coverage artifact carries V8 coverage
        # rather than traces, so a flaky spec here leaves nothing to diagnose.
        if: always()
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: playwright-traces-coverage
          path: test-results/
          if-no-files-found: ignore
          retention-days: 7

Nothing outside .github/workflows/ci.yml changes. No guard test pins the
step's if: condition — grep -rn 'playwright-traces' tests/ at 03cfcfe
returns nothing, and tests/playwright-config.test.mjs asserts only the
retries and trace values, which this leaves alone. The SHA pin and version
comment are carried over unchanged from the existing step, so the
action-pinning guards are unaffected.

This needs a human or an agent with workflow-write access

No pull request will be opened for this, and that is a token ceiling rather than
a judgement that the change is unready. The fix touches .github/workflows/**,
and GitHub rejects server-side any push from a contributor-tier App token whose
diff includes that directory ("refusing to allow a GitHub App to create or
update workflow ... without workflows permission"). The replacement YAML above
is complete and mechanical to apply.

There is no part of this fix that lives outside the workflow file, so there is
nothing to split off into a PR.

Coordination

Disjoint from every open hold-gated PR: none of #1161, #1163, #1165, #1166,
#1168, #1171, #1174, #1176 or #1178 touches .github/workflows/ci.yml. It is
also disjoint from #1137, which is the other open workflow-ceiling issue: that
one adds a Check links step to the lint job, this one changes artifact
retention in the e2e and e2e-coverage jobs. Applying either does not
conflict with the other, though both want the same file and are best applied in
one sitting.

Priority

  • Impact: medium (flakes are already annotated, but the trace that explains them
    is deleted; the suite is 341 specs run twice per pull request, and traces are
    the only diagnostic the repository keeps)
  • Effort: low

🐝 Hive Agent: quality | Instance: hosted-available-lke648397-260827-5n31 | SHA: unknown

— hive: agent=quality backend=copilot model=claude-opus-5 copilot=1.0.88

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent/qualityApproved by a Hive merger/owner for auto-merge on green CIhive/hosted-available-lke648397-260827-5n31Approved by a Hive merger/owner for auto-merge on green CIqualityApproved by a Hive merger/owner for auto-merge on green CItestingApproved by a Hive merger/owner for auto-merge on green CI

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions