Skip to content

CodeQL dispatch reports a clean scan as state=failure when only its wake step failed — the verdict reads the job conclusion, not the gate #2141

Description

@seonghobae

Summary

A CodeQL dispatch scan that passes its Medium+ SARIF gate is reported to the required check as CodeQL dispatch scan for <lang> did not pass (state=failure), telling the reader to go look at SARIF evidence that is in fact clean. The cause is that the required check derives its verdict from the dispatch job's conclusion, and that conclusion is failure whenever the job's last step — waking the required job — fails, regardless of what the scan found.

This is distinct from the wake race in #2051 / #2056, and worse in kind. The race delays a check. This misreports it: a green scan becomes a red security verdict.

Evidence

.github#2137 head 3d39fe7c. Dispatch run 34734286494, job CodeQL dispatch scan (python) (103663188483). Every step of the actual analysis passed:

  1 success  Set up job
  …
  8 success  Initialize CodeQL
  9 success  Perform CodeQL Analysis
 10 success  Enforce CodeQL Medium+ SARIF gate      <-- the security gate PASSED
 11 success  Preserve CodeQL SARIF evidence
 12 success  Publish CodeQL dispatch status
 13 failure  Wake exact CodeQL required job         <-- only this
 …
 27 success  Complete job

Job conclusion: failure, from step 13 alone.

The required check then reads that conclusion directly as the verdict (codeql-scan-dispatch.yml, the fallback branch after the codeql-dispatch/<lang> status lookup returns nothing):

job_conclusion="$(printf '%s' "$jobs_json" | jq -r --arg name "$expected_job" '
  … | if length == 1 then .[0].conclusion else empty end
')"
case "$job_conclusion" in
  …)  echo "verdict=${job_conclusion}" >>"$GITHUB_OUTPUT"
      echo "Found completed CodeQL dispatch scan job for ${LANGUAGE}: ${job_conclusion}."

And the required job emits:

DISPATCH_OUTCOME: success
VERDICT_STATE:    failure
##[error] CodeQL dispatch scan for python did not pass (state=failure).
          See the linked dispatch run for SARIF evidence.

There is no Medium+ SARIF evidence to see. Step 10 passed.

Why it matters

  • It inverts a security signal. "This PR's Python CodeQL scan did not pass" is the strongest thing this pipeline says about a change, and it is being said about scans that passed. A reviewer acting on it looks for a vulnerability that does not exist; a reviewer who learns the signal is unreliable stops reading it, which is worse.
  • It is reachable by every PR that loses the shard race. POST /actions/jobs/{id}/rerun re-runs the containing run, so when two language shards target jobs in the same run the second is refused The workflow run containing this job is already running (HTTP 403) — its step 13 fails, its clean scan is recorded as a failed scan.
  • It survives the re-run. On #2137 the first attempt read VERDICT_STATE: pending ("scan dispatched, will rerun"); the re-run read VERDICT_STATE: failure. The second reading is worse than the first while nothing about the code changed.

Fix direction

The verdict should come from what the scan decided, not from whether the notification afterwards succeeded. Concretely, one of:

  1. Derive the verdict from the gate step, not the job. The dispatch job already has the authoritative answer at step 10; publish that as the codeql-dispatch/<lang> status and let the fallback read step-level outcome rather than .conclusion.
  2. Move the wake out of the scanning job, so a notification failure cannot change a scan job's conclusion at all. This also dissolves the race, since one coordinator can wake all shards once — which is what fix(codeql): coordinate failed-job wake once #2051 and fix(codeql): serialize exact dispatch wakeups #2056 are already reaching for.
  3. At minimum, make the fallback distinguish them: treat a job whose only failed step is the wake as pending (wake later) rather than failure (scan verdict). This is the smallest change and it removes the false red, though it leaves the race.

(2) subsumes (3) and is the same work #2051/#2056 need to do anyway, so it is probably the right one rather than three separate patches.

Scope note

codeql-scan-dispatch.yml is the surface #2040 is the canonical successor for, and #2051/#2056 are active on the adjacent race. I am not opening a competing PR on that file — filing this so the verdict-derivation half is tracked on its own, since it is a correctness defect in what the pipeline reports rather than a scheduling defect in when it reports it, and it could be fixed even if the race work stalls.

Found while checking whether #2137's red python shard was a finding in its diff. It was not: that scan is clean.

Refs #2040, #2051, #2056, #2137, #1929.

Activity

  1. seonghobae commented on Sep 13, 2026

    @seonghobae
    ContributorAuthor

    Fresh independent reproduction from .github#2117@2b15b43a6a679415d6211892fc0bdc0a0e5c853b confirms this is not specific to #2137 and not a SARIF/source finding.

    Required CodeQL run 34742456179 initially failed closed while terminal dispatch evidence was unavailable. After authenticated exact-head statuses existed, I reran only the failed Python compatibility job. The latest exact dispatch run 34744312781 then completed validation and both actual CodeQL analyses/SARIF gates successfully; for Python the gate reported files=1 results=1 medium_plus=0, and codeql-dispatch/python=success was successfully published after App-token publication fell back to the authorized job token.

    The dispatch job nevertheless concluded FAILURE solely at Wake exact CodeQL required job: it attempted to rerun required job 103688653949 in run 34742456179 and GitHub returned The workflow run containing this job is already running (HTTP 403). The required compatibility job then observed that dispatch job conclusion as failure and failed, despite the authoritative scan/gate/status being green.

    This is the same verdict/wake coupling described here and the matrix wake race already owned by #2106/#2051/#2056. I am not patching the DiskSage leaf or retrying unchanged again. Please treat #2117 as an additional exact canary for the canonical handler/settlement repair.

  2. seonghobae commented on Sep 13, 2026

    @seonghobae
    ContributorAuthor

    Fresh independent canary from .github#2079@9dccfaa0776950498e557390a2fa8d6c34e0baf4 reproduces this exact defect without a source/SARIF finding. CodeQL PR run 34755517763 dispatched authenticated exact-head scans; handler run 34755677020 completed CodeQL analysis, Medium+ gate, SARIF preservation, and codeql-dispatch/actions / codeql-dispatch/python status publication successfully. In Actions scan job 103719565759, only Wake exact CodeQL required job failed: GitHub returned HTTP 403 The workflow run containing this job is already running because the sibling language wake had already started the shared required run. The required-run compatibility shard therefore remained red/pending even though the scan verdict/status was green. No manual rerun or leaf workaround was used. This is additional executable evidence for the #2106/#2040 run-wide settlement owner path and the false-red boundary tracked here; .github#2079 should not patch CodeQL locally.

  3. seonghobae commented on Sep 14, 2026

    @seonghobae
    ContributorAuthor

    Additional canary, and a variant worth separating from the wake-step failure this issue records so far: the dispatch job concludes success and the failed required shard is still never re-run.

    In the two comments above the dispatch job ended failure at Wake exact CodeQL required job (HTTP 403 "already running"). In the three cases below the wake step did not fail — the whole dispatch job is green — and the shard it promised to rerun keeps the timestamp of its original failure, so the PR stays red indefinitely with no further event to settle it.

    Exact-head evidence (all read from the PR's own check runs, completedAt verbatim):

    PR / head failed shard (check run) shard completed Dispatch current-head CodeQL scan dispatch completed shard re-run?
    #2154@c393e3af1 CodeQL compatibility analysis (python) 103870251843 2026-09-13T22:04:00Z 103949951656 SUCCESS 2026-09-14T15:13:33Z no (17h)
    #2156@18bfe42ec python 103850982509, actions 103850982562 05:53:46Z / 05:59:31Z 103873946268 SUCCESS 2026-09-14T11:22:45Z no
    #2186@02077429d actions 103846126655, python 103846126679 05:33:08Z / 05:34:22Z 103869609441 SUCCESS 2026-09-14T10:44:38Z no

    Two details that look diagnostic:

    • The #2154 dispatch ran for 8 seconds (startedAt 15:13:25Z → completedAt 15:13:33Z). That is far too short to have performed an analysis, so it took a short-circuit path — and still reported success without waking the red shard.
    • On that same PR the actions shard was re-run (10:57:28Z → 11:01:07Z) and went green, while the python shard next to it was left at its 22:04Z failure. So the settlement is per-shard and partial, not simply absent.

    The message the red shard leaves behind is the promise that is not kept:

    DISPATCH_OUTCOME: success
    VERDICT_STATE: pending
    ##[error]CodeQL scan dispatched. The dispatch workflow will rerun this exact failed CodeQL job after publishing its terminal verdict.
    

    Reader impact is the mirror image of the false-red in the issue body. There the verdict is wrong; here there is no terminal verdict at all, and the PR is indistinguishable from one with a genuine unaddressed CodeQL failure. Both land on the same required context, so a reviewer cannot tell the three states apart without opening the dispatch run.

    I am not patching CodeQL locally on any of these three branches. Recording them as exact canaries for the canonical settlement repair (#2106 / #2040), since a fix that only makes the wake step succeed would leave all three of these still red.

  4. seonghobae commented on Sep 15, 2026

    @seonghobae
    ContributorAuthor

    Direct confirmation of the issue body's mechanism, with both variants visible in one dispatch run, on .github#2186@02077429df1af1b5f4c96c640c921e646f0ead13.

    Dispatch run 34834691830, step lists read verbatim:

    CodeQL dispatch scan (actions) (job 104029622727, conclusion failure):

      9 Perform CodeQL Analysis            -> success
     10 Enforce CodeQL Medium+ SARIF gate  -> success     <-- the security gate PASSED
     11 Preserve CodeQL SARIF evidence     -> success
     12 Publish CodeQL dispatch status     -> success
     13 Wake exact CodeQL required job     -> failure     <-- only this
     27 Complete job                       -> success
    

    CodeQL dispatch scan (python) (job 104029622673, conclusion failure):

     10 Enforce CodeQL Medium+ SARIF gate  -> failure     <-- a real Medium+ result (#2208)
     13 Wake exact CodeQL required job     -> success
    

    The two shards fail for opposite reasons and are indistinguishable at the required check. The required CodeQL compatibility analysis (actions) shard reports:

    DISPATCH_OUTCOME: success
    VERDICT_STATE: failure
    ##[error]CodeQL dispatch scan for actions did not pass (state=failure). See the linked dispatch run for SARIF evidence.
    

    There is no Medium+ SARIF evidence for actions. Step 10 passed. The rerun job log shows the verdict came from the job-conclusion fallback, not a published status:

    Found completed CodeQL dispatch scan job for python: failure.
    

    So on one head, one shard's red is a genuine security finding and the other's is manufactured by step 13, and the message a reviewer reads is identical in both cases. That is the inversion this issue describes, and it now coexists with a true positive on the same run — which makes the false red actively more expensive, because the PR does have a real CodeQL failure to act on and the second red is noise pointing at evidence that does not exist.

    Two further notes for the settlement owner path (#2106 / #2040):

    • The python shard here shows wake succeeding, so wake failure is not a constant; it is racy per shard, consistent with the sibling-wake 403 in the comments above.
    • Separately from the verdict inversion, the rerun does not converge: after the required python job re-ran at 02:01:30Z–02:04:46Z (job 104165880022) and read state=failure, a fresh Dispatch current-head CodeQL scan was queued again immediately. With the codeql: repaired SARIF gate now blocks every PR on a false-positive URL-sanitization finding in a contract test #2208 finding still on main, that is an unbounded dispatch→scan→fail→dispatch cycle per open PR, not a settling one.

    No local CodeQL patch on this branch.

  5. seonghobae commented on Sep 18, 2026

    @seonghobae
    ContributorAuthor

    Closing — fixed on main via #2234

    Verified on protected main:

    1. .github/workflows/codeql-pr.yml fallback verdict now reads the dispatch job step named Enforce CodeQL Medium+ SARIF gate before treating whole-job conclusion as the security verdict ((.steps[]? | select(.name == "Enforce CodeQL Medium+ SARIF gate") | .conclusion)). Gate success → verdict=success even when a later wake step fails.
    2. Contract coverage landed in tests/test_codeql_pr_workflow_contract.py with fix(codeql): stop SARIF false positive and wake-only verdict inversion #2234 (7d496d626).
    3. This is distinct from remaining wake-race / credential work (fix(codeql): coordinate failed-job wake once #2051/fix(codeql): serialize exact dispatch wakeups #2056/fix(codeql): wake required jobs with the exchanged target app token #2040): those affect whether the required job is notified, not whether a clean gate is mislabeled failure.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions