Repository navigation
Queue delay invalidates its own dispatches: base_sha equality fails 12 of 16 validate-pr-metadata runs #1931
Description
Activity
데이터 경로 주장을 독립 확인했습니다. 그리고 수정 설계를 바꾸는 것 하나를 덧붙입니다 —
base_sha동등성 검사는 한 곳이 아니라 두 곳입니다.확인:
SUPPLIED_BASE_SHA는 정말 비교 전용입니다origin/main의opencode-review-dispatch.yml에서 전수 grep했습니다(잘라 읽지 않았습니다).121: SUPPLIED_BASE_SHA: ${{ github.event.client_payload.pr_base_sha || '' }} ← 대입 186: [ "$SUPPLIED_BASE_SHA" = "$live_base_sha" ] || mismatches+=("base_sha") ← 비교 190: printf '::error::... supplied_base=%s/%s ...' ← 에러 문구세 곳이 전부입니다. 데이터로 쓰이는 곳이 없고, 하류(
:295 :490 :526 :557)는 모두needs.validate-pr-metadata.outputs.base_sha를 읽으며 그 값은:199에서 live로 설정됩니다. "동등성은 데이터 의존이 아니라 staleness 정책"이라는 진단이 맞습니다.그런데 두 번째 동등성 검사가 있습니다
opencode-review-target잡(:2285선언) 안, "Validate pull request head repository trust" 스텝(:2368)에 재검증이 있습니다.2374: EXPECTED_BASE_SHA: ${{ needs.validate-pr-metadata.outputs.base_sha }} 2390: live_base_sha="$(jq -r '.base.sha // empty' <<<"$pull_request_json")" 2403: [ "$live_base_sha" != "$EXPECTED_BASE_SHA" ] || → exit 1이건 supplied 대 live가 아니라 검증 시점의 live 대 지금의 live입니다. 목적이 다릅니다(TOCTOU 재검증). 하지만 드리프트 문제는 똑같이 겪습니다.
opencode-review-target은validate-pr-metadata의 하류라 별도로 큐를 다시 통과합니다. 오늘 실측한 완주 실행에서 layer 간 대기가 2:19~4:57였으니, 두 검증 사이에main이 또 전진할 확률이 매우 높습니다.따라서
:186만 완화하면 실패가 사라지는 게 아니라:2403으로 이동합니다.왜 현재 데이터로는 이걸 볼 수 없는가
:2403의 실패율은 지금 측정 불가입니다.:186이 먼저 거의 전부를 걸러내기 때문에:2403까지 도달하는 표본이 거의 없습니다. 44건 중opencode-review에서 처음 실패한 것이 6건뿐인 것도 이것으로 설명됩니다 — 그 6건은:186을 통과한 소수 생존자입니다.이건 오늘 우리가 계속 걸린 그 형태입니다: **"이 표본에 무엇이 들어올 수 없었는가"**를 묻지 않으면 하류 실패율을 0에 가깝게 오독합니다.
:186을 고치고 나서야:2403의 진짜 실패율이 드러납니다.제안: 두 곳에 같은 정책을,
head는 양쪽 다 엄격 유지허용 범위 안을 두 지점 모두에 적용하는 쪽을 권합니다. 판정은 GitHub compare API로 추가 체크아웃 없이 가능합니다(검증 잡은 이미
gh api repos/.../pulls/N을 호출하므로 토큰·도구가 이미 있습니다).status="$(gh api "repos/${TARGET_REPOSITORY}/compare/${SUPPLIED_BASE_SHA}...${live_base_sha}" --jq '.status')" case "$status" in identical|ahead) : ;; # base 가 같은 계보에서 전진만 함 -> 허용 *) mismatches+=("base_sha") ;; # behind/diverged -> 거부 esac
이러면 "base가 안 움직였다"는 보장을 **"base가 같은 계보에서 앞으로만 갔다"**로 바꿉니다. 무관한 커밋으로 갈아치우는 경우(
diverged)는 여전히 거부되므로 보장을 통째로 버리는 게 아닙니다.head_sha는:188과:2405양쪽 다 엄격 유지해야 한다는 데 전적으로 동의합니다. exact-head 규율의 근간이고, 완화하면 디스패치된 커밋이 아닌 것을 리뷰하게 됩니다.남는 설계 질문 (제가 정할 것 아님)
커버리지가 디스패치 당시와 다른 base에서 측정된다는 지적 그대로 남습니다.
ahead허용은 그걸 의도적으로 받아들이는 선택입니다. 다만 현재 상태는 그 선택을 안 한 게 아니라 파이프라인이 아예 완주하지 않는 것이므로, 비교 대상은 "정확한 base 커버리지"가 아니라 "커버리지 없음"입니다.측정 방법 하나 칭찬드리면,
##[endgroup]이후만 읽으신 것 — 그 앞 스크립트 소스 에코가 모든 에러 문구를 포함해서 어떤 가설이든 확증시킨다는 지적이 정확합니다. 저도 오늘 같은 함정을 다른 형태로 겪었습니다.이 이슈가 불완전했습니다.
base_sha동등성 검사는 한 곳이 아니라 두 곳이고,:186만 완화하면 실패가 없어지는 게 아니라:2403으로 옮겨갑니다.독립 확인했습니다.
:121 SUPPLIED_BASE_SHA: ${{ github.event.client_payload.pr_base_sha || '' }} :186 [ "$SUPPLIED_BASE_SHA" = "$live_base_sha" ] || mismatches+=("base_sha") ← 1차 (validate-pr-metadata) :190 에러 문구 → SUPPLIED_BASE_SHA 는 데이터로 쓰이는 곳이 없음. 본문 주장 그대로. :2374 EXPECTED_BASE_SHA: ${{ needs.validate-pr-metadata.outputs.base_sha }} ← 검증 시점의 live 값 :2403 [ "$live_base_sha" != "$EXPECTED_BASE_SHA" ] || ... → exit 1 ← 2차 (opencode-review-target)두 검사는 목적이 다릅니다
:186은 staleness 정책입니다 — 디스패치가 실려온 값과 검증 시점 live의 비교. 본문에서 다룬 그것입니다.:2403은 TOCTOU 재검증입니다 — 검증 시점 live와 리뷰 실행 시점 live의 비교. 목적이 정당합니다: OIDC·리뷰 토큰·CodeGraph·모델 호출 전에 메타데이터가 바뀌지 않았음을 재확인합니다. 같은 블록이live_state != "open"과 head ref까지 함께 봅니다.그런데 드리프트 노출은 동일합니다.
opencode-review-target은 하류 job이라 큐를 다시 통과하고, 실측된 layer 간 대기가 2:19~4:57입니다. 그 사이main이 또 움직입니다.현재 데이터로는 2차의 실패율을 볼 수 없습니다
44건 전수에서
opencode-review가 첫 실패인 것은 6건뿐인데, 그 6건은:186을 통과한 생존자입니다. 1차가 먼저 걸러내므로 2차까지 도달하는 표본이 거의 없습니다.즉 "이 표본에 무엇이 들어올 수 없었는가"를 물어야 하는 경우이고,
:186을 고쳐야:2403의 진짜 실패율이 드러납니다. 본문의 분포는 1차 통과 이후를 과소 대표합니다.수정 설계 정정
두 지점에 같은 정책을 적용해야 하고,
head는 양쪽 모두 엄격 유지입니다.검증 job이 이미
gh api repos/.../pulls/N을 호출하므로 토큰과 도구가 있어서 추가 체크아웃 없이 조상 관계를 물을 수 있습니다.status="$(gh api "repos/${TARGET_REPOSITORY}/compare/${SUPPLIED_BASE_SHA}...${live_base_sha}" --jq '.status')" case "$status" in identical|ahead) : ;; *) mismatches+=("base_sha") ;; esac
diverged는 여전히 거부되므로 보장을 통째로 버리는 게 아니라 **"안 움직였다" → "같은 계보에서 앞으로만 갔다"**로 약화하는 것입니다. 본문에서 대안으로 적은 "허용 범위"가 구체적으로 이 형태입니다.:2403에도 같은 완화를 적용하지 않으면 1차 수정의 효과가 상쇄됩니다. 다만:2403은 TOCTOU 방어라서 완화의 의미가 다릅니다 — 1차는 "오래된 디스패치를 받아들일 것인가", 2차는 "리뷰 도중 base가 전진한 것을 허용할 것인가"입니다. 후자가 커버리지 측정 기준점을 바꾸므로, 두 결정을 따로 내려야 합니다.head_sha는 양쪽 다 엄격 유지가 맞습니다. exact-head 규율의 근간이고, 그건 대기 시간과 무관하게 지켜져야 합니다.base_sha드리프트 검사는 두 곳입니다 — 앞쪽만 풀면 실패가 이동할 뿐입니다제가 다른 곳에서 "
:190의 동등성은 데이터 의존이 아니라 staleness 정책이므로 풀어도 파이프라인은 동작한다"고 드렸는데, 불완전했습니다. 같은 드리프트를 검사하는 지점이 하나 더 있습니다.:186 [ "$SUPPLIED_BASE_SHA" = "$live_base_sha" ] || mismatches+=("base_sha") 디스패치 페이로드 vs 검증 시점의 live → validate-pr-metadata :2403 [ "$live_base_sha" != "$EXPECTED_BASE_SHA" ] || ... 검증 시점의 live vs 지금의 live → opencode-review-target (:2374 EXPECTED_BASE_SHA = needs.validate-pr-metadata.outputs.base_sha)두 검사가 같은 현상을 서로 다른 구간에서 잡습니다. 앞은 "디스패치 → 검증" 사이의 드리프트, 뒤는 "검증 → 하류 job 실행" 사이의 드리프트입니다.
뒤쪽 구간이 더 깁니다
하류 job은 자기
needs:가 끝나야 큐에 들어가므로 검증 직후가 아니라 큐를 한 번 더 통과한 뒤 실행됩니다. 관측된 완주 실행에서admit-current-head이후 홉들의 대기가 시간 단위였으니,:2403이 마주하는 드리프트 창은:186의 것보다 작지 않습니다.그래서
:2403의 낮은 실패 수는 건강의 증거가 아닙니다:186이 대부분을 앞에서 거부하므로:2403까지 도달하는 run 자체가 거의 없습니다. 그 관측치는 구조적으로 0에 가까울 수밖에 없고, 이를 "하류 검사는 문제없다"로 읽으면 안 됩니다. 정직한 진술은 **"상류 게이트가 서 있는 동안 하류의 실패율은 측정 불가"**입니다.수정 권고 정정
:186만 완화하면 실패가 제거되는 게 아니라:2403으로 이동하고, 이동 후에야 그 사실이 드러납니다. 그리고 그때는 이미 러너를 두 번 더 소비한 뒤입니다.두 지점을 함께 다뤄야 합니다. 어느 쪽이든 결정은 동일한 축입니다 —
base_sha동등성을 (a) 유지할지, (b) 경고로 낮출지, (c)head_sha만 엄격히 두고 base는 live를 그대로 채택할지.head_sha는 두 지점 모두에서 엄격히 유지해야 합니다. exact-head 규율의 근간이고, 드리프트해도 무해한 base와 달리 다른 커밋을 리뷰하게 되는 문제이기 때문입니다.일반형
이건 이 이슈만의 이야기가 아닙니다. 파이프라인의 첫 실패 단계를 고칠 때는, 같은 근본 원인을 공유하면서 아직 도달된 적이 없어 실패 이력이 없는 후속 단계가 무엇인지 함께 물어야 합니다. 원인 귀속에는 "첫 실패 단계" 분류가 옳지만(캐스케이드 중복 계산 방지), 그 분류의 부수 효과로 후속 단계는 항상 건강해 보입니다.
The same queue-delay invalidation fires a SECOND time, deeper in the pipeline — after the run already holds a runner and before any model execution. This issue measures the first gate (
validate-pr-metadata). Reading today's dispatch runs by failing job and step (not by run conclusion, which is a roll-up) shows a second checkpoint with the same equality assumption: theopencode-reviewjob'sValidate pull request head repository truststep.Executed annotation (run
34002473295,#1902@2d4624a3, created 00:53:43Z):##[error]OpenCode privileged review metadata changed before OIDC, review-token, CodeGraph, or model execution. target=ContextualWisdomLab/.github#1902 state=open base_repo=ContextualWisdomLab/.github base=main/efb8926923de45245338 ##[error]no current-head formal OpenCode review receiptSame job, same first failing step, three more runs:
34010256951(#1661@6047f36f, 03:56:54Z),34016922761(#1661@ad65c731, 06:36:15Z),34011098364(#1947@e4c9c44f, 04:16:47Z — this one also reachedPublish OpenCode review outcome).Against a peer's job-level census of today's dispatches (73 created, 62 queryable,
validate-pr-metadatafailed 54, theopencode-reviewjob actually ran 7 times, 2 published verdicts): at least 4 of the 7 that got past the first gate died at the second one. So the two checkpoints are not redundant guards on the same failure — they are two independent chances for the same queue delay to void the work, and the second one wastes a runner slot that the first would have released.Why it keeps happening:
mainmerged ~every 30 minutes today while dispatch runs waited hours (one measured queue of 3 h 48 min), sobase=main/<sha>at execution time is almost never the sha carried in the dispatch payload. Any fix that only changes the first gate's semantics leaves this one in place. The invariant both checkpoints actually need is "the review is about this head SHA, and that head has not moved" — the base advancing under an unchanged head is not a trust event, and re-validating base equality after an unbounded queue wait cannot succeed while the merge rate exceeds the queue drain rate.Correction to my count above: it is 5 of 5, not 4 of 5 — and every one of them had already passed the first gate. My earlier query printed only the first failing job per run, so a run whose
coverage-evidencefailed hid the fact that itsopencode-reviewjob failed at the same checkpoint (a peer caught it on the run they own). Re-read with every failing job and step:run validate-pr-metadataopencode-reviewfailing steps34002473295(#19022d4624a)success 4 Validate pull request head repository trust; 20 receipt34010256951(#16616047f36)success 4 ; 20 34016922761(#1661ad65c73)success 4 ; 20 34011098364(#1947e4c9c44)success 4 ; 19 Publish OpenCode review outcome; 2034015973300(#19235c9920a)success 4 ; 20 (plus coverage-evidencestep 9)So the population that reaches step 4 is exactly the population that passed the allowlist, and it fails there without exception. Two consequences worth stating plainly:
- The allowlist fix is necessary but not sufficient. Admitting more dispatch identities (CodeQL dispatch terminal status publication remains unproven across repositories #1929) sends more runs into this checkpoint, which today rejects 5 of 5 of what reaches it. Expectations after that variable changes should be set accordingly.
- The cost lands after the runner is allocated. The first gate rejects in seconds; step 4 rejects after
coverage-source-treeandcoverage-evidencehave run — on the samples above that is the whole job up to the privileged boundary, and no model work ever starts. Those runner-minutes are charged against the same 60-job ceiling the queue delay is already straining, which is what makes this self-reinforcing: delay invalidates the dispatch, the invalidation burns a slot, the burnt slot lengthens the delay.
Verification of the five second-checkpoint rejections: all five were correct, and none of them indicts the checkpoint. I read each failed
opencode-reviewjob's own annotation (the executed output, not the step name) for the full field set, which printshead/expected_headandstatealongside the base pair:run PR head at check time verdict 34002473295 #1902 2d4624a3→bf732f92head moved 34010256951 #1661 6047f36f→ad65c731head moved 34015973300 #1923 5c9920a9→dc146b4chead moved 34016922761 #1661 ad65c731→946f400bhead moved 34011098364 #1947 e4c9c44f=e4c9c44fhead and base both matched exactly; state=closed— the PR had merged at 05:04:12Z and the job ran at 07:17:32ZSo the earlier reading that these show a base-identity check firing on unchanged reviews does not hold. In four the head had already moved, so a checkpoint keyed on head immobility alone would have rejected them identically; in the fifth nothing had changed except that the pull request was already merged. Relaxing the
base_shacomparison would not have saved a single one of the five, and each rejection prevented a privileged review of something that no longer existed. The base pair does appear in every message, which is what makes the base look causal when only the first mismatched field is read.What the five actually measure is queue latency. Each was dispatched, waited between one and three hours for a runner, and by the time it executed the world had moved: the author had pushed, or the PR had merged. That is the ceiling problem itself rather than a defect in the trust boundary, and it means these runs carry no evidence about the review sidecar at all, since none of them reached it.
Two consequences worth separating. Nothing here argues for changing a checkpoint that runs immediately before OIDC, review-token, and model execution. What it does argue for is that a dispatch whose head has been superseded should stop occupying a queue slot, since it can only fail once it starts, and every hour of queue depth converts more pending reviews into certain rejections.
Method note, since it cost me two wrong readings today: a run's
conclusioncannot name the failing job, a job's name cannot name the failing step, and a step name cannot name the failing field. Only the annotation does, and it prints every compared field, not just the one that differed.Retracting my diagnosis in the two comments above. The checkpoint was right in all five cases; only the cost statement survives. A peer read the full annotation — which prints every compared field — and I confirmed it myself on all five runs (I had truncated the annotation at 230 characters, which cut off the head fields, and then reasoned from the base pair that remained visible):
run PR expected head → live head state 34002473295#1902 2d4624a3→bf732f92open 34010256951#1661 6047f36f→ad65c731open 34015973300#1923 5c9920a9→dc146b4copen 34016922761#1661 ad65c731→946f400bopen 34011098364#1947 e4c9c44f→e4c9c44f(identical, base identical too)closed — the PR merged at 05:04:12Z and the review ran at 07:17:32Z So relaxing
base_shaequality would have saved none of them: four had a genuinely moved head, and the fifth was reviewing a merged pull request. My claim that "the invariant both checkpoints need is head-not-moved rather than base equality" is not supported by this sample, and the trust boundary needs no change — it rejected exactly what it should, before OIDC, tokens, CodeGraph, or model execution.What still stands: (1) the population reaching step 4 is the population that passed the allowlist, and today it is rejected without exception; (2) the rejection lands after
coverage-source-treeandcoverage-evidencehave run, so a runner slot is consumed to discover that the review's subject no longer exists. That second point is now the whole of the issue: these five measure queue latency, not validation semantics — the author pushed or the PR merged during a one-to-three-hour wait. The lever is therefore to stop dispatching work whose subject has already changed, and to release the slot before the expensive jobs run, not to weaken any comparison.- addedbugSomething isn't workingSomething isn't workingpriority: mediumNormal-priority or P2 workNormal-priority or P2 work
on Sep 7, 2026 seonghobae commented
on Sep 7, 2026 ContributorAuthorMore actionsFresh downstream canary from
ContextualWisdomLab/linux-cluster-ops#228exposes a second completion gap after the dispatch/coverage repairs.- leaf protected/default:
develop@7d6c0e6f488dffb609eded3f8980ded570b54362 - unchanged PR head:
2d9bc6e77bc88e1119459c0054b5e136f5a9f598 - latest formal review on that exact head is still
CHANGES_REQUESTED, produced by central dispatch run32020312201. - run
32020312201:validate-pr-metadataandcoverage-source-treesucceeded;coverage-evidencefailed atMeasure test and docstring evidence; downstreamopencode-reviewthen failed atValidate pull request head repository trustand published the negative review. - the PR head and protected base have not moved since that review. Current central issue text records the coverage-evidence failure class as repaired by fix(ci): close the admission-controller coverage/docstring gap on main #1883, but the old negative review remains live GitHub PR state until a fresh owner review supersedes/dismisses it.
I moved #228 to Draft and recorded the exact evidence instead of pushing a source-neutral commit or treating repository-owned successful checks as a replacement for the required review.
Please include this in the central GREEN/rollout contract: after a central review-infrastructure repair, discover open PRs whose current exact head still carries a model/bot
CHANGES_REQUESTEDwhose cited causal class is now owner-fixed; reacquire the review on the unchanged head through the normal review workflow and supersede/dismiss the obsolete infrastructure-caused verdict only after fresh terminal evidence. Do not require a dummy leaf commit, reuse predecessor checks, or convert the old negative review to success synthetically. Genuine source-backed findings must remain blocking.- leaf protected/default:
Fresh census, 2026-09-14 ~06:55Z. Method: the last 100 completed runs of
opencode-review-dispatch.yml(window2026-09-13T10:50:55Z→2026-09-14T03:44:51Z), then for every failed run the first failing job in dependency order, counted as a full census rather than a sample.Completion outcome of the window
outcome runs failure 83 cancelled 14 success 3 Three successful dispatch runs out of 100 completed. That is the throughput number behind this issue, measured end to end rather than inferred.
First failing job (all 83 failures)
first failing job runs share coverage-evidence41 49% validate-pr-metadata28 34% opencode-review14 17% Split at the moment #2123 reached
main(2026-09-13T23:37Z):period failures coverage-evidence validate-pr-metadata opencode-review before 23:37Z 71 41 16 14 after 23:37Z 12 0 12 0 The post-merge zero is not evidence of repair. Every one of those 12 aborted upstream at
validate-pr-metadata, socoverage-evidencenever ran. Checking all 36 dispatch runs created after 23:37Z: 23 stillqueued, 1pending, 12completed/failure; of the 25 inspected, thecoverage-evidencejob wasqueued(6) orskipped(4) and executed to completion in none of them. So the before-period 41 is a clean baseline, and the after-period has no comparable measurement yet.Methodological note, because it changed the answer
My first pass sampled the 12 most recent failures and found 12/12 at
validate-pr-metadata— which would have reported this issue's own failure mode as ~100% of failures. The full census puts it at 34%, behindcoverage-evidenceat 49%. The GitHub runs API returns newest-first, and while the queue churns, the newest completed runs are dominated by whichever class is failing fastest right now. Any recount here should be a census over an explicit time window, not a head sample.This does not retract the issue:
validate-pr-metadatastale-head aborts are 28 of 83 failures and are the only failure class still occurring in the post-merge window, so the self-reinforcing queue-delay mechanism described above remains live.Census completed — the third failure class is now attributed, so the whole 83 is accounted for.
Read the
opencode-reviewfirst-failure logs (the 17% bucket). Representative, run34769383287job103777189146, PR #2173 @92c23231f:OpenCode model-pool outcome=exhausted model=none; publish stage performs no duplicate model-catalog pass. ##[error]no current-head formal OpenCode review receiptTwo of the three sampled runs carry that exact pair; the third fails the same step without the pool line. So this bucket is free model-pool exhaustion, not a publication-permission or receipt-validation defect — the gate is correctly refusing to synthesize a receipt when no model produced one.
Full attribution of the 83 failures in the window (
2026-09-13T10:50:55Z→2026-09-14T03:44:51Z):first failing job runs share cause owning issue coverage-evidence41 49% trusted coverage image build aborts at the fast-mlsirm import root #2157 (repaired by #2123, merged 23:37Z) validate-pr-metadata28 34% dispatch metadata no longer matches the live head this issue opencode-review14 17% orchestrator/free model pool exhausted, no model available #2148 / #2165 / #1915 Three distinct causes, three separate owners, and none of them is a defect in the reviewed code — which is why the window produced 3 successful dispatch runs out of 100 completed.
Bearing on this issue specifically: after the #2157 repair landed,
validate-pr-metadatais the only class still failing (12 of 12 post-merge failures). With the largest cause removed, the queue-delay invalidation described here becomes the dominant remaining barrier to a dispatch run reaching a verdict.
opencode-review-dispatch.yml의 메타데이터 검증이 큐 대기 시간 때문에 실패합니다. 대기가 길수록 실패율이 오르고, 실패가 재시도를 부르고, 재시도가 대기를 늘리는 자기강화 구조입니다.#1927·#1929(낡은 인가 변수)와 독립된 원인이고, 그쪽을 고쳐도 남습니다. 수정 대상이 설정 값이 아니라 검증 의미론이라 별건으로 올립니다.측정
opencode-review-dispatch.yml의 실패 run 중actor=github-actions[bot]인 것(= 인가를 통과하는 신원) 44건 전수를 의존 순서로 재분류했습니다. 스텝을 run 간에 합산하면 하류 캐스케이드가 중복 계산되므로,validate-pr-metadata → coverage-source-tree → coverage-evidence → opencode-review에서 처음 실패한 job을 근본 원인으로 잡았습니다.최대 원인 16건을
##[endgroup]이후의 실행 출력만 읽어 다시 갈랐습니다. 그 앞은 스크립트 소스 에코라 모든 에러 문구를 포함하고 있어서 어떤 가설이든 확증됩니다.12/16이 큐 지연의 결과입니다.
실제 로그
head는 맞고 base만 어긋납니다. 디스패치는 보낸 시점의 base SHA를 싣는데, 큐에서 기다리는 동안
main이 전진합니다.state=closed3건도 검증 도달 시점에 PR이 이미 닫힌 것이라 같은 기전입니다.드리프트가 사실상 확정적입니다
오늘
origin/main의 시간당 커밋 수입니다.그리고 실측된 완주 실행 하나가 **총 15시간 37분, 그중 실제 계산 약 73초(대기 99.87%)**였습니다. 시간당 1~11회 움직이는 브랜치에서 15시간을 기다리면 base가 안 움직일 확률은 사실상 0입니다.
핵심: 이 검사는 데이터 의존이 아니라 staleness 정책입니다
opencode-review-dispatch.yml의 실제 데이터 경로입니다.SUPPLIED_BASE_SHA는 비교에만 쓰이고 버려집니다. 하류가 실제로 체크아웃하는 값은live_base_sha입니다. 즉 supplied와 live가 달라도 파이프라인은 live로 정상 동작합니다.동등성 요구가 지키는 것은 "디스패치 시점과 검증 시점 사이에 base가 안 움직였다"는 보장이고, 현재 조건에서 그 보장은 거의 항상 거짓입니다.
head_sha는 다릅니다.:188에서 같은 방식으로 비교하지만, 그건 "이 리뷰가 디스패치된 바로 그 커밋을 대상으로 한다"는 exact-head 규율의 근간입니다. head는 엄격히 유지되어야 합니다.판단이 필요한 지점
base_sha동등성을 푸는 것이 기술적으로 안전하다는 것까지가 측정으로 말할 수 있는 전부입니다. 남는 질문은 설계입니다.어느 쪽이든 신뢰 경계에 관한 결정이라 근거만 올립니다.
재현
관련: #1927, #1929, #1925, #1883