Repository navigation
Strix run ended "cancelled" at 6h with no verdict: 909 provider attempts, 92% on one route, 0 exhaustion events, request identity absent from evidence #2000
Description
Activity
Sharpening my own "runner occupancy" point before anyone acts on it
The consequences section above pairs "18
in_progress, 14 of them Strix" with "150queued, oldest waiting 39 minutes." Those numbers are what the API returned at18:1x, but putting them side by side invites the inference that the long Strix runs are what makes the queue deep. A second sample says they are not, and I would rather correct that here than have it steer the fix.Measured
2026-09-07T00:07Zon this repository:2026-09-06 18:1x 2026-09-07 00:07 in_progress18 (14 Strix) 6 (5 Strix + 1 Noema) queued150 161 oldest queued 39 min 32 min The queue is being served normally throughout: median queued age 14.1 min, newest 0.4 min, last
successconcluded 6.6 min ago and lastfailure0.6 min ago. Queue depth is dominated by arrival rate, and it grew while concurrent Strix runs fell from 14 to 5 — so it is not a backlog behind the long jobs.What the long jobs actually cost is unchanged and is what this issue is about: each of the six currently running has held a slot for 81–205 minutes, and on the #1187 evidence a scan on a degraded pool holds one until the 6-hour ceiling kills it with no verdict at all. The waste is the slot-hours and the missing gate decision, not queue latency for everything else.
This does not change the defect, the reproduction, or the requested remedy — attempt/route exhaustion rather than an elapsed-time cap. It only removes a claim the evidence does not support.
Generated by Claude Code
seonghobae commented
on Sep 7, 2026 ContributorMore actions독립 집계에 따른 주장 범위 정정
CO 담당 작업이 기존 artifact 9997372950을 다운로드해 집계했고, 이 작업은 그 감사 receipt를 직접 읽었습니다. 설치 CO는 414f22973658c4ddc3d4320fcf7acd9b4e8ba991, stderr SHA256은
ce5ed2846ecf284674cf5c6d0d80e674e79bfc4e58c27f40c878fde3bc25d70a입니다.909 starts / 603 recorded failures / 541 TimeoutError / 62 HTTPError / 536 internal_error / provider_exhausted 0 및 circuit 55·3·43 집계는 일치합니다. 다만 주된 route는 837 starts와 564 recorded failures입니다. 따라서 본문의 '837 consecutive failures'는 입증되지 않았습니다. 나머지 273 starts를 성공이나 실패로 임의 분류해서는 안 됩니다.
stderr request_id 필드는 0개이고 gate-console은 2줄, stdout은 비어 있습니다. 이 자료만으로 caller retry, 독립·동시 logical request, 요청 내부 fallback을 구분할 수 없습니다. 소진 이벤트 부재만으로 무한 caller loop를 확정할 수도 없습니다. tail의 30초 circuit reset 뒤 재허용은 관측되지만 전체 반복의 원인이나 올바른 reset 정책을 증명하지는 않습니다.
취소 시각 6h00m17s 또한 플랫폼 취소 주체의 직접 증거와 구분해야 합니다. 다음 조사는 exact-source의 caller/route exhaustion 계약과 요청 정체성을 보존한 runtime evidence입니다. 경과 시간 상한 추가, 공급자 재호출, 재실행 또는 기존 실행 취소는 하지 않았습니다.
Independently reproduced, and three of this issue's claims retracted
Re-derived from the same artifact on this side. The audit boundary matches exactly:
stderr SHA256 = ce5ed2846ecf284674cf5c6d0d80e674e79bfc4e58c27f40c878fde3bc25d70aquantity this issue's body your count re-derived now total provider_attemptstarts909 909 909 total provider_attempt_failed603 603 603 main-route starts 837 837 837 main-route recorded failures — 564 564 main-route starts − failures — 273 273 request_idoccurrences— 0 0 You are right and the issue body is wrong. Retracting three claims:
-
"837 consecutive failures" is not supported. 837 is the start count; 564 are recorded failures on that route, and 273 starts have no recorded outcome. The log emits no success-shaped event at all (
provider_attempt,provider_attempt_failed,request_failed,circuit_*,provider_backoff,provider_discovery_failed,discovery_diagnostics_completeare the only event kinds present), so those 273 cannot be classified in either direction. I should not have written a failure count I had not counted. -
"the repetition is the caller re-selecting a dead route rather than in-call retry" is not established.
attempt=1/1bounds the per-call retry budget and nothing more. With zerorequest_idfields there is no way to group 909 starts into logical requests, so caller retry, independent concurrent requests, and in-request fallback are indistinguishable from this evidence. The title's "has no exhaustion condition" inherits the same defect:provider_exhausted: 0is an absence, and an absence of exhaustion events is not proof of an unbounded loop. -
"killed by the 6-hour Actions ceiling" overstates attribution.
15:19:48Z → 21:20:05Z= 6 h 00 m 17 s andconclusion: cancelledare consistent with the platform ceiling, but I have no direct evidence naming the canceller and should have said so.
I have retitled the issue accordingly.
What the evidence still supports, unchanged
- One job produced 909 provider attempt starts and 603 recorded failures, of which 541 are
TimeoutErrorand 62HTTPError, alongside 536request_failed status=500 code=internal_error. - 92% of starts (837/909) went to one
agent_id. That is a distribution fact about starts and needs no outcome classification. provider_exhaustedappears 0 times;circuit_failure55,circuit_opened3,circuit_cleared43,circuit_reset3.- The preflight for this run:
candidate_count: 24,ready_count: 5,target_ready: 8, all five ready routes onnvidia_nim/nvidia_nim_sub([CO owner] eliminate orchestrator/free provider-family SPOF before gateway release #1915). - The outcome, which holds regardless of who cancelled it: a required check ended
cancelledafter six hours with no findings and no gate decision.cancelledis not a verdict, and Sweep pre-#1669 PRs stuck with a wrongly-cancelled current-head Strix/OpenCode/Noema check #1756 is already sweeping PRs stuck in that state.
On the next investigation
Agreed on exact-source caller/route exhaustion contracts plus runtime evidence that preserves request identity. The concrete blocker is the
request_idrow above: zero occurrences is exactly what makes this undiagnosable from artifacts. Until a correlation id is emitted per logical request and carried onto everyprovider_attempt/provider_attempt_failedline, any future run produces the same ambiguity and the same argument. That seems worth doing before, not after, the contract work — it is what would let the next run answer the question this one cannot.Noting for the record that neither side has added an elapsed-time cap, re-called providers, re-run, or cancelled anything. This issue does not ask for a time cap; if what remains after these retractions is judged too thin to act on, closing it in favour of a narrower evidence-instrumentation issue would be a reasonable outcome.
Generated by Claude Code
-
- changed the title
[-]Strix/Noema request path has no exhaustion condition: 837 attempts on one dead route, killed by the 6-hour Actions ceiling with no verdict[/-][+]Strix run ended at 6h with no verdict: 909 provider attempts, 92% on one route, 0 exhaustion events, request identity absent from evidence[/+]on Sep 7, 2026 - changed the title
[-]Strix run ended at 6h with no verdict: 909 provider attempts, 92% on one route, 0 exhaustion events, request identity absent from evidence[/-][+]Strix run ended "cancelled" at 6h with no verdict: 909 provider attempts, 92% on one route, 0 exhaustion events, request identity absent from evidence[/+]on Sep 7, 2026 - added a commit that references this issue
on Sep 7, 2026 - addedbugSomething isn't workingSomething isn't workingpriority: highHigh-priority or P1 workHigh-priority or P1 work
on Sep 7, 2026 A third sample, and it is a different regime from the second
My earlier comment here said the queue was being served normally and that its depth was arrival-rate dominated rather than a backlog behind the long Strix runs. That was true of the
00:07Zsample. It is not true now, and I would rather post the change than leave a reader with a stale reading of my own correction.Measured
2026-09-07T08:43Z:09-06 18:1x 09-07 00:07 09-07 08:43 in_progress18 (14 Strix) 6 5 — all five Strix queued150 161 163 oldest queued 39 min 32 min 161 min median queued age — 14.1 min 32.7 min queued > 60 min — — 34 last successconcluded— 6.6 min ago 53.8 min ago The five in-progress runs are Strix Security Scans aged 228, 229, 251, 270 and 324 minutes. Nothing else is running, and nothing has concluded successfully in the last 54 minutes.
A concrete cost, on a PR I drive.
.github#1884@767e740c's Required OpenCode Review, run 34082931427, was created at04:23:46Z. Its stages started at05:51:49Z(bootstrap, +88 min),05:53:02Z(cancel-superseded), and07:38:27Z(admit-current-head, +105 min after that). Its three remaining jobs —coverage-evidence,coverage-source-tree,opencode-review— have beenqueuedsince07:38:31Zwith zero steps recorded. Four hours and twenty minutes after the push, the review job has not begun. The PR itself has 17 successful checks, no failures, and an APPROVE fromcwl-noema-review[bot]on the exact head.What I am and am not claiming. Five slots occupied by multi-hour scans while 163 runs queue is a much stronger association than the second sample showed, and it is the shape this issue predicts. But three samples do not establish the concurrency ceiling's value, whether these five slots are the whole capacity, or that Strix is the cause rather than a co-occupant — I have not measured any of that, and the earlier retraction stands as written for the sample it described. What has changed is that "the queue drains normally" no longer describes the system, so the slot-hours this issue is about now plausibly cost throughput outright rather than latency.
That is the part worth acting on if the exhaustion condition lands: on the
#1187evidence a scan on a degraded pool holds its slot until the six-hour ceiling and produces no verdict, and there appear to be only a handful of slots.
Generated by Claude Code
seonghobae commented
on Sep 10, 2026 ContributorMore actionsLifeOS current-head canary (read-only consumer evidence):
ContextualWisdomLab/life-os#249@09d1430ec0f3266740266519af36e78907482ca1, required Strix run34448869518, job102780611243has now reached a terminal SUCCESS. Exact-head admission, scope detection, hardened runner setup, trusted Strix checkout, target-head fetch, contract self-test, secret gate, contextual-orchestrator sidecar provisioning, Strix install, and model/API input preparation all succeeded.Run Strix (quick)ran from2026-09-10T07:41:30Zto09:11:53Zand succeeded; report collection/upload then succeeded and the job completed at09:11:58Z. The enclosing required workflow completed SUCCESS at09:11:59Z.This sample therefore does not reproduce #2000's no-verdict failure. It is a useful healthy counterexample: a long-running current-head Strix/CO path can terminate normally without an elapsed-time/model timeout. I did not cancel or cap it. Please keep the issue focused on typed attempt/route exhaustion, request identity, and fail-closed terminal evidence that can distinguish route exhaustion/provider end from user/admin/platform cancellation, rather than treating elapsed duration itself as failure.
seonghobae commented
on Sep 12, 2026 ContributorMore actions새 unchanged-head consumer canary를 추가합니다. 이 증거만으로 provider-attempt loop를 재확정하지는 않으며, 현재 단계에서는 #2000의 동일한 외형(장시간
Run Strix (quick)에 머물고 terminal verdict가 없음)이 재현되고 있다는 사실만 기록합니다.- consumer:
ContextualWisdomLab/LineageWeave#1039 - exact head:
9da817da8306c8ac8c631ff1c19f6c74dd6e35a1 - required Strix run:
34671908661 - strix job:
103494782412 - admission / changed-scope / cancellation jobs: terminal SUCCESS
- sidecar provisioning:
2026-09-12T04:02:16Z → 04:14:42Z, SUCCESS - Strix install / model-input preparation: SUCCESS
Run Strix (quick): started2026-09-12T04:15:11Z; fresh read at 2026-09-12 07:46Z에도in_progress; report collection / artifact upload / same-head status publication은 아직 pending- 같은 exact head의 repository Tests/PROV-O/Security/SAST는 terminal SUCCESS이고, GHAS changed-source CodeQL도
No new alerts를 반환합니다. 따라서 이 장시간 상태를 LineageWeave source defect나 leaf no-op retrigger로 우회하지 않습니다.
현재 job metadata만으로는 #2000에서 관측한
provider_attempt횟수, route 재선택,provider_exhausted부재를 증명할 수 없습니다. job이 terminal이 되어 artifact가 생기면 그 증거를 기준으로 attempt/route exhaustion 여부를 판별해야 합니다. 그 전에는 elapsed time 자체를 모델 실패 판정으로 쓰지 않고, exact-head scan evidence가 아직 미완성이라는 상태만 유지하는 것이 맞습니다.이 canary는 LineageWeave 쪽에서 취소하거나 provider/model fallback을 추가하지 않고 그대로 보존 중입니다.
- consumer:
seonghobae commented
on Sep 12, 2026 ContributorMore actionsFresh unchanged-head canary from
LineageWeave#1039@9da817da8306c8ac8c631ff1c19f6c74dd6e35a1is now terminal and should be added to the exhaustion/control-plane evidence set.- Strix run
34671908661, job103494782412 Run Strix (quick)started2026-09-12T04:15:11Zand failed2026-09-12T08:39:53Zafter ~4h24m42s.- Admission/current-head verification, target checkout, required-workflow self-test, secret gate, CO sidecar provisioning, Strix installation and model-input preparation all completed successfully before the failing step.
- Report collection and artifact upload completed successfully after the failure; the same-head manual status-publication step was skipped.
- This run failed before GitHub's 6-hour hard ceiling, so it is distinct from the original
cancelledsample. It confirms that the current path can occupy a runner for multi-hour execution and eventually produce a terminal failure without any LineageWeave head movement.
I am deliberately not inferring an unbounded provider loop from job duration alone. The workflow/job metadata does not expose request identity, provider-attempt counts, route exhaustion, or circuit-breaker evidence. Please correlate the uploaded Strix artifact/CO sidecar logs for this run with this issue's bounded-attempt hypothesis before attributing cause.
LineageWeave has returned #1039 to Draft because the long-running validation lane is now terminal and CodeQL/OpenCode/Noema/Strix are not all authoritative GREEN. No no-op retrigger, provider override, or leaf-side gate shim was introduced.
- Strix run
seonghobae commented
on Sep 12, 2026 ContributorMore actionsAdditional unchanged-head consumer canary from
ContextualWisdomLab/LineageWeave#1046@f80c0ec5f35fd4fc0a867125bc46dfd2d2dee9a9:- Strix run
34692420821, main job103550422240, is now terminal failure rather than platform-cancelled. - Admission, changed-scope detection, target materialization, contextual-orchestrator sidecar provisioning, Strix install, and model-input preparation all succeeded.
Run Strix (quick)ran from2026-09-12T12:18:04Zto13:17:44Zand failed; report collection/upload succeeded afterward; same-head manual status publication was skipped.- One retained
strix-reportsartifact exists: id10298596767, 33,038 bytes, digestsha256:0ab02471f2d81667914ae48281aaa0808bd239330bb5b8188a51cf2725b5fd8b.
This is useful as a post-#2000 consumer canary because the scan now terminates before the six-hour platform ceiling, but workflow/job metadata alone is not sufficient to claim that #2000's attempt/route-exhaustion defect is fixed or to classify this terminal failure as provider exhaustion versus source-backed Strix findings. The retained artifact must be inspected for request identity, provider-attempt/exhaustion/circuit evidence and report findings before making that causal claim. No LineageWeave-local provider/status shim or elapsed-time timeout was added.
- Strix run
seonghobae commented
on Sep 12, 2026 ContributorMore actionsArtifact follow-up for the LineageWeave #1046 canary above: the retained evidence is now inspected, so the attempt/exhaustion comparison can be made precisely.
contextual-orchestrator-sidecar.stderr.logcontains 116 provider-attempt starts, 53 attempt failures, 41request_failed status=500 code=internal_error, 0provider_exhausted, and 0request_id=fields.nvidia_nim_deepseek_ai_deepseek_v4_flash_0731received 99/116 attempts (85.3%), with 42 failed attempts. The scan nevertheless completed successfully and emitted internally consistent zero-finding terminal evidence (run.json: status=completed, scan_completed=true, success=true; SARIF results empty; report says no exploitable vulnerability).The wrapper then failed closed with
STRIX_PROVIDER_UNAVAILABLE: ... exhausteddespite the sidecar recording zeroprovider_exhaustedevents. So the original six-hour non-termination did not reproduce in this consumer — Strix reached a terminal report in about an hour — but two #2000 observability/termination-contract defects remain visible: request identity is still absent, and the terminalexhaustedlabel is not backed by aprovider_exhaustedreceipt. The complete-consistent-vs-recovered-fault verdict collapse is separately recorded on #2026, where this artifact is a direct acceptance fixture.seonghobae commented
on Sep 12, 2026 ContributorMore actionsLineageWeave unchanged-head Noema canary adds a sharper exhaustion/reselection case.
Consumer:
ContextualWisdomLab/LineageWeave#1042@7381233b12b7160a0c0c08d9749334a9e05eb862
Required Noema run/job:34699867159 / 103570052890
Result: terminal FAILURE at 2026-09-12T15:21:47Z. Exact-head admission, trusted gate materialization, credential mint, head validation, and contextual-orchestrator sidecar provisioning all succeeded.Prepare Noema model verdictalone failed; publication was skipped. Failure artifact:noema-sidecar-evidenceid10300004597, digestsha256:3f65b2fe7b72153bc5f7c25eee1a568a0ff253c3d7f7e1b9f44dc6617af95bdc.Artifact observations, without inferring logical-request identity that is not recorded:
- preflight: candidate_count=24, ready_count=6, target_ready=8; every ready route is NVIDIA NIM/NIM_SUB; OpenRouter free candidates were deferred on 429; several NVIDIA discovery candidates were rejected on 404/timeout.
- stderr:
provider_attempt69,provider_exhausted1,request_idoccurrences 0. - both
nvidia_nim_deepseek_ai_deepseek_v4_flash_0731and its SUB sibling opened circuits at threshold=3 with reset_seconds=30.0. - the primary flash route then reached an actual
attempt=1/3 -> 2/3 -> 3/3, emittedprovider_exhaustedat15:20:14.239Z, and 59 ms later the same agent_id was selected again at15:20:14.298Zas a newattempt=1/1. - that post-exhaustion same-route attempt timed out at
15:21:44.411Z; the request then endedstatus=502 code=provider_connection_error.
This does not prove whether the reselection is a new logical request or caller retry because request identity is absent. It does prove that an emitted per-call/provider exhaustion receipt does not by itself keep the exhausted route out of immediate subsequent selection. That is directly relevant to this issue's route-exhaustion/termination contract. Please preserve request identity in the evidence and make exhaustion state observable across selection boundaries; do not solve this with a wall-clock model timeout.
seonghobae commented
on Sep 12, 2026 ContributorMore actionsFresh downstream canary from
ContextualWisdomLab/Orgmetra#64@c8d1c3993ce1eb4e8bdabfe4b666c61eda57fcffadds a different timeout shape that should be covered by this owner path before any blind rerun.Required Strix run/job:
34704247956/103581586116. Trusted reusable-workflow source was protected.github@fb17ef556f94f673234aa557254ae52779e9a7b0; vendored contextual-orchestrator was414f22973658c4ddc3d4320fcf7acd9b4e8ba991.Admission and setup were healthy: exact-head/current-PR admission passed, trusted source materialization passed, CO sidecar provisioning passed, Strix installation/configuration passed. CO discovery selected 24 free candidates with
priced_selected_count=0; runtime preflight found 2 ready routes, and gatewaychat/completionspreflight succeeded on attempt 1 withfinish_reason=stop.The workflow/compat boundary explicitly requested unbounded inference:
LLM_TIMEOUT=0,STRIX_MEMORY_COMPRESSOR_TIMEOUT=0,STRIX_PROCESS_TIMEOUT_SECONDS=0,STRIX_TOTAL_TIMEOUT_SECONDS=0, withCWL_STRIX_UNBOUNDED_INFERENCE=1. Nevertheless the liveRun Strix (quick)path ran for about 1880 seconds and terminated with Strix'sLLM CONNECTION FAILED,Could not establish connection to the language model,Error: Request timed out.The gate then recorded modelorchestrator/freeexit code 1 and correctly failed closed asSTRIX_PROVIDER_UNAVAILABLEbecause no authoritative vulnerability report was produced. Retained artifactstrix-reportsid10302600731, GitHub digestsha256:b00d613097ac0cd838d08bc1a95200697e4ad23ec4f644f8ffc9699d4d78029econtains the runtime evidence/logs but no completed structured vulnerability verdict.This is not evidence for adding a wall-clock model timeout, a direct provider/model fallback, or treating 1880 s itself as failure. It instead shows that the requested null/unbounded timeout semantics are not yet demonstrably preserved end-to-end through the Strix client / LiteLLM-compatible boundary / CO request / selected provider path. The next owner repair/evidence should preserve request identity and timeout provenance well enough to identify which layer generated the terminal timeout, while keeping
orchestrator/freeand fail-closed behavior. If the provider itself terminated the request, record that distinctly from an administrative/client timeout; if an intermediate client imposed a hidden timeout despite the zero/null contract, that is the causal defect to remove.I did not rerun, cancel another run, add a provider fallback, or change the Orgmetra security gate.
seonghobae commented
on Sep 13, 2026 ContributorMore actionsFresh bounded-but-still-unattributed request/exhaustion sample from
ContextualWisdomLab/Orgmetra#64@d9cc516d54b4642f59fe126c331a19945fdf75f1, Strix run/job34731009530/103653717441, artifact10311173510(sha256:3cb9623b44ba03143a0955b96644251d81f8902514664708d4e9d9fb78679878).The sidecar ran from the first provider attempt at
2026-09-13 01:43:14Zthrough the last at04:06:12Zand records 297 provider-attempt starts, 169provider_attempt_failed, 0provider_exhausted, with 268/297 starts (~90%) selectingnvidia_nim_sub_deepseek_ai_deepseek_v4_flash_0731. Repeated failures includeTimeoutErrorandrequest_failed status=500 code=internal_error; circuit telemetry iscircuit_failure=11,circuit_opened=1,circuit_cleared=9. The log contains norequest_id=,trace_id=,correlation_id=, orinvocation_id=fields.Unlike the original six-hour sample, this scan eventually completed, so this is not evidence for adding a wall-clock cap. It is fresh evidence that route-level repetition and request-identity/exhaustion attribution are still weak: the dominant route was selected repeatedly, no provider-exhaustion event was emitted, and the final consumer gate nevertheless diagnosed
STRIX_PROVIDER_UNAVAILABLE ... exhaustedeven though terminal Strix evidence was produced. #2026 now has the corresponding terminal-evidence classification fixture; this issue should continue to own request identity / route exhaustion mechanics.seonghobae commented
on Sep 13, 2026 ContributorMore actionsLineageWeave current-head Strix canary를 추가합니다. 이 증거는 경과 시간만으로 provider loop 또는 모델 실패를 판정하지 않습니다. 현재 확인 가능한 사실은 exact-head required scan이 multi-hour
Run Strix (quick)상태를 유지하면서 아직 terminal verdict/artifact를 만들지 못했다는 점입니다.- consumer:
ContextualWisdomLab/LineageWeave#1055 - exact head:
50c4935eef1029467595f7004818643598b737c9 - required Strix run/job:
34746057545 / 103694153476 - admission / changed-scope / superseded-run cancellation: terminal SUCCESS
- trusted source/workspace materialization, required-workflow self-test, credential gate: SUCCESS
- contextual-orchestrator sidecar provisioning:
2026-09-13T07:47:26Z → 08:02:53Z, SUCCESS - Strix install/model-input preparation: SUCCESS
Run Strix (quick): started2026-09-13T08:03:24Z; fresh read at approximately2026-09-13T12:45Zstillin_progress- report collection / artifact upload / same-head status publication: still pending
- same exact head's repository Tests (including full PostgreSQL), Security, and SAST are terminal SUCCESS. Required CodeQL/OpenCode/Noema failures are separately owned central/CO settlement/readiness boundaries and are not being converted into a LineageWeave source finding.
This canary has now remained in the actual scan step longer than the prior LineageWeave #1039 sample did before its terminal failure (~4h24m), but that comparison still does not establish attempt/route exhaustion. While the job is live, workflow metadata exposes no request identity, provider-attempt count, route reselection, or
provider_exhaustedevidence. Do not cancel it or add a wall-clock model timeout merely because it is long-running; preserve the exact-head run until terminal evidence/artifact exists. If it terminalizes with an artifact, correlate that artifact with #2000's request-identity/attempt-exhaustion acceptance before assigning cause.No leaf no-op retrigger, provider/model override, fallback, synthetic status, or source mutation was introduced.
- consumer:
seonghobae commented
on Sep 13, 2026 ContributorMore actionsFresh LineageWeave canary resolves the previously live case without weakening the issue's attempt-based acceptance.
ContextualWisdomLab/LineageWeave#1055@50c4935eef1029467595f7004818643598b737c9Strix run34746057545is now terminal SUCCESS. Main job103694153476ran 07:47:12Z→13:31:08Z;Run Strix (quick)itself ran 08:03:24Z→13:31:03Z (about 5h27m39s) and completed successfully. Report collection/upload also succeeded. Artifact10318318320(strix-reports) has digestsha256:45c82c7f1f044185a4795b4b4db37d2ce30346e44e516ce027f2bf53fe89581e.This is a useful counterexample to any elapsed-time-only inference: a >5 h Strix quick scan can still complete normally and produce inspectable evidence. In this case the artifact contains one Medium CWE-862 product finding, so workflow SUCCESS is execution success, not 'zero findings'. The finding has been verified independently against protected LineageWeave main and is tracked in LineageWeave#1078.
I am not using this success to weaken #2000. It instead sharpens the boundary: wall-clock duration alone is non-diagnostic; the actionable failure signal remains observed request/route identity, repeated failed attempts, circuit/exhaustion state, and whether a terminal gate verdict can be emitted before the platform ceiling. Keep the issue's attempt-/route-bounded acceptance rather than adding a shorter wall-clock model timeout.
What happened
strixon.github#1187@541cadd1ran from15:19:48Zto21:20:05Zand was cancelled at exactly 6 h 00 m 17 s — GitHub Actions' hard per-job ceiling, not any policy timeout. The required check is nowcancelled, which is not a verdict at all: no findings, no gate decision, and a required context that branch protection cannot interpret.Job 101503665803; every step succeeded through
Prepare Strix model input file, andRun Strix (quick)is the cancelled one. Evidence artifactstrix-reports(id 9997372950) uploaded successfully, so the whole run is inspectable.The run never had a termination condition
From
contextual-orchestrator-sidecar.stderr.log(2,691 lines,15:25:29Z → 21:19:46Z):provider_attempt(starts)provider_attempt_failederror_type=TimeoutErrorerror_type=HTTPErrorrequest_failed status=500 code=internal_errorprovider_exhaustedcircuit_failurecircuit_openedcircuit_clearedThe final log line is a new
provider_attemptat21:19:46, 19 seconds before the platform killed the job. Nothing in the run was converging; it was stopped from outside.92% of all attempts went to one route. Attempts per
agent_id:nvidia_nim_deepseek_ai_deepseek_v4_flash_0731nvidia_nim_sub_deepseek_ai_deepseek_v4_pro_0813nvidia_nim_deepseek_ai_deepseek_v4_pro_0813The 13 single-attempt rows are the preflight probes. After preflight, the request path re-selected one already-failing route 837 times over six hours.
This is not in-call retry:
attempt=is1/1on 904 of 909 attempts (the other five are one1/3→2/3sequence). The per-call retry budget is 1. The repetition comes from the caller re-selecting the same route after each failure, andprovider_exhaustednever firing once in 837 consecutive failures on it.Sustained, not a tail-end artifact — attempts per hour:
15:0051 (partial hour),16:00258,17:00196,18:00204,19:00146,20:0041,21:0013.The circuit breaker did not keep the dead route out
55
circuit_failureproduced only 3circuit_opened, against 43circuit_cleared. An earlier run on this same pool loggedcircuit_opened … threshold=3 reset_seconds=30.0. A 30-second reset against a failure cycle whose dominant mode is a ~90-second socket idle timeout means the breaker re-admits the same route roughly every cycle, so it records failures without ever excluding the route for a meaningful interval. I am reporting the observed counts and that outcome; I have not read the selection code closely enough to assert which specific branch decides re-admission.Why this is not the same issue as #1915
#1915 tracks the free pool having no provider-family diversity, and this run reproduces that again — third independent confirmation today:
with all five
readyroutes onnvidia_nim/nvidia_nim_sub,provider_discovery_failed provider=bytez code=http_status_500, and both OpenRouter free routes deferred.But the two are separable. Restoring provider diversity would change which routes get hammered, not the fact that a degraded pool produces an unbounded retry loop. An upstream-wide outage with a perfectly diverse pool yields the same six-hour burn. Conversely, a bounded exhaustion condition would turn this run into a fail-closed verdict in minutes regardless of pool composition.
Consequences already observable
cancelledrequired checks. Sweep pre-#1669 PRs stuck with a wrongly-cancelled current-head Strix/OpenCode/Noema check #1756 is already sweeping PRs stuck behind wrongly-cancelled Strix/OpenCode/Noema checks. This is a live production path that creates them.2026-09-06T18:1xon this repository:status=in_progress= 18 runs, 14 of themStrix Security Scan, aged 97–381 minutes, againststatus=queued= 150 with the oldest waiting 39 minutes. If each stuck scan runs to the 6-hour ceiling, that is a large fraction of the runner pool held by runs that will produce no verdict. Related but distinct from ops: three required workflows each boot a runner for the same "Detect changed scope" job #1976.What this issue is not asking for
Not a wall-clock timeout on the model path.
docs/product-goal-directive.md§8 accepts that central OpenCode/Strix/Noema may take more than two hours per model, and #1889/#1890/#1892 each added a 900-second cap on genuine multi-hour-hang evidence and were all reverted (#1891, #1895). Elapsed inference time must not become a model-failure verdict, and nothing here argues otherwise: a single 90-second timeout is a normal event, and 837 of them is not a slow model.The lever is attempt/route exhaustion, which is orthogonal to elapsed time:
provider_exhaustedfiring after a bounded number of consecutive failures on the sameagent_id, so the selector stops re-picking it.readyroute has been exhausted, so the gate fails closed with evidence instead of being killed with none.All three bound attempts, not duration. A route that is genuinely slow but progressing is untouched by any of them.
Reproduction
strix-reportsid 9997372950 on run 34036172117 (contextual-orchestrator-preflight.json,contextual-orchestrator-sidecar.stderr.log,gate-console.log).noema-reviewon docs: confirm review pipeline already routes through orchestrator/free, not NIM directly #1884@396b4dee, artifactnoema-sidecar-evidenceid 9994541963 —ready_count: 6, allnvidia_nim/nvidia_nim_sub, endsrequest_failed status=502 code=provider_connection_error.Generated by Claude Code