Skip to content

[Security] Fail closed when Strix produces no authoritative scan evidence #891

Description

@seonghobae

Problem

A Strix provider/backend outage can currently leave the required workflow in a skipped or neutral-success shape without authoritative scan evidence. GitHub may treat successful, skipped, or neutral required-check conclusions as satisfying the check name, so transport/provider availability can be confused with security success.

Required contract

A mandatory security gate passes only when authoritative Strix evidence is present for the exact source head and relevant live-base context. Provider/tool unavailability is a typed deferred or failing state, never scan success.

Acceptance criteria

  • Add a terminal always-running gate that evaluates every prerequisite and exact-head Strix receipt.
  • Missing, skipped, neutral, cancelled, action-required, predecessor-head, synthetic, or untrusted-producer evidence is non-passing when Strix is required.
  • Retry only classified transient provider failures within attempt and wall-clock budgets.
  • Preserve useful bounded diagnostics while applying publication-boundary credential redaction.
  • Tests cover provider outage, empty output, skipped job, neutral conclusion, timeout, stale head, fallback exhaustion, and valid current-head finding/no-finding evidence.
  • Ruleset documentation names the expected check source/producer.
  • Protected-main consumer evidence demonstrates both fail-closed outage and recovery.

Related work

Coordinate log redaction with #842 and retry policy with ADR-0003.

Activity

  1. seonghobae commented on Aug 17, 2026

    @seonghobae
    ContributorAuthor

    Fresh ScopeWeave exact-head evidence reproduces this fail-open class on a required Strix lane and should be treated as central-owner repair input, not leaf success.

    Target: ContextualWisdomLab/scopeweave#523

    • protected live base: develop@44e7903cf8891c65410f7fc6ca5144de3fdb5185
    • exact contributor head: 58b542103a3c6009d693d32410340c1262ee87a3
    • central Strix run/job: run 31978036107, job/check 95240357165
    • GitHub check state: status=completed, conclusion=success
    • annotation: Strix backend unavailable — "Strix could not complete because its LLM backend was unavailable (rate limit / token cap / connection or warm-up failure) before producing a vulnerability report. Treating as a neutral skip so an infrastructure outage does not block merges; genuine findings still fail the check."

    That result conflicts with this issue's acceptance contract that provider/backend unavailability, missing report evidence, neutral/skipped evidence, and other unavailable-review states are non-passing when Strix is required. ScopeWeave therefore treats the current strix result as non-authorizing despite the green check conclusion and will not weaken or patch its product branch to compensate.

    Required central acceptance/revalidation for this reproduction: make the exact same backend-unavailable/no-report condition non-passing (or a typed deferred state that cannot satisfy the required workflow), retain bounded classified retry/fallback, preserve genuine finding failure semantics, then rerun Strix against this unchanged ScopeWeave head and require authoritative exact-head report evidence before merge classification.

  2. seonghobae commented on Aug 17, 2026

    @seonghobae
    ContributorAuthor

    Fresh Context Fabric reproduction on the protected Enterprise Architecture Core synchronization root confirms the same fail-open Strix contract and blocks treating the lane as merge-authorizing.

    Target: ContextualWisdomLab/enterprise-architecture-core#14

    • protected live base: develop@1c0fa8b15ceb9e72186274aeb255d6777eb84ef4
    • exact contributor/head branch: main@ca6889497728e1a3f09d68790a9096576e13a3ff
    • GitHub merge candidate: 30ef39d345bb00fac57a511732a6dd4ddbdf2361
    • central Strix run/job/check: run 32036458615, job/check 95407832860
    • GitHub check state: status=completed, conclusion=success
    • exact annotation: Strix backend unavailable; the backend hit a rate-limit/token-cap/connection/warm-up failure before any vulnerability report and the workflow explicitly says it is treating that condition as a neutral skip so it does not block merges.
    • live org ruleset 18156473 requires .github/workflows/strix.yml on protected develop.

    This is non-passing under the issue acceptance contract and Context Fabric gate policy: rate-limited/backend-unavailable/neutral/no-report evidence cannot authorize integration merely because GitHub received a green conclusion. No leaf branch change or gate weakening is appropriate.

    Required central RED/GREEN acceptance for this exact reproduction: first preserve a regression proving the backend-unavailable/no-authoritative-report condition cannot satisfy the required Strix gate; then make that state terminal non-passing (or typed deferred in a way the required gate cannot accept), retain bounded classified retry/fallback and credential-safe diagnostics, and revalidate the unchanged exact #14 candidate with authoritative Strix report evidence. Context Fabric will re-evaluate #14 only from fresh exact-head/base evidence after the central repair.

  3. seonghobae commented on Aug 20, 2026

    @seonghobae
    ContributorAuthor

    DiskSage has a concrete leaf-repository reproduction that should be included in this central acceptance matrix rather than patched locally: ContextualWisdomLab/disksage#212 exact head be1696159b97c9714b39c0133e0b64e42860e48b, required Strix run 31689301907, attempt 4, job 94632920593, central workflow source ContextualWisdomLab/.github@6eb06cdd08c79a06f7b390069d4ffa49e2eb7dba. The PR changed Rust production paths under src-tauri, but the Strix target materialized through the central __PR_SCOPE__ path contained workflow/support material rather than the changed product source. Strix explicitly reported that no source code was available for assessment, then the required workflow still concluded success/zero findings. DiskSage tracks the leaf evidence in #221 and must not work around this by synthesizing source or weakening the required check.

    Please include this as a distinct RED acceptance fixture for #891: when exact diff discovery contains at least one scannable product/config/workflow blob, PR-scope materialization must preserve every admitted blob at exact head bytes (execute bits stripped), emit a bounded path + head blob SHA + materialized digest/source-class manifest, require at least one scannable materialized file, and reject a Strix result that says no source/no code/no assessable target or otherwise lacks structured coverage of that manifest. The terminal receipt must bind repository, exact head, live-base context, manifest digest, and completed scan result. Missing/mismatched/stale/empty/no-source evidence stays non-passing. A repaired central canary should be replayed against an unchanged DiskSage head containing a real src-tauri change before this leaf blocker is considered resolved.

  4. seonghobae commented on Aug 20, 2026

    @seonghobae
    ContributorAuthor

    Fresh Inkspan exact-head reproduction for the existing Strix owner contract:

    • Consumer: ContextualWisdomLab/inkspan PR feat(cloudflare): DNS-as-code + reusable Pages deploy for 6 org zones #362
    • Exact head: 11d5cfecdcc0949ec98e6ca110d482124bff00c4
    • Required Strix run/job: 32344526380 / 96350330274
    • The primary and fallback NVIDIA NIM attempts returned HTTP 429 rate-limit failures.
    • The fallback then emitted a HIGH CVE-2022-25314 for golang.org/x/oauth2, attributed to SECURITY.md and a Strix workspace path. Inkspan and the PR changed-file set contain zero Go files, go.mod, or go.sum files and no golang.org/x/oauth2 reference.
    • The run therefore has no source-backed Inkspan dependency finding; it remains correctly non-authorizing because the provider/model and attribution evidence are not trustworthy.

    Please include this as a central RED fixture for #891: provider rate-limit/fallback exhaustion must remain terminal non-passing, and any fallback finding must be bound to an exact materialized changed-file manifest before it can fail or authorize a leaf PR. Inkspan will not add a local dependency rewrite, suppress the finding, rerun unchanged evidence, or weaken the required gate. The leaf can be revalidated only after the central repair and a fresh exact-head scan.

  5. seonghobae commented on Aug 20, 2026

    @seonghobae
    ContributorAuthor

    Inkspan follow-up: PR #373 (head 724b45b8d824be9581f0f29cc6eb07e6aeceb70f) raises the workspace override floors for fast-uri, nanoid, and postcss and refreshes pnpm-lock.yaml. Local pnpm audit, typecheck, package-config tests, build, and full package verification pass. The PR is open and not merged; protected-main review and required checks remain authoritative. No Strix suppression or protection bypass is being used.

  6. seonghobae commented on Aug 20, 2026

    @seonghobae
    ContributorAuthor

    Current central owner lane for this issue: PR #1153 (fix(strix): fail closed on incomplete provider scans), exact head 2e8e78271481a5ac0d7c6502973aa9ee8cfbc206 against .github/main@aa8503f4383e8328d89104796bc3e9f7da810376. It contains the typed non-passing provider evidence and exact scan-boundary repairs relevant to Inkspan #362. Current state: 38 checks completed, Strix in progress, 7 queued, no failures, and no formal approval. Inkspan will not rerun unchanged Strix or bypass the central gate; revalidation remains downstream of this protected central repair.

  7. added
    area: apiAPI, protocol, event, or external contract
    area: authAuthentication, authorization, identity, or tenant isolation
    area: ci-cdCI, GitHub Actions, checks, release, or supply chain
    area: securitySecurity boundary, hardening, or vulnerability prevention
    status: triagedOpen issue has an organization taxonomy assignment
    type: featureNew or expanded product capability
    on Aug 22, 2026
  8. seonghobae commented on Aug 23, 2026

    @seonghobae
    ContributorAuthor

    Fresh owner repair: #1255 exact head a1fc43a13085b4dbe9367a871e33461c4eea9ae0 addresses the new authoritative false-negative boundary observed in #897 run 32633538138 job 97191695465 and #941 run 32631414731 job 97191719493. Both complete NVIDIA NIM scans reported Vulnerabilities 0 but were rejected because Strix's exact MODEL QUALITY WARNING advisory was classified as provider failure; #941 also reproduced the report-log form. Two deterministic RED fixtures precede the narrow shared classifier fix. Exact-head hosted quality run 32640211595 passed 1,393 tests, 16 subtests, and the full quick-gate. Keep #891 open through reviewed protected-main integration and unchanged #897/#941 operational reruns; genuine provider exhaustion remains non-passing.

  9. seonghobae commented on Aug 27, 2026

    @seonghobae
    ContributorAuthor

    Fresh OriginWeave owner-route evidence (2026-08-28 KST), preserving fail-closed semantics and without any leaf workaround:

    Two independent exact OriginWeave heads are currently blocked by the central required Strix workflow at the provider/fallback boundary rather than by a demonstrated leaf finding.

    • OriginWeave#37 exact head 2aafc9893f7d18a4b62d90d55f77687be39c5bb4: required strix check 98537456818 is completed failure in central run 33046673523.
    • OriginWeave#40 exact head 3c6bcd47773132fbd6584390a2e57988b2539454: required strix check 98454281982 is completed failure in central run 33048578455 attempt 2. The job reaches Run Strix (quick) and then fails after the trusted central workflow/materialization/setup steps succeed.

    The #40 exact-head strix-reports artifact (9639384592, digest sha256:7519499d5ba59e482274415ce5e3e8a4d3478345a5ae8865fcf3c5a09547fa05) shows a provider cascade rather than a target-code vulnerability verdict: NIM primary repeatedly receives HTTP 429, the NIM fallback returns 404 Function Not Found, the OpenRouter fallback returns API 502 / Invalid URL from provider Stealth, and the OpenAI-direct fallback ends in insufficient_quota / credit_balance_exhausted. An intermediate scan path reports 0 vulnerabilities, but the required workflow correctly refuses to promote incomplete/provider-failed evidence to success.

    Please treat this as a fresh consumer reproduction for #891's authoritative-evidence/retry contract. The needed repair belongs in the central Strix provider/fallback/retry path: preserve bounded retries and fail-closed publication, but ensure currently configured fallback providers are actually usable/authority-equivalent and classify the repeated provider failures without requiring leaf source churn. After central repair, rerun/revalidate the unchanged exact OriginWeave heads rather than manufacturing a manual success status.

  10. 24 remaining items

  11. added
    bugSomething isn't working
    type: bugDefect or incorrect behavior
    on Sep 7, 2026
  12. seonghobae commented on Sep 8, 2026

    @seonghobae
    ContributorAuthor

    Fresh protected-consumer Strix runtime evidence from Naruon #1612 exact 3da3ae8e60e1bb049f59ae86bfe82db12b7e3cc7: Strix run 34271416220, job 102214101801, trusted reusable workflow .github@7fd571dbcdbae6acf29d8f4ee704d7ba6297e4db.

    The orchestrator/free gateway sidecar itself reached ready state (61 admitted / 24 selected / 3 ready). strix-agent==1.5.3 then failed before an authoritative scan because its Caido sandbox guest login could not connect to 127.0.0.1:48080 after 10 attempts. The workflow performed its one classified sandbox-specific same-model retry; attempt 2 failed with the same loginAsGuest/48080 signature. Final typed result was STRIX_PROVIDER_UNAVAILABLE: STRIX_SANDBOX_UNAVAILABLE, explicitly naming the Strix sandbox rather than the LLM gateway, and the required check failed closed.

    This matches the contract retained when #1185 was closed as a duplicate of #891: the signature must remain non-passing until authoritative scan evidence exists. It is now a fresh recovery/RCA datum for this canonical integration lane, not a Naruon security finding and not a CO routing failure. Please repair the Strix/Caido bootstrap path (or its owned runtime packaging/lifecycle) with a real RED->GREEN and then replay the unchanged Naruon head; do not neutral-success this signature or mask it with a leaf retry loop.

  13. seonghobae commented on Sep 9, 2026

    @seonghobae
    ContributorAuthor

    Fresh Naruon consumer reproduction on ContextualWisdomLab/naruon#1619@8cb1157d7b30e79b5531c2b4917bda6c3bbae9db is terminal and matches the Strix sandbox-unavailable owner path.

    Required Strix run 34338934606, job 102425631244 completed failure. Scope/admission, exact-head fetch, Strix contract self-test, secrets gate, contextual-orchestrator sidecar provisioning, Strix 1.5.3 install, gateway inputs and artifact publication all succeeded. Failure is isolated to Run Strix (quick).

    The uploaded strix-reports artifact (10099472459, sha256 1f3434e452ea290dc450051e0e41f87ea026bc731ef1a25e40e8fd44ca31ab78) contains two failed scan attempts. Both created ghcr.io/usestrix/strix-sandbox:1.3.0 successfully and resolved a host endpoint, but Caido never became reachable inside the sandbox: loginAsGuest failed 10/10 times with curl exit 7 against 127.0.0.1:48080. Both run.json receipts report status=failed, llm_usage.requests=0, zero tokens and zero findings. The second attempt started at 10:27:20Z and ended at 10:28:31Z with the same failure. This is therefore not an LLM/provider timeout and not a Naruon source-analysis finding.

    Keep fail-closed semantics. Canonical acceptance remains: repair the Strix/Caido sandbox lifecycle in .github, prove current-head authoritative finding/no-finding evidence after recovery, then replay the unchanged consumer head. Naruon will not add retries/model/provider pins, synthetic statuses, or local workflow copies for this defect.

  14. seonghobae commented on Sep 9, 2026

    @seonghobae
    ContributorAuthor

    Fresh exact-head Naruon reproduction on ContextualWisdomLab/naruon#1623@09cb87a25e59b0b5e737f915f77b404cafe245ab is now terminal and independently confirms the same Strix/Caido sandbox-unavailable owner defect as the later #1619 specimen.

    Required Strix run 34330773245, job 102400192673 completed failure. Exact-head admission/scope, target materialization, Strix workflow self-test, secrets gate, contextual-orchestrator sidecar provisioning, Strix install, gateway input preparation, report collection and artifact publication all succeeded. Failure is isolated to Run Strix (quick).

    Artifact strix-reports id 10096915829, SHA-256 8f8970ddb2fd8ecf20ab6901f943dd0026cf04d4fa0c32999b7a83b2a4a81bb1 contains two failed quick-scan attempts (strix-pr-scope-wf4gfm_8a45, strix-pr-scope-wf4gfm_4824). In both attempts Strix 1.5.3 resolves openai/orchestrator/free, successfully creates ghcr.io/usestrix/strix-sandbox:1.3.0, resolves the host endpoint (127.0.0.1:32768 / :32769), then fails loginAsGuest 10/10 times because Caido is unreachable at container 127.0.0.1:48080 (curl exit 7). Both run.json receipts are status=failed with llm_usage.requests=0, zero input/output tokens and zero findings.

    This exact-head failure therefore occurs before any LLM request and is not a contextual-orchestrator/provider failure or a Naruon security finding. Keep fail-closed semantics and repair the sandbox/Caido lifecycle in the canonical .github Strix owner. Acceptance remains unchanged-head authoritative finding/no-finding evidence after recovery; Naruon must not add local retries, provider/model pins, synthetic statuses, workflow copies, or source mutations just to retrigger the gate.

  15. seonghobae commented on Sep 9, 2026

    @seonghobae
    ContributorAuthor

    Naruon #1619 current-head replay materially narrows the prior Caido evidence and should be retained as a recovery/RCA specimen, not merged with the predecessor sandbox-only failure.

    Exact consumer head: 0b679c8a9c473be0469cc3866c106d95062a2ac9. Required Strix run/job: 34341866828 / 102434298195. Artifact: 10102755364, sha256:691bed119f65d786c0a08ed1f98d84c726a1206b1f0699139d696e721774292b.

    The same sandbox startup signature appeared first: Caido loginAsGuest could not reach 127.0.0.1:48080 on attempts 1-4. Unlike the earlier Naruon specimens, attempt 5 recovered, a Caido project was selected, the sandbox became ready, and the security scan actually ran. run.json records 86 LLM requests / 7,183,925 input tokens / 5,423 output tokens. The terminal failure was later in the contextual-orchestrator provider path: HTTP 400 invalid_request_error from meta/llama-3.2-11b-vision-instruct, followed by STRIX_PROVIDER_UNAVAILABLE. SARIF contained zero findings, but the scan status was failed, so the required gate correctly remained non-passing.

    For #891 this is positive evidence that transient Caido startup failure can recover without a leaf retry loop, and that zero findings plus substantial partial execution still must not count as authoritative clean evidence when the scan terminates failed. The remaining request-scoped routing/capability defect is handed to contextual-orchestrator#1106 (comment 5602918992). Do not neutral-success this run or add a Naruon/source workaround.

  16. seonghobae commented on Sep 9, 2026

    @seonghobae
    ContributorAuthor

    Fresh Orgmetra exact-head canary reproduces the existing placeholder/non-substantive report => green Strix failure class on the current central protected workflow; this is central-owner evidence, not an Orgmetra source finding.

    Consumer: ContextualWisdomLab/Orgmetra#96

    • protected live base: develop@eb9757f8649aaad026a9865508d9aad50c1a7a4f
    • exact contributor head: fc3c5111c2a6a4b9bef15954c483126dcd771def
    • required Strix run/check: 34381793320 / 102568747529, GitHub conclusion SUCCESS
    • immutable strix-reports artifact: id 10116773992, archive digest sha256:86d5433424b7b76e58a847cbc961068a57f2b7ccb8b7dbfb84881aadbdafd18a
    • central gate materialized 4 scannable changed files and ran openai/orchestrator/free; Strix run.json records status=completed, scan_completed=true, success=true, 3 LLM requests, 129,736 input tokens, 510 output tokens, cost 0
    • SARIF is structurally valid but contains 0 results
    • the generated penetration report is only 307 bytes and its four required sections are literal generic template text: Business-level summary for leadership., Frameworks, scope, and approach., Consolidated findings + systemic themes., Prioritized, actionable remediation.
    • report/run target identity is only the ephemeral /tmp/strix-runtime.../strix-pr-scope... path; the published artifact contains no machine-readable receipt binding Orgmetra repository, PR, exact head, live base and materialized scope digest to the terminal finding set
    • scanner logs show the first finish_scan tool call failed and was retried; the second call completed and Strix reported zero vulnerabilities. That eventual tool success does not cure the non-substantive report/receipt problem.

    This is the same semantic acceptance defect already captured by the Orgmetra #55 placeholder canary in this issue, now reproduced on current Orgmetra #96 after correct exact-head scope materialization and a successful contextual-orchestrator route. Therefore #96 treats the green strix check as non-authorizing rather than as proof of zero exploitable vulnerabilities; no leaf retry, provider override, local workflow fork, synthetic status, or gate weakening is appropriate.

    RED acceptance to retain: exact scope + exit 0 + zero-finding SARIF, but template/non-substantive analysis or missing target/provenance receipt => required Strix evidence remains non-passing. GREEN remains: validate a machine-readable exact target/base/head/scope/provenance receipt and minimum substantive report semantics, cross-check the terminal finding set against that report, then replay an unchanged/current Orgmetra head through protected central truth.

  17. seonghobae commented on Sep 9, 2026

    @seonghobae
    ContributorAuthor

    Orgmetra #96에서 같은 exact-head Strix 증거 계약 결함이 다시 재현됐습니다. Leaf workaround가 아니라 #891의 중앙 acceptance fixture로 다뤄야 합니다.

    • consumer: ContextualWisdomLab/Orgmetra#96
    • protected live base: develop@eb9757f8649aaad026a9865508d9aad50c1a7a4f
    • exact head: ff16a08dc035faf05dc1b921e947b1cf82cfb36e
    • Required Strix: run 34396347513, job 102617064104
    • GitHub job: terminal SUCCESS; trusted target materialization, CO sidecar, Strix install, orchestrator/free input, Run Strix (quick), report collection and artifact upload all completed successfully
    • immutable artifact: strix-reports id 10122149854, digest sha256:91b165b34cf2ecf645cfb80bdefbaf70ad4238095494894c5f8c04ae3fe0c550

    Artifact 검증 결과는 merge-authorizing zero-finding evidence가 아닙니다. findings.sarif는 Strix 1.5.3의 빈 results만 담고 있지만, 307-byte penetration_test_report.md는 Executive Summary/Methodology/Technical Analysis/Recommendations 네 섹션이 각각 Business-level summary for leadership., Frameworks, scope, and approach., Consolidated findings + systemic themes., Prioritized, actionable remediation.라는 일반 템플릿 문장뿐입니다. run.json도 status=completed, scan_completed=true, success=true를 기록하면서 동일한 placeholder 문구를 terminal analysis로 저장합니다.

    또한 artifact 안에는 repository/PR/exact head/live base/materialized-scope digest를 terminal finding set과 결속하는 machine-readable receipt가 없습니다. 따라서 빈 SARIF + workflow SUCCESS를 취약점 없음으로 승격할 수 없습니다.

    #891 acceptance에 이 case를 포함해 주세요. Exact target/base/head/scope/provenance receipt와 substantive report semantics가 모두 검증돼야 zero-finding SUCCESS가 passing이 되며, template-only/no-receipt evidence는 fail closed해야 합니다. 중앙 repair가 protected .github/main에 정상 통합된 뒤 unchanged Orgmetra#96@ff16a08...를 다시 검증하는 것이 leaf의 acceptance boundary입니다.

  18. seonghobae commented on Sep 10, 2026

    @seonghobae
    ContributorAuthor

    Fresh Orgmetra consumer canary after #289 source/test repair:

    ContextualWisdomLab/Orgmetra#96@6573a6293db49632c525f0979973377869b4817c Strix run 34427956795, job 102719245996, is now terminal SUCCESS. Run Strix (quick) ran from 02:48:18Z through 04:28:26Z; report collection and artifact upload then succeeded. Artifact 10136631500, digest sha256:ee4f6d58e682783cd388a32c9d7fb6af79c05e1613e4d5622118a248eac3434e, is 22,837 bytes and is materially better than the prior 307-byte template-only Orgmetra canaries.

    Direct immutable-artifact inspection shows:

    • penetration_test_report.md is 2,645 bytes and contains package-specific methodology/analysis of OrganizationHierarchyChangeReviewPacket, exact-type validation, closure-owned issuance state, digest/export integrity, requester/reviewer separation and PII minimization. This is no longer a hollow generic four-section template.
    • findings.sarif is well-formed SARIF 2.1.0 with tool.driver.name="Strix", version 1.5.3, and zero results.
    • run.json records a completed successful scan, 25 LLM requests / 589,978 total tokens, local-code target path, scope_mode="auto", diff_scope.active=false, and diff_base=null.
    • The gate console says merge-base computation failed and it fell back to direct base/head diff, materializing 4 scannable changed files for findings attribution.

    The remaining #891 evidence defect is narrower but still material: the uploaded terminal evidence does not contain a machine-readable receipt binding repository + PR number + exact source head + exact live base + the four materialized file identities/content digests + scanner/run identity + report/SARIF digests/finding set. GitHub artifact metadata externally binds this archive to head 6573a629..., but it does not bind the relevant live-base context or materialized-scope digest; run.json only names an ephemeral target path, and the human report narrows its stated scope to the package directory without enumerating the four gate-materialized files.

    Under #891's existing acceptance contract, this should be treated as substantive scan output but still non-authorizing exact-head/live-base evidence, not as proof that the required gate's provenance contract is complete. Please use this current canary in the canonical #891/#1563 repair: terminal success should emit/validate an exact provenance receipt and cross-check it against the SARIF/report finding set before the required gate can authorize merge. Do not regress the newly substantive report semantics while adding the receipt.

  19. seonghobae commented on Sep 10, 2026

    @seonghobae
    ContributorAuthor

    Fresh Orgmetra consumer canary on exact ContextualWisdomLab/Orgmetra#96@f2c6e70edcf5c62ae0f848c29efd7bd3b8d45b97 shows the original hollow-success class is still possible under protected central Strix source main@cb0872c9a20d5584703dffacca65c096fc034c6c.

    Required Strix run 34441832410, scan job 102760977777, completed SUCCESS. Artifact 10139235414 (strix-reports, 14,754 bytes, digest sha256:c56b080f7357b5e9187f2d5d55396f61bbaa223e0e1cd27ea5acfe0fccde8c3e) is bound by GitHub metadata to that exact PR head. The scan itself used strix-agent==1.5.3, orchestrator/free, 7 LLM requests, 129,938 input tokens and only 630 output tokens.

    However the terminal evidence is not authoritative no-finding evidence:

    • penetration_test_report.md is only 437 bytes and is the generic template: No vulnerabilities were found, OWASP WSTG/gray-box boilerplate, No technical analysis is required, and General hardening recommendations.
    • findings.sarif is structurally valid SARIF 2.1.0 with tool.driver.name=Strix, version=1.5.3, and results=[], but it carries no target/head/base/materialized-scope binding.
    • run.json says status=completed, scan_completed=true, success=true, yet records only an ephemeral /tmp/.../pr-scopes/... target, diff_scope.active=false, diff_base=null, and no repository/PR/exact-head/live-base/content-digest terminal receipt.
    • The workflow therefore exits GREEN even though the only narrative evidence is the same template-only form that [Security] Fail closed when Strix produces no authoritative scan evidence #891/fix(strix): require authoritative report artifacts on success #1563 are intended to reject.

    This is especially useful because predecessor Orgmetra #96 exact 6573a629... previously produced a substantive long-form Strix report, while this unchanged product lane/current exact head produced the generic 437-byte template. That demonstrates report prose quality is nondeterministic and cannot itself be the no-finding authority.

    Required central repair remains: terminal gate must bind repository + PR + exact head + relevant live base + materialized scope/content digest + scanner/version + attempt/run identity + report/SARIF digests + finding-set count/digest, and template-only/semantically hollow success must be non-passing even when run.json.success=true and SARIF is empty. Preserve orchestrator/free; do not solve this with paid/model pinning, elapsed-time cutoffs, leaf reruns, or synthetic statuses.

  20. seonghobae commented on Sep 10, 2026

    @seonghobae
    ContributorAuthor

    Fresh Orgmetra protected-consumer persistence canary for the Strix/Caido sandbox failure class.

    Consumer: ContextualWisdomLab/Orgmetra#295@63a1103135f7f98a55cf284e6e2ff706c60deea9, direct protected base develop@eb9757f8649aaad026a9865508d9aad50c1a7a4f.

    Required Strix run 34506401099, job 102969699765 is terminal FAILURE. Exact-head admission, trusted Strix source checkout, target workspace materialization, required-workflow self-test, secret gate, and contextual-orchestrator sidecar provisioning all succeeded. The downloaded artifact 10164581579 (strix-reports, digest sha256:e674803d9bdd0a44a0ab67dac5ac73f1677c693753c0cb944e7db7de459126d9) isolates the failure before model inference:

    • contextual-orchestrator preflight reached gateway-ready state (candidate_count=24, probe_budget=16, ready_count=2, target 8);
    • Strix 1.5.3 / sandbox image ghcr.io/usestrix/strix-sandbox:1.3.0 created two distinct sandbox containers across the bounded same-model sandbox retry;
    • in each attempt, loginAsGuest could not connect to the sandbox-internal Caido endpoint 127.0.0.1:48080 for all 10 bootstrap attempts;
    • both attempt run.json records are failed, both report llm_usage.requests=0 / zero tokens, and the empty SARIF results therefore are not vulnerability-scan success evidence;
    • the gate correctly emitted STRIX_SANDBOX_UNAVAILABLE after the one sandbox-specific retry and remained non-passing.

    This matches the previously observed Naruon/central loginAsGuest/48080 signature, so it is not a new CO routing defect and not an Orgmetra source vulnerability. It is a fresh source-change consumer confirming that the sandbox runtime/lifecycle defect persists on the current central protected workflow generation.

    RCA refinement requested on the canonical owner path: before increasing retry counts or sleeps, capture and classify the sandbox container's startup logs, exit state, health/readiness, and actual listener state for port 48080 (including the distinction between Docker's allocated host port and the service's internal listener). Add a deterministic RED fixture for "container created but Caido never becomes reachable" and prove GREEN with authoritative exact-head scan evidence. Then integrate through normal protected main and replay this unchanged Orgmetra head. Do not neutral-success this class, mutate the leaf merely to retrigger, transfer predecessor evidence, or add leaf-local retries.

  21. seonghobae commented on Sep 10, 2026

    @seonghobae
    ContributorAuthor

    Fresh exact-head Strix canary from ContextualWisdomLab/Orgmetra#295@d15a741c10685208b124477e9d7820cdf17339e8, base develop@eb9757f8649aaad026a9865508d9aad50c1a7a4f.

    Run 34515071750, scan job 102998800862, is terminal SUCCESS and the workflow's pre-scan path says it materialized the current PR head and retained 3 scannable changed files. Artifact 10168136999 (strix-reports, sha256 9093fb1aa46882c60ab101708c401fed939a6c6b7ebc53ce942309d047b5b34c) is bound to that GitHub run/head. However, the artifact's own run.json reports diff_scope.active=false and diff_base=null; its methodology/scope is generic https://app.acme.example / /api/v1/, not the materialized Orgmetra PR source scope. It records only 2 LLM requests. penetration_test_report.md is a short generic no-finding report, while SARIF is structurally valid and empty.

    Therefore the successful job is not admissible authoritative no-finding evidence for the stated exact-head source scan: the execution receipt does not bind the materialized repository/PR/base/diff scope to the report/SARIF/finding set, and the report describes a different generic target. This is the hollow-success class #891 is meant to fail closed on. Please require terminal evidence to bind repository, PR, exact head, live base, materialized scan scope, report/SARIF and finding set before a Strix success can satisfy the security gate. No leaf synthetic status or rerun workaround should be used.

  22. seonghobae commented on Sep 11, 2026

    @seonghobae
    ContributorAuthor

    Fresh protected-consumer canary from ContextualWisdomLab/Orgmetra#295 exact head c7d39a6702ed2149c79dc2aa7c2c4be983679f9c confirms the remaining evidence-integrity defect after the Strix job itself succeeds.

    • Required Strix run: 34549296455; job 103110710471 completed SUCCESS at 2026-09-11T02:14:44Z.
    • Artifact: strix-reports id 10181773992, digest sha256:3865ee8df2cd4f9d9819d046290f808ef56c7fde17a7a888ba90cedf5afd5e74, bound by GitHub artifact metadata to the exact head above.
    • run.json: target is a local PR-scope directory, but diff_scope.active=false, diff_base=null, and the receipt contains no repository/PR/live-base/exact-head/materialized-file-set binding.
    • penetration_test_report.md is generic prose (report complete, generated using the Orgmetra Keyverse adapter) and does not enumerate scanned files, base/head, test actions, or a concrete no-finding basis.
    • SARIF contains 0 results but no invocation/provenance record tying that no-finding set to the exact PR materialization.
    • The run consumed openai/orchestrator/free as required and did complete, so this is not provider unavailability; it is a successful check whose published receipt is still insufficient to prove what exact source was scanned.

    Please keep the gate fail-closed until the authoritative receipt cryptographically/logically binds at least repository, PR, exact head, live base, materialized scope/file inventory, scan mode/tool version, report/SARIF, and resulting finding set. Orgmetra will not promote this SUCCESS to admissible security evidence and will not patch the leaf workflow around the canonical owner.

  23. seonghobae commented on Sep 11, 2026

    @seonghobae
    ContributorAuthor

    Fresh exact-head fail-closed Strix infrastructure canary from ContextualWisdomLab/fast-mlsirm#1816@432765ccf633c9802e0f796ceeb4d6d572059acf.

    Required Strix run/job: 34606293205 / 103285730814.

    The job passed exact-head admission, checkout, contextual-orchestrator sidecar/model preparation, and sandbox launch, but failed before any authoritative Strix scan could start. Wait for Strix sandbox bootstrap repeatedly attempted POST http://127.0.0.1:48080/auth/loginAsGuest and terminated with:

    STRIX_SANDBOX_BACKEND_CONTROL_FAILED: Caido guest bootstrap failed with status 000

    The Caido-side log also records an asyncio BrokenPipeError; cleanup ran afterward. There is therefore no exact-head security finding/no-finding receipt to promote, and the current required failure is correct fail-closed behavior rather than a fast-mlsirm source defect.

    Please retain this canary under #891's authoritative-evidence contract: sandbox/tool bootstrap unavailability must remain non-passing, but the terminal receipt should identify the bootstrap/backend cause distinctly from provider exhaustion or a product vulnerability. The leaf will not bypass Strix, fabricate a clean status, copy the central sandbox workflow, or weaken the required context.

  24. seonghobae commented on Sep 12, 2026

    @seonghobae
    ContributorAuthor

    Current fail-closed sandbox outage specimen: ContextualWisdomLab/fast-mlsirm#1829@3be36281c80ca5a9d9275c7def69b5cdd9928429, run 34700510941, job 103571721118.

    Strix 1.5.3 with sandbox image tag ghcr.io/usestrix/strix-sandbox:1.3.0 failed both workflow attempts during caido_bootstrap.py::_login_as_guest: curl exit 7, connection refused at 127.0.0.1:48080. Each bootstrap reports 10 login attempts; workflow attempt durations were 179s and 84s. At 2026-09-12T15:09:32Z the terminal classification is STRIX_SANDBOX_UNAVAILABLE, exit 1. This is not an LLM gateway error or a completed vulnerability scan; no security pass is claimed. The captured log shows the image tag, not its resolved digest.

    This confirms the fail-closed outage case, but not recovery. The next owner action is to capture sandbox container state, resolved image digest and Caido startup logs; reproduce the bootstrap failure with the pinned runner/tool configuration; then repair the actual startup/compatibility cause and validate a successful exact-head scan. Closed #1185 classified this condition; classification alone does not repair it. No extra model fallback, suppression or consumer workflow copy is proposed.

    Investigation checkpoint: 2026-09-13 UTC, next review 2026-09-14 UTC, because this currently prevents authoritative security evidence for an active consumer PR. Those are internal investigation dates, not promised resolution dates. Full consumer job logs are available at the linked run/job; source changes and an owner repair remain outstanding.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: apiAPI, protocol, event, or external contractarea: authAuthentication, authorization, identity, or tenant isolationarea: ci-cdCI, GitHub Actions, checks, release, or supply chainarea: securitySecurity boundary, hardening, or vulnerability preventionbugSomething isn't workingpriority: mediumNormal-priority or P2 workstatus: triagedOpen issue has an organization taxonomy assignmenttype: bugDefect or incorrect behaviortype: featureNew or expanded product capability

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions