Skip to content

fix: seven newcomer passes — a false-positive audit event, a hardcoded isolation claim, and eleven other defects found by installing this from scratch - #23

Merged
convee merged 10 commits into
mainfrom
bugfix/20260905-onboarding/main
Sep 6, 2026
Merged

convee merged 10 commits into
mainfrom
bugfix/20260905-onboarding/main

Conversation

@convee

@convee convee commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Seven newcomer passes over this repository — clone, read the docs, install, deploy, use it — fixing what each pass turned up and re-running from a fresh clone until a pass found nothing. Then a benchmark run and an isolation-backend assessment, which turned up three more.

Every fix has a test, and every new assertion was mutation-checked: the old behaviour was put back and the test had to go red. Where a claim could be settled by measurement rather than reasoning, it was.

Onboarding, as measured

From a fresh clone on a 4-core Linux host: make bootstrap 13s, make verify 122s, 0 manual interventions, 870 tests. The source-level experience is genuinely good. The wall is the cluster: make up-local needs Lima + /dev/kvm + 10.5 GiB free memory + 35 GiB free disk, and make doctor stops honestly before wasting anyone's time. There is no lighter shipped profile, so a machine without nested virtualisation cannot run the product at all.

For the record, overlays/rwo-single-node does work on a single-node kind cluster with three adjustments — empty SANDBOX_RUNTIME_CLASS, the sandbox.hullwork.com/node-role=runtime label, and a restricted-PSS-compliant PostgreSQL. Everything below was verified against that deployment with a real PostgreSQL and a real S3 (MinIO).

Correctness

A denial audit event that could only ever be a false positive. /v1/workspaces/resolve derives its Workspace ID from the caller's own tenant (ws:v2:<tenant>:…), so ownership can never fail there for the reason the audit exists to record — yet a miss filed workspace.access / denied, field-for-field identical to a real cross-tenant probe. One per Workspace ever created, burying the signal. Live evidence after the fix: five Workspaces created, one denied row — the genuine cross-tenant read.

sandbox_status asserted gVisor isolation from a constant. mcp.py merged the literal "runtime": "gvisor" into the agent-facing status. On a deployment with SANDBOX_RUNTIME_CLASS empty the Control Plane reports cluster-default and the Console renders it, while the one surface an agent uses to reason about its own confinement claimed isolation regardless. The lease already carries runtime_class; the SDK was discarding it.

Checkpoint restore rejected the absent body its contract declares optional. The OpenAPI says requestBody: required: false; the handler called read_json(), so no body answered 400 Content-Length is required and Content-Length: 0 answered 400 request body must be valid JSON. A missing checkpoint was reported as 400 rather than 404 for the same reason. Chunked bodies are refused explicitly rather than read as absent — treating an unreadable body as "no sha256 requested" would skip the integrity check the caller asked for.

The E2E announced isolation it had skipped. scripts/test.sh asserts the gVisor self-report only when SANDBOX_RUNTIME_CLASS is gvisor — correctly, because on the cluster default runtime that dmesg line belongs to the host kernel. The closing line said + gVisor unconditionally, so a run against an empty runtimeClass, which the AI-LOCK in core.py requires on nodes without runsc, announced a property nothing checked.

A cross-component timing invariant neither side knew about. ACTIVITY_PROBE_TIMEOUT is 2s and a timeout means "deletable"; shell_sessions.py records a measured 2.60s for that read while reaping a SIGKILLed child, with the comment "The 2 second probe line has been crossed". The lock discipline that holds it together is guarded by a real behavioural test, but that test's PROBE_BUDGET was a constant local to runtime/ and tied to nothing. Both drift directions failed silently: lower the timeout and the guarded bar stops being the real bar; raise it and the test's induced slowness no longer trips the probe, so passing proves nothing.

A bare KeyError where the codebase's own rule says otherwise. core.py states it: "SystemExit instead of KeyError: what the container wants is an instruction to follow, not a stack trace." KubeClient.__init__ raised KeyError: 'KUBERNETES_SERVICE_HOST' outside a cluster, and FileNotFoundError when automountServiceAccountToken was off.

Evidence chain

Benchmark environment captures were silently truncated into invalid JSON. output[:20_000] applied to every captured command with no marker; a kubectl get pods -o json measured 71,099 bytes on a real run. Anyone parsing environment.json to confirm what a report measured got a decode error.

And it described the cluster, not the run. runtimeClass was captured as the cluster's gvisor RuntimeClass object, snapshotted before the first Runtime exists. A report produced on a cluster with no gVisor read as a gVisor measurement. Each lease already reports the class it got; summary.json now carries observedRuntimeClasses, with unreported rather than a guess when a Control Plane is too old to send it.

Documentation that contradicted the code

docs/ARCHITECTURE.md said the Runtime provider "is selected independently by SANDBOX_RUNTIME_DRIVER" while configured_runtime_driver() never reads it — advertising exactly the plug-in surface README.md promises does not exist. The guard added for it is conditional: it reads the selector's presence out of the AST, so the prose may not promise a selector today and must mention one the day a registry lands.

docs/TROUBLESHOOTING.md told a reader debugging Runtime placement to run kubectl get nodes --show-labels | grep sandbox-node. That string appears nowhere else in the repository; the label is sandbox.hullwork.com/node-role. The command returned nothing on a correctly labelled cluster and nothing on an unlabelled one — no discriminating power, in the one place a reader is already lost.

docs/DEPLOYMENT.md said all four cluster dependencies are configured through values.yaml; the gVisor RuntimeClass has no value at all, and the chart pins both isolation variables as literals in a block where everything else templates. The same file's operator recipe never mentioned that both namespaces enforce PodSecurity restricted, so a stock PostgreSQL is rejected by the ReplicaSet controller while kubectl apply looks like it worked. And installing the chart in the documented order is impossible: it creates both namespaces and no Secrets, so creating the namespaces first to hold the Secrets makes helm install refuse them.

The object namespace — which path roots each scope accepts, which locator fields each requires, and the opposite constraints on workspace import and export — existed only in control_plane/core.py. It took six 400s to assemble one working call. The errors name the allowed set, which makes the API usable by trial; the table makes it usable by reading.

README.md's prerequisites omitted qemu-system-<arch> and shasum, which make doctor hard-requires on Linux — so the first documented command failed on a machine that had followed the documentation. The test suite already knew both, with a comment explaining why.

CLI and Console

run spells the Workspace name --name while create, exec and stop take it positionally, and no argument carried help text, unlike sandboxctl. sandbox stop --name demo — the natural next command after the README — was rejected with the top-level usage.

The Console rendered 1 minutes ago. The relative-time messages now use the plural mechanism the catalogs already had, with the compound form assembled from two independently pluralised halves, because in 5 hours 1 minutes is what a single template cannot avoid.

Ratchets

Nine of these could recur silently, so each is now pinned by a test that fails when it does: the README prerequisites against what dev-doctor.sh actually requires; the troubleshooting selector against the deployed label; the Helm install order against every workload that mounts a pre-provisioned Secret; the object namespace against the allowed_roots literals, read by AST; the architecture prose against the selector's presence, also by AST; the eviction bar against the probe timeout that enforces it; and the E2E verdict against the branch lifted out of scripts/test.sh rather than restated in the test.

Reported, not changed

The Console offers "Suspend" and "View keys" on the reserved management tenant. The backend refuses both with 403, so nothing breaks, but an operator is asked to confirm a destructive-sounding action that cannot succeed. Fixing it means either duplicating a backend constant in the frontend — the drift that produced the sandbox-node bug — or adding a reserved field to the public Tenant schema. That is a maintainer's call.

Not covered

gVisor isolation itself, OIDC, the Lima profile and scale-workers, the 1,000-file collection export, 30-minute steady load, and Control Plane restart recovery. The host has no /dev/kvm.

🤖 Generated with Claude Code

https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh

convee and others added 10 commits September 5, 2026 18:11
…, and two CLI/Console defects

## 5. 修复过程
问题现象:以新人身份从零 clone 走安装-部署-试用全流程,发现六处问题:
  1. 每次租户首次使用一个新 Workspace,审计表都会落一条
     `workspace.access / denied`,与真正的跨租户越权尝试逐字段同形;
  2. `sandbox stop --name demo`(README 教了 `run --name`)被拒,且 CLI
     所有参数无任何 help 文本;
  3. 控制台英文渲染 `1 minutes ago` / `in 1 minutes`;
  4. 控制面在集群外启动只抛 `KeyError: 'KUBERNETES_SERVICE_HOST'`;
  5. README 前置表未列 Linux 上 `make doctor` 硬要求的 qemu 与 shasum,
     文档第一条命令在合规机器上即失败;
  6. DEPLOYMENT.md 说 chart 的四项集群依赖都由 values.yaml 配置,
     而 gVisor RuntimeClass 没有对应 value;同文档让运营者自带
     PostgreSQL,却未提两个 namespace 都是 enforce=restricted。

根因分析:
  1. `/v1/workspaces/resolve` 的 workspace_id 由调用方自己的 tenant_id
     经 HMAC 派生(`ws:v2:<tenant>:...`),结构上不可能属于别的租户,
     该路径上的所有权失败只可能是"这个 session 还没有 Workspace"。
     它却复用了 by-id 路由的审计分支,于是每建一个 Workspace 就伪造一条
     入侵信号,把这张表唯一要暴露的东西淹掉。
  2. `run` 的位置参数是待执行命令,只能把名字写成 `--name`;其余三个
     子命令用位置参数。差异真实存在,但既无 help 也未进文档。
  4. `KubeClient.__init__` 直接 `os.environ[...]`,与 core.py:655 自己写下的
     "给容器一条能照做的指令,不是堆栈"相悖。

修复方案:
  1. `require_workspace_tenant` 增加 `audit_denial` 形参,仅 resolve 站点
     传 False;所有权校验本身照旧执行。by-id 路由的审计不变。
  2. 补齐全部参数 help(与 sandboxctl 一致),USAGE.md 增补 `sandbox stop`
     示例并说明为何只有 run 用 `--name`。
  3. 按仓库既有的 Intl.PluralRules 复数机制重构 relative.* 文案:
     单位各自成 one/other 家族,复合形式由两个已变复数的片段拼装,
     消除 "in 5 hours 1 minutes" 这类单一模板无法回避的错误。
     中英渲染结果除 zh 复合形式由「分」统一为「分钟」外保持不变。
  4. 两处失败改抛 SystemExit 并说明该怎么做。
  5. README 补两项,并新增闸门:README 前置表必须覆盖 dev-doctor.sh
     解析出的每个命令(变异验证:还原旧 README 即报出 qemu/shasum 两项)。
  6. 按事实改写两处文档。

验证:make verify 全绿(848 用例,原 844 + 4);四条新断言逐条做过变异
验证;kind 单节点真机复验:resolve 404 后审计计数不变(10→10),
而 beta 跨租户读 alpha Workspace 仍照常记账。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…Helm install order

## 5. 修复过程
问题现象:以新人身份第二轮验证(从修复后的树全新 clone)发现两处:
  1. TROUBLESHOOTING「Runtime Pod stays Pending」给的检查命令是
     `kubectl get nodes --show-labels | grep sandbox-node`,在正确打了标签的
     集群上和完全没打标签的集群上**都返回空**;
  2. 按 DEPLOYMENT 的顺序装 Helm chart 装不上:chart 自己创建两个
     namespace,又不创建任何 Secret,而 Secret 必须先于工作负载存在。
     先建 namespace 以便放 Secret,`helm install` 就报
     `invalid ownership metadata`。

根因分析:
  1. `sandbox-node` 这个串在整个仓库里只出现在这一行,不是任何地方写入的
     标签键(真实键是 `sandbox.hullwork.com/node-role`)。读者正在排查
     调度失败时,这条命令恒定答"标签不存在"。
     test_label_domain 的扫描范围有意排除 docs/,所以没人挡住它。
  2. 文档的 Pre-provisioned resources 一节是按 kustomize(`kubectl apply`)
     写的,Helm 一节只说"直接 helm install",两者的先后顺序没有交代。

修复方案:
  1. 改成 `kubectl get nodes -l <真实标签>`,空列表本身就是结论;
     新增闸门:该小节必须按 k8s/control-plane.yaml 里
     SANDBOX_RUNTIME_NODE_SELECTOR 的实际值来选(变异验证:还原旧命令即红)。
  2. 新增「Install order」小节,给出实测走通的顺序(install → 建 Secret →
     重启三个挂载它们的工作负载),并说明 GitOps 侧不需要重启这步;
     新增闸门:渲染结果里凡挂载这四个 Secret 的工作负载都必须在该小节被
     点名(变异验证:从文档里去掉 sandbox-volume 即红)。

顺带核实、确认**不是**缺陷的:chart 的 workspace PVC 是 RWX,kind 的
local-path 拒绝——文档已明确要求 RWX StorageClass;控制面启动时那条
WORKSPACE_IDLE_TTL_SECONDS 告警是 core.py 里写明的有意设计,base 与 chart
默认值一致。

验证:make verify 全绿(850 用例);两条新断言各做过变异验证;
kind 上实测走通 helm install 的正确顺序,并确认错误顺序会复现原始报错。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…e under

## 5. 修复过程
问题现象:第三轮以新人身份试用 MCP bridge 时,在一个**没有 gVisor** 的集群
(SANDBOX_RUNTIME_CLASS 置空,Runtime 跑在集群默认运行时)上调用
`sandbox_status`,返回 `"runtime": "gvisor"`。同一时刻控制面对同一个
sandbox 回的是 `runtime_class: "cluster-default"`,控制台也照实显示
`Isolation: cluster-default`。

根因分析:`sandbox_platform/mcp.py` 把 `"runtime": "gvisor"` 作为**字面量**
合进结果,与部署实际配置无关。这是 agent 用来判断自身被如何隔离的那个面,
而仓库在控制面一侧本就有明确不变量守着同一件事
(`test_empty_runtime_class_reports_cluster_default_not_gvisor_isolation`),
README 也写着「runtime-driver 抽象拒绝假装」。SDK 其实在建箱应答
(SandboxLease.runtime_class,契约里已声明)里收到了真值,只是丢掉了。

修复方案:Lease 记录控制面报告的 runtime_class(建箱与恢复两条路径写入,
runtime 被回收/释放的两处清空),`SandboxManager.status()` 增加同名字段,
MCP 直接透传,不再注入常量。工具描述改为说明该字段的三种取值
(gvisor / cluster-default / null)。

surface 变化:`sandbox_status` 结果中的 `runtime` 键被 `runtime_class` 取代。
键名与控制面、契约、控制台一致;旧键的值在半数部署上是错的,保留它没有意义。

验证:make verify 全绿(851 用例);新断言做过变异验证(把常量塞回去即红);
kind 真机复验:MCP 报 cluster-default,与 `GET /v1/sandboxes` 对同一
sandbox id 的回答逐字一致。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…al on checkpoint restore

## 5. 修复过程
问题现象:第五轮把对象存储真正接上(kind 里起 MinIO,用仓库自带的
bucket-init Job 建桶)后,checkpoint 全路径首次可测。按契约不带 body 调用
`POST /v1/workspaces/{id}/checkpoints/{id}/restore` 返回
`400 Content-Length is required`;带 `Content-Length: 0` 返回
`400 request body must be valid JSON`;只有 `-d '{}'` 才走得通。
而 OpenAPI 对该路由写的是 `requestBody: required: false`,且该 body 只承载
一个可选的 `sha256`。副作用:checkpoint 不存在时也被这条挡在前面,报的是
400 而不是 404。

根因分析:处理器无条件调 `read_json()`,它在缺 Content-Length 时即抛错。
兄弟路由 `POST /checkpoints` 不读 body,所以同一资源上的两个动作行为不一致。
全契约扫描确认只有这一个路由声明了可选 body。

修复方案:新增 `read_optional_json()`,仅此站点使用。
🔴 **不能把「读不到 body」当成「没要求 sha256」**:chunked 编码同样没有
Content-Length,而本服务从不解码 chunked——那样会在调用方明确要求完整性
校验时静默跳过它。因此显式拒绝 Transfer-Encoding,只有真正缺席才当空。

验证:新增 6 条断言,两个方向各做变异(改回 read_json、去掉 chunked 守卫,
都能转红);make verify 全绿(857 用例)。kind 真机复验六种情况:
无 body/CL:0/`{}` 均 200,错 sha256 与坏 body 仍 400,缺失的 checkpoint
现在正确报 404。

本轮同时核实、**未发现问题**的面:checkpoint 建/列/恢复往返(恢复是替换
语义,与 LIFECYCLE_AND_DATA.md 一致)、sha256 完整性闸门、checkpoint 的
跨租户 404、对象 put/stat/list 的 owner 分区(同租户异 subject、异租户同
subject 均读不到)、`/healthz` 的 object_storage 由 unchecked 转 ok。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…onment captures

## 5. 修复过程
问题现象:按 docs/BENCHMARKS.md 对 kind 部署跑基准(100 轮 0 失败)后核对证据,
`environment.json` 有两处让报告说不清自己测的是什么:
  1. `systemPods` / `workloadPods` 是被 `output[:20_000]` 截断的
     `kubectl get pods -o json`。本次实测 71,099 与 47,045 字节,
     也就是**每次运行**这两块都是从对象中间截断的非法 JSON,且不留任何标记——
     解析的人只拿到 JSONDecodeError,看不出是被截断还是命令半路失败。
  2. `runtimeClass` 记的是集群里那个 `gvisor` RuntimeClass 对象,而快照在
     run 开始**之前**采集,此时被测 Runtime 一个都还没建。于是一份在
     `SANDBOX_RUNTIME_CLASS` 为空(Pod 跑在集群默认运行时)的部署上跑出来的
     报告,读起来像 gVisor 测量——而发布报告的头条数字正是「gVisor 冷启动」。

根因分析:runner 的 docstring 写着「summary 与 environment 只从观测派生」,
但隔离这一项是从集群配置这个**代理量**推出来的,不是从发生的事。而
`runtime_cold_start` 每次拿到的 SandboxLease **本来就带** `runtime_class`
(契约里早有该字段),收到后被丢弃。

修复方案:
  1. 超过上限时记 `truncated` 与 `fullLength`;上限本身不动。
  2. 每次冷启动把 lease 的 `runtime_class` 收进集合,summary 增
     `observedRuntimeClasses`。缺字段记 `unreported` 而不是猜——老版本控制面
     不发这个字段,那既不是 gVisor 的证据也不是反证。
  3. BENCHMARKS.md 说明这两个字段,并点明 environment.json 描述的是
     「集群提供什么」而非「这些数字跑在什么上面」。

验证:新增 3 条断言,两处各做变异(去掉 lease 记录、去掉截断标记,都能转红);
make verify 全绿(860 用例);打过补丁的 runner 真机跑 20 轮,
summary 出 `observedRuntimeClasses: ["cluster-default"]`,
environment 标出两处 71,099→20,000 与 47,045→20,000 的截断。

本次基准读数(kind 单节点 / 无 gVisor / 4 核 EC2 且经 port-forward,
**不可与发布报告的数字合并**):100 轮 0 失败,成功率 100%;
冷启动 p50 3.12s / p95 3.63s,warm exec p50 141ms / p95 194ms,
1KiB 读 p50 132ms、写 p50 148ms;excellent 未达标三项:
health.p95 100.6ms(阈值 100ms)、workspace_create.p95 172ms 与 p99 279ms。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
## 4. 功能描述
功能概要:docs/API.md 增「Object namespace」一节,说明对象定位器两个 scope
各自必需的字段、path 首段允许的根、生成的 key 形状,以及 workspace
import/export 两条路由方向相反的收窄规则。
用户可感知变化:调用对象存储不必再靠试错拼参数。

背景:第六轮以新人身份走 objects 全路径时,我连撞五次 400 才拼出一条能走通
的调用——依次是 `agent_id must be a lowercase DNS-style identifier`、
`object path must start with one of: artifacts, inputs, logs, meta, outputs`、
`workspace path must start with artifacts/`、`only upload objects can be
imported`、`object path must start with one of: derived, meta, source`、
`upload destination must start with data/uploads/`。每条报错都点名了允许集合,
所以这套 API **试得出来**;但这些前缀规则在 docs/ 与 contracts/ 里零命中,
只存在于 `control_plane/core.py` 的 `object_location`,也就是**读不出来**。

同时新增闸门 `tests/test_object_namespace_doc.py`:用 AST 取出
`object_location` 里两处 `allowed_roots={...}` 字面量,要求 docs/API.md 逐个
列出;并从 api.py 正则取出 import 目标前缀,要求文档写的是同一个值。
变异验证两个方向:代码加一个根而文档不加 → 红;文档删一个根 → 红。

本轮同时核实、**未发现问题**:workspace export(`artifacts/` 源)与 upload
import(`data/uploads/` 目标)双向往返字节完全一致;export 返回完整对象记录
(嵌在 `object` 键下)。

验证:make verify 全绿(862 用例)。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…orces it

## 5. 修复过程
问题现象:调研「架构能否换用 Kata/Firecracker 等其它虚拟化」时发现一条
**跨组件的时序不变量**,两端互不知道对方:
  · `control_plane/core.py:501` `ACTIVITY_PROBE_TIMEOUT` 默认 2 秒,
    `:3854-3857` 超时即 `return False` —— 判定「不忙」,也就是可删;
  · `runtime/shell_sessions.py:772-779` 记着实测(2026-08-19):Runtime 在回收
    被 SIGKILL 的子进程时,`activity_snapshot` 要 2.60s,注释原文
    「The 2 second probe line has been crossed」,后果是「一个正在工作的
    Runtime 连同它所有会话被一起删掉」。

守住它的是一条编码纪律(绝不在持锁时 `close()`),
`tests/test_shell_sessions.py:632` 已有一条**真行为测试**守着这一半
(注入 2.0s 慢回收 + 断言 `activity_snapshot` < `PROBE_BUDGET=1.0`)。

根因分析:`PROBE_BUDGET` 是 runtime 侧测试文件里的本地常量,
`ACTIVITY_PROBE_TIMEOUT` 在 control_plane 里,两者**零关联**(各自只在自己
文件中出现)。于是有两个静默失效方向:把控制面超时调到 1.0 以下,既有测试
守的那条线就不再是真线;把它调到 2.0 以上,`SLOW_REAP` 就不足以触发真探针,
那条测试**通过也证明不了任何东西**。两种情况全套测试都是绿的。
`docs/CONFIGURATION.md` 把这个旋钮列了出来,但没说调小它会删掉工作中的
Runtime。

修复方案:
  1. 新增 `ActivityProbeBudgetTests`(不依赖 bash,用正则从 core.py 源码取
     默认值,避免为一个数字去满足 control_plane 的导入环境):
     断言 `PROBE_BUDGET < 默认探针超时`,且 `SLOW_REAP >= 默认探针超时`。
  2. `docs/CONFIGURATION.md` 该行补上后果、那条 2.60s 实测,
     并点明**这条余量是隔离后端 syscall 成本的属性,不是这个设置的属性**,
     换 RuntimeClass 前必须重测。

验证:两个方向各做变异(探针改 0.8 / 改 5),都能转红;
make verify 全绿(864 用例)。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…t exist

## 5. 修复过程
问题现象:`docs/ARCHITECTURE.md:146` 写「the Runtime provider is selected
independently by `SANDBOX_RUNTIME_DRIVER`」,而同一仓库的 `README.md:88-92`
写的是「`SANDBOX_RUNTIME_DRIVER` accepts only `gvisor` ... **There is no
provider plug-in surface advertised that does not exist.**」——两句话直接打架,
而代码站在 README 这边。

根因分析:`configured_runtime_driver()`(`core.py:2573-2581`)不读这个变量,
无条件 `return GVisorRuntimeDriver(...)`;该变量唯一作用是 `core.py:247` 的
相等校验。也就是说 ARCHITECTURE 那句宣传的正是 README 声明「不存在」的那个
插件面。评估「能不能换 Kata/Firecracker」的人会先读架构文档,这句话会让他
以为只要改个环境变量。

修复方案:按事实改写该段(变量是启动期校验的闸门,不是选择器),并新增
`tests/test_runtime_provider_claims.py`。该闸门**是条件式的**:用 AST 检查
`configured_runtime_driver` 函数体内是否真的引用了 `SANDBOX_RUNTIME_DRIVER`,
没有选择器时禁止文档承诺选择器,将来选择器落地后同一条测试反过来要求文档
补上。变异验证两向:还原旧措辞 → 红;让函数体引用该变量而文档不改 → 也红。

验证:make verify 全绿(867 用例)。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…check

## 5. 修复过程
问题现象:`scripts/test.sh:381` 的收尾行无条件打印
`Sandbox E2E passed: Runtime MCP files/shell + shared Workspace + gVisor`,
而同一脚本 `:141` 的隔离断言是**有条件的**:
`assert sys.argv[2] != "gvisor" or "gVisor" in data["stdout"]`。
于是在 `SANDBOX_RUNTIME_CLASS=""` 下——这是文档支持的配置,
`control_plane/core.py:257` 的 AI-LOCK 明说节点没有 runsc 时必须置空——
断言被短路跳过,报告却照样宣布验过了 gVisor。

根因分析:`:138-139` 的注释把条件化的理由讲得很对(跑集群默认运行时时那行
dmesg 是宿主内核的启动日志,断言它必然误报),但收尾那行没有跟着同一个条件。
断言与结论各走各的,于是「跳过检查」被渲染成「检查通过」。

修复方案:收尾行按 `SANDBOX_RUNTIME_CLASS` 分支——是 gvisor 就说 gVisor,
否则报出实际 runtimeClass 并明写 `isolation not asserted`。
新增 `E2EVerdictTests` 三条断言:条件化的前提仍在、收尾行不是常量、
且三种取值下输出正确。
🔴 第三条**从 scripts/test.sh 里把该分支取出来跑**,不是在测试里抄一份——
抄一份的话改了脚本而没改拷贝,测试会继续验旧代码并照常绿。

变异验证两向:还原无条件那行 → 红;让分支对所有 class 都答 gVisor → 红。

背景:本条是评估「架构能否换用 Kata/Firecracker」时顺带发现的。
`:141` 的守卫写法本身在换后端时也会**静默丢覆盖**(断言整条跳过、测试照常
绿),所以那条守卫不宜当作模板照抄;正确做法是按 handler 分派,且没有对应
取证手法时显式失败。

验证:make verify 全绿(870 用例);`bash -n scripts/test.sh` 通过。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
…refuses

## 5. 修复过程
问题现象:PR #23 上四个 `images` job 全红,构建本身成功,红在第 5 步
「Scan built image with the release gate's parameters」。Trivy 报
`Total: 7 (HIGH: 7, CRITICAL: 0)`,全部落在 `libuuid`:
CVE-2026-53612 / 53613 / 53614(mount TOCTOU)、CVE-2026-76642、
CVE-2026-78408 / 78409 / 78410。main 昨天的 CI 四个 job 都是 success。

根因分析:与本分支无关。本分支不动任何 Dockerfile、`.lock`、`requirements`、
`package-lock`、`pyproject`,镜像内容与 main 的差异只有 .py/.ts 源码。
四个镜像的两个基础镜像(`python:3.14-alpine` 与
`nginxinc/nginx-unprivileged:1.31-alpine3.24`,均 Alpine 3.24.1)都装着
`libuuid-2.42.1-r0`,而这批公告是在 main 上次跑 CI 之后才披露的。
实测确认 Alpine v3.24 main 已发布修复版 `libuuid-2.42.3-r1`。

修复方案:沿用 `console/Dockerfile` 已有的先例(该文件本就为 libexpat 做了
同样的事,并写明「Remove once a newer base digest passes trivy on its own」),
对四个镜像加 `apk upgrade --no-cache libuuid`:console 那行扩成
`libexpat libuuid`,其余三个各加一条,注释同样点明这是临时措施、
以及它是 FROM 之后唯一一处访问包索引的地方。

验证:四个镜像本地重建后逐个查证 `libuuid-2.42.3-r1`(改动前均为 2.42.1-r0);
make verify 全绿(870 用例)。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015s137CobDdyvryGDawa5Rh
@convee
convee merged commit 332dcd0 into main Sep 6, 2026
12 checks passed
@convee
convee deleted the bugfix/20260905-onboarding/main branch September 6, 2026 02:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant