ci: gate the exec path on a real container, no KVM required - #205
Conversation
The exec e2e test has existed since the exec dispatch fix but only ran by hand
(`LANTERN_RUNTIME_E2E=1`), so the regression it guards was unprotected: for a
long time `grpcSchedulerClient.Exec` returned a canned string with exit code 0,
making a MISSING implementation indistinguishable from a command that ran and
printed nothing. Mocks cannot catch that — only a real workload can.
This runs on a stock GitHub runner. The runtime-manager's `docker` backend
spawns a real container and hosted runners have Docker, so no nested virt is
needed; that is why this can be an ordinary PR gate while the Firecracker boot
job still waits on a KVM runner. The job builds the manager and scheduler,
starts both, and runs the test against them.
No Postgres service is needed: with DATABASE_URL unset the scheduler uses its
in-memory store and runs always-leader, which is exactly a single-node test.
The gate asserts the test RAN. `go test` reports a skipped test as a pass, so
without this the job would go green while proving nothing — the same failure
mode as the boot assertion that reported PASS for kernel-panicking VMs. The
step fails on any `--- SKIP` and requires the PASS line for the real-VM test.
Verified both directions locally before pushing:
- env unset -> test SKIPs -> gate correctly FAILS
- env set -> 2/2 subtests PASS against a live manager + scheduler -> gate
passes, and the test's own cleanup left no containers behind
Scheduler build command and workflow YAML both checked.
🔎 Codex cross-audit (agent:audit)Blocking
No other PR-introduced security/correctness issues found in the added workflow. |
🔷 Gemini cross-audit (agent:audit-gemini)Audit Report:
|
The runtime-manager's build.rs compiles the runtime proto with prost-build, which shells out to protoc. Hosted runners do not ship it, so the build failed with 'Could not find protoc' — a dev Mac has it, which is exactly why this only appeared once the job ran in real CI rather than in local rehearsal. Same step sdlc-qa already uses.
The test failed with 'placement failed: no nodes registered in cluster'. The manager self-registers with the scheduler via SCHEDULER_URL (POST /v1/nodes/heartbeat), and the workflow never set it — the LaunchAgent setup this was rehearsed against already had it, which is exactly why the gap only surfaced in real CI. Scheduler now starts FIRST so registration lands immediately, the manager gets SCHEDULER_URL and a stable NODE_NAME, and a readiness step waits for the node instead of racing it. That step watches the manager's own log rather than GET /v1/cluster: the endpoint requires an Authorization header, so polling it would 401 for the full timeout and report nothing useful. It also fails fast on 'heartbeat rejected' / 'heartbeat send failed' instead of waiting out the timeout. Verified locally that 'heartbeat ok' really is emitted under RUST_LOG=...scheduler_heartbeat=debug before relying on it.
Item #2 from the remaining-work list. The exec e2e test has existed since the exec dispatch fix but only ran by hand (
LANTERN_RUNTIME_E2E=1), so the regression it guards was unprotected.That regression is worth remembering:
grpcSchedulerClient.Execreturned a canned string with exit code 0, making a missing implementation indistinguishable from a command that ran and printed nothing. Mocks can't catch that — only a real workload can.Why this needs no KVM
The runtime-manager's docker backend spawns a real container, and hosted runners have Docker. So this is an ordinary PR gate, while the Firecracker boot job still waits on a nested-virt runner. The job builds the manager and scheduler, starts both, and runs the test against them.
No Postgres service needed — with
DATABASE_URLunset the scheduler uses its in-memory store and runs always-leader, which is exactly what a single-node test wants.The gate asserts the test actually RAN
go testreports a skipped test as a pass. Without this the job would go green while proving nothing — the same failure mode as the boot assertion that reportedPASSfor kernel-panicking VMs all session. So the step:--- SKIP--- PASS: TestRuntimeExecE2E_RealVMlineVerified both directions before pushing
The test's own cleanup left no containers behind. Scheduler build command and workflow YAML both checked separately.
What this covers
The 6 subtests: stdout is real command output, the command runs inside the container, nonzero exit propagates, stderr stays separate, filesystem writes persist across execs, and an unknown vm_id errors instead of reporting success.
Workflow-only change.