Skip to content

Worker silently loses GPU access mid-run, and every deployment after that fails with a CUDA init error #186

Description

@nilsmechtel

A BioEngine worker started with Docker's --gpus=all can silently lose access to the GPU while it is running. The host GPU stays perfectly healthy and the worker keeps reporting itself as fine, because the applications already running hold the graphics device open and go on working. The damage only becomes visible the next time the worker has to start a fresh process on the GPU — a replica restart, a redeploy, a new app — and from that moment every GPU deployment fails at startup with a CUDA initialisation error. On a long-running worker this reads as "the GPU randomly disappears every couple of days"; in reality the access was lost much earlier, and a routine restart merely revealed it.

Restarting the container restores it, which is why this has been papered over repeatedly rather than diagnosed. The fix is a one-line change to how the container is launched, and it is already in the deployment guide as the vulnerable form.

--- detail below ---

Signature

Inside the container, while the host is healthy:

$ nvidia-smi
Failed to initialize NVML: Unknown Error

$ python -c "import os; os.open('/dev/nvidiactl', os.O_RDONLY)"
PermissionError: [Errno 1] Operation not permitted: '/dev/nvidiactl'

$ python -c "import ctypes; print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
100          # CUDA_ERROR_NO_DEVICE

On the host at the same moment nvidia-smi is normal and other containers are unaffected. The device nodes still exist in the container and have correct ownership and mode — open(2) is refused by the cgroup device allowlist, not by the filesystem.

In BioEngine this surfaces as DEPLOY_FAILED with GPU deployment was assigned a GPU but could not initialize a CUDA context, raised by the health gate in bioengine/_app/mixin.py:300 when CudaMemorySampler (bioengine/_app/gpu_memory.py) fails to cuInit through libcuda.so.1. That gate is behaving correctly — it is reporting a container-level fault, not an app fault.

Mechanism

--gpus=all does not put the NVIDIA device nodes in the container's OCI spec. It records a DeviceRequest, and the nvidia-container-runtime prestart hook injects the nodes and patches the device cgroup afterwards, out of band:

$ docker inspect <container> --format '{{json .HostConfig.DeviceRequests}} {{json .HostConfig.Devices}} {{.HostConfig.Runtime}}'
[{"Driver":"","Count":-1,"DeviceIDs":null,"Capabilities":[["gpu"]],"Options":{}}] [] runc

Devices is empty. So whenever anything causes runc to re-apply the container's cgroup configuration, the allowlist is rebuilt from the OCI spec — which never mentioned the NVIDIA majors — and the hook's patch is wiped. The container keeps running; it just can no longer open a graphics device.

Preconditions: cgroup v2 with Docker's systemd cgroup driver. Classic triggers are systemctl daemon-reload (including reloads fired by package managers and by snapd), and anything else that makes runc rewrite the cgroup for a live container.

This is a known interaction of --gpus + cgroup v2 + systemd cgroup driver, not a BioEngine bug. BioEngine is only the thing that notices.

Why it looks intermittent

The device cgroup gates open(2) and nothing else. A process that already holds an open /dev/nvidia* fd keeps its CUDA context and keeps computing after the wipe — inference continues to succeed, GPU memory reporting continues to work, the worker reports RUNNING/HEALTHY. Nothing fails until a new process opens a device node.

The general form is worth stating on its own, because it governs how any timeline of this fault should be read: when a fault gates acquisition of a resource rather than use of it, every process already holding the resource is immune, so the system looks healthy for an unbounded period and the eventual restart is the revealer, not the cause.

On the affected worker, both incidents were revealed by Ray OOM-killing replicas: the Serve controller recovers from checkpoint, the fresh replicas cannot init CUDA, and the app flips to DEPLOY_FAILED — potentially days after the access was actually lost. Dating the breakage from the OOM is the one reading the evidence rules out.

This also caveats the observed ~1.5–3 day failure cadence: that number is an estimate of replica-restart frequency and only a lower bound on wipe frequency. They are different quantities that happen to have been sampled by the same events.

Is a given container exposed? (read-only, safe on production)

Devices empty while DeviceRequests is populated is the vulnerable shape:

docker inspect <container> --format '{{json .HostConfig.Devices}} {{json .HostConfig.DeviceRequests}}'

Both the bare --gpus all / --gpus "device=N" form and the compose deploy.resources.reservations.devices form produce it. Add cgroup v2 + the systemd cgroup driver (docker info --format '{{.CgroupDriver}} {{.CgroupVersion}}') and the host is exposed.

Reproduction (≈5 seconds, no root)

Do not run this on a production container. On a container in the vulnerable shape it is not a probe, it is the trigger — it will break GPU access for every process started afterwards. Use a throwaway container.

Any cgroup-re-apply works; docker update is the cheapest one that doesn't need privileges:

docker run -d --name gputest --gpus all <image> sleep 600
docker exec gputest python -c "import ctypes,os; os.close(os.open('/dev/nvidiactl',os.O_RDONLY)); print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
# -> 0

docker update --cpu-shares 512 gputest          # forces runc to re-apply the cgroup

docker exec gputest python -c "import ctypes,os; os.close(os.open('/dev/nvidiactl',os.O_RDONLY)); print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
# -> PermissionError: [Errno 1] Operation not permitted: '/dev/nvidiactl'

Fix

Name the device nodes explicitly so they are in the OCI spec and survive every re-apply. Keep --gpus=all — it is still what mounts the driver libraries and sets up the container:

docker run ... \
  --gpus=all \
  --device /dev/nvidia0 \
  --device /dev/nvidiactl \
  --device /dev/nvidia-uvm \
  --device /dev/nvidia-uvm-tools \
  ...

Validated on the affected worker: four consecutive docker update re-applies, device open OK and cuInit 0 after each, cuDeviceGetCount 1, nvidia-smi normal inside the container; an unpatched control container on the same host broke on the first re-apply. No root, no daemon.json edit, no Docker restart required.

Caveats:

  • /dev/nvidia0 is per-GPU — a multi-GPU host needs one --device per board (/dev/nvidia0 … /dev/nvidiaN).
  • /dev/nvidia-uvm-tools is not always present; drop it if absent.
  • MIG or vGPU setups need their additional nodes (/dev/nvidia-caps/*) listed too.

The generic alternative is CDI — nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml and then --device nvidia.com/gpu=all — which puts the devices in the spec by construction and needs no per-GPU enumeration. It requires root to generate the spec, which is why the explicit --device list was used here.

Suggested changes

  • docs/deployment-guide.md:24 documents the vulnerable invocation (--gpus=all alone) as the recommended way to start a worker. It should carry the --device flags, or CDI, plus a short note on why. The Podman line immediately below already recommends --device nvidia.com/gpu=all, which is not vulnerable to this.
  • Worth considering: when the health gate raises the CUDA-context error, say in the message that a container-level device-access loss is the common cause and point at the fix. As written, the message reads like a scheduling or app problem and sends people looking in the wrong place.

Scope

Observed and reproduced on a single-machine Docker worker: Ubuntu, cgroup v2, Docker 29.3.1 with the systemd cgroup driver, nvidia-container-toolkit 1.19.0, driver 555.42.06, RTX 3080. It should apply to any Docker host with that cgroup driver combination. Kubernetes deployments are a different launch path and should not be assumed to share this cause without the open(2)/EPERM check above — that check is the discriminator: EPERM on open('/dev/nvidiactl') with a healthy host GPU means the device cgroup, anything else means something else.

Two cautions for anyone matching a GPU failure against this one:

  • nvidia-smi is not the discriminator. It is a strong positive signal here, but a virtualised GPU substrate (vGPU / mediated devices, where the guest driver talks to a host-side manager) can report a perfectly healthy card and full free memory while CUDA context creation fails. A GPU probe should be cuInit plus an allocation attempt, with the device-node open(2) check to localise the cause — not nvidia-smi output.
  • Probing only from the affected container cannot tell a container-level fault from a host-level one. Both look identical from inside. Re-probe from a fresh container on the same host before concluding the host or the GPU is dead; if the fresh one is clean, the fault is the container's, and restarting it is the cure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions