A BioEngine worker started with Docker's --gpus=all can silently lose access to the GPU while it is running. The host GPU stays perfectly healthy and the worker keeps reporting itself as fine, because the applications already running hold the graphics device open and go on working. The damage only becomes visible the next time the worker has to start a fresh process on the GPU — a replica restart, a redeploy, a new app — and from that moment every GPU deployment fails at startup with a CUDA initialisation error. On a long-running worker this reads as "the GPU randomly disappears every couple of days"; in reality the access was lost much earlier, and a routine restart merely revealed it.
Restarting the container restores it, which is why this has been papered over repeatedly rather than diagnosed. The fix is a one-line change to how the container is launched, and it is already in the deployment guide as the vulnerable form.
--- detail below ---
Signature
Inside the container, while the host is healthy:
$ nvidia-smi
Failed to initialize NVML: Unknown Error
$ python -c "import os; os.open('/dev/nvidiactl', os.O_RDONLY)"
PermissionError: [Errno 1] Operation not permitted: '/dev/nvidiactl'
$ python -c "import ctypes; print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
100 # CUDA_ERROR_NO_DEVICE
On the host at the same moment nvidia-smi is normal and other containers are unaffected. The device nodes still exist in the container and have correct ownership and mode — open(2) is refused by the cgroup device allowlist, not by the filesystem.
In BioEngine this surfaces as DEPLOY_FAILED with GPU deployment was assigned a GPU but could not initialize a CUDA context, raised by the health gate in bioengine/_app/mixin.py:300 when CudaMemorySampler (bioengine/_app/gpu_memory.py) fails to cuInit through libcuda.so.1. That gate is behaving correctly — it is reporting a container-level fault, not an app fault.
Mechanism
--gpus=all does not put the NVIDIA device nodes in the container's OCI spec. It records a DeviceRequest, and the nvidia-container-runtime prestart hook injects the nodes and patches the device cgroup afterwards, out of band:
$ docker inspect <container> --format '{{json .HostConfig.DeviceRequests}} {{json .HostConfig.Devices}} {{.HostConfig.Runtime}}'
[{"Driver":"","Count":-1,"DeviceIDs":null,"Capabilities":[["gpu"]],"Options":{}}] [] runc
Devices is empty. So whenever anything causes runc to re-apply the container's cgroup configuration, the allowlist is rebuilt from the OCI spec — which never mentioned the NVIDIA majors — and the hook's patch is wiped. The container keeps running; it just can no longer open a graphics device.
Preconditions: cgroup v2 with Docker's systemd cgroup driver. Classic triggers are systemctl daemon-reload (including reloads fired by package managers and by snapd), and anything else that makes runc rewrite the cgroup for a live container.
This is a known interaction of --gpus + cgroup v2 + systemd cgroup driver, not a BioEngine bug. BioEngine is only the thing that notices.
Why it looks intermittent
The device cgroup gates open(2) and nothing else. A process that already holds an open /dev/nvidia* fd keeps its CUDA context and keeps computing after the wipe — inference continues to succeed, GPU memory reporting continues to work, the worker reports RUNNING/HEALTHY. Nothing fails until a new process opens a device node.
The general form is worth stating on its own, because it governs how any timeline of this fault should be read: when a fault gates acquisition of a resource rather than use of it, every process already holding the resource is immune, so the system looks healthy for an unbounded period and the eventual restart is the revealer, not the cause.
On the affected worker, both incidents were revealed by Ray OOM-killing replicas: the Serve controller recovers from checkpoint, the fresh replicas cannot init CUDA, and the app flips to DEPLOY_FAILED — potentially days after the access was actually lost. Dating the breakage from the OOM is the one reading the evidence rules out.
This also caveats the observed ~1.5–3 day failure cadence: that number is an estimate of replica-restart frequency and only a lower bound on wipe frequency. They are different quantities that happen to have been sampled by the same events.
Is a given container exposed? (read-only, safe on production)
Devices empty while DeviceRequests is populated is the vulnerable shape:
docker inspect <container> --format '{{json .HostConfig.Devices}} {{json .HostConfig.DeviceRequests}}'
Both the bare --gpus all / --gpus "device=N" form and the compose deploy.resources.reservations.devices form produce it. Add cgroup v2 + the systemd cgroup driver (docker info --format '{{.CgroupDriver}} {{.CgroupVersion}}') and the host is exposed.
Reproduction (≈5 seconds, no root)
Do not run this on a production container. On a container in the vulnerable shape it is not a probe, it is the trigger — it will break GPU access for every process started afterwards. Use a throwaway container.
Any cgroup-re-apply works; docker update is the cheapest one that doesn't need privileges:
docker run -d --name gputest --gpus all <image> sleep 600
docker exec gputest python -c "import ctypes,os; os.close(os.open('/dev/nvidiactl',os.O_RDONLY)); print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
# -> 0
docker update --cpu-shares 512 gputest # forces runc to re-apply the cgroup
docker exec gputest python -c "import ctypes,os; os.close(os.open('/dev/nvidiactl',os.O_RDONLY)); print(ctypes.CDLL('libcuda.so.1').cuInit(0))"
# -> PermissionError: [Errno 1] Operation not permitted: '/dev/nvidiactl'
Fix
Name the device nodes explicitly so they are in the OCI spec and survive every re-apply. Keep --gpus=all — it is still what mounts the driver libraries and sets up the container:
docker run ... \
--gpus=all \
--device /dev/nvidia0 \
--device /dev/nvidiactl \
--device /dev/nvidia-uvm \
--device /dev/nvidia-uvm-tools \
...
Validated on the affected worker: four consecutive docker update re-applies, device open OK and cuInit 0 after each, cuDeviceGetCount 1, nvidia-smi normal inside the container; an unpatched control container on the same host broke on the first re-apply. No root, no daemon.json edit, no Docker restart required.
Caveats:
/dev/nvidia0 is per-GPU — a multi-GPU host needs one --device per board (/dev/nvidia0 … /dev/nvidiaN).
/dev/nvidia-uvm-tools is not always present; drop it if absent.
- MIG or vGPU setups need their additional nodes (
/dev/nvidia-caps/*) listed too.
The generic alternative is CDI — nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml and then --device nvidia.com/gpu=all — which puts the devices in the spec by construction and needs no per-GPU enumeration. It requires root to generate the spec, which is why the explicit --device list was used here.
Suggested changes
docs/deployment-guide.md:24 documents the vulnerable invocation (--gpus=all alone) as the recommended way to start a worker. It should carry the --device flags, or CDI, plus a short note on why. The Podman line immediately below already recommends --device nvidia.com/gpu=all, which is not vulnerable to this.
- Worth considering: when the health gate raises the CUDA-context error, say in the message that a container-level device-access loss is the common cause and point at the fix. As written, the message reads like a scheduling or app problem and sends people looking in the wrong place.
Scope
Observed and reproduced on a single-machine Docker worker: Ubuntu, cgroup v2, Docker 29.3.1 with the systemd cgroup driver, nvidia-container-toolkit 1.19.0, driver 555.42.06, RTX 3080. It should apply to any Docker host with that cgroup driver combination. Kubernetes deployments are a different launch path and should not be assumed to share this cause without the open(2)/EPERM check above — that check is the discriminator: EPERM on open('/dev/nvidiactl') with a healthy host GPU means the device cgroup, anything else means something else.
Two cautions for anyone matching a GPU failure against this one:
nvidia-smi is not the discriminator. It is a strong positive signal here, but a virtualised GPU substrate (vGPU / mediated devices, where the guest driver talks to a host-side manager) can report a perfectly healthy card and full free memory while CUDA context creation fails. A GPU probe should be cuInit plus an allocation attempt, with the device-node open(2) check to localise the cause — not nvidia-smi output.
- Probing only from the affected container cannot tell a container-level fault from a host-level one. Both look identical from inside. Re-probe from a fresh container on the same host before concluding the host or the GPU is dead; if the fresh one is clean, the fault is the container's, and restarting it is the cure.
A BioEngine worker started with Docker's
--gpus=allcan silently lose access to the GPU while it is running. The host GPU stays perfectly healthy and the worker keeps reporting itself as fine, because the applications already running hold the graphics device open and go on working. The damage only becomes visible the next time the worker has to start a fresh process on the GPU — a replica restart, a redeploy, a new app — and from that moment every GPU deployment fails at startup with a CUDA initialisation error. On a long-running worker this reads as "the GPU randomly disappears every couple of days"; in reality the access was lost much earlier, and a routine restart merely revealed it.Restarting the container restores it, which is why this has been papered over repeatedly rather than diagnosed. The fix is a one-line change to how the container is launched, and it is already in the deployment guide as the vulnerable form.
--- detail below ---
Signature
Inside the container, while the host is healthy:
On the host at the same moment
nvidia-smiis normal and other containers are unaffected. The device nodes still exist in the container and have correct ownership and mode —open(2)is refused by the cgroup device allowlist, not by the filesystem.In BioEngine this surfaces as
DEPLOY_FAILEDwithGPU deployment was assigned a GPU but could not initialize a CUDA context, raised by the health gate inbioengine/_app/mixin.py:300whenCudaMemorySampler(bioengine/_app/gpu_memory.py) fails tocuInitthroughlibcuda.so.1. That gate is behaving correctly — it is reporting a container-level fault, not an app fault.Mechanism
--gpus=alldoes not put the NVIDIA device nodes in the container's OCI spec. It records aDeviceRequest, and the nvidia-container-runtime prestart hook injects the nodes and patches the device cgroup afterwards, out of band:Devicesis empty. So whenever anything causes runc to re-apply the container's cgroup configuration, the allowlist is rebuilt from the OCI spec — which never mentioned the NVIDIA majors — and the hook's patch is wiped. The container keeps running; it just can no longer open a graphics device.Preconditions: cgroup v2 with Docker's systemd cgroup driver. Classic triggers are
systemctl daemon-reload(including reloads fired by package managers and by snapd), and anything else that makes runc rewrite the cgroup for a live container.This is a known interaction of
--gpus+ cgroup v2 + systemd cgroup driver, not a BioEngine bug. BioEngine is only the thing that notices.Why it looks intermittent
The device cgroup gates
open(2)and nothing else. A process that already holds an open/dev/nvidia*fd keeps its CUDA context and keeps computing after the wipe — inference continues to succeed, GPU memory reporting continues to work, the worker reportsRUNNING/HEALTHY. Nothing fails until a new process opens a device node.The general form is worth stating on its own, because it governs how any timeline of this fault should be read: when a fault gates acquisition of a resource rather than use of it, every process already holding the resource is immune, so the system looks healthy for an unbounded period and the eventual restart is the revealer, not the cause.
On the affected worker, both incidents were revealed by Ray OOM-killing replicas: the Serve controller recovers from checkpoint, the fresh replicas cannot init CUDA, and the app flips to
DEPLOY_FAILED— potentially days after the access was actually lost. Dating the breakage from the OOM is the one reading the evidence rules out.This also caveats the observed ~1.5–3 day failure cadence: that number is an estimate of replica-restart frequency and only a lower bound on wipe frequency. They are different quantities that happen to have been sampled by the same events.
Is a given container exposed? (read-only, safe on production)
Devicesempty whileDeviceRequestsis populated is the vulnerable shape:Both the bare
--gpus all/--gpus "device=N"form and the composedeploy.resources.reservations.devicesform produce it. Add cgroup v2 + the systemd cgroup driver (docker info --format '{{.CgroupDriver}} {{.CgroupVersion}}') and the host is exposed.Reproduction (≈5 seconds, no root)
Any cgroup-re-apply works;
docker updateis the cheapest one that doesn't need privileges:Fix
Name the device nodes explicitly so they are in the OCI spec and survive every re-apply. Keep
--gpus=all— it is still what mounts the driver libraries and sets up the container:Validated on the affected worker: four consecutive
docker updatere-applies, device open OK andcuInit0 after each,cuDeviceGetCount1,nvidia-sminormal inside the container; an unpatched control container on the same host broke on the first re-apply. No root, nodaemon.jsonedit, no Docker restart required.Caveats:
/dev/nvidia0is per-GPU — a multi-GPU host needs one--deviceper board (/dev/nvidia0 … /dev/nvidiaN)./dev/nvidia-uvm-toolsis not always present; drop it if absent./dev/nvidia-caps/*) listed too.The generic alternative is CDI —
nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yamland then--device nvidia.com/gpu=all— which puts the devices in the spec by construction and needs no per-GPU enumeration. It requires root to generate the spec, which is why the explicit--devicelist was used here.Suggested changes
docs/deployment-guide.md:24documents the vulnerable invocation (--gpus=allalone) as the recommended way to start a worker. It should carry the--deviceflags, or CDI, plus a short note on why. The Podman line immediately below already recommends--device nvidia.com/gpu=all, which is not vulnerable to this.Scope
Observed and reproduced on a single-machine Docker worker: Ubuntu, cgroup v2, Docker 29.3.1 with the systemd cgroup driver, nvidia-container-toolkit 1.19.0, driver 555.42.06, RTX 3080. It should apply to any Docker host with that cgroup driver combination. Kubernetes deployments are a different launch path and should not be assumed to share this cause without the
open(2)/EPERM check above — that check is the discriminator: EPERM onopen('/dev/nvidiactl')with a healthy host GPU means the device cgroup, anything else means something else.Two cautions for anyone matching a GPU failure against this one:
nvidia-smiis not the discriminator. It is a strong positive signal here, but a virtualised GPU substrate (vGPU / mediated devices, where the guest driver talks to a host-side manager) can report a perfectly healthy card and full free memory while CUDA context creation fails. A GPU probe should becuInitplus an allocation attempt, with the device-nodeopen(2)check to localise the cause — notnvidia-smioutput.