skip to content

A Docker GPU container fails with `CUDA driver version is insufficient for CUDA runtime version`. How do you diagnose and fix it?

level: seniorimportance: should knowfreq 44%

answer

  1. Two stacks meeting at one seam
  2. Which side is old, host or image
  3. Ask the container what driver it got
  4. Majors matter, minors usually do not
  5. Error 35 is cudaErrorInsufficientDriver

basics

~20 s

The host's NVIDIA driver is older than the CUDA runtime in the image. Compare the driver version nvidia-smi reports inside the container with the image's CUDA version, then upgrade the host driver or rebuild on an older CUDA base.

solid answer

~50 s

GPU containers straddle two version stacks: the driver, which lives on the host and is injected into the container, and the CUDA runtime, which lives in the image. This error (CUDA error 35) means the injected driver is too old for the runtime the application was built against. Diagnose by running `nvidia-smi` inside the container — the version it prints is the *host's* driver, since the binary and libraries were mounted in — and comparing it with the image's CUDA version from its tag or `/usr/local/cuda/version.json`. Since CUDA 11, minor-version compatibility means a same-major mismatch usually works, so a real error 35 almost always means a major-version gap. The fix is a host action: upgrade the driver on the node, or pin the image to a CUDA major the fleet's driver supports. Never install a driver inside the image.

code

bash · 8 lines
bash
# What driver did the container actually receive? (this is the HOST's)
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

# What CUDA runtime is in the image?
docker run --rm my-rasteriser:3.2 cat /usr/local/cuda/version.json

# Same request, no GPU-aware runtime on the host:
#   could not select device driver with capabilities: [[gpu]]

go deeper

for a junior

Take away the core rule: the driver comes from the host, CUDA comes from the image, and a version error means those two disagree. Knowing which side to look at is enough here.

for a middle

Explain what nvidia-smi inside a container is really reporting and why, and be able to name the image's CUDA version from its tag or version file when comparing the two halves.

for a senior

Show a clean triage that separates error 35, the NVML driver/library mismatch, a missing GPU runtime and a missing --gpus flag, and say who owns the fix in each case before touching anything.

for a principal

Own driver lifecycle across the fleet: a documented minimum driver version per image, controlled driver rolls with drain and smoke test, and a policy on forward-compatibility packages as time-boxed debt.

### The two stacks Every GPU container has a seam running through it. Below the seam, on the host: the NVIDIA kernel module and the matching user-space driver libraries, which a container toolkit hook mounts into the container at start. Above the seam, in the image: the CUDA runtime (`libcudart`), any CUDA libraries such as cuBLAS or cuDNN, and the application. The image never contains a driver, and the host never contains your CUDA runtime. Every version failure in this area is a statement about that seam. `CUDA driver version is insufficient for CUDA runtime version` is CUDA error 35, `cudaErrorInsufficientDriver`. It says: the driver I found is older than the minimum this CUDA runtime demands. ### Diagnosis, in order **1. Confirm the container sees a GPU at all.** `docker run --rm --gpus all <image> nvidia-smi`. A failure here is a different problem — `could not select device driver with capabilities: [[gpu]]` means no GPU-aware runtime is installed on the host, and `no CUDA-capable device is detected` usually means the `--gpus` flag was forgotten or the visible-device list was empty. **2. Read the driver version — from inside the container.** The number `nvidia-smi` prints inside the container is the host's driver version, because both the binary and `libnvidia-ml.so` were bind-mounted from the host. That is the whole point: you do not need to shell onto the node to learn what driver the container actually got. **3. Read the image's CUDA version.** The base-image tag usually says it (`nvidia/cuda:12.4.1-runtime-ubuntu22.04`); otherwise `cat /usr/local/cuda/version.json` inside the container, or `nvcc --version` if the toolkit is present. Frameworks bundle their own CUDA build, so also check what the framework was compiled against rather than assuming the base image is the whole story. **4. Compare majors, not minors.** Since CUDA 11, minor-version compatibility means an application built with any 12.x toolkit runs on a driver that meets the *major* version's minimum, so a 12.4 image on a 12.0-era driver is normally fine. A genuine error 35 therefore nearly always means a major gap — for example a node still on a CUDA-11-era driver running an image built on CUDA 12.4. That framing narrows the investigation immediately. ### The fixes, best first * **Upgrade the host driver.** The canonical answer. It is a node-level operation: drain the node, upgrade the driver package, reload the kernel module (or reboot), re-run the smoke test. Nothing about the image changes. * **Pin the image down.** When the fleet's driver cannot move — a leased host, a vendor-supported image, a change window weeks away — rebuild on a CUDA base image whose major the existing driver supports, and pin that tag so a later rebuild cannot silently step forward. * **Forward compatibility, narrowly.** NVIDIA ships `cuda-compat` packages that place a newer user-space driver alongside the old one so a newer CUDA runtime can run on an older kernel module. It is supported only on data-center GPUs and requires the library path to be set so the compat build is found first. Treat it as an escape hatch for a specific fleet, not as a default. ### Failures that look similar and are not * `Failed to initialize NVML: Driver/library version mismatch` — the host driver package was upgraded while the old kernel module was still loaded, so user space and kernel space disagree. Containers started before and after the upgrade both hit it. The fix is on the host: unload and reload the module, or reboot the node. No image change helps. * `nvidia-container-cli: initialization error` complaining the driver is not loaded — the node has the toolkit but no working driver at all. * `could not select device driver with capabilities: [[gpu]]` — the toolkit itself is missing or not registered with the daemon. Separating these four is most of the value a senior engineer adds here, because the remediation owner is different in each case: image team, node operator, node operator, node bootstrap. ### Worked incident A batch job attached to a PDF-signing service does GPU-accelerated rasterisation before signing. It runs happily on the build fleet and dies on three older inference nodes with error 35. `docker run --rm --gpus all <image> nvidia-smi` on a failing node prints driver `470.223.02`; the image is `nvidia/cuda:12.4.1-runtime-ubuntu22.04`. That is a major gap — a CUDA-11-era driver against a CUDA 12 runtime — so no amount of rebuilding the application layer will help. The three nodes were drained and their drivers upgraded to a 12-capable branch; the image was left alone, and a `nvidia-smi` smoke test was added to node bootstrap so a node whose driver is behind the fleet's floor never becomes schedulable in the first place. ### Prevention Record a minimum driver version alongside each GPU image and check it at container start or at node bootstrap; pin CUDA base-image tags exactly rather than floating on `latest`; make the driver a fleet-wide property that moves in a controlled roll rather than per-node drift; and never, under any circumstances, `apt-get install` a driver inside an image to "make the versions match".

  • How does `Failed to initialize NVML: Driver/library version mismatch` differ from error 35?
    Error 35 is a mismatch across the container seam: host driver older than the image's CUDA runtime. The NVML mismatch is entirely host-side — the driver's user-space packages were upgraded while the old kernel module stayed loaded, so the two halves of the driver disagree. Rebuilding or repinning the image changes nothing; the node needs its module reloaded or a reboot. Recognising which side owns the fix is the point of the distinction.
  • Why is installing the NVIDIA driver inside the image never the fix?
    The driver has a kernel-module half that must match the running host kernel, and containers share that kernel — you cannot ship one in an image. The toolkit hook also mounts the host's driver libraries over the image's, so an in-image install is at best ignored and at worst produces a library set that disagrees with the loaded module. It bloats the image and makes it non-portable across nodes for no gain.
  • The fleet's driver cannot be upgraded for another month. What do you do with an image that needs a newer CUDA?
    Rebuild on a CUDA base image whose major the current driver supports and pin the tag, which is the low-risk option. If that is impossible because a framework build demands the newer CUDA, the narrow escape hatch is a forward-compatibility (`cuda-compat`) package that places a newer user-space driver alongside the old kernel module — supported only on data-center GPUs and requiring the library path to prefer it. Record it as debt tied to the upgrade date.

saying these in an interview costs you the question

  • Proposes installing the GPU driver inside the container image
  • Thinks `nvidia-smi` in the container shows the image's CUDA version
  • Confuses the NVML library/kernel mismatch with a host-versus-image mismatch
  • Reaches for `--privileged` to fix a driver version error
  • Assumes any CUDA image runs on any host with a GPU
  • Treats a driver upgrade as a per-container rather than per-node action

context