skip to content

What does `docker run --gpus all` require on the host, and what does it actually do to the container?

level: middleimportance: should knowfreq 62%

answer

  1. The flag only asks; something else answers
  2. Who supplies the driver, who supplies CUDA
  3. A hook that runs before the process
  4. Nodes injected, host libraries mounted
  5. CDI devices arrive through a different flag

basics

~20 s

--gpus only records a device request; a GPU-aware runtime must be installed on the host. The NVIDIA Container Toolkit's hook then creates /dev/nvidia* nodes in the container and mounts the host's driver libraries and nvidia-smi into it.

solid answer

~40 s

`--gpus all` is not implemented by the kernel or by `dockerd` alone — it stores a device request on the container and hands it to a runtime that understands GPUs. On an NVIDIA host that is the NVIDIA Container Toolkit, which registers a runtime wrapping `runc` and a hook that runs before the container's process starts. The hook creates `/dev/nvidiactl`, `/dev/nvidia-uvm` and the per-GPU `/dev/nvidia0…` nodes, then bind-mounts the *host's* user-space driver libraries (`libcuda.so`, `libnvidia-ml.so`) and `nvidia-smi` into the container. The image therefore ships the CUDA runtime and your application, never the driver. You can narrow the grant with `--gpus 2` or `--gpus '"device=0,1"'` (indices or GPU UUIDs) and select driver capabilities with `--gpus 'all,capabilities=compute,utility'`. If the toolkit is missing you get `could not select device driver with capabilities: [[gpu]]`.

code

bash · 9 lines
bash
# Smoke test: proves toolkit + driver + injection all work
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

# Two specific cards, only the CUDA and NVML driver slices
docker run --rm --gpus '"device=0,1"' \
  --gpus 'all,capabilities=compute,utility' my-inference:2.9

# What the daemon recorded for the container
docker inspect --format '{{json .HostConfig.DeviceRequests}}' infer-7

go deeper

for a junior

Know that --gpus all exists, that it needs a container toolkit installed on the host, and that the smoke test is running nvidia-smi inside a CUDA base image.

for a middle

Explain the split cleanly: the host supplies the kernel module and driver libraries, injected by a pre-start hook, while the image supplies the CUDA runtime and the application. Know the device= and capabilities= forms.

for a senior

Be ready to bootstrap and verify a GPU host end to end, read the device request out of docker inspect, recognise the could not select device driver failure as host configuration, and pin containers to specific cards.

for a principal

Own the fleet story: how driver versions are rolled across nodes, whether you standardise on CDI or the legacy hook, and where the boundary sits between single-host device grants and cluster-level GPU scheduling.

### `--gpus` is a request, not a mechanism Unlike `--device`, which the daemon can satisfy on its own, `--gpus` is a *device request*. Docker records it on the container (`docker inspect --format '{{json .HostConfig.DeviceRequests}}'` shows `Driver`, `Count` or `DeviceIDs`, and `Capabilities: [[gpu]]`) and expects some component on the host to know what a GPU is. If nothing does, the run fails immediately with: ``` docker: Error response from daemon: could not select device driver with capabilities: [[gpu]]. ``` That error is a host-configuration message, not an image problem — it appears just the same with a perfectly good CUDA image. ### What the NVIDIA Container Toolkit installs On an NVIDIA host the fulfilling component is the NVIDIA Container Toolkit. Installing it does three things: it puts an OCI runtime on the host that wraps `runc`, it registers that runtime with the daemon (an entry under `runtimes` in `/etc/docker/daemon.json`), and it installs a hook that runs after the container's filesystem is prepared but before its first process starts. The hook does the injection: * creates the control nodes `/dev/nvidiactl` and `/dev/nvidia-uvm` and one `/dev/nvidiaN` per requested GPU inside the container; * bind-mounts the host's user-space driver libraries — `libcuda.so.<driver version>`, `libnvidia-ml.so.<driver version>` and friends — into the container and runs `ldconfig`; * bind-mounts host binaries such as `nvidia-smi`. This is the single most important fact about GPU containers and the source of most interview follow-ups: **the driver comes from the host, the CUDA runtime comes from the image.** A container image must never install the NVIDIA driver; it installs the CUDA toolkit/runtime and the application. That is also why `nvidia-smi` run *inside* the container prints the host's driver version — it is literally the host's binary talking to the host's kernel module. ### Selecting GPUs and capabilities ``` docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi docker run --rm --gpus 2 ... # any two GPUs docker run --rm --gpus '"device=0,1"' ... # specific indices docker run --rm --gpus '"device=GPU-0e1f…"' ... # by GPU UUID docker run --rm --gpus 'all,capabilities=compute,utility' ... ``` The quoting on `device=` is awkward because the value itself must reach the daemon quoted. Under the hood these map onto the environment variables the toolkit reads — `NVIDIA_VISIBLE_DEVICES` and `NVIDIA_DRIVER_CAPABILITIES` — which the vendor's CUDA base images already set (typically `compute,utility`). Capabilities decide which slices of the driver get mounted: `compute` for CUDA, `utility` for `nvidia-smi`/NVML, `video` for hardware encode/decode. A container that only needs NVENC does not need `compute`. ### The older spellings Before Docker 19.03 there was no `--gpus`; you ran `--runtime=nvidia`, or used a `nvidia-docker` wrapper, or set the NVIDIA runtime as the daemon's default. Those still work where they are configured, and you will meet them in old runbooks, but `--gpus` is the current spelling and the one to give in an interview. ### CDI: the vendor-neutral successor The Container Device Interface (CDI) replaces per-vendor hooks with a declarative spec file. A tool on the host (for NVIDIA, `nvidia-ctk cdi generate`) writes a JSON/YAML spec into `/etc/cdi` or `/var/run/cdi` describing named devices such as `nvidia.com/gpu=0` and exactly which nodes, mounts and environment variables each one implies. Docker Engine gained CDI support in 25.0, initially behind a `cdi` entry in the daemon's `features` map, and CDI devices are requested through **`--device`, not `--gpus`**: ``` docker run --rm --device nvidia.com/gpu=0 my-inference:2.9 ``` That detail — GPUs arriving through `--device` once CDI is in play — is a common gotcha for people who learned only the `--gpus` spelling. ### What `--gpus` does *not* do It controls *visibility*, not partitioning. Two containers each started with `--gpus all` see the same physical GPU and compete for its memory; one filling the card makes the other fail to allocate. There is no per-container GPU memory limit equivalent to a memory cgroup. Placement across many GPUs and many hosts is a scheduling problem that belongs to an orchestrator, not to `docker run`. The single-host answer is discipline: hand each container an explicit `device=` list rather than `all`, so two jobs cannot silently land on the same card. Finally, verify rather than assume. `docker run --rm --gpus all <cuda-base-image> nvidia-smi` is the one-line smoke test that proves the toolkit is installed, the driver is loaded, the nodes were injected, and the container can see the cards — and it is worth running as part of node bootstrap before any real workload is scheduled onto a host.

  • Why does `nvidia-smi` inside a container report the host's driver version rather than anything from the image?
    Because it is the host's binary and the host's libraries. The toolkit hook bind-mounts `nvidia-smi` and the matching `libnvidia-ml.so` from the host into the container, and they talk to the host's kernel module. Nothing in the image contributes a driver version. That is why the correct fix for a driver problem is always a host action, and why an image that installs the driver itself is a bug.
  • Two containers on one host are both started with `--gpus all`. What isolation do they get?
    Visibility only. Both see every card and both allocate from the same GPU memory, so whichever grows first can starve the other into allocation failures; there is no memory cgroup equivalent for GPU RAM. On a single host the mitigation is to give each container an explicit `--gpus '"device=N"'` list so they never share a card, and to treat placement across cards as something you decide, not something the engine decides for you.
  • You inherit a host where GPU containers run with `--runtime=nvidia` and no `--gpus` flag. Is that broken?
    No, just older. Before Docker 19.03 the toolkit was selected by naming its runtime, with `NVIDIA_VISIBLE_DEVICES` in the environment choosing the cards; that path still works where the runtime is registered. It is worth migrating to `--gpus` so the request is visible in `docker inspect` as a device request rather than hidden in an environment variable, and so the host is not depending on a default-runtime setting.

saying these in an interview costs you the question

  • Says the container image must install the NVIDIA driver
  • Thinks `--gpus all` works on any host with a GPU installed
  • Believes `--gpus` gives each container its own slice of GPU memory
  • Confuses the CUDA runtime in the image with the host driver
  • Suggests `--privileged` as the way to expose GPUs

context