skip to content

A GPU inference service genuinely needs host device access; how do you contain the blast radius rather than reaching for --privileged?

level: principalimportance: should knowfreq 38%

answer

  1. Challenge 'it needs --privileged'
  2. Narrow --device and --cap-add, not the superset
  3. Assume breakout, bound what it reaches
  4. Dedicated host, minimal cloud identity
  5. Patch the runtime; accept residual risk

basics

~20 s

Grant the minimum, not the superset: expose only the specific device and the few capabilities the workload needs, never blanket --privileged. Then assume breakout is possible and shrink what a breakout reaches -- a dedicated low-trust host or node pool, no other tenants or secrets, a minimal cloud identity, network segmentation, and a patched runtime.

solid answer

~50 s

Start by challenging the premise: 'needs privilege' almost always means it needs one device or one capability, which you can grant narrowly (`--device /dev/nvidia*` or the GPU vendor's device plugin, plus a specific `--cap-add`) instead of `--privileged`, which is the lazy all-caps, all-devices, unconfined superset. Then reason about blast radius by **assuming the container can reach root on its host** -- because a raw device or a runtime CVE might let it. So bound what that would find: run it on a **dedicated node pool** with no other tenants, no secrets, and no docker.sock; give that host a **minimal cloud identity** so a breakout inherits little; segment its network egress; and keep runc/containerd patched, since runtime escapes ignore every config control. Accept and document the residual risk with an owner and detection around it. The tradeoff is real: dedicating hosts lowers utilization, and that is the price of containing a workload you cannot fully confine.

code

bash · 6 lines
bash
# minimum grant for CUDA inference, not the all-caps superset
docker run --rm --gpus '"device=0"' \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  --read-only --tmpfs /tmp \
  ml-inference:cuda

go deeper

for a junior

Recall that --privileged is dangerous and that a container usually needs only a specific device or capability, which can be granted narrowly instead of turning everything on.

for a middle

Explain how to replace --privileged with a narrow --device and specific --cap-add, and why that shrinks the escape surface without breaking a GPU workload.

for a senior

Design the containment: dedicated host/pool, no colocated secrets or docker.sock, least-privilege cloud identity, network segmentation, and a patched runtime -- justified by assuming the container can reach host root.

for a principal

Own the tradeoff between isolation cost and blast radius across the estate: decide which workloads get dedicated hardware, encode the policy in fleet-wide controls, and record an explicit, owned, reviewed risk acceptance for the privilege you cannot remove.

### First, refuse the false 'needs --privileged' Most requests for `--privileged` are wrong about their own requirements. `--privileged` is a superset -- **all** capabilities, **all** host devices, unmasked `/proc` and `/sys`, and no seccomp/AppArmor confinement. Almost no workload needs that whole bundle. A **Python ML inference image with CUDA wheels** needs the GPU, which you provide with the vendor's GPU runtime / device plugin or a narrow `--device /dev/nvidia0 --device /dev/nvidiactl`, not the world. A workload that wants one capability (say `NET_ADMIN` for a tunnel) gets exactly that with `--cap-add NET_ADMIN` over `--cap-drop ALL`. So step one is engineering the grant down to the minimum that makes the workload run. That alone removes most of the escape surface. ### Then plan for the grant you cannot remove Some grants are irreducibly dangerous -- a raw block device, a vendor runtime that insists on broad device access, a capability like `CAP_SYS_ADMIN` the app genuinely uses. Here config hardening has run out of road, so you switch from *prevention* to **blast-radius containment**: design as if the container will reach root on its host, and make that outcome cheap. **1. Dedicate the host / node pool.** Run the risky workload only on hosts that run *nothing else you care about*. No other tenants' containers, no shared secrets on disk, no control-plane components, no `docker.sock` mounted anywhere on that host. If the container escapes, it lands on a machine whose only job was to run it. **2. Minimize the host's own privilege.** Give that host (or its node pool) a **least-privilege cloud identity** -- a role scoped to exactly the buckets/queues the workload uses. A breakout inherits the host's credentials; if those credentials are narrow, the breakout is narrow. Do not put broad admin roles or long-lived org tokens on a host that runs a low-trust workload. **3. Segment the network.** Default-deny egress from that pool, allow-list only what the inference service must reach, and firewall it away from your control plane, secret stores and other subnets so a foothold cannot pivot laterally. **4. Patch the runtime.** Keep runc/containerd/Docker current. Runtime escapes (the runc CVEs) bypass every capability, seccomp and user setting, so an unpatched engine makes all the above moot for a determined attacker -- patching is *part of* the containment plan, not separate from it. **5. Detect and accept.** Put monitoring on the pool (unexpected egress, new processes on the host, credential use) and record an **explicit, owned risk acceptance**: what is granted, why, who owns it, and when it is reviewed. An unreviewed 'temporary' privileged exception is how these become permanent. ### The tradeoff you own Containment costs money and utilization. Dedicating a GPU node pool to one class of workload means you cannot bin-pack other services onto that expensive hardware; least-privilege identities and segmentation add operational friction. That is the real principal-level decision: how much isolation to buy for a workload you cannot fully confine, weighed against the cost of a host compromise blast-radiating into the rest of the estate. For anything processing untrusted input or third-party code, buy the isolation. Where an orchestrator is in play, its admission and policy controls can enforce these choices across the fleet, but the reasoning above is what those controls encode. ### The one-line stance Grant the minimum; assume the minimum you must grant is still enough to escape; make the escape land somewhere it cannot hurt you; and keep the runtime patched so the assumption stays a hypothetical.

  • The vendor's GPU runtime still wants broad device access -- how do you bound what it can reach?
    Pin the workload to a dedicated node pool that runs only trusted GPU jobs, give those nodes a minimal cloud identity, keep no other tenants' data, secrets or control-plane components on them, segment their egress, and monitor. When you cannot narrow the grant itself, you bound the damage by controlling what a breakout would find on that host.
  • How does keeping runc/containerd patched fit a blast-radius plan?
    It is part of it, not separate. Config hardening cannot stop a runtime-level escape -- the runc CVEs escaped from fully unprivileged containers -- so a current engine is a load-bearing control. An unpatched runtime makes dedicated hosts, dropped capabilities and least-privilege identities all bypassable, so patching is what keeps 'assume breakout' a hypothetical rather than a certainty.

You cannot make a fireworks workshop fireproof, so you put it in its own shed at the edge of the lot with nothing valuable inside and a fence around it -- if it goes up, only the shed is lost.

saying these in an interview costs you the question

  • Reaches for --privileged because the app 'needs a device'
  • Colocates the privileged workload with other tenants and secrets
  • Thinks config hardening alone contains a runtime escape
  • Gives the risky host a broad cloud role a breakout would inherit
  • Treats blast-radius planning as unnecessary once flags are set

context