How do `docker run --pids-limit` and `--ulimit nproc` differ as defences against a fork bomb in a container?
answer
- Two different kernel mechanisms, one goal
- One is a cgroup, one is a rlimit
- What exactly does each one count
- Per container, or per UID host-wide
- Threads are tasks too
basics
~20 s--pids-limit sets the container cgroup's pids.max, capping every process and thread inside that one container; further clone() calls fail with EAGAIN. --ulimit nproc sets a per-UID kernel limit counted across the whole host, so it is the weaker and more surprising control.
solid answer
~40 s`docker run --pids-limit 384` writes `pids.max` for the container's cgroup, so the kernel counts every task — processes *and* threads — in that container and refuses to create the next one with `EAGAIN`. It is a true per-container cap, which is what you want against a fork bomb or a runaway thread pool: the container degrades, the host does not run out of PIDs. `--ulimit nproc=384` sets `RLIMIT_NPROC`, which the kernel counts per UID across the whole host — two containers running as the same UID share one budget, a container that can change UID escapes it, and processes on the host under that UID count too. Use `--pids-limit` as the control and treat `--ulimit nofile` (open file descriptors) as a separate, unrelated knob.
code
bash · 7 linesdocker run -d --name queue-worker \
--pids-limit 384 \
--ulimit nofile=9216:16384 \
queue-worker:3.2
docker inspect -f '{{.HostConfig.PidsLimit}}' queue-worker
docker stats --no-stream --format '{{.Name}} {{.PIDs}}' queue-workergo deeper
Know that a container can be told how many processes it may create, and that the flag on docker run is --pids-limit. Recognise Resource temporarily unavailable as a fork failure rather than a disk or permission problem.
Explain the mechanics: the pids cgroup counts tasks including threads and fails clone with EAGAIN, whereas RLIMIT_NPROC is counted per UID across the host. Say why that difference makes one a container boundary and the other not.
Show you can operate it: pick a cap from measured pids.current peaks with headroom, recognise the JVM and runtime error strings it produces, and diagnose from outside when docker exec itself can no longer start.
Frame it as host-resource protection rather than per-service tuning. Be ready to justify a fleet default, the size of the headroom, and why exhausting a host-wide resource such as the PID space is a different class of incident from one container degrading.
## Two limits that look alike and are not Both flags cap "how many", both are set on `docker run`, and both are reached for when someone says "stop a container forking forever". They are enforced by completely different kernel machinery. **`--pids-limit N`** configures the pids cgroup controller for that container's cgroup — `pids.max`. The controller counts *tasks*, and in Linux a thread is a task, so a JVM with 200 threads consumes 200 of the budget just as 200 separate processes would. When the cgroup is at its limit, `fork()` and `clone()` fail with `EAGAIN`. The scope is exactly the container: its own processes, nothing else, no matter what UID they run as. `--pids-limit -1` (or 0) means unlimited. **`--ulimit nproc=N`** sets the POSIX resource limit `RLIMIT_NPROC` on the container's initial process, which children inherit. The kernel enforces it **per real UID, counted across the whole host**, not per container. That produces three surprises worth naming in an interview: 1. Two containers whose processes run as the same UID — very common, since UID 1000 or UID 0 is the default in countless images — draw from a single shared budget. One noisy container makes the other fail to fork. 2. Processes running as that UID *outside* any container count toward the same total. 3. A container process that can change UID (the default container root holds `CAP_SETUID`) simply moves to a different UID and gets a fresh budget. So `nproc` is not a container boundary. `--pids-limit` is. ## What hitting the limit looks like The symptom is rarely "fork bomb". It is a service that stops making progress in a way that reads like an application bug: - Shell and libc: `fork: Resource temporarily unavailable`. - A JVM: `java.lang.OutOfMemoryError: unable to create native thread` — despite plenty of heap. - A Node.js process: `spawn EAGAIN` from `child_process`. - A Go service: `runtime: failed to create new OS thread`. - `docker exec` into the container fails, because the exec's own process cannot be created inside the same cgroup. That last one is the practical trap: at exactly the moment you want to look inside, you cannot get a shell in. Read the numbers from the outside instead — the `PIDS` column of `docker stats`, or `pids.current` and `pids.max` in the container's cgroup. ## Why a memory limit is not a substitute A common wrong answer is "`--memory` already covers it". Each task costs kernel memory for its stack and structures, so a fork bomb does eventually bump into a memory limit — but slowly, noisily, and after it has already consumed a large share of a host-wide resource. The number of PIDs on a Linux host is capped by `kernel.pid_max`; exhaust it and *nothing* on the host can start a new process, including your monitoring agent and your SSH session. `--pids-limit` fails the bomb at task N and leaves the rest of the machine untouched. It is the cheapest containment control Docker offers. ## Picking a value Measure rather than guess. Run the workload under realistic load and watch `pids.current`; take the peak and give it real headroom, because startup, a GC-heavy moment, or a connection storm can double the steady-state task count. A small event-loop service might peak at 25 tasks; a thread-per-request server or anything that shells out per job can be in the hundreds. Set the cap above the peak, not at it — a limit that trips only under production load is worse than no limit, because it looks like an application fault. ## The unrelated neighbour: nofile `--ulimit nofile=9216:16384` sets the soft and hard limits on **open file descriptors**, which is a different exhaustion story — sockets and files, not processes. It is genuinely per-process and it is the right tool for its own job; the failure mode is `EMFILE: too many open files` under connection load. Do not conflate the two: `nofile` for descriptors, `--pids-limit` for tasks. Daemon-wide defaults for ulimits live under the `default-ulimits` key in `daemon.json`; in Compose the per-service keys are `pids_limit:` and `ulimits:`. ## Putting it together The defensible answer is: cap tasks with `--pids-limit` because it is per-container, counts threads, and is enforced by the cgroup regardless of UID; use `--ulimit nofile` for descriptor exhaustion; and treat `--ulimit nproc` as a legacy knob whose per-UID, host-wide accounting makes it unreliable as container isolation.
- Do threads count toward `--pids-limit`, or only processes?Threads count. The pids cgroup controller counts tasks, and every thread is a task on Linux, so a JVM or a Go service with a large thread pool consumes the budget just as separate processes would. That is why a cap chosen from `ps` process counts is often far too low: measure `pids.current` under load instead, and leave headroom for startup and connection storms.
- A container has hit its pids limit and `docker exec` into it fails. How do you investigate?From outside. The exec needs to create a process inside the same cgroup, so it fails for the same reason the app does. Read the `PIDS` column of `docker stats`, or `pids.current` and `pids.max` from the container's cgroup on the host, and read the container's logs — `fork: Resource temporarily unavailable` or `unable to create native thread` confirms it. Then raise the cap deliberately or fix the leak.
- Why is `--memory` not an adequate defence against a fork bomb?It works eventually and badly. Each task costs kernel memory, so the bomb does hit a memory limit, but only after consuming a large share of the host's PID space, which is capped by `kernel.pid_max` and shared by every container and by the host itself. Once that is exhausted, nothing on the machine can start a process. `--pids-limit` fails the bomb at task N and leaves the host usable.
saying these in an interview costs you the question
- Thinks --ulimit nproc is a per-container process cap
- Believes only processes count and threads are free
- Sets --pids-limit 0 or -1 believing it hardens the container
- Assumes a memory limit already stops a fork bomb
- Confuses nofile (open descriptors) with nproc (processes)
- Picks a cap from guesswork, then blames the app when it trips