skip to content

How do `docker run --pids-limit` and `--ulimit nproc` differ as defences against a fork bomb in a container?

level: middleimportance: should knowfreq 38%

answer

  1. Two different kernel mechanisms, one goal
  2. One is a cgroup, one is a rlimit
  3. What exactly does each one count
  4. Per container, or per UID host-wide
  5. Threads are tasks too

basics

~20 s

--pids-limit sets the container cgroup's pids.max, capping every process and thread inside that one container; further clone() calls fail with EAGAIN. --ulimit nproc sets a per-UID kernel limit counted across the whole host, so it is the weaker and more surprising control.

solid answer

~40 s

`docker run --pids-limit 384` writes `pids.max` for the container's cgroup, so the kernel counts every task — processes *and* threads — in that container and refuses to create the next one with `EAGAIN`. It is a true per-container cap, which is what you want against a fork bomb or a runaway thread pool: the container degrades, the host does not run out of PIDs. `--ulimit nproc=384` sets `RLIMIT_NPROC`, which the kernel counts per UID across the whole host — two containers running as the same UID share one budget, a container that can change UID escapes it, and processes on the host under that UID count too. Use `--pids-limit` as the control and treat `--ulimit nofile` (open file descriptors) as a separate, unrelated knob.

code

bash · 7 lines
bash
docker run -d --name queue-worker \
  --pids-limit 384 \
  --ulimit nofile=9216:16384 \
  queue-worker:3.2

docker inspect -f '{{.HostConfig.PidsLimit}}' queue-worker
docker stats --no-stream --format '{{.Name}} {{.PIDs}}' queue-worker

go deeper

for a junior

Know that a container can be told how many processes it may create, and that the flag on docker run is --pids-limit. Recognise Resource temporarily unavailable as a fork failure rather than a disk or permission problem.

for a middle

Explain the mechanics: the pids cgroup counts tasks including threads and fails clone with EAGAIN, whereas RLIMIT_NPROC is counted per UID across the host. Say why that difference makes one a container boundary and the other not.

for a senior

Show you can operate it: pick a cap from measured pids.current peaks with headroom, recognise the JVM and runtime error strings it produces, and diagnose from outside when docker exec itself can no longer start.

for a principal

Frame it as host-resource protection rather than per-service tuning. Be ready to justify a fleet default, the size of the headroom, and why exhausting a host-wide resource such as the PID space is a different class of incident from one container degrading.

## Two limits that look alike and are not Both flags cap "how many", both are set on `docker run`, and both are reached for when someone says "stop a container forking forever". They are enforced by completely different kernel machinery. **`--pids-limit N`** configures the pids cgroup controller for that container's cgroup — `pids.max`. The controller counts *tasks*, and in Linux a thread is a task, so a JVM with 200 threads consumes 200 of the budget just as 200 separate processes would. When the cgroup is at its limit, `fork()` and `clone()` fail with `EAGAIN`. The scope is exactly the container: its own processes, nothing else, no matter what UID they run as. `--pids-limit -1` (or 0) means unlimited. **`--ulimit nproc=N`** sets the POSIX resource limit `RLIMIT_NPROC` on the container's initial process, which children inherit. The kernel enforces it **per real UID, counted across the whole host**, not per container. That produces three surprises worth naming in an interview: 1. Two containers whose processes run as the same UID — very common, since UID 1000 or UID 0 is the default in countless images — draw from a single shared budget. One noisy container makes the other fail to fork. 2. Processes running as that UID *outside* any container count toward the same total. 3. A container process that can change UID (the default container root holds `CAP_SETUID`) simply moves to a different UID and gets a fresh budget. So `nproc` is not a container boundary. `--pids-limit` is. ## What hitting the limit looks like The symptom is rarely "fork bomb". It is a service that stops making progress in a way that reads like an application bug: - Shell and libc: `fork: Resource temporarily unavailable`. - A JVM: `java.lang.OutOfMemoryError: unable to create native thread` — despite plenty of heap. - A Node.js process: `spawn EAGAIN` from `child_process`. - A Go service: `runtime: failed to create new OS thread`. - `docker exec` into the container fails, because the exec's own process cannot be created inside the same cgroup. That last one is the practical trap: at exactly the moment you want to look inside, you cannot get a shell in. Read the numbers from the outside instead — the `PIDS` column of `docker stats`, or `pids.current` and `pids.max` in the container's cgroup. ## Why a memory limit is not a substitute A common wrong answer is "`--memory` already covers it". Each task costs kernel memory for its stack and structures, so a fork bomb does eventually bump into a memory limit — but slowly, noisily, and after it has already consumed a large share of a host-wide resource. The number of PIDs on a Linux host is capped by `kernel.pid_max`; exhaust it and *nothing* on the host can start a new process, including your monitoring agent and your SSH session. `--pids-limit` fails the bomb at task N and leaves the rest of the machine untouched. It is the cheapest containment control Docker offers. ## Picking a value Measure rather than guess. Run the workload under realistic load and watch `pids.current`; take the peak and give it real headroom, because startup, a GC-heavy moment, or a connection storm can double the steady-state task count. A small event-loop service might peak at 25 tasks; a thread-per-request server or anything that shells out per job can be in the hundreds. Set the cap above the peak, not at it — a limit that trips only under production load is worse than no limit, because it looks like an application fault. ## The unrelated neighbour: nofile `--ulimit nofile=9216:16384` sets the soft and hard limits on **open file descriptors**, which is a different exhaustion story — sockets and files, not processes. It is genuinely per-process and it is the right tool for its own job; the failure mode is `EMFILE: too many open files` under connection load. Do not conflate the two: `nofile` for descriptors, `--pids-limit` for tasks. Daemon-wide defaults for ulimits live under the `default-ulimits` key in `daemon.json`; in Compose the per-service keys are `pids_limit:` and `ulimits:`. ## Putting it together The defensible answer is: cap tasks with `--pids-limit` because it is per-container, counts threads, and is enforced by the cgroup regardless of UID; use `--ulimit nofile` for descriptor exhaustion; and treat `--ulimit nproc` as a legacy knob whose per-UID, host-wide accounting makes it unreliable as container isolation.

  • Do threads count toward `--pids-limit`, or only processes?
    Threads count. The pids cgroup controller counts tasks, and every thread is a task on Linux, so a JVM or a Go service with a large thread pool consumes the budget just as separate processes would. That is why a cap chosen from `ps` process counts is often far too low: measure `pids.current` under load instead, and leave headroom for startup and connection storms.
  • A container has hit its pids limit and `docker exec` into it fails. How do you investigate?
    From outside. The exec needs to create a process inside the same cgroup, so it fails for the same reason the app does. Read the `PIDS` column of `docker stats`, or `pids.current` and `pids.max` from the container's cgroup on the host, and read the container's logs — `fork: Resource temporarily unavailable` or `unable to create native thread` confirms it. Then raise the cap deliberately or fix the leak.
  • Why is `--memory` not an adequate defence against a fork bomb?
    It works eventually and badly. Each task costs kernel memory, so the bomb does hit a memory limit, but only after consuming a large share of the host's PID space, which is capped by `kernel.pid_max` and shared by every container and by the host itself. Once that is exhausted, nothing on the machine can start a process. `--pids-limit` fails the bomb at task N and leaves the host usable.

saying these in an interview costs you the question

  • Thinks --ulimit nproc is a per-container process cap
  • Believes only processes count and threads are free
  • Sets --pids-limit 0 or -1 believing it hardens the container
  • Assumes a memory limit already stops a fork bomb
  • Confuses nofile (open descriptors) with nproc (processes)
  • Picks a cap from guesswork, then blames the app when it trips

context