How do you choose `--pids-limit` and file-descriptor ulimit defaults for a whole container fleet?
answer
- What resource are you really protecting
- Host-wide, not per-service, exhaustion
- Measure the peak before choosing
- A small number of tiers, one default
- No audit mode for a cgroup ceiling
basics
~20 sTreat them as host-resource protection, not per-service tuning. Measure peak task and descriptor use across the estate, define two or three tiers rather than per-service numbers, set the tier as the platform default, and make a higher tier an explicit, reviewed choice.
solid answer
~50 sStart from what you are protecting: PIDs are a host-wide kernel resource capped by `kernel.pid_max`, so one container's runaway forking can stop every other container and the host's own agents from starting a process. The cap exists to bound one container's blast radius, not to right-size it. Measure first — collect peak `pids.current` and descriptor use per workload under real load — then group the estate into a small number of tiers (an event-loop service, a thread-per-request server, a build agent) rather than negotiating a number per service. Ship the tier as the default in the platform's run template, with headroom well above the observed peak, because a cap that trips only under production load reads as an application bug. Higher tiers should be a reviewed opt-in with a named owner. Revisit the numbers when runtimes change, not never.
code
json · 9 lines{
"default-ulimits": {
"nofile": {
"Name": "nofile",
"Soft": 9216,
"Hard": 16384
}
}
}go deeper
Know that these are ceilings set on each container, and that hitting one shows up as a failure to create a process or open a file rather than as a normal application error.
Explain what each limit counts and where it is enforced, and why headroom matters: startup and reconnect storms push task counts well above steady state.
Show you can pick and defend a value from measured peaks, recognise the runtime error strings a tripped cap produces, and diagnose from outside the container when exec itself can no longer start.
Own the framing and the process: these caps bound one container's share of a host-wide kernel resource, so express them as a small set of reviewed tiers with a platform default, and refuse the request to delete rather than raise one.
## Decide what the limit is for The first thing to get right is the purpose, because it determines the number. A per-container task cap is **not** capacity planning for the service. It is a bound on how much of a shared, host-wide kernel resource a single failure can consume. Linux caps the total number of processes by `kernel.pid_max`; exhaust it and nothing on that host can fork — not the other containers, not the log shipper, not the operator's shell. The same reasoning applies to descriptors: the host has its own ceilings, and a container that leaks sockets without a `nofile` limit competes with everything else on the machine. Once framed that way, the target changes. You are not looking for the smallest value the service can survive; you are looking for a value comfortably above anything healthy and comfortably below anything ruinous. That is a much easier number to agree on, and it is why per-service haggling is the wrong process. ## Measure before you legislate Collect, across the estate under real load: peak `pids.current` per container (the `PIDS` column of `docker stats` or the cgroup file), and peak open descriptors per process. Two things usually fall out. First, the distribution is bimodal rather than smooth — a Node.js worker consuming a queue sits around 20 to 40 tasks, while a thread-per-request JVM or a build agent that shells out per job runs in the hundreds. Second, the peaks are not at steady state: startup, a GC storm, a reconnect stampede after a dependency blips, and a graceful shutdown all spike the task count. That is the argument for headroom. If a fraud-scoring endpoint peaks at 118 tasks, a cap of 128 is a trap: it will hold for months and then trip during the exact incident when everything else is also going wrong, and it will present as an application failure — `unable to create native thread`, `spawn EAGAIN` — sending the on-call engineer down the wrong path for an hour. A multiple of the observed peak, not a margin of a few percent, is the right instinct for a control whose job is blast radius rather than efficiency. ## Tiers, not snowflakes The operationally sustainable design is two or three named tiers, each with a task cap and a descriptor limit, and a default tier applied by the platform's run template. Something like: a standard tier that covers the large majority of services, a high-concurrency tier for thread-heavy servers and connection-fanout proxies, and a build/agent tier for workloads that spawn processes as their job. Every service gets the default unless it asks. This matters more than the numbers. Per-service values become unowned folklore within a year: nobody remembers why one team's worker runs with 4096 and the identical service next to it runs with 512. Tiers are reviewable, comparable, and can be changed for everyone at once when the runtime underneath changes. ## Rolling it out without breaking things A cgroup limit has no audit mode — it either allows the task or it does not. So the observation phase has to happen before enforcement, using measurement rather than a dry-run flag: instrument `pids.current` and descriptor counts across the fleet for long enough to include a weekly peak and at least one incident, and place the tier ceilings above what you saw. Then apply the default to new workloads first, and to existing ones on their next deploy so the change lands with a rollback path attached. When a service does trip a cap, the diagnostic path must be short. The error strings are distinctive; make sure they are searchable, and make sure the platform emits the container's task count alongside CPU and memory so the answer is one dashboard away. Raising a cap should be routine and reviewed; deleting the cap should not. ## Descriptors deserve their own thought `nofile` is a different failure story — sockets and files, presenting as `too many open files` under connection load — and it is worth setting deliberately rather than inheriting whatever the daemon defaults to. Very large hard limits are not free: some programs still walk the descriptor table on startup or before exec, and an enormous ceiling makes that measurably slower. Pick a limit sized to the workload's connection fanout with headroom, set the fleet default in `daemon.json` under `default-ulimits`, and let a tier override it. ## What you own as a lead Three things. The framing — these caps protect the host, so they are not negotiable away entirely. The process — measured tiers with an owner and a review, rather than per-service numbers nobody can justify. And the honesty — say plainly that these controls bound resource exhaustion and nothing else, so nobody mistakes a pids cap for a security boundary or for a substitute for fixing the leak that keeps hitting it.
- Why not simply set the cap to the observed peak plus ten percent?Because the peak you measured is not the peak that matters. Startup, reconnect storms after a dependency blips, and graceful shutdown all spike task counts, and they cluster during incidents. A cap that tight will trip exactly when everything else is failing and will present as an application bug, costing an hour of misdirected diagnosis. The control bounds blast radius, so generous headroom costs you nothing.
- A team asks for their task cap to be removed entirely because they keep hitting it. What do you say?Raise the tier, do not remove the cap — and treat repeated hits as a defect signal rather than a limit problem. An uncapped container can exhaust the host's PID space, which takes down every neighbour and the host's own agents. If the workload genuinely spawns thousands of tasks, it belongs in a tier that says so, with an owner, so the next person can see the decision instead of finding an unexplained absence.
- How do you sanity-check that a fleet default is not simply too generous to matter?Compare it against what the host can absorb: the sum of caps for the containers you actually pack onto a node, against `kernel.pid_max` and the host's descriptor ceilings. The cap has to be low enough that one container hitting it leaves the host able to schedule work and an operator able to log in. If a single container's ceiling approaches the host's, you have written down a number, not a control.
saying these in an interview costs you the question
- Treats the cap as capacity planning for the service
- Negotiates a bespoke number with every service team
- Sets caps at the observed peak with no headroom
- Removes the limit whenever a team complains
- Assumes memory limits already bound process explosion
- Calls a pids cap a security boundary