skip to content

On a shared host running dozens of PHP-FPM pools, how would you set process manager modes and limits so one busy site cannot starve the rest?

level: principalimportance: nice to knowfreq 18%

answer

  1. sum of max_children vs RAM
  2. ondemand for mostly idle pools
  3. process.max as a global cap
  4. reserved capacity for important sites
  5. overcommit is a deliberate bet

basics

~20 s

Give each site its own pool with a pm.max_children it may never exceed, use ondemand for mostly idle sites, keep the sum of caps within RAM or overcommit deliberately, and add process.max in php-fpm.conf as a global ceiling.

solid answer

~50 s

There is no single right answer; it is a capacity policy. One pool per site gives each a hard `pm.max_children`, so a busy site can occupy only its own workers. Most sites on a shared host are idle most of the time, so `pm = ondemand` with a modest `pm.process_idle_timeout` keeps their memory near zero, while the few busy or latency-sensitive sites get `dynamic` or `static` with reserved workers. The key decision is **overcommit**: if the sum of all caps times worker memory exceeds RAM, you are betting that not every site peaks at once. `process.max` in `php-fpm.conf` puts a global ceiling on workers across all pools, so the bet fails as queueing rather than swapping, but it is shared first-come-first-served, so reserve capacity for important sites by giving them static workers or a separate FPM instance.

code

ini · 14 lines
ini
; php-fpm.conf (global): hard ceiling across all pools
[global]
process.max = 120

; pool.d/shop.conf: critical client, reserved warm workers
[shop]
pm = static
pm.max_children = 20

; pool.d/blog-042.conf: one of many mostly idle sites
[blog-042]
pm = ondemand
pm.max_children = 6
pm.process_idle_timeout = 20s

go deeper

for a junior

Recall that each PHP-FPM pool has its own worker cap, so one site per pool stops a busy site from taking every worker.

for a middle

Explain ondemand for idle sites, why the sum of caps times worker memory matters, and what process.max limits.

for a senior

Show the failure modes: overcommit turning into swapping without a global cap, the global cap starving good sites, and caps that do not cover a shared database.

for a principal

Own the policy: which tier gets reserved capacity, how much overcommit is acceptable, and when a site moves to its own FPM instance or host.

## The problem A shared host runs many sites — agency client sites, a hosting plan's customers, internal tools — each with its own PHP code. In PHP-FPM each worker serves one request at a time and holds its own memory, so the resources that one site can take from the others are **workers** and **RAM**. A site with a traffic spike or a slow database, left unbounded, would fork workers until the host swaps, and every site slows down together. The goal is **isolation with good utilisation**, and the two pull in opposite directions. ## Building blocks - **One pool per site.** Each pool has its own `pm` style and `pm.max_children`, so a busy site can saturate only its own workers; other pools keep theirs. (Separate pools also allow separate Unix users and sockets, which is the pool configuration's concern.) - **`pm = ondemand`** for sites that are idle most of the time: zero workers when quiet, forks on demand, and `pm.process_idle_timeout` (default 10s) reaps them afterwards. - **`pm = dynamic` or `static`** for the few busy or latency-sensitive sites, so they keep warm workers. - **`process.max`** in the global `php-fpm.conf`: the maximum number of processes FPM forks across **all** pools (default `0`, no limit). The sample file says it was designed for controlling the global number of processes when using dynamic PM within many pools, and to use it with caution. ## The central decision: overcommit or not Let *W* be the memory of a busy worker. Two policies: | Policy | Rule | Outcome | |---|---|---| | **No overcommit** | Σ `pm.max_children` × *W* ≤ RAM for PHP | Any mix of peaks fits; most RAM idle most of the time; each site's cap is small | | **Overcommit** | Σ caps × *W* > RAM, bounded by `process.max` × *W* ≤ RAM | Better use of RAM; sites can burst higher; simultaneous peaks queue at the global cap | With overcommit, `process.max` is what turns "everyone peaks at once" from an out-of-memory event into queueing. But the global cap is **first come, first served**: the site that spikes first takes the headroom, and a well-behaved site can find FPM unable to fork a worker for it. ## Protecting what matters A principal-level answer names which sites deserve guarantees and gives them capacity nobody else can take: 1. **Tier the sites.** Paying or critical sites get `pm = static` (or `dynamic` with a solid `pm.min_spare_servers`) so their workers already exist and are not part of the global race — except when a worker is recycled by `pm.max_requests`, because its replacement counts against the same global cap; the long tail gets `ondemand` with small caps. 2. **Cap each site below its "blast radius".** A per-site cap of, say, 8 workers means one site can take at most 8 × *W* of memory, whatever its traffic. 3. **Separate FPM instances** when tiers must not share a global cap: a second master with its own `process.max` for the long tail, or a separate host or container for the heavy sites. 4. **Account for everything outside FPM:** database servers, cron jobs and CLI workers use RAM that FPM's caps do not count. ## Operating it - **Watch per-pool saturation.** For dynamic and ondemand pools, the log line `server reached pm.max_children setting` (or `max_children` for ondemand) identifies the site hitting its cap; static pools never log it, so watch their listen queue. - **Review caps with data.** Busy-worker memory differs per application; a large content-management site and a small API should not share a number. - **Mind cold starts.** Ondemand sites pay a fork after each idle spell; for a customer who complains about the first page load, a small `dynamic` pool with one spare worker may be worth its memory. - **Plan for the noisy neighbour you cannot cap:** a slow shared database makes every pool's workers wait longer, so per-site caps protect memory but not latency. ## Summary of the trade-off Strict per-site caps without overcommit give the strongest isolation and the worst utilisation; aggressive overcommit gives the best utilisation and turns simultaneous peaks into shared queueing. Most hosts land in between: ondemand small caps for the tail, reserved static capacity for the few that matter, and `process.max` as the safety net that keeps the host out of swap.

  • With process.max set, a critical site on a dynamic pool cannot grow during another site's spike. Why, and what fixes it?
    `process.max` is a single global limit across pools, taken first come, first served. The spiking site used the headroom, so FPM cannot fork for the critical pool even below its own `pm.max_children`. Give the critical site workers that already exist — `pm = static` or a higher `pm.min_spare_servers` — or run it under a separate FPM instance.
  • Why not make every pool on the shared host static?
    Static workers exist permanently, so every site holds its full `pm.max_children` in memory even when idle. On a host with dozens of mostly idle sites that wastes most of the RAM and forces small caps for everyone; ondemand lets idle sites release their memory so the busy ones can have more.

saying these in an interview costs you the question

  • One big pool for all sites is the most efficient setup
  • process.max reserves capacity for each pool
  • Per-pool caps also protect sites from a shared slow database
  • Making every pool static gives the best isolation at no cost