skip to content

During a ticket on-sale, PHP-FPM logs 'server reached pm.max_children setting (40), consider raising it'; what happens to incoming requests, and should you just raise the limit?

level: seniorimportance: should knowfreq 45%

answer

  1. every worker busy at once
  2. connections wait in the listen backlog
  3. web server times out: 502 or 504
  4. logged once until the condition clears
  5. static pools never log it

basics

~20 s

The pool is at pm.max_children with idle workers exhausted or below the spare minimum, so new connections queue in the listen backlog until a worker frees up or the web server times out. Raise the cap only with spare memory, CPU and database capacity.

solid answer

~50 s

The warning means FPM wanted another worker — in a dynamic pool because idle workers fell below `pm.min_spare_servers`, in an ondemand pool because a connection found no idle worker — but the pool already had `pm.max_children`. Once the last idle workers are taken, new FastCGI connections are not rejected by FPM; they wait in the kernel's **listen backlog** of the pool socket until a worker frees up. If they wait too long, the web server's FastCGI timeout returns a 504, and when the backlog itself is full new connections fail, typically showing up as 502s. Raising the limit helps only if the server has free memory and CPU and the database can take more connections; otherwise it trades queueing for swapping or database overload. First ask why workers are busy: slow queries, external calls or locks make each request hold a worker longer. Static pools never log this warning.

go deeper

for a junior

Recall that the warning means the PHP-FPM pool hit its worker cap, so new requests soon had to wait for a free worker.

for a middle

Explain the listen backlog, how waiting turns into 504 and 502 at the web server, and how dynamic, ondemand and static pools differ in reporting it.

for a senior

Diagnose before tuning: check memory, CPU and database headroom, find why workers are held, and use pool splits or warm spares for a planned spike.

for a principal

Decide on capacity for known traffic events — waiting rooms, pre-scaling and pool isolation — rather than reacting to a log line mid-sale.

## What the warning means A PHP-FPM worker handles one request at a time, and `pm.max_children` caps the number of workers in a pool. The master logs, at WARNING level: ```text WARNING: [pool tickets] server reached pm.max_children setting (40), consider raising it ``` when it wants another worker and the pool already has the maximum. When exactly depends on the process manager: - **`pm = dynamic`** — the once-a-second maintenance pass sees fewer idle workers than `pm.min_spare_servers` while the pool is already at `pm.max_children`. - **`pm = ondemand`** — a connection arrives with no idle worker and the pool is at its cap. The message text here reads `server reached max_children setting`, without the `pm.` prefix. - **`pm = static`** — never logs it: the pool is always at its cap and there is no spawning logic that could hit it. Saturation of a static pool is visible only through the listen queue and response times. The message is logged **once** and not repeated until FPM has been able to fork a worker for that pool again, so one line can stand for minutes of saturation. FPM also counts these events in the status page's `max children reached` field. ## What happens to the requests FPM does not reject anything itself. Each pool listens on a socket (Unix or TCP), and a connection that no worker has accepted sits in the kernel's **listen backlog** for that socket, whose size the pool's `listen.backlog` setting requests. 1. A worker finishes a request and accepts the next queued connection, so users see **added latency**. 2. If a connection waits longer than the web server's FastCGI read timeout, the web server gives up and returns **504 Gateway Timeout**; the PHP request may still run later. 3. If the backlog is full, new connections fail or stall, and the web server typically reports **502 Bad Gateway**. On a ticket site this is the classic on-sale failure: the page hangs, then errors, and users retry — adding more load. ## Should you raise pm.max_children? Only when all of these hold: | Check | Why | |---|---| | Free memory at peak ≥ extra workers × busy-worker memory | Otherwise the server swaps or the OOM killer strikes | | CPU is not saturated | CPU-bound requests do not get faster with more processes | | The database can accept more connections and load | Each extra busy worker may add a query stream | | Downstream services can take more concurrent calls | Payment or seat-reservation APIs have their own limits | If memory and CPU are idle while workers are all busy, the workers are **waiting** — on slow queries, row locks during seat reservation, or outbound HTTP calls — and more workers may help, up to the database's limits. If CPU is maxed out, more workers only lengthen every request. ## Better fixes than a bigger number - **Shorten the time each request holds a worker:** index the slow queries, add timeouts to outbound calls, move e-mail and PDF generation to a queue. - **Split pools:** give checkout its own pool so a flood of browsing requests cannot occupy every worker. - **Warm up:** with `pm = dynamic`, a higher `pm.min_spare_servers` before a planned sale avoids spawning delays; static removes them entirely. - **Protect the queue:** rate-limit or put a waiting room in front of the sale at the web tier, so the backlog does not grow without bound. - **Scale out:** add application servers when one host's memory is the real limit. Finding which requests are slow belongs to FPM's slow log and status page; the decision here is whether the cap or the request duration is wrong. ## A related warning from dynamic pools A dynamic pool can also log `seems busy (you may need to increase pm.start_servers, or pm.min/max_spare_servers), spawning N children…`. That one means the master has had to spawn at a rate of 8 or more workers per second to keep up with falling idle counts — the spare settings are too small for the burst — while the cap has **not** yet been reached.

  • A static pool serves slow responses at peak, but the FPM log has no max_children warning. Is the pool unsaturated?
    Not necessarily. With `pm = static` the pool always runs `pm.max_children` workers and has no spawning logic, so FPM never logs that warning. Saturation of a static pool shows up as a growing listen queue on the status page and as rising response times, not as a log line.
  • Users see 504s while PHP-FPM keeps running. Where did those requests go?
    They waited in the pool socket's listen backlog because all workers were busy, and the web server's FastCGI timeout expired first, so it answered 504. A queued request may still be accepted by a worker later and run to completion even though nobody is waiting for the answer.

saying these in an interview costs you the question

  • FPM returns an error to the client as soon as all workers are busy
  • The warning appears once per rejected request
  • A static pool logs the warning like dynamic pools do
  • Doubling pm.max_children always fixes the on-sale slowdown
  • The seems busy warning means pm.max_children has been reached