During a flash sale, a PHP site behind PHP-FPM returns intermittent 502s; how do you use FPM's status page and logs to find the cause?
answer
- 502 means the upstream failed, not PHP errors
- queue full vs worker killed vs crash
- status: active, queue, queue len
- FPM log: timed out, exited on signal
- slow log shows what workers wait on
basics
~20 sA 502 means the web server got no valid FastCGI response. Line up its timestamps with the FPM status page (all workers busy, queue near its length) and the FPM log (terminations, crashes), then read the slow log for what workers wait on.
solid answer
~50 sA 502 means the web server could not get a valid FastCGI response: it could not connect, or the connection was closed before a complete response. With PHP-FPM that usually has one of four causes. **Backlog overflow:** every worker busy and the socket's queue full, visible on the status page as `active processes` equal to `total processes` at the cap and `listen queue` near `listen queue len`, with a `server reached pm.max_children setting` warning for dynamic pools. **Killed requests:** `request_terminate_timeout` terminating workers, logged as `execution timed out (… sec), terminating`. **Crashes:** workers exiting on SIGSEGV or SIGBUS, logged as `exited on signal`. **FPM down or restarting**, or the kernel's OOM killer taking workers. Line these up with the 502 timestamps. Then open the slow log: its backtraces show what the busy workers were waiting on — a locked stock row, a slow payment call — which is usually the real fix.
code
bash · 6 lines# sample the pool every 2 seconds during the sale, for later correlation
while sleep 2; do
curl -s 'http://127.0.0.1/fpm-status?json' \
| jq -c '{t: (now | todate), active: ."active processes", total: ."total processes",
queue: ."listen queue", qlen: ."listen queue len", maxed: ."max children reached"}'
done >> /var/log/fpm-status-samples.loggo deeper
Recall that a 502 means the web server got no valid answer from PHP-FPM, and that FPM's log and status page are where to look.
Map each 502 cause — full backlog, terminated requests, crashes, FPM down — to the status field or log line that reveals it.
Correlate timestamps across web-server, FPM and slow logs, read status caveats correctly, and fix the waiting call rather than only raising limits.
Plan sale readiness: continuous FPM metrics, load tests on the hot path, timeout alignment between tiers, and capacity decided before the event.
## What a 502 means here In a web-server-plus-PHP-FPM setup, **502 Bad Gateway** comes from the web server, not from PHP. It means the server could not obtain a valid FastCGI response: it could not connect to the pool, or the connection closed before a complete response arrived. PHP errors in application code produce 500s (or whatever the application sends); a 502 points at FPM's process side. A 504 is different again: the web server waited longer than its own timeout. During a flash sale, four causes dominate: | Cause | Where you see it | |---|---| | Socket backlog full while every worker is busy | Status page: `active processes` = `total processes` at `pm.max_children`, `listen queue` near `listen queue len`; FPM log: `server reached pm.max_children setting` (dynamic/ondemand) | | Requests killed by `request_terminate_timeout` | FPM log: `child N, script '…' (request: "…") execution timed out (… sec), terminating` | | Worker crashes | FPM log: `child N exited on signal 11 (SIGSEGV …)` or SIGBUS, at WARNING level | | FPM restarting, stopped, or workers killed by the OS | FPM log gaps or reload notices; kernel log for the OOM killer | ## Step 1: line up the timestamps Take the 502 times from the web server's access log and put them next to: 1. **FPM's error log** (the global `error_log` in `php-fpm.conf`) — the warnings above carry the pool name, PID, script and request URI. 2. **Status page samples**, ideally scraped every few seconds as JSON or OpenMetrics during the sale. A single manual look after the fact shows only counters since start. 3. **The web server's error log**, which says whether it failed to connect or the upstream closed the connection early. Typical FPM log lines around a burst of 502s: ```text WARNING: [pool shop] server reached pm.max_children setting (40), consider raising it WARNING: [pool shop] child 8790, script '/srv/shop/public/index.php' (request: "POST /checkout/confirm") execution timed out (30.001204 sec), terminating WARNING: [pool shop] child 8790 exited on signal 15 (SIGTERM) after 612.310201 seconds from start ``` ## Step 2: read the status page correctly - `active processes` equal to `total processes`, with `total processes` at `pm.max_children`, means every worker was busy. - `listen queue` above 0 means connections were waiting; if it reaches `listen queue len`, new connections are refused and the web server reports 502. - `max children reached` rising (for dynamic and ondemand pools) confirms the cap was hit repeatedly. - **Caveats:** a Unix-socket pool reports `listen queue` 0 regardless, because FPM measures the queue only on TCP listeners; and the status page is served by the same busy workers unless `pm.status_listen` gives it its own hidden pool. ## Step 3: find out why workers are busy Saturation is a symptom; the cause is requests holding workers too long. The **slow log** answers that: with `request_slowlog_timeout` set (for example `3s`) and a `slowlog` file, FPM writes a PHP backtrace of every request still running after that time. During a sale the backtraces typically cluster on one call — a `SELECT … FOR UPDATE` on the stock row, a payment-provider HTTP request without a timeout, a session file lock — which is the thing to fix. ## Step 4: fix the right thing 1. **Waiting on a lock or remote call** → shorten it: add timeouts to outbound calls, narrow the locked section, move work after the response or into a queue. 2. **Genuinely CPU-bound with free memory** → raise `pm.max_children`, within memory and database limits. 3. **Terminations** → a `request_terminate_timeout` shorter than legitimate slow requests turns slowness into 502s; align it with real request times and with the web server's timeout. 4. **Crashes** → find the faulting extension; `emergency_restart_threshold` and `emergency_restart_interval` can make FPM reload automatically after a burst of SIGSEGV/SIGBUS exits, which masks the crash but can restore service. ## Preparing for the next sale - Keep `pm.status_path` scraped continuously and alert on sustained `listen queue > 0`. - Keep the slow log on in production with a threshold above normal page times. - Load-test the checkout path and watch the same fields before the event, not during it.
- What do emergency_restart_threshold and emergency_restart_interval do?They are global settings in `php-fpm.conf`, both 0 (off) by default. If `emergency_restart_threshold` workers exit with SIGSEGV or SIGBUS within `emergency_restart_interval`, FPM logs 'failed processes threshold … is reached, initiating reload' and reloads itself. The sample configuration mentions working around accidental corruption of an accelerator's shared memory; it restores service but does not fix the crash.
- Why can request_terminate_timeout itself produce 502s during a sale?When a request runs longer than `request_terminate_timeout`, the FPM master kills the worker with SIGTERM. The web server's connection closes without a complete response, which it reports as 502. If the limit is shorter than checkout requests take under load, slowness turns into errors; align it with realistic request times.
saying these in an interview costs you the question
- A 502 means the PHP code threw an uncaught exception
- Raising pm.max_children is always the fix for 502s under load
- One look at the status page after the sale shows what happened during it
- A Unix-socket pool with listen queue 0 cannot be saturated
- Emergency restart settings fix segmentation faults