An Apache httpd server stops keeping up: requests queue for seconds and the error log repeats 'server reached MaxRequestWorkers setting, consider raising the MaxRequestWorkers setting', while the machine's CPU is mostly idle. How do you work out whether raising it is actually the right fix?
answer
- saturation is a symptom, not a cause
- idle CPU means waiting, not computing
- classify what the workers are waiting on
- the ceiling multiplied by per-worker memory
- more workers means more load downstream
basics
~20 sIdle CPU with every worker busy means workers are blocked waiting, not computing. Find what they wait on — a slow backend, or idle keep-alive connections under prefork or worker — before raising the ceiling, because raising it multiplies memory use and downstream load.
solid answer
~50 sThat message means every worker slot is occupied, so new connections wait in the listen queue. Idle CPU tells you the workers are not computing — they are blocked. So the diagnostic question is what they are blocked on. Look at the worker states: if most are in keep-alive rather than processing, you are on `mpm_prefork` or `mpm_worker` and are paying a worker for every idle persistent connection, which `mpm_event` or a shorter `KeepAliveTimeout` fixes without any extra memory. If most are actively processing, they are waiting on something downstream — a database, an upstream service, a slow disk — and the honest fix is there, or in a timeout that stops one slow dependency from consuming the whole pool. Raising `MaxRequestWorkers` is only right when you have headroom for it: the ceiling multiplied by resident memory per worker must fit in RAM without swapping, and the backend must survive the higher concurrency you are about to send it.
go deeper
Understand that the message means every worker is busy and new connections are queueing, and that the config value alone does not tell you why they are busy.
Explain the occupancy arithmetic — workers times requests per second is bounded by how long each request holds its worker — and why idle CPU implies blocking rather than compute.
Demonstrate the diagnostic order: classify worker states, identify the wait, fix that, and treat raising the ceiling as a last step constrained by memory and downstream capacity.
Frame the worker ceiling as deliberate admission control for everything behind it, and set the policy for timeouts and isolation so one slow dependency cannot consume a shared web tier.
## Read what the message actually says "Server reached MaxRequestWorkers setting" is not a diagnosis, it is a symptom: every worker slot Apache is allowed to have is currently in use, so newly accepted connections sit in the kernel's accept queue until a worker frees up. Clients experience that as seconds of latency before any byte of the response, with the application logging nothing wrong — because from the application's point of view, the requests have not started yet. The second half of the observation does most of the work. If CPU is near idle while every worker is occupied, the workers are not doing computation. They are waiting. Everything else follows from identifying what they wait on, and there are only a few candidates. ## Candidate one: idle keep-alive connections Under `mpm_prefork` and `mpm_worker`, a worker stays bound to its connection for the connection's whole life, including the gaps between requests on a persistent connection. With a five-second `KeepAliveTimeout` and browsers holding connections open, a large share of your workers can be occupied by clients that are not asking for anything. Occupancy looks like saturation and CPU looks idle, because nothing is happening. The check is worker state: Apache's status page distinguishes workers that are sending a reply from those sitting in keep-alive and those merely waiting for a connection. If keep-alive dominates, the fix is not more workers — it is `mpm_event`, which parks idle connections on a listener thread instead of a worker, or a shorter `KeepAliveTimeout` as a cruder version of the same idea. Both cost nothing in memory. ## Candidate two: a slow dependency If most workers are actively processing a request, they are blocked inside your handler — on a database query, an upstream HTTP call, a lock, or slow storage. Concurrency arithmetic explains the cliff precisely: at 400 workers and 20 ms per request you serve about 20,000 requests per second; when a dependency slows to 2 seconds, the same 400 workers serve 200 per second. Nothing about Apache changed. The pool drained because occupancy time grew a hundredfold. Here raising `MaxRequestWorkers` is close to the worst available move: you send more concurrency into the thing that is already too slow, and you convert a partial outage into a total one. The right responses are fixing or bounding the dependency — a timeout so a hung backend releases its worker, and enough isolation that one slow endpoint cannot consume the entire pool. ## Candidate three: slow clients A worker also stays occupied while it reads a request body or writes a response over a slow link. Large uploads over poor connections behave exactly like a slow backend from the pool's point of view. This is the classic argument for putting a buffering reverse proxy in front, so the slow byte-shuffling happens somewhere connections are cheap and Apache only sees complete, fast requests. ## If raising it *is* right, size it honestly Sometimes the answer really is that the ceiling is too low for the traffic. Then it is bounded by two things, not by optimism. Memory: `MaxRequestWorkers × resident memory per worker` must fit in RAM with room to spare. Under prefork with an in-process interpreter, a child can be tens of MB, so the ceiling is often far lower than people expect and the moment you exceed RAM, swapping makes every request slower and the queue grows faster. Under a threaded MPM the per-worker cost is a thread stack rather than a process image, which is precisely why the threaded MPMs allow much higher ceilings on the same machine. Downstream capacity: whatever number you choose is the maximum concurrent load you will place on the database and every upstream service. A web tier ceiling is also a backend admission control, whether or not anyone designed it that way. Raising it without checking the backend's connection limits simply moves the queue somewhere with worse failure behaviour. ## The order to work in Confirm the effective ceiling from the startup log rather than the config file, since an inconsistent `ServerLimit` may already have lowered it. Then classify worker occupancy — keep-alive, reading, writing, processing. Then fix the class you found: MPM or `KeepAliveTimeout` for idle connections, timeouts and dependency work for slow backends, buffering for slow clients. Raise the ceiling last, and only with the memory arithmetic and the downstream limit written down next to the new number.
- How would you tell an idle keep-alive problem from a slow-backend problem in this situation?Look at what the workers are doing rather than how many are busy. Apache's status page separates workers sending a reply from workers sitting in keep-alive. Keep-alive dominance points at the MPM and KeepAliveTimeout; a pool full of workers actively processing points downstream, and you confirm it by correlating with backend latency for the same window.
- Why can raising MaxRequestWorkers make the outage worse rather than better?Two ways. If the ceiling times per-worker memory exceeds RAM, the box swaps and every request slows, so the queue grows faster than the extra workers drain it. And if workers are blocked on a saturated backend, more workers means more concurrent pressure on the thing already failing, turning a degraded service into a fully broken one.
- What role does a timeout play in stopping one slow dependency from consuming the whole worker pool?Without a bound, a hung dependency holds each worker indefinitely and the pool drains until nothing else can be served, including endpoints that do not touch that dependency. A timeout converts an unbounded hold into a bounded one: workers return, unaffected traffic keeps flowing, and the failure stays confined to the callers of the slow path.
saying these in an interview costs you the question
- Raises MaxRequestWorkers immediately because the log suggested it
- Reads idle CPU as proof the server has spare capacity
- Ignores memory per worker when choosing a new ceiling
- Forgets the ceiling also caps concurrent load on the database
- Assumes switching to mpm_event fixes a slow backend