What does HAProxy do with a request when every server in a backend has reached the `maxconn` set on its `server` line, and how does that limit differ from `maxconn` in the global section?
answer
- the same word at three scopes
- excess waits rather than failing
- the wait needs its own bound
- overload becomes latency you can see
- cap it near the worker count
basics
~20 sHAProxy holds the request in the backend's queue until a server slot frees, bounded by timeout queue, after which the client gets 503. Server-line maxconn caps concurrent connections to one server; global maxconn caps them for the whole process.
solid answer
~50 s`maxconn` on a `server` line caps how many connections HAProxy will hold to that server at once. When every server in the backend is at its cap, HAProxy does not reject the request and does not overshoot the cap — it parks the request in the backend's queue and dispatches it as soon as a slot frees. `timeout queue` bounds that wait, falling back to `timeout connect` when it is not set, and a request that times out in the queue gets a 503. `maxconn` in `global` is a different thing entirely: a process-wide ceiling on concurrent connections, related to the file-descriptor limit, above which HAProxy stops accepting and connections pile up in the kernel's accept queue. The point of the server-level limit is to turn backend overload into bounded, measurable waiting at the proxy instead of thread-pool exhaustion and timeouts inside the application.
go deeper
Know that maxconn on a server line limits concurrent connections to that server, and that the same keyword in the global section is a process-wide ceiling, not the same thing.
Explain the queue: excess requests wait rather than fail, timeout queue bounds the wait and falls back to timeout connect, and expiry produces a 503. Know where queue depth and queue time are visible.
Justify capping concurrency in front of a bounded worker pool, size the cap from the pool, and read queue time in the logs to tell proxy-side waiting apart from slow application responses.
Own admission control as a platform policy: where load is shed versus queued, how the queue timeout fits inside the caller's timeout budget, and what a queue tells you about capacity rather than about configuration.
## Three different maxconn settings The keyword appears at three scopes and means something different at each: - **`global maxconn`** — the whole process. It bounds total concurrent connections and is tied to the file-descriptor limit HAProxy computes for itself. When it is reached, HAProxy stops accepting; new connections sit in the kernel's listen backlog and clients see latency rather than an error, until the backlog itself fills. - **`maxconn` in a frontend** — a per-frontend ceiling, useful for stopping one entry point from consuming the whole process budget. - **`maxconn` on a `server` line** — a per-server concurrency limit, and the one that produces queueing. ``` backend app balance leastconn timeout queue 5s server app1 10.0.0.11:8080 check maxconn 60 server app2 10.0.0.12:8080 check maxconn 60 ``` ## What happens at the limit With all servers at their cap, HAProxy holds the request in the **backend queue**. It is not rejected, and the cap is not exceeded. When any server finishes a request and frees a slot, the oldest queued request is dispatched to it. Because the wait happens at the proxy, it is visible: the stats output exposes the current and maximum queue depth per server and per backend, and the log's queue-time field shows how long each request actually waited. There is a second, per-server queue for requests that are pinned to one server by persistence, since those cannot be served by anyone else. `maxqueue` on the `server` line caps how many may wait there; beyond it, requests are sent elsewhere rather than waiting indefinitely, which only matters when persistence is in play. ## The timeout that bounds the wait `timeout queue` limits how long a request may sit in the queue. If it is not set, HAProxy falls back to `timeout connect`, which is usually far too short to be a deliberate choice — set `timeout queue` explicitly whenever you set a server `maxconn`. On expiry, the client gets a 503 and the log's termination state records that the request died waiting rather than at a server. That pairing is the whole design. `maxconn` decides how much concurrency the backend may see; `timeout queue` decides how much waiting you are willing to sell to a client before admitting defeat. Setting one without the other gives you either unbounded queueing or accidental instant failure. ## Why limit concurrency at the proxy at all Most application servers have a bounded worker pool and degrade badly past it: threads pile up, memory grows, garbage collection or context switching eats the CPU, and response time climbs for *every* in-flight request rather than just the excess. Capping concurrency in front of the pool converts that collapse into a queue. The servers keep running at the concurrency they handle best, the excess waits somewhere you can measure, and `timeout queue` puts a ceiling on how long anyone waits before being told no. A reasonable starting point is the backend's own worker count — if an application server runs 50 workers, `maxconn 50` on its `server` line means HAProxy never asks it to hold more work than it has workers for. Then watch queue depth: persistently non-zero means the pool is genuinely undersized, and brief spikes during traffic bursts are exactly what the mechanism is for. ## Dynamic limits `minconn` and `fullconn` scale the per-server limit with load rather than pinning it: with `minconn` and `maxconn` both set on a server, the effective limit slides between them in proportion to how close the backend's total session count is to `fullconn`. That keeps per-server concurrency low when traffic is light — better for connection reuse and for latency — and lets it rise under load. It is a refinement, not a starting point; get the static cap and `timeout queue` right first. ## Reading the symptoms When users report slowness and the application's own latency looks fine, queue time is the first thing to check: the request spent its life waiting at the proxy, and the server never saw it until late. When users get 503s during bursts with healthy servers, look at whether `timeout queue` expired. Both are proxy-side facts that the application's metrics cannot show you, which is exactly why the queue is a feature rather than a hidden cost.
- What should timeout queue be set to, and what happens if you leave it out?Set it explicitly to the longest wait you are willing to sell a client — often a second or two, well inside the caller's own timeout. Left out, it falls back to `timeout connect`, which is typically a few seconds chosen for TCP setup rather than for queueing, so requests fail sooner or later than anyone intended. On expiry the client gets a 503 and the termination state records that it died in the queue.
- How would you pick a value for a server line's maxconn?Start from the backend's real concurrency limit — its worker or thread-pool size — so HAProxy never hands a server more simultaneous work than it can process. Then watch queue depth and queue time under load: persistent queueing means the pool is undersized, while short bursts of queueing are the mechanism doing its job.
- How does this differ from what global maxconn does when it is reached?Global maxconn stops HAProxy accepting new connections at all; they wait in the kernel's listen backlog and there is no HAProxy-level queue, no queue timeout and no 503 — clients just see connect latency, then failures once the backlog fills. The server-level limit is a graceful admission control with a visible queue; the global one is a hard process ceiling tied to file descriptors.
saying these in an interview costs you the question
- Thinks requests beyond server maxconn are rejected immediately
- Treats global and server-line maxconn as the same limit
- Sets a server maxconn without setting timeout queue
- Assumes HAProxy will exceed maxconn under pressure
- Says queueing at the proxy is always worse than passing load through