You need to take one server out of an HAProxy backend for a deploy without dropping in-flight requests and without reloading HAProxy. What does the Runtime API let you do, what must the config expose for it, and how do the DRAIN and MAINT states differ?
answer
- change state without touching the config file
- the control channel must be declared first
- admin level, and file permissions are the authz
- one state honours persistence, one does not
- in-memory only unless you persist it
basics
~20 sHAProxy's Runtime API, exposed through a stats socket at admin level, accepts set server <backend>/<server> state drain to stop new sessions while existing ones finish. MAINT forces the server fully out and stops its health checks; DRAIN keeps checking and still honours persistence.
solid answer
~50 sI declare a `stats socket` in `global` at `level admin`, then drive it with socat: `echo "set server web/app1 state drain" | socat stdio /var/run/haproxy/admin.sock`. DRAIN means the server takes no new load-balanced sessions, but connections already established keep running to completion and requests carrying persistence are still served; health checks continue, and the stats page shows the server as DRAIN. I watch the current-session count fall to zero with `show stat`, then deploy, then `set server web/app1 state ready`. MAINT is the harder state: the server is forced out administratively, health checks stop entirely, and persistence no longer keeps traffic on it — but note that neither state kills established connections, so `shutdown sessions server web/app1` is what you use if you must cut them. The trap is that everything set through the Runtime API is in-memory only: without `server-state-file` plus `load-server-state-from-file`, the next reload silently puts the drained server straight back into rotation.
code
bash · 17 linesSOCK=/var/run/haproxy/admin.sock
# take app1 out of rotation for new traffic, keep in-flight work alive
echo "set server web/app1 state drain" | socat stdio "$SOCK"
# wait until it has no current sessions left (scur is field 5 of show stat)
until [ "$(echo 'show stat' | socat stdio "$SOCK" | awk -F, '$1=="web" && $2=="app1" {print $5}')" = "0" ]; do
sleep 1
done
# ... deploy the new build on app1 ...
# put it back and let the health checks confirm it
echo "set server web/app1 state ready" | socat stdio "$SOCK"
# persist runtime state so an unrelated reload does not undo the above
echo "show servers state" | socat stdio "$SOCK" > /var/lib/haproxy/server-statego deeper
Know that HAProxy has a Runtime API on a Unix socket and that a server can be taken out of rotation with set server <backend>/<server> state drain rather than by editing the config.
Explain what the config must declare — a stats socket at level admin — and describe precisely what DRAIN keeps doing: existing connections finish, persistence is still honoured, health checks continue.
Walk a full rolling deploy: drain, poll show stat until sessions reach zero, deploy, restore to ready, and name the reload trap plus the server-state-file and load-server-state-from-file fix that closes it.
Own the operating model — who or what is allowed to drive the socket, how that authority is granted and audited, and whether server lifecycle belongs to imperative runtime commands or to a declarative pipeline that regenerates config.
## Exposing the socket Nothing works until the config offers a control channel. In the `global` section: ```haproxy global stats socket /var/run/haproxy/admin.sock mode 660 level admin expose-fd listeners stats timeout 30s ``` `level admin` is what permits state-changing commands; the lower `user` and `operator` levels allow progressively less. File permissions on that socket *are* the authorization boundary — anyone who can write to it can drain your entire fleet — so keep it in a directory owned by the HAProxy user and group-readable only by the operators or automation that need it. `expose-fd listeners` matters for reloads and is covered separately. The HTML stats page can also carry a control panel, with `stats admin if LOCALHOST` inside a `listen` or `frontend` block, but the socket is what automation should target. ## Draining one server ```bash echo "set server web/app1 state drain" | socat stdio /var/run/haproxy/admin.sock ``` The server immediately stops being a candidate for new load-balanced sessions. What continues: - **Established connections** run to completion. HAProxy does not tear them down. - **Persistence-bound requests** still reach it — a client holding the persistence cookie or stick entry for `app1` is still served there. This is the defining property of DRAIN: it drains the *new* traffic while honouring commitments already made. - **Health checks keep running**, so the server's real health is still visible in the stats. You then wait for the current-session counter to reach zero. `show stat` returns CSV including the `scur` (current sessions) and `status` columns: ```bash echo "show stat" | socat stdio /var/run/haproxy/admin.sock | cut -d, -f1,2,5,18 ``` When it is quiet, deploy, let the health checks confirm the new build, and restore with `set server web/app1 state ready`. ## DRAIN versus MAINT `set server <b>/<s> state maint` (equivalently `disable server <b>/<s>`) is the stronger administrative removal: | | DRAIN | MAINT | |---|---|---| | New load-balanced sessions | no | no | | Persistence-bound requests | still served | not served | | Health checks | keep running | stopped | | Stats status | `DRAIN` | `MAINT` | Use DRAIN for a graceful rolling deploy where sticky sessions must be allowed to finish naturally. Use MAINT when the node must be *out*, regardless of persistence — hardware maintenance, a suspected bad build, or a machine you are about to terminate. Because checks stop in MAINT, a server left there never comes back on its own; it is an explicit, sticky decision. Neither state kills sessions already in flight. If a WebSocket or a long poll would otherwise hold the node for hours: ```bash echo "shutdown sessions server web/app1" | socat stdio /var/run/haproxy/admin.sock ``` ## Weight zero is not the same thing `set server web/app1 weight 0` also stops new balanced traffic and is sometimes used interchangeably. It is closer to DRAIN than MAINT — persistence still applies — but it is a weighting decision rather than an explicit state, which makes it harder to read in the stats and easier to overwrite the next time an automation adjusts weights. Prefer the explicit state verbs for lifecycle operations, and reserve weights for capacity shaping. ## The reload trap Everything above lives in the running process's memory. A `systemctl reload haproxy` spawns a fresh process from the config file on disk, where `app1` is a perfectly ordinary server — so a node you drained twenty minutes ago rejoins the pool the moment an unrelated config change is deployed. Mid-deploy, that is exactly the wrong moment. The fix is HAProxy's server-state file. Dump state before stopping or reloading, and tell the new process to load it: ```bash echo "show servers state" | socat stdio /var/run/haproxy/admin.sock > /var/lib/haproxy/server-state ``` ```haproxy global server-state-file /var/lib/haproxy/server-state defaults load-server-state-from-file global ``` Wire the dump into the unit's `ExecReload`/`ExecStop` so it happens automatically. This carries over the administrative state, the operational UP/DOWN status and the weight — which also removes a second reload hazard: without it, every server starts fresh and receives traffic before its first health check has had a chance to fail. ## What good operational practice looks like Drain, verify sessions reached zero, deploy, wait for checks to pass, restore — scripted, not typed, so it is identical at 3 a.m. Keep the socket's permissions tight, keep the state file wired into reload, and prefer explicit states over weight fiddling so that anyone reading the stats page can tell *why* a server is not taking traffic.
- You drained a server but its session count never reaches zero. What is going on and what are your options?Something is holding long-lived connections — WebSockets, server-sent events, long polling, or an idle keep-alive connection a client refuses to close. DRAIN does not tear those down. Either wait with a bounded deadline, or cut them explicitly with `shutdown sessions server web/app1` and accept that those clients see a reset and must reconnect.
- How do you make a runtime drain survive a config reload?Dump the running state with `show servers state` into the path named by `server-state-file` in `global`, and set `load-server-state-from-file global` so the new process reads it at startup. Wire the dump into the service unit's reload and stop hooks. It carries administrative state, health status and weight, so a drained server stays drained.
- Why prefer `set server ... state drain` over `set server ... weight 0`?Both stop new balanced traffic, but weight 0 is a capacity knob, not a lifecycle statement: the stats page still shows the server as UP, and the next automation that recalculates weights can silently undo it. The explicit DRAIN state is visible, self-documenting, and survives into the server-state file as an intent rather than a number.
- What stops an unauthorized user from draining your whole backend through that socket?Only the filesystem. The `stats socket` is a Unix socket whose `mode`, owner and group are the access control, alongside the declared `level` — `admin` permits state changes, `operator` and `user` progressively less. Put it in a directory owned by the HAProxy user, restrict the group to the operators and automation that need it, and never expose an admin-level control channel over TCP without a separate authenticated path.
saying these in an interview costs you the question
- Says you must edit the config and reload to remove a server
- Thinks DRAIN immediately kills established connections
- Confuses DRAIN and MAINT on whether health checks keep running
- Forgets runtime state is lost on the next reload
- Exposes an admin-level stats socket without tight permissions