skip to content

A `systemctl reload haproxy` runs `haproxy -f haproxy.cfg -p /run/haproxy.pid -sf $(cat /run/haproxy.pid)`. What does the `-sf` flag do to the old process and its in-flight connections, and what can still be lost across that reload?

level: middleimportance: should knowfreq 50%

answer

  1. a reload is a new process, not a re-read
  2. one flag finishes gently, one kills
  3. the listening socket is handed over, not reopened
  4. memory does not travel with the sockets
  5. old workers linger on long-lived connections

basics

~20 s

HAProxy's -sf means soft-finish: the new process takes over the listening sockets, then tells the listed old processes to stop accepting and finish their existing sessions before exiting. Runtime state, server health history and in-memory counters do not carry across.

solid answer

~50 s

`-sf` ("soft finish") starts the new process, hands it the listeners, and then signals the old processes to stop accepting new connections while letting sessions already in flight run to completion; `-st` is the brutal counterpart that terminates them at once, killing those sessions. In master-worker mode with `expose-fd listeners` on the stats socket, the new process inherits the actual listening file descriptors, so no client SYN is refused during the swap; without that, HAProxy pauses the listeners so the kernel keeps queueing in the backlog, and can un-pause them if the new config fails to parse. What does *not* survive is anything held in memory: Runtime API state such as a drained server, per-server health history — every server starts fresh and is dispatched to before its first check can fail — and in-memory counters. `server-state-file` with `load-server-state-from-file global` fixes the first two, and `hard-stop-after` bounds how long old workers may linger holding long-lived connections.

code

bash · 12 lines
bash
# validate before touching production
haproxy -c -f /etc/haproxy/haproxy.cfg

# persist runtime server state so drains and health status survive the swap
echo "show servers state" | socat stdio /var/run/haproxy/admin.sock \
  > /var/lib/haproxy/server-state

# soft-finish reload: new process takes over, old ones drain their sessions
haproxy -f /etc/haproxy/haproxy.cfg -p /run/haproxy.pid -sf $(cat /run/haproxy.pid)

# old workers linger while long-lived sessions are still open
pgrep -a haproxy

go deeper

for a junior

Know that reloading HAProxy starts a new process rather than re-reading the file, and that -sf is the flag that lets the old process finish its current sessions before exiting.

for a middle

Explain the handover mechanics: -sf versus -st, the listener transfer under master-worker with expose-fd listeners, and the fact that paused listeners keep the socket bound so nothing is refused mid-swap.

for a senior

Name what does not survive — runtime drains, health-check history, in-memory counters — and close those gaps with a server-state file dumped on reload plus hard-stop-after so long-lived connections cannot stack old workers indefinitely.

for a principal

Decide how configuration reaches the edge at all: reload frequency as an operational budget, whether long-lived connection workloads belong behind a process that must be replaced to change config, and what clients are required to tolerate when it is.

## What actually happens on reload HAProxy has no in-place config re-read. A "reload" is a *new process* started from the config on disk, plus an orderly handover from the old one. The flags decide how brutal the handover is: ```bash haproxy -f /etc/haproxy/haproxy.cfg -p /run/haproxy.pid -sf $(cat /run/haproxy.pid) ``` - **`-sf <pids>`** — *soft finish*. The new process signals the listed PIDs to stop accepting new connections and to finish the sessions they already have, then exit on their own. - **`-st <pids>`** — *stop*. The listed processes are terminated immediately, dropping whatever they were serving. Reserve this for when a config emergency outweighs the in-flight requests. Because the old workers linger until their sessions end, you will normally see two (or more) HAProxy processes for a while after a reload. That is expected, not a leak — up to a point, discussed below. ## Why no connection is refused The risky moment is between "old process stops listening" and "new process is listening". Two mechanisms close that window. In **master-worker** mode (`master-worker` in `global`, or `-W`) with `expose-fd listeners` declared on the stats socket, the new process asks the old one for its listening file descriptors and inherits them directly. The socket is never closed, so the kernel's accept queue is continuous and clients see nothing at all. ```haproxy global master-worker stats socket /var/run/haproxy/admin.sock mode 660 level admin expose-fd listeners hard-stop-after 60s ``` Even without descriptor transfer, HAProxy does not simply close and reopen. On Linux it *pauses* the listeners — the socket stays bound, so incoming connections accumulate in the backlog rather than being refused — and if the new process fails to start, it un-pauses them and carries on. A config syntax error therefore does not take the site down, which is why `haproxy -c -f haproxy.cfg` before reloading is a cheap and worthwhile check. ## What is lost anyway The handover moves sockets and lets sessions drain. It does not move memory. Three things routinely surprise people: **1. Runtime API state.** A server you drained with `set server web/app1 state drain` is an ordinary configured server as far as the new process is concerned; it rejoins the pool the instant the reload lands. Mid-deploy this is exactly wrong. **2. Health-check history.** A freshly started process has never checked anything. Servers begin in their configured state and are eligible for traffic before the first probe result arrives, so a reload during an incident can send requests to a machine the previous process had already evicted. This is the mechanism behind "we get a burst of 5xx every time we reload". Both are solved by the server-state file: ```bash echo "show servers state" | socat stdio /var/run/haproxy/admin.sock > /var/lib/haproxy/server-state ``` ```haproxy global server-state-file /var/lib/haproxy/server-state defaults load-server-state-from-file global ``` The dump belongs in the service unit's `ExecReload`/`ExecStop` so it is never forgotten. With it, administrative state, operational UP/DOWN status and weights carry across. **3. Anything else counted in memory.** In-memory tables and counters held by the old workers are not merged into the new process; whatever is not persisted or replicated is simply gone with the old PIDs. ## Old processes that will not leave `-sf` waits for sessions to end — and some sessions do not end. WebSocket connections, server-sent event streams and long-polling clients can hold an old worker for hours. Deploy every twenty minutes and you accumulate a stack of old processes, each with its own memory footprint and each still serving the *previous* configuration. ```haproxy global hard-stop-after 60s ``` `hard-stop-after` sets the maximum time a worker may keep running after being told to stop; when it expires, remaining sessions are cut. Choosing the value is the usual trade — long enough for normal requests to finish, short enough that reload frequency does not produce process pile-up. Frontends serving long-lived streams should expect their clients to reconnect, and clients should be written to do so. ## The checklist worth reciting Validate with `-c` first; reload with `-sf`; run master-worker with `expose-fd listeners` so the socket handover is seamless; dump and load the server-state file so drains and health status survive; set `hard-stop-after` so old workers cannot accumulate; and expect long-lived connections to be the thing that makes a "hitless" reload visible to somebody.

  • What is the difference between `-sf` and `-st`?
    `-sf` tells the named old processes to stop accepting and finish the sessions they already hold, so in-flight requests complete. `-st` terminates them immediately, dropping those sessions. Normal reloads use `-sf`; `-st` is for when a bad configuration is actively causing harm and cutting connections is the lesser evil.
  • Why do teams sometimes see a burst of 5xx right after a reload even with `-sf`?
    The new process has no health-check history. Servers start in their configured state and are eligible for traffic before the first probe completes, so requests can land on a node the previous process had already evicted. Loading a server-state file at startup carries the previous UP/DOWN status across and removes that window.
  • What does `hard-stop-after` protect you from?
    Old workers that never exit. `-sf` waits for sessions to end, and WebSocket or long-poll connections may not end for hours, so frequent deploys stack up processes that each hold memory and still run the old configuration. `hard-stop-after` caps that lifetime and cuts the stragglers, at the cost of those clients having to reconnect.
  • What does `expose-fd listeners` on the stats socket contribute to a reload?
    It lets the incoming process request the old one's listening file descriptors over the socket in master-worker mode, so the bound socket is inherited rather than closed and reopened. The kernel's accept queue is never interrupted, which makes the handover invisible to clients rather than merely brief.

saying these in an interview costs you the question

  • Thinks a reload makes the running process re-read its config in place
  • Believes -sf drops existing connections like -st does
  • Assumes runtime drains and health state survive the reload
  • Treats leftover old haproxy processes as a memory leak
  • Reloads without validating the config with -c first

context