skip to content

Health Checks, Stats & Runtime Control

Here I keep the pool honest: active `httpchk` probes, passive `observe` checks that mark a server down on real traffic errors, and the stats socket / runtime API that lets me drain a node without a reload. Interviewers like it because it exposes whether I understand flapping, `rise`/`fall` tuning, redispatch on retry, and how to deploy config changes with zero dropped connections.

on this pageshow

questions

6

In HAProxy, a backend contains the line `server app1 10.0.0.1:8080 check`. What does the `check` keyword probe on its own, and how do `option httpchk` and `http-check expect` change what HAProxy accepts as a healthy server?

level: juniorimportance: must knowfreq 74%

answer

  1. which layer the probe actually reaches
  2. a listening socket is not a working app
  3. the default request is OPTIONS / HTTP/1.0
  4. any 2xx or 3xx passes unless narrowed
  5. http-check expect pins status or body

basics

~20 s

HAProxy's bare check keyword only opens a TCP connection to the server port. option httpchk makes the probe send an HTTP request instead, and http-check expect narrows success from any 2xx or 3xx reply to a specific status, regex or body string.

solid answer

~50 s

`check` on its own is a layer-4 probe: HAProxy connects to the server's address and port on every interval and treats a completed TCP handshake as healthy — which is why a process that is still listening but internally broken keeps showing UP. Adding `option httpchk` to the backend upgrades it to a layer-7 probe. With no arguments it sends `OPTIONS / HTTP/1.0`; on 2.2 and later you describe the request explicitly with `http-check send meth GET uri /healthz ver HTTP/1.1 hdr Host app.example.com`. Without an explicit expectation, any 2xx or 3xx reply passes and everything else fails, so I pin it with `http-check expect status 200` or `http-check expect string OK`. The classic trap is the default HTTP/1.0 request carrying no Host header: a name-based virtual host answers it with 400 or a redirect, and a perfectly good server is marked DOWN.

code

bash · 5 lines
bash
# Reproduce exactly what a bare `option httpchk` sends to the server
printf 'OPTIONS / HTTP/1.0\r\n\r\n' | nc 10.0.0.1 8080

# Reproduce the HTTP/1.1 form with an explicit Host, as http-check send builds it
printf 'GET /healthz HTTP/1.1\r\nHost: app.example.com\r\nConnection: close\r\n\r\n' | nc 10.0.0.1 8080

go deeper

for a junior

Know that check on the server line is what enables probing at all, that alone it is only a TCP connect, and that option httpchk is what makes HAProxy send an actual HTTP request.

for a middle

Be ready to write the modern form — http-check send plus http-check expect — and to explain the default acceptance rule of any 2xx or 3xx, plus why the missing Host header on an HTTP/1.0 probe breaks virtual hosts.

for a senior

Show judgment about what the endpoint should assert: cheap enough for fleet-wide polling, failing for the same reasons traffic fails, and pinned with an expectation rather than trusting a bare 200.

for a principal

Own the standard across the estate: one health-endpoint contract every service implements, a rule about probing the traffic port versus an admin port, and a position on whether a probe is allowed to reflect downstream dependencies at all.

## What `check` alone actually does In HAProxy, health checking is opt-in per server. A `server` line without `check` is never probed at all — HAProxy assumes it is up forever and will happily keep dispatching to a dead machine until the connection attempt fails at request time. Adding `check` enables a *layer-4* probe. HAProxy periodically opens a TCP connection to the server's address and port, and if the handshake completes it closes the connection and records a success. That is the entire test. The check status you see in the stats page for this mode is `L4OK`, or `L4CON` / `L4TOUT` when the connection is refused or times out. ```haproxy backend web server app1 10.0.0.1:8080 check ``` A layer-4 check answers exactly one question: is something accepting connections on that port? It does not know whether the application behind the socket can reach its database, has deadlocked its thread pool, or is returning 500 to every request. This is why a TCP-only check is the single most common reason a load balancer keeps traffic on a broken node. ## Upgrading to an HTTP check `option httpchk`, declared in the backend (or in `defaults`), turns the probe into a real HTTP request over that same connection. ```haproxy backend web option httpchk http-check send meth GET uri /healthz ver HTTP/1.1 hdr Host app.example.com http-check expect status 200 server app1 10.0.0.1:8080 check ``` With `option httpchk` and nothing else, HAProxy sends `OPTIONS / HTTP/1.0` — a deliberately cheap request that most servers answer without executing application code. On HAProxy 2.2 and later, `http-check send` is the supported way to describe the request: `meth` sets the method, `uri` the path, `ver` the HTTP version, and `hdr` adds header lines. The older single-line form, `option httpchk GET /healthz`, still parses but is superseded. One subtlety catches nearly everyone: `HTTP/1.0` requests carry no `Host` header by default. If the backend serves several name-based virtual hosts, the probe lands on whatever the default vhost is — frequently producing a 400, a 301 to canonical HTTPS, or the wrong application entirely. Either send `ver HTTP/1.1` with an explicit `hdr Host ...`, or make sure the probe target answers correctly without a host name. ## What counts as healthy If you declare no expectation, HAProxy's rule is simple: **2xx and 3xx are healthy; everything else is a failure.** That is often too loose. A service that returns `200 {"status":"degraded"}` passes, and so does a `302` to a login page from a misrouted probe. `http-check expect` replaces that default with something you chose: ```haproxy http-check expect status 200 http-check expect rstatus ^2[0-9][0-9]$ http-check expect string "\"status\":\"ok\"" http-check expect ! rstring (degraded|draining) ``` `status` matches an exact code, `rstatus` a regular expression over the status line, `string` a literal substring of the response body, and `rstring` a regex over it. Prefixing with `!` inverts the match, which is how you express "healthy unless the body says otherwise". Body matching only inspects what fits in the check buffer — the first `tune.chksize` bytes, 16 KB by default — so keep health responses small and put the verdict near the top. A related keyword, `http-check disable-on-404`, is worth knowing: when the probe returns 404 the server moves to NOLB rather than DOWN, meaning it stops receiving *new* load-balanced sessions while still serving requests that carry persistence. That is a graceful-shutdown signal the application itself can raise. ## Where the probe goes By default the check targets the same address and port as production traffic, which is what you want — it exercises the same listener. The server keywords `port` and `addr` override that: ```haproxy server app1 10.0.0.1:8080 check port 9000 ``` This is useful when the app exposes a dedicated management port, but it weakens the signal: the admin port can answer perfectly while the traffic port is wedged. Prefer checking what clients actually use unless you have a concrete reason not to. ## Practical shape of a good check Make the endpoint cheap enough to survive being called every couple of seconds by every proxy in the fleet, make it return a machine-readable status you can pin with `http-check expect`, and make sure it fails for reasons that traffic would also fail for. A probe that only proves the HTTP listener is alive has recreated the layer-4 check with more moving parts.

  • How would you health-check a server on a different port from the one that serves traffic, and why might you not want to?
    The server keywords `port` and `addr` retarget the probe, as in `server app1 10.0.0.1:8080 check port 9000`. It is handy when the app exposes a management listener, but it decouples the signal from reality: the admin port can answer while the traffic listener is saturated or wedged. Prefer probing the production port unless you have a specific reason.
  • Your health endpoint returns HTTP 200 with a body of {"status":"degraded"}. How do you make HAProxy treat that as DOWN?
    Add a body expectation rather than relying on the status: `http-check expect ! rstring degraded`, or positively match the good state with `http-check expect string "\"status\":\"ok\""`. HAProxy only matches inside the check buffer — `tune.chksize`, 16 KB by default — so keep the response small and put the status field early in it.
  • What does `http-check disable-on-404` give you that a plain failure does not?
    A 404 from the probe moves the server to NOLB instead of DOWN: it stops receiving new load-balanced sessions but still serves requests that carry persistence, and it stays visible in the stats. That lets the application itself signal a graceful wind-down before a deploy, rather than being yanked out abruptly.

saying these in an interview costs you the question

  • Thinks `check` alone verifies the application, not just the port
  • Assumes the default HTTP check sends GET / and requires 200
  • Believes any HTTP response at all means the server is UP
  • Sends an HTTP/1.0 probe with no Host to a name-based vhost
  • Confuses http-check expect with an ACL applied to live traffic

context

open as a page

You need to take one server out of an HAProxy backend for a deploy without dropping in-flight requests and without reloading HAProxy. What does the Runtime API let you do, what must the config expose for it, and how do the DRAIN and MAINT states differ?

level: seniorimportance: must knowfreq 56%

basics

~20 s

HAProxy's Runtime API, exposed through a stats socket at admin level, accepts set server <backend>/<server> state drain to stop new sessions while existing ones finish. MAINT forces the server fully out and stops its health checks; DRAIN keeps checking and still honours persistence.

open as a page

A `systemctl reload haproxy` runs `haproxy -f haproxy.cfg -p /run/haproxy.pid -sf $(cat /run/haproxy.pid)`. What does the `-sf` flag do to the old process and its in-flight connections, and what can still be lost across that reload?

level: middleimportance: should knowfreq 50%

basics

~20 s

HAProxy's -sf means soft-finish: the new process takes over the listening sockets, then tells the listed old processes to stop accepting and finish their existing sessions before exiting. Runtime state, server health history and in-memory counters do not carry across.

open as a page

An HAProxy server line reads `server app1 10.0.0.1:8080 check inter 2s fall 3 rise 2`. What does each of `inter`, `fall` and `rise` control, roughly how long can it take HAProxy to notice that server has died, and what do `fastinter` and `downinter` add?

level: middleimportance: should knowfreq 62%

basics

~20 s

In HAProxy, inter is the interval between probes, fall the number of consecutive failed checks that mark a server DOWN, and rise the consecutive successes that bring it back. Detection costs roughly fall times inter, which fastinter shortens once a check has already failed.

open as a page

An HAProxy backend keeps dispatching to a server whose /healthz probe still returns 200 while real requests to it fail or return 5xx. Which HAProxy server keywords let live traffic itself mark that server down, and what has to be in place for them to take effect?

level: seniorimportance: should knowfreq 42%

basics

~20 s

HAProxy's observe server keyword watches real traffic instead of probes: observe layer4 counts connection errors, observe layer7 also counts 5xx and malformed responses. error-limit sets how many errors trigger the on-error action, and active checks must be enabled for it to work.

open as a page

HAProxy's default HTTP log line ends each request with a four-character termination state such as `sH--` or `SC--`. What do the first two characters encode, and what do `sH`, `sQ`, `cD` and `SC` each tell you about a failed request?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

In HAProxy logs the first character says who or what ended the session and the second says which phase it was in. sH is a server-side timeout waiting for response headers, sQ a timeout in the queue, cD a client timeout during data transfer, and SC a server refusing or resetting the connection.

open as a page