skip to content

An HAProxy server line reads `server app1 10.0.0.1:8080 check inter 2s fall 3 rise 2`. What does each of `inter`, `fall` and `rise` control, roughly how long can it take HAProxy to notice that server has died, and what do `fastinter` and `downinter` add?

level: middleimportance: should knowfreq 62%

answer

  1. one interval, one down counter, one up counter
  2. consecutive results, not a rolling ratio
  3. detection is about fall multiplied by inter
  4. the failing probe's own timeout adds to it
  5. different intervals for transition and down states

basics

~20 s

In HAProxy, inter is the interval between probes, fall the number of consecutive failed checks that mark a server DOWN, and rise the consecutive successes that bring it back. Detection costs roughly fall times inter, which fastinter shortens once a check has already failed.

solid answer

~50 s

`inter` is the delay between successive checks — 2 seconds by default. `fall` is how many consecutive failures are needed before the server is marked DOWN, defaulting to 3; `rise` is how many consecutive successes bring a DOWN server back, defaulting to 2. So the shipped defaults detect a dead server in roughly `fall × inter`, about six seconds, plus however long the failing check itself takes to give up — bounded by `timeout check` once the connection is established, and by `timeout connect` before that. Rather than shrinking `inter` globally and probing every server hard forever, HAProxy gives you two state-specific intervals: `fastinter` applies while the server is in transition — after the first failure, or while it is rising — and `downinter` applies while it is fully DOWN, so a dead node can be polled less often. `spread-checks` in `global` jitters check start times so a large backend is not probed in lockstep.

go deeper

for a junior

Recall the three defaults — check every 2 seconds, three consecutive failures to go DOWN, two successes to come back — and be able to point at which keyword on the server line sets each.

for a middle

Explain the detection arithmetic out loud, including that the failing check's own timeout adds to fall × inter, and describe what fastinter and downinter change about the polling rate in each state.

for a senior

Show that you have sized these against a real fleet: probe load across all proxy instances, spread-checks to avoid synchronised spikes, an explicit short timeout check, and slowstart so a recovered node is not knocked straight back down.

for a principal

Own the tradeoff as policy: how much stale-traffic exposure the platform accepts versus how much probe load services must absorb, and whether check tuning is a per-service knob or a standard every backend inherits from defaults.

## The three numbers Every HAProxy active check is governed by a small state machine on the server line: ```haproxy server app1 10.0.0.1:8080 check inter 2s fall 3 rise 2 ``` - **`inter`** — the interval between checks. Default 2000 ms. Accepts time units (`2s`, `500ms`). - **`fall`** — consecutive *failed* checks required to move an UP server to DOWN. Default 3. - **`rise`** — consecutive *successful* checks required to move a DOWN server back to UP. Default 2. Both counters require consecutive results. One success in the middle of a failing streak resets the failure counter, which is the mechanism behind the asymmetry you usually want: quick to notice trouble, deliberate about declaring recovery. ## Arithmetic of detection time With the defaults, a server that goes dark is removed after roughly `fall × inter` = 3 × 2 s ≈ 6 s. That figure is a floor, not the whole story, because the failing check's own duration adds to it. A refused connection returns instantly (`L4CON`); a black-holed one hangs until a timeout fires. Which timeout depends on the phase. Before the connection is established, `timeout connect` applies. Once it is established, `timeout check` — if declared — bounds the rest of the exchange; without it, HAProxy falls back to `timeout server`. In practice this matters enormously: a backend with `timeout server 60s` and no `timeout check` can leave each probe hanging for a minute, so three failures take three minutes rather than six seconds. ```haproxy defaults timeout connect 3s timeout check 2s timeout server 60s ``` Declaring an explicit, short `timeout check` is one of the highest-value lines in an HAProxy config. ## `fastinter` and `downinter` The naive way to detect failures faster is to lower `inter`. That is expensive: every proxy instance probes every server at that rate, forever, whether or not anything is wrong. HAProxy instead lets the interval depend on the server's state: ```haproxy server app1 10.0.0.1:8080 check inter 5s fastinter 1s downinter 10s fall 3 rise 2 ``` - **`inter`** applies in the steady, healthy state. - **`fastinter`** applies while the server is *transitioning* — after a check has failed but before `fall` is reached, and while a DOWN server is accumulating successes toward `rise`. - **`downinter`** applies while the server is confirmed DOWN, letting you stop hammering a machine that is already out of rotation. That combination gives cheap steady-state polling (5 s) with fast confirmation once something looks wrong: the first failure at 5 s, then two more at 1 s each, so about 7 s to DOWN while costing a fifth of the steady-state probe traffic of a flat `inter 1s`. ## Fleet effects Two global-level concerns show up as soon as the backend is large. First, **check load**. Probes are requests. A 200-server backend at `inter 1s`, fronted by four HAProxy instances, is 800 health requests per second before a single client shows up. Health endpoints get written as if they were free; at that rate they are not. Second, **synchronisation**. If every check starts on the same tick, the backend sees a periodic spike. `spread-checks <0..50>` in the `global` section adds a random offset of up to that percentage of `inter` to each server's schedule: ```haproxy global spread-checks 5 ``` ## Coming back gracefully `rise` decides *when* a server returns, not *how much traffic it gets on return*. A server that has just restarted with cold caches and an empty JIT profile can be knocked straight back down by receiving its full share instantly. The server keyword `slowstart <time>` ramps the effective weight from zero to its configured value over that period after the server transitions to UP: ```haproxy server app1 10.0.0.1:8080 check rise 2 slowstart 30s ``` ## Choosing values Ask what you are protecting against. Shorter `inter` and lower `fall` shorten the window in which clients hit a dead server; they also make the pool more sensitive to a single slow probe. A useful default shape is a moderate `inter`, an explicit short `timeout check`, `fastinter` for quick confirmation, `fall 3` so one blip does not evict a node, and `rise 2` or 3 with `slowstart` so recovery is not a cliff. Then verify against the stats page: the `chkfail` and `chkdown` counters tell you how often checks fail versus how often that actually caused a state change.

  • Why is setting `inter 100ms` across a 200-server backend a bad way to detect failures faster?
    It multiplies out: every server is probed ten times a second by every proxy instance, permanently, so a few proxies generate thousands of health requests per second against endpoints nobody budgeted for. Use `fastinter` to accelerate only once something looks wrong, keep a moderate steady-state `inter`, and add `spread-checks` so the probes are not synchronised into spikes.
  • A server passes `rise` and returns to rotation, then immediately falls over again under load. What would you change?
    Add `slowstart <time>` to the server line so its weight ramps from zero to full over that window instead of taking its full share on the first request — that gives caches, pools and the JIT time to warm. Raising `rise` helps confirm stability but does nothing about the traffic cliff itself.
  • Why does an explicit `timeout check` matter so much for detection time?
    Without it, an established check falls back to `timeout server`, which is usually tens of seconds. A hung backend then takes `fall` × that timeout to be evicted rather than `fall` × `inter`. A short explicit `timeout check` — a couple of seconds — keeps the probe's own duration from dominating the arithmetic.

saying these in an interview costs you the question

  • Thinks fall and rise count a percentage rather than consecutive checks
  • Ignores the failing probe's own timeout in detection time
  • Lowers inter globally instead of using fastinter
  • Sets fall 1 so a single blip evicts a healthy server
  • Assumes rise controls how much traffic a recovered server gets

context