On a liveness check, how do the failure threshold, interval and per-attempt timeout trade detection speed against false restarts?
answer
- three knobs, one detection time
- interval times threshold, plus one timeout
- a short timeout fails a merely slow instance
- a stall is not yet a verdict
- restart verdicts want patience, routing verdicts speed
basics
~20 sDetection takes roughly the interval times the failure threshold, plus one timeout. Shrinking any of the three catches a hang sooner and also turns an ordinary stall into a restart, so all three should be set against the longest pause the service legitimately has.
solid answer
~50 sThree numbers set one behaviour. The **interval** is how often the platform asks; the **timeout** is how long one attempt may take before it counts as a failure; the **failure threshold** is how many consecutive failures make a verdict. Worst-case detection is about `interval x threshold + timeout`, so every reduction buys faster detection of a genuine hang and simultaneously shortens the stall the service is allowed to have without being restarted. The timeout is the subtlest of the three: set below the service's own slow-path response time, it turns *slow* into *dead*, and slowness is exactly what happens under load — so a tight timeout can restart the instances you can least afford to lose. The practical rule is to make the restarting verdict slow and sure, and let the routing verdict be the twitchy one, because unrouting is cheap and reversible.
code
pseudocode · 17 lines# the verdict loop the platform runs per instance, per check
consecutiveFailures = 0
every check.interval seconds:
result = ask(check.endpoint, giveUpAfter = check.timeout)
if result == success:
consecutiveFailures = 0 # one success clears the count
else:
consecutiveFailures = consecutiveFailures + 1
if consecutiveFailures >= check.maxFailures:
act(check.verdict) # restart, or leave the routing set
consecutiveFailures = 0
# with interval 10, timeout 2, maxFailures 3, a total hang is declared
# between (3 - 1) * 10 + 2 = 22s and 3 * 10 + 2 = 32s after it begins,
# depending on where in the interval the hang startedgo deeper
Recall what each of the three numbers means on its own: how often the platform asks, how long one attempt may take, and how many consecutive failures make a verdict.
Do the arithmetic out loud — interval times threshold plus one timeout — and explain why shrinking the numbers buys faster detection at the price of restarting healthy instances.
Demonstrate the asymmetry of cost under load: a tight timeout on a saturated fleet converts slowness into restarts, and you would rather be late to a rare hang than early to a common stall.
Argue for the defaults other teams inherit. Whatever numbers land in the shared template become the fleet's behaviour under stress, so they should be justified by measured tail latency rather than chosen for how fast they sound.
## The three knobs - **Interval** — how often the platform asks. It sets the resolution of the whole mechanism: nothing can be detected sooner than the next scheduled attempt. - **Per-attempt timeout** — how long a single attempt may take before it is recorded as a failure. It bounds one attempt, never the sequence. - **Failure threshold** — how many *consecutive* failures are needed before the platform acts. Under the common model a single success clears the count, which is what makes an isolated blip harmless. They are not independent. Together they define a single quantity you actually care about: how long a stuck instance keeps its place before anything happens, which is the same quantity as how long a healthy instance is allowed to stall before it is ended. ## What the numbers actually buy Take an interval of 10 seconds, a timeout of 2 seconds and a threshold of 3 consecutive failures, and suppose an instance hangs completely. - If it hangs *just before* a scheduled attempt, that attempt times out at +2s (failure 1), the next at +12s (failure 2), the third at +22s (failure 3) — verdict at about **22 seconds**. - If it hangs *just after* a success, the first attempt that can see it is a full interval away, so the same three attempts land at +12s, +22s and +32s — verdict at about **32 seconds**. So detection lands between roughly `(threshold - 1) x interval + timeout` and `threshold x interval + timeout`. Those are the numbers to state out loud; a single figure is a sign the model has not been thought through. ## The per-attempt timeout is the subtle one A timeout is often set to a small round number by reflex, and it is the knob most likely to cause a wrong verdict. If the check is answered on the same path that serves requests, then when the instance is saturated the check waits behind the same backlog. A timeout tighter than the service's own slow-path latency converts *busy* into *failed*, and the platform answers load by ending processes — removing capacity from a fleet that is already short of it. The cost is asymmetric, which is what makes the tuning decision straightforward: being a few seconds late to catch a hang costs a few seconds; restarting a healthy, loaded instance costs its warm state, its in-flight work and a whole start-up. ## Restarting and routing verdicts want different numbers | | restarting verdict | routing verdict | |---|---|---| | cost of being wrong | high and irreversible | low and self-correcting | | sensible threshold | higher — several consecutive failures | lower — react quickly | | sensible interval | longer | shorter | | sensible timeout | generous, above the slow-path latency | closer to the latency callers will accept | | what it should catch | a process that will never recover | an instance that cannot serve right now | Copying one set of numbers into both is a common defect. It either makes the routing verdict too sluggish, so callers keep hitting an instance that cannot serve, or makes the restarting verdict too eager, so ordinary stalls become restarts. ## A way to pick the numbers 1. **Measure the service's slow-path latency** — the tail, not the median — and set the check timeout comfortably above it. 2. **Decide the longest stall that is honestly acceptable** and make `interval x threshold` at least that long for the restarting verdict. 3. **Set the routing verdict independently**, tuned to how quickly you want traffic pulled from an instance that is briefly unable to serve. 4. **Sanity-check the extremes.** Detection that is nearly instant means the fleet is one latency spike away from restarting itself; detection measured in many minutes means a hang is indistinguishable from a healthy instance for that whole time. ## Why 'make it instant' is the wrong instinct The attraction of tiny numbers is that a hung instance is caught immediately. The hidden term is that *every* momentary condition is now a verdict: a pause while a large allocation completes, a burst of slow requests, a brief loss of a connection. Those are common and self-resolving; a permanently hung process is rare. Tuning should reflect that ratio, not the drama of the rare case.
- Why can the routing verdict afford twitchier settings than the restarting one?Because its verdict is cheap and reversible. Pulling an instance out of the routing set costs a little capacity for a few seconds and undoes itself on the next success, so reacting early is nearly free. A restart destroys state and costs a whole start-up, so the same eagerness is expensive.
- Why does an overloaded but perfectly healthy instance tend to fail its checks first?Because the check usually shares the request path, so it queues behind the same backlog and hits its timeout while the service is merely slow. Load then arrives as a health verdict, and restarting removes capacity from a fleet that is already short.
- If a single failed attempt is followed by a success, what has the platform recorded?Under the common consecutive-failure model, nothing that matters: the success resets the count and the threshold is never approached. That is deliberate, because isolated failures are ordinary and are not evidence of anything on their own.
saying these in an interview costs you the question
- Sets the check timeout below the service's slow-path response time
- Says one failed attempt should be enough to restart an instance
- Gives the restarting and routing verdicts identical numbers by default
- Claims a shorter interval costs nothing but a little extra traffic
- Believes failures accumulate across intervening successes
- Thinks the per-attempt timeout bounds the total detection time