skip to content

questions

4

When a running instance fails a liveness check, versus when it fails a readiness check, what does each failure cost?

level: juniorimportance: must knowfreq 72%

answer

  1. one verdict is fatal, one reversible
  2. what the platform does, not what it asks
  3. restart discards memory; unrouting keeps it
  4. routing-set membership versus process life
  5. one signal driving both is the trap

basics

~20 s

Failing a liveness check costs the instance its life: the platform ends the process and starts a fresh one, discarding everything in memory. Failing a readiness check costs it only traffic: it leaves the routing set, keeps running, and rejoins when it passes again.

solid answer

~50 s

The two checks differ in what the platform *does* when one fails, not in how the question is asked. A failed liveness check is a verdict that the process is stuck and unrecoverable, so the platform ends it and starts a new one — in-memory state, warm caches and in-flight requests all go with it, and the instance is unavailable for a whole start-up. A failed readiness check is a far cheaper statement: *do not send me work right now*. The platform takes the instance out of the routing set, leaves the process running, and puts it back the moment the check passes. So a readiness failure is reversible and self-correcting, while a liveness failure is destructive. That asymmetry is why wiring both to the same answer is dangerous: a condition that should merely have parked an instance restarts it instead.

go deeper

for a junior

Recall the two consequences and keep them the right way round: one check's failure restarts the instance, the other's only takes it out of the routing set. Say which one is destructive.

for a middle

Explain the asymmetry mechanically: a restart discards in-memory state and costs a full start-up, while unrouting is reversible the instant the check passes again.

for a senior

Show the judgment: decide which conditions deserve a restart at all, and refuse to answer a transient or external problem with a verdict that ends the process.

for a principal

Frame the trade-off across a fleet. Aggressive restarting trades a rare stuck instance for routine loss of capacity and warm state, and that bargain moves sharply as start-up gets more expensive.

## What a health check actually is A **health check** is a question the platform asks a running instance on a fixed schedule — over the network, by running a command inside the container, or by opening a connection to a port. The answer is only ever success or failure. The platform counts **consecutive failures**, and when that count reaches a configured **failure threshold** it takes an action. Everything interesting is in the action, not the question. Two checks can call an identical endpoint with identical settings and still be completely different instruments, because the platform does completely different things when each one fails. ## The liveness verdict: end this process A check wired to the restart action — a **liveness check** — asserts exactly one thing: *this process is no longer worth keeping*. Past its threshold, the platform ends the container and starts a fresh one. That is a destructive verdict, and its costs all follow from one fact: a new process starts empty. - **In-memory state is gone.** Caches, loaded indexes, open pools and whatever warm-up the process had accumulated since it started are discarded. - **In-flight requests are lost.** Work the instance had accepted and not finished dies with it. - **Capacity is gone for a whole start-up.** If the service takes minutes to become useful, the fleet is short one instance for minutes. - **Nothing is diagnosed.** A restart clears the symptom and destroys most of the evidence in the same move. Liveness exists for one narrow case: a process that is running but permanently stuck — deadlocked, spinning, or wedged on something it will never get — where nothing short of a restart recovers it. If the condition would clear by itself, restarting is a strictly worse way to wait. ## The readiness verdict: do not send me work A check wired to the routing action — a **readiness check** — asserts something much weaker: *right now, do not send me requests*. Past its threshold, the platform removes the instance from the **routing set**: the set of instances the service-level address is allowed to hand traffic to. Nothing else happens. - The process keeps running, with all its memory intact. - It keeps the slot it was placed into, so its reservation on the node is still held. - It can keep doing background work — loading, reconnecting, compacting, draining. - The moment the check passes again, the platform returns it to the routing set. Readiness is therefore reversible and self-correcting, and it is the right answer to every condition that is temporary: still warming up, a local queue too deep to accept more, a connection this instance alone needs to re-establish. ## The cost of failing each | | liveness failure | readiness failure | |---|---|---| | what the platform does | ends the process, starts a fresh one | removes it from the routing set | | the process | killed | keeps running | | in-memory state | lost | kept | | in-flight requests | lost | allowed to finish; no new ones arrive | | can the instance undo it? | no — that instance is gone | yes — pass once and it rejoins | | time to full recovery | a complete start-up | roughly one check interval | ## Why answering both from the same signal is the trap If one signal drives both verdicts, every condition gets answered with the most expensive action available. A twenty-second stall that should have parked the instance restarts it instead, throwing away state that took minutes to build. Worse, conditions like these are often identical on every instance at once, so the expensive verdict fires fleet-wide rather than on one bad instance. The useful mental test before wiring anything to a restart: *if I restarted this instance right now, would the condition be gone?* If the honest answer is no — because the cause is outside this process — the condition belongs to the routing verdict, or to how individual requests are answered, not to liveness. ## The third job, in one line Neither verdict knows that an instance has never served yet, which is why platforms also offer a **startup check**: a bounded allowance that holds the other two off while a slow start completes. It buys time; it does not restart and it does not route. ## What an interviewer is listening for Not the definitions — the consequences, stated in the right direction: restart versus unroute, destructive versus reversible, and the judgment that a cheap reversible verdict should be preferred wherever it is honest.

  • If an instance fails its readiness check but is otherwise fine, is there any cost at all?
    Yes, an indirect one. The instance still occupies its slot and its reservation on the node while serving nothing, so the fleet is paying for capacity it cannot use, and the remaining instances absorb its share of traffic. If enough instances unroute at once, the survivors can be pushed over.
  • Can an instance pass its liveness check and still be useless to callers?
    Easily. A liveness check answers only whether the process is worth keeping, so a thin one passes while the instance cannot actually serve — still loading, or without the connection it needs. That gap is exactly what a readiness check exists to close.
  • When is a restart genuinely the right response?
    When the process is stuck in a state it cannot leave on its own — deadlocked, spinning, or holding a resource it will never get — and a fresh start demonstrably recovers it. If the cause lives outside the process, a restart only destroys warm state and buys nothing.

A checkout lane's 'next lane, please' sign takes the lane out of service until the cashier is ready again. Sending the cashier home ends the shift, and whoever replaces them starts cold.

saying these in an interview costs you the question

  • Says a failing readiness check restarts the instance
  • Says a failing liveness check only stops traffic reaching the instance
  • Believes an instance removed from routing has been stopped or replaced
  • Thinks both verdicts must be driven by the same signal
  • Says a restart is harmless because the instance comes back
open as a page

A search indexer that needs four minutes to load its index is killed at the same point on every attempt — why?

level: middleimportance: must knowfreq 62%

basics

~20 s

The liveness check starts counting before the index is loaded, so the same failure count is reached at the same elapsed time on every attempt and the platform restarts it — forever. The fix is a start-up allowance, not a looser restart threshold.

open as a page

On a liveness check, how do the failure threshold, interval and per-attempt timeout trade detection speed against false restarts?

level: middleimportance: should knowfreq 45%

basics

~20 s

Detection takes roughly the interval times the failure threshold, plus one timeout. Shrinking any of the three catches a hang sooner and also turns an ordinary stall into a restart, so all three should be set against the longest pause the service legitimately has.

open as a page

Every instance of a service fails its readiness check at once because a shared downstream store is slow — what does the platform do?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It empties the routing set — every instance leaves at once, so even requests that never touch the slow store now fail, turning a partial degradation into a total outage. The processes keep running, unless the liveness check reads the same signal too.

open as a page