skip to content

Liveness, Readiness, and Startup Probes

Liveness restarts the container, readiness pulls it from Service endpoints, and startup holds the other two off while a slow app boots. Misconfigured probes turn a slow dependency into an outage, so interviewers ask which probe to use and what each one costs.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

In a Kubernetes Pod spec, what is the difference between livenessProbe and readinessProbe, and what action does the platform take when each one fails?

level: juniorimportance: must knowfreq 85%

answer

  1. liveness fails -> restart container
  2. readiness fails -> pulled from endpoints
  3. readiness reversible, liveness destructive
  4. no probe = ready as soon as running
  5. readiness gates rolling updates and PDBs

basics

~20 s

A failing livenessProbe makes the kubelet restart that container. A failing readinessProbe leaves the container running but removes the Pod from Service endpoints so it stops receiving traffic. Liveness asks 'is it broken beyond recovery?', readiness asks 'can it serve right now?'.

solid answer

~50 s

Both are declared per container and executed by the **kubelet**, but the consequence of failure is completely different: - **livenessProbe** - "is this process wedged?" After `failureThreshold` consecutive failures the kubelet **kills and restarts the container** in place (restartCount increments, Pod IP unchanged, backoff applies on repeated failure). - **readinessProbe** - "can this instance take traffic right now?" On failure the Pod's `Ready` condition flips false and the endpoint controllers **remove its address from the Service's EndpointSlices**, so proxies stop sending new connections. Nothing is restarted, and the Pod is added back as soon as the probe passes again. So readiness handles the temporary and recoverable - warming caches, a full work queue, a brief dependency blip - while liveness handles only the unrecoverable, like a deadlock. Readiness also gates rolling updates: a Deployment waits for new Pods to become Ready before removing old ones. Getting this backwards - restarting on a transient dependency failure - is the classic mistake.

code

yaml · 17 lines
yaml
containers:
  - name: api
    image: registry.example.com/api:2.3.1
    ports:
      - containerPort: 8080
    livenessProbe:
      httpGet:
        path: /healthz
        port: 8080
      periodSeconds: 10
      failureThreshold: 3
    readinessProbe:
      httpGet:
        path: /readyz
        port: 8080
      periodSeconds: 5
      failureThreshold: 2

go deeper

for a junior

State the two outcomes clearly: liveness failure restarts the container, readiness failure removes it from Service endpoints.

for a middle

Add that the kubelet runs both per container, that thresholds control how many consecutive failures are needed, and that readiness also gates rolling updates.

for a senior

Reason about consequences: restarts destroy warm state, dependency checks belong in readiness at most, and missing readiness probes make rollouts drop traffic.

for a principal

Set a fleet policy - readiness mandatory, liveness only for provably unrecoverable local states - and connect probe semantics to rollout safety, PodDisruptionBudgets and blast radius.

## The mental model Every probe is a periodic health question the **kubelet** asks a container on its own node. The probe type does not change how the question is asked - it changes **what Kubernetes does with a 'no'**. - **liveness = restart me.** The kubelet kills the container and restarts it per the Pod's `restartPolicy`. This is a hammer: it discards the process, its in-memory state and its warm caches. - **readiness = stop sending me traffic.** The container keeps running untouched; only its membership in Service endpoints changes. ## What actually happens on failure **Liveness failure.** After `failureThreshold` consecutive failed checks the kubelet terminates the container and starts it again inside the same Pod sandbox. The Pod keeps its name, UID, node and IP; `restartCount` increments and you see `Unhealthy` and `Killing` events. Repeated failures fall into exponential backoff (CrashLoopBackOff). Nothing moves to another node. **Readiness failure.** The `ContainersReady` and `Ready` conditions go false. The EndpointSlice controller removes the Pod IP from the Service's slices; kube-proxy and ingress controllers reconcile and stop routing new connections there. Existing connections are *not* killed. The container keeps running, keeps its state, and returns to rotation as soon as a probe succeeds (`successThreshold`, default 1). ## When to use which Use **readiness** for anything temporary and self-correcting: startup warm-up, cache loading, a saturated request queue, a config reload, a brief downstream blip. Use **liveness** only for states the process cannot escape on its own - a deadlocked event loop, a thread pool that never drains, a wedged native library. If a restart would not actually fix the condition, a liveness probe is the wrong tool and will only amplify the incident. A useful reflex: *readiness is cheap and reversible; liveness is destructive and irreversible.* Many mature teams run readiness everywhere and liveness sparingly or not at all. ## Readiness beyond traffic routing Readiness is also the signal that rollouts and disruption logic depend on: - A Deployment's rolling update waits for new Pods to become Ready, subject to `maxUnavailable`/`maxSurge`, so a broken release stalls instead of replacing every healthy Pod. - `minReadySeconds` requires a Pod to stay Ready for a period before it counts as available. - PodDisruptionBudgets count Ready Pods, so readiness affects whether a node drain may proceed. That is why a missing readiness probe is worse than it looks: without one, a Pod counts as Ready the instant its containers start, so a rollout can replace every old Pod with new ones that cannot yet serve, and every request in that window fails. ## Endpoints and terminating Pods Readiness also flips false when a Pod starts terminating, which is what begins the endpoint-removal side of graceful shutdown. Removal is asynchronous across the cluster, so a Pod can keep receiving connections briefly after going unready - which is why draining relies on the Pod continuing to serve for a few seconds rather than slamming the door. ## Defaults worth knowing If no probe is declared, the kubelet treats the container as alive and ready as soon as it is running - the process merely existing is the health check. Probes are per container: in a multi-container Pod *all* containers must be ready for the Pod to be Ready, but a liveness failure restarts only the offending container. ## The failure to avoid The most common production mistake is a liveness probe that checks something external - the database, an auth service. When that dependency blips, every replica fails liveness at once and the whole fleet restarts, turning a partial degradation into a full outage and losing all warm state. Dependency checks, if they belong anywhere, belong in readiness - and even there they need care.

  • What happens if you define no probes at all?
    The kubelet considers the container alive and ready as soon as the process is running, so the Pod receives traffic immediately after start even while warming up, and a hung process is never restarted. That is why a readiness probe is the minimum for anything behind a Service.
  • Should a liveness and a readiness probe point at the same endpoint?
    Usually no. Readiness may legitimately include required dependencies and warm-up state, while liveness should test only local, unrecoverable failure. Sharing one endpoint means any condition that makes the Pod temporarily unready also restarts it, destroying state for no benefit.

Readiness is a waiter stepping off the floor for a minute; liveness is firing the waiter. Use the first for a busy moment, the second only when they have genuinely stopped working.

saying these in an interview costs you the question

  • Saying a failing readiness probe restarts the Pod
  • Believing a failing liveness probe reschedules the Pod to another node
  • Pointing a liveness probe at a database or downstream service
  • Assuming existing connections are dropped the moment readiness fails
  • Thinking one container failing liveness restarts every container in the Pod

context

open as a page

What problem does the Kubernetes startupProbe solve, and how does it interact with the liveness and readiness probes while it is still running?

level: middleimportance: must knowfreq 60%

basics

~20 s

It protects slow-starting containers: while a startupProbe is defined and has not yet succeeded, the kubelet suspends the liveness and readiness probes and the container is not ready. Once it succeeds it never runs again. Its budget is failureThreshold x periodSeconds.

open as a page

A service's Kubernetes livenessProbe calls an endpoint that verifies the database connection. Why is that dangerous, and what would you do instead?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A shared dependency failing makes every replica fail liveness simultaneously, so the kubelet restarts the whole fleet - which does not fix the database, destroys warm state, and creates a thundering herd on recovery. Liveness should test only local, unrecoverable failure; dependency checks belong in readiness, selectively.

open as a page

How do you choose periodSeconds, timeoutSeconds, failureThreshold and initialDelaySeconds for a Kubernetes probe? Work through what those numbers mean for detection time and for false restarts under load.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Worst-case reaction time is roughly initialDelaySeconds + periodSeconds x failureThreshold, plus timeoutSeconds. Defaults are period 10s, timeout 1s, failureThreshold 3. Keep liveness generous and cheap so latency spikes do not restart healthy Pods; keep readiness tighter, since removing traffic is reversible.

open as a page

Kubernetes probes support exec, httpGet, tcpSocket and grpc handlers. How does each one work, and what are the failure modes of choosing the wrong one?

level: middleimportance: should knowfreq 50%

basics

~20 s

httpGet: the kubelet calls the container's IP and port; status 200-399 is success. tcpSocket: success if a TCP connection opens - it proves only that a listener exists. exec: runs a command inside the container, exit 0 is success, but forks a process every period. grpc: calls the standard gRPC health checking service.

open as a page