What is a liveness probe versus a readiness probe in a container orchestrator like Kubernetes, and what does each one control?
answer
- kill vs remove-from-service
- liveness=restart, readiness=traffic gate
- don't reuse same handler for both
- deep checks belong on readiness, not liveness
- startup probe covers slow boot separately
basics
~20 sA liveness probe checks if an app is stuck and needs restarting. A readiness probe checks if it's ready to handle traffic right now. Failing liveness means restart the container; failing readiness means stop sending it requests, but leave it running.
solid answer
~50 sLiveness and readiness probes are two independent health checks Kubernetes runs against a pod, and they trigger different actions on failure. A liveness probe answers 'is this process alive and not deadlocked?' — if it fails repeatedly past its failureThreshold, the kubelet kills and restarts the container. A readiness probe answers 'can this pod currently serve traffic?' — if it fails, the pod is removed from the Service's endpoint list, no restart, so no new requests reach it until it passes again. They're deliberately separate: a pod can be alive but not ready — warming a cache, temporarily overloaded, a downstream dependency degraded — and you don't want to restart a healthy process just because it's briefly not ready. Reusing the same handler for both is a classic mistake that turns transient unreadiness into unnecessary restarts.
go deeper
Should state the core distinction — liveness restarts, readiness gates traffic — even if fuzzy on exact Kubernetes field names.
Should know the actual probe types (HTTP/TCP/exec), the key timing fields, and that readiness affects Service endpoints, not restart count.
Should proactively flag the anti-pattern of deep-checking dependencies in liveness and connect readiness to graceful shutdown and rolling-deploy correctness.
Should discuss this as a fleet-wide reliability control — probe design as a lever for blast radius during dependency outages and deploys, not just pod-level plumbing.
## What a probe actually is A Kubernetes probe is a periodic check the kubelet runs against a container using one of three mechanisms: - an **HTTP GET** to a path that must return a 2xx/3xx status; - a raw **TCP** connection attempt; - an **exec** command inside the container whose exit code must be zero. Every probe is governed by the same timing knobs: - `initialDelaySeconds` (wait before the first attempt); - `periodSeconds` (interval between attempts); - `timeoutSeconds` (per-attempt deadline); - `failureThreshold` (consecutive failures before the probe is declared Failed); - and `successThreshold` (consecutive successes required to flip back to healthy, mostly relevant for readiness). ## What differs is the remedy What differs is what happens once a probe is declared Failed. | Probe | What a Failed verdict triggers | |---|---| | **Liveness** | A failed liveness probe causes the kubelet to kill the container and restart it according to the pod's `restartPolicy` — the same lifecycle event as a crash. | | **Readiness** | A failed readiness probe does not touch the container at all; it only removes the pod from the Service's `Endpoints` (or `EndpointSlice`) object, which is what `kube-proxy` and any Endpoints-aware Ingress controller use to build their routing tables, so the pod simply stops receiving new traffic until it passes again. | ## Why the separation exists This separation exists because 'the process is running' and 'the process can usefully do work right now' are genuinely different questions, and production systems need to answer both without conflating the remedies. A pod can be perfectly alive — no deadlock, no crash, event loop healthy — while being temporarily unable to serve requests: - it might be finishing a **cache warm-up**; - briefly **overloaded** past its concurrency limit; - or waiting on a **downstream dependency** that is itself degraded. In none of those cases does restarting the process help; restarting just adds cold-start latency on top of an already-degraded situation. Readiness exists precisely to let a pod say 'don't send me work right now, but don't kill me either — I'll recover on my own.' ## The trade-off in tuning The trade-off surfaces in how aggressively each probe is tuned. 1. **A liveness probe tuned too aggressively** — short `periodSeconds`, low `failureThreshold` — risks restarting a process during a brief, self-recoverable hiccup (a garbage-collection pause, a momentary thread-pool saturation), paying a full cold-start cost for something that would have resolved itself in seconds. 2. **Tuned too leniently**, a genuinely hung process can sit unresponsive for a long time before Kubernetes notices and replaces it, serving errors or timeouts to anyone still routed to it via stale connections. 3. **Readiness probes carry an analogous but lower-stakes trade-off**: too aggressive, and a pod flaps in and out of the routing table on minor blips, causing connection churn at the load balancer; too lenient, and a genuinely broken pod keeps receiving traffic it can't handle. ## The failure modes that show up in production The most common production failure mode is misassigning depth or duty to the wrong probe. Pointing a liveness probe at a handler that checks a downstream dependency — a database, a cache, another service — is dangerous: if that dependency has a shared outage affecting every replica at once, every replica's liveness probe fails simultaneously, and the orchestrator restarts the entire fleet in near-lockstep. Restarting does nothing to fix the external dependency, so the replicas come back, immediately fail liveness again, and enter `CrashLoopBackOff` — turning a transient dependency blip into a total, self-inflicted outage. A second common failure mode is omitting the readiness probe altogether: without one, Kubernetes assumes a pod is ready the moment its container starts, so during any rolling deploy, freshly started pods that haven't finished booting can receive traffic immediately, producing a burst of connection-refused or 5xx errors on every release. ## Seeing it in one deploy window A concrete illustration: a Java service behind a Kubernetes Deployment takes several seconds to finish loading its routing table and warming an in-memory cache after the process starts. With only a liveness probe configured and no readiness probe, a rolling update replaces old pods with new ones, and the Service immediately routes a share of traffic to each new pod as soon as it's Running — even though the app can't yet answer requests correctly — producing visible errors during every deploy window. Adding a readiness probe that only flips healthy once the cache is warm eliminates that error burst entirely, because the Service simply withholds traffic from the pod until it says it's actually ready, while the separate liveness probe continues to watch only for genuine process-level hangs during the pod's ongoing lifetime.
- What happens if you configure only a liveness probe and no readiness probe on a pod that takes 30 seconds to load its routing table after the process starts?Without a readiness probe, the pod is added to the Service's endpoints as soon as the container starts running, not when the app can actually serve requests. Traffic gets routed to it during that 30-second warm-up and clients see connection refused or 5xx errors until initialization finishes. This is a common cause of brief error spikes during rolling deploys.
- Why is it dangerous to point a liveness probe at an endpoint that checks the database connection?If the database has a brief outage, every pod's liveness probe fails at roughly the same time and the orchestrator kills and restarts all of them at once. Restarting doesn't fix the database, so the pods come back, fail liveness again, and get restarted repeatedly, turning a database blip into a full crash-loop across the fleet. Liveness should only ever check the process's own internal health.
- What's a startup probe and why was it added separately from liveness?A startup probe covers slow-booting containers by disabling the liveness and readiness probes until the startup probe first succeeds. Before it existed, teams had to set a large initialDelaySeconds on liveness itself, which meant a genuinely hung process during normal operation wouldn't be caught until that same long delay elapsed. Separating the two lets boot be slow while runtime liveness checks stay tight.
Liveness is like checking a worker still has a pulse — if not, replace them. Readiness is like checking whether they're currently at their desk able to take a call — if not, route calls elsewhere without firing them.
saying these in an interview costs you the question
- Says liveness failing removes the pod from load balancing
- Uses the identical handler/endpoint for both liveness and readiness
- Doesn't know liveness failure causes a container restart, not traffic rerouting
- Thinks readiness probes affect container restart count
- Can't explain why deep dependency checks are dangerous in a liveness probe