skip to content

Health Endpoint Monitoring

Expose liveness, readiness and startup endpoints so orchestrators and load balancers can make correct decisions about your instance. You will learn what belongs in a deep dependency check, why a bad liveness probe causes restart loops, and how synthetic monitoring complements it.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

6

What is a liveness probe versus a readiness probe in a container orchestrator like Kubernetes, and what does each one control?

level: juniorimportance: must knowfreq 85%

answer

  1. kill vs remove-from-service
  2. liveness=restart, readiness=traffic gate
  3. don't reuse same handler for both
  4. deep checks belong on readiness, not liveness
  5. startup probe covers slow boot separately

basics

~20 s

A liveness probe checks if an app is stuck and needs restarting. A readiness probe checks if it's ready to handle traffic right now. Failing liveness means restart the container; failing readiness means stop sending it requests, but leave it running.

solid answer

~50 s

Liveness and readiness probes are two independent health checks Kubernetes runs against a pod, and they trigger different actions on failure. A liveness probe answers 'is this process alive and not deadlocked?' — if it fails repeatedly past its failureThreshold, the kubelet kills and restarts the container. A readiness probe answers 'can this pod currently serve traffic?' — if it fails, the pod is removed from the Service's endpoint list, no restart, so no new requests reach it until it passes again. They're deliberately separate: a pod can be alive but not ready — warming a cache, temporarily overloaded, a downstream dependency degraded — and you don't want to restart a healthy process just because it's briefly not ready. Reusing the same handler for both is a classic mistake that turns transient unreadiness into unnecessary restarts.

go deeper

for a junior

Should state the core distinction — liveness restarts, readiness gates traffic — even if fuzzy on exact Kubernetes field names.

for a middle

Should know the actual probe types (HTTP/TCP/exec), the key timing fields, and that readiness affects Service endpoints, not restart count.

for a senior

Should proactively flag the anti-pattern of deep-checking dependencies in liveness and connect readiness to graceful shutdown and rolling-deploy correctness.

for a principal

Should discuss this as a fleet-wide reliability control — probe design as a lever for blast radius during dependency outages and deploys, not just pod-level plumbing.

## What a probe actually is A Kubernetes probe is a periodic check the kubelet runs against a container using one of three mechanisms: - an **HTTP GET** to a path that must return a 2xx/3xx status; - a raw **TCP** connection attempt; - an **exec** command inside the container whose exit code must be zero. Every probe is governed by the same timing knobs: - `initialDelaySeconds` (wait before the first attempt); - `periodSeconds` (interval between attempts); - `timeoutSeconds` (per-attempt deadline); - `failureThreshold` (consecutive failures before the probe is declared Failed); - and `successThreshold` (consecutive successes required to flip back to healthy, mostly relevant for readiness). ## What differs is the remedy What differs is what happens once a probe is declared Failed. | Probe | What a Failed verdict triggers | |---|---| | **Liveness** | A failed liveness probe causes the kubelet to kill the container and restart it according to the pod's `restartPolicy` — the same lifecycle event as a crash. | | **Readiness** | A failed readiness probe does not touch the container at all; it only removes the pod from the Service's `Endpoints` (or `EndpointSlice`) object, which is what `kube-proxy` and any Endpoints-aware Ingress controller use to build their routing tables, so the pod simply stops receiving new traffic until it passes again. | ## Why the separation exists This separation exists because 'the process is running' and 'the process can usefully do work right now' are genuinely different questions, and production systems need to answer both without conflating the remedies. A pod can be perfectly alive — no deadlock, no crash, event loop healthy — while being temporarily unable to serve requests: - it might be finishing a **cache warm-up**; - briefly **overloaded** past its concurrency limit; - or waiting on a **downstream dependency** that is itself degraded. In none of those cases does restarting the process help; restarting just adds cold-start latency on top of an already-degraded situation. Readiness exists precisely to let a pod say 'don't send me work right now, but don't kill me either — I'll recover on my own.' ## The trade-off in tuning The trade-off surfaces in how aggressively each probe is tuned. 1. **A liveness probe tuned too aggressively** — short `periodSeconds`, low `failureThreshold` — risks restarting a process during a brief, self-recoverable hiccup (a garbage-collection pause, a momentary thread-pool saturation), paying a full cold-start cost for something that would have resolved itself in seconds. 2. **Tuned too leniently**, a genuinely hung process can sit unresponsive for a long time before Kubernetes notices and replaces it, serving errors or timeouts to anyone still routed to it via stale connections. 3. **Readiness probes carry an analogous but lower-stakes trade-off**: too aggressive, and a pod flaps in and out of the routing table on minor blips, causing connection churn at the load balancer; too lenient, and a genuinely broken pod keeps receiving traffic it can't handle. ## The failure modes that show up in production The most common production failure mode is misassigning depth or duty to the wrong probe. Pointing a liveness probe at a handler that checks a downstream dependency — a database, a cache, another service — is dangerous: if that dependency has a shared outage affecting every replica at once, every replica's liveness probe fails simultaneously, and the orchestrator restarts the entire fleet in near-lockstep. Restarting does nothing to fix the external dependency, so the replicas come back, immediately fail liveness again, and enter `CrashLoopBackOff` — turning a transient dependency blip into a total, self-inflicted outage. A second common failure mode is omitting the readiness probe altogether: without one, Kubernetes assumes a pod is ready the moment its container starts, so during any rolling deploy, freshly started pods that haven't finished booting can receive traffic immediately, producing a burst of connection-refused or 5xx errors on every release. ## Seeing it in one deploy window A concrete illustration: a Java service behind a Kubernetes Deployment takes several seconds to finish loading its routing table and warming an in-memory cache after the process starts. With only a liveness probe configured and no readiness probe, a rolling update replaces old pods with new ones, and the Service immediately routes a share of traffic to each new pod as soon as it's Running — even though the app can't yet answer requests correctly — producing visible errors during every deploy window. Adding a readiness probe that only flips healthy once the cache is warm eliminates that error burst entirely, because the Service simply withholds traffic from the pod until it says it's actually ready, while the separate liveness probe continues to watch only for genuine process-level hangs during the pod's ongoing lifetime.

  • What happens if you configure only a liveness probe and no readiness probe on a pod that takes 30 seconds to load its routing table after the process starts?
    Without a readiness probe, the pod is added to the Service's endpoints as soon as the container starts running, not when the app can actually serve requests. Traffic gets routed to it during that 30-second warm-up and clients see connection refused or 5xx errors until initialization finishes. This is a common cause of brief error spikes during rolling deploys.
  • Why is it dangerous to point a liveness probe at an endpoint that checks the database connection?
    If the database has a brief outage, every pod's liveness probe fails at roughly the same time and the orchestrator kills and restarts all of them at once. Restarting doesn't fix the database, so the pods come back, fail liveness again, and get restarted repeatedly, turning a database blip into a full crash-loop across the fleet. Liveness should only ever check the process's own internal health.
  • What's a startup probe and why was it added separately from liveness?
    A startup probe covers slow-booting containers by disabling the liveness and readiness probes until the startup probe first succeeds. Before it existed, teams had to set a large initialDelaySeconds on liveness itself, which meant a genuinely hung process during normal operation wouldn't be caught until that same long delay elapsed. Separating the two lets boot be slow while runtime liveness checks stay tight.

Liveness is like checking a worker still has a pulse — if not, replace them. Readiness is like checking whether they're currently at their desk able to take a call — if not, route calls elsewhere without firing them.

saying these in an interview costs you the question

  • Says liveness failing removes the pod from load balancing
  • Uses the identical handler/endpoint for both liveness and readiness
  • Doesn't know liveness failure causes a container restart, not traffic rerouting
  • Thinks readiness probes affect container restart count
  • Can't explain why deep dependency checks are dangerous in a liveness probe

context

open as a page

When you design a service's health-check endpoint, should it verify that its database and downstream APIs are reachable, or just confirm the process itself is running? What are the trade-offs of 'deep' versus 'shallow' checks?

level: middleimportance: must knowfreq 80%

basics

~20 s

A shallow check just says 'the app process is up.' A deep check also verifies things like the database connection work. Deep checks catch more real problems, but can make a shared outage, like the database going down, take down every instance at once if used carelessly.

open as a page

How does an external load balancer's health check, such as an AWS ALB target group health check, differ semantically from a Kubernetes readiness probe, and what production problem can arise from treating them as interchangeable during a rolling deploy or scale-down?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The load balancer's health check and Kubernetes' internal readiness check run on separate schedules and separate systems, so there's a lag between Kubernetes deciding a pod is gone and the load balancer noticing. If a pod shuts down before the load balancer catches up, some requests get sent to a pod that no longer exists, causing failed requests during deploys or scale-downs.

open as a page

A Kubernetes deployment defines livenessProbe with initialDelaySeconds: 5, periodSeconds: 10, and failureThreshold: 3 for a Java service whose JVM takes about 45 seconds to finish class-loading and warm its caches before it can respond. What happens when this deployment starts, and how would you fix it?

level: middleimportance: should knowfreq 65%

basics

~20 s

The health check starts too early and fails because the app isn't ready yet, so Kubernetes thinks it's broken and restarts it before it ever finishes starting up. Fix: give it a separate 'still starting' probe with a longer allowance, rather than just delaying the regular checks.

open as a page

A platform team wants to add synthetic monitoring — scripted checks that periodically exercise a service's health endpoints and key user flows from outside the cluster — and wire failures directly into automated remediation such as auto-restart, auto-scale, or traffic failover. What can go wrong if this is built without safeguards, and how do you guard against it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Synthetic monitoring means scripts periodically pretend to be a user and check the app still works from the outside. If a single failed check triggers an automatic fix, like restarting servers, a flaky check or a small real problem can trigger a big automatic overreaction, like restarting everything at once, and make things worse instead of better.

open as a page

A globally distributed service uses DNS-based health checks, such as Route 53 health checks, to automatically fail traffic away from an unhealthy region. What makes this pattern riskier than an in-cluster readiness probe, and how would you design against flapping and split-brain?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

DNS-based failover switches which region's address gets handed out based on health checks, but DNS answers get cached everywhere, in browsers, ISPs, and apps, for minutes, so recovery is slow and uneven. If the health check itself flaps, users can end up split across two different regions at once, which is riskier than a fast, uncached, in-cluster check.

open as a page