skip to content

A Kubernetes-deployed microservice exposes both a liveness probe and a readiness probe, plus a /metrics endpoint scraped by Prometheus. What's the functional difference between the two probes, and what happens to the service's traffic and pod status if only one of them is misconfigured to always report healthy?

level: principalimportance: should knowfreq 55%

answer

  1. liveness = restart, readiness = traffic gate
  2. startup probe for slow boot
  3. RED method: rate/errors/duration
  4. pull-based scrape vs push gateway
  5. restart storm from aggressive liveness

basics

~20 s

Liveness checks whether the app should be restarted because it's stuck; readiness checks whether it should currently receive traffic. If readiness is broken and always says "yes," a struggling instance keeps getting new requests it can't handle while it's failing.

solid answer

~40 s

Liveness answers "should this instance be restarted" - failing it kills and restarts the container. Readiness answers "should this instance get traffic right now" - failing it pulls the pod from load-balancing without restarting, letting it recover (cache warm-up, DB reconnect) before rejoining. If readiness is hardcoded healthy, an overloaded or DB-disconnected instance keeps receiving full traffic and erroring on it, instead of draining out - defeating the whole point. If liveness never fails even when truly deadlocked, a stuck pod never restarts and just sits erroring. Beyond probes, metrics aggregation (Prometheus scraping /metrics) rolls per-instance rate/error/duration (RED method) into fleet-wide dashboards and alerts, revealing degradation trends across many instances that a single instance's binary health check can't show.

go deeper

for a junior

Should know there's a check for 'is it alive' and a check for 'is it ready for traffic' and that they do different things, even if fuzzy on exact orchestrator behavior.

for a middle

Should correctly state liveness triggers restart and readiness gates traffic, and know metrics get scraped and aggregated centrally (e.g., Prometheus) rather than read per-instance.

for a senior

Should explain restart storms from misconfigured liveness, the need for startup probes on slow-boot services, and name a metrics methodology like RED to explain what fleet aggregation reveals beyond individual health checks.

for a principal

Should discuss cardinality management and cost at scale, designing probes to avoid correlated failure, and how metrics/tracing/logs correlate to give the full observability picture, not just probes.

## Health checks and what each probe answers Health checks and metrics aggregation serve two different, complementary jobs in keeping a fleet of microservice instances operating correctly, and conflating them is a common source of production incidents. - A **liveness probe** answers a narrow question about a single instance: "is this process in a state where restarting it is likely to help?" It's typically a lightweight endpoint (or even just a TCP connect check) that an orchestrator like Kubernetes polls on an interval; if it fails enough consecutive times, the orchestrator kills and restarts that specific container, on the assumption that whatever's wrong (a deadlock, a memory leak that's degraded the process beyond recovery, an unresponsive event loop) is best fixed by starting fresh. - A **readiness probe** answers a different question: "should this specific instance currently be sent new traffic?" Failing readiness doesn't kill or restart anything - it simply removes that pod's IP from the set of endpoints the load balancer/service routes to, so no new requests reach it, while existing in-flight requests can still finish and the process itself keeps running. This distinction exists because there are states an instance can legitimately be in where it needs to stop receiving new work without needing to be killed: still warming up a cache after startup, temporarily unable to reach a downstream dependency it needs, or intentionally draining before a planned shutdown. Many orchestrators also support a third probe, a **startup probe**, specifically to give a slow-booting service a generous grace period before liveness checks even begin being evaluated, since without it a service that legitimately takes 60 seconds to load a large in-memory index would otherwise get killed repeatedly by an impatient liveness probe before it ever finishes starting. ## What metrics aggregation adds Metrics aggregation operates at a different level entirely: instead of a binary per-instance signal, it collects continuous numeric time series - request counts, error counts, latency histograms - from every instance and rolls them up centrally, most commonly via **Prometheus**, which pulls (scrapes) a `/metrics` endpoint on each instance at a regular interval and stores the resulting time series, queryable and alertable via PromQL and typically visualized in Grafana. A widely used framework for what to actually measure is the **RED method** from Google's SRE practice: - **Rate** (requests per second), - **Errors** (rate of failing requests), - and **Duration** (latency distribution, usually as percentiles like p50/p95/p99). Aggregated across the whole fleet and over time, RED metrics reveal gradual degradation trends - a creeping rise in p99 latency, an error rate ticking up on one specific endpoint - well before enough individual instances would actually be unhealthy enough to fail a binary health check; this is precisely the gap probes can't fill, because a probe only tells you about one instance's current pass/fail state, not about trends or about problems that are real but not yet severe enough to flip that binary signal. ## When the mechanism causes the outage The trade-off and failure-mode space here is significant because these mechanisms sit directly in the availability control loop and misconfiguring them actively causes outages rather than merely failing to prevent them. 1. **Readiness that always reports healthy.** If a readiness probe is misconfigured to always report healthy regardless of the instance's real internal state - a common mistake is hardcoding it to return 200 instead of actually checking a dependency like the database connection - an instance that's genuinely failing (lost its DB connection, exhausted a resource pool) stays fully in the load balancer's rotation and keeps receiving its normal share of new traffic, which it then fails to serve, producing errors for real users instead of being quietly drained out to recover. 2. **Liveness too aggressive.** Conversely, a liveness probe that's too aggressive - too short a timeout, or one that depends on a slow downstream call like a database query as part of its own check - can cause what's often called a **restart storm**: during a genuine traffic spike, if response latency including the database's rises fleet-wide, the liveness probe can start failing across many or most instances simultaneously purely because of load, not because those processes are actually stuck. The orchestrator then restarts a large fraction of the fleet at once, which sheds real serving capacity and adds cold-start cost (re-establishing connections, warming caches) across the remaining and restarting instances right as the fleet is already under pressure - the health-checking mechanism itself becomes the proximate cause of a worse outage, a self-inflicted denial of service triggered by the orchestrator's own automation. ## The cardinality pitfall A second operational pitfall specific to metrics is **cardinality**: each unique combination of label values on a metric (e.g., endpoint, status code, and a per-user or per-request identifier) creates a distinct time series the backend must store and index. Using a high-cardinality value as a label - a raw user ID or request ID, rather than a bounded set like endpoint name or status code class - can generate millions of near-permanent time series, which is a well-documented way to blow up storage and query cost or badly degrade a Prometheus deployment, entirely independent of whether the underlying monitoring intent was reasonable. ## Different tools for different jobs A well-designed operational setup treats probes and metrics as different tools for different jobs: - **liveness** is a narrow, dependency-free self-check meant to catch a genuinely stuck process; - **readiness** gates traffic based on real internal state, including downstream dependency health, without being so strict it flaps under normal load; - and fleet-wide **RED metrics** with alerting catch the gradual degradation trends that individual binary checks are structurally unable to see.

  • What is a 'restart storm' and how can an overly aggressive liveness probe cause one?
    If liveness probe timeouts are too tight relative to normal load-induced latency, instances get killed and restarted simultaneously across the fleet during a traffic spike or slowdown. Restarting sheds even more capacity and adds cold-start load right when the remaining instances are already struggling, compounding the outage - effectively a self-inflicted denial of service triggered by the orchestrator's own health checks.
  • Why do you need a separate startup probe (or generous initial delay) in addition to liveness/readiness for a service with a slow boot sequence?
    Because a liveness probe with a short timeout would kill the container repeatedly during normal slow startup, mistaking 'still starting' for 'stuck', before it ever finishes booting. A startup probe gives a longer grace period before liveness checks even begin being evaluated.
  • What does the RED method (Rate, Errors, Duration) tell you that a single service's health-check status cannot?
    RED metrics aggregated across the whole fleet show request rate, error rate, and latency distribution over time and across instances, revealing degradation trends like creeping p99 latency well before enough individual instances would actually start failing health checks. Health checks are binary and instance-local; RED metrics are continuous and fleet-wide.
  • Why is unbounded label cardinality, such as adding a raw user ID as a Prometheus metric label, dangerous for a metrics aggregation system?
    Each unique label combination creates a new time series the metrics backend must store and index; a high-cardinality label like a user ID can create millions of near-permanent series, exploding storage and query cost and potentially degrading or crashing the metrics backend - a well-known operational pitfall independent of the monitoring intent.

Like a restaurant host (readiness) deciding whether to seat new customers at a table that might be backed up in the kitchen, separate from building security (liveness) deciding whether to physically remove someone who's collapsed and unresponsive - one manages incoming flow, the other manages whether the worker itself needs replacing.

saying these in an interview costs you the question

  • conflates liveness and readiness as the same check
  • thinks failing readiness restarts the pod
  • doesn't know about restart storms from aggressive liveness timeouts
  • unaware of metric cardinality cost
  • thinks a single instance's health check tells you fleet-wide health

context