skip to content

When you design a service's health-check endpoint, should it verify that its database and downstream APIs are reachable, or just confirm the process itself is running? What are the trade-offs of 'deep' versus 'shallow' checks?

level: middleimportance: must knowfreq 80%

answer

  1. shallow=process alive, deep=can-do-real-work
  2. deep on liveness = restart storm risk
  3. readiness can be deep, cautiously
  4. cache/debounce dependency checks
  5. degraded status, not just binary

basics

~20 s

A shallow check just says 'the app process is up.' A deep check also verifies things like the database connection work. Deep checks catch more real problems, but can make a shared outage, like the database going down, take down every instance at once if used carelessly.

solid answer

~40 s

It depends which probe you're answering with. For readiness — the check that decides whether to send this instance traffic — some depth is valuable: if the instance can't reach its primary database, it genuinely can't serve most requests, so marking it not-ready and letting the load balancer route elsewhere is correct. For liveness — the check that decides whether to kill and restart the process — depth is dangerous: restarting doesn't fix an external outage, and if every instance's liveness check depends on the same downstream service, that one failure triggers a synchronized restart storm across the whole fleet. The safe pattern is: liveness checks only internal process health, and readiness checks dependencies, ideally with caching and circuit breakers so a flaky dependency doesn't flap readiness on every single check.

go deeper

for a junior

Should say shallow checks just prove the process runs while deep checks confirm it can actually serve requests, and name one downside of deep checks.

for a middle

Should correctly map depth to probe type — shallow for liveness, selectively deep for readiness — and know at least one mitigation like caching results.

for a senior

Should design the caching and circuit-breaker layer for dependency checks and explain the restart-storm and probe-amplification failure modes with mechanism.

for a principal

Should reason about this as a fleet-wide correlated-failure problem, discuss degraded-status modeling, and connect probe design to blast-radius and SLO impact during real dependency incidents.

## Shallow and deep, defined A **shallow health check** verifies only that the process itself can respond — the HTTP server thread accepts a connection and returns a fixed 200, with no interaction beyond the request handler itself. A **deep health check** actively exercises the service's real capabilities: - it pings its **database connection pool**; - calls a **downstream API**; - checks **disk space** or **queue depth**; - and aggregates the results into a composite status, often returning per-dependency detail in the response body rather than a bare pass/fail. The two answer very different questions: | Check | The question it settles | |---|---| | Shallow | shallow answers 'is the process alive,' | | Deep | deep answers 'can this instance actually do useful work right now.' | ## Why deep checks exist Deep checks exist because a process can be technically running while functionally broken. The HTTP server thread can accept connections and return 200 on a shallow check while the connection pool to the database is exhausted, or a required downstream auth service is down, meaning every real request the instance receives will fail. A shallow check would report healthy the entire time, so nothing in the traffic-routing layer knows to avoid that instance, and real users keep hitting an instance that can't help them. Deep checks close that gap by testing actual capability rather than mere existence. ## The trade-offs The trade-offs run in both directions. - **Deep checks cost load**: if a probe runs every few seconds across dozens or hundreds of instances, and each check makes a fresh call to a shared dependency, the aggregate polling traffic against that dependency can itself become significant, sometimes enough to be the thing that tips an already-stressed database over. - **They also introduce correlated-failure risk**: if a shared downstream dependency degrades, every instance's deep check can fail at nearly the same moment, and depending on which probe carries the deep check, that means either a synchronized crash-loop restart storm (liveness) or the entire fleet marked not-ready simultaneously (readiness) — total outage instead of the intended graceful degradation. - **Shallow checks avoid both costs but pay in false positives**: the load balancer keeps sending traffic to an instance that looks fine but can't actually complete requests, hurting real users while every health dashboard stays green and alerting stays silent. ## The pattern that resolves it The pattern that resolves this in practice is to match depth to purpose and to decouple the check's cost from its frequency. 1. **Liveness should stay shallow** — process-only — always, because its remedy (restart) can never fix an external dependency problem. 2. **Readiness can be selectively deep**, but the dependency check itself should run on a background timer independent of probe frequency, with the readiness endpoint simply returning the cached result instantly; this keeps the probe response fast and cheap while the actual dependency polling happens at a much lower rate, often paired with a circuit breaker so a hanging dependency doesn't pile up slow calls. 3. Where possible, **health status should be modeled as more than binary** — a 'degraded' state for a non-critical dependency lets an instance keep serving the majority of its traffic while a specific feature fails gracefully, rather than pulling the whole instance out of rotation over one non-essential dependency. ## A realistic failure mode, and the fix A realistic production failure mode illustrating why this matters: a fleet of application pods runs a liveness probe that executes a live SQL query against a shared primary database on every check. A brief database failover event, lasting well under two minutes, causes every pod's liveness probe to fail at nearly the same moment; the orchestrator restarts the whole fleet, and because the database is still failing over when the pods come back up, many hit `CrashLoopBackOff` and fail liveness again. The database recovers in under two minutes, but the fleet-wide outage, compounded by exponential restart backoff, lasts substantially longer than the actual database event that caused it. The fix teams apply after such an incident is exactly the pattern above: - move the SQL check out of liveness entirely; - keep liveness as a pure process check; - and put the database check behind a cached, background-refreshed readiness signal so a shared dependency blip degrades traffic routing gracefully instead of triggering a synchronized restart storm.

  • How can you get the safety of a deep readiness check without paying the cost of hammering the database with a fresh query on every single probe call?
    Run the dependency check on a background timer, say every 10 to 30 seconds, independent of probe frequency, cache the result, and have the readiness endpoint just return the cached status instantly. This decouples probe frequency, which can stay fast for quick traffic-routing reactions, from dependency-check frequency, which stays low to avoid load, and a circuit breaker around the background check prevents it from piling up slow calls if the dependency is hanging.
  • If a service has three downstream dependencies and only one is degraded, should its readiness probe report unhealthy?
    It depends on whether that dependency is on the critical path for most requests. If it's required for the majority of traffic, marking not-ready is correct so the load balancer stops sending work it can't complete. If it only backs a small feature, a better design returns a degraded-but-ready status so the instance keeps serving what it can, and that one feature fails gracefully instead of taking the whole instance out of rotation.

A shallow check is like asking a shop 'are your lights on?' A deep check is like asking 'can you actually ring up a sale right now?' — more useful, but if the answer depends on a shared card network being up, every shop in the mall reporting 'no' at the same moment doesn't mean you should fire all the cashiers.

saying these in an interview costs you the question

  • Recommends checking the database from inside the liveness probe
  • Treats health status as pure binary with no notion of degraded
  • Doesn't consider the load a deep check itself puts on the dependency
  • Assumes a deep readiness check protects against correlated fleet-wide failures automatically
  • No mention of caching or debouncing dependency checks

context