A service's Kubernetes livenessProbe calls an endpoint that verifies the database connection. Why is that dangerous, and what would you do instead?
answer
- a restart must be able to fix it
- shared dependency = correlated fleet-wide failure
- lost caches, crash loops, thundering herd
- /healthz local vs /readyz dependencies
- no liveness probe is a valid choice
basics
~20 sA shared dependency failing makes every replica fail liveness simultaneously, so the kubelet restarts the whole fleet - which does not fix the database, destroys warm state, and creates a thundering herd on recovery. Liveness should test only local, unrecoverable failure; dependency checks belong in readiness, selectively.
solid answer
~60 sLiveness means "restart me", and a restart can only fix problems that live **inside this process**. A database outage is not one of them, so the probe converts a partial degradation into a self-inflicted outage: 1. The dependency blips; every replica's probe fails at the same moment. 2. The kubelet restarts all of them at once, cluster-wide. 3. Warm caches, connection pools and in-flight work are lost; the Pods restart, fail again, and enter CrashLoopBackOff. 4. When the database recovers, every replica reconnects simultaneously - a thundering herd that can knock it over again. Instead: - **Liveness**: a trivial, local `/healthz` that touches no dependency; it should fail only on a deadlock or unrecoverable internal state, with generous timeouts. Omitting the liveness probe entirely is a legitimate choice for services that crash cleanly. - **Readiness**: may include *strictly required* dependencies so a Pod that genuinely cannot serve leaves rotation - but if every replica goes unready you lose the ability to serve cached or degraded responses, so prefer degrading in place, with circuit breakers and jittered backoff handling the dependency itself.
code
yaml · 14 lineslivenessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3go deeper
Say that a liveness failure restarts the container and that restarting cannot fix a database outage, so the check does not belong there.
Describe the correlated failure across replicas and the split between a local /healthz and a dependency-aware /readyz.
Narrate the full incident chain including crash loops and the recovery thundering herd, and prescribe circuit breakers, jittered backoff, cached checks and graceful degradation.
Make it policy - a defined health-endpoint contract, liveness restricted to local unrecoverable states, dependency handling pushed into resilience patterns - and reason about when degrading in place beats leaving rotation fleet-wide.
## Why it is dangerous The rule behind all probe design: **a liveness probe should fail only for a condition that a restart fixes.** Restarting resets process-local state. It has no effect whatsoever on a failing database, a slow auth service, or a saturated broker. When liveness checks a shared dependency, the probe becomes *correlated across every replica* - they all query the same thing and all fail within the same few seconds. The result is a fleet-wide restart: - **Capacity collapse.** Every instance is killed while the dependency is degraded. Requests that could still have been served from cache, or that never touch the database, now fail too. - **State loss.** Warm caches, JIT-compiled code, connection pools, in-memory sessions and queued work are discarded, so recovery is slower than the outage that caused it. - **Crash loops.** The restarted containers still cannot reach the dependency, fail again and back off exponentially - so even after the database returns, Pods can sit in CrashLoopBackOff for minutes. - **Thundering herd.** When they do come back, hundreds of instances reconnect and replay work simultaneously, which frequently re-breaks the freshly recovered dependency. The original incident was "the database is slow". The incident you get is "the entire service is down and will not come back". ## What liveness should check Only conditions that are **local and unrecoverable**: a deadlocked main loop, an exhausted thread pool that never drains, a wedged native library, a watchdog counter that has stopped advancing. In practice a good liveness endpoint returns a constant 200 while the event loop can serve it, and nothing more. If it is hard to name a failure a restart would fix, that is a strong sign the service does not need a liveness probe at all - a common and defensible position, since a process that exits on fatal errors is already restarted by the restartPolicy. ## What readiness should check - carefully Readiness means "do not send me traffic", which is reversible, so including dependencies is more defensible. Still, if a dependency is shared, **all** replicas go unready at once, the Service loses every endpoint, and clients get connection failures instead of a controlled error. Guidance: - Include a dependency only if the instance genuinely cannot serve **any** useful request without it. - Exclude optional or degradable dependencies - a recommendations service being down should not take the product page out of rotation. - Cache the dependency check result briefly so the probe does not amplify load on a struggling backend. - Distinguish "unready" from "degraded": returning healthy while serving reduced functionality is often better for users. ## The proper tools for dependency failure Dependency problems are handled in the application and the mesh, not by the kubelet: connection retry with exponential backoff and jitter, circuit breakers that fail fast and recover gradually, timeouts and bulkheads so one slow dependency cannot consume every thread, caching and graceful degradation, and queueing when work can be deferred. The platform's job is to keep the process alive so those mechanisms can work. ## How to answer in an interview State the rule ("a restart must be able to fix it"), narrate the correlated-failure chain including the recovery thundering herd, then give the concrete fix: split `/healthz` (local only, generous timeout) from `/readyz` (required dependencies only, cached), and put backoff, circuit breaking and degradation in the app. Mentioning that no liveness probe at all is often the right answer signals real production experience. ## Detecting it in an existing system Look for restarts spiking across all replicas of a service at the same timestamp, `Unhealthy` liveness events correlated with a dependency's error rate, and health handlers that open database connections. Any of the three justifies splitting the endpoints.
- Is it acceptable to check the database in the readiness probe instead?It is defensible only when the instance cannot serve any useful request without that database. Even then all replicas go unready together and the Service loses every endpoint, so cache the check result, exclude optional dependencies, and consider serving degraded responses rather than leaving rotation entirely.
- When is it reasonable to define no liveness probe at all?When the process reliably exits on unrecoverable errors, since the container runtime already restarts it via restartPolicy. If you cannot describe a specific hang that a restart would fix, a liveness probe adds risk without adding recovery, and many mature teams ship readiness only.
It is like rebooting every branch office's computers because head office lost power - none of the reboots restore power, and now nobody can even work offline.
saying these in an interview costs you the question
- Reusing one health endpoint for both liveness and readiness
- Believing a restart can resolve an external dependency outage
- Ignoring that all replicas fail a shared-dependency check simultaneously
- Adding optional dependencies to readiness and taking the whole Service out of rotation
- Treating a liveness probe as mandatory for every workload