How do you decide whether a Go service's liveness handler may check a shared dependency like its database?
answer
- one lever, and it is destructive
- ask what a restart actually repairs
- every replica checks the same thing at once
- shared dependency, correlated failure, fleet restart
- a process-local watchdog instead of a dependency ping
basics
~20 sDecide by what a restart can fix. A dependency check makes every replica fail liveness at the same moment during one shared outage, so the default is no dependency I/O in liveness; put shared state in readiness or in alerts instead.
solid answer
~50 sThe test I apply is simple: would restarting this process fix the condition the check reports? A database blip fails that test - a restart does not repair the database, it just discards warm caches and connections, and because every replica probes the same dependency they all fail together and the fleet restarts at once, hammering the recovering dependency. So the posture I set is that liveness performs no I/O and touches nothing shared; it reports process-local wedging only. Dependency state goes into readiness if traffic should be shed, and into metrics and alerts otherwise. The exception worth honouring is a genuine in-process wedge - a stuck work loop, an exhausted internal pool - which I detect with a process-local watchdog, an `atomic.Int64` timestamp the main loop updates and the handler compares against, not with a dependency ping. I write that down as policy, and I accept that SRE can overrule it after an incident; my job is to make the tradeoff explicit rather than to win it.
code
go · 9 linesvar lastLoop atomic.Int64 // UnixNano, stored by the work loop each pass
func livez(w http.ResponseWriter, r *http.Request) {
if time.Since(time.Unix(0, lastLoop.Load())) > 2*time.Minute {
w.WriteHeader(http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
}go deeper
Remember the rule of thumb rather than the policy: liveness does not talk to the database, because failing it restarts the process and a restart cannot repair something outside it.
Be able to explain the correlated-failure mechanism concretely - one shared blip, every replica failing within a probe period, caches and pools rebuilt at the worst moment.
Show the alternative you would actually build: a process-local watchdog with a conservative staleness threshold, dependency health moved to metrics, and readiness reserved for warm-up and draining.
Own it as a posture: ship the safe default, write the rationale down, name who can overrule it, and revisit with incident data instead of defending the rule on principle.
## The decision, stated plainly A liveness probe has exactly one lever: kill the process. So the only question worth asking about anything you put in a liveness handler is *would a restart fix this?* Everything else is noise. Run the candidate checks through it: - The database is unreachable. Restart does not fix it. **Not liveness.** - A downstream service is timing out. Restart does not fix it. **Not liveness.** - The cache is cold. Restart makes it worse. **Not liveness - readiness.** - A background work loop has been stuck for ten minutes with no progress. Restart plausibly fixes it. **Candidate for liveness.** ## Why the dependency check is not merely useless but actively dangerous The failure mode is correlation. Every replica probes the same shared dependency, so a single blip fails the check on all of them within one probe period. The platform obediently kills them all. Now you have: warm caches gone, connection pools re-established all at once against a dependency that was already struggling, cold start latency added to an ongoing incident, and - if enough replicas restart together - not enough capacity to serve the traffic that was still working. A degradation becomes an outage, and the outage is self-inflicted by the health check. The same argument applies, more weakly, to putting a dependency check in *readiness*: it takes the whole fleet out of rotation simultaneously, which turns a partly-working service into a completely-unreachable one. Readiness at least does not destroy the processes, so it is a recoverable mistake, but the correlation reasoning is identical. If the service can still serve anything useful when the dependency is down - cached reads, a subset of endpoints, a degraded mode - then failing readiness on the dependency throws that away. ## What I do instead 1. **Liveness is process-local and dependency-free.** It writes 200 and touches nothing that another goroutine locks. 2. **Readiness reflects only this instance's own willingness to take traffic**: warm-up finished, and not draining. Shared dependency health goes in only after an explicit argument that shedding all traffic is better than serving degraded traffic. 3. **Dependency health becomes a metric and an alert.** A human or an automated policy with a wider view decides what to do about it; the process-killer does not get that decision. 4. **If there is a real in-process wedge to detect, detect it locally.** A watchdog is the honest mechanism: the main work loop stores a timestamp on each pass, and liveness fails only if that timestamp is stale by an order of magnitude more than the loop's normal period. It reports on this process, cannot be triggered by a dependency's latency, and a restart genuinely does clear it. ```go var lastLoop atomic.Int64 // UnixNano, stored by the work loop each pass func livez(w http.ResponseWriter, r *http.Request) { if time.Since(time.Unix(0, lastLoop.Load())) > 2*time.Minute { w.WriteHeader(http.StatusServiceUnavailable) return } w.WriteHeader(http.StatusOK) } ``` Even this deserves suspicion: pick the staleness threshold so far above the normal loop period that a slow-but-progressing service can never trip it, and be honest that a watchdog which fires during a legitimate slow period is worse than no watchdog at all. ## The organisational half This is a posture, not a code review comment, and it only holds if it is written down and owned: - **Make it a default that ships**, ideally as a shared health package with a dependency-free liveness handler and a readiness flag, so that a team has to write extra code to get the dangerous behaviour rather than getting it by accident. - **Name the escalation path.** SRE can and will overrule this after an incident - typically an incident where a wedged process sat in rotation serving errors and someone concludes that a dependency check would have caught it. That conclusion is usually wrong (readiness or an alert would have caught it too, without the fleet restart), but the argument is legitimate and it is not settled by fiat. What I owe is the counter-evidence: the blast radius of a correlated restart versus the blast radius of the wedge. - **Revisit it with data.** Keep a record of incidents caused by health checks alongside incidents that a better health check would have caught. If the second list ever gets longer than the first, the posture deserves to change. - **Accept asymmetry across services.** A stateless request-serving service and a single-instance batch worker have different answers, because the correlated-restart argument depends on there being many replicas sharing one dependency. Do not force one rule where the shapes genuinely differ. ## How I would answer a challenge If someone says "but then how do we notice the database is down?", the answer is: from the metric and the alert that already exist, and from the request error rate - both of which point at the database rather than at a hundred restarted pods, and neither of which makes the incident worse while you read them.
- SRE points at an incident where a wedged process stayed in rotation serving errors and asks for a dependency check in liveness. How do you respond?Grant the gap and argue about the mechanism. A wedged process should be caught by a process-local watchdog or by readiness, both of which fix that incident without adding a correlated fleet restart during the next dependency blip. I would bring the blast-radius comparison, propose the watchdog with a conservative threshold, and agree on a review date - and if they still want the dependency check, it is their call to own, documented as such.
- Does the same argument apply to putting a database check in readiness rather than liveness?The correlation argument is identical - every replica fails at once - but the consequence is recoverable, since the processes survive and rejoin when the dependency returns. The real question is whether the service can still do something useful without that dependency. If it can serve cached reads or a subset of endpoints, failing readiness throws away working capacity, so the dependency belongs in an alert instead.
- How do you make this posture stick across many teams rather than re-arguing it per service?Ship it as the default: a shared health package whose liveness handler is dependency-free and whose readiness is a flag, so the dangerous version requires deliberate extra code. Pair it with a short written rationale, a named owner, and the escalation path for overruling it. Defaults plus a documented exception process outlast a review comment.
saying these in an interview costs you the question
- Adds a database ping to liveness so failures are noticed sooner
- Cannot say what a restart would actually repair
- Ignores that every replica checks the same dependency simultaneously
- Treats liveness and readiness as two names for one check
- Sets a watchdog threshold close to the normal loop period
- Declares the policy without naming who can overrule it