What does a dependency-free liveness http.HandlerFunc in a Go service do, and what must it not touch?
answer
- the smallest handler in the whole service
- it answers a narrower question than you think
- failing it does not shed traffic
- anything it waits on can kill the process
- no locks, no I/O, just 200
basics
~20 sA liveness handler reports only that the process is alive and still serving HTTP. In Go it is a tiny http.HandlerFunc that writes 200 OK and touches nothing else: no locks, no database, no outbound calls.
solid answer
~40 sLiveness answers one narrow question: is this process still able to run code and answer HTTP? So the handler is deliberately trivial - `func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(http.StatusOK) }`, registered on a `ServeMux` as something like `/healthz`. It must not take a mutex that request handlers hold, must not query a database or call another service, and must not build a report by fanning out to subsystems. Every one of those turns a slow or contended dependency into a failed liveness check, and a failed liveness check means the process gets killed and restarted. The question "can this instance usefully serve traffic right now?" is a different endpoint - readiness - because failing readiness only takes the instance out of rotation instead of destroying it.
code
go · 5 linesmux := http.NewServeMux()
mux.HandleFunc("/healthz", func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
io.WriteString(w, "ok")
})go deeper
Be ready to write the handler on the spot: a func taking http.ResponseWriter and *http.Request that calls w.WriteHeader(http.StatusOK), registered on a ServeMux. Then say in one sentence why it checks nothing else.
An interviewer expects you to explain the consequence asymmetry: a failed liveness check restarts the process while a failed readiness check only removes traffic, which is why dependency state never belongs in liveness.
Show that you have seen the failure mode - a shared dependency blip failing liveness on every replica at once and restarting the fleet - and describe how you keep the health path out of the middleware chain and the access log.
Own the posture across services: what may appear in a probe at all, what belongs in metrics and alerts instead, and how you keep teams from encoding "degraded" into a check whose only lever is killing a process.
## Two endpoints, two questions A long-running Go HTTP service normally exposes two health endpoints, and the whole design follows from the fact that the platform reacts to them differently. - **Liveness** asks: *is this process wedged in a way that only a restart can fix?* A failing liveness check gets the process killed and started again. - **Readiness** asks: *should this instance receive traffic right now?* A failing readiness check only removes the instance from routing; the process keeps running. Because the consequences differ that sharply, the two handlers should not share an implementation. Reusing one handler for both means every condition that should merely shed traffic instead restarts the process. ## What the liveness handler looks like in Go The honest implementation is a few lines: ```go mux := http.NewServeMux() mux.HandleFunc("/healthz", func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(http.StatusOK) }) ``` `http.HandlerFunc` is the adapter that lets a plain `func(http.ResponseWriter, *http.Request)` be used where an `http.Handler` is wanted, and `ServeMux.HandleFunc` does that wrapping for you. Nothing else belongs in the body. If you want a body at all, a literal `io.WriteString(w, "ok")` is enough; note that writing a body implicitly sends 200 anyway, so the explicit `WriteHeader` is only for clarity. ## Why "touches nothing" is the rule, not a style preference The handler runs on a goroutine like any other request. Anything it blocks on becomes a reason to kill the process: - **A mutex.** If the handler takes a `sync.Mutex` that the hot request path also holds, then under contention - a slow query holding the lock, say - the health handler simply waits. The probe times out, the check fails, and the platform restarts a process whose only problem was that it was busy. Restarting it drops in-flight work and moves the load onto the remaining instances, which makes contention worse. - **A database or a downstream call.** Now every instance's liveness depends on one shared thing. When that thing has a blip, every replica fails liveness at the same moment and the whole fleet restarts together. This is the classic self-inflicted outage, and it is why the default posture is that liveness performs no I/O at all. - **Aggregating a status report.** Building a JSON document out of "database: ok, cache: ok, queue: ok" is useful for a human debugging endpoint, but it is the wrong body for a liveness probe, because every subsystem it touches is a new way to fail. ## Practical details that come up - **Give it its own path and keep it cheap.** Probes run every few seconds, forever, on every instance. A handler that allocates or logs on each call is measurable noise in profiles and in the access log; many services skip logging the health path entirely so real traffic stays readable. - **Handle it before your middleware chain, or exempt it.** If authentication, rate limiting or request-body reading wraps the health path, then the probe's success depends on that middleware too - which reintroduces exactly the coupling you were avoiding. - **Do not report per-request failures through it.** A handler that starts failing liveness after N 500s is a restart loop waiting to happen: the errors are usually caused by something outside the process, and restarting will not fix them. - **The status code is the payload.** Probes look at the status code. 200 means alive; anything else, or no answer, means not. Do not encode "degraded" as a 200 with a body field that nothing parses. ## The mental model Liveness is a heartbeat, not a diagnosis. It answers "is the runtime still scheduling my goroutines and can the HTTP server still write a response" - which is precisely the condition where a restart genuinely helps. Everything richer than that (is the database reachable, is the cache warm, is the queue backed up) is either readiness or a metric, and belongs where a false alarm costs a routing change or a page, not a process kill.
- What is the difference between what a liveness handler and a readiness handler tell the platform?Liveness says "this process is wedged, restart it"; readiness says "do not send me traffic right now, but leave me running". The consequences are completely different, so a condition that should only shed traffic - a cold cache, a draining instance, a slow dependency - belongs in readiness. Sharing one handler between the two means every transient problem becomes a process restart.
- Why is taking the service's main sync.Mutex inside the liveness handler dangerous?The handler then blocks whenever the hot path holds that lock. Under contention the probe times out, liveness fails, and the platform kills a process that was merely busy - dropping in-flight requests and pushing more load onto the remaining instances. A liveness handler must be unable to block, which in practice means it reads no shared state that anything else locks.
- Should the liveness handler write a JSON body describing each subsystem?No. Probes act on the status code, so the body is unread, and every subsystem you touch to build it is a new way for the check to fail or hang. A rich status document is a useful separate debug endpoint for humans, but it must not be what decides whether the process gets restarted.
saying these in an interview costs you the question
- Says the liveness handler should ping the database to prove health
- Uses one handler for both liveness and readiness
- Takes the service's main mutex inside the health handler
- Thinks a failed liveness check just removes the instance from routing
- Fails liveness after a run of 500s from request handlers
- Builds a JSON report by calling every downstream on each probe