How do you gate a Go readiness handler on an atomic.Bool so it answers 503 once the process starts draining?
answer
- one bit, many readers, one writer
- a plain bool here is a reported race
- the zero value is already the right start
- headers must be set before the status
- Load in the handler, Store at the two transitions
basics
~20 sKeep one package-level atomic.Bool for readiness. Store true when warm-up finishes and false as soon as the process begins draining; the handler calls Load and writes 200 when true, or Retry-After plus 503 Service Unavailable when false.
solid answer
~40 sReadiness is a single boolean that many goroutines read and one goroutine writes, so it lives in a `sync/atomic` `atomic.Bool`. Startup calls `ready.Store(true)` once caches, migrations and connections are in place; the signal-handling goroutine calls `ready.Store(false)` the instant the process starts draining. The handler is then `if !ready.Load() { w.Header().Set("Retry-After", "5"); w.WriteHeader(http.StatusServiceUnavailable); return }`. A plain `bool` would be a genuine data race - the race detector flags it, and without synchronisation there is no guarantee a handler goroutine ever observes the write. Note the ordering: every header must be set before `WriteHeader`, because that call is what commits the status line and headers to the wire. And crucially, failing readiness does not stop the server - in-flight and newly arriving requests are still served normally while the flag is false.
code
go · 10 linesvar ready atomic.Bool // false until warm-up finishes
func readyz(w http.ResponseWriter, r *http.Request) {
if !ready.Load() {
w.Header().Set("Retry-After", "5")
w.WriteHeader(http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
}go deeper
Know that readiness is a separate endpoint from liveness, that it answers 503 when the instance should not receive traffic, and that response headers must be set before the status code is written.
Be able to write the flag and both transitions from memory, and explain why a plain bool read by handler goroutines while another goroutine writes it is a data race the -race detector reports.
Show the drain-time split state - readiness 503, liveness 200, real requests still served - and explain why flipping the flag deliberately changes nothing about accepting connections.
Decide what is allowed to move this flag at all. Every dependency wired into readiness makes the fleet fail together, so the default is a process-local flag and dependency health goes to metrics instead.
## The shape of the state Readiness is one bit of shared state with a very lopsided access pattern: written a handful of times over the process's life (once at the end of warm-up, once when draining starts), read on every probe. That is the textbook case for `sync/atomic`: ```go var ready atomic.Bool ``` `atomic.Bool` has `Load() bool`, `Store(bool)`, `Swap(bool) bool` and `CompareAndSwap(old, new bool) bool`. You almost always need only `Load` and `Store` here. ## Why not a plain bool One goroutine writes the variable while handler goroutines read it concurrently. In Go's memory model that is a data race - not a theoretical one: build with `-race` and the detector will report it the first time a probe overlaps the flip. Two things go wrong without synchronisation. First, the program is simply undefined under the memory model; there is no guarantee that a reading goroutine ever observes the write, so the instance could keep answering 200 forever after it began draining. Second, races are exactly the kind of bug that hides in testing and appears under production load. A `sync.RWMutex` around a plain bool would also be correct, but it is more code and more contention for one word - `atomic.Bool` is the idiomatic answer. ## The handler ```go func readyz(w http.ResponseWriter, r *http.Request) { if !ready.Load() { w.Header().Set("Retry-After", "5") w.WriteHeader(http.StatusServiceUnavailable) return } w.WriteHeader(http.StatusOK) } ``` Two mechanical details matter here. **Header ordering.** `w.Header()` returns the map that will be written; mutating it after `w.WriteHeader` (or after the first `w.Write`) has no effect, because by then the status line and header block are already committed. Set `Retry-After` first, then the status. Getting this backwards is a common bug that silently drops the header without any error - net/http will only complain if you call `WriteHeader` twice, logging a "superfluous response.WriteHeader call". **503 is the right code.** 503 Service Unavailable is the "I am up but cannot take this right now" answer, which is precisely what a draining or still-warming instance means. `Retry-After: 5` on it is a hint, in seconds, that the condition is transient. ## The two transitions **False to true, once, at the end of warm-up.** The zero value of `atomic.Bool` is false, which is the correct starting state: an instance that has not finished loading configuration, running migrations or filling a cache should not be sent traffic. Store true only after the last of those completes. This is the piece people forget - they register the handler and never flip it, or they flip it before the warm-up goroutine has finished, and the platform sends traffic into an instance that immediately errors. **True to false, once, when draining starts.** The termination signal arrives on a channel from `os/signal`: ```go sig := make(chan os.Signal, 1) signal.Notify(sig, syscall.SIGTERM, os.Interrupt) go func() { <-sig ready.Store(false) }() ``` The channel is buffered with capacity 1 because `signal.Notify` never blocks delivering a signal - it drops the signal if the receiver is not ready, so an unbuffered channel can miss it. ## What failing readiness does and does not do This is the part interviewers probe. Storing false changes exactly one thing: the readiness endpoint's answer. The server keeps listening, keeps accepting, and keeps serving every request that arrives - which is not a bug but the entire point. Requests are still in flight, and requests will keep arriving for a while because routing is updated asynchronously by whatever sits in front of the instance. If flipping the flag also stopped the server, you would be dropping exactly the traffic you were trying to protect. So during a drain the instance is in a deliberate split state: readiness says 503, liveness still says 200, and real traffic is served normally. Liveness must keep answering 200 throughout, or the platform will kill the process in the middle of the drain. ## Extensions people reach for, and when they are worth it - **A reason string.** Storing an `atomic.Pointer[string]` or an `atomic.Value` alongside the flag so the 503 body can say "draining" versus "warming" is cheap and helps whoever reads the logs. - **More than one gate.** If several subsystems must be ready, prefer one flag flipped by the code that waits for all of them over a handler that checks several flags - it keeps the readiness decision in one place. - **Dependency checks.** Any check that touches something shared makes readiness correlated across the fleet: one dependency blip takes every instance out of rotation at once. That is a policy decision, not a mechanical one, and it should be made deliberately.
- What happens if you call w.Header().Set("Retry-After", "5") after w.WriteHeader(http.StatusServiceUnavailable)?The header never reaches the client. `WriteHeader` commits the status line and the header block, so later mutations of the map returned by `w.Header()` are ignored - silently, with no error. The fix is ordering: set every header first, then call `WriteHeader` exactly once.
- Once the readiness flag is false, should the server stop accepting new connections?No. Flipping the flag only changes what the readiness endpoint answers. The routing layer in front of the instance is updated asynchronously, so requests keep arriving for a while and must still be served normally. Refusing them at the moment the flag flips is precisely the dropped traffic the readiness signal exists to avoid.
- Why is the os/signal channel created with make(chan os.Signal, 1) rather than unbuffered?`signal.Notify` never blocks when delivering: if the channel is not ready to receive, the signal is dropped. With an unbuffered channel the notification is lost unless a goroutine happens to be parked on the receive at that instant. A buffer of one guarantees the signal is held until the handler goroutine reads it.
saying these in an interview costs you the question
- Uses a plain bool shared between the signal goroutine and handlers
- Sets Retry-After after calling WriteHeader
- Initialises the flag to true before warm-up completes
- Returns 500 rather than 503 while draining
- Stops accepting connections the moment readiness turns false
- Guards one boolean with a full RWMutex and calls it necessary