skip to content

Why do requests still reach a Go instance after its readiness handler starts returning 503 during a rollout?

level: seniorimportance: should knowfreq 40%

answer

  1. the flag flip is local, the routing is not
  2. health signals are polled, not pushed
  3. one failed probe is usually not enough to act
  4. the sleep is the part that saves the requests
  5. compare probe-failure events with access log timestamps

basics

~20 s

Because the routing layer learns about the 503 only after its next probe fails and that decision propagates to every proxy. The instance must keep serving normally for longer than that lag before it stops accepting, or it drops in-flight traffic.

solid answer

~50 s

Flipping the readiness flag changes one thing: what the `/readyz` handler answers. Nothing in front of the instance knows yet. Whatever routes traffic has to poll the endpoint, see enough consecutive failures to act, and then push the updated endpoint set to every proxy - all of which takes seconds, not microseconds. Requests dispatched before that convergence are already on the wire. So the correct code after `ready.Store(false)` is a deliberate wait - keep accepting and serving for a drain delay comfortably longer than the probe period times its failure threshold plus propagation - and only then begin termination. Liveness must keep returning 200 for the whole window, or the process is killed mid-drain. You confirm the sizing by reading the platform's probe-failure event timestamps next to the access log: if the log shows requests served after the first failed probe, that gap is your minimum delay.

code

go · 4 lines
go
<-sig
ready.Store(false)     // /readyz now answers 503
time.Sleep(drainDelay) // keep accepting and serving while routing converges
// only after this delay does termination begin

go deeper

for a junior

Understand that returning 503 from readiness is only an announcement - nothing upstream reacts to it until it next polls the endpoint and decides to act.

for a middle

Be able to lay out the ordered sequence in code: flip the flag, wait a configurable delay while still serving, then begin termination - and say what each step buys.

for a senior

Diagnose it from evidence: read the probe-failure events next to the access log, measure how much traffic arrived after the first failed probe, and size the delay from that rather than from intuition.

for a principal

Set this as a platform-wide default so every service inherits a correct drain sequence, and make sure the delay fits inside the termination grace period the platform enforces.

## What the flag flip actually changes When the process receives its termination signal and stores false into its readiness flag, exactly one thing becomes true: the readiness handler now writes 503. The listener is unchanged, connections are unchanged, and - most importantly - everything upstream of the process still believes this instance is a valid target. ## Why the routing layer is behind Health signals are pulled, not pushed. Whatever fronts the instance discovers the change by making the next request to the readiness endpoint, which happens on a period. It usually requires more than one consecutive failure before acting, so that a single dropped probe does not flap an instance out of rotation. Then the decision has to reach the component that actually picks a backend for each request, which may be several proxies. Each of those steps costs time, and they add: ``` lag ~= time to the next probe + (failureThreshold - 1) * probe period + propagation to every proxy ``` During the whole of that window, requests are being dispatched to this instance in good faith. Some are already in flight on established connections. If the process stops accepting or exits the moment the flag flips, every one of those becomes a connection error or a 502 that the client sees - and during a rolling deploy that happens once per instance, which is why the symptom shows up as "a burst of errors on every release" rather than as a steady failure. ## The fix: a deliberate drain delay The pattern is three ordered steps inside the process: 1. Store false into the readiness flag, so the endpoint starts answering 503. 2. **Sleep.** Keep the server accepting and serving exactly as before, for a delay longer than the worst-case lag above. 3. Only then begin shutting the process down. ```go <-sig ready.Store(false) // readiness now answers 503 time.Sleep(drainDelay) // let routing converge; keep serving normally // ... only now begin terminating ``` The delay looks wasteful and is not. It is the only part of the sequence that has any effect on the client-visible error rate, because the later stages only wait for requests the instance already accepted - they cannot help with requests that have not arrived yet. Make `drainDelay` configurable, and set it from the platform's actual probe settings rather than from a guess. A common mistake is to make it shorter than a single probe period, which guarantees the instance stops serving before the first failed probe is even observed. ## Liveness during the window The instance is now in a split state that must be deliberate: readiness answers 503, liveness answers 200, and real traffic is served normally. If liveness is wired to the same flag - a tempting simplification, since both are "health" - then the platform sees a dead process during a normal drain and kills it, cutting the drain short and producing exactly the errors you were trying to prevent. Keep the two handlers independent. ## Confirming it, rather than guessing The diagnostic is a timestamp comparison, and it is convincing in a postmortem: - Take the platform's probe-failure events for the instance: when did the readiness endpoint first fail, and when was the instance actually removed from the routing set? - Take the instance's own access log: what is the timestamp of the last request it served, and how many requests did it serve after the first failed probe? If the access log shows real traffic arriving after the first probe failure, that interval is a direct measurement of the convergence lag on your platform, and your drain delay must exceed it with margin. If instead the log stops abruptly at the moment of the signal while the routing set still lists the instance, you have found the dropped traffic and its cause in one step. A second, complementary check: the errors clients see should be connection-level (refused or reset) rather than application 5xx. Application errors point at something else entirely; connection errors at rollout boundaries are the signature of this problem. ## Common variants of the mistake - **Flipping readiness but exiting immediately.** The delay is missing entirely; the flag flip is cosmetic. - **Using the same handler for liveness and readiness.** The drain gets cut short by a restart. - **Waiting only for in-flight requests to finish.** That handles requests already accepted but not the ones still being routed, which is the larger population during a rollout. - **Tuning the delay by trial and error in production.** The probe period and threshold are knowable; read them, then add margin for propagation. - **Assuming a graceful signal is always received.** If the platform's grace period expires the process is killed outright, so the whole sequence - delay included - has to fit inside it.

  • How would you size the drain delay rather than guessing at it?
    Read the probe period and the failure threshold the platform is configured with, multiply them for the worst-case detection time, add the propagation time to the proxies, then add margin. Validate it empirically: after a rollout, compare the first probe-failure event for an instance against the last line in its access log - traffic served after that first failure is the lag you must cover.
  • During the drain window, what should the liveness handler return, and why does it matter?
    200, throughout. The process is healthy - it is deliberately shedding future traffic, not wedged. If liveness is wired to the same flag as readiness, the platform sees a dead process during every normal drain and kills it, truncating the drain and causing the dropped requests the sequence exists to avoid.
  • Is waiting for in-flight requests to finish enough on its own?
    No. That covers requests the instance already accepted, but during a rollout the larger group is requests still being routed to it because the routing layer has not converged. Without the delay after the readiness flip, those arrive at a socket that is no longer accepting and become connection errors the client sees.

saying these in an interview costs you the question

  • Flips readiness and exits in the same breath
  • Thinks the routing layer is notified of the 503 immediately
  • Wires liveness and readiness to the same flag
  • Assumes waiting for in-flight requests covers the rollout gap
  • Blames client retry logic for errors that appear only at deploys
  • Sets the drain delay shorter than one probe period