skip to content

When a replica is asked to stop, why does closing its listening socket first still drop requests, and what order avoids that?

level: middleimportance: must knowfreq 60%

answer

  1. order, not just speed
  2. refusing is the second step
  3. the routing set updates with delay
  4. keep serving while removal propagates
  5. a deliberate pause before closing the listener

basics

~20 s

Because removal from the routing set is not instant. For a few seconds after the stop request, callers are still being sent here, and a closed listener turns each one into a connection failure. Leave the routing set first, keep serving while that propagates, then refuse new work.

solid answer

~50 s

Refusing new work is the second step, not the first. At the moment a replica is asked to stop, everything that routes traffic still believes it is a valid target, and requests are already on their way to it. Closing the listener at that instant does not politely decline them - it turns them into connection failures the caller must deal with, and it does so for exactly the callers that had no reason to expect it. The correct order is: announce that you are no longer a valid target, keep serving normally for a short drain pause while that removal spreads, then stop accepting new connections, finish what is in your handlers, release any claims or leases, and exit. The pause between step one and step three is the only part of shutdown you pay on every single replacement, and it is what turns a drain into a clean one.

code

pseudocode · 15 lines
pseudocode
on stopRequest:
    routingEligible = false            # 1. leave the routing set first
    sleep(drainDelay)                  # 2. let that removal propagate

    stopAcceptingNewConnections()      # 3. only now refuse new work

    deadline = now() + gracePeriod - drainDelay - safetyMargin
    while inFlightCount > 0 and now() < deadline:
        sleep(100ms)

    if inFlightCount > 0:
        log("abandoning", inFlightCount, "requests at deadline")

    releaseHeldClaims()                # 4. leases and claimed messages
    exit()                             # 5. do not wait out the window

go deeper

for a junior

Remember the order as a sentence: stop being chosen, then stop accepting, then finish, then exit. The common mistake is doing the second of those first.

for a middle

Explain why the order exists - routing membership propagates with delay, so requests dispatched in that gap arrive at a replica that has already closed its door. Name what the caller sees when it does.

for a senior

Show that you would measure the drain pause rather than pick a round number, and connect the burst of transport-level failures during deploys back to a shutdown sequence that refuses work too early.

for a principal

Argue about where the sequence belongs. A drain pause implemented per service is inconsistent and a platform-provided one is not portable, so decide which way the estate standardises and what it costs per replacement.

## Two things must stop, and the order matters A shutting-down replica has to stop two different flows: **traffic arriving at it** and **work running inside it**. The instinct is to stop the first one immediately - close the listening socket, nothing new can arrive, now finish what is left. That instinct produces dropped requests on every deployment, and the reason is timing. Routing is not a single fact held in one place. Whatever decides which replicas are valid targets - the layer in front, the callers' own client-side view, anything caching a copy of that set - learns about a change *after* it happens, not at the same instant. So at the moment the stop request arrives, the replica is still a published, advertised, actively selected target, and some number of requests are already in flight toward it or about to be. ## Why closing the listener first is the wrong move Those in-flight callers get a connection failure rather than a response. That is materially worse than a slow response, for three reasons: - **It is not an application error**, so it arrives at the caller as a transport failure with no body, no status and no idempotency information. For a ledger write, the caller genuinely does not know whether the write happened. - **It happens on every replacement**, which means every rollout step and every scale-down produces a small burst of them - the classic "we get errors during deploys and nobody knows why". - **It is invisible in the shutting-down replica's own logs**, because the request never reached a handler. You see it only on the caller's side. | Order | What the caller experiences | |---|---| | Close the listener first, then leave the routing set | Requests dispatched in the propagation gap fail at connection time | | Leave the routing set first, then close the listener after a pause | Requests dispatched in the gap are served normally by a replica that is still healthy | ## The ordering that works 1. **Announce that you are no longer a valid target.** How membership of the routing set is decided is a separate subject with its own mechanism; what matters here is that you trigger it first, before anything else changes. 2. **Keep serving, unchanged, for a drain pause.** This is a deliberate sleep. The replica is healthy, it answers everything normally, and it is simply waiting for the world to stop pointing at it. 3. **Stop accepting new connections.** Now the listener can close, because nothing should be arriving. 4. **Finish the work already in your handlers**, up to a deadline you compute from what is left of the grace period. 5. **Release what you hold** - claimed queue messages, an exclusive lease, a reserved slot. 6. **Exit.** Do not wait out the rest of the window for its own sake. Some platforms will run something on your behalf between the stop request and the workload seeing it, which gives you a place to put step one if the application cannot do it itself. Platforms differ in whether they offer that, so a shutdown sequence that only works when the platform runs a pre-stop step for you is not portable; putting the whole sequence inside the application is. ## How long the pause should be Long enough to cover the propagation delay of whatever holds a copy of the routing set, plus the longest time a caller might hold a connection it opened just before the change. A few seconds is typical, and it should be **measured, not guessed**: take a replica out of the routing set on a quiet system and watch how long requests keep arriving. If they stop after four seconds, a ten-second pause is honest and a one-second pause is theatre. The drain pause is the one part of the sequence you pay on every shutdown, even when there is no work in flight at all. It is therefore the part that shows up as "why does a rollout of forty replicas take so long" - and the wrong fix is to delete it, which simply moves the cost onto callers as failures instead. ## The failure this prevents, stated plainly A replica that closes its listener first is **refusing work it was still being sent**. A replica that leaves the routing set first is **being sent less work until it is sent none**, and then refuses nothing at all. Same total shutdown, same window, completely different experience for the caller - which is why interviewers ask for the order rather than the list.

  • Why keep serving requests after you have announced you are leaving the routing set?
    Because the announcement has not taken effect anywhere yet. For the length of the propagation delay, callers are still choosing this replica, and those requests have no other home. Serving them normally is the only outcome that costs nobody anything; refusing them turns a planned shutdown into caller-visible failures.
  • How would you measure the drain pause rather than guessing it?
    Mark one replica ineligible on a quiet system and keep counting requests that still arrive at it. The time until that count reaches zero is your propagation delay. Set the pause somewhat above that, and re-measure when the routing layer or the caller population changes.
  • What if the application cannot announce its own departure from the routing set?
    Some platforms can run a step of your choosing between issuing the stop request and the workload acting on it, which is a reasonable place to do it. It is not universal, so a sequence that depends on it is not portable - implementing the whole ordering inside the application is the version that works everywhere.

saying these in an interview costs you the question

  • Closes the listening socket as the first action on shutdown
  • Assumes routing changes take effect everywhere instantly
  • Thinks a connection failure is equivalent to an error response
  • Treats the drain pause as wasted time to be deleted
  • Believes finishing in-flight work is enough without leaving routing first