skip to content

An instance deregisters itself from the service registry before it exits, yet callers keep sending it requests for tens of seconds afterwards. What determines the length of that window, and how do you shut the instance down without dropping those requests?

level: seniorimportance: should knowfreq 42%

answer

  1. a write, not a stop
  2. every hop keeps its own copy
  3. deregister first, keep serving
  4. the pause must still answer requests
  5. a crash waits for the lease

basics

~20 s

Deregistration is a write, not a stop: the removal must replicate, every caller must refresh its cached member list, and open connections outlive the entry. Keep serving after deregistering, for longer than that whole window, then exit.

solid answer

~60 s

Removing yourself from a registry only changes a record; it does not reach into the callers. The window is the sum of several independent delays: the registry applying and replicating the write to whichever replica each caller reads, the caller's own refresh — a poll interval, or watch-delivery latency — its local cache of the member list, and finally the connections it already holds open to your address, which no membership change closes. Add in-flight requests on top. The fix is ordering, not speed: deregister **first**, then keep serving normally for longer than that worst-case window, then stop accepting new work, finish what is in flight, and only then exit. The classic bug is the inverse — the process closes its listener on the shutdown signal and deregisters on the way out, so every caller that had not yet refreshed gets connection failures. And note the asymmetry: this clean path only exists for a planned exit. An instance that crashes is removed when its heartbeat lease expires, which is usually far longer.

go deeper

for a junior

Know that a service tells the registry it is leaving, and that callers do not find out instantly — so a shutdown has to give them time rather than exiting the moment it is told to stop.

for a middle

Explain what the delay is made of: the registry write and its replication, the caller's poll or watch, the caller's cached member list, and connections that are already open to the instance.

for a senior

Show the ordering that fixes it in production — deregister, keep serving through a lame-duck pause longer than the convergence window, stop accepting, finish in flight, exit — and name the inverted-order bug you have seen cause deploy-time errors.

for a principal

Own the window as a platform constant. Decide the registry's convergence targets and the callers' refresh defaults together, publish one shutdown budget for all services, and accept residual staleness rather than tightening heartbeats until healthy instances get evicted.

## Deregistration is a write, not a stop The instinct behind the bug is that removing yourself from the registry is like closing a valve. It is not. It is an update to a piece of shared state that a number of independent parties read on their own schedules. Nothing about that write reaches out and stops a caller mid-request, and nothing about it closes a socket. Until every reader has observed it, you are still, from their point of view, a member. ## The window, term by term Write down the terms and the window stops being mysterious: 1. **Apply and replicate.** The registry accepts the write and makes it durable. If it is replicated, the caller may be reading a replica that has not converged yet, so the relevant figure is not "when the write returned" but "when the replica this caller reads reflects it". 2. **Caller refresh.** How the caller learns. A polling client learns on its next poll — worst case, one full interval. A watching or streaming client learns when the push is delivered, which is fast but not instantaneous and not guaranteed while it is reconnecting. 3. **Caller-side cache.** Many discovery clients keep a local snapshot of the member list, deliberately, so that a registry blip does not black-hole traffic. That cache has its own refresh period, which stacks on top of the previous term. 4. **Connections already open.** Even after the caller drops you from its member list, a connection pool may hold an established socket to your address. Membership changes generally do not tear down live connections, so requests keep arriving on them until the pool rebuilds. 5. **In-flight work.** Requests you have already accepted and not yet answered. The total is the number your shutdown sequence has to respect. It is a property of the platform — of the registry's convergence and the callers' refresh settings — and not of your application code, which is exactly why teams that tune only the application keep seeing the errors. ## The shutdown protocol that works The correct order is: 1. **Deregister** (or otherwise mark yourself as leaving) while still fully serving. 2. **Pause** — keep accepting and answering requests normally for longer than the window above. This lame-duck period is doing the real work: it exists purely so callers can converge while you are still able to answer them. 3. **Stop accepting new work** — stop advertising, refuse new connections. 4. **Finish in flight**, bounded by a deadline so a stuck request cannot block the exit indefinitely. 5. **Exit.** Two details separate an implementation that works from one that only looks right. The pause must be *serving*, not sleeping with the listener closed — a paused-but-closed instance produces exactly the connection failures you were trying to avoid. And a shutdown must never answer arriving requests with errors as a form of "draining": that converts a clean handover into visible failures and, worse, into failures that look like the dependency being broken. ## Planned versus unplanned removal The entire sequence above only exists when the process gets to run it. If an instance is killed hard, loses its host, or wedges, it never deregisters, and its entry is removed only when the registry decides the lease has lapsed — heartbeat interval multiplied by a missed-beat threshold, plus reaping. That is typically far longer than the planned path, and it is why callers must still tolerate connection failures to instances that are currently listed. The two paths are complementary: the lease is a safety net for the case where nobody told the registry anything, not the mechanism you rely on for a rolling deploy. ## Why you cannot simply tune the window to zero Each term resists shrinking. Shorter caller poll intervals multiply read load on the registry and, at scale, turn every deploy into a thundering read. Shorter heartbeat intervals raise write load and make the registry more likely to remove a healthy instance during a brief network blip — you trade staleness for stability, and evicting healthy members is much worse than serving one for a few extra seconds. Caller-side caching, which lengthens the window, exists on purpose so that a registry outage degrades gracefully rather than emptying every member list at once. So the practical target is not zero. It is: know the number, make the lame-duck pause exceed it with margin, and make callers resilient to the residual. ## How to present this in an interview Say "deregistration is a write, not a stop", enumerate the terms, and then give the ordering. If you add one more sentence, make it the asymmetry between a clean exit and a crash — that is the detail that shows you have watched a deploy do this to a production service.

  • How would you choose the length of the pause between deregistering and refusing new work?
    Measure rather than guess: take the registry's worst-case convergence, add the slowest caller's refresh interval and any caller-side cache period, then add margin. It is a platform-wide constant, so publish it as one number every service uses instead of letting each team invent its own.
  • What happens instead when the instance crashes rather than exiting cleanly?
    No deregistration occurs, so the entry survives until its heartbeat lease expires and the registry reaps it — usually much longer than the planned window. Callers must therefore treat a connection failure to a listed member as normal and be able to move on from it.
  • Why can shortening the heartbeat interval to shrink that window backfire?
    It multiplies write traffic against the registry and narrows the margin for a transient network blip, so healthy instances start getting removed. Losing good members costs far more availability than briefly serving a dead one, which is why heartbeat intervals are deliberately conservative.

saying these in an interview costs you the question

  • Deregistering removes the instance from traffic immediately
  • Handling the shutdown signal by closing the listener is graceful shutdown
  • If the registry entry is gone, no caller can still hold the address
  • Shortening the heartbeat interval is a free way to cut the window
  • Returning errors during shutdown counts as draining

context