skip to content

A volatile tier's promotion finished in seconds, yet one service kept addressing the dead primary for an hour — what did it get wrong?

level: seniorimportance: should knowfreq 52%

answer

  1. promotion moves the address, not the caller
  2. retrying is not rediscovering
  3. resolved once at start-up
  4. cached name resolution outlives the failover
  5. restart-to-recover is the diagnosis

basics

~20 s

Promotion only moves which node accepts writes; it does not reach into callers. A service that resolved the write address once at start-up and never again keeps talking to the dead node, however clean the failover was.

solid answer

~50 s

A failover changes the tier's answer to "who is the primary?" — it changes nothing inside a caller that never asks the question again. The broken service had pinned its address: a literal in configuration, a name resolved once at start-up, or a client that reconnects only to the endpoint it was constructed with. The fix is on the caller side and is about *when* it re-resolves, not how hard it retries: treat a connection failure as a signal to look the write address up again, cap how long a resolved address is trusted, and confirm the runtime is not caching name resolution forever. Where the fix lives depends on the arrangement — a caller holding its own address must refresh it, a caller that asks the deciding party must be configured with that party rather than with a node, and a caller behind a stable re-pointed address has nothing to do.

go deeper

for a junior

Remember that promoting a replica changes the tier, not your application. If your service was told an address once and never asks again, it keeps calling a node that is gone.

for a middle

Explain the three ways a caller can learn the write address, and why retrying the same address forever is not rediscovery. Name cached name resolution as a common cause.

for a senior

Diagnose it: a service that recovers only on restart has pinned its address. Prescribe re-resolution on failure, a bounded trust period for a resolved address, and a real failover drill with every consumer watched.

for a principal

Make it a standing requirement rather than a per-incident fix. Decide whether callers discover addresses at all or are kept behind one stable address, and require a periodic failover drill as the evidence that the answer still holds.

## The window nobody configured After a promotion, three things are true at once: the new primary is accepting writes, the old node is not, and every caller is still sending wherever it was sending before. Closing that gap is **address discovery**, and it is the only one of the failover intervals that lives outside the tier entirely. A textbook-clean promotion plus a caller that never asks again equals an outage of unbounded length, which is exactly the hour in the scenario. ## How a caller can learn the write address There are three shapes, and a deployment uses one of them. Confusing them is why advice about this subject so often does not apply to the system in front of you. | Shape | How the caller resolves | What goes stale | Where the fix lives | |---|---|---|---| | Caller holds the address or a map | Resolves once, then refreshes on some trigger | The caller's own copy | In the caller: refresh on failure, and bound how long a copy is trusted | | Caller asks the deciding party | Queries the observers or controller for the current primary | Little, if it asks on failure | In the configuration: point the client at the deciding party, never at a node | | Stable address in front of the tier | Always uses one address; a proxy or controller re-points it behind the scenes | Nothing the caller sees | In the component that re-points, which you now operate | In the first shape a node may also answer that the address it holds is no longer the right one — the store telling the caller that what it wants lives elsewhere — but only a caller that is still able to reach some node gets that hint, and a caller talking to a corpse gets nothing but connection errors. ## The four ways a service gets stuck 1. **An address literal in configuration.** Someone put the primary's address in a deployment file. Nothing will ever update it. 2. **Resolution cached for the process lifetime.** The service used a name, but resolved it once at start-up, or the runtime caches name resolution indefinitely by default. The name was re-pointed; the process never noticed. 3. **A pool that reconnects to the same endpoint.** The connections failed and were re-established — to the node they were configured with, which is dead. 4. **Retry without re-resolution.** The client retries diligently, with backoff, against the same address. Retrying is not rediscovery; how the retry itself should be shaped is a caller-side concern of its own. A fifth case is worth naming because it is the nastiest: if the old node comes back and has not been demoted, a caller that never re-resolved does not even fail. It succeeds, against a node that is no longer the primary. ## What to do about it - **Re-resolve on failure, not only at start-up.** A connection error should invalidate the cached write address, not just trigger another attempt at it. - **Bound the trust.** Give a resolved address a maximum age, so even a caller that never errors eventually asks again. - **Check the runtime's name-resolution caching.** This is a per-platform default that is frequently far longer than anyone assumes, and it silently defeats the re-pointing shape. - **Configure clients with the deciding party, where that is the arrangement.** A client that is given the address of a node has no way to discover a new one; a client given the observers or the controller does. - **Drill it.** Fail the primary over deliberately and watch every caller. The rediscovery window is the only failover interval you cannot read out of configuration, so it is the only one you have to measure. A service that needs a restart to recover has just told you it has this bug. ## The point to make in an interview Separate the two halves cleanly: promotion is the tier's business, rediscovery is the caller's, and a failover report that only covers the first half is not a failover report. Then name which of the three address-discovery shapes the system uses, because the fix is in a different place for each: in the caller, in the client's configuration, or in the component that owns the stable address.

  • Why is a client that retries with backoff still stuck here?
    Because retrying repeats the same destination. Recovery needs the address to be looked up again, so the failure has to invalidate the cached address rather than only schedule another attempt. Retry shaping and rediscovery are separate concerns, and a client can be excellent at one and have none of the other.
  • Which failover arrangement makes this failure impossible for the caller?
    The one where callers use a single stable address that something in front re-points — an intervening proxy or a managed controller's endpoint. The caller never learns a node address, so it has nothing to hold stale. The trade is that the re-pointing component is now in your path and is the thing that has to be right.
  • How do you find services with this bug before an incident does?
    Run a deliberate failover in a non-production environment with every consumer running, and watch which ones recover on their own. Any service that only works again after a restart has pinned its address. Doing this on a schedule is what keeps the answer current as new consumers appear.

saying these in an interview costs you the question

  • Thinks a completed promotion means callers are already talking to the new primary
  • Says retry with backoff is enough to recover from a failover
  • Resolves the write address once at start-up and caches it for the process lifetime
  • Configures clients with a node's address where a deciding party exists
  • Never ran a failover drill against real consumers