skip to content

Everything in your checkout path is deployed across three zones, yet losing one zone took checkout down — what class of dependency explains that?

level: seniorimportance: should knowfreq 49%

answer

  1. the path is only as spread as its worst hop
  2. identity against interchangeable copies
  3. look for the leader, the lock, the volume
  4. count instances per zone, not configuration
  5. walk the write path separately from reads

basics

~20 s

Something in the path had exactly one home in the lost zone — a single writable database, a broker node, a lock or leader, a session cache, an outbound path, or a fleet whose spread had quietly drifted into one zone. The path is only as multi-zone as its least-spread hop.

solid answer

~50 s

A request path inherits the worst spread of anything it touches, so one single-homed component makes the whole path single-zone regardless of how well the rest is distributed. The usual suspects are things with an identity rather than interchangeable copies: the writable copy of a database, a broker or coordination node, whatever holds a distributed lock or elects a leader, a session store that is not replicated, a shared volume mounted in one zone, and the outbound path everything egresses through. A second, quieter cause is drift — a scale-in or a backfill during an earlier event leaves most instances in one zone, so the deployment is nominally spread and actually is not. You find both the same way: walk every hop of the request and write path and ask how many zones can serve it and where its state lives.

go deeper

for a junior

Take away the idea that a chain is only as spread as its least-spread link, and that state — a database, a cache, a lock — is where the single home usually is. You are not expected to have hunted one down yet.

for a middle

Be able to list the usual single-homed components and explain why a leader, a lock holder or a mounted volume is different from a group of interchangeable instances that any zone can replace.

for a senior

Describe the walk: every hop of the request path and separately the write path, counting what runs per zone and naming where each hop's state lives. Bring the drift case, because it is the one that survives design review.

for a principal

Decide which single homes are worth removing and which are accepted with a written promotion window, and make per-zone distribution a standing signal rather than something rediscovered in the next post-mortem.

## Why one component decides the outcome Availability composes by multiplication down a dependency chain, not by averaging. If nine hops in the checkout path are spread across three zones and the tenth exists in one, then the path survives a zone loss only when the lost zone is not the tenth one's home. The fleet-level dashboards look excellent throughout, which is what makes this failure so common and so annoying: the thing you measured was never the thing that broke. The useful mental filter is **identity against interchangeability**. Anything with interchangeable copies survives a zone loss automatically once the failed copies are health-checked out. Anything with an identity — *the* writable copy, *the* leader, *the* lock holder, *the* mounted volume, *the* fixed address — has exactly one home at a time, and needs a mechanism to acquire a new one. ## Where the single home usually hides - **The writable copy of a database**, with read copies elsewhere that nobody had arranged to promote. - **A message broker or coordination service** running as a single node, or as a cluster whose members all landed in one zone. - **Whatever elects a leader or holds a lock** for scheduled work, so background jobs stop even though the API is fine. - **Session or cart state in a single cache**, which turns a zone loss into a visible loss of customer work rather than a blip. - **A shared file volume**, which is attached in one zone by construction and does not follow the workload to another. - **The outbound path**, if every private workload egresses through one address-translating gateway in one zone. - **A supporting service nobody calls a dependency** — configuration, feature evaluation, a licence check, an internal certificate issuer — deployed once because it 'is not in the request path', and then synchronously called by everything that is. - **A third-party call** that resolves to a single point in the region; your spread does not extend to somebody else's. ## The quieter cause: the spread is only nominal Even components designed to span zones can end up concentrated: 1. **Backfill drift.** During an earlier event, capacity was added wherever it was available; nothing rebalanced afterwards, so the fleet has been lopsided for months. 2. **Scale-in bias.** Repeated scale-in removed instances in the order the group offered them, and the surviving distribution drifted. 3. **Quorum placement.** A three-member cluster spread across two zones puts two members in one of them, so losing that zone loses the majority and the cluster stops accepting writes even though one member survives. 4. **Family availability.** A particular machine shape is not offered in every zone, so a fleet that requires it silently narrows to the zones that have it. The first three are found by counting instances per zone rather than trusting the configuration. The fourth is found by reading what the placement actually resolved to. ## How to find them before an outage does Walk the path rather than the inventory. For every hop of a request — and separately for every write — ask three questions: 1. **How many zones can serve this hop right now?** Count what is running, not what the configuration permits. 2. **Where does its state live, and is there a second copy that can take over?** If the answer names one place, you have found one. 3. **What happens to this hop's callers while it moves?** A hop that can fail over in ninety seconds still fails every request in those ninety seconds unless the caller retries. The write path deserves the separate pass because read traffic frequently survives an event that writes do not, and a checkout that can browse but not buy is still an outage. ## Fixing them in order Rank by how much of the path each one gates, not by how easy it looks. A configuration service everything calls synchronously is worth more attention than a rarely used admin tool. Some of these are cheap — spread a broker's members over an odd number of zones so a majority survives; replicate the session store; rebalance the fleet and then alarm on the per-zone distribution so the drift is visible next time. Some are expensive and deserve an explicit decision: a stateful engine with one writable copy is a real piece of work to make zone-redundant, and the honest outcome is sometimes 'we accept a promotion window here, and it is written down'. ## How to answer this out loud Name the class first — one hop with a single home makes the whole path single-homed — then give two or three concrete examples from systems you have run, then describe the walk you would do to find the rest. Mentioning the drift case will separate you from candidates who only check the deployment configuration.

  • Why does a three-member cluster spread across two zones behave badly when a zone is lost?
    Because one of the two zones holds two members. Losing that zone leaves a single member, which is short of the majority the cluster needs to keep accepting writes, so it goes read-only or stops entirely. Spreading an odd number of members across an odd number of zones keeps a majority alive whichever single zone is lost.
  • The dashboards showed every service healthy across three zones. How was that consistent with the outage?
    Because service-level health is an average and availability composes by multiplication. Nine well-spread hops and one single-homed hop produce a single-homed path, while the aggregate view still looks green. Measure the path end to end, and count capacity per zone per component rather than per service.
  • How do you stop the spread drifting back over time?
    Make the distribution a monitored property rather than an initial setting: alarm when any zone holds materially more than its share of a fleet, and check it after every event that added capacity in a hurry. Configuration that requests a balanced spread is a request, not a guarantee.

saying these in an interview costs you the question

  • Judging zone spread from configuration rather than from what is running
  • Assuming a supporting service is not a dependency because it is not user-facing
  • Spreading a three-member cluster across two zones and calling it redundant
  • Checking the read path and never walking the write path
  • Believing that a service-level health dashboard proves the path survives
  • Treating a shared volume as if it moved with the workload