skip to content

A dependency review finds several platform services with a single global home sitting under all three of your regions — what posture do you set?

level: principalimportance: should knowfreq 38%

answer

  1. independence is capped by the shared floor
  2. regions still cover regional failures
  3. posture per dependency, not one answer
  4. bootstrap dependencies are not harmless
  5. the honest sentence beats partial engineering

basics

~20 s

Independence is capped by the least independent thing underneath, so decide per dependency: accept and document it, move it off the request path, define a deliberate degraded mode, or genuinely replace it. The lead's deliverable is that decision plus an honest availability story, not the removal of a dependency that is not yours.

solid answer

~50 s

Start by refusing the two easy answers. 'Regions make us safe' is false, because three regions cannot be more independent than a service all three consult that has one home. 'Multi-region is therefore pointless' is also false, because the regions still cover every failure that is genuinely regional, which is most of them. What you actually set is a **per-dependency posture**: accept and document where the exposure is small, take the dependency off the request path so only bootstrap needs it, define what the product does in a deliberate degraded mode, or replace the capability with something you or another party operates — in roughly that order of cost. Then write the honest availability story down, including which outages the design does not cover, and make naming single-home dependencies a required part of design review.

go deeper

for a junior

Know that running in several regions does not make you independent of the platform itself. Some of what the platform provides exists in one place conceptually, and everything you run depends on it.

for a middle

Explain why independence is bounded by the most shared dependency underneath, and distinguish a dependency consulted on every request from one consulted only while a process is starting up.

for a senior

Show the per-dependency reasoning: where in the lifecycle it is consulted, how long a held answer can be trusted, and what the product does without it. Pick the cheapest posture that changes the outcome.

for a principal

Own the honest availability sentence and the standard behind it — every design claiming multi-region resilience names its single-home dependencies and the posture chosen for each. Re-read it as the estate takes on new managed capabilities.

## What a single global home means Large platforms are not uniformly regional. Some capabilities are deliberately global — there is conceptually one of them, it is authoritative, and although it is served from many places its control and its source of truth live in one home. Providers differ in which capabilities they build this way and in how much of the serving path is replicated outwards, but every large platform has some, and a tenant cannot see the internal topology from outside. The consequence is blunt and worth stating in one sentence to whoever commissioned the multi-region design: **three regions cannot be more independent than the least independent thing underneath them.** If all three consult a service with one home, that service's failure is not a regional failure at all. It is a correlated failure across every region you run in, and across every other tenant at the same time. Two overcorrections follow and both are wrong. The first is to conclude that the regions are worthless — they still protect you from power, network, cooling and capacity failures confined to one place, which is the large majority of what actually happens. The second is to conclude that the global dependency only matters if it sits on the request path. A bootstrap dependency is what stops a replaced process from returning, and during a long incident that is the difference between degraded and gone. ## The four postures There are only four honest responses, and the lead's job is to assign one to each dependency rather than to find a single answer for all of them. | Posture | What it costs | When it is right | |---|---|---| | Accept and document | Nothing to build; a written, honest exposure | The capability is rarely used, or the product survives without it | | Move off the hot path | Modest engineering; holding values and bounded staleness | The dependency is consulted often but its answer changes slowly | | Deliberate degraded mode | Product decisions and real work | The capability is central and a reduced product is genuinely useful | | Replace the capability | The most expensive, and ongoing | A small number of dependencies where the exposure is unacceptable | A fifth option — asking the provider to make it regional — is a roadmap request, not a posture. It leaves the estate exactly as the review found it. ## Choosing per dependency Run each finding through the same three questions: - **Where in the lifecycle is it consulted?** On every request, on connect, or only at bootstrap. Bootstrap-only is much cheaper to live with, and is also the one people wrongly dismiss. - **How long can a held answer be trusted?** If the answer changes slowly, holding it with a bounded staleness window converts a hard dependency into a soft one for the length of that window. That is the single highest-value move available and it is usually cheap. - **What does the product do without it?** If the honest answer is 'nothing', you have found either a candidate for replacement or a limit you must state out loud. ## What you owe the business The deliverable that matters is not a mitigation plan; it is an accurate sentence about what the design buys. Something close to: *we survive the loss of a region, we do not survive the loss of these named shared platform services, and for those our exposure is a reduced product for the duration.* That sentence is worth more than any partial engineering, because it is what other people's plans get built on. The failure mode this review exists to prevent is an organisation that believes it is insulated from provider failure because it paid for three regions. Be equally honest about the limits of your own knowledge. You cannot see the provider's internal dependency graph, so the inventory is necessarily incomplete and is built from observed behaviour — what each process contacts before it serves its first request, and what stopped working during the last incident. Treat every provider incident as free discovery and add what it revealed to the list. ## The standard you set The durable output is a rule rather than a fix: a design claiming multi-region resilience must name the shared platform services it depends on and state, for each, what happens when that service is unavailable and which posture was chosen. It is a short table, it takes an hour, and it converts an assumption that nobody wrote down into a decision somebody owns. Pair it with a periodic re-read, because the answer changes as the estate takes on new managed capabilities — each one arriving with its own dependencies that nobody has looked at yet.

  • Which of the four postures usually gives the most resilience for the least money?
    Moving the dependency off the hot path. If the looked-up answer changes slowly, holding it and trusting it for a bounded window converts a hard dependency into a soft one for the length of that window, which is often longer than the incident. It costs a cache with an explicit staleness bound and loud logging, and it does not require any product decision or any second supplier.
  • How do you build the inventory when you cannot see the provider's internal topology?
    From behaviour and from history. List what every process contacts before it serves its first request, including what the platform client libraries do that no application code mentions, and record which capability each supports. Then add what every past provider incident actually took out — those are free observations of the real dependency graph, and they are the only ones you can fully trust.
  • How do you answer an executive who asks why three regions did not protect you?
    Directly: the regions protect against failures confined to one place, and this was not one. A capability all three depend on has a single home, so its loss is correlated across all of them. Then give the decision rather than the excuse — which dependencies you are accepting, which you are moving off the hot path, and what the product does during the window.

saying these in an interview costs you the question

  • Claims multi-region deployment removes exposure to provider-wide failures
  • Overcorrects to declaring the extra regions worthless after the review
  • Dismisses a bootstrap-only dependency as harmless during a long outage
  • Offers a provider roadmap request as if it were a posture
  • Presents a partial mitigation instead of an honest statement of exposure
  • Assumes the dependency inventory is complete because it came from a diagram