Your product runs an independent volatile tier per region and one region is lost, with its traffic shifted to the survivor - what should the team expect?
answer
- separate what refills from what is gone
- sole-copy state has no second source
- ask where the writes were being accepted
- one tier now holds two populations
- blast radius is decided months earlier
basics
~20 sThree separate losses, not one: derived entries for the arriving population are absent and get rebuilt; sole-copy state held only in the lost tier is unrecoverable; and if that region accepted the writes, the write path is a second, larger failure.
solid answer
~40 sSeparate what refills from what is gone. The survivor's tier holds only its own population's working set, so the arriving traffic is served from the underlying system of record until the tier has seen it - extra load, temporary, and it must be sized. Sole-copy state that existed only in the lost region's tier is unrecoverable: claims released, deduplication records forgotten, sessions invalid, quota counts reset. Third, ask where writes landed. If the lost region accepted writes to the system of record, the survivor can keep serving derived reads while nothing can be written anywhere - a very different incident from a read-path shortfall. Capacity is the fourth: a tier sized for one population now holds two and will evict its way down unless that was planned for.
go deeper
The surviving region's tier only holds what its own callers asked for, so the arriving users are unknown to it. Their requests go to the underlying system of record until the tier has seen them.
Distinguish the entries that rebuild themselves from the entries that do not exist anywhere else. The first group is a slow period; the second is lost claims, forgotten deduplication records and logged-out users.
Add the two questions that change the incident's shape: where writes to the system of record were being accepted, and whether the survivor is sized for two populations or expected to degrade.
Cap the blast radius before the day. Rule on which state classes may live sole-copy in a regional tier, give each a named consequence for losing it, and rehearse the shift at combined load rather than trusting the plan.
## A region loss is several losses, and they are usually reported as one When a design runs **an independent tier per region**, losing a region removes a tier, a population's worth of state, and possibly a write path. Treating it as "the tier failed over" hides the decisions that matter. Take the four apart: 1. **Derived entries for the arriving population are simply absent.** The survivor's tier was filled by the survivor's traffic, so the incoming keys were never looked up there. Those requests fall through to the underlying system of record until the tier has seen them. This is recoverable and temporary - but it is load, and the load arrives at the same moment as the extra traffic itself. 2. **Sole-copy state in the lost region's tier is gone permanently.** A claim on a resource, a record marking a request as already handled, a quota count, a session - if the lost tier was the only place it existed, nothing anywhere can reproduce it. This is the loss nobody writes down in advance. 3. **The write path may have gone with the region.** If writes to the system of record were accepted only in the lost region, the survivor's tier can keep answering derived reads perfectly while the product is unable to accept a single write. That is a far larger incident than a cold tier, and it has its own recovery, on its own timeline. 4. **Capacity on the survivor has silently halved per user.** A tier sized for one population is now asked to hold two. It will evict its way back down to its ceiling, which means the arriving population and the existing one compete, and the existing population's hit ratio falls too. ## What is actually gone, in user-visible terms | what the lost tier held | after the shift | user-visible effect | |---|---|---| | derived entries with a shared source | rebuilt on demand in the survivor | a slower period, extra load on the engine | | deduplication records | gone, unreproducible | in-flight retries can be executed twice | | leases and claims | gone, unreproducible | work re-picked up, or briefly held by nobody | | sessions, if held only there | gone, unreproducible | that population is logged out | | quota counters | gone, unreproducible | allowances effectively reset for those users | The first row is an operational inconvenience. Every other row is a correctness or trust event, and the difference is decided months earlier, by whether anyone allowed sole-copy state into a regional tier. ## The questions to have answered before the day - **Which state classes are permitted to live sole-copy in a regional tier at all, and who approved each one?** This is the decision that sets the blast radius; everything else is consequence. - **Where are writes to the system of record accepted, and is that the same region for every product surface?** If the answer varies by surface, the region loss is several different incidents at once. - **Is the survivor sized for both populations, or sized for its own and trusted to degrade?** Both are legitimate answers. Only one of them is a decision. - **What does the product do while the arriving population is being served without the tier?** Shedding some load deliberately is often better than letting the system of record take everything. - **Which of the lost sole-copy state has a safe re-creation story?** A session can be re-established by asking the user to sign in again. A payment taken twice cannot be undone by anything technical. ## What varies Products differ on whether cross-region propagation is even an option, and where it exists it changes the first row of the table and possibly the last few - at the price of a cross-region path on writes during normal operation, every day, to improve one bad hour. That is a real trade and not an obvious one. It is also worth being clear that propagation between regions makes copies agree; it does not by itself give a claim exclusivity or make a counter a single counter, so it repairs less of this table than it appears to. ## The posture worth writing down Cap the blast radius rather than trying to make the failure invisible. In practice that means: keep regional tiers holding derived state only, by rule; give each class of sole-copy state a single explicit home with a named consequence for losing it; know which region accepts writes for each surface; and rehearse the shift at both populations' load, because the survivor's capacity and the engine's behaviour under the rebuild are the two things that are never as predicted.
- Does propagating copies between regions make this failure go away?It improves the first row and costs a cross-region path on writes every ordinary day, to improve one bad hour. It also repairs less than it appears to: propagation makes copies agree, but it does not give a claim exclusivity or turn two counters into one. Sole-copy state still needs a single home and a named consequence.
- The survivor's tier is sized for its own population. Size it for both, or let it degrade?Both are defensible; the failure is not choosing. Sizing for both means paying for idle memory in normal operation. Letting it degrade means the existing population's experience drops too, because the tier evicts its way back to its ceiling. Decide it in advance, write down which one you picked, and rehearse at that load.
saying these in an interview costs you the question
- Reports the region loss as one failure rather than several
- Expects sole-copy state to come back with the region
- Ignores where writes to the system of record were accepted
- Assumes the survivor's tier absorbs both populations unchanged
- Thinks cross-region propagation would have made it a non-event
- Plans the shift without rehearsing at the combined load