A per-user quota counter and a deduplication record live only in each region's own volatile tier while callers can be routed to either region - what breaks?
answer
- ask what else could rebuild this entry
- no other home, no independence
- two regions means two counters
- the retry lands where the record is not
- one home, or one cross-region anchor
basics
~20 sNeither entry is a copy of anything, so two independent tiers hold two different facts about one user: the quota is granted once per region, and a retry landing in the other region finds no record and is treated as new.
solid answer
~50 sIndependence between regional tiers is safe only for entries either region can rebuild from a shared source. A quota counter and a deduplication record are not copies of anything - they *are* the state - so running one per region silently multiplies them. The user gets roughly one full quota in each region, and a payment retried into the other region finds no record marking it handled and can be taken a second time. The same failure hits leases and claims: two regions can each believe they hold one. There are three honest repairs: give that state a single home and route it there, move that one class of key to a store authoritative across regions and accept the hop for its small volume, or decide deliberately that per-region semantics are correct for it - which is defensible for a rate limit and never for "charged once".
go deeper
Ask one question of any entry before you accept it being per-region: if this disappeared, could anything else rebuild it? A quota count and a handled-request marker cannot be rebuilt.
Explain how the failure shows up: independent counters sum into more allowance than intended, and a retry routed to the other region finds no record and repeats the work.
Separate the key classes rather than the regions. Keep derived entries local and independent, and give the small set of sole-copy keys a single home or a cross-region anchor, paying the hop only for them.
Own the standing rule for which state classes may live sole-copy in a regional tier, and require an owner and a review for each one. This is where a regional design quietly turns into double-charged customers.
## Derived state and sole-copy state are different problems The case for **an independent tier per region** rests on a single premise: the entry is a locally derived view of something with another home, so two regions holding different values costs nothing. Sole-copy ephemeral state breaks that premise outright. A quota counter, a record marking a request as already handled, a lease on a resource, a session - these are not copies. **They are the state itself**, and the tier is the only place it exists. Run one tier per region and you have not distributed that state. You have created two of it. ## What each kind of sole-copy entry does wrong across regions - **Quota and rate counters.** Each region counts only what it saw. A user routed across both regions gets roughly the sum of the per-region allowances, and the overshoot grows with how evenly traffic is spread. Nothing detects it, because from each region's point of view the count is correct. - **Deduplication records.** A request marked handled in region A leaves no trace in region B. A retry - a client retry, a queue redelivery, an impatient user - that lands in the other region sees a key that was never written and does the work a second time. This is the expensive one: it is how a payment gets taken twice. - **Leases and claims.** Two regions can each hand out the same exclusive claim, because neither asked the other. Note that this is not the partition-time failure of one tier with two writable sides; it is the ordinary, healthy steady state of a design that never intended the two regions to talk. - **Sessions.** A user served in the other region is simply unknown there and is treated as unauthenticated, which is a visible outage for that user even though every component is healthy. The common shape: **the same mechanism produces an annoyance for derived entries and an incident for sole-copy state**, and nothing in the design distinguishes them. Only you can. ## Three repairs, and how to choose 1. **Give the state one home and route to it.** Pin the identity - user, account, resource - to a home region, and send requests that touch its sole-copy state there. This keeps the tier cheap and local for everything else, and its cost is that a home-region loss takes that state with it, which you must decide about in advance. 2. **Move that key class to a store authoritative across regions.** Pay the cross-region hop deliberately for the small volume of keys that need it - the claim, the deduplication record - while ordinary derived entries stay local. Here the hop buys correctness, not speed, so the arithmetic that rejected cross-region lookups elsewhere comes out the other way. 3. **Redefine the requirement as per-region.** Sometimes the honest answer is that per-region semantics are fine. A rate limit expressed as "per region" is a legitimate product decision if someone makes it knowingly. "Charged once" and "held by exactly one worker" never survive that treatment. | state | safe as one tier per region? | why | |---|---|---| | derived view of a stored row | yes | either region rebuilds it from the shared source | | rate counter | only if per-region limits are an accepted decision | independent counters sum | | deduplication record | no | the retry that lands elsewhere sees nothing | | lease or exclusive claim | no | two regions can each grant it | ## Where writes land decides more than it looks One fact usually settles repair 2 cheaply. If writes to the underlying system of record are accepted in only one region, then every write already crosses a region, and the correctness anchor can live on that same path at little extra cost - a uniqueness constraint or a conditional write in the durable engine, with the tier used only to keep the common case off it. If instead each region writes its own system of record, you do not have one shared truth to anchor to, and the sole-copy state has to be designed for explicitly rather than assumed into place. ## What varies across stores Do not assume a product-level way out. Some stores in this class offer copies propagated between regions, a few of them accepting writes in more than one region with machinery to make copies converge; others offer nothing. Even where it exists, converging two counters is not the same as enforcing one limit, and converging two claims is not the same as granting one. **Convergence makes copies agree eventually; exclusivity requires that only one grant ever happened.** Those are different guarantees, and the second is the one a claim needs.
- Is a per-region rate limit ever the right answer rather than a defect?Yes, when someone decides it knowingly and the limit exists to protect capacity rather than to promise the user a number. Capacity is per region, so a per-region limit matches what it protects. It is wrong whenever the limit is a commitment - a plan allowance, a free-tier count - because the user experiences the sum.
- Why is this not simply the split-brain problem in a new setting?Split brain is a failure: two sides of one tier both accept writes during a partition, and it ends when the partition heals. This is the healthy steady state of a design whose regions were never meant to talk, so nothing detects it, nothing recovers from it, and it is present on the best day the system ever has.
saying these in an interview costs you the question
- Treats every entry as reproducible from somewhere else
- Thinks a shorter lifetime makes sole-copy state safe across regions
- Assumes two regional counters sum themselves on read
- Believes a retry always returns to the region that served it
- Confuses eventual convergence with exclusive ownership
- Calls this split brain rather than a design that never shared state