skip to content

A service expands into a second region but leaves its volatile tier in the first: what does each lookup now cost?

level: middleimportance: must knowfreq 54%

answer

  1. compare the hop against the work avoided
  2. distance sets the floor, not the store
  3. tens of milliseconds versus a fraction of one
  4. a miss pays the hop and the work
  5. one tier per region, filled locally

basics

~20 s

Each lookup turns into a cross-region round trip - tens of milliseconds instead of a fraction of one - which can exceed the cost of the work the tier was avoiding, making the remote path slower than having no tier.

solid answer

~50 s

A volatile tier earns its place by arithmetic: the round trip to it plus the lookup must cost decisively less than the work it replaces. Distance is the term a second region changes. Inside one data centre the round trip is a fraction of a millisecond; between regions it is tens of milliseconds, set by the route and the speed of light in fibre, not by the store. So a caller a region away can often answer its own question locally, from the durable engine, faster than the remote tier answers it - and on a lookup that finds nothing it pays the hop and then the work anyway. That is why the usual design is an independent tier per region, each filled from its own region's traffic, with the underlying system of record as the only shared truth.

go deeper

for a junior

Remember that a volatile tier is fast partly because it sits next to the caller. Put a region between them and the network, not the store, decides how long a lookup takes.

for a middle

Be able to state the comparison out loud: round trip plus lookup against the cost of the work being replaced. Then show that a cross-region round trip can lose that comparison to a plain local query.

for a senior

Decide this on the arithmetic rather than on a preference for consistency. Show the miss path too, where the remote caller pays the hop and then the work, and name the new cross-region availability dependency you just created.

for a principal

Set the rule for the organisation: which state classes may ever be reached across a region, and what a team must measure before adding a cross-region dependency to a request path that was previously contained in one region.

## What the tier is buying, stated as arithmetic A store whose medium is memory is not valuable because memory is fast in the abstract. It is valuable because a whole exchange - the caller's request, a lookup that involves no disk seek and no query planning, and the reply - completes inside one data centre. Running one is therefore a bet with a very specific shape: **the round trip to the tier, plus the lookup, must cost decisively less than the work it replaces** - a query against the durable engine, a computation, a call to another service. A tier that is marginally cheaper than the work is not worth its operational weight. A tier that is an order of magnitude cheaper is the reason this component exists at all. ## What a second region does to the bet Adding a region changes exactly one term, and it is the term the bet rests on. Inside a data centre a round trip is a fraction of a millisecond. Between regions it is tens of milliseconds, rising with distance, and it is set by the route packets actually take and the speed of light in fibre - not by the store, the protocol, the client library, or the amount of memory involved. No product choice removes it. | the caller does | rough order of magnitude | what sets it | |---|---|---| | lookup against a tier in the same data centre | a fraction of a millisecond | the store and the local network | | query the durable engine in the same data centre | single-digit milliseconds, workload dependent | the engine and the query | | lookup against a tier in another region | tens of milliseconds, growing with distance | geography | Put those rows together and the premise inverts: **a caller a region away from the tier can often answer its own question locally, from the durable engine, faster than the remote tier can answer it**. The tier has stopped being an optimisation and become an extra, slower dependency. Three further effects sharpen it: - On a lookup that finds nothing, the remote caller pays the cross-region round trip **and then** the local work. The hop is pure loss. - A request path that issues several lookups multiplies the hop by that count, which is why paths that were comfortable locally are the first to fall apart. - The remote path now fails whenever the link between regions does, so a request that could have been served entirely inside one region has acquired a cross-region availability dependency. ## The design the arithmetic forces The common answer is **an independent tier per region**: 1. Each region runs its own tier, sized for its own traffic and reached only by callers in that region. 2. Each tier fills from its own region's traffic, so the first caller in each region pays the underlying work once. 3. The underlying system of record is the only thing the two regions share; each tier is a locally derived view of it. The two tiers are not kept in agreement, and that is deliberate. The moment a write to one region's tier has to reach the other, the cross-region hop is back on the path - now on the write side, where it is harder to hide behind a fast local read. ## What varies, and what does not The geography does not vary. Almost everything else does: - Some products in this class offer copies propagated between regions, a few of them accepting writes in more than one region with machinery to make the copies converge; others offer one-way propagation only; others offer nothing and expect separate deployments. **Do not assume the option exists, and do not assume what it guarantees.** - Where propagation does exist, it changes what each tier holds. It does not change what the hop costs, and a caller still reads its own region. - Whether the durable engine behind the tier is itself present in both regions is a separate decision with its own arithmetic, and it is usually the one that settles where writes land. ## When crossing a region is still the right call Arithmetic decides this, not a preference for locality. Reaching across a region is defensible when: - the work being replaced is genuinely expensive - a multi-second computation, a costly third-party call, an aggregation over a large table - so that even tens of milliseconds is a large saving; or - the entry is not a copy of anything and has to be unique across all regions - a claim, a deduplication record, a quota - where the hop buys correctness rather than speed, and the alternative is simply being wrong. Outside those two cases, the honest answer to "how do we reach the tier from the new region?" is that you do not: you give the new region its own. ## The same lookup on three paths. The figures are illustrative orders of magnitude, not measurements, and the point is only the ordering: once the tier is a region away, the local engine is the faster path ``` caller in region A, tier in region A t0 caller -> tier: lookup key t0+0.4ms tier -> caller: value total: about 0.4 ms caller in region B, tier still in region A t0 caller -> tier: lookup key (crosses a region) t0+38ms tier -> caller: value total: about 38 ms same caller in region B, no tier, durable engine in region B t0 caller -> engine: query t0+6ms engine -> caller: rows total: about 6 ms ```

  • When is reaching across a region for the tier actually the right call?
    When the hop is small next to what it replaces, or when it buys correctness rather than speed. A multi-second computation or an expensive third-party call still comes out far ahead at tens of milliseconds. So does an entry that must be unique across all regions - a claim, a quota, a deduplication record - where a local copy per region would simply be wrong.
  • Does putting a copy of the tier in the new region fix the arithmetic?
    It fixes the read hop only. Keeping the copies in agreement puts a cross-region exchange back on the write path, and the entries a region reads may still be behind. Products differ sharply here: some offer propagation between regions, some accept writes in more than one, some offer neither - so this is a purchasing question before it is a design question.

A reference book on your own desk is worth keeping because looking something up beats working it out. Phone a library on another continent to read one line and the call takes longer than working it out yourself - unless what you are working out would have taken all afternoon.

saying these in an interview costs you the question

  • Thinks memory speed survives the distance between regions
  • Prices the cross-region hop as a few extra milliseconds
  • Assumes the tier is worth reaching for no matter what it replaces
  • Forgets that a lookup finding nothing pays the hop and then the work
  • Treats the remote tier as free availability rather than a new dependency
  • Never compares the hop against the local query it was avoiding