skip to content

A chatty checkout tier is now split across zones — what does each cross-zone hop add, and what sets that floor?

level: middleimportance: should knowfreq 55%

answer

  1. distance sets a floor you cannot tune
  2. count round trips per request
  3. a metro hop, not a continental one
  4. chattiness multiplies the hop
  5. batch, or keep reads zone-local

basics

~20 s

A cross-zone hop costs a metro round trip — propagation over the physical distance plus switching, typically well under a few milliseconds. The floor is distance and cannot be tuned away; the damage comes from chattiness, one request paying that hop many times.

solid answer

~50 s

Splitting a tier across zones turns some in-zone calls into metro calls. The added cost per call is small and largely fixed: signal in fibre covers tens of kilometres in a fraction of a millisecond each way, plus switching, so a cross-zone round trip lands well under a few milliseconds while a same-zone one is smaller still. What makes it visible is **fan-out**: a request that makes forty sequential calls pays the difference forty times, and a difference under a millisecond becomes tens of milliseconds of user-visible latency. So the number to look at is round trips per request, not the hop. The fixes are to reduce round trips — batch, cache in the calling zone, keep a read replica zone-local — rather than collapsing the tiers back into one zone. The same arithmetic, at tens to hundreds of milliseconds per hop, is why the equivalent split across regions is a different design altogether.

code

pseudocode · 12 lines
pseudocode
# assumed, qualitative figures; every call below crosses a zone boundary
sameZoneHopMillis  = 0.2
crossZoneHopMillis = 0.8

callsPerRequest = 40                                  # sequential, each waits
before = callsPerRequest * sameZoneHopMillis          # 8.0 ms of network time
after  = callsPerRequest * crossZoneHopMillis         # 32.0 ms of network time
added  = after - before                               # 24.0 ms per request

# same work, batched into fewer round trips
callsPerRequest = 4
after = callsPerRequest * crossZoneHopMillis          # 3.2 ms of network time

go deeper

for a junior

Know that crossing into another zone is a real network hop with a real cost, small but not zero, and that the cost is paid once per call rather than once per deployment.

for a middle

Explain the composition of the hop — propagation over a metro distance plus switching — and show that sequential round trips per request are what turn a sub-millisecond cost into a visible one.

for a senior

Diagnose from evidence: latency split by caller and callee zone, sequential depth per request, tail rather than mean, and a fix that reduces round trips instead of undoing the spread.

for a principal

The judgment here is which call paths are allowed to cross a zone boundary at all, and whether the team's service interfaces are coarse enough that a placement change does not become a latency incident.

## Where the latency comes from A cross-zone hop is made of three things, and only one of them is under your control: - **Propagation.** Signal in fibre travels at roughly two-thirds the speed of light. Zones sit a metro distance apart, so each direction costs a fraction of a millisecond. This is physics: no amount of tuning, protocol choice or instance sizing reduces it. - **Switching and queuing.** Each device on the path adds a small, mostly stable amount, growing under load. - **Work at the far end.** Whatever the callee actually does, which is the same wherever it runs. The first two together are the **floor** — the smallest possible cost of putting a boundary between caller and callee. Same-zone calls pay a smaller floor; cross-region calls pay a floor of tens to hundreds of milliseconds, because the distance is three or four orders of magnitude greater. ## Why chattiness is the multiplier The hop is small. The pattern is not. If a checkout request makes many **sequential** calls — each waiting for the previous answer — every one of them pays the floor: 1. Count the round trips a single request makes to the tier that moved. 2. Multiply by the per-hop difference between a same-zone and a cross-zone call. 3. Compare the product to the request's latency budget. With an assumed same-zone hop of a fifth of a millisecond, an assumed cross-zone hop of eight tenths, and forty sequential calls, the tier goes from about eight milliseconds of network time to about thirty-two — twenty-four milliseconds added, all of it before anyone does any work. Batch the same work into four calls and the cross-zone total falls to roughly three milliseconds. **The chattiness moved the number, not the zones.** Those figures are assumed and qualitative; the ratio is the point. Calls made in **parallel** behave very differently: ten concurrent cross-zone calls cost roughly one hop, not ten, though they raise the tail because the slowest of the ten decides the answer. That is why the honest measurement is *sequential depth per request*, not total call count. | Boundary crossed | Typical round trip | What it permits | |---|---|---| | Same zone | The smallest floor available | Chatty patterns survive, badly but survivably | | Another zone, same region | A metro hop, well under a few milliseconds | Synchronous replication and a spread tier | | Another region | Tens to hundreds of milliseconds | Asynchronous copies with a lag window | ## What to measure before changing anything - **Round trips per request** to the tier that moved — the single highest-leverage number. - **Latency broken down by caller zone and callee zone.** If every zone pair looks alike, the spread is not the problem; if one pair is an outlier, the path is. - **Tail latency, not the mean.** Fan-out raises the tail first, because a request waits for its slowest dependency. - **How much of the request's time is network at all.** A tier where work dominates will barely notice the move, and a tier that was already chatty will notice enormously. ## What not to conclude **Do not conclude that the hop is free.** Zones are separate facilities with real distance between them; "it is all one region" does not make the network cost zero, and cross-zone traffic is also metered on its own schedule, which is a separate subject from latency. **Do not conclude that the hop is a region-sized cost.** Treating a metro hop as if it were a continental one leads teams to abandon zone spread entirely, which is a large loss for a small and usually fixable number. **Do not conclude that tuning timeouts helps.** A propagation floor is not a timeout problem. Shortening a timeout below the floor produces failures, not speed. **Do not conclude that collapsing back into one zone is the fix.** The chatty pattern was already costing you inside the zone; the zone boundary only made it legible. Reducing round trips helps in both layouts, and it is the change that survives the next move. ## Where this answer stops The subject here is what the cross-zone hop costs in **time** and what governs that cost. What those same bytes add to the bill is the charge model's material, and how you design the system to survive losing the zone at the far end belongs to the reliability material. The latency story is the one this placement decision owns.

  • Why does the same reasoning rule out a synchronous version of this split across regions?
    Because the floor grows by orders of magnitude. A metro hop under a millisecond can be paid a few times inside a request budget; a continental hop of tens to hundreds of milliseconds cannot be paid even once on a synchronous write path without changing what the product feels like. That is why cross-region copies are asynchronous and carry a lag window.
  • Ten cross-zone calls per request are issued concurrently rather than one after another. What changes?
    The added latency collapses to roughly one hop instead of ten, because the calls overlap. What worsens is the tail: the request waits for the slowest of the ten, so a single slow path now decides the response time. Measure sequential depth for the mean and fan-out width for the tail.

saying these in an interview costs you the question

  • Says cross-zone latency is zero because it is all one region
  • Treats a cross-zone hop as a cross-region one, tens of milliseconds
  • Blames the hop rather than the round trips per request
  • Thinks shorter timeouts remove a propagation floor
  • Concludes the only fix is collapsing the tiers back into one zone