skip to content

Cross-AZ data transfer between chatty internal services is the largest line on your AWS bill. An engineer proposes making every service call only same-Availability-Zone instances of its dependencies. How would you evaluate that proposal?

level: principalimportance: nice to knowfreq 30%

answer

  1. measure the pairs before pinning anything
  2. locality spends the smoothing, not spare change
  3. prefer local, but always spill over
  4. small pools have worse tails
  5. fewer bytes beats fewer zones

basics

~20 s

Zone-local routing removes a real per-gigabyte charge but trades away the load spreading and headroom that multi-AZ deployment buys. Quantify the saving first, then require per-zone capacity headroom and an automatic fallback to remote zones before accepting it.

solid answer

~50 s

It is a legitimate lever with a real availability cost, so I would treat it as a design change rather than a config tweak. First, **quantify**: attribute cross-AZ bytes to specific service pairs from flow logs, because the saving is usually concentrated in two or three chatty paths and the rest is not worth touching. Second, **check what breaks**: zone-local routing means a zone's capacity must absorb that zone's traffic, so any imbalance — uneven instance counts, an unlucky hash, one large tenant — now shows up as latency instead of being smoothed across the fleet. Third, **insist on fallback**: preferring local while still failing over to remote zones on health-check failure keeps the resilience story intact; hard pinning does not. I would apply it to the few heaviest paths, keep per-zone headroom, and consider the cheaper alternative first — sending fewer bytes through compression, caching, or moving a chatty pair into the same process.

go deeper

for a junior

Know that traffic crossing Availability Zones is billed per gigabyte in both directions, so where services sit relative to each other has a cost, not just a latency, consequence.

for a middle

Be able to explain why random cross-zone spreading exists — it smooths load across a larger pool — and why removing it couples each zone's capacity to that zone's traffic.

for a senior

Show the working method: attribute the bytes to specific service pairs, reduce the traffic before restricting the routing, and implement locality as a preference with automatic spill-over rather than a hard pin.

for a principal

Own the trade explicitly. Put the saving in currency against the availability property being spent, refuse it for durability and quorum flows, and make sure the organisation does not end up single-AZ in practice while claiming multi-AZ on paper.

## What the proposal is really trading Cross-AZ traffic inside a Region is billed per gigabyte and billed on both ends. For a service mesh where every request fans out to several dependencies chosen at random across three zones, roughly two-thirds of internal traffic crosses a zone boundary — so the charge scales with request volume and fan-out, not with fleet size, and it grows every time someone adds a hop. Zone-local routing removes that charge by preferring dependency instances in the caller's own zone. What it removes along with the charge is the **statistical smoothing** that random cross-zone spreading provides. That smoothing is not free insurance you can cancel; it is doing work. ## What breaks, concretely **Capacity coupling.** With zone-local routing, zone A's dependency fleet must be able to serve all of zone A's traffic. If autoscaling produced 4 instances in A, 7 in B and 6 in C — a routine outcome — zone A is now hot while spare capacity sits in B and C, unreachable. The effective utilisation ceiling drops, and you buy back part of the saving in extra instances. **Imbalance becomes visible.** Uneven traffic per zone, a large tenant hashed to one zone, or a slow instance in a small zone-local pool now shows up as tail latency instead of being diluted across a larger pool. Small pools have worse tails than large ones for the same reason. **Failure behaviour changes.** If the zone-local dependency pool is unhealthy, a hard pin means failure; a *preference* with fallback means degraded cost, not degraded service. Anything that pins must define what happens when the local pool is empty or failing, and the answer must be "spill to another zone", not "error". **Deployment blast radius.** A bad deploy that rolls zone by zone now takes down all of one zone's traffic instead of a third of everyone's — which can be better (contained) or worse (a whole zone's users see it), depending on how the front door routes. ## The order I would actually work in 1. **Measure before designing.** Attribute cross-AZ bytes to service pairs from flow logs. The distribution is almost always heavily skewed: a replication stream, a cache-fill path, or one fan-out call is most of the bill. Fixing three paths gets most of the saving with a fraction of the risk. 2. **Try sending fewer bytes first.** Compression on internal RPC, a local read-through cache, returning less per call, batching, or collapsing a chatty pair of services into one process all reduce the charge without touching the availability model. These are strictly better trades when available. 3. **Then apply locality selectively**, as a preference with automatic spill-over, on the heaviest paths only — not as a fleet-wide default. 4. **Fund the headroom.** Require per-zone capacity balance and enough headroom that losing one zone's local pool degrades cost rather than service. 5. **Re-measure.** Confirm the saving landed and that per-zone tail latency did not move. ## Where locality is usually the wrong answer For **stateful** systems the calculation inverts. A database's synchronous replica in another zone exists precisely because you want the data to survive that zone; the cross-AZ transfer is the price of durability and is not negotiable. The same goes for a quorum that must span zones. Cost-optimising those flows is optimising away the reason you are multi-AZ at all. Some flows are also priced differently by the service in front of them, so check the specific path rather than assuming: load balancers, for example, differ in whether spreading traffic across zones carries a transfer charge and in whether cross-zone spreading is on by default. Verify how the actual path is billed before predicting a saving. ## The judgment being tested The interviewer wants to see whether you treat a cost lever as an architecture decision. The weak answer is "yes, pin everything, cross-AZ is expensive." The strong answer names the saving in currency, names what resilience property is being spent, proposes a preference-with-fallback rather than a pin, and reaches for reducing the bytes before reducing the redundancy. Cost is a legitimate constraint; it is not a reason to quietly become a single-AZ system with a multi-AZ diagram.

  • How would you decide the saving is worth the engineering and risk at all?
    Put a number on it. Attribute cross-AZ bytes per service pair, convert to a monthly figure at the current rate, and compare it against the extra per-zone headroom you will have to fund plus the engineering time. If the answer is a few hundred dollars a month against a meaningful availability trade, the correct decision is to leave it alone.
  • What would make you refuse the proposal for a specific flow outright?
    When the cross-zone traffic is the durability or quorum mechanism — synchronous replication to a standby in another zone, or a consensus group that must span zones. That traffic exists so the system survives losing a zone. Removing it does not optimise the design, it deletes the property the design was bought for.
  • How do you keep zone-local routing from silently degrading into single-zone dependence?
    Make the fallback exercised rather than theoretical: run regular zone-evacuation drills, alarm on per-zone capacity skew, and keep the routing a weighted preference so a small share of traffic always crosses zones and proves the path works. A failover path that has never carried traffic is a hypothesis, not a capability.

saying these in an interview costs you the question

  • Pins traffic zone-locally with no fallback path
  • Assumes cross-AZ transfer is negligible at any scale
  • Applies locality fleet-wide instead of to the heaviest paths
  • Optimises a synchronous replication flow that exists for durability
  • Claims a saving without attributing bytes to service pairs first

context