A transcoder's cross-zone traffic charge now exceeds its compute bill after the tiers were split across zones — what happened, and what do you measure first?
answer
- the bill followed the calls, not the machines
- a boundary appeared inside the request path
- requests times hops times payload
- a small rate on an enormous internal volume
- zone-aligned paths, replicas still everywhere
basics
~20 sEvery inter-tier call now crosses a metered zone boundary, so the charge scales with calls multiplied by payload, not with machine count. Measure gigabytes crossing a boundary per request before changing any part of the design.
solid answer
~40 sSplitting the tiers across zones did not add machines; it added a **priced boundary** to the middle of every request. With callers landing on any zone, most inter-tier hops now cross it, so the charge tracks `requests x hops per request x payload per hop`, a number that has nothing to do with how the compute is sized. A media pipeline shuttling intermediate frames is the worst possible shape for that formula: the payloads are large and the hops are many. Measure first, in this order: gigabytes crossing a zone boundary per request, which service pair produces them, and what fraction of hops cross at all. Only then choose between **zone-aligned request paths**, fewer hops, smaller payloads and compression. Right-sizing the machines changes nothing, because the same bytes still cross.
code
pseudocode · 20 lines// illustrative only - no rate here comes from any price list
rateCrossZone = r // charge per GiB moved across a zone boundary
rateInternetOut = manyTimes(r) // internet egress costs far more per GiB
requestsPerMonth = 40000000
innerCallsPerRequest = 6
payloadPerCallGiB = 0.002 // ~2 MiB of intermediate frames
zones = 3
fractionCrossingZones = (zones - 1) / zones // caller lands anywhere = 0.66
crossZoneGiB = requestsPerMonth * innerCallsPerRequest
* payloadPerCallGiB * fractionCrossingZones // = 316800 GiB
transferCharge = crossZoneGiB * rateCrossZone
// machineCount and machineSize appear nowhere in this expression,
// so right-sizing cannot move transferCharge.
// Only fewer calls, smaller payloads, or fractionCrossingZones -> 0 can.
if routeInnerCallsToSameZone then
fractionCrossingZones = 0
transferCharge = 0 // replicas stay in every zone; only the path is pinnedgo deeper
Recall that traffic between zones can be charged even though it never leaves the region, and that the charge counts gigabytes rather than requests or machines.
Explain the formula behind the line: requests, times hops per request, times payload per hop, times the fraction of hops that cross. Then say why changing machine size leaves every term of it untouched.
Demonstrate the measurement order — bytes per crossing per request, then the service pair, then the crossing fraction — and pick a fix whose resilience cost you state out loud. Zone-aligned routing over collapsing to one zone is the answer that shows judgment.
Make it a standing rule rather than an incident: chatty inter-tier paths that fan out across failure domains are a recurring priced decision, so the default routing posture and a unit-cost-per-request target belong in the platform standard other teams build against.
## What actually changed Nothing about the workload got bigger. The tiers were spread across zones for resilience, and a **metered boundary** appeared in the middle of a path that used to be local. Before the split, a request entered the front tier and called the transcode tier over traffic that stayed inside one zone — normally unmetered. After it, a caller can land in any zone and its downstream hops land in any zone, so most hops cross. That matters because the transfer charge is not a function of anything you usually tune: - It does **not** track machine count, machine size or utilisation. - It does **not** track CPU seconds, which is what the compute line tracks. - It **does** track bytes across a boundary — requests, multiplied by hops per request, multiplied by payload per hop, multiplied by the fraction of hops that cross. A media pipeline is the pathological case: intermediate renditions are large, and a job touches them several times. ## The arithmetic, with the rates left qualitative Take 40 million requests a month, six inner calls per request, about 2 MiB of intermediate payload per call, three zones and callers landing anywhere, so roughly two thirds of hops cross a boundary: - 40,000,000 x 6 = 240,000,000 inner calls - 240,000,000 x 0.002 GiB = 480,000 GiB moved between tiers - 480,000 x 0.66 = **about 316,800 GiB — roughly 317 TiB crossing a zone boundary every month** The cross-zone rate is a small fraction of the internet rate per gigabyte, and it still produces a line larger than the compute for this service, because the volume is two orders of magnitude above what the service delivers to users. **That is the whole lesson: a low rate on an enormous internal volume beats a high rate on a small external one.** ## Measure before you change anything 1. **Gigabytes per boundary crossing, per request.** Divide the transfer line by request count. This single number tells you whether you have a payload problem or a hop-count problem. 2. **Which service pair produces them.** Attribute the flow to a pair of tiers, not to a region total. Without this you will optimise the wrong hop. 3. **What fraction of hops actually cross.** If callers land randomly across three zones, expect roughly two in three; if it is near 100%, something is pinning traffic across the boundary and that alone may be the fix. 4. **What the same bytes would cost if the path stayed in one zone.** That is the size of the prize, and it bounds how much redesign is worth doing. ## The fixes, in order of what they cost you - **Zone-aligned request paths.** Keep replicas of every tier in every zone, but route a request's inner hops to the tier instance in the same zone. Cheapest win by far, and it keeps the failure domains you bought. - **Fewer hops.** Six inner calls per request is a design choice; collapsing chatty round trips into one call removes crossings outright. - **Smaller payloads.** Pass a reference to the intermediate in shared storage instead of the intermediate itself, where the storage read is on the cheaper side of the boundary question. - **Compression on the hot pair.** The meter counts bytes on the wire, so compressing a compressible payload reduces the charge directly, at the price of CPU on both ends. - **Collapsing to one zone.** It removes the charge and it removes the resilience you split for. This is a trade, not an optimisation, and it should be stated as one. ## What does not fix it - **Right-sizing or fewer machines.** The same bytes cross the same boundary; you have changed a different meter. - **A term commitment on compute.** A discount on the compute line does nothing to a transfer line, and buying one to "deal with the bill" locks in spend on the part that was not the problem. - **Moving one tier to another region.** That replaces a cross-zone rate with a dearer cross-region one for the same flow. - **Caching at the edge of the system.** It helps what you deliver outward; the flow here never leaves the region. ## The judgment being tested The interviewer is checking whether you reason from the **meter** rather than from habit. Bill surprises in a distributed service are usually charged at a boundary a diagram does not draw, and the correct first move is measurement that attributes bytes to a crossing and a service pair. The second thing being checked is honesty about the trade: zone-aligned routing keeps most of the resilience, collapsing to one zone does not, and saying which one you chose and why is the answer a senior engineer is expected to give.
- Pinning every request's inner hops to one zone removes the charge. What does that cost you?In-flight requests in a failing zone are lost rather than served across the boundary, so you trade graceful degradation for cost. You keep the important part — replicas of every tier in every zone, so the zone's share of traffic fails over — but the request path no longer spans zones, and that difference has to be stated explicitly rather than assumed away.
- Why doesn't right-sizing the machines reduce this bill?The charge is per gigabyte crossing a boundary, and fewer or smaller machines move exactly the same bytes across exactly the same boundary. Right-sizing moves the compute line. Only fewer crossings, fewer bytes per crossing, or a path that stays inside one zone moves the transfer line.
- The team proposes compressing the inter-tier payload. When is that the wrong call?When the payload is already compressed media, where the CPU is spent for almost no byte reduction and both tiers pay for it, or when the hop count rather than the payload size is the driver. Measure gigabytes per request first: a hop-count problem is fixed by removing round trips, not by shrinking each one.
saying these in an interview costs you the question
- Blames the compute line when the traffic line is what grew
- Says spreading across zones is free because it is one region
- Proposes collapsing to a single zone without naming the resilience lost
- Measures only total transfer, never gigabytes per request per crossing
- Claims compression cannot help because the charge is per request