skip to content

Which requests does an edge tier actually take off your region's fleet, and which does it merely forward?

level: middleimportance: should knowfreq 50%

answer

  1. latency for all, capacity for some
  2. absorbed only if answerable locally
  3. weight the share by cost to serve
  4. a spike of personalised traffic passes through
  5. proximity multiplies, compute does not

basics

~20 s

It removes requests it can answer itself - a shared cached response, a redirect, a rejected malformed call. Everything personalised, transactional or freshly computed is forwarded, so the region must still be sized for that arrival rate.

solid answer

~40 s

The thin tier is an answer to distance, not a source of capacity. Requests it can satisfy from what it already holds genuinely never reach the region, and for a read-heavy shared workload that can be most of the traffic. Requests that need the authoritative dataset arrive at the region exactly as before, one for one. So the sizing question is not `how big is the tier` but `what share of live traffic is answerable without the region`. A team that plans to shrink the region fleet because the tier is in front must first show that the share it absorbs covers the requests that were driving the fleet size. If the expensive traffic is the personalised traffic, the fleet does not shrink at all.

go deeper

for a junior

Remember that only requests the tier can answer itself disappear from the region's load; everything else is passed along unchanged. Proximity is the thing the tier always provides.

for a middle

Explain the two classes and the asymmetry between them, and state that the sizing question is what share of traffic is answerable locally when weighted by the cost of serving it.

for a senior

Show the measurement you would run before promising a capacity reduction, and describe how a spike of uncacheable traffic reaches the region at full strength despite the tier.

for a principal

Weigh whether the tier is the right spend at all against a nearer region or cheaper region-side work, and set the expectation other teams should hold about what a front layer does and does not absorb.

## The claim to be careful with `Put it behind an edge tier and the region gets quieter` is true for one class of traffic and false for another, and the difference is worth being precise about, because sizing decisions get made on the strength of it. The tier answers **distance**. Its locations are numerous and small, so they add proximity cheaply. They do not add a meaningful pool of general-purpose compute, and they hold no authoritative data. Whatever cannot be answered locally is forwarded, and the region does exactly the work it would have done anyway. ## Requests that truly leave the region's load - **A response valid for many users.** Served from what the location already holds; the region never sees the request. - **A redirect or a rewrite** decidable from the request alone. - **A malformed or clearly unauthorised request** rejected at the front door on a cheap check. - **A repeat of something the location has recently been given**, for as long as that copy remains valid. For a read-heavy workload with a large shared surface, this is a real and sometimes dramatic reduction in arrivals. It is not a trick or a rounding error, and denying it is its own mistake. ## Requests that are merely forwarded - **Personalised responses**, because the answer differs per user. - **Anything transactional**, because the write has to land where the data durably lives. - **Freshly computed answers**, because the inputs are in the region. - **Anything requiring an exact view of shared state**, because no location has one. For these, the arrival rate at the region is unchanged by the tier's existence. Adding more locations does not change it either - proximity multiplies, capacity does not. | Traffic class | Effect on region load | Effect on user-visible latency | |---|---|---| | Shared, cacheable response | Removed entirely while valid | Large improvement | | Redirect or cheap rejection | Removed entirely | Large improvement | | Personalised read | Unchanged, one for one | Setup cost improves only | | Durable write | Unchanged, one for one | Setup cost improves only | | Heavy computation | Unchanged, one for one | Setup cost improves only | ## The sizing conversation The useful question is a measurement, not an opinion: **what share of live requests, weighted by the cost of serving them, is answerable without the region?** Three outcomes follow. 1. **A large share, and it is the expensive share.** The fleet can genuinely shrink, and the saving is real. 2. **A large share, but it is the cheap share.** Arrival counts fall and the fleet does not, because the remaining requests were always the ones setting the size. 3. **A small share.** Latency improves for everyone through connection termination, and the fleet is untouched. The second outcome is the one that surprises teams. A dashboard showing a high proportion of requests answered at the tier feels like an enormous win, and it may still leave the region's capacity requirement exactly where it was, because the requests that dominate the fleet's sizing are precisely the ones nothing else could answer. ## What this means during a traffic spike The tier is also a partial shock absorber, with the same asymmetry. A spike of shared, cacheable traffic is largely absorbed. A spike of personalised or transactional traffic passes straight through to the region, arriving at full strength, and what protects the region then is the region's own scaling and admission controls, not the tier's width. Designing on the assumption that the front layer will flatten every spike is how a team discovers this during the spike rather than before it. ## A check you can run today Take an hour of live traffic and label each request with the class it belongs to, then weight it by the cost of serving it. The labels are usually available already: whether the response varied per caller, whether it wrote anything, and whether it was served without the region being asked at all. The result is a single number - the share of serving cost the tier can remove - and it is the only input that makes a fleet-sizing claim defensible rather than hopeful. ## Saying it precisely The clean formulation is: the thin tier answers latency for everything and capacity only for what it can answer itself. Both halves of that sentence matter. Drop the first half and you claim the tier is useless for dynamic work, which is wrong. Drop the second and you plan a fleet reduction that the traffic will not support.

  • A dashboard shows most requests answered at the tier, yet the region's fleet has not shrunk. Why?
    Because share of requests is not share of cost. The absorbed requests were the cheap, shareable ones, while the personalised and transactional requests that actually set the fleet size are still arriving one for one. Re-weight the measurement by cost to serve and the picture usually explains itself.
  • Does adding more edge locations increase how much load is taken off the region?
    Only marginally. More locations mean more users are near one, which improves latency, but each location starts with nothing and has to be given a copy of what it serves. Width buys proximity; it does not buy a larger share of answerable traffic, which is a property of the workload.
  • Can the tier protect the region during a traffic spike?
    Partly, and asymmetrically. A spike concentrated on shared, cacheable responses is largely absorbed. A spike of authenticated, personalised or transactional requests arrives at the region at full strength, so the region still needs its own scaling and admission controls rather than relying on the front layer.

saying these in an interview costs you the question

  • Plans to shrink the region fleet because personalised traffic now passes through a tier
  • Says the tier adds compute capacity for work that must return
  • Claims the tier absorbs nothing because requests are forwarded anyway
  • Counts absorbed requests without weighting them by cost to serve
  • Assumes the front layer will flatten any traffic spike