skip to content

Edge Presence

A wide, thin tier of small locations in front of a few large regions: what can run there, and what still has to go back to a region. Asked because teams reach for it to fix the wrong problem.

on this pageshow

questions

5

An edge location sits in front of your single region for far-away users - what can it absorb, and what must travel back?

level: juniorimportance: must knowfreq 68%

answer

  1. wide and thin, in front of few and large
  2. proximity, not capability
  3. answerable from the request alone
  4. truth and durability stay in the region
  5. replicated to many sites, so keep it small

basics

~20 s

An edge location can terminate the connection and TLS, serve a cached response, issue a redirect, and make a cheap decision from the request itself. The authoritative dataset, durable writes and heavy computation still travel back to the region.

solid answer

~40 s

A provider's footprint has a few large regions and, in front of them, a wide thin tier of many small locations close to users. The thin tier is built to end the client's connection nearby and hand back something it already has or can decide in microseconds: a cached response, a redirect, a rewritten path, a cheap check on a header or token. What it deliberately does not hold is the authoritative state - the database, the search index, the object store in the region - or the heavy compute that reads them. So a request that needs a fresh, personalised or transactional answer is forwarded, and the region still does that work. The useful mental model is `many small locations, few large regions`: the small ones absorb distance, the large ones own truth.

go deeper

for a junior

Remember the two layers and their shapes: a few large regions that hold the data, and many small locations in front that hold connections and cached answers. Name three things the small layer can do without asking anyone.

for a middle

Explain the property that decides the split - work answerable from the request plus locally-held material stays; work needing the authoritative dataset, a durable write or heavy compute goes back. Say why replication cost, not just CPU, enforces that.

for a senior

Show that you check which class a real request falls into before proposing a move, and that you can say what fraction of live traffic is genuinely edge-answerable rather than assuming it.

for a principal

Frame it as a placement standard other teams can apply: what an application is allowed to assume exists at the thin tier, and what must remain a region responsibility so the estate does not accumulate state nobody can reason about.

## Two layers on the same map A large cloud provider's footprint is not one uniform thing. A **region** is a large installation - many buildings, several independent failure domains inside it, and the full catalogue of services, including the ones that store data durably. A provider operates a comparatively small number of them, because each is expensive and each is a serious commitment of land, power and staff. In front of those regions sits a second layer with the opposite shape: a **wide, thin tier** of many small locations, often called edge locations or points of presence, placed close to concentrations of users. There are far more of them than there are regions, and each one is small. Its job is proximity, not capability. The practical consequences of being small and numerous: - Very little durable storage, and none you should treat as authoritative. - Modest compute, sized for short work measured in milliseconds, not minutes. - A partial, local view of the world - each location knows what it has seen, not what the others have. - A footprint that is replicated across a large number of sites, so every capability you add there is a capability you pay for many times over. ## What the thin tier can genuinely absorb 1. **Connection and transport-security termination.** The client's handshake completes against a machine a few milliseconds away instead of one on the other side of an ocean. 2. **A response it already holds.** If the same bytes are valid for many users, the tier can return them without involving the region at all. 3. **A redirect or a rewrite.** Sending the client somewhere else, or reshaping a path, needs nothing but the request. 4. **A cheap check on the request itself.** Verifying that a signature or token is well formed and unexpired, or rejecting an obviously malformed call, uses only what arrived plus a key the tier already has. 5. **A brief request-time decision.** Short logic that reads the request, decides, and either answers or forwards. The common property is that each of these is answerable from the request plus something the location already holds. None of them needs to agree with any other location, and none of them writes anything that has to survive. ## What still has to go back to a region - **The authoritative dataset.** A database, a search index, an object store - anything whose current contents are the answer. - **Durable writes.** An acknowledgement that must still be true after the machine that gave it has gone. - **Heavy computation.** Ranking, aggregation, model inference over a large dataset, report generation. - **Anything needing a consistent view of shared state.** An exact balance, an exact limit, a lock. | Work | Where it lands | Why | |---|---|---| | Connection and TLS termination | Edge location | Needs only a certificate and key, and the win is proximity | | A response valid for many users | Edge location | Already held locally; no shared state consulted | | Redirect, rewrite, malformed-request rejection | Edge location | Decidable from the request alone | | Personalised or freshly-computed response | Region | Needs the authoritative dataset | | Durable write | Region | The acknowledgement has to survive | | Heavy ranking or aggregation | Region | Too much CPU and too much data to replicate widely | ## Why the split falls exactly there Two forces put the line where it is. The first is **physics**: the only thing a location near the user can give you is a shorter round trip, so it helps precisely with work that can finish without asking anyone far away. The second is **economics of replication**: whatever you place at the thin tier, you place in every site in it. A large mutable dataset cannot be copied to hundreds of small locations and kept correct without the coordination traffic costing more than the distance you were trying to avoid. ## How to say it in an interview Say what the tier is for and what it is not for, in one breath: it ends the connection close to the user and answers from what it has; the region keeps the truth and does the heavy work. Then name the failure mode that matters - a team that expects the thin tier to carry their application will discover that every request that needs real data was forwarded anyway, and that the region is doing exactly as much work as before.

  • Does the thin tier replace the region for users who are close to it?
    No. It fronts the region, it does not substitute for it. A user close to an edge location still reaches the region for anything that needs the authoritative dataset or a durable write - the location simply makes the parts that can finish locally finish locally, and shortens the setup cost of everything else.
  • Why not simply run more regions instead of a thin tier?
    A region is a large, expensive installation with durable storage and a full service catalogue, so a provider can only justify a limited number of them. The thin tier exists because proximity is cheap to buy in small increments and capability is not. They solve different problems: distance against capability.
  • If a decision at the edge needs a key or a small rule set, is that still edge-shaped work?
    Yes, provided the material is small, changes rarely, and is the same everywhere. Verifying a signature against a public key, or matching a path against a compact rule set, is answerable from the request plus something already present. The moment the decision needs a per-user or per-object lookup, it is region work.

The thin tier is the front counter of a shop and the region is the warehouse behind it. The counter can hand over what is already on the shelf and answer a quick question, but anything else is a trip to the back.

saying these in an interview costs you the question

  • Thinks the application database can be hosted at an edge location
  • Describes the edge tier as replacing the region rather than fronting it
  • Expects a durable write acknowledged at an edge location to survive
  • Treats an edge location as a smaller region with the same services
  • Believes heavy ranking or reporting work can run at every location
open as a page

Why does terminating the client connection at a nearby edge location cut time to first byte for an uncacheable response?

level: middleimportance: must knowfreq 58%

basics

~20 s

Connection setup costs several round trips before any request bytes move. Terminating nearby pays those against the short hop, while the forwarded request crosses the long distance once over a connection the edge already holds warm.

open as a page

Which requests does an edge tier actually take off your region's fleet, and which does it merely forward?

level: middleimportance: should knowfreq 50%

basics

~20 s

It removes requests it can answer itself - a shared cached response, a redirect, a rejected malformed call. Everything personalised, transactional or freshly computed is forwarded, so the region must still be sized for that arrival rate.

open as a page

A search tier in one region is slow for distant users - what part of a search request can move to the edge, and what cannot?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Connection termination, a cheap token check, a redirect and a repeat of an identical popular result can move. The index, the ranking that reads it and any write path cannot - so measure setup, crossing and region service time first.

open as a page

Why is an exact per-user counter a poor fit for a tier of many small edge locations?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Each location sees only its own traffic, so an exact shared count needs the locations to agree - and agreeing costs the round trip the tier existed to avoid. Exact, durable counting belongs in a region.

open as a page