skip to content

questions

6

A user in Tokyo requests a static image from a website whose origin server sits in Virginia, USA, and the site is served through a CDN with points of presence (PoPs) worldwide. Explain what happens on the very first request for that image versus the hundredth request from a different Tokyo user, and why the CDN makes the site faster.

level: juniorimportance: must knowfreq 75%

answer

  1. PoP = point of presence
  2. cache hit vs cache miss
  3. TTL governs freshness
  4. distance drives round-trip latency
  5. origin offload as a side benefit

basics

~20 s

A CDN stores copies of files on servers (PoPs) close to users everywhere. The first request has to fetch and cache the file from the real origin server, which is slow. Later requests are served from the nearby copy, which is fast because the data travels a much shorter distance.

solid answer

~40 s

The CDN runs many PoPs distributed globally, each caching content near end users. The Tokyo user's request is routed to the nearest PoP. On the first request that PoP has no copy of the image (a cache miss), so it pulls the object from the origin in Virginia, stores it locally according to its TTL rules, and returns it — paying the full origin round-trip latency once. Every subsequent request that hits that PoP within the TTL window is a cache hit: the PoP serves the bytes directly from local storage, cutting round-trip time from 200ms+ trans-Pacific/trans-continental down to a few milliseconds, and simultaneously taking load off the origin.

go deeper

for a junior

Should know what a PoP is, the difference between a cache hit and a cache miss, and that CDNs mainly help static content.

for a middle

Should be able to explain TTL, name a header that reveals cache status, and describe pull-based caching end to end.

for a senior

Should discuss the cold-PoP penalty, cache-warming implications, and origin offload as a scaling and cost lever, not just a latency trick.

for a principal

Should reason about global footprint design trade-offs — PoP count, tiering, and freshness SLAs — as an architectural decision tied to cost and availability goals, not just a vendor feature toggle.

## What a CDN actually is A **Content Delivery Network (CDN)** is a globally distributed set of caching servers, called **points of presence (PoPs)**, positioned in many cities so that a copy of frequently requested content lives physically close to the people requesting it. The site's real server, the **origin**, holds the authoritative copy of everything; the CDN sits in front of it. When a client resolves the site's hostname, DNS (or, for some CDNs, anycast routing at the IP layer) directs the request to whichever PoP is closest to that client by network path, not necessarily by straight-line distance. ## First request versus the hundredth 1. **Cache miss.** The first time any PoP is asked for a given object it has never seen, it has nothing cached — this is a cache miss. 2. **Origin round trip.** The PoP itself acts as a client to the origin, fetches the object over the CDN's own backbone (often faster and more direct than the public internet path a browser would take), stores a copy according to the Cache-Control/TTL rules the origin supplied, and finally returns the object to the requester. That first request therefore still pays close to the full origin round-trip cost. 3. **Cache hit.** Every later request that lands on the same PoP while the copy is still fresh is a cache hit: the PoP answers directly from local storage without contacting the origin at all. ## Why it exists This exists because network latency is bounded by physics — light (and therefore any signal) takes a minimum amount of time to cross a given distance, and a request/response cycle over TCP/TLS typically needs several round trips (connection setup, TLS handshake, then the actual HTTP exchange) before the first byte comes back. A round trip between Tokyo and Virginia might cost 150–250ms; a round trip between a Tokyo user and a Tokyo-area PoP might cost single-digit milliseconds. Multiplying several round trips by that difference is the entire reason CDNs exist for latency-sensitive content: moving the data physically closer to the reader removes most of that unavoidable travel time. As a side effect, it also protects the origin — instead of every single request worldwide hitting one server, only the relatively rare cache misses do, so the origin can be sized for a fraction of total traffic. ## The trade-off The core trade-off is **freshness versus speed**. - A **longer TTL** (time-to-live, the duration a cached copy is considered valid before the PoP must re-check with origin) means more cache hits and less origin load, but it also means changes at the origin take longer to reach users, since a PoP serving a stale-but-not-yet-expired copy has no way to know the underlying content changed. - **Very short TTLs** keep content fresh but push traffic back toward the origin, eroding much of the latency benefit. - There's also a **cold PoP** cost: a CDN with, say, 200 PoPs doesn't magically have every object everywhere — content only gets cached at a PoP after that PoP's own first miss, so a brand-new region or a rarely-visited object can still pay full origin latency on its first hit even though the object has been cached elsewhere for hours. - Running a large PoP fleet is also **operationally expensive**: more infrastructure to monitor, secure, and keep in sync with cache-invalidation logic. ## Failure modes Failure modes show up in a few recognizable ways. - **Misconfigured TTLs or cache keys.** A PoP can serve outdated or even wrong content (e.g., an old version of a page) well past when it should have refreshed, and different users in different regions can legitimately see different versions of the same page during a rollout, which is confusing if not anticipated. - **A PoP itself goes down.** The routing layer needs to redirect its users to the next-nearest healthy PoP or fall back toward origin; if that fallback path isn't well designed, users nearest the failed PoP see the worst latency of anyone. - **A very popular new object.** When it is requested for the first time simultaneously by users spread across many regions, each region's PoP independently misses and calls origin at roughly the same moment — a smaller, per-object version of a thundering herd, which is one of the reasons larger CDN deployments add an origin-shielding tier (a separate concern from this one). ## Where it shows up A concrete real-world example: an e-commerce or media site puts its product images, CSS, and JS behind a CDN such as Cloudflare, Akamai, or Fastly. A shopper in Tokyo loading the homepage for the first time triggers a handful of cache misses at the Tokyo PoP, each one paying a one-time trip to the US origin; every other shopper who visits from the Tokyo area afterward, for as long as the TTL holds, gets those same assets served in single-digit milliseconds from the local PoP instead of crossing the Pacific. Video-heavy platforms take this even further — Netflix's Open Connect appliances are placed directly inside ISP networks so that video segments barely leave the user's own network at all.

  • What's the difference between a pull-based and a push-based CDN?
    A pull-based CDN (the common model) caches lazily — a PoP only fetches an object from origin the first time it's requested, as described above. A push-based CDN has the origin proactively upload content to PoPs ahead of time, guaranteeing it's already warm everywhere before the first user request, which is useful for content you know will spike immediately (like a scheduled release) but requires more upfront coordination and storage.
  • How would you tell, from an HTTP response, whether a request was served from cache or went to origin?
    Many CDNs add a diagnostic header such as X-Cache: HIT or MISS, or an Age header showing how many seconds the response has sat in cache; the absence or freshness of these headers is the practical way to debug whether a slow response is a cache miss or something else entirely.
  • What happens if the nearest PoP to the Tokyo user is unavailable?
    The routing layer (DNS-based or anycast) detects the outage and steers the request to the next-nearest healthy PoP; that PoP may itself have a cold cache for the object, so the user pays a cache-miss penalty even though the underlying content was already cached at the now-unavailable PoP, until it warms up again.

A CDN is like a chain of neighborhood bakeries restocked from one central factory: the first customer at a brand-new branch might have to wait for a delivery truck from the factory, but every customer after that just grabs bread off the shelf instead of driving across the country to the factory.

saying these in an interview costs you the question

  • Says the CDN 'stores the whole website forever' without mentioning TTL or expiration
  • Thinks the CDN replaces the origin server entirely
  • Doesn't realize a cache miss still has to contact the origin
  • Believes a CDN speeds up dynamic, uncacheable API responses the same way it speeds up static assets
  • Can't explain why physical/network proximity reduces latency

context

open as a page

Your team ships a new JS bundle under the same URL /app.js on every deploy, and users keep complaining they're stuck on stale versions after a release even though the file changed on the server. What CDN caching mistake is likely happening, and how would you fix Cache-Control and the deployment strategy to solve it for both this JS bundle and images that rarely change?

level: middleimportance: must knowfreq 80%

basics

~20 s

The file's cache setting is probably telling browsers and the CDN to keep the old copy for too long, and reusing the same filename gives no signal that anything changed. Fix: give changed files new filenames whenever the content changes (like app.abc123.js) and cache those forever, while keeping the URL that points to the latest version very short-lived.

open as a page

A CDN advertises the exact same IP address, say 203.0.113.10, from data centers in Frankfurt, Singapore, and Sao Paulo simultaneously via BGP, and relies on the internet's normal routing to send each user to the nearest one. Explain how this works, and describe a failure scenario where a user's connection to that IP breaks mid-session even though all three data centers are healthy.

level: seniorimportance: must knowfreq 50%

basics

~20 s

Anycast means many servers around the world share one IP address, and the internet's normal routing (BGP) automatically sends each user to whichever one is 'closest' by network path. The problem: if the network path changes mid-connection, a user can suddenly get routed to a different server than before, breaking anything that depended on talking to the same machine.

open as a page

A news website publishes a breaking-news article, then five minutes later the editor fixes a factual error in the headline. Readers in different countries keep seeing the old, wrong headline for varying amounts of time after the fix. What CDN mechanisms would you use to get the correction out quickly and reliably, and what's the trade-off of using them aggressively for every edit?

level: middleimportance: should knowfreq 55%

basics

~20 s

Tell the CDN to throw away its cached copy of that page (purge/invalidate) so the next request fetches the fixed version from the real server. Doing this for every tiny edit is slow to spread everywhere and dumps extra load back on the server, so it should be used sparingly and precisely, not for every small change.

open as a page

A CDN with 200 edge PoPs serves a video-on-demand site whose new episode just dropped. At the moment of release, hundreds of PoPs simultaneously experience their first cache miss for the same file within the same second. Without any extra configuration, what happens to the origin, and what's the standard CDN feature that prevents it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

All those PoPs would hit the real server at once and could overwhelm it. The fix is 'origin shielding': one extra layer of caching servers sits between the edge PoPs and the origin, so only that shield fetches from origin once, and every edge PoP gets the file from the shield instead.

open as a page

A team wants to run personalized A/B test bucketing logic on every request at the CDN edge (for example, using Cloudflare Workers or a similar edge-compute platform) instead of in their origin application server. What kinds of logic are a good fit for that edge-compute layer, and what would make you push back and say 'this belongs in the origin, not the edge'?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Edge compute runs small pieces of code on the CDN's servers near the user, so simple, fast decisions like which test group to show, or rewriting a URL, can happen instantly without a trip to the main server. It's a bad fit for anything needing a big database, heavy computation, or long-running work, because edge servers are deliberately limited and don't reliably keep state.

open as a page