A user in Tokyo requests a static image from a website whose origin server sits in Virginia, USA, and the site is served through a CDN with points of presence (PoPs) worldwide. Explain what happens on the very first request for that image versus the hundredth request from a different Tokyo user, and why the CDN makes the site faster.
answer
- PoP = point of presence
- cache hit vs cache miss
- TTL governs freshness
- distance drives round-trip latency
- origin offload as a side benefit
basics
~20 sA CDN stores copies of files on servers (PoPs) close to users everywhere. The first request has to fetch and cache the file from the real origin server, which is slow. Later requests are served from the nearby copy, which is fast because the data travels a much shorter distance.
solid answer
~40 sThe CDN runs many PoPs distributed globally, each caching content near end users. The Tokyo user's request is routed to the nearest PoP. On the first request that PoP has no copy of the image (a cache miss), so it pulls the object from the origin in Virginia, stores it locally according to its TTL rules, and returns it — paying the full origin round-trip latency once. Every subsequent request that hits that PoP within the TTL window is a cache hit: the PoP serves the bytes directly from local storage, cutting round-trip time from 200ms+ trans-Pacific/trans-continental down to a few milliseconds, and simultaneously taking load off the origin.
go deeper
Should know what a PoP is, the difference between a cache hit and a cache miss, and that CDNs mainly help static content.
Should be able to explain TTL, name a header that reveals cache status, and describe pull-based caching end to end.
Should discuss the cold-PoP penalty, cache-warming implications, and origin offload as a scaling and cost lever, not just a latency trick.
Should reason about global footprint design trade-offs — PoP count, tiering, and freshness SLAs — as an architectural decision tied to cost and availability goals, not just a vendor feature toggle.
## What a CDN actually is A **Content Delivery Network (CDN)** is a globally distributed set of caching servers, called **points of presence (PoPs)**, positioned in many cities so that a copy of frequently requested content lives physically close to the people requesting it. The site's real server, the **origin**, holds the authoritative copy of everything; the CDN sits in front of it. When a client resolves the site's hostname, DNS (or, for some CDNs, anycast routing at the IP layer) directs the request to whichever PoP is closest to that client by network path, not necessarily by straight-line distance. ## First request versus the hundredth 1. **Cache miss.** The first time any PoP is asked for a given object it has never seen, it has nothing cached — this is a cache miss. 2. **Origin round trip.** The PoP itself acts as a client to the origin, fetches the object over the CDN's own backbone (often faster and more direct than the public internet path a browser would take), stores a copy according to the Cache-Control/TTL rules the origin supplied, and finally returns the object to the requester. That first request therefore still pays close to the full origin round-trip cost. 3. **Cache hit.** Every later request that lands on the same PoP while the copy is still fresh is a cache hit: the PoP answers directly from local storage without contacting the origin at all. ## Why it exists This exists because network latency is bounded by physics — light (and therefore any signal) takes a minimum amount of time to cross a given distance, and a request/response cycle over TCP/TLS typically needs several round trips (connection setup, TLS handshake, then the actual HTTP exchange) before the first byte comes back. A round trip between Tokyo and Virginia might cost 150–250ms; a round trip between a Tokyo user and a Tokyo-area PoP might cost single-digit milliseconds. Multiplying several round trips by that difference is the entire reason CDNs exist for latency-sensitive content: moving the data physically closer to the reader removes most of that unavoidable travel time. As a side effect, it also protects the origin — instead of every single request worldwide hitting one server, only the relatively rare cache misses do, so the origin can be sized for a fraction of total traffic. ## The trade-off The core trade-off is **freshness versus speed**. - A **longer TTL** (time-to-live, the duration a cached copy is considered valid before the PoP must re-check with origin) means more cache hits and less origin load, but it also means changes at the origin take longer to reach users, since a PoP serving a stale-but-not-yet-expired copy has no way to know the underlying content changed. - **Very short TTLs** keep content fresh but push traffic back toward the origin, eroding much of the latency benefit. - There's also a **cold PoP** cost: a CDN with, say, 200 PoPs doesn't magically have every object everywhere — content only gets cached at a PoP after that PoP's own first miss, so a brand-new region or a rarely-visited object can still pay full origin latency on its first hit even though the object has been cached elsewhere for hours. - Running a large PoP fleet is also **operationally expensive**: more infrastructure to monitor, secure, and keep in sync with cache-invalidation logic. ## Failure modes Failure modes show up in a few recognizable ways. - **Misconfigured TTLs or cache keys.** A PoP can serve outdated or even wrong content (e.g., an old version of a page) well past when it should have refreshed, and different users in different regions can legitimately see different versions of the same page during a rollout, which is confusing if not anticipated. - **A PoP itself goes down.** The routing layer needs to redirect its users to the next-nearest healthy PoP or fall back toward origin; if that fallback path isn't well designed, users nearest the failed PoP see the worst latency of anyone. - **A very popular new object.** When it is requested for the first time simultaneously by users spread across many regions, each region's PoP independently misses and calls origin at roughly the same moment — a smaller, per-object version of a thundering herd, which is one of the reasons larger CDN deployments add an origin-shielding tier (a separate concern from this one). ## Where it shows up A concrete real-world example: an e-commerce or media site puts its product images, CSS, and JS behind a CDN such as Cloudflare, Akamai, or Fastly. A shopper in Tokyo loading the homepage for the first time triggers a handful of cache misses at the Tokyo PoP, each one paying a one-time trip to the US origin; every other shopper who visits from the Tokyo area afterward, for as long as the TTL holds, gets those same assets served in single-digit milliseconds from the local PoP instead of crossing the Pacific. Video-heavy platforms take this even further — Netflix's Open Connect appliances are placed directly inside ISP networks so that video segments barely leave the user's own network at all.
- What's the difference between a pull-based and a push-based CDN?A pull-based CDN (the common model) caches lazily — a PoP only fetches an object from origin the first time it's requested, as described above. A push-based CDN has the origin proactively upload content to PoPs ahead of time, guaranteeing it's already warm everywhere before the first user request, which is useful for content you know will spike immediately (like a scheduled release) but requires more upfront coordination and storage.
- How would you tell, from an HTTP response, whether a request was served from cache or went to origin?Many CDNs add a diagnostic header such as X-Cache: HIT or MISS, or an Age header showing how many seconds the response has sat in cache; the absence or freshness of these headers is the practical way to debug whether a slow response is a cache miss or something else entirely.
- What happens if the nearest PoP to the Tokyo user is unavailable?The routing layer (DNS-based or anycast) detects the outage and steers the request to the next-nearest healthy PoP; that PoP may itself have a cold cache for the object, so the user pays a cache-miss penalty even though the underlying content was already cached at the now-unavailable PoP, until it warms up again.
A CDN is like a chain of neighborhood bakeries restocked from one central factory: the first customer at a brand-new branch might have to wait for a delivery truck from the factory, but every customer after that just grabs bread off the shelf instead of driving across the country to the factory.
saying these in an interview costs you the question
- Says the CDN 'stores the whole website forever' without mentioning TTL or expiration
- Thinks the CDN replaces the origin server entirely
- Doesn't realize a cache miss still has to contact the origin
- Believes a CDN speeds up dynamic, uncacheable API responses the same way it speeds up static assets
- Can't explain why physical/network proximity reduces latency