skip to content

You're designing a multi-tier caching architecture for a global e-commerce site: CDN edge nodes, a reverse-proxy cache (e.g., Varnish or Nginx), an application-level cache (e.g., Redis), and the database's own buffer pool. What does each tier optimize for, and how do you reason about propagating an invalidation across all four when a product price changes?

level: principalimportance: should knowfreq 40%

answer

  1. CDN=latency+global fanout, proxy=app compute, app cache=derived data, DB buffer pool=disk I/O
  2. blast radius grows and reliability drops moving outward
  3. trigger invalidation from one event/CDC, not scattered calls
  4. TTL backstop sized to staleness tolerance per tier
  5. bypass caching entirely for strict-correctness values like checkout price

basics

~20 s

Each caching layer solves a different part of the speed problem: the CDN cuts distance to users, the reverse proxy protects app servers, the app cache stores computed data, and the database's own memory speeds up its queries. When something like a price changes, you have to update or clear the copy at every layer, in the right order, or some users will keep seeing the old price.

solid answer

~50 s

Each tier optimizes a different bottleneck: the CDN minimizes network latency and origin traffic by caching near end users; the reverse proxy shields application compute by caching full rendered responses; the application cache (Redis) stores derived/computed data cheaply reusable across requests; and the DB buffer pool minimizes disk I/O for whatever queries do reach the database. For a price change, invalidation must propagate outside-in from the source of truth: write the new price to the database, invalidate/update the Redis entry (so any app logic reads the new price), purge the specific reverse-proxy cache entries for pages showing that product, and issue a CDN purge for the corresponding URLs — all ideally triggered from a single event (a domain event or CDC stream off the database write) rather than four separate manual calls, since a missed step at any tier leaves a stale price visible to some users, and the tiers closer to the user (CDN, proxy) have the widest fan-out and slowest, least reliable purge propagation.

go deeper

for a junior

Should recognize that there are multiple caching layers and that changing data means all of them need to somehow find out.

for a middle

Should name what each of the four tiers roughly caches and give a basic sense that CDN purges reach more places than an app-cache update.

for a senior

Should explain the blast-radius/reliability asymmetry across tiers and propose centralizing invalidation via an event rather than scattered calls.

for a principal

Should design the end-to-end invalidation pipeline (source-of-truth event/CDC to fan-out consumer), size TTL backstops per tier by business staleness tolerance, and make the judgment call of which values should bypass outer caching tiers entirely for correctness.

## Four genuinely different bottlenecks A multi-tier caching architecture exists because no single cache location can simultaneously minimize network latency to a globally distributed user base, protect application compute from repeated rendering work, avoid redundant expensive data lookups, and avoid disk I/O inside the database — these are four genuinely different bottlenecks, and each tier is purpose-built to attack one of them. ## What each tier optimizes for - **The CDN's job is minimizing physical network latency** and shielding the origin infrastructure from the full weight of global traffic. Edge nodes are placed in many geographic points of presence, so a user in São Paulo gets a response from a nearby edge server instead of round-tripping to wherever the origin data center actually is; at scale, this also means the origin only ever sees cache-miss traffic, which is a small fraction of total requests for popular content. - **The reverse-proxy cache's job is protecting application compute**: it sits immediately in front of the app servers (usually within the same data center or region) and caches full rendered HTTP responses keyed by URL, so an identical request for the same page doesn't have to re-run application logic, template rendering, or authorization checks — this matters even for traffic the CDN didn't catch, e.g. because the CDN's TTL was very short or the response varied per-region. - **The application-level cache's job is avoiding redundant expensive computation** or data assembly inside the app tier itself — not full HTTP responses, but arbitrary derived data: a product's price and inventory count fetched from `Redis` instead of re-querying and re-joining several database tables on every single request, even for requests whose rendered HTML differs (e.g., a personalized page that embeds the same product price). - **Finally, the database's buffer pool job is minimizing disk I/O** for whatever query volume does reach the database after all the layers above have already absorbed most of the traffic — it keeps hot pages/rows resident in RAM so even a 'full miss' through every application-level cache is still relatively cheap. ## Blast radius grows outward, reliability drops Because each tier catches a successively smaller, more filtered slice of traffic, the tiers closer to the user have both the largest blast radius (many more cached copies, spread across the globe) and the least reliable, slowest invalidation propagation, while the tiers closer to the database have the smallest blast radius and the fastest, most reliable invalidation. This asymmetry is the central fact to reason about when propagating a price change: the database write is instantaneous and authoritative by definition (it is the change), but reflecting that change everywhere a stale copy might exist gets progressively harder and slower moving outward. - **Redis** is typically a single logical store (or small cluster) reachable with one delete/update call and near-instant propagation. - **The reverse-proxy layer** may have several nodes but is still within one region, so a purge is a small, fast fan-out. - **The CDN** may have dozens to hundreds of globally distributed edge nodes, each independently caching, so a purge call has to fan out worldwide and can take anywhere from sub-second to tens of seconds depending on the provider, and individual edge purges can silently fail. ## Fan out from one event at the source of truth The disciplined way to reason about this is to trigger invalidation from a single event at the source of truth rather than four separate manual or ad hoc calls scattered through application code. A common pattern: 1. The price-update transaction, on commit, publishes a **domain event** (or is captured via change-data-capture off the database's write-ahead log) carrying the product ID and new price. 2. A dedicated invalidation consumer/service subscribes to this event and is solely responsible for fanning out: delete/update the `Redis` key, purge the specific reverse-proxy cache entries for that product's URLs, and call the CDN's purge API for the same URLs. Centralizing this avoids the failure mode where a developer remembers to invalidate Redis but forgets the CDN purge (or vice versa), which is exactly how 'the price shows differently depending on which server/region you hit' bugs happen in production. ## TTL backstops, and what to leave uncached Given that CDN and reverse-proxy purges are the least reliable link, every tier should also carry its own **TTL as a backstop**, sized according to the business tolerance for a stale price being shown — e.g., a CDN TTL of 60 seconds means that even if a purge is silently dropped, the worst case is a stale price shown for at most a minute, which is often an acceptable trade against building a bulletproof, low-latency global purge system. For strict correctness needs (e.g., final checkout price must never be wrong even briefly), the pragmatic answer is often to not cache that specific value at those outer tiers at all — bypass the CDN and reverse-proxy cache for the final price-confirmation step and read it live from `Redis` or the database, while still caching the merely-informational 'display price' on the product listing page with a short TTL and best-effort purge. This illustrates the principal-level judgment call: not every value needs the same caching aggressiveness, and the right architecture assigns each piece of data to the tier and staleness tolerance appropriate to its actual correctness requirement, rather than uniformly caching (or uniformly not caching) everything.

  • Why is it risky to invalidate a CDN, reverse proxy, and application cache with four separate manual calls scattered across application code, instead of one centralized event-driven process?
    Scattering the calls means every code path that changes the price has to remember all four invalidations correctly, and any missed or reordered call leaves a stale copy at exactly that tier indefinitely (or until its TTL happens to expire). Centralizing on a single domain event or CDC-triggered consumer means there's one place to get the fan-out logic right, one place to monitor for failures, and new write paths automatically get correct invalidation for free.
  • For a value like a final checkout price that must never be shown stale, what's the principal-level design decision regarding which cache tiers to use?
    The right call is often to deliberately bypass the outer, least-reliable tiers (CDN, reverse proxy) for that specific value and read it live from Redis or the database, while still caching the merely-informational 'display price' elsewhere with a short TTL. This assigns each piece of data a caching aggressiveness matched to its actual correctness requirement rather than uniformly caching everything the same way.
  • Why does CDN purge propagation take meaningfully longer than an application-cache (Redis) invalidation?
    Redis invalidation is typically a single call to one logical store (or small cluster) with near-instant propagation, whereas a CDN purge has to fan out to dozens or hundreds of independently-operated edge nodes distributed globally, each of which processes the purge on its own schedule and can fail independently. That larger, more distributed blast radius is inherently slower and less reliable to fully propagate than a single-region cache update.

It's like a company's information flowing from HQ (the database) through regional offices (Redis), branch offices (reverse proxy), and local kiosks (CDN edges) — updating the master record at HQ is instant, but getting every kiosk worldwide to update their printed flyers takes progressively longer and is progressively more likely for at least one kiosk to miss the memo.

saying these in an interview costs you the question

  • Treats all four tiers as interchangeable/equally reliable to invalidate
  • Proposes invalidating each tier with separate uncoordinated manual calls with no mention of centralizing on an event
  • Doesn't recognize that CDN purges are slower/less reliable than app-cache invalidation
  • Assumes every piece of data should be cached identically at every tier regardless of correctness needs
  • Can't name what specific bottleneck each tier addresses

context