skip to content

questions

6

A web app serves the same product page thousands of times per minute. What is caching, and how do different cache layers (CDN, reverse proxy, application, database) each help reduce load on the origin?

level: juniorimportance: must knowfreq 85%

answer

  1. hit vs miss
  2. closer = faster
  3. CDN=edge, proxy=full HTTP response, app=arbitrary data, DB=buffer pool
  4. each layer shields the next
  5. Zipfian hot-key traffic

basics

~20 s

Caching means saving a copy of data somewhere fast so you don't have to redo slow work every time. Different layers (CDN, reverse proxy, app memory, database) each keep a copy closer to where it's needed, so fewer requests reach the slow original source.

solid answer

~40 s

A cache stores a copy of expensive-to-produce data in a faster, closer location so repeated requests are served without redoing the original work. A CDN caches static/semi-static content at edge locations near users, cutting latency and origin traffic globally. A reverse proxy (e.g., Varnish, Nginx) caches full HTTP responses in front of app servers, absorbing traffic before it hits application code. An application-level cache (e.g., Redis, in-process memory) stores computed results, session data, or query results inside the service layer. The database itself caches via a buffer pool/query cache, keeping hot pages in memory. Each layer trades cost and staleness risk for reduced load and latency further upstream.

go deeper

for a junior

Should know what a cache hit/miss is and name at least two layers (e.g., CDN and application cache) with a rough sense of what each stores.

for a middle

Should describe all four layers, explain roughly what each caches (static assets vs full HTTP responses vs arbitrary app data vs DB pages), and articulate the latency/load benefit at each.

for a senior

Should reason about which layer is appropriate for a given piece of data, flag the personalization/cache-key-vary risk at the proxy/CDN layer, and connect layer choice to Zipfian traffic patterns.

for a principal

Should design the end-to-end caching topology for a system, including which layers to invest in given the traffic shape, cost model, and staleness tolerance, and anticipate cross-layer invalidation complexity.

## What a cache is A cache is a smaller, faster storage layer that sits in front of a larger, slower, authoritative source of data — typically a database, a filesystem, or an expensive computation — and keeps a copy of recently or frequently requested results. When a request arrives, the system first checks the cache: - A **'hit'** means the data is found and returned immediately without touching the slow source. - A **'miss'** means the cache doesn't have it, so the system falls through to the origin, fetches the data, returns it, and usually stores a copy in the cache for next time. The core mechanism is always the same trade — spend memory and some risk of staleness in exchange for speed and reduced load on whatever sits behind the cache. ## The layers of the stack Real systems rarely have a single cache; they stack several layers, each solving a different part of the latency/load problem. 1. **A Content Delivery Network (CDN)** caches content at edge servers physically distributed close to end users around the world. When a user in Tokyo requests a product image or a mostly-static HTML page, the CDN edge node in or near Tokyo can serve it directly, avoiding a round trip to the origin data center possibly on another continent. This cuts both latency (physics — the speed of light over distance) and the total request volume that ever reaches the origin. 2. **A reverse proxy cache** (`Varnish`, `Nginx`, `Squid`) sits directly in front of the application servers, typically within the same data center or cloud region. It caches full HTTP responses keyed by URL (and sometimes headers/cookies), so identical requests for the same resource are served from the proxy's memory without ever invoking application code, database queries, or business logic. This protects the app tier from load spikes and reduces compute cost. 3. **An application-level cache** (`Redis`, `Memcached`, or in-process caches like `Caffeine`) lives inside or next to the service layer and stores arbitrary computed data — query results, rendered fragments, session state, feature flags, rate-limit counters — keyed however the application chooses. Unlike the reverse proxy, it isn't limited to whole HTTP responses; it can cache a single expensive join result or an API call to a third party. 4. **Finally, the database itself** maintains an internal cache — a buffer pool (e.g., InnoDB's buffer pool in MySQL) or a query cache — that keeps frequently accessed pages/rows in memory so disk I/O is avoided even when a query does reach the database. ## Why stack them at all This layering exists because each tier removes load from everything behind it, and each tier has a different reach and blast radius. Without caching, every single page view would require a fresh trip to the origin: rendering HTML, running database queries, hitting disk. At even moderate scale (thousands of requests per minute for the same product page), this is wasteful — the underlying data for that page probably hasn't changed since the last request a few seconds ago. Caching exploits the fact that read traffic for popular items is highly repetitive: a small number of **hot keys** account for a large fraction of traffic (a Zipfian/power-law distribution), so caching just the hot set yields an outsized reduction in backend load. ## The trade-offs The trade-offs run in both directions. Each additional layer adds infrastructure cost (memory, cache servers, CDN bills) and operational complexity: - more places for a bug to hide, - more moving parts to monitor, - and — critically — more places where stale data can be served if the underlying data changes and the cache isn't told. A CDN caching a price that just changed, or a reverse proxy caching a personalized response for the wrong user, are classic production incidents. There's also a **cost/latency curve**: caching too aggressively saves money but risks serving wrong data; caching too little (or with very short TTLs) keeps data fresh but doesn't save much load. ## Failure modes, and the path through the stack Failure modes at this layer typically show up as either 'why is the site showing old data' (a caching layer didn't get invalidated) or 'why did the database fall over' (a cache layer was accidentally bypassed, misconfigured with a 0-second TTL, or evicted its working set). A concrete example: an e-commerce product page request first checks the CDN edge (serves static assets/images instantly); if it's a dynamic page fragment, it may hit a reverse proxy cache for the rendered HTML; the app server, on a cache miss, checks `Redis` for the product's price and inventory before falling back to a full database query, which itself benefits from the DB's buffer pool keeping the product table's hot pages in RAM. Each layer shaves off a large fraction of the traffic that would otherwise reach the next slower layer.

  • What determines whether a piece of data is a good candidate for caching versus one that should always be fetched fresh?
    Good candidates are read-heavy, relatively expensive to compute or fetch, and tolerant of some staleness — for example, a product description or a rendered page fragment. Poor candidates are highly personalized, change on nearly every request, or require strict real-time correctness, like a bank account balance mid-transaction or a live auction bid count. The decision is really a tolerance-for-staleness question weighed against the cost savings of caching.
  • Why might caching at the reverse-proxy layer be dangerous for personalized or authenticated content?
    A reverse proxy typically caches by URL, so if it isn't configured to vary the cache key by session, cookie, or Authorization header, it can serve one user's personalized response (or another user's private data) to a completely different user. This is a well-known class of production incident, so proxies must be explicitly configured with a cache-key vary policy for anything user-specific.
  • How does a CDN know when to refetch content from the origin instead of serving its cached copy?
    It relies on cache-control directives set by the origin — headers like Cache-Control: max-age, ETag, or Last-Modified — plus an explicit purge/invalidation API the origin can call when content changes. Without correct headers or explicit purges, a CDN will happily keep serving stale content for its configured TTL.

Like a chain of pantries between a farm and your kitchen: a corner store (CDN) near your house, a neighborhood warehouse (reverse proxy), your own kitchen cupboard (app cache), and the farm's own loading dock storage (DB buffer pool) — each stop saves you from driving all the way to the farm.

saying these in an interview costs you the question

  • Treats caching as a single layer instead of a stack with different responsibilities
  • Assumes caching always makes data fresher instead of trading freshness for speed
  • Doesn't mention cache hit/miss as the basic mechanism
  • Would cache highly personalized/private data at a shared layer like a CDN or reverse proxy without varying the key
  • Can't explain why hot-key/read-heavy data benefits most from caching

context

open as a page

A cache with limited memory must decide which entries to evict when it's full. Compare LRU (Least Recently Used), LFU (Least Frequently Used), and TTL (Time To Live) eviction policies — how does each decide what to remove, and when would you pick one over another?

level: middleimportance: must knowfreq 75%

basics

~20 s

When a cache runs out of space, it has to throw something away. LRU throws out the item nobody's touched in the longest time. LFU throws out the item used the fewest times overall. TTL just deletes items after a fixed amount of time, whether or not they're popular.

open as a page

In caching, compare the cache-aside, write-through, and write-behind (write-back) strategies for keeping a cache and its backing database in sync: describe how each handles reads and writes, and the consistency/durability trade-off of each.

level: middleimportance: must knowfreq 80%

basics

~20 s

Cache-aside: the app checks the cache first, and only saves to it after reading from the database, while writes go straight to the database and the cache entry is cleared. Write-through: every write goes to the cache and the database together, right away. Write-behind: writes go to the cache immediately but reach the database later, in the background.

open as a page

Phil Karlton famously said cache invalidation is one of the two hard problems in computer science. What are the main strategies for invalidating stale cache entries — TTL expiration, explicit invalidation/purge on write, and versioned or keyed cache-busting — and what can go wrong with each in a distributed system with multiple cache nodes?

level: seniorimportance: must knowfreq 70%

basics

~20 s

When the real data changes, the cached copy becomes wrong until it's removed or updated. You can let it expire on its own after a set time, tell the cache directly to delete it when the data changes, or change the cache's key/version so old copies are never looked up again.

open as a page

A popular cache key expires, and thousands of concurrent requests all miss the cache at the same instant, hammering the database simultaneously. What is this failure mode called, and what techniques prevent it?

level: seniorimportance: must knowfreq 65%

basics

~20 s

It's called a thundering herd or cache stampede. When a popular cached item expires, every request that arrives at that moment finds the cache empty and rushes to the database at once, which can overload it. Fixes include locking so only one request refetches, spreading out expiration times, or serving a slightly old copy while refreshing in the background.

open as a page

You're designing a multi-tier caching architecture for a global e-commerce site: CDN edge nodes, a reverse-proxy cache (e.g., Varnish or Nginx), an application-level cache (e.g., Redis), and the database's own buffer pool. What does each tier optimize for, and how do you reason about propagating an invalidation across all four when a product price changes?

level: principalimportance: should knowfreq 40%

basics

~20 s

Each caching layer solves a different part of the speed problem: the CDN cuts distance to users, the reverse proxy protects app servers, the app cache stores computed data, and the database's own memory speeds up its queries. When something like a price changes, you have to update or clear the copy at every layer, in the right order, or some users will keep seeing the old price.

open as a page