A web app serves the same product page thousands of times per minute. What is caching, and how do different cache layers (CDN, reverse proxy, application, database) each help reduce load on the origin?
answer
- hit vs miss
- closer = faster
- CDN=edge, proxy=full HTTP response, app=arbitrary data, DB=buffer pool
- each layer shields the next
- Zipfian hot-key traffic
basics
~20 sCaching means saving a copy of data somewhere fast so you don't have to redo slow work every time. Different layers (CDN, reverse proxy, app memory, database) each keep a copy closer to where it's needed, so fewer requests reach the slow original source.
solid answer
~40 sA cache stores a copy of expensive-to-produce data in a faster, closer location so repeated requests are served without redoing the original work. A CDN caches static/semi-static content at edge locations near users, cutting latency and origin traffic globally. A reverse proxy (e.g., Varnish, Nginx) caches full HTTP responses in front of app servers, absorbing traffic before it hits application code. An application-level cache (e.g., Redis, in-process memory) stores computed results, session data, or query results inside the service layer. The database itself caches via a buffer pool/query cache, keeping hot pages in memory. Each layer trades cost and staleness risk for reduced load and latency further upstream.
go deeper
Should know what a cache hit/miss is and name at least two layers (e.g., CDN and application cache) with a rough sense of what each stores.
Should describe all four layers, explain roughly what each caches (static assets vs full HTTP responses vs arbitrary app data vs DB pages), and articulate the latency/load benefit at each.
Should reason about which layer is appropriate for a given piece of data, flag the personalization/cache-key-vary risk at the proxy/CDN layer, and connect layer choice to Zipfian traffic patterns.
Should design the end-to-end caching topology for a system, including which layers to invest in given the traffic shape, cost model, and staleness tolerance, and anticipate cross-layer invalidation complexity.
## What a cache is A cache is a smaller, faster storage layer that sits in front of a larger, slower, authoritative source of data — typically a database, a filesystem, or an expensive computation — and keeps a copy of recently or frequently requested results. When a request arrives, the system first checks the cache: - A **'hit'** means the data is found and returned immediately without touching the slow source. - A **'miss'** means the cache doesn't have it, so the system falls through to the origin, fetches the data, returns it, and usually stores a copy in the cache for next time. The core mechanism is always the same trade — spend memory and some risk of staleness in exchange for speed and reduced load on whatever sits behind the cache. ## The layers of the stack Real systems rarely have a single cache; they stack several layers, each solving a different part of the latency/load problem. 1. **A Content Delivery Network (CDN)** caches content at edge servers physically distributed close to end users around the world. When a user in Tokyo requests a product image or a mostly-static HTML page, the CDN edge node in or near Tokyo can serve it directly, avoiding a round trip to the origin data center possibly on another continent. This cuts both latency (physics — the speed of light over distance) and the total request volume that ever reaches the origin. 2. **A reverse proxy cache** (`Varnish`, `Nginx`, `Squid`) sits directly in front of the application servers, typically within the same data center or cloud region. It caches full HTTP responses keyed by URL (and sometimes headers/cookies), so identical requests for the same resource are served from the proxy's memory without ever invoking application code, database queries, or business logic. This protects the app tier from load spikes and reduces compute cost. 3. **An application-level cache** (`Redis`, `Memcached`, or in-process caches like `Caffeine`) lives inside or next to the service layer and stores arbitrary computed data — query results, rendered fragments, session state, feature flags, rate-limit counters — keyed however the application chooses. Unlike the reverse proxy, it isn't limited to whole HTTP responses; it can cache a single expensive join result or an API call to a third party. 4. **Finally, the database itself** maintains an internal cache — a buffer pool (e.g., InnoDB's buffer pool in MySQL) or a query cache — that keeps frequently accessed pages/rows in memory so disk I/O is avoided even when a query does reach the database. ## Why stack them at all This layering exists because each tier removes load from everything behind it, and each tier has a different reach and blast radius. Without caching, every single page view would require a fresh trip to the origin: rendering HTML, running database queries, hitting disk. At even moderate scale (thousands of requests per minute for the same product page), this is wasteful — the underlying data for that page probably hasn't changed since the last request a few seconds ago. Caching exploits the fact that read traffic for popular items is highly repetitive: a small number of **hot keys** account for a large fraction of traffic (a Zipfian/power-law distribution), so caching just the hot set yields an outsized reduction in backend load. ## The trade-offs The trade-offs run in both directions. Each additional layer adds infrastructure cost (memory, cache servers, CDN bills) and operational complexity: - more places for a bug to hide, - more moving parts to monitor, - and — critically — more places where stale data can be served if the underlying data changes and the cache isn't told. A CDN caching a price that just changed, or a reverse proxy caching a personalized response for the wrong user, are classic production incidents. There's also a **cost/latency curve**: caching too aggressively saves money but risks serving wrong data; caching too little (or with very short TTLs) keeps data fresh but doesn't save much load. ## Failure modes, and the path through the stack Failure modes at this layer typically show up as either 'why is the site showing old data' (a caching layer didn't get invalidated) or 'why did the database fall over' (a cache layer was accidentally bypassed, misconfigured with a 0-second TTL, or evicted its working set). A concrete example: an e-commerce product page request first checks the CDN edge (serves static assets/images instantly); if it's a dynamic page fragment, it may hit a reverse proxy cache for the rendered HTML; the app server, on a cache miss, checks `Redis` for the product's price and inventory before falling back to a full database query, which itself benefits from the DB's buffer pool keeping the product table's hot pages in RAM. Each layer shaves off a large fraction of the traffic that would otherwise reach the next slower layer.
- What determines whether a piece of data is a good candidate for caching versus one that should always be fetched fresh?Good candidates are read-heavy, relatively expensive to compute or fetch, and tolerant of some staleness — for example, a product description or a rendered page fragment. Poor candidates are highly personalized, change on nearly every request, or require strict real-time correctness, like a bank account balance mid-transaction or a live auction bid count. The decision is really a tolerance-for-staleness question weighed against the cost savings of caching.
- Why might caching at the reverse-proxy layer be dangerous for personalized or authenticated content?A reverse proxy typically caches by URL, so if it isn't configured to vary the cache key by session, cookie, or Authorization header, it can serve one user's personalized response (or another user's private data) to a completely different user. This is a well-known class of production incident, so proxies must be explicitly configured with a cache-key vary policy for anything user-specific.
- How does a CDN know when to refetch content from the origin instead of serving its cached copy?It relies on cache-control directives set by the origin — headers like Cache-Control: max-age, ETag, or Last-Modified — plus an explicit purge/invalidation API the origin can call when content changes. Without correct headers or explicit purges, a CDN will happily keep serving stale content for its configured TTL.
Like a chain of pantries between a farm and your kitchen: a corner store (CDN) near your house, a neighborhood warehouse (reverse proxy), your own kitchen cupboard (app cache), and the farm's own loading dock storage (DB buffer pool) — each stop saves you from driving all the way to the farm.
saying these in an interview costs you the question
- Treats caching as a single layer instead of a stack with different responsibilities
- Assumes caching always makes data fresher instead of trading freshness for speed
- Doesn't mention cache hit/miss as the basic mechanism
- Would cache highly personalized/private data at a shared layer like a CDN or reverse proxy without varying the key
- Can't explain why hot-key/read-heavy data benefits most from caching