skip to content

In the cache-aside (lazy-loading) caching pattern, walk through step by step what happens when application code requests a key that is not currently in the cache.

level: juniorimportance: must knowfreq 85%

answer

  1. read-through is app, not cache
  2. miss then populate
  3. cache is a dumb store
  4. lazy loading
  5. cold key penalty

basics

~20 s

The app checks the cache first. If the value isn't there (a miss), the app reads it from the real database, saves a copy in the cache, then returns it to the caller. Next time it's a fast cache hit.

solid answer

~40 s

Cache-aside puts the application in charge of loading: on a read, the app queries the cache directly. On a hit, it returns the cached value with no database involvement. On a miss, the app itself queries the database (or origin store), gets the current value, writes that value into the cache — usually with a TTL — and then returns it to the caller. The cache is purely a passive key-value store; it never talks to the database on its own. This means the cache and database can go out of sync since two independent operations connect them, and the application owns both keeping them consistent and handling misses.

go deeper

for a junior

Should describe the two-branch hit/miss flow correctly and know that a miss means the app itself queries the database and populates the cache.

for a middle

Should additionally explain why the cache is described as 'passive' / 'dumb', and mention TTL as the mechanism that eventually expires stale entries.

for a senior

Should discuss the cold-start/self-healing property, the latency cost of the first read, and how this pattern decouples cache technology choice from the database.

for a principal

Should connect this base mechanism to system-level consequences: cache warm-up strategy after deploys, cache outage blast radius on the database, and when cache-aside is the wrong default for an entire service.

## What cache-aside is **Cache-aside**, also known as **lazy-loading**, is a caching pattern in which the application code — not the cache server — owns the logic for keeping the cache populated and in sync with the underlying data store. The cache itself (`Redis`, `Memcached`, an in-process LRU map, etc.) is a completely passive key-value store: it can only answer "do you have this key, and if so what's the value" or accept "store this value under this key." It has no idea what a database is, no query capability against one, and no mechanism to refresh itself. ## The read path has two branches Every read against a cache-aside cache follows the same two-branch flow. 1. First, the application issues a `GET` for the key. 2. **Hit** — if the cache has the key, the cached value is handed straight back to the caller, and the request never touches the database. This is the fast path that gives caching its latency and load-reduction benefit. 3. **Miss** — if the cache does not have the key, the application falls through to a second step: it queries the system of record — a relational database, another microservice, an external API — for the current value, gets the result back, writes that result into the cache with a `SET` (typically attaching a TTL so the entry expires and self-cleans over time), and only then returns the value to the original caller. ## Why the design exists This design exists because it lets you cache exactly the working set that's actually being read, with no advance warming step and no coupling between the cache technology and the write path of the system of record. - **Cheap and swappable.** Because the cache is dumb, it's cheap, horizontally scalable, and swappable — you can point the same read-then-populate logic at `Memcached` today and `Redis` tomorrow without touching the database schema or write path at all. - **A cold cache self-heals.** It also means a cold cache (freshly deployed, just flushed, or a brand-new node joining a cluster) self-heals automatically: the first request for any key simply takes the miss path and populates it, no separate warm-up job required. ## The costs of that simplicity The cost of that simplicity is that the cache and the database are two independently-updated stores connected only by application logic, so they can and will drift out of sync for short windows — cache-aside offers no built-in consistency guarantee, only "eventually correct, cache expires or gets invalidated." Two more costs follow: - **The "cold key" penalty.** A second cost is that the very first read of any key (or any read after expiry) pays the full latency of a database round trip plus a cache write, on top of the actual work — there's no way to make that first read cheap under pure cache-aside. - **Duplicated populate logic.** A third cost is that every application (or service) that reads this data has to correctly implement the get-then-populate logic; if that logic is duplicated across services without a shared client library, subtle bugs (missing TTLs, wrong key formats) creep in. ## How it shows up in production In production this shows up in a few characteristic ways. - **A mass expiry event** — many keys with the same TTL expiring at once, e.g. after a cache flush or a deploy that resets TTLs — produces a burst of concurrent misses that all hit the database simultaneously for the same or related keys, which can look like a self-inflicted spike on the database tier. - **A benign failure.** If the application crashes or times out after reading from the database but before writing to the cache, the failure is benign: the value just never gets cached this round, and the next reader takes the miss path again — no corruption results. - **The riskier failure mode** is when the cache goes down entirely: cache-aside degrades to "every read hits the database," and if the database was sized assuming a high cache hit rate, this can cascade into a full outage rather than a graceful slowdown, so production systems typically pair cache-aside with circuit breakers or load shedding on the database side. ## The canonical example at scale The canonical large-scale example is Facebook's use of `Memcached` in front of `MySQL`, described in their well-known "Scaling Memcache at Facebook" engineering writeup: web-tier servers read from Memcached first, and on a miss query MySQL directly and populate Memcached with the result before answering the request, while writes go to MySQL and then delete (rather than update) the corresponding Memcached key. That delete-on-write choice, and the miss-then-populate read path, is textbook cache-aside operating at a scale of many millions of reads per second, and it's the pattern most engineers reach for by default whenever they bolt caching onto an existing read path without redesigning the write path.

  • What happens on a cache hit versus a cache miss in terms of database load?
    On a hit, the database is never touched at all — the request is served entirely from the cache, which is the whole point of the pattern. On a miss, the database is queried exactly once, and the result is cached so subsequent requests for the same key become hits until the TTL expires or the key is invalidated.
  • Does cache-aside require the cache to know anything about the database schema?
    No — that's one of its main selling points. The cache just stores opaque key-value pairs; all knowledge of how to fetch and shape the data lives in the application code, so the same cache technology can sit in front of completely different data stores.
  • What happens if the application crashes after reading from the database but before writing to the cache?
    Nothing dangerous — the cache simply stays cold for that key, and the next request takes the miss path again, re-reading from the database and retrying the populate. It's a safe, self-correcting failure, unlike a crash mid-write to the database itself.

Like a barista who only makes a latte when someone orders it, then keeps a cup on the counter for the next person who asks for the same drink — nobody pre-brews drinks nobody's ordered yet.

saying these in an interview costs you the question

  • Says the cache automatically fetches from the database on a miss
  • Doesn't mention that the app has to write the value back into the cache after a miss
  • Thinks a miss means the key definitely doesn't exist in the database
  • Can't explain why the first read of a key is always the slowest

context