You put a cache such as Redis or memcached in front of the database to cut read load. What failure modes and consistency issues must the design handle?
answer
- cache-aside race: stale write-back after invalidate
- always TTL; delete after commit
- stampede ⇒ single-flight + TTL jitter + early refresh
- can the DB survive a cold cache?
- skewed repeated reads only; never the source of truth
basics
~20 sPlan for: stale entries after writes (invalidate after commit, keep TTLs short), stampedes when a hot key expires or the cache restarts (single-flight plus jittered TTLs), the database having to survive a cold cache, and a hit-rate collapse from eviction. A cache is not a source of truth.
solid answer
~60 sWith cache-aside, a read checks the cache, falls back to the database on a miss, and populates the cache. The failure modes: - **Staleness and write races.** A common bug: reader misses, loads the old row, then a writer commits and deletes the key, then the reader writes its stale value back — the cache is now wrong until TTL. Mitigate with delete-after-commit, short TTLs, versioned keys, or compare-and-set on write-back. - **Stampede / thundering herd.** A hot key expires and thousands of concurrent requests all miss and hit the database at once. Fix with single-flight (one loader per key, others wait), jittered TTLs, and early/probabilistic refresh before expiry. - **Cold cache.** A restart or eviction storm sends 100% of traffic to the database. The database must have enough headroom, or you need request coalescing, admission control/load shedding, and warm-up. - **Hot keys** concentrating on one cache node, and eviction quietly degrading hit rate as the dataset grows. - **Cache-only data loss.** Anything not durably in the database is gone when the cache is. And caches only help skewed, repeated reads — uniform random access or write-heavy workloads gain almost nothing.
code
text · 5 linesR: miss user:1
R: SELECT ... -> v1
W: UPDATE ... -> v2 ; COMMIT
W: DEL user:1 (key absent)
R: SET user:1 = v1 <-- stale until TTLgo deeper
Describe cache-aside and know that entries can go stale, so TTLs and invalidation on write are needed.
Add the stampede problem and the invalidate-after-commit rule, and explain why caches only help repeated, skewed reads.
Cover the cache-aside write-back race, single-flight and TTL jitter, cold-start survivability, hot keys, and hit-rate monitoring.
Treat the cache as a capacity and availability dependency: state the staleness budget per data class, require the database to survive a cold cache or design explicit shedding, and decide when caching is the wrong lever entirely.
## The patterns - **Cache-aside (lazy loading).** The application checks the cache, and on a miss reads the database and stores the result. Writes go to the database and invalidate (or update) the key. Most common; the cache never sits in the write path's critical correctness. - **Read-through / write-through.** The cache library handles the fallback and the write, keeping cache and database in step at the price of write latency and a harder failure story. - **Write-behind.** Writes land in the cache and are flushed to the database asynchronously — fast, and a genuine data-loss risk if the cache dies. Only for data you can afford to lose or reconstruct. ## Consistency: the race that actually bites Cache-aside has a well-known interleaving: 1. Reader R misses on key `user:1`. 2. R reads the row from the database — value **v1**. 3. Writer W updates the row to **v2** and commits. 4. W deletes `user:1` from the cache (nothing there yet). 5. R writes **v1** into the cache. The cache now serves v1 until its TTL expires, potentially forever if TTL is infinite. Mitigations, in increasing strength: - **Always set a TTL.** Turns an unbounded error into a bounded one. This alone rescues most systems. - **Delete rather than update on write**, and delete *after* the transaction commits — deleting before commit lets a concurrent reader repopulate with the pre-commit value. - **Versioned keys**: include a row version or updated-at in the key so a new version never collides with an old one; old entries fall out by eviction. - **Compare-and-set write-back**, storing the version with the value and refusing to overwrite a newer one. - **Delayed double delete** (delete, wait a beat, delete again) — a pragmatic hack widely used; it narrows the window rather than closing it. Also remember the cache is not transactional with the database. If the write commits but the invalidation fails (cache down, network blip), you serve stale data. Make invalidation retryable — for example by publishing invalidations through a durable channel or by relying on a change-data-capture stream — or accept the TTL as your correctness bound. ## Availability: stampedes and cold starts **Stampede.** A key with a 60-second TTL backing a page served 5,000 times a second: at expiry, up to 5,000 requests miss simultaneously and all query the database with the same expensive query. The database's capacity was sized for the cached rate, so it falls over, timeouts cascade, and the retry storm makes it worse. Defences: - **Single-flight / request coalescing:** the first miss takes a per-key lock and loads; the rest wait for its result. - **Jittered TTLs:** randomise expiry (e.g. 60 s ± 10 s) so keys populated together don't expire together. - **Early refresh:** refresh probabilistically as the entry approaches expiry, so it is renewed by one request while the old value is still being served. - **Serve-stale-on-error:** keep the expired value and return it if the loader fails or is slow. **Cold cache.** After a cache restart, deploy, flush, or failover, hit rate is zero. If the database can only serve 10% of the request rate, the system cannot self-recover. Every design that leans on a cache must answer: *can the database survive a cold cache?* If not, you need admission control (shed or queue a fraction of traffic), a warm-up procedure, or a replicated cache tier that never loses everything at once. **Hot keys.** One celebrity key can saturate a single cache node's CPU or network. Defences are local in-process caching in front of the shared cache for a very short TTL, or key splitting across replicas. **Eviction drift.** Under memory pressure the cache evicts; hit rate degrades silently as the dataset grows and database load creeps up. Monitor hit ratio and database QPS together; alert on hit-rate drops, not just on cache errors. ## Where caching doesn't help - **Write-heavy workloads** — a cache offloads reads, not writes, and heavy invalidation traffic can even reduce the hit rate to near-uselessness. - **Uniform random access over a huge key space** — no reuse means no hits; you have added latency and a component. - **Data that must be authoritative** — balances checked before debits, uniqueness checks, anything with FOR UPDATE semantics. Read those from the primary. - **Cheap queries** — caching a 0.2 ms indexed point lookup behind a 0.3 ms network round trip is a loss; the database's own buffer cache is already doing the job. A cache is a good answer when reads are repeated, skewed, and either expensive to compute or expensive to fetch. Otherwise fix the query or the index first. ## Also worth mentioning **Negative caching** (remembering misses) prevents a stream of lookups for non-existent keys from hitting the database every time — a real availability issue when the key space is attacker-controlled. Cache misses briefly and with a short TTL.
- On a write, should you update the cached value or delete the key?Delete, in almost all cases, and delete after the transaction commits. Updating requires the writer to compute the exact cached representation and creates lost-update races when two writers reorder; deleting means the next reader repopulates from the authoritative source. The cost is one extra miss per write, which is negligible unless the key is extremely hot — in which case single-flight plus early refresh keeps that miss from becoming a stampede.
- Your cache cluster fails completely at peak traffic. What should happen?The database sees the full read rate, so the honest questions are whether it has the headroom and what sheds first. A resilient design has request coalescing so identical queries collapse, per-endpoint concurrency limits or load shedding so the database stays responsive for the most important traffic, timeouts short enough that clients do not pile up, and circuit breakers so cache errors degrade to direct reads rather than to errors. If none of that exists, the cache is not an optimisation — it is a single point of failure.
- Which workloads gain almost nothing from a cache?Write-heavy ones, because caches offload reads only and constant invalidation destroys the hit rate; uniform random reads over a very large key space, where reuse and therefore hits are rare; and workloads whose queries are already sub-millisecond indexed lookups, where the network round trip to the cache costs as much as the database. In those cases the effort belongs in schema, indexing, or write-path work instead.
saying these in an interview costs you the question
- Caching without a TTL, so a missed invalidation is permanent.
- Deleting the cache key before the transaction commits.
- Not accounting for a cold cache — assuming the database can absorb 100% of traffic.
- Treating the cache as a source of truth for data not durably written.
- Believing a cache helps a write-bound or uniformly-random-read workload.