A nightly warm-up job writes tens of thousands of Redis cache keys, all with the same TTL. Explain the failure this sets up and how randomising the TTL per key changes it.
answer
- Batch write + same TTL = synchronised expiry
- Jitter = base + rand(0, 20% of base)
- Window must exceed refill time (size ÷ refill QPS)
- Alignment sneaks in via deploys and midnight EXPIREAT
- Jitter fixes cross-key, not one hot key
basics
~20 sAll the keys expire in the same second, so the origin sees one huge burst of misses and Redis a spike of expiry work. Add jitter — a random offset of roughly 10-25% of the base TTL per key — so expiries spread out over a window wider than the time to refill.
solid answer
~60 sWriting 50,000 keys within a few minutes with an identical TTL means they all become unavailable within the same few minutes, one TTL later. The result is a correlated miss burst: the cache hit ratio collapses to near zero for that key family, the origin database receives the full uncached read rate at once, latency climbs, and the slower origin makes the refill window longer — which is how a cache-caused outage starts. Redis itself also feels it: the active-expiry cycle and the resulting deletes and refill writes spike CPU and can add latency. The fix is **TTL jitter**: give each key `base ± random`, e.g. `EX base + rand(0, base * 0.2)`. Expiries then spread over a window instead of a point, so refill demand becomes a plateau the origin can absorb. Rules of thumb: the jitter window should be wider than the time it takes to refill the whole family; apply jitter in the shared cache client so nobody forgets; and beware sources of accidental alignment — deployment-time warm-ups, `EXPIREAT` on a wall-clock boundary such as midnight or the top of the hour, and identical TTLs derived from one config value.
code
text · 14 lines# without jitter: every key in the warm-up dies together
SET cat:1001 <json> EX 3600
SET cat:1002 <json> EX 3600
...
# with one-sided jitter: base 3600s + up to 20% random
# (computed per key by the client, illustrated as redis-cli calls)
SET cat:1001 <json> EX 3712
SET cat:1002 <json> EX 3388
SET cat:1003 <json> EX 4021
# spotting alignment after the fact: expiry spikes one TTL apart
INFO stats # watch expired_keys as a rate
INFO keyspace # db0:keys=..,expires=..,avg_ttl=..go deeper
Explain the core idea: identical TTLs mean everything expires at once and the database gets hit all at once, so add a random amount to each TTL.
Add the mechanics — one-sided versus two-sided jitter, applying it in the shared client — and mention that the origin burst feeds back into longer refill time.
Size the jitter window from origin refill capacity, enumerate the hidden alignment sources (deploys, midnight EXPIREAT, shared constants), and describe verifying with expired_keys periodicity.
Position jitter as one control in a cache-availability strategy alongside single-flight refill, controlled warm-up rate limits, and serve-stale fallbacks, and make it a platform default so no service can opt out by accident.
## The failure: correlated expiry Redis expiry is per key, but the *distribution* of deadlines is decided by your write pattern. If keys are created in a batch and given identical TTLs, their deadlines inherit the batch shape. One TTL later, the whole family disappears within roughly the same interval it took to create. Three things then happen at once: 1. **Origin burst.** Every request for that family becomes a miss. The database receives its full uncached load in a spike, often ten to a hundred times steady state. 2. **Refill amplification.** Each miss triggers a recompute and a cache write. Under load, many concurrent requests for the *same* key recompute in parallel — the classic herd on a single key, which is handled by a separate mechanism (single-flight locking or early recompute); jitter alone does not fix that. What jitter fixes is the *cross-key* correlation. 3. **Redis-side cost.** Expiry is not free: the active-expiry cycle samples keys with TTLs many times per second and deletes those that have passed, and a mass expiry makes that cycle work hard, adding CPU and latency; the deletes themselves free a large amount of memory at once (mitigated but not eliminated by `lazyfree-lazy-expire yes`), and the refill writes arrive as a burst too. Because the origin gets slower under the burst, refill takes longer, and traffic keeps arriving — a self-reinforcing loop that can outlive the original trigger. ## Jitter: spreading the deadlines Jitter means adding a per-key random component to the TTL. Two common shapes: - **One-sided:** `ttl = base + random(0, base * 0.2)` — never shorter than the base, so the staleness bound stays intact and only lengthens slightly. - **Two-sided:** `ttl = base * (1 + random(-0.1, +0.1))` — keeps the mean at the base, but some keys are slightly fresher than the nominal bound. Fine when the bound is soft. Sizing rule: the jitter window must be **wider than the time your origin needs to refill the family at a tolerable rate**. If refilling 50,000 keys at a sustainable 500/s takes 100 seconds, a jitter window of a few seconds accomplishes nothing; you need a window of minutes. Work backwards from origin capacity: window ≥ family size ÷ acceptable refill QPS. Apply jitter **once, centrally**, in the cache client wrapper, not at each call site. It is exactly the sort of discipline that erodes when a new service copies an old snippet. ## Sources of accidental alignment Jitter helps only if you notice the alignment. Watch for: - **Batch warm-up jobs and imports** — the case in the question. - **Deploys.** A rolling restart that drops an in-process cache, or a deploy that changes the key prefix (a common cache-busting trick), re-populates everything in a narrow window. The next expiry inherits that shape even months later if TTLs are long. - **Wall-clock deadlines.** `EXPIREAT` pinned to midnight, or a TTL computed as "seconds until the top of the hour", aligns not only your keys but every instance in the fleet, and often several services at once. If a business rule genuinely wants a daily boundary, jitter the boundary by a random number of seconds per key. - **A single shared TTL constant** used by many key families that were all populated on the same schedule. - **Cluster-wide events.** A failover or a `FLUSHALL`-and-rebuild produces a synchronised cold start; jitter on rewrite reduces the echo one TTL later. ## What jitter does not solve Jitter decorrelates *different* keys. It does nothing for a single very hot key whose expiry is followed by hundreds of simultaneous recomputes, and nothing for a cold start where everything is missing at once. Those need per-key single-flight or early/probabilistic recomputation, and a controlled warm-up with rate limiting respectively. Say this out loud in an interview — it shows you know which tool solves which shape of problem. ## Verifying it worked Graph `expired_keys` from `INFO stats` as a rate: correlated expiry shows as periodic spikes at the TTL period, jittered expiry as a broad plateau. Also watch origin QPS for the same periodicity and the hit ratio for periodic dips. If the spikes recur exactly one TTL apart, you still have alignment somewhere.
- How wide should the jitter window be?Wider than the time your origin needs to refill that key family at a rate it can sustain. Divide the family size by the acceptable refill QPS: 50,000 keys at 500/s needs a window of at least 100 seconds, so a jitter of a few seconds on an hour-long TTL would be useless. Ten to twenty-five percent of the base TTL is a reasonable starting point precisely because it scales with the family's size in most systems.
- Does jitter solve the case where one extremely popular key expires and hundreds of requests recompute it at once?No. Jitter decorrelates deadlines across different keys; it cannot help when the contention is on a single key, because that key still has exactly one expiry moment. That case needs a single-flight mechanism — one recomputing caller, the rest waiting or serving the previous value — or recomputation started before the deadline. Jitter and single-flight solve different shapes of the same symptom.
saying these in an interview costs you the question
- Assuming identical TTLs are harmless because 'expiry is per key anyway'
- Adding a tiny jitter (a second or two) to an hour-long TTL and calling it fixed
- Claiming jitter also solves contention on a single hot key
- Pinning TTLs to a wall-clock boundary such as midnight and not seeing it aligns the whole fleet
- Ignoring that mass expiry costs Redis CPU and memory-free work, not just the origin