How do you choose the TTL for a value you are caching in Redis? Walk through what actually drives the number, rather than defaulting to five minutes.
answer
- TTL = staleness bound AND load knob
- Origin refetch ≈ N keys / T seconds
- Invalidation on write → TTL is a backstop, can be long
- Hit-ratio-vs-TTL curve has a knee
- Classes per prefix, not one global number
basics
~20 sStart from how stale the data may be for the business, then check it against origin load: with N cached keys and TTL T, expiries alone push roughly N/T refetches per second. Shorten for volatile or sensitive data, lengthen when you also invalidate on write.
solid answer
~60 sA TTL is two things at once: a **staleness bound** and a **load knob**, so I pick it from both directions. From the data side: how long may a user see an old value? Prices and permissions get seconds; a product description gets minutes to hours; immutable, content-addressed blobs need no TTL at all. If the value changes rarely, a long TTL costs nothing in freshness. From the load side: with N distinct hot keys and TTL T, expiry alone generates roughly N/T origin fetches per second in steady state. Halving the TTL doubles that. So I size T against the origin's spare capacity, not against intuition. Then: if I can invalidate explicitly on write (`DEL`/`UNLINK` the key, or `SET ... EX`), the TTL becomes a safety net for missed invalidations and can be much longer. If I cannot, the TTL *is* my only correctness mechanism and must be short. Finally I verify empirically: hit ratio from `INFO stats` (`keyspace_hits`/`keyspace_misses`), `expired_keys`, origin QPS, and memory. Hit ratio versus TTL flattens quickly — beyond the knee, longer TTL buys staleness and no hits.
go deeper
Say that TTL trades freshness for load: shorter means fresher but more database hits, longer means fewer hits but staler data, and that different data deserves different numbers.
Bring the N/T arithmetic and the staleness budget together, and name the metrics you would check — hit ratio, expired_keys, origin QPS.
Distinguish TTL-as-contract from TTL-as-backstop behind explicit invalidation, and show you would measure the hit-ratio knee rather than pick a round number.
Define TTL classes as policy across teams, tie them to origin capacity and memory budgets, and make the values runtime-tunable so TTL becomes an incident lever rather than a redeploy.
## The two forces Every TTL is a negotiation between **freshness** and **origin load**. Long TTL: fewer misses, less database traffic, more memory held, more staleness. Short TTL: fresher data, more misses, more origin work, and — if the miss path is expensive — more tail latency for the unlucky requests that hit an expired key. ## Force 1: the staleness budget Ask the owning team the blunt question: *how many seconds may a user see the old value before it is a bug?* Answers cluster into classes: - **Immutable / content-addressed** (rendered assets, versioned config, `sha256`-keyed blobs): no staleness possible; TTL exists only to reclaim memory, so hours or days, or no TTL plus an LRU eviction policy. - **Reference data** (country lists, catalogues, feature-flag snapshots): minutes to hours, because a late update is annoying, not wrong. - **User-visible mutable state** (profile, cart summary, counters): tens of seconds to a few minutes. - **Security-relevant state** (permissions, entitlements, revocation, ban lists): seconds, or event-driven invalidation only — a stale *allow* is a security incident, and no TTL is short enough to make caching a revocation safe by itself. - **Negative entries** ("this ID does not exist"): seconds, far shorter than positives. ## Force 2: the arithmetic of origin load For a working set of N distinct keys that are all requested more often than they expire, the steady-state refetch rate from expiry alone is approximately **N / T** per second. Ten thousand keys at a 300-second TTL is ~33 origin reads/s; the same keys at 30 seconds is ~333/s. That single formula turns "pick a TTL" into capacity planning: decide how much origin QPS the cache tier is allowed to generate, then solve for T. Add the miss traffic from cold keys, evictions, and deploys on top. ## Force 3: do you have invalidation? If the write path can delete or overwrite the cache entry when the source of truth changes, TTL stops being the freshness mechanism and becomes a **backstop** against the invalidations you will inevitably miss (a batch job that writes straight to the database, a failed delete, a bug). Backstop TTLs can be long — tens of minutes to hours — and the freshness guarantee comes from invalidation. If there is no invalidation path, the TTL is the entire contract and must equal the staleness budget. ## Force 4: memory and the working set TTL also decides how much of the key space is resident. `keys × average size` must fit comfortably under `maxmemory` with headroom; otherwise eviction starts making the decision for you, which is a worse, less predictable version of the same choice. If a longer TTL would blow the memory budget, the honest answer is a shorter TTL, a smaller cached representation, or more memory — not hoping eviction sorts it out. ## Measuring rather than guessing - `INFO stats` → `keyspace_hits` / `keyspace_misses` gives the hit ratio; track it per deployment, and per prefix if you shard namespaces. - `INFO stats` → `expired_keys` shows how much churn expiry is causing; `evicted_keys` above zero means memory, not TTL, is bounding residency. - Plot hit ratio against TTL from real traffic. The curve rises steeply and then flattens: past the knee, extra TTL buys staleness, not hits. Pick just past the knee, then trim toward the staleness budget. - Watch origin QPS attributable to cache misses, and the latency of the miss path — that is what a too-short TTL actually costs users. ## Practical defaults Do not carry one global number. Set a small table of TTL classes per key prefix (immutable / reference / mutable / sensitive / negative), make the value configurable without a deploy so you can react during an incident, apply jitter centrally so a batch of keys does not expire together, and treat any key written without a declared class as a defect.
- Your product owner says the data may never be stale. What do you do?Point out that any cache implies some staleness, then move the guarantee from TTL to invalidation: the write path deletes or rewrites the key as part of the same operation, and a modest TTL stays as a backstop for missed invalidations. If even that is unacceptable — typically for permissions or revocation — the honest answer is not to cache the read, or to cache only a version token you validate against the source.
- Does doubling the TTL double the hit ratio?No. Hit ratio versus TTL is a saturating curve: once the TTL comfortably exceeds the typical inter-arrival time for a key, further increases add almost no hits because the key was going to be hit anyway. Past that knee, extra TTL buys only staleness and memory residency, which is why measuring the curve beats guessing.
saying these in an interview costs you the question
- Using one global TTL for every key pattern in the system
- Treating TTL purely as a freshness setting and ignoring the origin load it generates
- Claiming a longer TTL always improves hit ratio, with no knee in the curve
- Caching authorization or revocation data on a TTL and calling it safe
- Setting no TTL because 'eviction will handle it', without checking the eviction policy or memory budget