skip to content

A Redis cache is configured with `maxmemory-policy volatile-lru`, but the application only sets a TTL on a small fraction of its keys. What happens when the instance fills up, and how do the `allkeys-*` and `volatile-*` eviction policy families differ in general?

level: middleimportance: must knowfreq 55%

answer

  1. prefix = candidate pool, suffix = victim choice
  2. volatile-* touches only keys with a TTL
  3. no candidates ⇒ acts like noeviction
  4. INFO keyspace: keys= vs expires= ratio
  5. allkeys = whole instance is disposable

basics

~20 s

volatile-* policies only consider keys that have a TTL. If too few keys are volatile, Redis runs out of eviction candidates and behaves like noeviction — writes fail with OOM errors even though most of the keyspace is untouched. allkeys-* policies consider every key.

solid answer

~50 s

Redis's eviction policies come in two families that differ only in the **candidate pool**: - **`allkeys-lru` / `allkeys-lfu` / `allkeys-random`** — any key in the database may be evicted. Treats the whole instance as disposable cache. - **`volatile-lru` / `volatile-lfu` / `volatile-random` / `volatile-ttl`** — only keys with an expiry set are candidates. Keys without a TTL are permanently protected. With `volatile-lru` and few TTL'd keys, Redis evicts the handful of volatile keys and then has nothing left to free, so it falls back to the noeviction behavior: memory-increasing commands are rejected with `OOM command not allowed when used memory > 'maxmemory'`. The instance is full of un-evictable data and effectively read-only, while `evicted_keys` sits at some small plateau. The fix is a decision, not a tuning knob: either give cache keys TTLs so the volatile pool is large enough, or admit everything is disposable and switch to `allkeys-lru`/`allkeys-lfu`. `volatile-*` is the right choice **only** when the TTL genuinely marks the disposable subset.

code

text · 9 lines
text
127.0.0.1:6379> INFO keyspace
# Keyspace
db0:keys=12000000,expires=430000,avg_ttl=0
# only ~3.6% of keys are eviction candidates under volatile-*

127.0.0.1:6379> INFO stats
evicted_keys:430000      # plateaued: pool exhausted
127.0.0.1:6379> SET page:new "html"
(error) OOM command not allowed when used memory > 'maxmemory'.

go deeper

for a junior

Know the two families by name and the one-line rule: volatile-* only ever deletes keys that have a TTL, allkeys-* can delete anything.

for a middle

Explain the exhausted-pool fallback to OOM errors and how to check the pool size with INFO keyspace and INFO stats.

for a senior

Diagnose the plateaued-evictions signature, and fix it by deciding what the instance actually is — pure cache, mixed with TTL discipline, or something that should be split.

for a principal

Treat policy as a data-classification decision: which data in this instance is regenerable, who owns enforcing TTLs at the write path, and when the answer is a second instance rather than a policy change.

## The two families When an instance is at `maxmemory`, Redis must free space before running a memory-increasing command. `maxmemory-policy` answers two separate questions: *which keys may I delete* (the family prefix) and *which of those do I pick first* (the suffix). The prefix defines the candidate pool: - **`allkeys-…`** — the pool is the entire keyspace. Every key is fair game, TTL or not. - **`volatile-…`** — the pool is only keys that have an expiry set, i.e. keys tracked in the database's `expires` table. Keys with no TTL cannot be evicted, ever. - **`noeviction`** — the pool is empty by definition; writes fail instead. The suffix picks the victim from the pool: `-lru` (least recently used, approximated by sampling), `-lfu` (least frequently used, approximated with per-object counters), `-random` (uniform pick), and `-ttl` (volatile family only: nearest expiry first). The approximation details are a topic of their own; what matters here is that all of them operate strictly inside the pool the prefix defines. ## The failure mode in the question With `volatile-lru` and only, say, 5% of keys carrying a TTL, the pool is 5% of the keyspace. As memory fills, Redis evicts those volatile keys — and it will happily evict *hot* volatile keys, because they are the only ones it is allowed to touch. Once the pool is exhausted (or the remaining volatile keys are too small to free enough bytes), Redis has no candidates and returns the same OOM error as `noeviction`: `(error) OOM command not allowed when used memory > 'maxmemory'.` The telltale signature: `INFO stats` shows `evicted_keys` climbing and then flatlining while OOM errors appear in the application; `INFO memory` shows `used_memory` pinned at `maxmemory`; `DBSIZE` barely moves. Candidates commonly miss this because they assume "an eviction policy is set, so Redis will always make room". It will not — the policy grants *permission*, and `volatile-*` grants very narrow permission. A second, subtler symptom is cache-quality collapse before the hard failure: because eviction pressure concentrates entirely on a small volatile subset, that subset churns constantly and its hit rate falls toward zero, while non-volatile keys occupy memory indefinitely regardless of how cold they are. ## When each family is correct **Use `allkeys-*` when the instance is a pure cache.** Every value can be recomputed from a source of truth, so losing any key costs a cache miss, nothing more. This is the common case for a read-through cache in front of a database, and it is the configuration that degrades gracefully: memory pressure turns into a slightly lower hit rate rather than errors. **Use `volatile-*` when one instance genuinely mixes disposable and non-disposable data**, and TTL is the honest marker of "disposable". Example: sessions and rate-limit counters (TTL'd, evictable under pressure) alongside a small set of feature flags or config blobs that must never vanish. This only works if the TTL'd portion is large enough to absorb growth — otherwise you are back at the failure above. **Use `volatile-ttl`** when nearest-to-expiry is the best proxy for "least valuable" — for example short-lived, uniformly accessed tokens where recency carries no signal. It is a niche choice. **Keep `noeviction`** when nothing in the instance may be dropped: a job queue, a dedupe set, a stream you have not consumed. Then OOM errors are the *correct* outcome — a loud signal to scale, not a bug to configure away. ## Practical checks Before trusting a `volatile-*` policy, measure the pool. `INFO keyspace` reports both key count and expires count per database, e.g. `db0:keys=12000000,expires=430000,avg_ttl=0` — that ratio is the eviction pool as a fraction of the keyspace. If `expires` is a small minority and the keyspace is unbounded, the configuration is a time bomb. Either enforce a TTL at the write path (make the cache wrapper require one) or switch families. Finally, remember these are per-instance settings, not per-key. If you actually need two different eviction disciplines, the answer is two Redis instances (or two databases on separate deployments), not a cleverer policy.

  • How would you tell this apart from an instance that is simply undersized?
    Compare the eviction pool with the keyspace: `INFO keyspace` shows `keys=` and `expires=` per database. If `evicted_keys` plateaus while OOM errors continue and `expires` is a small fraction of `keys`, the problem is the policy family, not the size. A genuinely undersized `allkeys-lru` instance shows the opposite signature: `evicted_keys` rising continuously with no OOM errors, and a falling hit rate.
  • Is `volatile-lru` ever the right choice?
    Yes, when a single instance intentionally mixes disposable and non-disposable data and TTL truthfully marks the disposable part — sessions and rate-limit counters that may be dropped, alongside a small always-resident set that must not be. It requires that the TTL'd portion is large enough to absorb all growth, and that the write path enforces TTLs. If either condition is shaky, split the workload into two instances instead.
  • What does the `-ttl` suffix pick, and why does it exist only in the volatile family?
    `volatile-ttl` evicts the key whose expiry is nearest, approximated by sampling like the other policies. It exists only in the volatile family because non-volatile keys have no expiry to compare — the ordering is undefined for them. It is useful when TTL encodes value better than recency does, for example uniformly accessed short-lived tokens, but for most caches LRU or LFU is a better predictor of future access.

saying these in an interview costs you the question

  • Believing that setting any eviction policy guarantees writes never fail.
  • Thinking `volatile-lru` evicts non-TTL keys once the volatile ones run out — it never does.
  • Assuming volatile policies evict the *expired* keys; they evict live keys that merely happen to have an expiry.
  • Treating `allkeys-lru` as unsafe by default without asking whether the data is regenerable.
  • Trying to get per-key eviction behavior out of a per-instance setting instead of splitting instances.

context