One Redis instance holds both a disposable read cache and data the application cannot regenerate — login sessions, a dedupe set, and a pending-job list. How do you set `maxmemory` and `maxmemory-policy` for it, and what would you change about the architecture?
answer
- one instance, one policy, two data classes
- allkeys ⇒ loses sessions; noeviction ⇒ write outage
- volatile-lru + TTL enforced in the client wrapper
- split: cache (allkeys-lfu) vs state (noeviction + AOF)
- maxmemory ≈ 60–70% of RAM: frag, buffers, fork COW
basics
~20 sNo single policy is safe for both: allkeys-* can delete sessions and queued jobs, noeviction turns cache growth into write outages. Split into separate instances. If you cannot, use volatile-* with enforced TTLs on cache keys only, and keep 25–30% RAM headroom.
solid answer
~60 sThe honest answer is that the instance has two incompatible data classes and one per-instance policy knob. - `allkeys-lru`/`allkeys-lfu` make sessions and queued jobs silently evictable — a data-loss bug that surfaces as random logouts and lost work under load. - `noeviction` protects them, but then unbounded cache growth turns into `OOM command not allowed…` on every write path, including session creation. Cache pressure becomes a full outage. - `volatile-lru` is the least-bad single-instance compromise: TTL every cache key, never TTL the durable keys, and enforce that at the write path (a cache wrapper that requires a TTL argument). It fails if the volatile pool is too small — Redis then reverts to OOM errors. **Architecturally, split the instances**: a cache instance sized for hit rate with `allkeys-lfu`, and a state instance with `noeviction`, persistence enabled, alerting on memory well before the limit. Different data classes deserve different eviction, persistence, failover and capacity decisions. Size `maxmemory` to roughly 60–70% of container RAM: RSS exceeds it through fragmentation, client and replication buffers, and copy-on-write during an RDB fork.
code
text · 10 lines# cache.conf
maxmemory 8gb
maxmemory-policy allkeys-lfu
save "" # no RDB; regenerable data
# state.conf
maxmemory 4gb
maxmemory-policy noeviction # OOM error is the correct signal here
appendonly yes
appendfsync everysecgo deeper
Recognize the core conflict: eviction policy is per instance, so protecting sessions and bounding the cache cannot both be done with one setting.
Give the volatile-lru compromise with enforced TTLs, and explain why allkeys-* and noeviction each fail one of the two data classes.
Add sizing discipline — headroom for fragmentation, buffers and fork copy-on-write — plus monitoring of the volatile pool and per-class memory attribution.
Lead with the split by data class and the blast-radius and durability arguments behind it, including whether the job queue belongs in Redis at all, and define the migration path and its capacity model.
## Why this is a design question, not a config question `maxmemory-policy` is a property of the **instance**, not of a key. When one instance holds two data classes with opposite requirements, no value of that single setting is correct: | Policy | Cache under pressure | Sessions / queue | |---|---|---| | `allkeys-lru` / `allkeys-lfu` | degrades gracefully (misses) | **silently deleted — data loss** | | `noeviction` | writes fail, cache cannot grow | protected, but **every write path errors** | | `volatile-*` | evicts only TTL'd keys | protected **if** they truly have no TTL | Recognizing this trilemma is most of the answer. The rest is what you do about it. ## The single-instance compromise If you must stay on one instance today: 1. **TTL discipline.** Every cache key gets an expiry; durable keys never do. This is not a convention to document, it is a rule to enforce in code — a cache client whose `set` signature *requires* a TTL, so an untagged cache write is impossible to express. One forgotten TTL is one un-evictable key growing forever. 2. **Policy `volatile-lru`** (or `volatile-lfu` if batch scans pollute recency). Only the TTL'd cache is evictable. 3. **Monitor the eviction pool.** `INFO keyspace` reports `keys=` and `expires=` per database. If the volatile share shrinks, you are drifting toward the failure where Redis has no candidates and reverts to OOM errors — the same outage `noeviction` would have given you, but arriving later and less predictably. 4. **Alert on the durable set's absolute size**, not just total memory. Sessions and queue depth growing is a capacity signal that eviction will never relieve. 5. **Consider key-prefix separation plus `--bigkeys`/memory sampling** so you can attribute memory to each class and forecast when the split becomes mandatory. ## The architecture change Split by data class: - **Cache instance:** `allkeys-lfu` (or `-lru`), persistence often off or RDB-only, sized for hit rate, replicas optional, cheap to lose entirely. Losing it costs backend load, not correctness. - **State instance:** `noeviction`, AOF persistence with a chosen fsync policy, replicas plus failover, alerting long before the memory limit. Here an OOM error is the *correct* behavior: a loud back-pressure signal, not a bug. The split also decouples operations: you can flush, resize, restart or reshard the cache freely; you cannot do any of that to sessions. Blast radius separation is often a stronger argument than the eviction policy itself. A further question worth asking: does the pending-job list belong in Redis at all? A list used as a work queue with at-most-once semantics is a durability decision disguised as a data-structure choice. That may be fine — with AOF and replicas, and with jobs that are cheap to lose — but it should be a decision someone made, not a consequence of the cache instance being convenient. ## Sizing `maxmemory` `maxmemory` is compared against Redis's own `used_memory` accounting, which excludes replica output buffers and the AOF buffer, and it cannot know about allocator fragmentation. Real RSS is routinely 1.2–1.5x the configured limit, and more transiently: - **Fragmentation** — check `mem_fragmentation_ratio`; workloads with many differently-sized values fragment worse. `activedefrag` can help but costs CPU on the command-processing thread. - **Client output buffers** — a slow consumer of a big `LRANGE` or a Pub/Sub subscriber can accumulate hundreds of megabytes; `client-output-buffer-limit` bounds it. - **Replication backlog** — sized by `repl-backlog-size`, held whether or not a replica is attached. - **Fork copy-on-write** — during an RDB save or AOF rewrite, pages modified by the parent are duplicated. A write-heavy instance can add a large fraction of the dataset transiently; this is the classic cause of an OOM-kill on a box that "had plenty of room". A workable rule: set `maxmemory` around 60–70% of the container limit for a write-heavy persisted instance, closer to 75–80% for a read-heavy cache with persistence off. Then verify with `used_memory_rss` under real traffic including a background save, rather than trusting the arithmetic. ## What good sounds like in an interview Name the trilemma explicitly; refuse the false premise that a policy exists which protects sessions and bounds the cache simultaneously; give the compromise you would ship today and the split you would drive to; and show you know `maxmemory` is not the process's memory ceiling. Mentioning the fork copy-on-write spike and the `expires`-share monitor is what separates a designed answer from a recited one.
- You cannot split instances this quarter. What single change buys the most safety?Make TTLs unforgeable at the write path: a cache client whose set method requires an expiry, so no cache key can be created without one, combined with `volatile-lru`. That guarantees the eviction pool tracks the cache rather than slowly shrinking, and it keeps every durable key outside the candidate set. Pair it with an alert on the `expires`-to-`keys` ratio from `INFO keyspace` so you see drift before it becomes an OOM incident.
- Why set `maxmemory` well below the container's memory limit?Because the limit is checked against Redis's own accounting, which excludes replica output buffers and the AOF buffer and cannot see allocator fragmentation. On top of that, an RDB save or AOF rewrite forks the process, and every page the parent modifies afterwards is copied — a write-heavy instance can transiently add a large fraction of its dataset. Leaving 30–40% headroom is what keeps the kernel's OOM-killer out of your incident timeline.
- Is a Redis list an acceptable home for pending jobs in this design?It can be, but it is a durability decision, not just a data-structure one: with `appendfsync everysec` a crash loses up to a second of writes, and replication is asynchronous so a failover can lose recently pushed jobs. If those losses are acceptable and jobs are cheap to regenerate, a list (or a Stream with consumer groups for acknowledgement) is fine. If they are not, the queue belongs in a store whose durability guarantees match, and the point is that someone should have made that call deliberately.
saying these in an interview costs you the question
- Claiming some `maxmemory-policy` value protects sessions while still bounding cache growth — no per-key policy exists.
- Setting `allkeys-lru` on an instance holding queued jobs and calling it a cache configuration.
- Treating `maxmemory` as the process's memory ceiling and sizing the container to match it exactly.
- Ignoring fork copy-on-write during RDB/AOF rewrite when budgeting headroom.
- Answering only with tuning knobs and never questioning whether the two data classes belong together.