You operate three Redis deployments: a page-fragment cache, a user session store, and a small ledger of business events that no other system holds. For each, decide whether to run snapshots only, the append-only file only, or both — and justify the configuration you would set.
answer
- Where does the data come back from?
- Cache: snapshots only or none
- Sessions: both, everysec, preamble on
- Ledger: both + replicas + off-box + a real SoR
- AOF sets the window, RDB is the backup artifact
basics
~20 sCache: snapshots only, or nothing — data is rebuildable. Sessions: both, with the append-only file on everysec, since losing minutes of logins hurts but is survivable. Ledger: both, with a real replica and off-box backups, because Redis alone should not be the system of record.
solid answer
~60 sDecide by what a loss costs and where the data can be re-derived from. - **Page-fragment cache** — the source of truth is elsewhere, so persistence buys only warm-restart time. Snapshots only (`save` points, `appendonly no`), or persistence off entirely if a cold cache is tolerable and you want to avoid fork spikes. - **Session store** — losing sessions logs everyone out; annoying, not fatal. Enable both: `appendonly yes` with `appendfsync everysec` and the RDB preamble on, so the loss window is about a second and restarts are fast, plus save points to keep a snapshot artifact you can back up. - **Event ledger** — Redis is the only holder, so both, plus the honest caveat: `appendfsync always` only tightens the local window, and replication is asynchronous, so a failover can still lose acknowledged writes. Anything genuinely irreplaceable belongs in a store designed as a system of record, or is at minimum mirrored to one. The general rule: AOF choice sets the loss window, snapshots give you the portable backup, hybrid keeps recovery fast.
code
text · 17 lines# 1. page-fragment cache: warm restart optional, avoid fork pressure
appendonly no
save "" # or: save 900 1
stop-writes-on-bgsave-error no
# 2. session store: both mechanisms, ~1s window, fast reload
appendonly yes
appendfsync everysec
aof-use-rdb-preamble yes
save 3600 1
save 300 100
# 3. event ledger: tightest local window Redis offers, plus replicas/backups
appendonly yes
appendfsync always
aof-use-rdb-preamble yes
save 900 1go deeper
Say that a cache does not need persistence because the data can be rebuilt, while data that only lives in Redis needs the append-only file.
Give concrete settings per case — appendonly, appendfsync everysec, the preamble, and retained save points — and explain what each buys.
Reason from loss tolerance and recovery time, cover the fork/copy-on-write cost of snapshots, the MISCONF write-block behavior, and backing up from a replica.
Challenge the premise for irreplaceable data: name asynchronous replication and failover as the real loss modes, position Redis's guarantees honestly, and decide where a different system of record is required.
## The three questions that decide the configuration Before touching `redis.conf`, answer these: 1. **If this instance loses everything right now, where does the data come back from?** If the answer is "the database" or "the user just logs in again", you are choosing for convenience. If the answer is "nowhere", you are choosing for survival. 2. **How much loss is acceptable, in seconds or minutes?** That number picks between save points (minutes) and an fsynced command log (about a second). 3. **How fast must it come back?** Recovery time drives whether the append-only file carries an RDB body and whether you fail over to a replica instead of reloading. The mechanisms map cleanly onto those answers. Snapshots (`save <sec> <changes>`, `BGSAVE`) give you a compact, checksummed, point-in-time file that is trivial to copy off the box — a great *backup artifact*, a poor *durability guarantee*, because everything since the last one is gone. The append-only file gives you a fine-grained loss window set by `appendfsync`, at the cost of a bigger file and, historically, slow command-replay recovery — which is exactly what `aof-use-rdb-preamble yes` fixes by making the AOF base a binary snapshot. ## Deployment 1: page-fragment cache The origin service can regenerate every value. Persistence contributes only a warm start after a restart, which shortens the post-restart latency spike and origin load. Reasonable configuration: `appendonly no`, with modest save points (or none if you prefer a cold start). Note the real cost of keeping snapshots here is not disk, it is the fork: a write-heavy cache snapshotting frequently pays copy-on-write memory and a fork pause. Many teams run caches with persistence fully disabled (`save ""`) and accept the cold start, especially when `maxmemory` is set near the box limit and the copy-on-write headroom does not exist. If you do keep save points, also keep `stop-writes-on-bgsave-error` in mind: with it at the default `yes`, a failing background save makes Redis reject writes with `MISCONF`. For a pure cache that is usually the wrong trade — a cache should keep serving even if it cannot persist — so this is one of the few places where turning it off is defensible. ## Deployment 2: session store Sessions are re-creatable but only by inconveniencing every user at once. A restart that drops them is a visible incident; losing a few seconds of newly created sessions is not. Configuration: `appendonly yes`, `appendfsync everysec`, `aof-use-rdb-preamble yes`, plus save points retained so you still produce a `dump.rdb` for backups. This is the canonical "run both" case. The AOF bounds the loss window to roughly a second, the RDB body inside it keeps restart fast, and the separate snapshot is what your backup job ships off the machine. Add a replica so that a node failure is a failover rather than a reload, and remember that sessions usually carry TTLs — expirations are recorded in both formats, so a restore does not resurrect long-dead sessions. ## Deployment 3: business-event ledger Here the honest senior answer starts with a challenge to the premise. Redis's durability story is bounded: `appendfsync always` fsyncs per event loop iteration, which tightens the local window but costs throughput and still cannot survive disk loss; replication is **asynchronous**, so a primary can acknowledge a write and fail before any replica has it; and automatic failover promotes a replica that may be behind. `WAIT` can block until N replicas acknowledge, which raises confidence but is not a consensus commit. So: run both persistence mechanisms (`appendonly yes` with `everysec` or `always` depending on measured throughput, preamble on, save points on), take snapshots off-box on a schedule, keep replicas — and either mirror the events into a store designed to be a system of record, or accept and document the failure modes in which Redis loses acknowledged data. Choosing Redis as the sole home for irreplaceable data is a decision to defend explicitly, not a configuration detail. ## The decision heuristic to state out loud - Rebuildable data → snapshots only, or nothing. - Painful-but-survivable loss → both, `everysec`, preamble on. - Irreplaceable data → both, plus replicas, plus off-box backups, plus a second system — and be explicit about what Redis still cannot promise. Also keep the operational asymmetry in view: AOF-only is an unusual choice because you give up the cheap, portable snapshot artifact for backups and disaster recovery while gaining nothing on durability that both-enabled does not already give you. When someone proposes AOF-only, the usual real motivation is avoiding a second fork for save points on a memory-tight box — which is a capacity conversation, not a durability one.
- Why is running the append-only file alone, with no save points, an unusual configuration?Because it costs you the snapshot artifact without buying durability. The AOF already bounds the loss window; the separate `dump.rdb` is what you copy off the box, verify, and load elsewhere for backups and clones. The main honest reason to skip save points is to avoid a second fork and its copy-on-write memory spike on a memory-tight instance — and then you should be taking snapshots from a replica instead.
- Does `appendfsync always` mean an acknowledged write can never be lost?No. It bounds loss on that one node to essentially nothing for a clean process crash, but it does not survive the disk or machine dying, and replication is asynchronous — the primary can acknowledge a write and fail before any replica received it, after which a promoted replica serves a dataset missing that write. `WAIT` raises the bar by blocking until replicas acknowledge, but it is not a consensus protocol.
- How does adding a replica change the persistence configuration you would choose on the primary?It lets you move snapshot cost off the primary: run save points on the replica and back up from there, keeping the primary's fork pressure to just the AOF rewrite. It does not reduce the durability requirement on the primary, because replication is asynchronous and a replica can be behind at the moment of failure.
saying these in an interview costs you the question
- Enabling maximum durability everywhere without asking what the data costs to lose
- Treating `appendfsync always` as a guarantee that acknowledged writes survive failover
- Assuming persistence protects a cache's availability — a failing background save can actually block writes via MISCONF
- Recommending AOF-only as the modern default and dropping snapshot backups
- Ignoring the fork and copy-on-write cost of snapshotting on a memory-tight, write-heavy instance