skip to content

How would you take a backup of a Redis dataset that you can actually restore from, and what does the restore procedure look like?

level: seniorimportance: must knowfreq 44%

answer

  1. RDB is the backup artifact; AOF is the durability mechanism
  2. BGSAVE on a replica; child renames temp file, so dump.rdb is always complete
  3. Off-host, timestamped, retained, verified by real restore
  4. Restore with appendonly no first, then CONFIG SET appendonly yes
  5. Replicas are not backups: FLUSHALL replicates instantly

basics

~20 s

Use RDB as the backup artifact: trigger BGSAVE (ideally on a replica), copy dump.rdb off-host with a timestamped name and retention. Restore by stopping the server, placing the file in dir, starting with appendonly no, then enabling AOF at runtime. Verify by actually restoring.

solid answer

~50 s

The RDB is the backup artifact: one self-contained, checksummed, point-in-time file. The AOF is a durability mechanism, not a convenient backup, and in 7.x it is a directory plus manifest that must be copied as a consistent unit. Procedure: run `BGSAVE` (or rely on save points) on a **replica**, so the master never forks, then copy `dump.rdb` to another host. The child writes a temp file and renames it, so an existing `dump.rdb` is always complete and safe to copy. Timestamp the names, keep a retention window, and store off-host, since a backup on the same disk does not survive the failure it exists for. Restore: stop the server, put the file at `dir`/`dbfilename` with correct ownership, start with `appendonly no` so the RDB is loaded, check `DBSIZE`, then `CONFIG SET appendonly yes` and `CONFIG REWRITE`. Verify by restoring into a scratch instance on a schedule. A backup never restored is a hypothesis.

code

text · 15 lines
text
# --- backup (run against a replica) ---
BEFORE=$(redis-cli -h replica LASTSAVE)
redis-cli -h replica BGSAVE
while [ "$(redis-cli -h replica LASTSAVE)" = "$BEFORE" ]; do sleep 1; done
scp replica:/var/lib/redis/dump.rdb ./dump-$(date -u +%Y%m%dT%H%M%SZ).rdb
redis-check-rdb ./dump-*.rdb

# --- restore ---
systemctl stop redis
cp dump-20260814T031500Z.rdb /var/lib/redis/dump.rdb
chown redis:redis /var/lib/redis/dump.rdb
mv /var/lib/redis/appendonlydir /var/lib/redis/appendonlydir.old   # AOF must not shadow the RDB
redis-server /etc/redis/redis.conf --appendonly no
redis-cli DBSIZE
redis-cli CONFIG SET appendonly yes && redis-cli CONFIG REWRITE

go deeper

for a junior

Know that dump.rdb is the file you back up, that BGSAVE produces it, and that restoring means putting it in the data directory before starting the server.

for a middle

Add the atomic-rename detail, off-host copies with timestamps, and the appendonly-no step during restore.

for a senior

Run backups on a replica, verify with redis-check-rdb and real restores, and articulate why replicas do not replace backups.

for a principal

Define retention, RPO and restore-time objectives, treat snapshots as sensitive data requiring encryption and access control, and account for per-node snapshots and slot mapping in a clustered deployment.

## Choose the artifact: RDB For backups, use the RDB snapshot. It is a single file, self-contained, compact, checksummed with CRC64, and it represents one instant. Loading it is fast because it is a binary image rather than a command replay. The AOF is a poor backup artifact even though it is the better durability mechanism. On Redis 7.x it is a directory of base and incremental files governed by a manifest; a copy is only meaningful if the parts and the manifest are consistent with each other, and a rewrite completing mid-copy breaks that. On 6.x the single file is simply large and slow to load. Use AOF for the loss window, RDB for the backups; they answer different questions. ## Taking the snapshot `BGSAVE` forks a child that writes the dataset to a temporary file and then `rename()`s it over `dbfilename`. Two useful consequences: - An existing `dump.rdb` is always a complete, consistent file. You can copy it at any time without coordination, and you will never capture a half-written snapshot. - The fork has the usual cost, stalling the server proportionally to mapped memory and consuming copy-on-write memory. Which is why: **Take backups on a replica.** The replica holds the same data and its fork stalls nobody who is serving traffic. Note the replica lags by the replication delay, so the snapshot is that much older; for backups measured in hours of retention this is irrelevant. `SAVE` (no `BG`) performs the same work in the main process and blocks the server for the entire duration. It exists for scripted shutdown paths and small instances; it is not a backup strategy. `LASTSAVE` returns the unix time of the last successful save, which is how a script confirms a `BGSAVE` finished before it starts copying. ## Making it a real backup - **Off-host.** A copy on the same volume does not survive the disk failure or the instance termination it is supposed to protect against. Ship to object storage or a separate backup host. - **Timestamped names and retention.** `dump-<host>-<utc-timestamp>.rdb`, with a policy such as hourly for a day, daily for a month. Overwriting a single file means a corrupt or empty snapshot can silently replace your last good one. - **Immutability or versioning** where the threat model includes an operator or an attacker running `FLUSHALL`; replication and even a naive backup rotation will faithfully propagate that. - **Verification.** `redis-check-rdb dump-....rdb` catches structural damage and checksum failure. It does not prove the contents are what you expect. - **Restore rehearsal.** Periodically load the backup into a scratch instance and assert on `DBSIZE` and a few known keys. This is the only step that turns "we have backups" into a fact. ## Restoring The order matters because of Redis's load rule: with `appendonly yes`, the AOF is authoritative and `dump.rdb` is ignored entirely. So a restore that just drops the file in place and restarts an AOF-enabled server restores nothing. 1. Stop the server (`SHUTDOWN NOSAVE` if you must be certain it does not overwrite files on the way out; be equally certain that is what you want). 2. Confirm the target paths: `CONFIG GET dir` and `CONFIG GET dbfilename`. 3. Place the backup file there, with the correct owner and mode for the redis user. Wrong ownership is the most common reason a restore "does nothing". 4. Ensure AOF is not enabled for this start, either by starting with `appendonly no` or by moving the existing `appendonlydir` aside. 5. Start the server and check the log line `DB loaded from disk`, then `DBSIZE` and a few known keys. 6. Re-enable durability on the running server: `CONFIG SET appendonly yes` (which rewrites the AOF from memory) and `CONFIG REWRITE` to persist the setting. For a replicated deployment, restore one node and let replication reseed the others rather than restoring each independently; otherwise you can end up with nodes disagreeing about which snapshot is current. ## Things people get wrong - **"Replicas are our backup."** They are not. A replica applies `DEL`, `FLUSHALL` and an accidental mass expiry within milliseconds. Replication protects against node loss, not against a bad write. Backups protect against a bad write, not against staleness. You need both. - **Backing up the AOF directory file-by-file** without the manifest, or while a rewrite is running. - **Assuming the RDB is encrypted or access-controlled.** It is a plain data file containing everything, including anything sensitive stored in Redis; treat it with the same controls as a database dump, encrypted at rest and in transit. - **Forgetting the cluster case.** In Redis Cluster each master owns a slice of the keyspace, so a backup is a set of per-node snapshots taken close together, and a restore must line up with the slot assignment. Snapshots are per-node and not globally atomic. ## Also worth knowing `DUMP` and `RESTORE` serialize individual keys, and `MIGRATE` moves them between instances. Those are key-level tools for repair and rebalancing, not a backup strategy for a whole dataset, but naming them shows you know where the boundary is.

  • Why is a replica not a backup?
    Replication copies whatever the master does, including destructive commands. A `FLUSHALL`, a mass `DEL`, or an application bug that rewrites keys reaches the replica within milliseconds, so the replica is destroyed along with the master's data. Replication protects against node and host failure; only a point-in-time copy kept elsewhere protects against a bad write, and you need both.
  • Is it safe to copy dump.rdb while Redis is running?
    Yes. The snapshot child writes to a temporary file and atomically renames it over `dbfilename`, so the file you see on disk is always a complete, checksummed snapshot, never a partially written one. What you cannot control is which instant it represents, so scripts typically call `BGSAVE` and wait for `LASTSAVE` to change before copying.
  • How does backing up a Redis Cluster differ from backing up a single instance?
    Each master owns a subset of hash slots, so a backup is a set of per-node RDB files taken as close together as possible; there is no globally atomic snapshot across nodes. A restore must place each file on a node owning the matching slots, so the slot map has to be recorded alongside the snapshots. Backing up from replicas keeps the fork cost off the serving masters.

saying these in an interview costs you the question

  • Treating replicas or a replicated deployment as the backup strategy
  • Copying the AOF directory piecemeal without the manifest, or during a rewrite
  • Restoring an RDB onto a server with appendonly yes and expecting it to load
  • Never testing a restore, so retention and file integrity are assumed rather than known
  • Forgetting file ownership and permissions after copying the file into place

context