skip to content

Redis snapshot files are a compact binary format rather than a log of write commands. What operational advantages does that format buy, and where does it constrain you?

level: seniorimportance: should knowfreq 35%

answer

  1. stores state, not history
  2. internal encodings + LZF + CRC64
  3. one portable checksummed artifact
  4. full sync and DUMP/RESTORE use the same format
  5. opaque, whole-dataset only, version coupled

basics

~20 s

A snapshot is a compressed, checksummed image that loads by deserialization instead of replaying millions of commands, so restarts and replica seeding are fast and the file is small enough to copy off-box. The costs: not human-readable or editable, and version-coupled on downgrade.

solid answer

~50 s

The snapshot stores values in their internal encodings, optionally LZF-compressed, with a CRC64 checksum. That gives three operational wins: - **Fast load.** Restoring is bulk deserialization, not re-executing every command that built the data. A list assembled by a million `RPUSH` calls is one serialized key. - **Small, self-contained artifact.** One checksummed file per dataset — easy to copy off the host, verify, archive, and load into another instance, which is why it stays the backup format even on servers running the append-only file. - **Cheap replica seeding.** Full synchronization streams exactly this image to a replica before the command stream takes over. The constraints: it is opaque — you cannot grep it, edit out a bad command, or diff two snapshots; it is a whole-dataset artifact with no incremental form; and it carries an RDB version number, so a snapshot from a newer Redis will not load on an older binary, which shapes downgrade and cross-version restore plans.

code

text · 7 lines
text
$ head -c 9 dump.rdb
REDIS0011                 # magic + RDB format version

$ redis-check-rdb dump.rdb
[offset 0] Checking RDB file dump.rdb
[offset 26] AUX FIELD redis-ver = '7.2.4'
\o/ RDB looks OK! \o/

go deeper

for a junior

Say that the snapshot stores the data itself in a compact binary form, so loading it is faster than re-running every past command, and the file is easy to copy as a backup.

for a middle

Add the concrete properties — internal encodings, LZF compression, CRC64 checksum — and that the same format seeds replicas and backs DUMP/RESTORE.

for a senior

Balance the wins against opacity, whole-dataset granularity, and RDB version coupling, and connect it to why the hybrid append-only file uses a snapshot as its base.

for a principal

Position it as state-versus-history: choose the image for recovery-time objectives, portability, and seeding; keep a log where inspectability and fine-grained loss windows matter; and treat version coupling as a constraint on upgrade and rollback design.

## What is actually in the file An RDB file is a serialization of the live dataset. Its structure: a magic header `REDIS` plus a four-digit RDB version, auxiliary metadata (Redis version, creation time, memory used), then per database a sequence of entries — key, value, expiration if set — with the value written in a type-tagged form that mirrors Redis's internal encoding. Small hashes stored as listpacks are written as listpacks; a hash table is written as a hash table. Optional LZF compression (`rdbcompression yes`) shrinks strings, and a CRC64 checksum (`rdbchecksum yes`) at the end lets the loader detect truncation or corruption. Contrast that with the append-only file's default form: RESP text, one entry per write command the server accepted. Rebuilding from it means re-running that command stream against an empty dataset. ## Advantage one: load time The difference is structural, not incremental. Replaying commands pays the full command pipeline for every historical write: parse, look up the command table, execute, mutate structures, possibly rehash and re-encode as a collection grows through its encoding transitions. Deserializing a snapshot pays parse-and-build once per key, in its final encoding, with the sizes known up front. The practical effect is that a dataset built by a very large number of small operations loads dramatically faster from a snapshot than from a command log. This is exactly why hybrid persistence exists: enabling the RDB preamble makes the append-only file's base a snapshot image so it inherits this property. Fast load matters operationally in three places: process restarts after a deploy or crash, promoting or rebuilding a node, and any recovery where your time objective is measured in minutes. ## Advantage two: it is a real artifact A snapshot is one file, complete, checksummed, and internally consistent as of a single instant. That makes it a natural unit for operations that a log cannot serve as well: - **Backups.** Copy `dump.rdb` off the host on a schedule; verify it; keep versions. `redis-check-rdb` will validate one without loading it into a server. - **Cloning and staging.** Drop a production snapshot into a staging instance to reproduce a bug against real data shapes. - **Migration.** Move a dataset between hosts or clusters by moving one file. This is why "run both" is the standard recommendation for important data: the append-only file gives you a tight loss window, the snapshot gives you the artifact you can hold in your hand. ## Advantage three: replication seeding When a replica cannot resume from the replication backlog, it does a full synchronization, and what the primary sends is exactly this format — either produced to disk and then shipped, or, with `repl-diskless-sync` (default on in modern Redis), streamed straight from the forked child's socket. Compactness translates directly into shorter sync windows and less network transfer. The same format also underpins per-key transfer: `DUMP` returns a value in RDB serialization plus a version footer and checksum, `RESTORE` rebuilds it, and cluster key migration is built on that pair. ## Where the format constrains you **Opacity.** You cannot read it, grep it for a key, hand-edit a poisoned value out of it, or meaningfully diff two of them. When someone writes a bad value and you want to surgically remove it from the persisted state, a command log is inspectable and a snapshot is not. Tooling exists to dump RDB contents, but that is a conversion step, not a text file. **All-or-nothing granularity.** There is no incremental snapshot: every save writes the whole dataset. That is what makes frequent snapshots expensive in fork pauses, copy-on-write memory, and I/O, and it is why the loss window between snapshots is coarse. A log is naturally incremental; an image is not. **Version coupling.** The header carries an RDB version, and a file written by a newer Redis can contain encodings an older binary does not understand, so it will refuse to load. This constrains downgrades, cross-version restores, and any "restore last week's backup into the old build" plan. The same applies to `DUMP`/`RESTORE` payloads, which carry their own version footer and are rejected by an incompatible server. **Trust boundaries.** Loading a snapshot means letting the process build internal structures from bytes you were handed. Restoring an untrusted or third-party file — or accepting arbitrary `RESTORE` payloads from untrusted clients — is a real risk surface; Redis added `sanitize-dump-payload` for deep validation of `RESTORE` input precisely because of it. ## The summary an interviewer wants The binary format buys speed of load, small portable artifacts, and cheap replica seeding — all consequences of storing the *state* rather than the *history*. It costs inspectability, incrementality, and cross-version freedom, which is why serious deployments keep it as the backup and seeding artifact while relying on a command log for the fine-grained loss window.

  • Where else in Redis does this same serialization format appear besides the snapshot file?
    In replication, where a full synchronization ships the dataset in RDB form (streamed straight from the forked child when diskless sync is on), and in the `DUMP`/`RESTORE` command pair, which serializes and rebuilds a single value using the same encoding plus a version footer and checksum. Cluster key migration is built on `DUMP`/`RESTORE`, and hybrid persistence uses the format for the append-only file's base.
  • What practical problem does the format's version number cause?
    It blocks downward compatibility. A snapshot written by a newer Redis can contain encodings an older binary does not know, so the older server refuses to load it. That constrains rollbacks after an upgrade, restoring an old backup into a differently-versioned instance, and mixed-version tooling — the mitigation is to plan version-matched restores, or to migrate data through a live channel like replication or a logical export instead of a file.
  • Given the snapshot loads faster, why keep the append-only file at all?
    Because speed of load is not durability. A snapshot only reflects the last save point, so the exposure is minutes; the command log with `appendfsync everysec` narrows it to about a second. With the RDB preamble enabled the append-only file also carries a snapshot body, so you get the fast load and the tight window together, while the separate dump.rdb remains the artifact you back up.

A photograph of a finished building versus the complete construction diary. The photo is small and instantly tells you what stands there; the diary tells you how it got there and lets you find the day a wall went up wrong.

saying these in an interview costs you the question

  • Claiming the snapshot format is human-readable or hand-editable
  • Thinking Redis supports incremental snapshots that only write changed keys
  • Assuming any Redis version can load any snapshot file
  • Believing replication ships write commands rather than an RDB image during a full synchronization
  • Treating a snapshot as a durability mechanism rather than a point-in-time artifact

context