Your Redis server refuses to start, logging that the append-only file is bad after an unclean shutdown. Walk through how you diagnose and repair it, including what the redis-check-aof --fix tool actually does.
answer
- Copy first: every repair is destructive
- Truncated tail (tolerated by default) vs mid-file corruption (never)
- --fix truncates from first bad command to EOF, discarding the rest
- Byte offset vs file size tells you the cost of the repair
- redis-check-rdb verifies only, no repair exists
basics
~20 sCopy the file first. Distinguish a truncated tail (tolerated by default via aof-load-truncated) from mid-file corruption, which always blocks startup. redis-check-aof --fix truncates from the first invalid command to end of file, discarding everything after it. For RDB, redis-check-rdb only verifies; it cannot repair.
solid answer
~50 sFirst, **make a copy** of the AOF (in 7.x, the whole `appendonlydir`); every repair is destructive. Then classify the damage. A **truncated tail** means the last command was partially written, the normal outcome of a `kill -9` mid-write. With `aof-load-truncated yes` (default) Redis loads up to the last valid command, warns, and starts. If it is refusing to start, either that setting is `no` or the damage is **mid-file**, which is never tolerated and points at disk full, hardware or filesystem trouble. `redis-check-aof` verifies the file (in 7.x you point it at the manifest). With `--fix` it finds the first invalid command and truncates from there to EOF, then confirms interactively. That discards every valid command after the corruption too, so know what you are giving up. Redis 7 also offers `--truncate-to-timestamp` when `aof-timestamp-enabled` was on. For RDB, `redis-check-rdb` validates structure and checksum but has no repair. Restore from a backup or promote a replica instead.
code
text · 13 lines# 0. Never repair the only copy
cp -a /var/lib/redis/appendonlydir /var/lib/redis/appendonlydir.bak
# 1. Verify only: where is the first problem?
redis-check-aof /var/lib/redis/appendonlydir/appendonly.aof.manifest
# AOF analyzed: filename=appendonly.aof.2.incr.aof, size=1073741824, ok_up_to=1073741100, diff=724
# -> 724 bytes after 1 GB: a torn tail, cheap to fix
# 2. Repair (truncates from the first invalid command to EOF)
redis-check-aof --fix /var/lib/redis/appendonlydir/appendonly.aof.manifest
# RDB: verification only, there is no --fix
redis-check-rdb /var/lib/redis/dump.rdbgo deeper
Know that redis-check-aof exists, that --fix truncates the file, and that you copy the file before running it.
Distinguish a truncated tail from mid-file corruption, and connect the tail case to aof-load-truncated.
Use the reported offset to size the loss, weigh replica promotion and backups against repairing, and investigate the underlying cause afterwards.
Own the decision framework: repair is one recovery option among promotion and restore, chosen by comparing data loss and time to recover, and the incident should feed back into RPO and backup policy.
## Step zero: copy before you touch anything Every repair path here is destructive and irreversible. Before running any tool, copy the artifact aside: on Redis 7.x that is the entire `appendonlydir` including the manifest, on 6.x the single `appendonly.aof`. If the repair discards more than you expected, the copy is the only way back. ## Read the log and classify the damage Redis tells you which case you are in. **Case 1: truncated tail.** The file ends in the middle of a command. This is the *expected* result of the process being killed while writing, and it is benign: the missing bytes are the last fraction of a second of traffic. With `aof-load-truncated yes`, which is the default, Redis loads everything valid, logs a warning such as "AOF ... was truncated" and starts normally, then continues appending after truncating the partial tail. If the server refuses to start on a truncated file, someone set `aof-load-truncated no`, which converts this into a manual decision on purpose. **Case 2: corruption in the middle.** An invalid byte sequence appears somewhere other than the tail. Redis refuses to start regardless of `aof-load-truncated`, because everything after the bad point is unreadable, and silently discarding it would be data loss the operator never approved. This is not a normal consequence of a crash; suspect a full disk (a partial write in the middle can occur when combined with later appends), a failing device, a filesystem bug, or a file edited or truncated by a well-meaning human. **Case 3: the RDB is bad.** A checksum mismatch (`rdbchecksum yes` is the default) or a structural error. There is no repair. ## redis-check-aof Verification mode, which changes nothing: ``` redis-check-aof appendonlydir/appendonly.aof.manifest # Redis 7.x redis-check-aof appendonly.aof # Redis 6.x and earlier ``` It reports either that the AOF is valid, or the byte offset of the first problem and whether the file is truncated or corrupt. That offset is the important output: compare it with the file size to judge how much sits after the damage. A problem 40 bytes from the end is a torn tail; a problem at 12 percent of the file means the repair would throw away 88 percent of the log. Repair mode: ``` redis-check-aof --fix appendonlydir/appendonly.aof.manifest ``` What it does is deliberately simple: it locates the first invalid command and **truncates the file from that point to the end**, after asking for confirmation. It does not reconstruct, skip, or splice around the bad region. Consequently every valid command after the corruption is discarded as well. That is why the byte offset matters, and why the copy in step zero matters. Redis 7.0 added timestamp annotations to the AOF (`aof-timestamp-enabled`, off by default), which periodically write a timestamp comment into the log. When they are enabled you additionally get: ``` redis-check-aof --truncate-to-timestamp <unix-time> appendonlydir/appendonly.aof.manifest ``` which truncates the log at a chosen point in time. That is the closest Redis comes to point-in-time recovery, and it is also useful for undoing a bad bulk write, provided you enabled the annotations *before* the incident. ## redis-check-rdb ``` redis-check-rdb dump.rdb ``` It walks the file, validates the structure and verifies the CRC64 checksum, and reports what it found. There is no `--fix`. An RDB is a binary image with no redundancy, so a damaged one is simply not restorable. Your options are a previous backup, a replica's copy, or the AOF if one exists. This asymmetry is worth stating explicitly, because candidates often assume the tools are symmetric. ## Deciding what to do, not just how Repairing is one option among several, and often not the best: 1. **Is there a healthy replica?** Promoting it is usually faster and loses less than a truncation, since the replica has everything up to its replication lag. Compare that lag with the amount the truncation would discard. 2. **Is there a recent backup?** If the corruption is early in the file, a backup plus a short gap may beat a repair that keeps 12 percent of the log. 3. **Only then repair**, from the copy, verify with `redis-check-aof` (no `--fix`) that the result is clean, start the server, check `DBSIZE` and spot-check known keys. ## Afterwards A mid-file corruption is a signal, not just an incident. Check disk free space and the `aof_last_write_status` and `aof_last_bgrewrite_status` fields, look for ENOSPC in the log, check the device's health, and confirm nobody is editing or rotating files under the server. Then re-examine whether the durability posture matched the expectation: if losing the truncated tail hurt, `appendfsync` and the backup cadence deserve another look.
- redis-check-aof reports the file is fine up to 20 percent of its length. Would you run --fix?Almost certainly not as a first move, because `--fix` truncates from that point to the end and would discard 80 percent of the log. I would first look for a healthy replica to promote or a recent backup, and compare what each option loses. If a repair is still the best available outcome, I would run it against a copy and verify the result before putting it into service.
- Why does redis-check-rdb have no repair mode when redis-check-aof does?An AOF is a sequence of independent commands, so cutting it at a boundary still leaves a valid, replayable prefix. An RDB is a single binary image with internal structure and one CRC64 over the whole file; there is no prefix that is independently meaningful and no redundancy to reconstruct from. A damaged RDB is therefore restored from another copy, never repaired.
saying these in an interview costs you the question
- Running --fix directly on the only copy of the file
- Assuming --fix repairs or skips the corrupt region rather than truncating everything after it
- Expecting redis-check-rdb to have a repair option
- Treating a mid-file corruption as a normal crash artifact rather than a hardware, disk-space or human cause
- Ignoring a healthy replica or a recent backup and going straight to a lossy truncation