skip to content

You operate a 60 GB Redis instance with append-only-file persistence, and append-only-file rewrites cause periodic latency spikes and near-out-of-memory events. How would you decide when and how rewrites should happen?

level: principalimportance: should knowfreq 26%

answer

  1. Two budgets: fork-stall latency and restart time
  2. Host first: overcommit=1, THP never, maxmemory headroom
  3. percentage 0 + scheduled BGREWRITEAOF off-peak
  4. Persist on a replica so the master never forks
  5. Sharding fixes fork, COW, rewrite duration and RTO together

basics

~20 s

Decide with two budgets: acceptable fork-stall latency and acceptable restart time. Then choose among fewer rewrites (raise the percentage or disable auto and schedule off-peak), moving persistence to a replica, and sharding so no node maps 60 GB. Fix the host first: overcommit on, huge pages off.

solid answer

~50 s

I would frame it as two budgets. **Latency budget**: how long a fork stall clients tolerate, measured from `latest_fork_usec`, and how much copy-on-write headroom the box has. **Recovery budget**: how long a restart may take, which is what a larger AOF buys you in exchange for fewer rewrites. Sequence: 1. Fix the host: `vm.overcommit_memory=1`, transparent huge pages off, `maxmemory` set so dataset plus copy-on-write fits physical RAM. 2. Take rewrites off the automatic path during peak: raise `auto-aof-rewrite-percentage` (100 to 200 or 300) and `auto-aof-rewrite-min-size`, or set the percentage to `0` and drive `BGREWRITEAOF` from a scheduler in a quiet window, with an alert on `aof_current_size` so a broken schedule is visible. 3. Structurally, stop forking on the serving node: persist on a replica, or shard so each node maps 8 to 15 GB, which shrinks fork stall, copy-on-write, rewrite duration and restart time at once. And I would question whether a pure cache needs AOF at all.

code

text · 9 lines
text
# Stop Redis from choosing the moment
redis-cli CONFIG SET auto-aof-rewrite-percentage 0
redis-cli CONFIG REWRITE

# Drive it yourself in the quiet window (crontab)
17 3 * * *  redis-cli -h 127.0.0.1 BGREWRITEAOF

# ... and alert, because a silent schedule failure fills the disk
redis-cli INFO persistence | grep -E 'aof_current_size|aof_last_bgrewrite_status'

go deeper

for a junior

Recognize that a very large instance makes background rewrites expensive and that they can be scheduled or made less frequent.

for a middle

Name the concrete settings and the direct trade: fewer rewrites means a bigger file and a slower restart.

for a senior

Diagnose from metrics, fix host settings, schedule rewrites off-peak, and propose replica-side persistence with its caveats.

for a principal

Lead with explicit RPO and RTO budgets, treat the 60 GB single instance as the underlying design problem, and sequence stabilization before the structural change to sharding or replica-side persistence.

## Name the budgets before touching knobs There is no correct rewrite frequency in the abstract. The decision is bounded by two explicit numbers, and a good answer starts by asking for them. **Latency budget.** Every rewrite begins with a `fork()` that stops the server for a period proportional to mapped memory. Measure it, do not estimate it: `latest_fork_usec` after a rewrite on this exact hardware. At 60 GB, several hundred milliseconds is a realistic figure, worse on virtualized hosts. Compare that to client timeouts and to what the calling services do when Redis stalls; a 300 ms stall behind a 100 ms timeout is a partial outage every time it happens. **Recovery budget.** The AOF's size is roughly your restart time, because loading it means replaying it (the base part loads faster, being RDB-format, but the incremental part is replay). Rewriting less often means a larger file means a longer cold start. If the service tolerates 10 minutes of restart, a large file is fine; if it must be back in 60 seconds, the file must stay small and you need frequent rewrites, which pushes you towards the structural fixes instead. The near-OOM part of the symptom adds a third constraint: **memory headroom**, since copy-on-write during a rewrite adds memory proportional to write rate times rewrite duration. ## Fix the host before tuning Redis These are not tuning choices; they are prerequisites, and skipping them makes everything else look worse than it is. - `vm.overcommit_memory=1`. Without it the kernel can refuse the fork outright, and Redis logs a background-save failure while the AOF quietly keeps growing. - Transparent huge pages set to `never`. With THP on, each copy-on-write fault duplicates 2 MB instead of 4 KB, which is very often the direct cause of "rewrites nearly OOM us". - `maxmemory` set so that the dataset plus the expected copy-on-write growth fits physical RAM. On a write-heavy instance the conservative rule is that the dataset should be around half of RAM. - Check `aof_delayed_fsync` and disk latency. If the disk cannot absorb the child's write of 60 GB plus the ongoing appends, the whole rewrite window degrades, not just its start. ## The knobs, and what each one really trades **Raise the thresholds.** `auto-aof-rewrite-percentage 200` or `300`, and `auto-aof-rewrite-min-size` raised well above the default 64 MB so it is meaningful next to a 60 GB dataset. Fewer rewrites, less frequent pain, larger file, slower recovery. Cheap and reversible. **Disable the automatic trigger and schedule it.** `auto-aof-rewrite-percentage 0` plus a cron or scheduler issuing `BGREWRITEAOF` in the quietest window. This is the highest-leverage tuning move because it moves the cost to when the cost is lowest: fewer writes during the window means less copy-on-write, and fewer clients means the fork stall matters less. The price is that you now own the schedule: alert on `aof_current_size` and on `aof_last_bgrewrite_status`, because a schedule that silently stops fills a disk. **`no-appendfsync-on-rewrite yes`.** Removes fsync stalls in the parent while a child runs, at the price of losing durability for the entire duration of the rewrite. Legitimate, but state the trade to stakeholders rather than slipping it in as a performance tweak. **Verify the base is RDB-format** (`aof-use-rdb-preamble yes`, the default since 4.0). If someone turned it off, the file is far larger and slower to load, and the rewrite itself is more expensive. ## The structural answers, which are usually the real ones **Persist on a replica.** Run the master with `appendonly no` and no save points, and let a replica carry AOF and RDB. The master never forks for persistence, so the latency spike disappears from the serving path. The caveats belong in the answer: the replica is asynchronous, so the loss window widens by replication lag, and a master restarting with no persistence comes back empty, so promotion and reseeding procedure must be written down and rehearsed. **Shard.** Ten nodes of 6 GB instead of one of 60 GB reduces fork stall, copy-on-write, rewrite duration and restart time simultaneously, and it removes the single-instance memory ceiling. It is more moving parts, and it is the answer most large deployments converge on. **Question the requirement.** If this instance is a cache whose contents can be rebuilt from a source of truth, the honest recommendation may be no AOF at all: RDB snapshots on a replica for warm restarts, or nothing. A great deal of persistence pain is paid by systems that do not need durability, only warm start. ## How I would sequence it Stabilize first (host settings, thresholds raised, rewrites scheduled off-peak) because those are hours of work. Then plan the structural change (replica-side persistence or sharding) because that is what actually removes the class of problem. Then define the RPO and RTO in writing, and pick the persistence configuration that meets them, rather than inheriting defaults and discovering the numbers during an incident.

  • If you move persistence to a replica, what have you given up?
    Replication is asynchronous, so the replica's files lag the master by the replication delay; on a master crash you lose everything not yet shipped. You also lose the ability to restart the master into a warm dataset, since a master with persistence disabled comes back empty and would replicate an empty dataset outward if it is restarted as master. Both require a written, rehearsed promotion and reseeding procedure.
  • How would you decide between raising the rewrite thresholds and sharding the instance?
    By whether the restart budget survives the larger file. Raising thresholds trades recovery time for fewer spikes, so it works when a long cold start is acceptable and the box has copy-on-write headroom. When neither is true, the instance is simply too large for forked persistence and sharding is the only move that improves fork stall, memory headroom and recovery time at once.
  • Your alert fires that the AOF has grown far beyond expectations. What do you check first?
    `aof_last_bgrewrite_status` and the log. The common causes are a rewrite failing repeatedly (disk full, or fork refused because overcommit is disabled) and a scheduled rewrite that stopped running after the automatic trigger was disabled. Both leave durability intact while the file grows, so the failure mode is disk exhaustion and a very slow restart rather than data loss.

saying these in an interview costs you the question

  • Answering only with config values and never asking for the latency and restart budgets
  • Recommending frequent rewrites to keep the file small without accounting for fork and copy-on-write cost
  • Proposing no-appendfsync-on-rewrite yes as a free performance win
  • Treating replica-side persistence as risk-free, ignoring replication lag and the empty-master restart hazard
  • Ignoring host settings (transparent huge pages, overcommit) and blaming Redis for the memory spikes

context