A Redis instance serves sub-millisecond responses most of the time but shows recurring multi-hundred-millisecond spikes. Walk through the causes you would investigate and the evidence that distinguishes them.
answer
- blocking command / big DEL / Lua → SLOWLOG
- fork stall → latest_fork_usec, THP off, save on replica
- swap = fatal; eviction-cycle at maxmemory
- TTL cohort → jitter TTLs, LATENCY expire-cycle
- periodic vs traffic-correlated vs constant = the routing question
basics
~20 sCheck four families: a blocking O(N) command or long Lua script (SLOWLOG); a fork for RDB/AOF rewrite (LATENCY fork, latest_fork_usec); memory pressure — swapping or eviction cycles at maxmemory; and bulk key expiry (LATENCY expire-cycle). Confirm the host floor with redis-cli --intrinsic-latency.
solid answer
~1 minI would separate causes that have a command to blame from ones that do not. **Command-attributable:** an O(N) command over a big collection (`KEYS`, `HGETALL`, `LRANGE 0 -1`, `SMEMBERS`, `ZUNIONSTORE`), a big `DEL` of a large aggregate, or a long Lua script/Function — all of which occupy the command-processing thread while every other client waits. Evidence: a SLOWLOG entry roughly the size of the spike, plus `INFO commandstats` for the cheap-but-frequent variant. **Not command-attributable:** - **Fork stalls** — RDB `bgsave` or AOF rewrite forks; the call's cost scales with page tables and is worsened by transparent huge pages. Evidence: `latest_fork_usec`, `rdb_bgsave_in_progress`, `LATENCY HISTORY fork`, periodicity matching `save` rules. - **Memory pressure** — swap means every access can hit disk (a fatal condition for Redis); at `maxmemory`, eviction cycles run on the write path. Evidence: OS swap usage for the process, `LATENCY eviction-cycle`, rising `evicted_keys`. - **Bulk expiry** — many keys sharing one TTL. Evidence: `LATENCY expire-cycle`, `expired_keys` bursts. - **AOF fsync** — `appendfsync always`, or `everysec` blocking behind a rewrite competing for disk. - **The host itself** — CPU steal, noisy neighbour. Evidence: `--intrinsic-latency` on the box. Periodicity is the fastest discriminator: regular spikes point to saves, cron jobs, or TTL cohorts.
code
text · 15 lines# periodic? correlate with saves
INFO persistence | grep -E 'rdb_bgsave_in_progress|aof_rewrite_in_progress|aof_delayed_fsync'
INFO stats | grep -E 'latest_fork_usec|expired_keys|evicted_keys'
# which subsystem stalled?
CONFIG SET latency-monitor-threshold 50
LATENCY LATEST
LATENCY HISTORY expire-cycle
# which command?
SLOWLOG GET 10
INFO commandstats | sort -t= -k3 -rn | head
# is the host itself the floor? (run ON the server)
redis-cli --intrinsic-latency 60go deeper
Name the obvious ones — an expensive command like KEYS on a big keyspace, and running out of memory — and say you would check SLOWLOG.
Separate blocking commands from background work, mention forks for saves, bulk expiry, and eviction, and cite the INFO fields you would read.
Drive an evidence-based triage: use periodicity to route, name the specific counters and LATENCY events per hypothesis, and give the mitigation for each (UNLINK, TTL jitter, THP off, snapshot on a replica, headroom).
Turn it into policy: memory headroom targets that make swap and eviction impossible, a banned-command list enforced by ACLs or review, shard sizing chosen partly for fork cost, and where the durability/latency line is drawn for appendfsync.
## Framing Redis processes commands on a single thread, so *any* pause becomes everyone's pause. Diagnosing spikes is therefore a matter of asking: was the thread busy doing work (a command), or was it prevented from running (the OS or a system call)? Each family has distinct evidence, and naming the evidence is what separates a senior answer from a list of guesses. ## Family 1 — a command occupied the thread **Big-O over big values.** `KEYS *` scans the whole keyspace; `HGETALL`/`SMEMBERS`/`LRANGE 0 -1` materialise an entire collection; `SORT`, `ZUNIONSTORE`, `SINTERSTORE` over large sets are worse. Ten milliseconds of CPU is a lifetime when the median command is 50 microseconds. **Freeing memory is work too.** `DEL` of a 5-million-field hash frees every element synchronously. `UNLINK` (and `lazyfree-lazy-*` settings, plus `FLUSHALL ASYNC`) hands that reclamation to a background thread instead. **Scripts.** Lua scripts and Functions execute atomically; the whole script is one blocking unit. A loop over a large set inside a script is indistinguishable from a slow command, except it may not even be obvious from the SLOWLOG argument list. **Big payloads and huge pipelines.** Multi-megabyte values cost in allocation, copying, and reply construction; a client that pipelines 100k commands hands the server one enormous unit of work and can also inflate the output buffer. Evidence: `SLOWLOG GET` shows an entry of the right magnitude; `INFO commandstats` (`usec_per_call`) shows the mildly-slow-but-constant case that SLOWLOG never records; `LATENCY HISTORY command`. ## Family 2 — fork stalls RDB snapshots and AOF rewrites are produced by forking the process and letting the child write while the parent keeps serving, relying on copy-on-write. The **`fork()` call itself** is not free: the kernel must copy page tables, and the cost grows roughly with the amount of memory mapped. On a 40GB instance it can be hundreds of milliseconds; with **transparent huge pages (THP)** enabled it is dramatically worse, and THP also inflates copy-on-write memory afterwards, which is why Redis logs a warning recommending it be disabled. Evidence: `INFO stats: latest_fork_usec`; `INFO persistence: rdb_bgsave_in_progress`, `aof_rewrite_in_progress`, `rdb_last_bgsave_time_sec`; `LATENCY HISTORY fork`; and periodicity that matches the `save` directives or an `auto-aof-rewrite-percentage` trigger. Mitigations: disable THP; shard so each instance holds less memory; take snapshots on a replica instead of the primary; tune or remove `save` points if RDB is not your durability story. ## Family 3 — memory pressure **Swapping is the worst case.** Redis assumes memory access is nanoseconds. If the OS pages part of the dataset to disk, a random access becomes a disk seek while the single thread waits, and latency goes from microseconds to tens of milliseconds unpredictably. Check the process's swap usage directly (on Linux, `/proc/<pid>/smaps`); `INFO memory` also reports `used_memory_rss` and `maxmemory_policy`. The fix is capacity, not tuning: keep the dataset comfortably below RAM, and give room for copy-on-write during saves and for replication buffers. **Eviction cycles.** Running at `maxmemory` means every write may first evict. With approximated LRU/LFU sampling and large values to free, that work lands on the request path. Evidence: `LATENCY eviction-cycle` / `eviction-del`, `evicted_keys` climbing in `INFO stats`. ## Family 4 — expiry cohorts Redis expires keys lazily (on access) and actively (a periodic sampling cycle). If a batch job writes two million keys with `EX 3600` in the same minute, an hour later the active cycle faces a mountain of deletions in a short window. Evidence: `LATENCY expire-cycle`, plus a spike in `expired_keys` at a regular interval one TTL after a known job. Fix: jitter TTLs with a random offset so the cohort dissolves. ## Family 5 — disk and durability `appendfsync always` makes writes wait for an fsync — correct if you need it, but it is a latency decision, not a free one. Even `everysec` can block when an AOF rewrite is saturating the same disk (`aof-write-active-child`, `aof-rewrite-diff-write` events, and `aof_delayed_fsync` in `INFO persistence`). Network-attached storage with variable latency amplifies all of this. ## Family 6 — the host CPU steal on a shared or burstable instance, a noisy neighbour, NUMA effects, or aggressive power management can pause the process for milliseconds with no Redis-side evidence at all. `redis-cli --intrinsic-latency 60` run on the server reports the floor. If the floor is already 5ms, stop tuning Redis. ## Family 7 — clients and connections Connection storms (no pooling, TLS handshakes per request), slow consumers filling client output buffers, or `MONITOR` left running in a terminal all show up as latency. `INFO clients` (`connected_clients`, `blocked_clients`, `client_recent_max_output_buffer`) and `CLIENT LIST` locate these. ## The discriminating move Ask first: **is it periodic?** Regular spikes strongly imply a save/rewrite fork, a TTL cohort, or a cron job. Irregular spikes correlated with traffic imply an expensive command or eviction. Constant elevation with no spikes implies the host floor or the network. That single question routes you to the right evidence in one step.
- Spikes occur every five minutes almost exactly. What is your first hypothesis?Something scheduled. The two usual suspects are a background save triggered by a `save` directive whose change-count threshold is being met on a regular cadence, and an application cron job issuing an expensive command or writing a large batch. Check `rdb_last_save_time` and `latest_fork_usec` against the spike timestamps, and look at `LATENCY HISTORY fork`; if fork is quiet, look for a SLOWLOG entry at those timestamps instead.
- Why is any swapping considered fatal for Redis rather than merely undesirable?Redis's data structures assume memory access costs nanoseconds and the single command-processing thread cannot proceed while a page fault is serviced. A swapped-out page turns a microsecond operation into a disk read of milliseconds, and because access is random the faults are unpredictable, so latency becomes wildly variable rather than uniformly worse. There is no configuration that fixes it — the answer is to fit the working set, plus headroom for copy-on-write and replication buffers, in RAM.
- How would you eliminate fork-induced spikes on a primary that must keep snapshots?Move snapshotting off the primary: let a replica perform the RDB save and AOF rewrite, and disable or loosen the primary's save points. Independently, disable transparent huge pages, which inflate both fork time and copy-on-write memory, and shard the dataset so each process maps less memory, since fork cost scales with page-table size. If snapshots must stay local, schedule them into low-traffic windows and monitor `latest_fork_usec` as a first-class metric.
One thread serving everyone is a single checkout lane: a customer with 400 items (big command), the cashier stepping away to photocopy the ledger (fork), or the floor being repaved under them (swap) all look the same to the queue — you need different evidence to tell them apart.
saying these in an interview costs you the question
- Reaching for 'add a replica' or 'more memory' before gathering any evidence
- Assuming an empty SLOWLOG exonerates Redis, ignoring fork, swap, and expiry causes
- Not knowing that transparent huge pages inflate fork time and copy-on-write memory
- Treating DEL of a huge collection as free, and never mentioning UNLINK or lazy-free
- Writing large key cohorts with identical TTLs and being surprised by periodic expiry stalls
- Believing a long Lua script is safe because it is 'server-side and therefore fast'