skip to content

Redis includes a built-in latency monitoring framework exposed through the LATENCY family of commands. How does it work, how do you turn it on, and what does it show you that a log of slow commands cannot?

level: seniorimportance: should knowfreq 32%

answer

  1. latency-monitor-threshold in MILLISECONDS, 0 = off
  2. events: fork, expire-cycle, eviction-cycle, aof-fsync-always, command, fast-command
  3. LATEST = name/time/latest/max; HISTORY = 160 samples
  4. DOCTOR = prose hypothesis; GRAPH = ASCII plot
  5. fast-command slow with nothing else → host stall (swap/steal/THP)

basics

~20 s

Set latency-monitor-threshold to a millisecond value; Redis then records time-series samples for named event types (fork, expire-cycle, command, aof-fsync-always, eviction-cycle...). LATENCY LATEST shows the newest and worst per event, LATENCY HISTORY <event> the samples, LATENCY RESET clears, LATENCY DOCTOR explains them in prose.

solid answer

~50 s

The latency monitor is off by default: `CONFIG SET latency-monitor-threshold 100` arms it at 100ms. Redis then instruments internal code paths and records a `(timestamp, milliseconds)` sample whenever a named **event** exceeds that threshold, keeping the last 160 samples per event. The event names are the point. `command` and `fast-command` cover command execution, but the rest cover things no command log can attribute: `fork` (the copy-on-write fork for RDB/AOF rewrite), `expire-cycle` (the active expiry pass), `eviction-cycle` / `eviction-del` (maxmemory work), `aof-write`, `aof-fsync-always`, `aof-rewrite-diff-write`, `rdb-unlink-temp-file`. `LATENCY LATEST` returns per event: name, timestamp of the latest sample, latest value, all-time max. `LATENCY HISTORY <event>` gives the sample series for graphing. `LATENCY RESET [event]` clears. `LATENCY DOCTOR` reads all of it and emits a human-readable report with likely causes and advice, and `LATENCY GRAPH <event>` draws an ASCII plot. So: SLOWLOG says *which command*; the latency monitor says *which subsystem*, including ones with no command attached.

code

text · 16 lines
text
CONFIG SET latency-monitor-threshold 100   # milliseconds; 0 disables

LATENCY LATEST
1) 1) "fork"
   2) (integer) 1723642155   # when the latest sample happened
   3) (integer) 412          # latest: 412 ms
   4) (integer) 980          # all-time max
2) 1) "expire-cycle"
   2) (integer) 1723642101
   3) (integer) 143
   4) (integer) 143

LATENCY HISTORY fork          # up to 160 (timestamp, ms) samples
LATENCY GRAPH fork            # ASCII plot
LATENCY DOCTOR                # prose report with likely causes
LATENCY RESET fork

go deeper

for a junior

Know it exists, that you enable it with latency-monitor-threshold in milliseconds, and that LATENCY DOCTOR prints a readable report.

for a middle

Name the main event types and the LATEST/HISTORY/RESET/DOCTOR commands, and explain that it covers subsystems with no command to blame.

for a senior

Drive a diagnosis from event names to causes — fork versus expire-cycle versus eviction-cycle versus fast-command-implies-host — and describe scraping LATENCY LATEST into monitoring.

for a principal

Position it in an observability strategy: which events become alerts, what the mitigations cost (sharding to cut fork time, TTL jitter, memory headroom to avoid eviction cycles), and the tradeoff against durability settings like appendfsync always.

## The gap this fills SLOWLOG answers "which command took too long". A large class of Redis latency has **no command to blame**: the process forked to write an RDB and the fork call itself stalled for 400ms; the active expiry cycle ran long because millions of keys expired at once; `appendfsync always` made every write wait on the disk; evicting under `maxmemory` freed a huge hash. During those windows every client is slow and SLOWLOG can be empty. The latency monitoring framework instruments those internal code paths by name and samples them. ## Turning it on ``` CONFIG SET latency-monitor-threshold 100 ``` The unit is **milliseconds**, and `0` (the default) disables it. Only events that exceed the threshold are recorded, so the runtime cost is negligible — a clock read around instrumented sections. A common production setting is 100ms for general safety, lowered to 10–25ms while actively investigating. Persist it in `redis.conf` if you want it after restart. ## Event types worth knowing - **`command`** — a normal command whose execution exceeded the threshold. Overlaps with SLOWLOG, and is the bridge between the two tools. - **`fast-command`** — an O(1)/O(log N) command that still went slow, which is a strong signal of a process-level stall (swap, CPU steal) rather than an expensive operation. - **`fork`** — how long the `fork()` system call itself took. Proportional to the process's page-table size, so it grows with dataset size; on some virtualised hypervisors it is dramatically worse. Also visible as `latest_fork_usec` in `INFO stats`. - **`expire-cycle`** — the active expiration pass. Long values mean a large population of keys expiring together, typically because many keys were written with the same TTL at the same moment. - **`eviction-cycle`** / **`eviction-del`** — time spent selecting and deleting keys under `maxmemory`. Long values mean you are running at the memory ceiling and paying for it on every write. - **`aof-write`**, **`aof-write-active-child`**, **`aof-write-alone`**, **`aof-fsync-always`**, **`aof-rewrite-diff-write`** — the append-only-file paths. `aof-fsync-always` appearing means the durability policy is costing you latency per write; the rewrite-related ones mean a rewrite is competing for disk. - **`rdb-unlink-temp-file`** — deleting a large temp RDB, which can stall on some filesystems. ## The commands - `LATENCY LATEST` — array of `[event, unix-timestamp-of-latest, latest-ms, all-time-max-ms]`. This is the one to scrape into monitoring: it is compact and includes the all-time worst. - `LATENCY HISTORY <event>` — up to 160 `(timestamp, ms)` samples for one event. Good for graphing and for spotting periodicity. - `LATENCY RESET [event ...]` — clears samples (returns how many event time series were reset). Use before a controlled test. - `LATENCY DOCTOR` — analyses everything recorded and returns prose: which events fired, how bad, and specific advice ("check your fork time, consider disabling transparent huge pages", "many keys expiring at once"). It is genuinely useful as a first read and as a teaching tool, but it is heuristic — treat it as a hypothesis generator, not a verdict. - `LATENCY GRAPH <event>` — an ASCII latency plot over time for that event, handy over SSH. ## Reading the results A practical loop: 1. Arm it at 100ms fleet-wide, scrape `LATENCY LATEST` into metrics with the event name as a label. 2. When client p99 spikes, look at which event has samples in that window. 3. `fork` spikes → correlate with `rdb_bgsave_in_progress` / `aof_rewrite_in_progress`; mitigate by disabling transparent huge pages, choosing hosts with faster fork behaviour, shrinking the dataset per instance (shard), or moving saves to a replica. 4. `expire-cycle` spikes → stop writing millions of keys with identical TTLs; jitter TTLs by a random spread. 5. `eviction-cycle` spikes → you are at `maxmemory`; add capacity or reduce the working set. Also check `evicted_keys` in `INFO stats`. 6. `aof-fsync-always` → the durability policy is the cost; decide whether `everysec` is acceptable. 7. `fast-command` spikes with nothing else → suspect the host: check swap (`INFO memory` and the process's swap usage), CPU steal, and run `redis-cli --intrinsic-latency` on the box. ## Limits Samples are in memory, capped at 160 per event, and lost on restart — so scrape them. The threshold is global, not per event. And the monitor tells you a subsystem was slow, not why the host made it slow; pairing it with OS-level facts (swap in use, transparent huge pages enabled, hypervisor steal time) is what closes the case.

  • `LATENCY LATEST` shows repeated `expire-cycle` samples above 150ms. What is the likely cause and the fix?
    A large cohort of keys is expiring at nearly the same instant, usually because a batch job wrote them all with an identical TTL. The active expiry cycle then has an enormous amount of work in one pass and holds the command-processing thread. The fix is to jitter TTLs — add a random spread of a few percent to each key's expiry — so deletion work is spread over time, and to prefer smaller, more frequent write batches.
  • You see `fast-command` events firing but `command` events are quiet and SLOWLOG is empty. What does that suggest?
    O(1)-class commands should never be slow on their own, so the delay is coming from outside the command: the process was not running. Suspect host-level stalls — memory swapped out (check the process's swap usage and `INFO memory`), CPU steal on a shared or burstable instance, transparent huge page compaction, or an ongoing fork. Confirm with `redis-cli --intrinsic-latency` on the server; if the floor is already milliseconds, the fix is host or hypervisor level, not Redis config.
  • How does LATENCY DOCTOR differ from the raw LATENCY output, and how much should you trust it?
    DOCTOR reads the recorded event series and returns a human-readable narrative naming the worst events, their magnitudes, and rule-of-thumb advice such as checking transparent huge pages when fork time is high. It is a hypothesis generator built on heuristics, not a measurement. Use it to orient quickly, then confirm with the raw HISTORY samples, INFO counters, and OS-level evidence before acting.

saying these in an interview costs you the question

  • Assuming the latency monitor is on by default — it is disabled until `latency-monitor-threshold` is set
  • Giving the threshold in microseconds by analogy with slowlog-log-slower-than; it is milliseconds
  • Treating LATENCY DOCTOR output as a definitive diagnosis rather than a heuristic hypothesis
  • Believing SLOWLOG would have caught a fork stall or an expire-cycle spike
  • Forgetting samples are in-memory and capped at 160 per event, so nothing is retained across a restart

context