skip to content

Before blaming Redis for a slow endpoint, how would you measure an individual Redis server's own response latency and its floor on that machine, using tools shipped with Redis?

level: juniorimportance: should knowfreq 40%

answer

  1. --latency = PING round trip, milliseconds
  2. --latency-history = one line per interval, spot periodicity
  3. --intrinsic-latency = run ON the server, no commands, OS floor
  4. INFO commandstats → usec_per_call; Redis 7 latencystats → percentiles
  5. measure from an app-like host, not from the Redis box

basics

~20 s

Run redis-cli --latency (or --latency-history) from a client host: it sends PING in a loop and reports min/avg/max round-trip in milliseconds. Run redis-cli --intrinsic-latency <seconds> on the server itself to measure the machine's own scheduling floor with no network involved.

solid answer

~50 s

Two different measurements, and mixing them up is the classic mistake. `redis-cli --latency -h <host>` loops `PING` and prints running min/avg/max/samples in milliseconds — that is **round-trip latency as a client sees it**, so it includes the network, the kernel, and any queueing behind other commands. `--latency-history` prints one 15-second summary line per interval so you can see drift; `--latency-dist` draws a coloured spectrum of the distribution. `redis-cli --intrinsic-latency 100` must be run **on the Redis host itself**. It runs no Redis commands at all: it busy-loops measuring how long the kernel takes to give the CPU back. That is the machine's floor — noisy neighbours, a throttled cloud vCPU, or transparent huge pages will show up here as milliseconds of intrinsic latency, and no Redis tuning can go below it. Compare the two: if round-trip is 5ms and intrinsic is 4ms, the host is the problem, not Redis. Then add `INFO commandstats` to see which command is burning time on average.

code

text · 14 lines
text
# 1) round-trip as a client sees it — run from an APP host
redis-cli -h redis-prod-1 --latency
min: 0, max: 41, avg: 0.31 (5122 samples)

# 2) is it periodic? one summary line per 15s
redis-cli -h redis-prod-1 --latency-history -i 5

# 3) the machine's own floor — run ON the redis host
redis-cli --intrinsic-latency 60
Max latency so far: 1200 microseconds.

# 4) which command burns time on average?
redis-cli INFO commandstats | head
cmdstat_hgetall:calls=91221,usec=8113442,usec_per_call=88.94

go deeper

for a junior

Name the two commands and what each measures: --latency for round trip in milliseconds, --intrinsic-latency for the machine's floor, run on the server.

for a middle

Add --latency-history for periodicity, the fact that PING round trip includes queueing behind other commands, and INFO commandstats for average cost per command.

for a senior

Present an ordered triage that exonerates layers — client metrics, round trip, host floor, SLOWLOG, commandstats — and interpret low-avg/high-max as a spike problem tied to fork or blocking commands.

for a principal

Argue for continuous measurement rather than ad hoc: client-side percentiles as the SLO signal, host intrinsic latency as an instance-selection criterion, and a policy about which commands may run at all.

## Why measure before blaming "Redis is slow" is almost always reported from the application's stopwatch, which contains connection pool waits, client-library serialization, TLS, network, and the server. Ruling layers out in order is faster than guessing. Redis ships tools for exactly this and they need no agent, no APM, and no restart. ## Round-trip latency: `redis-cli --latency` ``` redis-cli -h redis-prod-1 --latency min: 0, max: 6, avg: 0.28 (1284 samples) ``` It sends `PING` as fast as it can and reports **min / max / avg round-trip in milliseconds** with a sample count, updating live. `PING` costs the server essentially nothing, so whatever you see is the cost of *everything except the work*: client scheduling, the network path, the server's event loop availability, and time queued behind other clients' commands. Interpretation: - avg well under 1ms with a low max — the path is healthy; the slowness is in your application or in specific expensive commands. - low avg but a large max — **spikes**, not a slow baseline. That points at blocking commands, fork during a save, or an expire burst. Switch to `--latency-history` (prints a summary line every 15 seconds by default; `-i` changes the interval) to see whether spikes are periodic, which is a strong hint of a background save or a cron-driven job. - high avg overall — either the network path (run it from a second host to compare) or the machine itself. Run it **from a machine that resembles your application hosts**. Running it on the Redis box hides all network cost; running it from a laptop over VPN measures your VPN. `redis-cli --latency-dist` renders the distribution as a colour spectrum, which makes a bimodal "mostly 0.2ms, sometimes 40ms" pattern obvious in a way an average never will. ## The floor: `redis-cli --intrinsic-latency` ``` redis-cli --intrinsic-latency 100 Max latency so far: 3 microseconds. ... Max latency so far: 1200 microseconds. ``` This mode is different in kind. It connects to nothing and executes no Redis command: it runs a tight loop for the given number of seconds, repeatedly reading the clock, and reports the largest gap it ever observed between two consecutive reads. That gap is the time the operating system took the CPU away — scheduler preemption, hypervisor steal time, page faults, transparent-huge-page compaction, SMIs. Because it measures the host, it **must be run on the Redis server itself** (it is safe: it burns one core for the duration, so run it briefly and be aware of that). The result is the hard floor: a single-threaded process cannot respond faster than the OS lets it run. If intrinsic latency is 4ms, then a p99 of 5ms from clients is not a Redis problem at all — it is a noisy or oversubscribed host, and the fix is a different instance type, CPU pinning, disabling transparent huge pages, or moving off a burstable/throttled vCPU. ## Filling in the middle: `INFO commandstats` and `INFO latencystats` `INFO commandstats` gives, per command: `calls`, `usec` (total microseconds), `usec_per_call`, plus rejected/failed counts. This answers "is one command mildly slow a million times a second?" — a case SLOWLOG never shows because no individual call crosses the threshold. Redis 7.0 adds `INFO latencystats`, which reports per-command latency **percentiles** from a histogram, so you can see p50/p99 per command rather than only an average. `CONFIG RESETSTAT` zeroes these counters so you can measure a clean window. ## A workable order of operations 1. Confirm the symptom from the client's own metrics (p99, not average). 2. `redis-cli --latency` from an app-like host — is the baseline itself bad, or is it spiky? 3. `redis-cli --intrinsic-latency 60` on the server — is the machine's floor already at that level? 4. `SLOWLOG GET` — is a single blocking command causing the spikes? 5. `INFO commandstats` / `latencystats` — is a frequent command mildly expensive? 6. `INFO clients` (`blocked_clients`, `client_recent_max_input_buffer`) and `INFO stats` — is the problem connection churn, a huge pipeline, or output-buffer pressure? Each step either exonerates a layer or names the offender, which is exactly what an interviewer wants to hear instead of "I'd add more memory".

  • `--latency` reports avg 0.3ms but max 40ms. What does that pattern tell you, and what do you check next?
    A low average with a high maximum means the baseline path is healthy and you are chasing spikes, not slowness. Next, run `--latency-history` to see whether the spikes are periodic — regular spikes suggest a background RDB save fork, an AOF rewrite, or a scheduled job issuing an expensive command. Then check SLOWLOG for a blocking command in the same window, and INFO for `latest_fork_usec` and `rdb_bgsave_in_progress`.
  • Why must `--intrinsic-latency` run on the Redis server rather than from a client?
    It never talks to Redis. It busy-loops reading the clock and reports the largest gap between consecutive reads, which measures how long the operating system on *that* machine denied the process the CPU. Run from elsewhere it would measure the wrong host. The number is a floor: Redis on that box cannot respond faster than the scheduler allows, so a high intrinsic latency reframes the problem as host or hypervisor tuning rather than Redis tuning.

--latency times the whole delivery route; --intrinsic-latency measures how often the driver is forced to stop by traffic lights on that street. If the lights already cost four minutes, no faster van will help.

saying these in an interview costs you the question

  • Running `redis-cli --latency` on the Redis host itself and concluding the network is fine
  • Thinking `--intrinsic-latency` measures Redis command performance rather than the OS scheduling floor
  • Reporting averages only, hiding a bimodal distribution where p99 is the actual complaint
  • Jumping straight to 'add more memory' or 'add a replica' with no measurement
  • Ignoring INFO commandstats and so missing a cheap command that is called millions of times

context