You own a fleet of Redis instances shared by a dozen services and need a latency observability strategy rather than ad-hoc debugging. What signals would you collect, what would you alert on, and where would you accept blind spots?
answer
- client percentiles = SLO; server signals = explanation
- SLOWLOG by last-seen id + raised max-len; label by CLIENT SETNAME
- LATENCY LATEST per event; fast-command = host canary
- alert on rates and headroom, never single samples
- blind spots: victims, attribution, lossy rings, MONITOR cost
basics
~20 sTreat client-side percentiles as the SLO signal and server-side facts as explanations: scrape SLOWLOG entries by id, LATENCY LATEST per event, INFO commandstats/latencystats, fork and eviction counters. Alert on client p99 and on rates of slow entries or stall events — not on single samples. Accept that queueing victims and per-tenant attribution stay blind spots.
solid answer
~60 s**SLO signal comes from clients.** Only the application measures what users feel — connection-pool wait, network, queueing behind other tenants. Server-side numbers explain, they do not define. **Collect on the server:** `SLOWLOG GET` polled with last-seen-id bookkeeping (raise `slowlog-max-len` so bursts survive between polls) and exported as events labelled by command and client name; `LATENCY LATEST` per event (fork, expire-cycle, eviction-cycle, aof-fsync-always, fast-command) with `latency-monitor-threshold` armed fleet-wide at ~100ms; `INFO` counters — `latest_fork_usec`, `evicted_keys`, `expired_keys`, `blocked_clients`, `mem_fragmentation_ratio`, `aof_delayed_fsync`; and Redis 7's `INFO latencystats` percentiles per command. Baseline the host floor with `--intrinsic-latency` at provisioning time and treat it as an instance-selection criterion. **Alert on rates and trends:** new slow entries per minute, fork time crossing a budget, eviction rate above zero when it should be zero, memory headroom shrinking. Page on client p99 breach; ticket on the explanatory signals. **Blind spots I accept:** victims of a blocking command are invisible server-side; per-tenant attribution needs `CLIENT SETNAME` discipline; the 160-sample latency history and the slow-log ring are lossy across restarts.
go deeper
Say that you would watch SLOWLOG and basic INFO metrics like memory and connected clients, and that the application should also time its own Redis calls.
Distinguish the tools — SLOWLOG for commands, LATENCY events for subsystems, INFO for counters — and describe scraping them into a dashboard with sensible thresholds.
Separate SLO signals from explanatory signals, define what pages versus what tickets, and describe attribution via client names and prevention via TTL jitter and headroom.
Own the whole system: governance (ACL-enforced command bans, standardised client wrappers), capacity policy that makes swap and eviction structurally impossible, snapshotting off the serving path, and an explicit list of accepted blind spots with the reason each is accepted.
## The core division: symptom versus explanation A fleet-level strategy fails when it alerts on server internals and hopes users are happy. The disciplined split is: - **Symptom signals** define the SLO and page someone. They come from **clients**: request latency percentiles for the Redis calls the application makes, plus timeout and error rates, plus connection-pool wait time. Only the client sees pool exhaustion, TLS handshakes, DNS, network, and — crucially — time spent queued behind another tenant's expensive command. - **Explanation signals** are server-side and are what you look at once paged. They should be rich, cheap, and never page on their own unless they are a leading indicator of a hard failure. ## What to collect from each instance **Slow commands.** Poll `SLOWLOG GET` on a short interval. Track the highest entry id already exported so each entry is emitted exactly once; entries carry id, timestamp, microseconds, truncated arguments, client address and client name. Raise `slowlog-max-len` well above the 128 default so a burst is not lost between polls, and set `slowlog-log-slower-than` near the service's latency budget rather than the 10ms default. Export as events with `command` and `client_name` labels — that is what makes multi-tenant attribution possible. **Subsystem stalls.** Arm `latency-monitor-threshold` (milliseconds) fleet-wide — 100 is a reasonable standing value — and scrape `LATENCY LATEST`, which returns per event the latest timestamp, latest value, and all-time max. The events that matter operationally: `fork`, `expire-cycle`, `eviction-cycle`/`eviction-del`, `aof-fsync-always` and the AOF write/rewrite family, `command`, and `fast-command`. `fast-command` firing with everything else quiet is the fleet's canary for host-level stalls. **Aggregate command cost.** `INFO commandstats` gives calls, total microseconds, and microseconds per call per command — this catches the case SLOWLOG structurally cannot see: a command that is only mildly expensive but is issued a million times a second. Redis 7 adds `INFO latencystats` with per-command percentiles, which is strictly better for this purpose. **Capacity and background work.** From `INFO`: `used_memory` versus `maxmemory` (headroom is the single most predictive number), `used_memory_rss` and fragmentation ratio, `evicted_keys`, `expired_keys`, `latest_fork_usec`, `rdb_bgsave_in_progress`, `aof_rewrite_in_progress`, `aof_delayed_fsync`, `connected_clients`, `blocked_clients`, `client_recent_max_output_buffer`, `total_net_output_bytes`, `keyspace_hits`/`misses`, and replication offsets/lag. **Host facts.** Swap in use by the process (any value is an incident), CPU steal, and the intrinsic-latency floor measured at provisioning time. If a candidate instance type shows a 3ms floor, that is a procurement decision, not a tuning problem. ## What to alert on Alert on **rates and budgets**, never on single samples: - **Page:** client p99 (or the SLO percentile) breached for a sustained window; error/timeout rate; instance unreachable; replica lag beyond the failover budget; memory headroom below the copy-on-write reserve. - **Ticket / warn:** new slow-log entries per minute above a baseline; `fork` latency above a chosen budget; `evicted_keys` above zero on an instance that is supposed to be sized as a store rather than a cache; `aof_delayed_fsync` climbing; fragmentation ratio drifting; a single client name responsible for a disproportionate share of slow entries. - **Never page on:** one 200ms `LATENCY` sample, or one slow-log entry. They are noisy and self-resolving; the rate is the signal. ## Governance, not just telemetry On a shared fleet, the cheapest latency win is preventing the pathology: - **Ban the O(N)-over-the-keyspace commands** (`KEYS`, `FLUSHALL` in production paths, unbounded `HGETALL`/`SMEMBERS`) using ACL command rules per tenant, so the guardrail is enforced by the server rather than by review. - **Require `CLIENT SETNAME`** in every client library wrapper. Without it, slow-log entries name an IP from an autoscaling group and attribution is guesswork. - **Standardise TTL jitter** in the shared cache client so no team can create an expiry cohort. - **Set a memory headroom target** (commonly leaving room for copy-on-write plus replication buffers) so that swap and eviction cycles are structurally impossible rather than monitored. - **Decide where snapshots run** — typically on a replica — so fork stalls never touch the serving path. ## Blind spots to accept explicitly 1. **Victims are invisible.** A blocking command's queueing cost appears only in client metrics. No server-side signal will ever show it, which is precisely why client percentiles are the SLO. 2. **Attribution is best-effort.** Without client names, and with connection pooling and proxies, mapping a slow command to a team is approximate. 3. **Lossy buffers.** The slow-log ring and the 160-sample-per-event latency history are in-memory and lost on restart; polling frequency bounds what you can recover. 4. **Cost of precision.** `MONITOR` shows everything and costs real throughput; keyspace-level analysis needs `--bigkeys`/`--memkeys` or offline RDB analysis, which are sampling or point-in-time, not continuous. 5. **Cardinality.** Exporting per-key or per-argument labels will destroy your metrics system; keep labels at command and client-name granularity. Stating those limits, and choosing them deliberately, is what makes this a strategy rather than a dashboard.
- Why not simply alert on server-side latency events and skip client instrumentation?Because the server cannot see the cost it imposes on waiting clients. Slow-log durations exclude queueing, so a single blocking command that ruins p99 for every tenant produces exactly one server-side entry. Client-side percentiles also include pool waits, TLS, and network — layers that are frequently the real cause. Server signals explain an incident; client signals decide whether there is one.
- How do you attribute a slow command to a team on a shared fleet?Enforce `CLIENT SETNAME` in the shared client wrapper so every connection is tagged with service and instance; Redis 4.0+ records client address and name in each slow-log entry, and `CLIENT LIST` exposes them live. Export slow-log events with the client name as a label and keep argument text at the command/key-prefix level to control cardinality. Pair it with per-tenant ACL users, which additionally lets you deny expensive commands rather than merely observe them.
- What memory policy makes latency predictable on a shared instance?Keep committed usage well below the ceiling so that neither swapping nor continuous eviction cycles ever occur on the request path, and reserve extra headroom for copy-on-write during saves and for replication and client output buffers. Alert on shrinking headroom as a leading indicator rather than on eviction as a lagging one. Shard when an instance approaches the target instead of raising maxmemory, because fork cost and failover time also grow with instance size.
saying these in an interview costs you the question
- Building the entire SLO on server-side INFO metrics and never instrumenting clients
- Paging on single latency samples, producing alert fatigue and no diagnosis
- Leaving MONITOR running as a permanent tap, which measurably costs throughput
- Exporting per-key labels and destroying the metrics backend with cardinality
- Treating eviction as normal operation on an instance the business uses as a store
- No client naming, so every slow entry maps to an anonymous IP in an autoscaling group