skip to content

You must remove every key matching the prefix 'session:' from a live Redis instance holding tens of millions of keys, without causing a latency incident. How do you do it, and what makes the naive approach dangerous?

level: seniorimportance: should knowfreq 40%

answer

  1. SCAN + MATCH + COUNT 500, loop to cursor 0
  2. UNLINK (async free) not DEL (inline free)
  3. pipeline batches, sleep between them
  4. MATCH filters late → cost is total keys
  5. per-master in Cluster; TTL/own DB is the real fix

basics

~20 s

Loop SCAN with MATCH session:* and a moderate COUNT, deleting each batch with a pipelined UNLINK (async free) and pausing between batches. The naive KEYS session:* piped into DEL blocks the server twice: once for the O(N) enumeration and again for one huge delete.

solid answer

~60 s

Iterate, don't enumerate. Run a `SCAN cursor MATCH session:* COUNT 500` loop, and for each batch issue a pipelined `UNLINK` of those keys, sleeping a few milliseconds between batches so production traffic interleaves. `redis-cli --scan --pattern 'session:*' | xargs -L 500 redis-cli UNLINK` is the one-liner version. Why `UNLINK` rather than `DEL`: `DEL` frees the value inline, so deleting large collections costs real time on the execution thread; `UNLINK` unlinks the key immediately and hands the memory reclamation to a background thread. The naive `KEYS session:* | xargs redis-cli DEL` blocks twice — the O(N) enumeration and then one enormous multi-key delete — and materialises a huge key list in server memory. Caveats: `MATCH` filters after retrieval, so the scan costs O(total keys) no matter how few match; duplicates are harmless here because deletion is idempotent; in Cluster you must run the loop per master node. Better still, design so this is never needed — put disposable data in its own logical database (`FLUSHDB ASYNC`), or give keys TTLs.

code

text · 13 lines
text
# Quick one-off (no pacing)
redis-cli --scan --pattern 'session:*' | xargs -L 500 redis-cli UNLINK

# Paced loop (pseudocode)
cursor = "0"
do {
  (cursor, batch) = SCAN cursor MATCH "session:*" COUNT 500
  if batch: pipeline { UNLINK batch }   # idempotent; duplicates are fine
  sleep(5ms)                            # yield to production traffic
} while (cursor != "0")

# NEVER on a large instance:
# redis-cli KEYS 'session:*' | xargs redis-cli DEL

go deeper

for a junior

Give the safe recipe: loop SCAN with MATCH, delete each batch, never use KEYS on a big production instance.

for a middle

Add why the naive version blocks twice, why UNLINK beats DEL, and that the loop ends only at cursor 0 while MATCH does not reduce traversal cost.

for a senior

Cover pacing and pipelining, the replication-stream impact of a mass delete, keys created mid-sweep being missed, per-master sweeps in Cluster, and testing the pattern before running it.

for a principal

Argue the sweep should not exist: mandate TTLs or a dedicated logical database/instance for disposable data, maintain index sets where enumeration is a real requirement, and enforce guardrails (ACL/rename-command on KEYS and FLUSH*) so the blocking recipe is unavailable in production.

## Why the obvious command is the wrong one The one-liner everybody reaches for is `redis-cli KEYS 'session:*' | xargs redis-cli DEL`. It fails in three separate ways on a large instance. 1. **The enumeration blocks.** `KEYS` is O(total keys in the database) and Redis runs commands one at a time, so the instance stops serving anyone for the duration — seconds, on tens of millions of keys. 2. **The reply is enormous.** Every matching key name is assembled into a single server-side reply buffer before transmission; millions of names is tens of megabytes allocated inside a process that may already be near `maxmemory`. 3. **The delete blocks.** `DEL k1 k2 ... kN` is one command; it frees every value inline on the execution thread. If the values are large collections, freeing them can dominate — deallocating a ten-million-element hash is not instantaneous. ## The safe shape Iterate in bounded slices and delete asynchronously: - `SCAN cursor MATCH session:* COUNT 500` until the cursor comes back `0`. - Delete each batch with a **pipelined `UNLINK`** (available since Redis 4.0). `UNLINK` removes the key from the keyspace immediately and queues the value's memory reclamation onto a background thread, so a huge value never stalls command processing. `DEL` frees inline. - **Pace the loop.** Sleep a few milliseconds between batches, or run the job at low traffic. Pacing is what keeps a bulk operation from consuming the entire command budget of the instance; it is far better than shrinking `COUNT` to 1, which only multiplies round trips. - **Pipeline** the deletes so you are not paying one round trip per key. `redis-cli --scan --pattern 'session:*' | xargs -L 500 redis-cli UNLINK` implements this adequately for one-off work; a small script gives you pacing, progress logging and resumability. ## Properties of SCAN that matter here - **`MATCH` filters after retrieval**, so the traversal cost is proportional to the *total* number of keys, not to the matching ones. Deleting 1% of a 50M-key database still means walking 50M keys. Budget for that; it is a background job, not a click. - **Duplicates are harmless.** SCAN may return the same key twice; `UNLINK` on an already-deleted key is a no-op. Because deletion is idempotent you need no client-side dedupe — a rare case where SCAN's weak guarantee costs nothing. - **New keys may be missed.** SCAN promises only that keys present for the whole iteration are seen at least once. If the application is still creating `session:*` keys, the scan will not catch the ones created after it passed. Either stop the writer first, or run a second sweep, or accept that the remaining keys will expire. - **Keys may vanish under you.** Handle the ordinary case where a key returned by SCAN is already gone; no error handling beyond "ignore" is needed for deletes. ## Cluster and replicas In Redis Cluster the keyspace is split across masters, and a SCAN cursor is meaningful only for the node that issued it. `redis-cli --scan` talks to one node; a cluster-wide sweep means running the loop against **every master** (for example via `redis-cli --cluster call` or by iterating the node list). Deletes must go to the owning node, which a cluster-aware client handles for you. Also remember that every `UNLINK` propagates to replicas, so a huge sweep produces a correspondingly huge replication stream — another reason to pace it and to watch replica buffers. ## Cheaper alternatives, in rough order of preference 1. **TTLs.** If the data is inherently disposable, it should have had an expiry from the start; then there is no cleanup job at all. 2. **A dedicated logical database or a dedicated instance.** Isolate disposable data so it can be removed with `FLUSHDB ASYNC` / `FLUSHALL ASYNC`, which is one command and reclaims memory in the background instead of an O(total keys) traversal. 3. **A maintained index.** `SADD session:index <key>` on creation turns "find all session keys" into a set read plus batched deletes, with no keyspace traversal. 4. **Only then**, a SCAN sweep — the recovery tool for data that was designed without any of the above. ## Guardrails worth mentioning Disable `KEYS`, `FLUSHALL` and `FLUSHDB` on production instances via ACL rules or `rename-command`, so nobody performs step 1 of the naive recipe by reflex. And test the sweep on a replica or a staging copy first: a pattern typo like `session*` versus `session:*` deletes a different set of keys, and there is no undo. ## The answer in one breath SCAN with MATCH and a moderate COUNT, pipelined UNLINK per batch, paced with sleeps, per master node in Cluster; naive KEYS+DEL blocks on both the enumeration and the free, and the real fix is TTLs or a separate database so the sweep is never needed.

  • Why prefer UNLINK over DEL for this sweep?
    DEL frees each value synchronously on the execution thread, so deleting keys that hold large collections costs proportional time right there and adds latency for every other client. UNLINK removes the key from the keyspace immediately and hands the memory reclamation to a background thread, so the command returns fast regardless of value size. For tiny string values the difference is negligible, but a sweep rarely knows the sizes in advance.
  • The application is still creating session keys while your sweep runs. Will the sweep remove all of them?
    No. SCAN only guarantees that keys present for the entire iteration are returned at least once; keys created after the cursor has passed their bucket may never be seen. Either stop or reconfigure the writer first, run a second sweep afterwards, or rely on TTLs to clear the stragglers. Treat the sweep as best-effort unless the writer is quiesced.
  • How does this change under Redis Cluster?
    The keyspace is partitioned across masters and a SCAN cursor is only valid for the node that produced it, so you must run an independent scan loop against every master rather than a single cluster-wide scan. Deletes are routed by key to the owning shard, which a cluster-aware client does automatically. Pace each node's sweep separately, since each has its own replication stream to feed.

saying these in an interview costs you the question

  • "KEYS pattern | xargs DEL is fine if I run it at night" — it still blocks the instance and can trip failover detection.
  • "MATCH means I only pay for the matching keys" — the traversal always covers the whole keyspace.
  • "I must dedupe SCAN results before deleting" — deletion is idempotent, so duplicates cost nothing here.
  • "DEL and UNLINK are the same, UNLINK is just newer" — UNLINK reclaims memory in a background thread, DEL frees inline.
  • "One SCAN loop covers a whole cluster" — cursors are per node; you must sweep every master.

context