skip to content

What do the MATCH and COUNT options of the Redis SCAN command actually control, and why can one SCAN call return zero keys even though the iteration is not finished?

level: middleimportance: should knowfreq 46%

answer

  1. COUNT = buckets hint, default 10
  2. may return more, fewer, or zero than COUNT
  3. MATCH filters after retrieval → saves bytes not work
  4. TYPE option since 6.0
  5. empty batch normal; stop at cursor 0

basics

~20 s

COUNT is a hint for how much work (how many buckets) one call does — default 10 — not how many keys come back. MATCH filters the retrieved names afterwards, so it saves network, not scanning. A call can therefore return an empty batch; only cursor 0 means done.

solid answer

~50 s

**COUNT** is a *hint*, not a limit: it tells Redis roughly how many hash-table buckets to visit in this call. The default is 10. You may get back more elements than COUNT (one bucket can hold many, and small collections are returned whole) or far fewer — including none. **MATCH** applies a glob to the element names **after** they have been retrieved from the table, on the server. It never reduces the traversal cost; a full iteration with `MATCH user:*` still visits every key in the database. What it buys you is a much smaller reply and no client-side filtering. Put together, this is why empty batches are normal: SCAN visited its ten buckets, nothing matched, and it returned so the command stays short. Terminate only on cursor `0`. Tuning COUNT is a latency/round-trip trade: 1000 finishes an iteration in far fewer round trips but makes each command do more work in one go; a few hundred to a thousand is a common production choice.

code

text · 15 lines
text
# Default COUNT 10 with a selective pattern: mostly empty batches
> SCAN 0 MATCH 'order:2026:*'
1) "384"
2) (empty array)

# Bigger work budget per call: far fewer round trips
> SCAN 0 MATCH 'order:2026:*' COUNT 1000
1) "7168"
2) 1) "order:2026:118"
   2) "order:2026:119"

# Filter by value type (Redis 6.0+)
> SCAN 0 TYPE stream COUNT 100
1) "512"
2) 1) "events:audit"

go deeper

for a junior

Know the two options by name: COUNT is roughly how much work per call (default 10), MATCH filters the names; keep looping until the cursor is 0.

for a middle

Explain that COUNT is a hint over buckets, not a result count, that MATCH is applied after retrieval so traversal cost is unchanged, and why empty batches happen.

for a senior

Reason about the tuning trade — round trips versus per-command occupancy — and throttle with sleeps rather than tiny COUNTs; know the TYPE filter and that COUNT is ignored for small compact-encoded collections.

for a principal

Set an organisational default for maintenance scans (COUNT range, pacing, which replica or node they run on, how they yield to production traffic) and prefer maintained indexes over recurring scans in the design itself.

## COUNT: a work budget, not a page size Every `SCAN`, `HSCAN`, `SSCAN` and `ZSCAN` call accepts `COUNT n`, defaulting to **10**. The number describes *how much of the underlying table this call should traverse* — roughly how many hash-table buckets are examined — not how many elements are returned. Three consequences follow: - **You can get more than COUNT.** Buckets contain chains, so one bucket may yield many elements. When Redis is mid-rehash, the corresponding bucket is scanned in both tables. - **You can get fewer than COUNT, including zero**, when the visited buckets are empty or nothing matched the `MATCH` pattern. - **COUNT is not remembered.** Nothing forces you to keep it constant across calls of the same iteration; you may raise it as you go. ## MATCH: filtering happens after retrieval `MATCH pattern` applies the same glob-style matching as `KEYS` (`*`, `?`, `[...]`) to each element the traversal has already produced, and drops the non-matching ones before building the reply. This is a **server-side filter, not an index seek**. Redis has no prefix index over the keyspace, so a full iteration with `MATCH session:*` still touches every key in the database — the pattern only shrinks the reply. What that saves is real but limited: network bytes, reply memory, and client CPU. Because the filter runs after retrieval, a call with a selective pattern will very often return an empty array while the cursor keeps advancing. Client code that treats an empty batch as "finished" will silently process a fraction of the data — this is the single most common SCAN bug, and it is easy to miss in testing because with a tiny dataset the first call usually returns everything. ## The TYPE option Since Redis 6.0, keyspace `SCAN` also accepts `TYPE <type>` (`string`, `list`, `set`, `zset`, `hash`, `stream`), filtering by value type. Like `MATCH` it is applied after retrieval and does not make the traversal cheaper — but it removes a round trip per key compared to calling `TYPE` yourself, and it is the clean way to find, say, every stream in a database. ## Tuning COUNT: the actual trade A full iteration is O(N) whatever COUNT you pick; COUNT decides how that O(N) is packaged. - **Small COUNT (the default 10)**: each command is very short, so the impact on other clients' latency is minimal — but you need a large number of round trips. Iterating 10 million keys at 10 per call, with a 0.2 ms round trip, is minutes of wall-clock time spent almost entirely on network latency. - **Large COUNT (1000+)**: dramatically fewer round trips and a much faster iteration, at the cost of each command occupying the server longer and producing a bigger reply. Push it far enough (say COUNT 100000) and you have re-created `KEYS`, one blocking chunk at a time. A practical default is a few hundred to a thousand, tuned by watching latency. If a maintenance scan is competing with production traffic, keep COUNT modest and add a small sleep between batches rather than trying to finish quickly. ## COUNT is ignored for small collections For `HSCAN`/`SSCAN`/`ZSCAN`, when the collection is small enough to be stored in its compact flat encoding rather than a hash table, Redis returns **all** the elements in the very first call and a cursor of `0`, ignoring COUNT entirely. That is cheap for a small collection but it means you cannot rely on the type-scan commands to bound the reply size for arbitrary inputs — they bound it only once the collection is big enough to be a real hash table. ## Putting it together in client code The correct loop is: start with cursor `"0"`, call `SCAN cursor MATCH ... COUNT ...`, process the batch (which may be empty), and repeat **until the returned cursor is the string `"0"`**. Do not count iterations, do not break on empty, do not assume a fixed number of elements per call. If you are throttling, sleep between calls rather than shrinking COUNT to 1, which just multiplies round trips. ## Cost intuition to state in an interview "COUNT is a hint about work per call, MATCH filters after the fact so it saves bandwidth not scanning, empty batches are normal, and I stop only at cursor 0." That sentence contains everything the question is really asking about.

  • Would raising COUNT to 100000 be a good way to finish a scan faster?
    It defeats the purpose. Each call would then traverse a huge slice of the table in a single command and build a large reply, so you are back to KEYS-style stalls, just delivered in several instalments. Keep COUNT in the hundreds-to-low-thousands and, if the scan must go faster, run it against a less loaded node or accept a longer wall-clock time.
  • Does MATCH make a SCAN over a 50-million-key database cheaper?
    Not in traversal cost — every key is still visited and glob-matched, because Redis has no prefix index over the keyspace. MATCH reduces reply size, server memory for the reply, network bytes and client-side work, which is worth having, but the number of buckets walked is identical.

saying these in an interview costs you the question

  • "COUNT is the page size, so I'll always get exactly COUNT keys" — it is a work hint; you can get more, fewer, or none.
  • "MATCH lets Redis jump straight to matching keys" — there is no prefix index; filtering is post-retrieval.
  • "An empty array means the scan is complete" — only cursor 0 means complete.
  • "COUNT must stay the same across the whole iteration" — it may vary per call.
  • "Setting COUNT very high is a free speed-up" — it converts SCAN back into long blocking commands.

context