skip to content

A Redis Set holds several million members. What does SMEMBERS cost the server and the client, and how would you satisfy the common requests against that set — its size, whether specific values are in it, a random sample, the size of its overlap with another set — without ever materialising it?

level: middleimportance: must knowfreq 58%

answer

  1. O(N) reply = a pause for every other client
  2. client-output-buffer-limit kills the connection
  3. SCARD O(1), SMISMEMBER batch, SRANDMEMBER n
  4. SINTERCARD … LIMIT n stops counting early
  5. pages of members = wrong data structure

basics

~20 s

SMEMBERS is O(N): one giant reply that occupies the single command thread and can blow the client output buffer, killing the connection. Ask set-native questions instead — SCARD for size, SISMEMBER/SMISMEMBER for membership, SRANDMEMBER for a sample, SINTERCARD with LIMIT for overlap size.

solid answer

~60 s

`SMEMBERS` is O(N). Redis runs one command at a time, so a multi-million-member set means a long allocation-and-serialise pause during which *no other client is served*, plus a burst of network traffic. Server-side, the reply sits in that connection's output buffer; `client-output-buffer-limit` for normal clients can be exceeded and Redis closes the connection — the client retries and reproduces the stall. Client-side you now hold the whole set in memory and pay the parse cost. The fix is usually not "iterate more politely" but "stop enumerating": - **Size** → `SCARD`, O(1). - **Membership** → `SISMEMBER`, or `SMISMEMBER key m1 m2 …` (6.2+) for a batch — never pull the set to filter in the app. - **Sample** → `SRANDMEMBER key n`. - **Overlap size** → `SINTERCARD k1 k2 LIMIT n` (7.0+), which stops counting at the limit instead of building the intersection. If the access pattern really is "pages of members", the key is modelled wrong — shard it or use an ordered structure. Use `SSCAN` only when you genuinely must visit every member.

code

text · 16 lines
text
# DON'T: proportional to the set, on the single command thread
> SMEMBERS users:active:2026-08      # 4.2M members -> one huge reply

# DO: ask the narrow question instead
> SCARD users:active:2026-08
(integer) 4218773                    # O(1)

> SMISMEMBER users:active:2026-08 u:17 u:99 u:404
1) (integer) 1
2) (integer) 1
3) (integer) 0                       # membership for a batch, one round trip

> SRANDMEMBER users:active:2026-08 5  # sample, cost ~ n

> SINTERCARD 2 users:active:2026-08 users:paid LIMIT 1000
(integer) 1000                       # "at least 1000" answered without building the intersection

go deeper

for a junior

Know that SMEMBERS returns everything and is O(N), that this is dangerous on a big set, and that SCARD gives the size and SISMEMBER tests one member without fetching anything.

for a middle

Explain both halves of the cost — single-threaded occupancy on the server plus the client output buffer — and reach for the narrow command (SCARD, SMISMEMBER, SRANDMEMBER, SINTERCARD LIMIT) before reaching for iteration.

for a senior

Diagnose it live: SLOWLOG for the offending command, CLIENT LIST omem for the buffer, and the disconnect-retry loop that masquerades as flakiness. Name UNLINK vs DEL for the cleanup, and SINTERSTORE to keep large intermediates server-side.

for a principal

Frame it as a modelling question: a request for pages of members means the Set is the wrong structure or the wrong granularity. Weigh sharding the key, moving the enumerable projection to the primary store, and setting a deliberate client-output-buffer-limit so one bad query degrades a connection rather than the instance.

## Why one big reply is a systemic cost, not a personal one Redis executes commands one at a time on a single thread. That means the cost of a command is not just the latency the caller experiences — it is a pause imposed on every other client whose request happens to be queued behind it. `SMEMBERS` on a set of five million members must walk the entire underlying hash table, allocate a reply array of five million entries, and serialise all of them before a single byte is sent. A p99 that is normally well under a millisecond becomes hundreds of milliseconds for unrelated traffic. This is the same family of hazard as `KEYS`, `HGETALL` on a huge hash, or `LRANGE 0 -1` on a huge list: **any command whose reply size is proportional to the size of the data structure is only safe when you control that size.** ## The second failure: the client output buffer The serialised reply does not vanish once the command returns. It is buffered on the server, in that connection's **client output buffer**, until the socket drains it. Redis enforces `client-output-buffer-limit` for normal clients (hard and soft limits), and a multi-hundred-megabyte reply to a consumer that reads slowly — a GC pause, a slow network, a busy event loop — can exceed it. Redis then **closes the connection**. The application sees an abrupt disconnect, reconnects, retries the same `SMEMBERS`, and reproduces the stall. On dashboards this reads as "Redis is flaky", when it is a self-inflicted loop. Memory pressure is real too: that buffer counts toward the instance's RSS while it exists, so a handful of concurrent big replies can push an instance toward `maxmemory` behaviour or the OOM killer. On the client side you have now allocated the entire set in application memory and paid the protocol-parse cost for every element — frequently the more expensive half in a managed-runtime service, and a reliable source of GC pauses. ## Ask the question you actually have Almost every real use of `SMEMBERS` is a proxy for a narrower question that Redis can answer directly, without producing a reply proportional to the set: - **"How big is it?"** → `SCARD key`. O(1); the cardinality is maintained as metadata, not counted. - **"Is X in it?"** → `SISMEMBER key X`, O(1). For a batch of candidates, `SMISMEMBER key m1 m2 m3 …` (Redis 6.2+) answers all of them in one round trip, returning a 1/0 per member. Pulling the whole set into the application to run a `contains` check is the classic anti-pattern — it moves millions of members over the wire to answer a question about ten. - **"Give me a few examples."** → `SRANDMEMBER key n` returns up to n distinct members without removing them; a negative count allows repeats. Cost is proportional to n, not to the set. - **"How much do these two sets overlap?"** → `SINTERCARD numkeys k1 k2 LIMIT n` (Redis 7.0+) returns only the count and, with `LIMIT`, stops as soon as it reaches n. That matters when the honest requirement is "are there at least 1000 in common?" — the alternative, `SINTER` and then measuring the reply, materialises the whole intersection and reintroduces exactly the problem you were avoiding. If you need the intersection itself for further work, compute it server-side into a key (`SINTERSTORE`) and keep it there rather than shipping it out. ## When the right fix is to restructure the key If the requirement genuinely is "show the user page 3 of the members", no iteration command will make that pleasant: Sets are unordered, so there is no stable page boundary to resume from. That is a modelling signal, not a command-choice problem. Options: **shard** the logical set across many keys by a hash of the member (`items:{shard}` for shard 0..255), so any one key stays small and enumerable; keep the ordered/pageable view in a structure designed for ranges (a Sorted Set — that structure's own topic covers its range and pagination commands); or store the enumerable projection in the primary datastore and keep Redis for the O(1) membership and cardinality answers it is genuinely good at. Sharding also caps the blast radius of deletion — and when you do drop a multi-million-member key, prefer `UNLINK` over `DEL` so reclamation happens off the main thread. ## If you truly must visit every member Use `SSCAN` when you genuinely must visit every member; its cursor contract — at-least-once delivery, possible duplicates, `COUNT` as a work hint — is covered in the SCAN topic. ## Diagnosing it in production `SLOWLOG GET` surfaces the offending command with its execution time. `CLIENT LIST` shows `omem` and `obl`/`oll` per connection — a large output buffer on one client is the smoking gun. `LATENCY DOCTOR` and the `INFO` field `total_net_output_bytes` corroborate. The remediation order is: answer the narrower question, then restructure the key, and only then iterate.

  • The client-output-buffer-limit for normal clients defaults to 0 0 0 (unlimited). Doesn't that make the disconnect risk theoretical?
    It removes the disconnect, not the cost. With no cap, the buffer grows until the slow consumer drains it, so a few concurrent multi-hundred-megabyte replies inflate the instance's RSS and can push it into eviction, swapping, or an OOM kill — a worse outcome than one dropped connection. Many managed Redis providers set a non-zero normal limit precisely for that reason, and hosting your own is a good argument for setting one deliberately.
  • You need the members two large sets have in common in order to act on each one. SINTERCARD only gives a count — what do you do?
    Keep the result server-side: `SINTERSTORE tmp:job k1 k2` computes the intersection inside Redis and stores it as a key, with an EXPIRE so it self-cleans. Then decide based on `SCARD tmp:job` — if it is small, fetch it; if it is still huge, process it in place or iterate it with SSCAN. The win is that the large intermediate never crosses the wire.
  • Would pipelining or raising the client's socket timeout fix an SMEMBERS stall?
    No. Both are client-side knobs and neither changes the server's work: Redis still walks N members, allocates one reply of N entries, and blocks other commands for the duration. A longer timeout only means the client waits patiently for the same stall, and pipelining more commands behind it makes the queue worse. The fix has to reduce reply size — a narrower command, or a smaller key.

SMEMBERS is demanding a photocopy of the whole filing cabinet to check whether one folder exists. SCARD reads the label on the drawer, SMISMEMBER asks the clerk about your ten folders, SRANDMEMBER grabs a handful — nobody blocks the corridor.

saying these in an interview costs you the question

  • "SMEMBERS is one round trip, so it's efficient" — the round-trip count is irrelevant; the reply size and the single-threaded occupancy are the cost.
  • "Just increase the client timeout / use a bigger connection pool" — client-side settings don't reduce the server's O(N) work or the output buffer.
  • Fetching the whole set into the application to check whether a handful of values are present, instead of SISMEMBER/SMISMEMBER.
  • "SCARD has to count the members, so it's O(N) too" — cardinality is maintained as metadata and SCARD is O(1).
  • "Use SINTER and take the length of the reply" to get an overlap size — that materialises the entire intersection; SINTERCARD (with LIMIT) exists to avoid it.
  • Treating SSCAN as the answer to every large-set request, including size, membership and sampling, which Redis answers directly without iterating.

context