skip to content

A page render needs 200 objects that are individually cached in Redis. How do you fetch them without paying 200 round trips, and how do you handle the subset that comes back missing?

level: middleimportance: should knowfreq 50%

answer

  1. MGET = one round trip, positional nils
  2. batch the misses into ONE database query
  3. MSET has no TTL — pipeline SET ... EX
  4. cluster: CROSSSLOT unless same slot / hash tag
  5. MGET is O(N) on the main thread — chunk it

basics

~20 s

Use MGET with all 200 keys in one round trip. The reply is positional: nil in slot i means key i missed. Collect the missing ids, load them in one batched database query, then write them back — pipelined SETs or MSET plus expiries, since MSET cannot set TTLs.

solid answer

~60 s

Issue one `MGET k1 k2 ... k200`. The reply is an array in the same order as the arguments, with `nil` in each missed position, so you map slot back to id, gather the misses, and do **one** batched database query (`WHERE id IN (...)`) rather than N of them. Writing back is the awkward half: `MSET` is atomic but cannot attach TTLs, and a cache entry without a TTL is a liability. So write back with **pipelined `SET key value EX ttl`** commands — one round trip, N commands, each carrying its own expiry. Two constraints to name. In **Redis Cluster** a multi-key command requires all keys in the same hash slot, so a cluster-aware client splits the MGET per slot or per node and merges the results; some do it transparently. And MGET is **O(N) on the single-threaded server** — 200 keys is fine, 100k keys in one call is a latency incident for everyone else on that node, so chunk into batches of a few hundred.

code

text · 8 lines
text
MGET product:v1:1 product:v1:2 product:v1:3 product:v1:4
1) "{...id:1...}"
2) (nil)          <- miss, id 2
3) "{...id:3...}"
4) (nil)          <- miss, id 4

# one database query for the misses:
#   SELECT * FROM products WHERE id IN (2,4)

go deeper

for a junior

Know MGET exists, returns results in argument order, and puts nil where a key was absent.

for a middle

Explain batching the misses into one database query and why write-back uses pipelined SET ... EX rather than MSET.

for a senior

Add the operational limits: O(N) on the main thread so chunk the batch, CROSSSLOT in cluster mode and how clients split it, per-position hit-ratio measurement.

for a principal

Reason about fan-out shape overall: whether 200 keys should be 200 keys, when to cache the assembled page fragment instead, and how batch expiry synchronization is created here.

## The problem with a loop of GETs A cache lookup is microseconds of server work wrapped in a network round trip that is typically tens to hundreds of microseconds inside a datacenter, and milliseconds across zones. Two hundred sequential GETs means 200 round trips, and that latency is entirely additive: the request spends most of its time waiting, not computing. This is the cache-layer version of the N+1 query problem. There are two ways to collapse it, and they solve slightly different things. ## MGET `MGET key1 key2 ... keyN` returns an array of replies in **argument order**, with a `nil` in every position whose key was absent. That positional correspondence is the whole API: you keep your ordered list of ids, zip it against the reply, and split into hits and misses. MGET is one command, so it executes atomically with respect to other commands — you get a consistent snapshot of those keys at a single point in time. It is also `O(N)` in the number of keys, executed on Redis's single main thread. That is cheap for hundreds of keys and dangerous for hundreds of thousands: one giant MGET is a head-of-line blocker for every other client on that node, and it also produces one very large reply that must be allocated in the output buffer. Chunk large fan-outs into batches of a few hundred to a couple of thousand. ## Pipelining Pipelining is the more general tool: the client writes many independent commands to the socket without waiting for each reply, then reads all the replies. It works for heterogeneous commands (GET, HGETALL, ZSCORE mixed together), which MGET cannot do since MGET only reads string keys — calling it on a hash or list yields nil for those positions rather than an error. Pipelining is *not* atomic: other clients' commands can interleave between your pipelined commands. If you need atomicity you need MULTI/EXEC or a script, which is a separate topic. For cache reads, atomicity across the batch is almost never required. Rule of thumb: MGET when everything is a plain string key you can name up front; a pipeline when the shapes differ or when you are mixing reads and writes. ## Handling the miss subset The mistake that eats the benefit is fixing misses one at a time. If 30 of 200 keys missed, do **one** database query for those 30 ids, not 30 queries. The whole point of batching the cache read is to end up with a single batched load behind it. Then write back. This is where people reach for `MSET` and regret it: MSET sets multiple keys atomically but has **no** TTL option, so every entry it writes is immortal unless you follow with N `EXPIRE` calls — which is both more commands than a pipelined SET and non-atomic in exactly the dangerous way (crash between MSET and EXPIRE leaves immortal keys). Pipeline `SET key value EX ttl` instead: same single round trip, every key gets its expiry atomically with its value. A subtlety worth mentioning: giving every key in a batch the identical TTL means the whole batch expires simultaneously, which produces a synchronized miss wave later. Spreading expiries is TTL-policy territory, but knowing the batch write is where it originates is part of understanding this pattern. ## Cluster constraints Redis Cluster shards by hash slot (CRC16 of the key, or of the substring inside `{}` if a hash tag is present, modulo 16384). Any multi-key command must have all its keys in **one** slot, otherwise the node replies `CROSSSLOT Keys in request don't hash to the same slot`. So a naive 200-key MGET fails on a cluster. Options: (a) use a cluster-aware client that splits the MGET by slot, issues the parts concurrently to the owning nodes, and reassembles the result in argument order — most mature clients do this; (b) use hash tags to co-locate keys that are always fetched together, at the cost of skewing slot distribution; (c) fall back to a pipeline of single-key GETs routed per node, which is still one round trip per node rather than per key. Single-node and Sentinel deployments have no such restriction. ## Partial failure and result assembly With a pipeline, individual commands can fail independently (a WRONGTYPE, for example) — read every reply and check it rather than assuming success. With MGET, an error fails the whole command. Either way, the fallback stays the same as any cache-aside path: anything you did not get from Redis, you get from the database, and a Redis-level failure degrades to "everything missed". Finally, measure the hit ratio **per batch position**, not just globally: a page that reliably has one always-missing key still pays a database round trip on every render, and that is invisible in an aggregate 97% hit rate.

  • Why not use MSET to write back the missing entries?
    MSET has no expiry option, so every key it writes lives forever unless you issue separate EXPIRE commands afterwards. That is more commands than just pipelining SET ... EX, and it is non-atomic per key: a crash between the MSET and the EXPIREs leaves immortal cache entries. Pipelined SET with EX costs the same single round trip and keeps value and TTL atomic.
  • What breaks if you send a 200-key MGET to a Redis Cluster?
    The node returns CROSSSLOT unless every key hashes to the same slot, because multi-key commands must be served by a single shard. A cluster-aware client fixes this by splitting the key list per slot or per node, issuing the sub-requests concurrently, and reassembling the replies in the original order. Alternatively, hash tags can co-locate keys that are always read together, at the cost of uneven slot distribution.
  • When would you pipeline individual GETs instead of using MGET?
    When the commands are not all plain string GETs — for example mixing HGETALL, ZSCORE, and GET — since MGET only reads string keys and returns nil for other types. Pipelining also lets you interleave reads and writes in the same round trip. The tradeoff is that a pipeline is not atomic: other clients' commands can execute between yours.

MGET is handing the librarian a list of 200 call numbers and getting back a tray with gaps where books were checked out, instead of walking to the desk 200 times.

saying these in an interview costs you the question

  • Looping single GETs and calling it 'the cache is slow' when the cost is round trips.
  • Fixing the missed subset with one database query per missing id.
  • Using MSET for cache write-back and ending up with keys that never expire.
  • Assuming MGET works unchanged on Redis Cluster across arbitrary keys.
  • Sending one enormous MGET of tens of thousands of keys, blocking the single-threaded server for everyone.

context