skip to content

You want to show a random sample of fields from a large Redis hash without pulling the whole thing. How does the Redis command HRANDFIELD serve that, how does its reply shape change with the WITHVALUES option under RESP2 versus RESP3, and how does it compare with HSCAN for the same job?

level: middleimportance: should knowfreq 22%

answer

  1. O(count) — cost independent of hash size
  2. HSCAN bounded by buckets visited, full walk O(N)
  3. WITHVALUES: RESP2 flat array, RESP3 map
  4. order arbitrary — never a cursor
  5. sample vs enumerate — not substitutes

basics

~20 s

HRANDFIELD key count returns random field names at O(count) cost, independent of how big the hash is. WITHVALUES adds each value: a flat field,value,field,value array under RESP2, a map under RESP3. It samples; HSCAN enumerates the whole hash. Its order is not pagination.

solid answer

~1 min

`HRANDFIELD key [count [WITHVALUES]]` picks random fields from a hash without removing anything. Its cost is **O(count)** — the work is bounded by how many entries you ask for, not by the size of the hash, so sampling 5 fields out of a ten-million-field hash is as cheap as out of a ten-field one. That is the whole point versus `HGETALL`. `WITHVALUES` (valid only with a count) returns values alongside fields. Under **RESP2** the reply is a flat array — `field, value, field, value, …` — that the client must pair up. Under **RESP3** the server sends a proper map/array-of-pairs type, so most clients hand you a dictionary directly. Code that hard-codes flat-array parsing breaks when a client negotiates RESP3. The returned order is arbitrary: not stable, not a fair shuffle of the hash. **Never use it as a cursor or pagination** — repeated calls can return the same fields forever and give no coverage guarantee. `HSCAN` is the enumerator: a cursor walk in bounded chunks, cost bounded by buckets visited, guaranteeing every element present for the whole iteration is returned at least once (duplicates possible). Sampling and enumeration are not substitutes — pick by which guarantee you need. The `count` argument follows the same positive-distinct / negative-with-replacement convention as `SRANDMEMBER` and `ZRANDMEMBER`.

code

text · 18 lines
text
HSET colors a red b green c blue

# sample, cost bounded by the count not the hash
HRANDFIELD colors 2
1) "c"
2) "a"

# RESP2: flat field/value array
HRANDFIELD colors 2 WITHVALUES
1) "a"
2) "red"
3) "c"
4) "blue"

# RESP3 (after HELLO 3): a map
HRANDFIELD colors 2 WITHVALUES
1# "a" => "red"
2# "c" => "blue"

go deeper

for a junior

Recall that HRANDFIELD returns random field names, that adding WITHVALUES also returns the values, and that HSCAN is the command for walking a whole hash.

for a middle

Explain O(count) versus HSCAN's bucket-bounded traversal, the RESP2 flat-array versus RESP3 map reply for WITHVALUES, and why the returned order cannot be used as a cursor.

for a senior

Frame the choice by guarantee — sample versus coverage — and call out the operational failure modes: an HRANDFIELD 'pagination' loop that never terminates, and reply-parsing that breaks when a client negotiates RESP3.

for a principal

Discuss it as a protocol- and access-pattern decision: which reads may touch a hot key at all, whether sampled statistics are acceptable in place of full enumeration, and standardising RESP3 handling in the client layer so reply shape is not a per-call-site concern.

## What HRANDFIELD is for A Redis *hash* is a single key holding a set of field→value pairs, like a small dictionary stored under one name. When such a hash gets large, reading it whole with `HGETALL` costs O(N) in the number of fields and materialises the entire thing into one reply — expensive on the server and on the network. `HRANDFIELD` exists for the case where you do not want all of it: you want a handful of representative entries. Admin panels showing "a few example rows", debug endpoints, cache-warming that exercises random members, and approximate statistics over a very large hash are the everyday uses. Syntax: ``` HRANDFIELD key [count [WITHVALUES]] ``` With no count it returns one random field *name* (a bulk string), or a nil reply if the key does not exist. With a count it returns an array of field names, or an empty array for a missing key. Nothing is removed — it is a pure read, unlike `SPOP` on sets, which has no hash equivalent. The `count` argument uses the same positive-distinct / negative-with-replacement convention as `SRANDMEMBER` and `ZRANDMEMBER`; that convention and its reply-size implications belong to the sets material and are not re-derived here. ## Why the cost is independent of hash size The documented complexity is **O(count)** for the counted forms (effectively O(1) for the single-field form). Redis picks entries directly out of the hash's internal representation rather than walking it, so the amount of work scales with what you asked for, not with what is stored. This is the property that makes it safe on a huge key: `HRANDFIELD bigkey 10` on a ten-million-field hash does roughly the same work as on a ten-field hash, and does not block the single-threaded server the way an `HGETALL` over the same key would. Contrast that with `HSCAN`, whose per-call cost is bounded by the number of hash-table buckets visited for that cursor step, and whose *total* cost across a full iteration is proportional to the size of the collection. `HSCAN` is cheap per call and expensive in aggregate; `HRANDFIELD` is cheap because it never promises aggregate coverage at all. ## Reply shape: WITHVALUES under RESP2 vs RESP3 `WITHVALUES` may only be used together with a count. It makes Redis return each selected field's value as well, saving a follow-up `HMGET` round trip. The shape depends on the protocol the connection negotiated (RESP3 is opted into with the `HELLO 3` handshake, supported since Redis 6): - **RESP2** has only a generic array type, so the reply is a *flat* array: `[field1, value1, field2, value2, …]`. The client library — or your code — has to pair adjacent elements. - **RESP3** has a map type, so the server returns the pairs as a map (clients typically surface a dictionary, or an array of two-element pairs depending on the library). The practical consequence: any parsing that assumes the flat layout silently breaks if the client driver upgrades to or negotiates RESP3, and vice versa. Let the client library normalise the reply rather than indexing raw positions, and be aware that `ZRANDMEMBER … WITHSCORES` has exactly the same RESP2/RESP3 split. ## Ordering is not pagination The order of returned fields is unspecified. It is neither a stable ordering you can resume from nor a uniform shuffle of the whole hash. There is no cursor, no marker, and no guarantee that repeated calls will eventually cover every field — a given field may keep reappearing and another may never appear. Building "page through the hash by calling HRANDFIELD repeatedly and de-duplicating" produces a loop with unbounded runtime and no completion condition. If a UI needs *browse* semantics, that is `HSCAN`. ## HRANDFIELD samples, HSCAN enumerates The two commands answer different questions and are not interchangeable: | | HRANDFIELD | HSCAN | |---|---|---| | Purpose | sample | enumerate | | Cost | O(count), independent of hash size | bounded per call by buckets visited; O(N) over a full iteration | | Guarantee | none about coverage | every element present for the whole iteration returned at least once | | Duplicates | possible (with a negative count, by design) | possible across cursor steps | | Resumable | no cursor | cursor-based, resumable | | Randomness | random selection | traversal order, not random | Choose by guarantee: if the requirement contains the word *every*, use `HSCAN`; if it contains *a few* or *random*, use `HRANDFIELD`. A common design mistake is reaching for `HRANDFIELD` because it "feels lighter" when the actual requirement was full traversal — the code then quietly misses fields. One further limit: since `HRANDFIELD` does not remove anything, "take a random field and delete it" is `HRANDFIELD` followed by `HDEL`, which is two commands with a race window between them; making it atomic needs a small Lua script.

  • Your dashboard calls HRANDFIELD with WITHVALUES and started returning malformed data after a client-library upgrade. What is the likely cause?
    The upgraded client probably negotiates RESP3 with `HELLO 3`, so the WITHVALUES reply arrives as a map instead of the RESP2 flat `field, value, field, value` array. Code that pairs adjacent array elements by index now misreads it. The fix is to consume whatever the library returns as a mapping rather than indexing raw positions, or to pin the protocol deliberately.
  • Can you use repeated HRANDFIELD calls to eventually visit every field in a hash?
    No. There is no cursor and no coverage guarantee, so the same fields can keep reappearing while others never surface; the loop has no completion condition and unbounded runtime. Full traversal is `HSCAN`, which guarantees that any element present for the whole iteration is returned at least once, at the cost of possible duplicates across cursor steps.
  • Why is HRANDFIELD safe on a hash with millions of fields when HGETALL is not?
    HRANDFIELD's complexity is O(count) — the server does work proportional to how many entries you requested, not how many exist, because it picks entries directly rather than walking the structure. HGETALL is O(N) and materialises the entire hash into one reply, which on a single-threaded server blocks other clients and spikes the output buffer.

saying these in an interview costs you the question

  • Calling HRANDFIELD in a loop as a pagination or full-iteration strategy
  • Assuming WITHVALUES always yields a flat array, ignoring the RESP3 map shape
  • Believing HRANDFIELD's cost grows with the number of fields in the hash
  • Treating HRANDFIELD and HSCAN as interchangeable ways to read a big hash
  • Expecting the returned order to be a stable or fairly shuffled ordering of the hash

context