When storing an object in Redis, what do you gain and lose by modelling it as a hash with one field per attribute versus serialising the whole object to JSON and storing it in a single string key?
answer
- hash = partial read/write + HINCRBY atomic
- JSON string = one round trip, nesting, whole-object swap
- JSON update = GET-modify-SET = lost update race
- hash values are flat strings only
- field names stored per hash — cost × N objects
basics
~20 sA hash gives partial reads and writes (HGET/HSET one field), atomic per-field updates like HINCRBY, and compact memory for small objects. A JSON string gives one round trip, nesting and arrays, and atomic whole-object replacement — but every update is a read-modify-write that can lose concurrent changes.
solid answer
~50 s**Hash wins when fields are updated independently.** `HSET user:42 plan pro` touches one attribute; two concurrent writers changing different fields both succeed. `HINCRBY` increments a counter atomically with no read-modify-write. `HMGET` fetches only the attributes a call path needs, saving bandwidth and deserialisation. Small hashes are also stored in a compact encoding, so they are memory-cheap. **JSON string wins when the object is read and written whole.** One `GET`, one deserialisation, one `SET` — and that `SET` is an atomic swap of a consistent snapshot. JSON also nests: arrays and sub-objects are free, whereas a hash is strictly flat strings, so nested data has to be flattened into dotted field names or serialised into a field and become opaque. **The decisive loss on the JSON side is concurrency**: updating one attribute means `GET`, edit, `SET`, so two writers overwrite each other. Fixing that needs `WATCH`/`MULTI`/`EXEC` or a Lua script. Rule of thumb: independently mutated attributes or per-field counters → hash. Whole-object read/write with nesting → JSON string.
code
text · 16 lines# hash: one field, one command, no race
HINCRBY user:42 logins 1
HSET user:42 plan pro
HMGET user:42 name plan # fetch only what you need
# JSON string: read-modify-write, and two writers lose updates
GET user:42 # {"name":"Ada","plan":"free","logins":3}
# ... edit in the client ...
SET user:42 '{"name":"Ada","plan":"pro","logins":4}'
# making the JSON path safe requires optimistic concurrency
WATCH user:42
GET user:42
MULTI
SET user:42 '{...edited...}'
EXEC # nil if another client wrote first -> retrygo deeper
Say that a hash lets you read and write single attributes while a JSON string is fetched and replaced whole, and give one example of each.
Cover partial access, HINCRBY atomicity, memory of small hashes, JSON's nesting and single-round-trip codec reuse, and the lost-update problem.
Drive the choice from access pattern and concurrency: which attributes mutate independently, what needs to be atomic, what WATCH or Lua would cost, and how field-name overhead scales across millions of objects.
Frame it as a schema decision with downstream cost: server-visible structure buys partial access, atomic field ops and per-field TTL, while opacity buys codec freedom, compression and clean versioning — and hybrids are legitimate when parts of the object differ in access pattern.
## The two shapes ``` # hash HSET user:42 name Ada plan pro logins 3 # JSON in a string SET user:42 '{"name":"Ada","plan":"pro","logins":3}' ``` Both are one key. The difference is whether Redis can see inside it. ## What a hash gives you **Partial reads.** `HGET user:42 plan` or `HMGET user:42 name plan` returns just those bytes. On a wide object accessed on a hot path, this is the difference between shipping 40 bytes and shipping 4 KB per request — and it removes the client-side JSON parse too, which is often the larger cost. **Partial writes, and therefore concurrency.** `HSET user:42 plan pro` writes one field. A concurrent `HSET user:42 email …` writes a different field. Both land; neither clobbers the other. With JSON, each writer read the whole document before editing, so the second `SET` silently discards the first writer's change — the classic lost update. **Server-side field operations.** `HINCRBY user:42 logins 1` is a single atomic command. The JSON equivalent is read, parse, increment, serialise, write — three round trips, a race, and a need for `WATCH` or Lua to be correct. **Memory efficiency at small sizes.** A hash below the configured entry/value thresholds is stored in one compact flat structure rather than a hash table, which makes many small objects far cheaper than one key per attribute, and competitive with (often better than) a JSON blob whose keys are repeated as text in every value. **Introspection.** `HLEN`, `HEXISTS`, `HKEYS`, `HSTRLEN`, `HSCAN` and `HRANDFIELD` all work because Redis understands the structure. A JSON string is opaque bytes; the server can tell you its length and nothing else. **Per-field TTL** (Redis 7.4+). `HEXPIRE` can expire individual fields — impossible for attributes inside a serialised blob. ## What a JSON string gives you **Atomic whole-object semantics.** `SET` replaces the document in one step, so readers never observe a half-updated object. With a hash, a multi-field `HSET` is atomic, but two separate `HSET` calls are not — a reader between them sees a mixed state unless you wrap them in `MULTI`/`EXEC` or a script. **Nesting and types.** JSON has arrays, sub-objects, booleans, nulls and numbers. Hash fields and values are flat byte strings, full stop. Modelling `user.addresses[0].city` in a hash means either flattening (`addresses.0.city`) — which makes list mutation awkward — or serialising the sub-tree into one field, which is JSON-in-a-hash and loses the very partial-access benefit you chose the hash for. **One round trip, one codec.** `GET` + your existing serializer (JSON, MessagePack, protobuf) reuses the same object mapper the rest of your stack already uses, including schema evolution rules. Hashes force a hand-written mapping from your object to string fields, with explicit handling for absent fields, type coercion, and `null` versus empty string. **Compression.** A blob can be gzip'd or stored in a compact binary codec, which for large documents can beat a hash by a wide margin. Hash field names, by contrast, are stored per hash — 500,000 objects with a field literally named `last_login_timestamp` pay for that string half a million times. **Simpler versioning.** Embedding a `"v":3` in the document and migrating on read is a well-trodden pattern; the hash analogue is a `_v` field plus per-field migration logic. ## The decision Choose a **hash** when: - attributes are read or written independently (profile field edits, feature flags, session attributes); - you need atomic per-field counters (`HINCRBY`); - objects are small, numerous and memory matters; - you want per-field TTL on 7.4+. Choose a **JSON string** when: - the object is essentially always fetched and replaced whole (a cached API response, a rendered fragment); - the data is nested or list-shaped; - you want to reuse an existing serializer, schema and compression; - the value is a cache entry whose whole-object TTL and whole-object replacement are exactly the semantics you want. ## Practical middle grounds - **Hybrid**: hot, independently-mutated scalars as hash fields; a cold nested sub-tree serialised into one field. Explicit, and honest about which parts are opaque. - **Two keys**: a hash for mutable state and a string for an immutable blob, sharing a key prefix (and a hash tag in a cluster so they stay in one slot). - **A JSON module**: the RedisJSON module adds path-addressable JSON with server-side get/set on sub-paths, which recovers partial access for nested documents. It is a module, not core Redis — call that out explicitly, because assuming `JSON.SET` exists on a stock server is a common mistake. ## The answer that lands Name the concurrency consequence explicitly. Most candidates list "partial reads" and stop; the interviewer is usually listening for the recognition that JSON-in-a-string turns every field update into a read-modify-write, and that fixing it costs you `WATCH` or Lua — whereas a hash gets it for free.
- Two requests update different attributes of the same object at the same time. How does each representation behave?With a hash, `HSET key a …` and `HSET key b …` touch different fields and both survive — Redis executes commands one at a time and neither rewrites the other's field. With JSON in a string, both clients read the same document, each edits its own attribute, and both write the full document back; the later `SET` wins and the earlier change is silently lost. Making that safe requires `WATCH`/`MULTI`/`EXEC` with a retry loop, or a Lua script that does the edit server-side.
- How do you keep a hash object memory-efficient when you have millions of them?Keep the object small enough to stay in the compact small-hash encoding, and keep field names short — field name strings are stored in every hash instance, so a verbose name is paid for once per object. Store values in their most compact form (integers as integers rather than padded strings) and avoid stuffing large blobs into a field when a separate key would do.
- Does storing JSON in Redis mean you should reach for the RedisJSON module?Only if you genuinely need server-side access to sub-paths of nested documents — `JSON.GET`/`JSON.SET` with a path avoid shipping and rewriting the whole document. It is a module that must be present in the deployment, not part of core Redis, so it is a deployment decision, and on a plain server the only options are hashes or opaque serialised strings.
A hash is a form with separate boxes — two clerks can fill in different boxes at once. A JSON string is a single printed page: to change one line you retype the whole page, and whoever prints last wins.
saying these in an interview costs you the question
- Claiming a hash 'stores JSON' or supports nested objects — values are flat strings
- Ignoring that updating one attribute of a JSON string is a read-modify-write with a lost-update race
- Assuming JSON.SET / RedisJSON commands exist on a stock Redis server
- Believing a multi-command sequence of HSETs is atomic (only a single HSET call is)
- Choosing hashes purely 'because they are faster' with no argument about access pattern