What drives your choice of serialization format for values stored in a Redis cache, and when would you store an object as a Redis hash instead of one serialized string?
answer
- Redis stores opaque bytes; cost is client-side
- axes: CPU, size, schema tolerance, cross-language, debuggability
- string = whole object; hash = per-field HGET/HSET
- small hashes use the listpack encoding
- never pickle/Java Serializable — slow + RCE risk
basics
~20 sPick a format that is fast to encode/decode, compact, and tolerant of field additions — JSON for debuggability, a binary format like Protobuf/MessagePack for size and speed. Use a string when you always read the whole object; use a hash when you read or update individual fields with HGET/HSET.
solid answer
~1 minRedis stores opaque bytes, so serialization cost is entirely on the client — and at cache scale it usually dominates the Redis round trip itself. What I weigh: **encode/decode CPU** (this runs on every hit), **size** (memory and network; Redis charges you for both), **schema tolerance** (can new code read old bytes — and does the key carry a version segment so it never has to), **cross-language readability** if more than one runtime reads the key, and **debuggability** (`GET` on a JSON value tells you everything; on a language-native binary blob it tells you nothing). Default: JSON for most caches, a compact binary format when profiling shows serialization or bandwidth is the bottleneck. Compress only above a threshold (a few KB) — gzip on a 200-byte value costs more than it saves. **String vs hash:** a string is one blob — atomic whole-object read/write, simplest. A hash lets you `HGET`/`HSET` individual fields, which avoids read-modify-write of the whole object and cuts bandwidth on wide objects. Small hashes are also memory-efficient thanks to the listpack encoding. The catch: expiry is per-key, not per-field before Redis 7.4, so you cannot age fields independently. Avoid language-native serialization (Java `Serializable`, Python `pickle`) — slow, brittle across versions, and a deserialization-gadget security risk if anything untrusted can write to Redis.
code
text · 10 lines# string: one blob, atomic whole-object read/write, MGET-friendly
SET user:v2:42 "{\"id\":42,\"email\":\"[email protected]\",\"last_seen\":1723}" EX 600
GET user:v2:42
# hash: per-field access, no read-modify-write of the whole object
HSET user:v2:42 id 42 email [email protected] last_seen 1723
EXPIRE user:v2:42 600
HGET user:v2:42 email
HMGET user:v2:42 email last_seen
HSET user:v2:42 last_seen 1999 # updates one field onlygo deeper
Know that Redis stores opaque bytes, that JSON is the common default, and that a hash lets you read and write individual fields.
Compare formats on CPU, size, and schema tolerance; explain when a hash beats a string and mention the listpack encoding for small hashes.
Add compression thresholds, the deserialization-security argument against native formats, big-value effects on the single-threaded server, and expiry granularity limits.
Set the org-wide policy: which formats are allowed on shared infrastructure, how payload versions are governed, and how serialization cost is budgeted against hit-ratio and memory targets.
## Redis stores bytes; the format is your problem A Redis string value is an opaque byte array up to 512 MB. Redis will not parse it, index it, or migrate it. Every decision about how the object becomes bytes — and back — is made in your client, on your CPU, in your request path. That matters more than people expect. On a hit, the Redis round trip inside a datacenter is often 0.2-0.5 ms, while deserializing a fat object with a reflection-based JSON mapper can be comparable or worse. A cache whose hit path is dominated by deserialization has a real bug even though every Redis metric looks perfect. ## The five axes **1. CPU per operation.** Deserialization runs on every hit; serialization runs on every miss. Binary formats (Protobuf, MessagePack, Avro, FlatBuffers) are typically several times faster than text JSON, and a fast JSON library is several times faster than a naive reflective one. Measure with your objects, not a benchmark blog. **2. Size.** Size is charged twice: RAM on the Redis node (the expensive resource in an in-memory store) and network bytes per hit. A 3x smaller encoding is 3x more cache in the same instance, which raises the hit ratio, which is usually a bigger win than the CPU difference. **3. Schema tolerance.** New code will meet values written by old code during any rolling deploy. Formats differ: Protobuf and Avro have explicit compatibility rules; JSON tolerates unknown fields if your mapper is configured to ignore them; language-native serialization tends to explode. The robust answer is not to rely on format tolerance at all but to put a **version segment in the key**, so a shape change writes to a different key entirely. **4. Cross-language access.** If a Kotlin service writes and a Node service reads, the format must be neutral. Java serialization or pickle immediately fails this test. **5. Operability.** When someone is debugging at 3 a.m., `GET key` returning readable JSON is worth real money. Opaque binary requires a decoder tool. Many teams accept JSON's overhead for exactly this reason and only switch where profiling justifies it. ## Compression Compress conditionally, above a size threshold — a few kilobytes is a common cutoff. Below it, framing overhead and CPU exceed the savings. LZ4 and Snappy give large throughput with modest ratios; zstd offers a tunable, generally better frontier; gzip is slow for a hot path. Store a small marker (a magic byte or a key-name segment) so readers know whether a value is compressed, and remember that compression makes values opaque to `GET` inspection. Do not compress values that are already compressed (images, most binary blobs) — you pay CPU for nothing. ## String or hash? Both are legitimate; the deciding question is **access granularity**. Use a **string** when the object is read and written as a whole. It is the simplest thing: one `SET key value EX ttl`, one `GET`, atomic by construction, works with `MGET` for batch reads, and lets you apply any serialization or compression you like. Use a **hash** when callers want individual fields. `HGET user:v1:42 email` fetches one field without transferring the object; `HSET user:v1:42 last_seen <ts>` updates one field without a read-modify-write cycle that could clobber a concurrent update to another field. `HMGET` fetches a chosen subset. For wide objects where callers need two of thirty fields, this is a large bandwidth and CPU saving. Memory is a secondary consideration and it favors small hashes: a hash under `hash-max-listpack-entries` (128 by default) and `hash-max-listpack-value` (64 bytes) is stored as a **listpack** — a compact, contiguous, sequentially-scanned representation — rather than a hash table with per-entry overhead. Above either threshold it converts to a real hash table and never converts back, and memory per field jumps. This encoding is also why grouping many small related values into one hash can use dramatically less memory than the same data as thousands of separate string keys, each of which carries its own key object and dictionary-entry overhead. The cost of hashes: **expiry granularity**. Historically TTLs apply to the whole key, not to fields, so you cannot age fields independently — Redis 7.4 added per-field expiration (`HEXPIRE` and friends), but be sure your deployment and client support it before relying on it. Hashes also do not participate in `MGET`; batching multiple hashes means pipelining `HGETALL`s. And `HGETALL` on a very large hash returns everything and is O(N), so wide hashes need `HMGET` discipline. ## Two things to avoid **Language-native serialization.** Java `Serializable`, Python `pickle`, .NET `BinaryFormatter`: slow, bloated, tightly coupled to class definitions so ordinary refactors break the cache, unreadable across languages, and — critically — deserializing untrusted bytes is a remote-code-execution class of vulnerability. If any path lets an attacker write to Redis (an exposed instance, a compromised service, an injection into a key), native deserialization turns a cache into an execution primitive. Use a data format with no code semantics. **Giant values.** Multi-megabyte values slow the single-threaded server on both read and write, blow up client output buffers, and make replication lumpy. If a value is very large, either cache a slimmer projection or split it — and reconsider whether the caching unit matches the consumption unit at all.
- Why avoid Java serialization or Python pickle for cache values?They are slow and bulky, they couple the stored bytes to exact class definitions so an ordinary refactor breaks every cached entry, and they are unreadable from other languages. The decisive reason is security: deserializing attacker-influenced bytes with these formats is a known remote-code-execution class, so anyone who can write to Redis can potentially execute code in your service.
- When is compressing cache values actually worth it?Above a size threshold — typically a few kilobytes — where the bytes saved in RAM and on the wire exceed the CPU spent on every hit and miss. Below that, framing and CPU overhead dominate. Use a fast codec such as LZ4, Snappy, or zstd rather than gzip on a hot path, mark compressed values so readers know how to decode, and never compress data that is already compressed.
- What do you give up by storing an object as a hash instead of a serialized string?Historically, per-field expiry: a TTL applies to the whole key, so fields cannot age independently unless you are on Redis 7.4 or later with HEXPIRE support. Hashes also cannot be batched with MGET — you pipeline HGETALL or HMGET instead — and HGETALL on a wide hash is O(N) and returns everything, so callers need HMGET discipline.
A string value is a sealed envelope you must open entirely to read one line; a hash is a labelled folder where you can pull a single sheet out.
saying these in an interview costs you the question
- Assuming Redis parses or understands the stored value in any way.
- Using language-native serialization (Serializable/pickle) for cache values.
- Compressing every value regardless of size, so small values cost more CPU than they save.
- Claiming hashes are always more memory-efficient without mentioning the listpack thresholds and the one-way conversion to a hash table.
- Believing a TTL can be set on an individual hash field in any Redis version.