skip to content

How do you design the key names for a Redis cache so you can change the shape of the stored value, or drop a whole category of entries, without scanning or flushing the keyspace?

level: seniorimportance: should knowfreq 40%

answer

  1. key naming = the schema
  2. app:entity:vN:id[:variants]
  3. bump vN to change payload shape
  4. generation counter + INCR = bulk drop
  5. never KEYS, never FLUSHDB as invalidation

basics

~20 s

Use structured, prefixed keys with an embedded version: app:entity:v3:id. Bumping the version makes every old key unreachable and it dies on its own TTL, so a deploy that changes the serialized shape can never read an old-format value. For bulk drops, embed a generation counter you can increment.

solid answer

~50 s

Key naming is a schema. I use a consistent, colon-delimited convention: `app:entity:version:identifier`, for example `svc:product:v3:915`, plus any variant dimensions that change the value (`locale`, `currency`, feature flag). The version segment does the heavy lifting. When the serialized shape changes — a field added, a serializer swapped, a computation fixed — you bump `v3` to `v4`. Old keys are instantly unreachable, so no deploy ever deserializes an old payload with new code, and no rollback reads new payloads with old code. The stale keys cost only memory and disappear at their TTL or via eviction. That is strictly better than a migration or a flush. For bulk invalidation of a category (one tenant, one catalog), embed an **indirection**: store a generation counter, e.g. `gen:tenant:42` = 7, and build keys as `svc:product:v3:t42:g7:915`. `INCR` on the counter orphans the whole family at once. What I avoid: `KEYS pattern` (O(N) on the main thread), `FLUSHDB` as an invalidation tool, and unbounded key cardinality from embedding raw user input.

code

text · 8 lines
text
# before the deploy, code builds keys with v3
SET catalog:product:v3:915 "{\"id\":915,\"price\":2999}" EX 600

# deploy adds a field / changes the serializer -> bump the constant to v4
SET catalog:product:v4:915 "{\"id\":915,\"price\":2999,\"currency\":\"EUR\"}" EX 600

# v3 keys are unreachable by the new code and expire on their own.
# a rollback still finds its v3 keys intact.

go deeper

for a junior

Know that keys should follow a documented convention like app:entity:v1:id, and that the version segment exists so a payload change does not break readers.

for a middle

Explain the rolling-deploy overlap that a version bump removes, and that every input affecting the value must appear in the key.

for a senior

Add the generation-counter indirection for bulk invalidation, why KEYS/FLUSHDB are non-answers, and the memory and cardinality consequences of key design.

for a principal

Treat the keyspace as a governed schema: ownership prefixes on shared instances, enforced key builders, cluster hash-tag policy, and the memory-versus-cold-start cost of version bumps.

## Keys are a schema, not strings Redis has no tables, no column types, and no migrations. The only structure in the keyspace is the structure you put in the names. Treating key naming as an afterthought produces caches nobody can safely change, because there is no way to reason about what a given key contains or who wrote it. A workable convention is hierarchical and delimited by colons, which is the community norm and what tooling (`redis-cli --scan --pattern`, key-space analyzers, `MEMORY USAGE` reports grouped by prefix) expects: ``` <app-or-service>:<entity>:<schema-version>:<identifier>[:<variant dimensions>] ``` Example: `catalog:product:v3:915:locale=de`. Each segment earns its place: - **service/app prefix** so a shared Redis instance can be attributed, monitored, and evacuated per owner. - **entity** so you know what the payload is. - **schema version** so the payload's shape is self-describing. - **identifier** the primary key. - **variant dimensions** every input that changes the value. If the cached value depends on locale, currency, A/B bucket, or permission scope, that input must appear in the key. Omitting one is a cache-poisoning bug: user A sees user B's variant. ## Why the version segment is the important one The hardest cache bug in practice is not staleness of *data* — TTLs bound that — it is staleness of *format*. You deploy code that adds a field, changes an enum, or switches serializers, and for the length of the longest TTL your new code is reading values written by the old code. Symptoms are deserialization exceptions in the best case and silently wrong defaults in the worst. Rolling deploys make it bidirectional: old and new instances write the same keys simultaneously. Bumping a version segment eliminates the whole class. `v3` and `v4` are different keys; new code writes and reads only `v4`; old code only sees `v3`. A rollback is equally safe because the old keys are still there. There is no migration step, no dual-read logic, and no flush. The cost is transient double memory and a cold-cache period for the affected entity — the same cost as a flush, but scoped to one entity family and without touching anything else. If the entity is large enough that a cold period is a problem, you can pre-warm the new version before flipping readers. Make the version a **constant in code next to the serializer**, so changing the payload type and forgetting the bump is visible in the same diff. Some teams derive it from a hash of the serialized class shape so it cannot be forgotten at all. ## Bulk invalidation without scanning "Invalidate everything for tenant 42" is a genuinely awkward operation in Redis. The naive approaches are all bad: - `KEYS tenant:42:*` is O(N) over the **entire** keyspace on the single main thread. On a large instance it is a multi-second stall — a self-inflicted outage. - `SCAN` with a pattern is incremental and safe for the server, but it still walks the entire keyspace to find a small subset, and it is only eventually complete under concurrent writes. - `FLUSHDB` throws away every other tenant's work as collateral damage. - Maintaining a Redis set of "all keys for tenant 42" works, but that index itself must be kept correct, expired, and memory-bounded, and it can become a large key of its own. The cheap trick is **indirection through a generation counter**. Keep `gen:tenant:42`, read it (or cache it in-process briefly), and build data keys that include its current value: `catalog:product:v3:t42:g7:915`. To invalidate the whole family, `INCR gen:tenant:42`. Every subsequent read composes a `g8` key, which misses and reloads; every `g7` key is orphaned and disappears at its TTL or under eviction. The cost is one extra lookup for the generation (amortizable with a short local TTL) and a period of double memory. The same trick powers the `v` segment — a version bump is just a generation bump with a human-chosen value. ## Practical hygiene - **Bound cardinality.** Never build keys from raw unbounded input (full query strings, free-text search terms) without hashing and capping; you will otherwise fill memory with keys read exactly once. Hash long components to a fixed-length digest to keep keys short. - **Keep keys short but readable.** Key names live in RAM, and at tens of millions of keys the difference between a 30-byte and an 80-byte name is real memory — but do not compress them into unreadable codes, because debugging a cache with opaque key names is miserable. - **Document the convention** in one place and enforce it in a key-builder helper rather than string concatenation scattered across the codebase. That helper is also where the version constant and the variant dimensions live. - **Cluster awareness:** hash tags `{...}` force co-location of keys in the same slot. Use them deliberately for keys you must read together with a multi-key command, and not by accident — putting the tenant id in braces for everything creates a hot slot. - **Namespace, not database index.** Do not rely on numbered logical databases (SELECT 1, 2, …) to separate concerns: Redis Cluster supports only database 0, and the numbered databases share the same instance, memory, and blocking behavior anyway.

  • What actually goes wrong if you skip the version segment and change the serialized value's shape?
    For the length of the longest TTL, new code reads values written by old code — and during a rolling deploy both versions write the same keys concurrently. The result is deserialization errors or, worse, silently wrong values where a missing field becomes a default. A version segment removes the overlap entirely and makes rollbacks safe, at the price of a transient cold cache for that entity.
  • Why not just use KEYS or SCAN with a prefix pattern to delete a category of entries?
    KEYS is O(N) over the whole keyspace and runs on Redis's single main thread, so on a large instance it stalls every other client for as long as it takes. SCAN is incremental and safe for the server but still walks the entire keyspace to find a small subset and gives only eventual completeness under concurrent writes. A generation counter turns the same operation into one INCR.
  • What has to be included in the key besides the entity id?
    Every input that changes the cached value: locale, currency, feature-flag or A/B bucket, permission or tenant scope, and any query parameters that alter the result. If such a dimension is omitted, one caller's variant is served to another — a cache-poisoning bug rather than a staleness bug. Keep the composition in a single key-builder helper so the rule is enforced in one place.

A version segment is like putting an edition number on a filing-cabinet label: the new edition goes in a new drawer, and the old drawer is simply never opened again until the cleaner empties it.

saying these in an interview costs you the question

  • Concatenating key names ad hoc across the codebase with no shared convention or builder.
  • Assuming a payload shape change is safe because 'the TTL is short'.
  • Reaching for KEYS pattern, or FLUSHDB, to invalidate a subset.
  • Leaving out a variant dimension (locale, tenant, permission scope) so users can be served each other's cached values.
  • Building keys from unbounded user input, exploding key cardinality and memory.

context