A team adds the signed-in user's id to a shared cache key on the server; what does that fix, and what does it not?
answer
- the key defines the audience
- completeness, not just identity
- still shared storage underneath
- entry count versus hit rate
- never key on a credential
basics
~20 sPutting the user id in the key stops one user's entry reaching another, but only if every other input that varies the output is there too. It does not make the store private, and it trades hit rate for entry count.
solid answer
~50 sThe key is what defines a shared copy's audience, so adding identity narrows that audience to one person - as long as the key is complete. If the output also varies by role, tenant, entitlement, locale, or a feature bucket, then two requests with the same user id but different values of those inputs collide, and you are back to serving the wrong bytes. What per-user keying does not change is that the copy still lives in shared storage: a key normalisation bug, a collision, or a code path that forgets the id exposes it, and the key itself is data that ends up in logs and tooling, so it must never be a session token. It also inverts the economics: entries scale with the user base, each is hit rarely, and eviction churns. Often the better answer is to keep the personalised part out of the shared copy entirely.
go deeper
Take away the core idea: a shared cache finds entries by key, so whatever is missing from the key cannot affect who gets served that entry.
Be able to list the inputs that belong in a key beyond identity, and explain the collision each omission produces in concrete terms.
Demonstrate that you weigh the hit rate before adopting per-user entries, keep authorization on the request, and treat key strings as data that leaks into logs.
Push the conversation to structure: argue for splitting the personalised holes out of a shared response so correctness rests on placement rather than on a key everyone must remember to build right.
## What the key is actually doing A copy that outlives the request has no idea who asked for it. It maps a **key** to bytes. Everything about who may be served that copy is therefore encoded in the key, and nowhere else. Adding the signed-in user's identifier to the key says: only requests that compute this same identifier may be answered from this entry. That is a real fix for cross-user bleed, and it is the mechanism behind every legitimate per-user cache. It is also the point where most of the subtle bugs in this area live, because a key is only as good as its completeness. ## The key must contain every input that varies the output Identity is rarely the only thing the render depends on. Think about which of these change what is produced: - **role or permission set** - an admin view and a member view of the same record; - **tenant or organisation** - the same person acting in two workspaces; - **entitlement or plan** - gated fields, quotas, upgrade prompts; - **locale, currency, timezone** - formatting and, sometimes, content; - **experiment or feature-flag bucket** - the whole point is that two users see different output; - **anything read from a header or cookie** that the render branches on. | If this varies the output but is missing from the key | The failure you get | | --- | --- | | Role | A member is served the admin version, or the reverse | | Tenant | Data from the wrong workspace under the right name | | Entitlement | A paid view served to a free account, or a downgrade that never takes effect | | Locale or currency | Wrong formatting and, in a shop, wrong prices | | Experiment bucket | Everyone converges on whichever variant was rendered first | The discipline that works is to derive the key from the same values the render reads, in one place, rather than assembling it by hand at each call site. A key built by hand drifts away from the render the first time someone adds a branch. ## What per-user keying does not fix 1. **It is still shared storage.** Privacy now rests on a string being computed correctly on every path. A normalisation slip (trimming, case, a missing segment), a truncated key, or a code path that builds the key without the id puts personal bytes where a stranger's request can reach them. Compare that with a request-bound read, which is private because of *where it lives* rather than because of a value someone remembered to include. 2. **It is not an authorization check.** A key decides which entry is returned; it does not decide whether the caller is allowed to see the data. Authorization still has to run on the request. A cached copy that skips the check because "the key is per user" is an access-control hole waiting for a key bug. 3. **It does not track permission changes.** If access is revoked, an entry produced while access was still granted keeps existing under a key the same person still computes. Making a copy fresh again is a separate mechanism from placing it. 4. **The key is data.** Keys show up in logs, metrics labels, traces, and admin tooling. That makes a session token, an authentication cookie value, or an email address a poor key. Use a stable opaque identifier. 5. **The economics invert.** Entries now scale with the number of users rather than the number of distinct pages. Each entry is read far less often, so the hit rate collapses, memory or storage grows, and eviction starts throwing entries away before they are reused. A per-user cache with a five percent hit rate is pure cost - you pay the storage and the complexity and still render almost every request. ## When per-user keying is nevertheless right It earns its place when the read is genuinely expensive, the same user hits it repeatedly within a short window, and the value is stable over that window - an assembled dashboard, an expensive permission expansion, a report. In those cases the entry is reused enough to pay for itself. A useful sanity test before adopting it: estimate reads per entry per lifetime. Below roughly two, you are building a write-only cache. ## The usually better shape Most pages are not personalised end to end. They are a large shared surface - navigation, catalogue, article body, layout - with a few personal holes: a name, a cart count, an entitlement banner. Splitting on that line lets the shared part sit in a copy keyed by URL alone, with high reuse and no identity anywhere near it, while the personal parts are rendered per request or filled in from the browser. You get most of the cache benefit, and the personalised bytes never enter shared storage at all, so there is no key to get wrong. Frameworks differ in how directly they support that split - some provide boundaries that let part of a response be reused while another part is produced per request, others make it an all-or-nothing property of the route - but the reasoning is the same everywhere: prefer removing personal data from the shared copy over encoding identity into its key.
- Besides the user id, what else belongs in a per-user key?Every input the render branches on: role or permission set, tenant, entitlement or plan, locale and currency, and any experiment bucket. If an input changes the output but is absent from the key, two requests that differ on it collide and one of them is served the other's bytes. Deriving the key from the same values the render reads keeps the two in step.
- Why is a session token a bad component of a cache key?Keys travel: they appear in logs, metrics labels, traces, and admin tooling, so keying on a token spreads a live credential into places that are not treated as secret. It also churns - a rotated token strands the old entry and starts from cold. Use a stable opaque user identifier instead.
- Does a per-user key mean you can skip the authorization check on a cache hit?No. The key selects an entry; it does not establish that the caller is entitled to the data. Authorization belongs on the request, before the response is returned, so a key bug degrades to a cache miss rather than to an exposure. It also matters because access can be revoked after the entry was produced.
saying these in an interview costs you the question
- Believes a user id in the key makes shared storage private
- Leaves role, tenant or locale out of the key
- Uses a session token or raw cookie value as the key
- Treats a cache key as an authorization control
- Ignores what per-user entries do to hit rate and eviction
- Never considers splitting the personal part out of the shared copy