skip to content

What must the key of an application-cached transfer model carry beyond the record's identifier?

level: seniorimportance: should knowfreq 52%

answer

  1. the key is the argument list
  2. what was read but never named
  3. tenant, scope, parameters, shape
  4. a wrong hit looks like a right one
  5. redact after the hit to keep sharing

basics

~20 s

Every input that changed the stored answer: the tenant, the viewer's permission scope if the model is filtered or redacted per viewer, the query parameters, and a version for the model's shape. An identifier alone serves one caller's answer to another.

solid answer

~50 s

The key must name **everything the assembled answer depended on**, not just the row it came from. In practice that is the tenant, the permission scope of the viewer when the model is filtered or redacted per viewer, all query parameters (filters, sort, page, locale), and a version segment for the model's own shape so a redeploy that changes fields misses rather than mis-reads old entries. Anything omitted becomes a collision: a key of just the identifier hands one tenant's data to another, or an administrator's expanded view to an ordinary user. Where putting the actor in the key would shatter the hit rate, the alternative is to cache the **unredacted** model and apply the filtering after the hit — keeping authorization in code that runs every request, rather than baking one viewer's verdict into a shared entry.

go deeper

for a junior

Learn the rule: whatever the answer depended on must appear in the key. Identifier alone is almost never enough once tenants or filters exist.

for a middle

Enumerate the inputs concretely — tenant, permission scope, every query parameter, the model's shape version — and say what a collision returns in each case.

for a senior

Show the defences, not just the list: key-building helpers, per-tenant namespaces, a tenant check on the deserialized value, and a test that reads across tenants.

for a principal

Own the trade between safety and hit rate: per-actor keys are safe and nearly useless at scale, so decide when to cache unredacted and filter after the hit, and where that filter is enforced.

## The key is the function signature The cleanest way to think about it: the cached value is the return value of a function, and the key has to be the function's full argument list. Anything the assembly *read* but the key does not *name* is a hidden argument, and hidden arguments make different calls collide on one entry. For a transfer model assembled in a service, the real arguments are usually more than one identifier: - **The tenant or account boundary.** Almost always implicit, taken from ambient request context rather than passed as a parameter — which is exactly why it is the one most often left out of the key. - **The viewer's permission scope**, whenever the model is filtered, trimmed or redacted according to who is asking. - **Every query parameter**: filters, sort order, page number and size, the date window, the locale or currency the labels were rendered in. - **The shape of the model itself**, as a version segment, because the code that assembles it is redeployed independently of the data. ## What omission actually costs | Omitted from the key | What a hit returns | Severity | |---|---|---| | Tenant | Another tenant's data | A data breach, not a bug | | Permission scope | A privileged view served to an ordinary viewer | Privilege escalation | | A filter or page parameter | The wrong slice, confidently | Visible wrongness | | Locale or currency | The right data, mislabelled | Subtle and long-lived | | Shape version | Fields missing, misread, or a deserialization failure after a deploy | Breaks at release time | The first two rows are the reason this is a security topic and not a performance one. A cache with an under-specified key does not fail loudly — it returns a well-formed, plausible object belonging to somebody else, and it does so faster than the correct answer would have arrived. ## Tenant: make it structural, not optional A tenant segment that each call site remembers to add will eventually be forgotten by one call site. The defences that actually hold are structural: 1. Build keys through a single helper that takes the tenant from the same context the data access uses, so a key cannot be constructed without one. 2. Give each tenant its own namespace or key prefix in the store, so a missing segment produces a miss rather than a foreign hit. 3. Assert in the read path that the deserialized model's own tenant field matches the caller's, and treat a mismatch as a hard error. A cheap equality check turns a silent breach into an alert. ## Actor: in the key, or after the hit Putting the viewer in the key is correct but expensive: the entry is no longer shared, and for a large audience the hit rate approaches zero while the store fills with near-duplicates. Two designs avoid that, and choosing between them is the judgment being tested: - **Cache by scope, not by person.** If viewers fall into a small number of permission classes, key on the class. The trap is fail-open behaviour: an unrecognised scope must not fall back to the widest entry — it must miss. - **Cache the unredacted answer, redact after the hit.** The entry is shared by everyone, and the filtering runs per request in code you can review, test and audit. It costs a little work on every hit, and it demands that the unredacted model never escapes the redaction step — so the redaction belongs at the boundary the value must pass through, not somewhere a caller can bypass. The second is usually preferable when the audience is large and the redaction is cheap; the first when the model differs so much per class that redaction would mean reassembly. ## Shape version: the one nobody sees coming Mapper-level tiers cache rows whose shape the schema defines, and schema changes go through migrations. An application entry caches a serialized model whose shape is defined by ordinary code. Deploy a version that adds a field, and old entries — written minutes ago by the previous version — deserialize into a model missing it, if they deserialize at all. Worse, during a rolling deploy both versions are writing entries under the same keys. A version segment in the key solves it completely and cheaply: the new code reads and writes a different key space, the old entries are simply never read again, and expiry cleans them up. The alternative — teaching the deserializer to tolerate both shapes — is compatibility code that has to be written, tested, and then remembered and removed later.

  • Why is the tenant segment the one most often missing from a key?
    Because it is rarely a parameter. It arrives in ambient context and is applied automatically by the data access layer, so the developer writing the key never types it. Constructing keys through a helper that reads the same context, and namespacing the store per tenant, removes the chance to forget.
  • How would you detect that a key is under-specified before it leaks data?
    Store the discriminating fields inside the value and check them on read — assert the entry's tenant and scope match the caller's, and fail loudly on mismatch. Add a test that fills the cache as one tenant and reads as another, expecting a miss; that test catches the whole class.
  • When is putting the actor in the key the right call despite the hit rate?
    When the model genuinely differs per person rather than per class — a personalised feed, per-user pricing — so no shared unredacted form exists to filter. Then accept the low sharing, keep the entries small and short-lived, and reconsider whether caching pays at all.

saying these in an interview costs you the question

  • Keys by record identifier alone in a multi-tenant service
  • Assumes separate store credentials keep tenants' keys apart
  • Caches a per-viewer redacted model under a shared key
  • Falls back to the widest cached view for an unrecognised scope
  • Leaves the model's shape out of the key across deploys
  • Omits paging or filter parameters because the entry is short-lived