skip to content

Every object read costs an unwrap call to the key manager, so a service caches unwrapped data keys for five minutes — what does that cache cost?

level: seniorimportance: should knowfreq 40%

answer

  1. one manager call per read
  2. the cache removes the call
  3. and removes the check
  4. revocation lags by the entry lifetime
  5. trail records the unwrap, not the reads

basics

~20 s

The cache removes a manager call per read, and with it the manager's chance to refuse that read and its record of it. While an entry lives, withdrawing the service's right to unwrap changes nothing for those objects, and the decryption never appears in the manager's trail.

solid answer

~50 s

Without a cache, every read is: fetch the record, hand the wrapped data key to the manager, get the plaintext data key, decrypt. That is one round trip and one authorization decision per read, and a hard dependency — if the manager is unreachable, reads stop. Caching the unwrapped data key turns a cache hit into a purely local operation. What you give up is precisely what the call was doing: the manager no longer decides whether this caller may still unwrap, and no longer records that the object was decrypted. A five-minute entry means a revoked unwrap right keeps working for up to five minutes **per process**, and an investigator later asking which identity decrypted an object sees a gap. The plaintext data key is also now resident in the service's memory rather than only inside the manager. The bound is a policy choice: shorter entries narrow both windows and raise call volume proportionally.

code

pseudocode · 12 lines
pseudocode
readObject(objectId):
    record = metadata.get(objectId)
    entry  = dataKeyCache.get(objectId)

    if entry != null and now < entry.expiresAt:
        dataKey = entry.dataKey          # no manager call: not authorised, not recorded
    else:
        dataKey = manager.unwrap(record.wrappedDataKey)   # authorised and recorded here
        dataKeyCache.put(objectId, { dataKey: dataKey,
                                     expiresAt: now + 5 minutes })

    return decrypt(read(record.ciphertextLocation), dataKey)

go deeper

for a junior

The shape to remember: reading an encrypted object normally means asking the key manager to unwrap that object's data key first. Caching the unwrapped key skips that step on later reads.

for a middle

Explain what the skipped call was doing — an authorization decision and an audit record, not just latency — and be able to say that the cached data key stays correct across a wrapping-key rotation.

for a senior

Reason about the windows: revocation lag equal to the entry lifetime in every process, gaps in the decryption trail, and an explicit fleet-wide drop as the only faster lever. Size the lifetime against a containment target, not a round number.

for a principal

Separate the two objectives. Latency caching and surviving a manager outage want very different lifetimes and accept very different exposure; pick one deliberately, write down the window it concedes, and define what the service does when an entry expires during an outage.

## The read path, with and without the cache Uncached, a read of an encrypted object is four steps: load the metadata record, send the wrapped data key to the manager, receive the plaintext data key, decrypt the object's bytes locally. The bulk data never goes near the manager — that is the point of the envelope — but the *decision* does, once per read. Cached, a hit skips steps two and three entirely. The service already holds the plaintext data key for that object and decrypts locally. The manager is not contacted, so it neither authorises nor records anything. That is the whole trade, and it is worth stating as a symmetry rather than as a performance tip: **a cache hit is a read the manager did not see.** ## What the cache buys - **Latency.** One network round trip leaves the hot path. - **Throughput and cost.** At millions of reads an hour, an unwrap per read is a serious call volume, and managers meter and rate-limit. - **Partial independence.** A brief manager outage no longer stops reads of recently-touched objects — though only those, and only for the entry lifetime. ## What the cache silently gives up | property | uncached | cached, five-minute entries | |---|---|---| | authorization checked | on every read | on the first read of an entry | | decryption recorded | every read | first read of an entry only | | effect of withdrawing unwrap rights | immediate | up to five minutes, per process | | plaintext data key resides | inside the manager | also in the service's memory | Four consequences follow, and each is a real operational surprise: 1. **Revocation lags.** If the service is compromised and its unwrap right is withdrawn, it keeps decrypting cached objects until the entries expire. The window is the entry lifetime, and it repeats independently in every process holding a cache — a fleet of 200 processes does not revoke faster than one, it simply has 200 copies of the same lag. 2. **The audit trail develops holes.** The manager's record shows the unwrap that populated the entry, not the reads served from it. "Which identity decrypted this object, and when" is answerable for the first read and not for the next thousand. If that question matters, the decryption event has to be recorded by the service itself, and that record is only as trustworthy as the service. 3. **Key material spreads.** Plaintext data keys now live in ordinary process memory across the fleet, which is a different custody story from the one the envelope model started with. 4. **What the cache does *not* do** is also worth being exact about: it does not go stale when the wrapping key is rotated. Rotation changes the wrapped blob, not the data key, so a cached data key remains correct. The staleness that matters here is about *permission*, not about correctness. ## Choosing the bound honestly The lifetime is a direct trade between call volume and revocation lag, and it should be argued in those terms: - Compute the call volume the uncached path actually implies, and compare it against what the manager will serve. If the uncached rate is affordable, do not cache — this is more often true than teams assume, because reads are usually skewed to a small working set anyway. - If caching is needed, size the entry lifetime against how quickly you must be able to cut a compromised consumer off. Five minutes is a defensible answer if "five minutes" is an acceptable containment time; it is not defensible as a number nobody chose. - Cache **per object**, not per wrapping key, and never cache anything that would let a process unwrap a key it has not already legitimately read. - Add an explicit, cheap way to drop the cache fleet-wide, and test it. Without that, the entry lifetime is your only containment lever. ## The dependency question underneath There is a second reason teams reach for this cache: they want reads to survive a manager outage. That is a different objective and deserves its own decision, because a cache sized for latency gives almost no outage coverage, while a cache sized for outage coverage means holding plaintext key material for hours. Do not let a latency cache be quietly reclassified as an availability control — decide which one you are building, state the exposure window that choice accepts, and write down what the service does when an entry expires and the manager is still unreachable.

  • The service's unwrap right is withdrawn at 10:00 across a fleet of 200 processes. When does the last cached read happen?
    Up to five minutes after the last cache population that preceded the withdrawal, so by about 10:05 in the worst case — and each process runs its own clock, so the fleet does not converge any faster than a single process. Only an explicit fleet-wide cache drop shortens it; withdrawing the right at the manager does not reach memory the manager cannot see.
  • An investigator asks which identity decrypted one specific object yesterday afternoon. What can the manager's record answer?
    Only the unwraps: which identity asked for that object's data key and when. Reads served from a live cache entry produced no call, so they are absent. Answering the full question requires the service to log its own decryptions, which is a weaker record because the same compromise that abused the key can edit or silence it.
  • Does rotating the wrapping key invalidate cached data keys?
    No. A rotation replaces the wrapped blob stored beside each object and leaves the data key itself unchanged, so a cached entry stays correct and keeps decrypting. Nothing about outer-key rotation is a cache-invalidation event — the only levers over a cache are its lifetime and an explicit drop.

saying these in an interview costs you the question

  • Thinks a cached data key still passes through the manager's authorization on each read
  • Says withdrawing the unwrap right stops reads immediately
  • Claims the manager's trail shows every object decryption regardless of caching
  • Argues a longer entry lifetime costs nothing because the data key has not changed
  • Expects rotating the wrapping key to flush cached data keys
  • Treats a latency cache as an outage-survival control without sizing it for one