skip to content

For 20 million customer uploads across 5,000 tenants, how would you choose between one data key per object and one per tenant?

level: principalimportance: should knowfreq 30%

answer

  1. blast radius against call volume
  2. 4,000 objects behind a tenant key
  3. destroying a key is erasure
  4. rewrap cost scales with key count
  5. the wrapping key still covers everything

basics

~20 s

Granularity trades blast radius against volume. A key per object confines a leaked data key to one file and makes discarding that key an erasure; a key per tenant cuts wrapped keys and unwrap calls by a factor of 4,000 and makes one leaked key expose that tenant's entire set.

solid answer

~50 s

Both ends are defensible, and the choice is made by four numbers rather than a principle. **Blast radius**: with 20 million objects and 5,000 tenants, one data key per tenant averages 4,000 objects behind each key, so a single exposed data key exposes 4,000 objects instead of one. **Call volume and cache behaviour**: per-object keys mean an unwrap per cold read and a cache that almost never hits across a broad access pattern; per-tenant keys amortise both. **Erasure**: destroying a data key renders everything under it unreadable, so granularity *is* your erasure granularity — per-object keys let you erase one upload, per-tenant keys do not. **Rewrap cost**: a wrapping-key rotation is proportional to the number of data keys, so 20 million of them is a sweep and 5,000 is nothing. What granularity does not change is the level above: one compromised key-encryption key reaches everything wrapped under it either way.

go deeper

for a junior

The idea to hold: every object is encrypted under a data key, and how many objects share one data key is a choice. Fewer objects per key means a leaked key exposes less.

for a middle

Be able to state both costs: sharing a data key across a tenant's 4,000 objects widens what one leaked key exposes, while a key per object multiplies unwrap calls and the work a wrapping-key rotation has to sweep.

for a senior

Argue it with the read pattern and the rotation plan in hand — cache hit rate, unwrap rate under an export, and whether a 20-million-record rewrap sweep is something the store can run routinely and restart safely.

for a principal

Own the erasure requirement, because it is the constraint that cannot be retrofitted cheaply, and say plainly what the choice does not buy: nothing at the wrapping-key level, and nothing against a caller whose unwrap right spans the whole key.

## What the choice actually decides The number of data keys is not a performance knob with a security side effect; it is a security boundary with a cost side effect. Four properties move together with it: - **Blast radius of one exposed data key** — every object encrypted under it. - **Unwrap call volume and cache hit rate** — one unwrap per distinct key touched, not per read. - **Erasure granularity** — destroying a data key makes everything under it unreadable, permanently and cheaply, without rewriting any ciphertext. - **Rewrap cost when the wrapping key rotates** — proportional to the number of wrapped data keys, not to the volume of data. ## The two ends, measured on this estate 20,000,000 objects, 5,000 tenants — an average of 4,000 objects per tenant. | | key per object | key per tenant | key per tenant-day | |---|---|---|---| | data keys stored (wrapped) | 20,000,000 | 5,000 | 5,000 x days retained | | objects behind one data key | 1 | ~4,000 | ~4,000 / days retained | | unwrap calls for a 1,000-object scan | up to 1,000 | 1 | a handful | | cache hit rate on broad reads | near zero | high | high within a day | | erase one upload by destroying a key | yes | no | no | | erase one tenant by destroying keys | 4,000 destroys | 1 destroy | one per day retained | | rewrap sweep on outer rotation | 20,000,000 records | 5,000 records | thousands of records | The table makes the shape of the decision visible: the per-object column is better at exactly one thing — containment and erasure at the finest grain — and worse at everything operational, by roughly the 4,000x factor. ## How to actually decide 1. **Start from the erasure requirement, because it is the hardest to retrofit.** If a single upload must become unreadable on request without rewriting anything else, per-object keys are close to forced; coarser keys mean erasure has to be done by re-encrypting everything else under a new key, which at 4,000 objects per tenant is a real job to run on demand. 2. **Then check the exposure story.** Ask how a plaintext data key would escape at all — almost always from a process that had legitimately unwrapped it. If that process only ever handles one object at a time, per-object keys genuinely confine the damage. If it holds a tenant's whole working set in memory anyway, per-object keys are buying a boundary the runtime already crosses, and the containment is on paper. 3. **Then price the volume.** Model the read pattern, not the object count: skewed reads against a small working set make even per-object keys cheap, while a scanning or export workload against per-object keys means an unwrap per object and a manager sized for it. 4. **Then check the rotation you intend to run.** If the plan is to rotate the wrapping key quarterly, a 20-million-record rewrap sweep has to be something the store can do routinely and restartably, or the rotation will quietly stop happening. ## Where the middle usually lands Most estates end up between the two ends rather than at either. Common middles: one data key per tenant per time window, which bounds a leak to one window's worth of a tenant's uploads and keeps erasure possible at window granularity; one per sensitivity class within a tenant, so the small set that genuinely needs per-object treatment gets it while bulk uploads share; or per-object keys for objects above a size or classification threshold and shared keys below it. Each is defensible when the numbers behind it are stated. "Per object because finer is safer" is not a decision; it is a default wearing a justification. ## What granularity does not change Be explicit about this, because it is where the argument is often overstated: - **The level above.** Every data key here is wrapped under a key-encryption key. Compromise of that key material reaches everything wrapped under it whatever the data-key count, so granularity buys nothing against the top of the hierarchy — splitting the *wrapping* keys is the lever there, and it is a separate decision. - **The unwrap right.** A caller whose right covers the whole wrapping key can unwrap every data key under it. Fine data keys with a broad unwrap right give containment that any compromised caller walks straight through. - **Objects already copied.** Neither choice affects ciphertext or keys an attacker already took. The defensible answer in an interview is not a side; it is the four numbers, the erasure requirement, and an explicit statement of what the chosen granularity does not protect against.

  • A contract requires one named upload to become unreadable on request. Which granularity does that force?
    Effectively per-object, or at least a key no coarser than the unit you must erase. Destroying a data key makes everything under it unreadable at once, so with a per-tenant key the only route is re-encrypting that tenant's other 4,000 objects under a new key and then destroying the old one — a job you would have to run reliably, on demand, under live traffic.
  • Does moving from per-tenant to per-object data keys reduce the damage from a compromised key-encryption key?
    No. Every data key, however many there are, is wrapped under that key-encryption key, so compromising it opens all of them. Reducing damage at that level means splitting the wrapping keys themselves — by tenant, region or sensitivity — and scoping each caller's unwrap right to the subset it serves. Data-key granularity is a separate boundary.
  • What breaks first if you choose per-object keys for an export workload?
    The unwrap path. An export that walks a tenant's 4,000 objects needs 4,000 unwrap calls with almost no cache reuse, so the manager's rate limit and the per-call latency become the export's throughput ceiling. Either the export is batched into fewer, coarser keys by design, or the manager has to be sized for that call rate.

saying these in an interview costs you the question

  • Claims a key per object bounds the damage of a compromised key-encryption key
  • Says coarser data keys are always cheaper, so pick the coarsest
  • Thinks granularity affects only performance and not what can be erased
  • Assumes one key per object means the manager stores twenty million keys
  • Argues blast radius is set by the key manager rather than by how many objects share a data key
  • Picks per-object keys as a default without pricing the unwrap volume