For a multi-tenant platform on AWS, how would you decide KMS key granularity — one key for the account, one per service, or one per tenant — and what forces the answer?
answer
- count the keys, then multiply the bill
- a key is the unit of destruction
- erasure forces isolation, authorization does not
- one key plus per-tenant encryption context
- request quota is shared per Region
basics
~20 sGranularity follows the blast radius and erasure requirements, then is checked against cost and quota. Per-tenant keys buy isolation and cryptographic erasure; a shared key with per-tenant encryption context and grants buys the same authorization boundary far more cheaply.
solid answer
~50 sI start from what the key boundary has to guarantee. A key is the unit of revocation and of destruction, so if a contract says "delete this customer's data provably", that customer needs its own key — you cannot cryptographically erase one tenant out of a shared one. If the requirement is only "tenant A's worker must never decrypt tenant B's data", a shared key with per-tenant encryption context and per-tenant grants gives the same authorization boundary without minting thousands of keys. Then I check the constraints: each customer managed key carries a monthly charge per Region and per multi-Region replica, so per-tenant keys scale the bill linearly with tenants; key policies have a size limit so you cannot enumerate tenants in one document; and KMS request-rate quotas are shared per Region, which pushes you to data-key caching regardless. My default is per-service keys, with per-tenant keys as an opt-in tier for customers who require them.
go deeper
Recognise that a KMS key is a boundary, not just a setting, and that more keys means more cost and more administration. Knowing the question is a tradeoff rather than a best practice is the expected answer here.
Be able to contrast the two mechanisms: separate keys give separate policies and separate deletion, while a shared key with encryption context gives per-tenant authorization and audit without multiplying keys.
Show the constraint checks — per-key monthly charge across Regions, key policy size, the shared per-Region request quota, grant quotas — and argue why data-key caching is the throughput lever rather than key count.
Own the rule that decides it: the erasure and isolation contracts set the minimum key count, ownership boundaries set the natural count, and encryption context covers everything in between. Be ready to price a dedicated-key tier and to defend it against a reviewer who equates more keys with more security.
## Start from what a key boundary actually is A KMS key is three boundaries at once, and conflating them is why this decision goes wrong: 1. **An authorization boundary** — who may call `Decrypt` or `GenerateDataKey`, expressed in the key policy, IAM, and grants. 2. **A destruction boundary** — scheduling a key for deletion makes every data key wrapped under it unrecoverable, permanently, after a 7-to-30-day waiting period. This is the only mechanism for cryptographic erasure. 3. **A blast-radius boundary** — a mistake in one key's policy, or a compromise of a principal allowed on it, exposes exactly the data wrapped under that key. Only the second of these genuinely requires more keys. The first can be achieved on a shared key with condition keys; the third is a matter of degree. ## The requirement that forces per-tenant keys If a contract or regulation demands provable, irreversible deletion of one customer's data — or the customer supplies and controls its own key material — that customer needs a dedicated key. There is no partial erasure of a shared key: you cannot delete one tenant's backing material out of it. This is the one requirement I would never try to engineer around. Everything else usually can be. Per-tenant *authorization* is achievable on a shared key: encrypt each tenant's data keys with an encryption context naming the tenant, then either condition `kms:Decrypt` on `kms:EncryptionContext:tenant` in the key policy, or issue each tenant worker a grant carrying an `EncryptionContextEquals` constraint. A decrypt attempt with the wrong tenant's blob fails on the binding even before authorization is considered, and CloudTrail records which tenant each unwrap was for. ## The constraints that push the other way **Cost.** A customer managed key carries a per-key monthly charge in each Region, and a multi-Region replica is a separately charged key. Ten thousand tenants across two Regions is twenty thousand key-months. That is a real number in a design review, and it buys nothing if the requirement was only authorization. **Policy size.** A key policy is a bounded JSON document. You cannot enumerate thousands of tenants in one, which means a shared key must express tenancy through condition keys, policy variables, or grants — not through statements per tenant. Conversely, per-tenant keys mean thousands of small policies to author, review, and drift-check. **Quota.** KMS enforces a shared per-Region request-rate quota across cryptographic operations. It is shared across your keys, so splitting into more keys does not buy throughput. The lever that does is envelope encryption with bounded data-key caching, and that decision is independent of granularity. **Operational load.** Every key needs rotation configuration, monitoring, an owner, deletion protection, and a story for what happens when it is disabled. Multiply that by tenant count and you have built a key-management product nobody asked for. Grants have a per-key quota too, so a design that issues a grant per tenant on one shared key has its own ceiling to respect. ## The shape I would actually propose - **A key per service and per environment** as the baseline. It aligns the blast radius with an ownership boundary that already exists, keeps the policy readable, and keeps the count in the tens. - **Encryption context on every operation**, naming tenant and record type, with `kms:EncryptionContextKeys` enforcing that the fields are present. This gives per-tenant authorization and per-tenant audit on shared keys. - **A dedicated key as a product tier** for customers who contractually require isolation or erasure — priced accordingly, because it genuinely costs more to run. - **Multi-Region keys only where data actually crosses Regions**, since each replica is another billed key and another policy to keep in sync. - **Deletion protection and alarms**: `ScheduleKeyDeletion` is the most dangerous call in this whole area, so alarm on it in CloudTrail and keep the waiting period long enough for a human to notice. ## How I would justify it in a review The honest framing is that key count is not a security metric. A thousand keys with a permissive policy each is weaker than one key with a per-tenant condition and full CloudTrail attribution. What I would commit to is: the erasure requirement determines the minimum number of keys; the ownership boundary determines the natural number; encryption context supplies the granularity in between; and the bill and the quota are the checks that stop the design running away.
- A customer demands proof that their data is unreadable after offboarding. What is the minimum design that satisfies it?Their data must be wrapped under a key nothing else uses, so that scheduling that key for deletion destroys the only path to it. Then the offboarding runbook is: stop writes, schedule deletion with a waiting period, and keep the CloudTrail record of the deletion as the evidence. A shared key cannot produce this, however the data is partitioned.
- Does splitting into more KMS keys give you more KMS throughput?No. The cryptographic-operation request-rate quota is enforced per account and Region across keys, so more keys do not raise the ceiling. Throughput comes from making fewer KMS calls — envelope encryption with a bounded data-key cache — and, beyond that, from a quota increase request. Designing key count for performance reasons is a category error.
- Why not express per-tenant access with one key-policy statement per tenant?A key policy is a bounded document, so it stops scaling after a few dozen tenants and becomes a change-managed hotspot every time one is added. Condition keys on the encryption context, or a grant per tenant with an EncryptionContextEquals constraint, express the same rule without editing a shared document — though grants have their own per-key quota to respect.
saying these in an interview costs you the question
- Treats key count as a security metric in itself
- Claims per-tenant keys increase KMS throughput
- Believes one tenant can be erased from a shared key
- Ignores the per-key monthly charge across Regions
- Plans to list every tenant in a single key policy