A batch job that reads tens of millions of small SSE-KMS-encrypted objects from Amazon S3 starts failing with KMS throttling errors, although the same job ran fine at a tenth of the scale. What is causing it, and how do S3 Bucket Keys change the picture?
answer
- the bottleneck is not S3
- per object, not per byte
- small files are the aggravating factor
- a shared account-wide ceiling
- one short-lived key for many objects
basics
~20 sEach read of an SSE-KMS object costs a KMS decrypt request, so object throughput becomes KMS throughput and hits the account's per-Region KMS request quota. S3 Bucket Keys let S3 derive object keys from a short-lived bucket-level key, cutting KMS calls sharply.
solid answer
~50 sWithout S3 Bucket Keys, every operation on an SSE-KMS object puts a request to AWS KMS in the path — one to unwrap the data key on read, one to generate it on write. So a job doing 20,000 object GETs a second is also doing roughly 20,000 KMS requests a second, against a shared per-account, per-Region quota on KMS cryptographic operations. Small objects make this worse, because the KMS request is per object, not per byte: the same terabyte in a few large objects costs almost nothing. Enabling an **S3 Bucket Key** makes S3 obtain a short-lived bucket-level key from KMS and derive per-object data keys from it inside S3, so KMS is called a handful of times instead of once per object — AWS documents up to a 99% reduction in KMS requests and cost. Two caveats: it applies only to objects written after you enable it, and CloudTrail's KMS records then reference the bucket rather than each object.
code
bash · 11 linesaws s3api put-bucket-encryption \
--bucket my-bucket \
--server-side-encryption-configuration '{
"Rules": [{
"ApplyServerSideEncryptionByDefault": {
"SSEAlgorithm": "aws:kms",
"KMSMasterKeyID": "alias/reports-key"
},
"BucketKeyEnabled": true
}]
}'go deeper
Know that choosing SSE-KMS puts AWS KMS in the path of every object read and write, and that KMS requests are billed and rate-limited — unlike SSE-S3, which is free and involves no other service.
Explain that the KMS request is per object rather than per byte, that S3 Bucket Keys make S3 derive object keys from a short-lived bucket-level key, and that the setting applies only to objects written after it is enabled.
Diagnose it properly: recognise the failure tracks object count, check KMS throttling metrics and the objects' bucket-key status, then sequence the fixes — enable the bucket key, rewrite or compact the hot historical objects, and only then ask for a quota increase.
Own the account-level consequence: a shared KMS quota means one batch job can starve unrelated workloads, so decide which data classes justify KMS at all, where heavy data-plane traffic is isolated, and whether losing object-level key-use audit is acceptable in exchange.
## What one SSE-KMS object operation actually costs When an object is stored with SSE-KMS and no bucket key, its data key is individually wrapped by AWS KMS. That means: - **Write:** S3 asks KMS for a data key, encrypts the object with it, stores the wrapped copy. - **Read:** S3 asks KMS to unwrap that object's data key before it can decrypt. The crucial property is that this is **per object, not per byte**. A 5 GB object costs one KMS request. Five million 1 KB objects cost five million KMS requests. Workloads built on lots of small objects — event landing zones, page-per-file data lakes, model shards, image thumbnails — convert their object request rate one-for-one into a KMS request rate. ## Why scale turns this into an outage KMS enforces a shared quota on cryptographic operations per account per Region. The exact ceiling varies by Region and can be raised on request, but it is finite and it is shared across everything in the account that uses KMS, not just this bucket. A job that was comfortably under the ceiling at a tenth of the volume walks straight into it at full volume, and the failure surfaces as throttling — a KMS `ThrottlingException` propagated back through the S3 request, which the SDK retries with backoff until the job either slows to a crawl or gives up. Three things make this a classic senior diagnosis: 1. **It is not an S3 problem.** S3 request rates scale essentially without limit per prefix; the wall you hit is in a different service. 2. **It is shared blast radius.** Exhausting the KMS quota affects every other workload in that account and Region — an unrelated service's decrypt starts failing because your batch job saturated the quota. 3. **It scales with object count, not data size.** The instinct to look at bandwidth or object size sends you the wrong way. ## What an S3 Bucket Key does Enable a bucket key and S3 changes its relationship with KMS. Instead of one KMS call per object, S3 obtains a **short-lived, bucket-level key** from KMS and derives the per-object data keys from it internally, reusing that bucket-level key for a limited period across many objects. KMS is called on the order of a few times a period instead of once per object. AWS documents this as up to a **99% reduction** in KMS requests, with the cost falling by roughly the same factor — which is why the feature is usually introduced as a cost optimisation and only later appreciated as a quota and availability one. Enabling it is a property of the bucket's default encryption configuration, or a per-request flag: ``` aws s3api put-bucket-encryption --bucket my-bucket \ --server-side-encryption-configuration '{"Rules":[{"ApplyServerSideEncryptionByDefault":{"SSEAlgorithm":"aws:kms","KMSMasterKeyID":"alias/reports-key"},"BucketKeyEnabled":true}]}' ``` ## What enabling it costs you - **Only new objects benefit.** A bucket key is applied at write time. Objects written before you enabled it keep their individually wrapped data keys and still cost a KMS request on every read. If the throttling comes from reading a historical corpus, enabling the bucket key changes nothing until you rewrite that corpus. - **Audit granularity drops.** With per-object keys, the KMS encryption context identifies the specific object, so CloudTrail can tell you which object a principal decrypted. With a bucket key, the encryption context names the bucket, so you get bucket-level rather than object-level attribution. If an auditor is relying on object-level key-use records, that is a genuine trade to raise rather than a footnote. - **Not available everywhere.** Bucket keys are an SSE-KMS feature; dual-layer DSSE-KMS does not use them. ## The other levers, in the order you would reach for them 1. **Enable the bucket key** for new writes — the highest-leverage change and usually enough. 2. **Rewrite the hot historical objects** (an S3 Batch Operations copy) so they too are covered. 3. **Reduce object count** — compacting millions of tiny files into larger ones fixes the KMS request rate and the S3 request cost and the job's own overhead at the same time. 4. **Request a KMS quota increase**, which is legitimate but is a ceiling raise rather than a fix, and does nothing about the per-request bill. 5. **Reconsider the encryption choice for the bulk tier.** If a corpus does not need per-key control or auditable key use, SSE-S3 removes KMS from the path entirely at zero cost. ## Diagnosing it from cold Look for KMS throttling in the failing client's errors, then confirm at the source: KMS publishes request and throttling metrics in CloudWatch, and CloudTrail shows the volume and the calling principals. Check whether the affected objects carry a bucket key — HeadObject reports it — and check the bucket's default encryption configuration to see whether new writes are covered. The tell that separates a real diagnosis from a guess is noticing that the failure tracks **object count**, not gigabytes.
- You enable a bucket key on a bucket that already holds a hundred million SSE-KMS objects. Do the reads stop calling KMS?No. Bucket keys are applied when an object is written, so existing objects keep individually wrapped data keys and still cost a KMS request per read. Only new writes benefit. Covering the historical corpus means rewriting it — an S3 Batch Operations copy job — which is often the moment teams also compact small objects into larger ones.
- What do you give up in your audit trail by enabling S3 Bucket Keys?Object-level attribution of key use. Without a bucket key the KMS encryption context identifies the individual object, so CloudTrail records show which object a principal decrypted. With a bucket key the context names the bucket, so KMS records become bucket-scoped. If an auditor depends on per-object key-use evidence, raise that trade before enabling it fleet-wide.
- Why did the throttling affect an unrelated service in the same account?KMS request quotas are per account and per Region and are shared across all cryptographic operations, not scoped to one bucket or workload. A batch job that saturates the quota starves everything else using KMS in that account and Region. That shared blast radius is an argument for separating heavy data-plane workloads into their own account, not just for raising the quota.
saying these in an interview costs you the question
- Blames S3 request rates instead of the KMS quota
- Thinks KMS cost scales with object size rather than object count
- Assumes enabling a bucket key retroactively covers existing objects
- Says a quota increase is the complete fix
- Ignores that the KMS quota is shared account-wide