A DynamoDB table provisioned at 10,000 WCU is consuming about 2,000 WCU on average, yet writes are being throttled. What is going on, how do you confirm it, and what are your options?
answer
- the table total is not one pipe
- hash routing decides who gets hit
- aggregate looks fine, one shard does not
- metrics first, then the offending key
- spreading beats provisioning
basics
~20 sTable capacity is spread across partitions, so a skewed workload can exhaust one partition's share while the table looks idle. Confirm with throttle metrics and Contributor Insights to find the hot key, then relieve it with caching, write sharding or a key change.
solid answer
~50 sProvisioned throughput is a table-level budget, but it is served by many physical partitions, and each partition has its own hard ceiling on reads and writes per second. If traffic concentrates on one partition key — a single popular tenant, a counter, a date-stamped key everything writes to today — that partition saturates while the table's aggregate consumption stays far below the provisioned number. Adaptive capacity mitigates this automatically by shifting unused throughput toward the busy partition and, for sustained heat, isolating the hot item onto its own partition, but it works in the background and cannot exceed the per-partition ceiling. Confirm the diagnosis with the `WriteThrottleEvents` and `ThrottledRequests` CloudWatch metrics against `ConsumedWriteCapacityUnits`, then enable CloudWatch Contributor Insights for DynamoDB to see which keys dominate. The durable fix is spreading the writes; caching, buffering through a queue, or an on-demand switch buy time but do not remove the skew.
code
bash · 5 linesaws dynamodb update-contributor-insights \
--table-name Orders \
--contributor-insights-action ENABLE
aws dynamodb describe-contributor-insights --table-name Ordersgo deeper
Know that DynamoDB spreads a table across partitions and that uneven traffic can throttle one of them even when the table as a whole is far below its provisioned capacity.
Be able to name the metrics that prove it — throttle events alongside consumed versus provisioned capacity — and to explain that a per-partition ceiling exists independently of the table total.
Walk the full diagnosis: metrics, then Contributor Insights to name the key, then a ranked set of mitigations, and say clearly which of them are stopgaps and which change the traffic shape.
Own the prevention side: how tables are reviewed for key skew before launch, what throttle alarms exist by default, and how multi-tenant workloads are kept from letting one tenant's growth become everyone's incident.
## The mental model that resolves the paradox A DynamoDB table is not one thing that runs at 10,000 WCU. It is a set of physical **partitions**, and the table's throughput is divided among them. A request is routed to a partition by hashing its partition key, and that partition alone serves it. Aggregate capacity is therefore an upper bound that you can only reach if traffic is spread evenly; skewed traffic hits a per-partition wall long before the table-level number is touched. Each partition has a fixed ceiling — as of 2025 AWS documents 3,000 read capacity units and 1,000 write capacity units per partition per second, and a partition holds up to 10 GB of data. Those numbers are the source of the paradox: 1,000 WCU of concentrated writes will throttle no matter how large the table's provisioned total is, because no single partition can be given more than its ceiling. ## What adaptive capacity already does for you AWS does not leave this entirely to you. **Adaptive capacity** is always on and does two things. First, it lends unused throughput from cool partitions to a busy one, so uneven-but-not-extreme traffic is absorbed silently. Second, for sustained imbalance it will **isolate frequently accessed items**, splitting a hot partition so the heavy key gets a partition of its own. Both are background behaviours with real limits. They react over time rather than instantly, they cannot lift a partition above its hard ceiling, and item isolation cannot help when the heat is on a *single* item — one key writing at 2,000 WCU has nowhere to be split to. So adaptive capacity turns many hot-partition problems into non-events, and leaves the extreme ones for you. ## Confirming it rather than guessing The diagnosis is a comparison, not a single metric: - `ConsumedWriteCapacityUnits` well below `ProvisionedWriteCapacityUnits` — the table looks idle. - `WriteThrottleEvents` (or `ReadThrottleEvents`) non-zero, and `ThrottledRequests` climbing — requests are being rejected anyway. - Client-side, `ProvisionedThroughputExceededException` in application logs, often visible first as latency because the AWS SDKs retry throttles with exponential backoff before surfacing them. That combination is essentially diagnostic of skew. Note the `ThrottledRequests` metric also has an `Operation` dimension, which tells you whether it is the base table or a secondary index being throttled — an overloaded index is a distinct and easily missed cause, since a write to the base table also consumes capacity on every index it updates. To name the offending key, enable **CloudWatch Contributor Insights for DynamoDB**, which publishes rules for the most-accessed and most-throttled keys: ```bash aws dynamodb update-contributor-insights \ --table-name Orders \ --contributor-insights-action ENABLE ``` The resulting graphs rank partition keys by request count and by throttle count over time, turning "something is hot" into "tenant 42 is 80% of writes". ## The options, from tactical to structural **Absorb it.** Switching the table to on-demand does not raise the per-partition ceiling, but it removes the table-level provisioning question and lets the service scale partitions more aggressively; it is a reasonable emergency lever, not a fix. **Buffer it.** Put the writes behind a queue and drain at a controlled rate, or aggregate in the application — a counter that is incremented a thousand times a second can often be accumulated in memory and flushed once a second, cutting write traffic by three orders of magnitude with an accuracy tradeoff you can usually accept. **Cache the reads.** For read-side heat, an in-front cache removes repeated reads of the same item entirely. **Spread the writes.** The structural answer is that the traffic should not all land on one key. The standard technique is write sharding — appending a bounded suffix to the partition key so writes fan out across several keys, at the cost of having to read all shards to reassemble the whole. Choosing what the key should be is data-modelling work and belongs with the design of the table rather than with its operation, but you should be able to say plainly in an interview that the durable fix lives there, and that no amount of provisioning tuning substitutes for it. ## What to say Lead with the mechanism — table capacity is divided across partitions, each with its own ceiling — then the confirmation path (throttle metrics against consumed capacity, then Contributor Insights for the key), then the ladder of mitigations, ending honestly at "the real fix is that the traffic has to stop being concentrated". Mentioning adaptive capacity unprompted signals that you know AWS already solved the easy half of this problem.
- Would switching this table to on-demand capacity make the throttling go away?Not reliably. On-demand removes the table-level provisioning ceiling and lets DynamoDB scale partitions more freely, so mild skew often stops hurting. But the per-partition throughput ceiling applies in on-demand mode too, so a single key taking thousands of writes per second still throttles. Treat it as an emergency lever that buys time, not a cure for concentrated traffic.
- Throttling shows up on writes to a table whose base-table consumption looks healthy. What else consumes write capacity?Every global secondary index. A write to the base table also writes to each index whose attributes it touches, consuming that index's capacity, and an under-provisioned index throttles the base-table write. Check the ThrottledRequests and consumed-capacity metrics per index, not just per table — an index sized for old traffic is a common invisible bottleneck.
- How does adaptive capacity differ from DynamoDB auto scaling?They operate at different layers. Auto scaling changes the table's provisioned RCU and WCU values over minutes, based on CloudWatch metrics, and is something you configure. Adaptive capacity is always on inside the service: it redistributes a table's existing throughput toward busy partitions and can isolate a persistently hot item onto its own partition. Neither can exceed the per-partition ceiling.
saying these in an interview costs you the question
- Throttling can't happen while consumed capacity is below provisioned
- Adaptive capacity means hot partitions no longer exist
- Just raise the provisioned WCU until it stops throttling
- Throttled requests always surface immediately as errors
- Only the base table consumes write capacity, not its indexes