skip to content

When would you run a Kinesis Data Streams stream in on-demand capacity mode instead of provisioned, and what do you give up by doing so?

level: middleimportance: should knowfreq 52%

answer

  1. who owns the shard count
  2. billed per shard-hour vs per GB
  3. elastic, but with a warm-up
  4. skew survives either mode

basics

~10 s

On-demand suits unpredictable or new workloads: AWS manages shard count and you pay per stream-hour plus data volume. Provisioned is cheaper at steady, well-understood throughput but makes you size and reshard the stream yourself.

solid answer

~60 s

The two modes differ in who owns the shard count and how you are billed. In **provisioned** mode you declare the number of shards, pay per shard-hour plus per PUT payload unit, and you are responsible for resharding — `UpdateShardCount`, `SplitShard`, `MergeShards` — when traffic changes. In **on-demand** mode AWS scales the shards for you based on observed traffic and bills per stream-hour plus per GB written and read. I reach for on-demand when the traffic pattern is unknown, spiky, or new, and for anything I do not want to run a scaling loop for. What I give up is cost efficiency at steady high volume, where per-GB pricing loses to a well-sized set of shards, and instant elasticity: on-demand accommodates growth relative to recent peak over a warm-up window of minutes, so a genuine cold-start spike can still throttle. You can switch modes on a live stream, but only a small number of times per 24 hours, so it is not a knob to toggle per traffic peak.

code

bash · 14 lines
bash
# Create a stream that manages its own shard count
aws kinesis create-stream \
  --stream-name events-ondemand \
  --stream-mode-details StreamMode=ON_DEMAND

# Later, once the traffic shape is understood, move to provisioned
aws kinesis update-stream-mode \
  --stream-arn arn:aws:kinesis:eu-west-1:111122223333:stream/events-ondemand \
  --stream-mode-details StreamMode=PROVISIONED

aws kinesis update-shard-count \
  --stream-name events-ondemand \
  --target-shard-count 8 \
  --scaling-type UNIFORM_SCALING

go deeper

for a junior

Know that a Kinesis stream has two capacity modes, that provisioned means you pick the shard count and on-demand means AWS does, and that both still bill you for the traffic.

for a middle

Explain the billing dimensions on each side — shard-hour plus PUT payload units versus stream-hour plus per-GB in and out — and name the resharding APIs you are responsible for in provisioned mode.

for a senior

Show that you plan for the warm-up: describe how you would pre-warm or provision ahead of a known spike, and how producers should behave when they are throttled during a ramp.

for a principal

Frame it as an operations-versus-cost decision: on-demand buys away a scaling loop nobody has budget to build and test, and that capability is often worth more than the per-GB premium — but say at what volume the arithmetic flips.

## Two ways to own the same shards Underneath, both capacity modes are the same shard machinery described elsewhere in this topic: a stream is a set of shards, each with its own throughput and hash-key range. Capacity mode only decides **who chooses the shard count and how AWS charges for it**. ## Provisioned mode You call `CreateStream` with a `ShardCount` and that is what you get. Billing is per shard-hour plus per **PUT payload unit** — a 25 KB chunk of written payload, so a 60 KB record costs three units. Extended retention and enhanced fan-out add their own line items. Because the shard count is yours, so is the scaling. `UpdateShardCount` performs a uniform scale (roughly doubling or halving is the safe path), while `SplitShard` and `MergeShards` let you act surgically on one hash-key range. Each of these closes parent shards and opens children, which consumers must handle — the KCL does this for you by processing a parent shard to completion before its children so per-key order survives the reshard. Provisioned mode is the right answer when you know your traffic: a steady pipeline with a predictable diurnal curve, where you can size once and adjust on a schedule. At sustained high volume it is meaningfully cheaper than on-demand. ## On-demand mode You create the stream with `StreamModeDetails` set to `ON_DEMAND` and never think about shards again. AWS observes traffic and splits or merges shards for you. Billing switches to a per stream-hour charge plus a per-GB charge for data written and data read — which means a consumer added later increases your bill directly, something provisioned mode hides inside the shard-hour. The elasticity is real but not instantaneous. On-demand scales relative to the recent observed peak, over a warm-up window measured in minutes. The practical consequence: a stream that has been idle for a week and suddenly receives a hundred times its previous peak will throttle during the ramp. If you know a spike is coming — a product launch, a batch backfill, a Black Friday — either pre-warm by driving traffic ahead of time or use provisioned mode with a deliberately sized shard count. There is also a default per-stream quota on on-demand throughput. It is generous, it is adjustable through a service-quota request, and it is not infinite — treat "on-demand" as "managed", not "unbounded". ## Choosing The decision usually collapses to three questions: 1. **Do I know the traffic?** No → on-demand. A new product, an unpredictable SaaS tenant, a stream whose volume depends on customer behaviour you cannot forecast. 2. **Is the volume high and steady?** Yes → provisioned. At a constant multi-MB/s ingest with well-behaved keys, shard-hours beat per-GB pricing, and the operational cost of resharding is small when you reshard rarely. 3. **Who is going to run the scaling loop?** If nobody will build and test autoscaling around `UpdateShardCount`, on-demand is buying you an operational capability, not just capacity. That is often the real justification — an under-provisioned stream throttles producers, and dropped events are far more expensive than the pricing delta. ## The trap: on-demand does not fix key skew This is the answer that separates a good response from a recited one. On-demand scales the *stream*; it cannot make a single partition key exceed one shard's throughput, because a key hashes to exactly one hash-key range at a time. A stream where 90% of records carry `tenant=acme` will throttle in on-demand mode too. On-demand does react to a hot shard by splitting it, but the split only helps if the traffic behind it is spread across several keys. Skew is a data-modelling problem, not a capacity-mode problem. ## Switching modes `UpdateStreamMode` changes the mode of a live stream without recreating it, and consumers keep reading. The switch is rate-limited to a small number of changes per 24 hours, so use it as a migration step — for instance, launching on on-demand to learn the traffic shape and moving to provisioned once the curve is known — not as a scaling mechanism.

  • A stream has been idle for days and a marketing campaign will drive a 50x spike at a known time. Does on-demand cover you?
    Not reliably. On-demand scales relative to recent observed peak over a warm-up of minutes, so a cold 50x step can throttle producers during the ramp. Either pre-warm the stream by driving synthetic traffic ahead of the event, or switch to provisioned with an explicitly sized shard count for the window. Whichever you pick, make the producer retry throttled records with backoff.
  • How does adding a second consumer application change the bill in each mode?
    In provisioned mode a standard consumer adds nothing directly — it shares the shard's existing 2 MB/s read budget, which you already pay for by the shard-hour. In on-demand mode reads are billed per GB retrieved, so a second consumer roughly doubles the read charge. In either mode, registering that consumer for enhanced fan-out adds a per-consumer-shard-hour and a per-GB retrieval charge.

saying these in an interview costs you the question

  • Says on-demand removes throughput limits entirely
  • Assumes on-demand scales instantly with no warm-up
  • Believes on-demand fixes a skewed partition key
  • Treats mode switching as a per-spike scaling knob
  • Compares only shard-hour price and ignores per-GB read charges

context