In Kinesis Data Streams, how does an enhanced fan-out consumer differ from the default shared-throughput consumer, and when is the extra cost justified?
answer
- 2 MB/s is shared, not per app
- register the consumer to get your own
- push over HTTP/2, not polling
- billed per consumer-shard-hour
basics
~20 sA standard consumer polls GetRecords and shares one shard's 2 MB/s read budget with every other standard consumer. An enhanced fan-out consumer registers with the stream and gets its own 2 MB/s per shard, pushed over HTTP/2, at extra cost.
solid answer
~50 sBy default consumers poll `GetRecords`, and each shard offers 2 MB/s and 5 calls per second **shared across every standard consumer** on it. So the third application you attach does not get its own bandwidth — it takes a slice, adds polling contention, and pushes everyone's latency up. Enhanced fan-out changes both the allocation and the transport: you call `RegisterStreamConsumer`, and that consumer uses `SubscribeToShard`, an HTTP/2 push subscription with its own dedicated 2 MB/s per shard, typically delivering in around 70 milliseconds rather than the couple of hundred you get from polling. The cost is a charge per consumer-shard-hour plus a per-GB data-retrieval charge, and there is a default quota on registered consumers per stream. I justify it when several applications must read the same stream independently, or when one latency-sensitive consumer must not be starved by a batch consumer sharing the shard. For a single consumer that tolerates a couple of hundred milliseconds, standard polling is cheaper and simpler.
code
python · 20 linesimport boto3
kinesis = boto3.client("kinesis")
stream_arn = "arn:aws:kinesis:eu-west-1:111122223333:stream/events"
consumer = kinesis.register_stream_consumer(
StreamARN=stream_arn,
ConsumerName="realtime-alerts",
)["Consumer"]
# HTTP/2 push subscription: dedicated 2 MB/s on this shard for this consumer.
resp = kinesis.subscribe_to_shard(
ConsumerARN=consumer["ConsumerARN"],
ShardId="shardId-000000000000",
StartingPosition={"Type": "LATEST"},
)
for event in resp["EventStream"]: # lasts up to 5 minutes, then renew
records = event.get("SubscribeToShardEvent", {}).get("Records", [])
print(len(records), "records pushed")go deeper
Know that by default all consumers of a shard share one 2 MB/s read budget, and that enhanced fan-out is the paid option that gives a registered consumer its own.
Explain both differences — dedicated throughput and HTTP/2 push via SubscribeToShard instead of GetRecords polling — and give the rough latency figures each delivers.
Show the cost model in your reasoning: per consumer-shard-hour plus per-GB retrieval, and the registered-consumer quota. Argue for a mixed setup where only the latency-critical readers pay for EFO.
Own the fan-out strategy for the platform: how many independent readers a stream should ever have, when a second stream or an archive-and-replay path beats registering another consumer, and who pays for the standing per-consumer-shard-hour cost.
## The default: shared throughput A standard Kinesis Data Streams consumer gets a shard iterator and calls `GetRecords` in a loop. The shard grants **2 MB per second of reads and at most 5 `GetRecords` calls per second — in total, across every standard consumer reading that shard**. That budget is not per application. Consequences follow directly: - With one consumer, you comfortably out-read a shard that ingests at most 1 MB/s. Life is good. - With two, each averages about 1 MB/s and they begin competing for the five calls per second. - With three or more, throttling (`ReadProvisionedThroughputExceeded`) becomes routine, and every consumer's `GetRecords.IteratorAgeMilliseconds` starts to climb. A batch job doing a big backfill can starve a real-time consumer that was fine yesterday. Latency is the second cost. Polling means a record waits for the next poll; with one consumer polling at the maximum rate you land around 200 ms of propagation delay, and with several consumers sharing five calls per second each application must poll less often, so it gets worse the more consumers you add. ## Enhanced fan-out: dedicated throughput and push delivery Enhanced fan-out (EFO) changes two things at once. **Allocation.** You call `RegisterStreamConsumer` to create a named consumer on the stream, which gets its own ARN. That consumer receives a **dedicated 2 MB/s per shard**, independent of every other consumer. Ten registered consumers on a 4-shard stream each read up to 2 MB/s per shard; nobody's backfill starves anyone's alerting pipeline. **Transport.** Instead of polling `GetRecords`, the consumer opens a `SubscribeToShard` call — an HTTP/2 stream over which Kinesis **pushes** records as they arrive. Typical propagation is about 70 ms and does not degrade as you add consumers. The subscription is not permanent: a `SubscribeToShard` call lasts up to five minutes, after which the client renews it. Any real client — the KCL, the Lambda event source mapping in EFO mode, the Flink Kinesis connector — handles that renewal for you. ```bash aws kinesis register-stream-consumer \ --stream-arn arn:aws:kinesis:eu-west-1:111122223333:stream/events \ --consumer-name realtime-alerts aws kinesis list-stream-consumers \ --stream-arn arn:aws:kinesis:eu-west-1:111122223333:stream/events ``` ## What it costs EFO is billed on two dimensions: a charge per **consumer-shard-hour** (registered consumers multiplied by shards multiplied by hours) and a charge per **GB of data retrieved**. The first is the one that surprises people: the meter runs on shard count and consumer count regardless of how much data you actually read, so registering five consumers on a 64-shard stream is a standing cost even overnight. Standard consumers, in provisioned capacity mode, add no incremental charge at all — you already pay for the shard. There is also a default quota on the number of consumers registered for EFO per stream (a small number, raisable through a service-quota request). If your architecture calls for dozens of independent readers, that quota is a design constraint worth checking early. ## Choosing Reach for EFO when: - **Multiple applications read the same stream** and must not interfere with each other. This is the dominant reason. - **One consumer is latency-critical.** Roughly 70 ms versus 200 ms-plus matters for alerting, fraud scoring, or live personalisation; for a five-minute batch aggregate it does not. - **A backfill must run alongside production.** Registering the backfill as its own EFO consumer isolates its aggressive reads from the live path. Stay with standard consumers when there is one reader, or two that both tolerate polling latency, and the stream is wide enough that the per-consumer-shard-hour charge would dominate. A useful middle position: not every consumer has to make the same choice. Register the latency-sensitive application for EFO and leave the analytics job on shared polling — the analytics job then has the shared 2 MB/s largely to itself. ## Things candidates get wrong EFO does not increase **write** capacity, and it does not increase the shard's storage or retention. It also does not fix a hot shard: if one shard receives most of the traffic, a dedicated 2 MB/s on that shard may still not keep up with what the producers wish they could write, and the shard's 1 MB/s write ceiling is untouched. Finally, EFO is a per-consumer property, not a stream-wide mode — you register consumers individually, and a stream can serve both kinds at once.
- You add a third standard consumer to a stream and all three start lagging. What exactly ran out?The shard's shared read budget — 2 MB/s and five `GetRecords` calls per second, split across all standard consumers on that shard. Each application now averages under 1 MB/s and polls less often, so `ReadProvisionedThroughputExceeded` appears and `GetRecords.IteratorAgeMilliseconds` climbs for everyone. The fixes are registering some consumers for enhanced fan-out, or adding shards so each carries fewer bytes.
- Does every consumer on a stream have to use the same retrieval mode?No. Enhanced fan-out is a property of an individual registered consumer, not of the stream. A common pattern is registering the latency-sensitive application for EFO while leaving a batch analytics job on standard polling — which also hands that job most of the shared 2 MB/s, since it is now the only one using it.
- Would enhanced fan-out help a stream whose producers are being throttled?No. EFO only changes how data is read out. Write capacity is still 1 MB/s or 1,000 records per second per shard, decided by shard count and partition-key distribution. Producer throttling is answered by fixing key skew, adding shards, or moving to on-demand capacity mode — never by changing the consumer's retrieval mode.
saying these in an interview costs you the question
- Thinks each consumer already gets its own 2 MB/s by default
- Says enhanced fan-out increases write throughput
- Believes EFO is a stream-wide setting rather than per consumer
- Ignores the per-consumer-shard-hour charge on wide streams
- Assumes a SubscribeToShard subscription lives forever without renewal