When would you choose an Amazon Data Firehose delivery stream over a Kinesis Data Streams stream for a high-volume event feed, and what do you give up?
answer
- delivery, not a log
- no consumer code, no shards
- buffered seconds, not milliseconds
- nothing stored means nothing replayed
- chain a data stream in front
basics
~20 sAmazon Data Firehose is a managed delivery pipeline: it buffers records and writes them to destinations such as S3, Redshift or OpenSearch with no consumer code and no shards to manage. You give up replay, retention and sub-second latency.
solid answer
~60 sFirehose and Kinesis Data Streams solve different halves of ingest. Data Streams is a durable log you read from: it retains records (24 hours by default, up to 365 days), lets many independent consumers read the same records at their own pace, and can be replayed from an earlier position — but you own the consumer code, the checkpointing and, in provisioned mode, the shard capacity. Firehose is the opposite trade: you point it at a destination, it buffers records and delivers them, and there is nothing to consume and nothing to size. It scales on its own and bills per GB ingested. What you lose is real: Firehose stores nothing, so there is no replay and no second consumer; delivery is at-least-once, so duplicates can land; and latency is a buffering window of seconds to minutes, not milliseconds. Pick Firehose when the answer to "who reads this?" is "a bucket or an index". Pick Data Streams when you need low latency, multiple consumers, or the ability to reprocess. If you need both, use a Data Stream as the Firehose source.
go deeper
Be able to say what each service is for in one line: Firehose delivers records into a destination for you, Data Streams stores them so your own code can read them. Naming the common Firehose destinations is enough at this level.
Explain the mechanics behind the choice — retention and consumer positions on one side, buffering and managed delivery on the other — and state plainly that Firehose keeps nothing, so replay is not available.
Show production judgment: call out at-least-once duplicates, the buffering latency floor, and source record backup as the safety net when a transformation goes wrong. Justify the choice against a concrete workload rather than in the abstract.
Own the composition and the economics: when one Data Stream should feed several branches including a Firehose archive, when paying twice to ingest is cheaper than running a consumer fleet, and how the choice constrains reprocessing strategy for the whole platform.
## Two primitives, two different jobs The Kinesis family is easiest to reason about if you stop treating the two services as tiers of the same product. **Kinesis Data Streams** is a *durable, replayable log*. **Amazon Data Firehose** (formerly Kinesis Data Firehose) is a *delivery pipeline*: records go in one end and land in a destination at the other, with no place in the middle where they sit and wait to be read. Almost every distinction that matters follows from that one difference. ## What Data Streams gives you A stream retains every record for a retention period — 24 hours by default, configurable up to 365 days. Because the records are *stored*, a consumer has a position in the log, and that position is under your control. That buys three things: - **Replay.** If a downstream bug corrupted yesterday's aggregates, you rewind the consumer and reprocess, provided the data is still inside retention. - **Fan-out.** Several independent applications can read the same records without one interfering with another; each keeps its own position. - **Low latency.** A consumer can read within milliseconds of the write. The cost is operational. You write the consumer, you handle checkpointing and failure recovery, and in provisioned mode you decide how much capacity the stream has and adjust it as traffic grows. ## What Firehose gives you Firehose has no consumer API at all. You create a delivery stream, choose a destination — S3, Amazon Redshift, OpenSearch Service, Splunk, Snowflake, or a generic HTTP endpoint that many SaaS observability vendors expose — and write records with `PutRecord` or `PutRecordBatch` (or wire a Kinesis data stream or other AWS source in as the source). Firehose accumulates records into a buffer, and when the buffer fills or the buffer interval elapses it writes the batch out. On the way it can do the work you would otherwise write a consumer for: invoke a Lambda function to transform or filter each record, convert JSON to Parquet or ORC using a schema from the AWS Glue Data Catalog, compress the output, and derive the S3 prefix from fields inside the records. Capacity is not your problem — there is nothing to provision and nothing to reshard — and you are billed on volume ingested rather than on reserved capacity sitting idle. ```text producers -> Firehose (buffer, transform, convert) -> S3 / Redshift / OpenSearch producers -> Data Stream (retain N days) -> your consumers (replay, fan-out) ``` ## What you give up, precisely 1. **No retention, therefore no replay.** Once a batch is delivered, Firehose is done. If the transformation had a bug, the only copy is what was written — which is why the *source record backup* option, which mirrors the raw untransformed records to a separate S3 location, exists and is worth enabling. 2. **No second consumer.** One delivery stream, one destination. Two destinations means two delivery streams, each ingesting (and billing for) the data again. 3. **At-least-once delivery.** Retries can produce duplicate records in the destination. Anything downstream that must be exact needs to deduplicate — for example on an idempotency key during the query or load step. 4. **Buffered latency.** Delivery is bounded by the buffering window rather than by the network. It is near-real-time, not real-time; a trading or fraud path that needs millisecond reaction should not sit behind it. ## The combination answer interviewers are listening for The two are not exclusive, and saying so is usually what separates a good answer from a memorised one. A Kinesis data stream can be the *source* of a Firehose delivery stream. That gives you one ingest path with the log's retention and multi-consumer semantics for the applications that need them, plus a hands-off archival branch that lands everything in S3 as partitioned Parquet for analytics — without writing an S3 writer, a Parquet encoder, or a batching layer. The reverse framing is the decision rule: if the only reader of this data is a bucket, a warehouse or a search index, Firehose removes an entire application you would otherwise have to run. The moment a human or a service needs to read the feed itself — with its own position, its own pace, or the option to go back — you need the log. ## Where candidates go wrong The common failure is treating Firehose as "Kinesis for beginners" and assuming it inherits retention and replay. It does not. The second failure is the opposite: reaching for Data Streams for a pure archival path, then spending weeks building the batching, compression and partitioning that Firehose does as configuration.
- Can one ingest path give you both replay and hands-off delivery to S3?Yes — use a Kinesis data stream as the source of the Firehose delivery stream. Applications that need low latency, their own position, or reprocessing read the stream directly; Firehose acts as one more branch that archives everything to S3. You pay for both, but you build only the consumers that genuinely need to exist.
- How does the billing model differ between the two?Firehose bills on volume ingested, with extra per-GB charges for optional features such as format conversion and dynamic partitioning; there is no idle capacity to pay for. Data Streams bills for capacity — in provisioned mode you pay per shard-hour whether or not the shards are busy. That makes Firehose cheap for spiky archival feeds and Data Streams predictable for steady, heavily-consumed ones.
- Firehose is at-least-once. How do you keep duplicates from corrupting downstream numbers?Assume duplicates and design the consumer of the destination to be idempotent: carry a stable event id in every record and deduplicate on it during the query, the load step, or a compaction job. Do not try to make Firehose exactly-once — it is not, and no configuration makes it so.
saying these in an interview costs you the question
- Says Firehose is just Data Streams with a friendlier name
- Claims you can replay data already delivered by Firehose
- Thinks you size a Firehose delivery stream in shards
- Assumes Firehose delivers exactly once, so no dedupe
- Calls Firehose millisecond-latency real-time streaming