A service writes records to an Amazon Data Firehose delivery stream targeting S3, but objects only appear in the bucket every few minutes. Why, and which settings control it?
answer
- batches, not per-record writes
- two hints, first one wins
- quiet stream waits for the clock
- hints, because Firehose may raise them
- small files versus fresh data
basics
~20 sFirehose batches records before writing. Two buffering hints control the flush — a buffer size in MiB and a buffer interval in seconds — and whichever is reached first triggers delivery, so a low-traffic stream waits for the interval.
solid answer
~50 sNothing is broken: Firehose is a buffered delivery pipeline, not a per-record writer. It accumulates records until either the **buffer size** hint (in MiB) or the **buffer interval** hint (in seconds, up to 900) is reached, then writes one object to S3. For the S3 destination the defaults are 5 MiB and 300 seconds, so a stream that trickles a few KB a second will always be flushed by the interval, which is exactly the delay being observed. They are called *hints* rather than limits because Firehose may raise the buffer size on its own when delivery is falling behind, in order to catch up. Tuning them is a straight tradeoff: a long interval and large buffer produce fewer, larger objects — cheaper and much faster to scan with Athena — while a short interval gets data visible sooner at the cost of many small files. Watch the `DeliveryToS3.DataFreshness` CloudWatch metric to see the age of the oldest undelivered record.
go deeper
Know that Firehose batches before it writes, and be able to name the two controls — buffer size and buffer interval — and say that whichever is hit first triggers delivery.
Explain why a low-traffic stream is always flushed by the interval, why the settings are hints rather than limits, and that the size is measured before compression.
Reason about the tuning tradeoff in production terms: object size versus freshness, the small-file cost on Athena and Spark, and alarming on data freshness to distinguish waiting from a stalled destination.
Own the platform-wide convention: standard buffering and file-size targets across delivery streams, when a freshness requirement should be met by a separate low-latency branch instead of shrinking buffers, and the query-cost consequences of getting it wrong at scale.
## Firehose delivers batches, not records The first thing to understand about Amazon Data Firehose is that its unit of delivery is a *batch*, never an individual record. Producers call `PutRecord` or `PutRecordBatch` and the records go into a buffer inside the service. When that buffer is flushed, Firehose writes the accumulated records out as one object (for an S3 destination) or one bulk request (for OpenSearch). Seeing objects appear every few minutes is the service working exactly as designed. ## The two hints A delivery stream has two buffering settings: - **Buffer size** — how much data to accumulate, expressed in MiB. - **Buffer interval** — how long to wait, expressed in seconds, with a maximum of 900 (15 minutes). They are evaluated together, and **whichever condition is met first wins**. For the S3 destination the defaults are 5 MiB and 300 seconds as of 2025. That produces two very different behaviours depending on traffic: - A busy stream fills 5 MiB quickly, so the *size* hint dominates and objects appear frequently. - A quiet stream never reaches the size hint, so the *interval* dominates and you get one object every 300 seconds, however little data is in it. That is why the symptom in the question is usually a low-volume stream rather than a fault. The values differ by destination — OpenSearch and HTTP-endpoint destinations use their own defaults and ranges — so read the ranges for the destination you actually configured rather than assuming the S3 numbers apply everywhere. ## Why AWS calls them hints The settings are *hints*, not guarantees, and the reason is capacity. If the destination or the delivery path is slower than the incoming rate, Firehose may **raise the buffer size beyond the configured value** so it can catch up with fewer, bigger writes. You can therefore see objects larger than the size you configured; you should not treat the buffer size as a hard upper bound on object size, and downstream tooling should not depend on one. One more detail catches people out: the buffer size applies to the data **before compression**. If you enable GZIP, the objects that land in S3 are smaller than the configured buffer size, sometimes dramatically so for JSON. ## Tuning: latency against file size This is the tradeoff an interviewer wants named explicitly. **Short interval, small buffer** gets records visible sooner. The price is a large number of small objects. Small files are the classic data-lake pathology: Athena and Spark pay a fixed per-file overhead, so a partition of ten thousand 40 KB objects can be an order of magnitude slower to scan than the same bytes in a handful of large ones, and you also pay more S3 request charges and more per-object lifecycle work. **Long interval, large buffer** produces well-sized objects that query efficiently, at the cost of data being minutes old before anyone can see it. A practical starting point for an analytics sink is to leave the interval near the default and raise the buffer size, then look at the resulting object sizes in the bucket and adjust. If a consumer genuinely needs fresher data than the interval allows, that is usually a signal the requirement belongs on a Kinesis data stream branch rather than on the archival path. If the delivery stream uses **dynamic partitioning**, buffering gets tighter: each active partition holds its own buffer, so the same interval now produces one object per partition per flush, and the minimum buffer size is larger. High-cardinality partition keys multiply the small-file problem instead of solving it. ## Observing it Two CloudWatch metrics tell you whether buffering is behaving: - `DeliveryToS3.DataFreshness` — the age of the oldest record still not delivered. If this climbs well past your buffer interval, delivery is falling behind rather than merely waiting. - The incoming-bytes and delivery-success metrics for the stream, which tell you whether the buffer is filling at all and whether flushes are succeeding. Alarming on data freshness is the standard way to catch a stalled or throttled destination, because a stream that has stopped delivering looks identical, from the producer side, to one that is simply waiting for its interval. ## The wrong diagnoses Do not conclude that records are being dropped: Firehose does not discard records because a buffer is "full", it flushes. Do not add a manual flush call — there is no such API. And do not assume the delay is Lambda transformation latency unless a transformation is actually configured; the buffering window applies either way.
- Why are they called hints rather than limits?Because Firehose may increase the buffer size on its own when delivery is lagging, so it can catch up with fewer, larger writes. Objects can therefore exceed the size you configured, and downstream tooling must not assume the configured value is a hard cap on object size.
- How does enabling GZIP compression interact with the buffer size?The buffer size is measured on the uncompressed data, so compression happens after the flush decision. With JSON, GZIP commonly shrinks the object several-fold, which means a 5 MiB buffer can produce objects well under a megabyte — worth accounting for when you are sizing files for Athena.
- How would you tell a slow buffer apart from a stalled delivery stream?Watch `DeliveryToS3.DataFreshness` in CloudWatch: it reports the age of the oldest undelivered record. Sitting near the buffer interval and sawtoothing is healthy waiting; climbing steadily past it means deliveries are failing or being throttled, and you should look at the delivery-success metrics and the destination itself.
saying these in an interview costs you the question
- Assumes Firehose writes each record to S3 immediately
- Thinks records are dropped when the buffer fills
- Looks for a manual flush API on the delivery stream
- Treats the buffer size as a hard cap on object size
- Sets the interval to the minimum without considering small files