An Amazon Data Firehose delivery stream writes to Amazon OpenSearch Service and the domain starts rejecting writes. What does Firehose do with those records, and how do you recover?
answer
- retry window, then park in S3
- that's why a backup bucket is required
- error prefix is not a redrive queue
- producers keep succeeding regardless
- alarm on the prefix, not the producer
basics
~20 sFirehose retries for the configured retry duration, then writes the still-failing records to the backup S3 bucket under an error output prefix. They are not retried again automatically — recovery means reading those objects and re-ingesting them yourself.
solid answer
~60 sFirehose does not drop the data and does not block indefinitely. It retries delivery for the stream's configured **retry duration**, and any records still failing at the end of that window are written to the **backup S3 bucket** under an error output prefix, wrapped with metadata about the failure. That behaviour is why an S3 backup bucket is mandatory for the OpenSearch, Redshift, Splunk and HTTP-endpoint destinations — S3 is the fallback when the real destination will not accept writes. What Firehose will *not* do is replay those objects once the domain recovers: the error prefix is a parking lot, not a retry queue, so recovery is a job you run — read the failed records back out and re-ingest them through the delivery stream or a bulk load. Operationally that means two things: alarm on objects appearing under the error prefix, since a stream that is quietly diverting everything looks healthy from the producer side, and watch the delivery-success and `DeliveryToS3.DataFreshness` metrics to catch the problem before the retry window expires.
go deeper
Know that Firehose retries a failing delivery for a while and then writes the records it could not deliver into an S3 backup location instead of discarding them.
Explain the retry duration, why non-S3 destinations require a backup bucket, and that the error output prefix separates failure types rather than pooling them.
Show that you have operated this: the failure is invisible to producers, so alarm on the error prefix and on delivery freshness, and have a tested replay job because Firehose offers no redrive.
Own the recovery contract across the platform — retry durations chosen per feed's freshness-versus-completeness needs, standard replay tooling, and who is accountable for reprocessing when a destination rejects data for days.
## Firehose's failure philosophy Every managed pipeline has to answer one question: when the destination will not take the data, what happens to it? Firehose's answer is consistent across destinations — **retry for a bounded window, then park the data in S3 rather than lose it or stall forever.** Understanding that shape is more useful than memorising per-destination knobs, because it explains why the console makes you nominate an S3 bucket even when your destination is an OpenSearch domain. ## The retry window Each non-S3 destination has a configurable **retry duration**. During that window Firehose keeps attempting delivery of the failed batch. The duration is the trade you are making explicitly: - **Short** — failures reach S3 quickly, so nothing sits in flight, but a transient blip that would have healed in ten minutes creates a backfill job. - **Long** — more transient failures self-heal, but data freshness degrades and a genuinely broken destination is discovered later. Setting the duration to zero means "do not retry at all", which is occasionally right for a feed where freshness beats completeness. ## The error output destination When retries are exhausted, the records go to the delivery stream's S3 backup location under an error output prefix, tagged with metadata describing why they were routed there. Different failure classes are separated by that prefix, which is why the `ErrorOutputPrefix` supports the `!{firehose:error-output-type}` element — records that failed a Lambda transformation and records the destination rejected land in distinguishable paths rather than in one undifferentiated pile. This is *not* a dead-letter queue with redrive. Nothing reads the prefix, nothing retries from it, and the delivery stream will not notice when the domain recovers. Recovery is entirely yours: read the objects, unwrap the records from their error envelope, and re-ingest — through the same delivery stream, through a bulk load into the destination, or through a one-off job. The realistic operational answer in an interview is "I'd have a documented, tested replay job for this before I needed it", because writing one under incident pressure while data accumulates is how backfills get skipped. ## Why the OpenSearch case in particular Rejections from an OpenSearch domain are usually not "the domain is down". They are per-document failures — a mapping conflict where a field arrived as a string after being indexed as a number, a rejected bulk request because the write queue is saturated, or a cluster that has gone read-only on low disk. Two consequences follow. First, the failure can be **partial and persistent**: a subset of documents fails every retry because they will never be accepted with the current mapping. Retrying longer does not help; those belong in the error prefix so you can fix the shape and replay. Second, the fix often is not in Firehose at all. Index mapping is the search engine's own concern and belongs to whoever owns the domain; Firehose's part of the problem is the retry duration, the backup bucket, and the alarm. Keep that boundary clear in the answer. ## S3 as destination is the special case When S3 *is* the destination there is no fallback bucket to divert to, so Firehose simply retries delivery for an extended period. If it still cannot write — a bucket policy that denies the delivery role, a KMS key it cannot use, a deleted bucket — the data eventually cannot be delivered at all. Nearly every real occurrence is a permissions or key problem rather than an S3 outage, which makes the delivery role's policy the first thing to check. ## Detecting it The dangerous property of this design is that failure is *silent from the producer's perspective*. `PutRecord` keeps succeeding; ingest metrics look normal; the data is simply no longer arriving where anyone reads it. Three signals close that gap: 1. **An alarm on objects appearing under the error output prefix.** The most direct signal, and the one most often missing. S3 event notifications into a metric or a Lambda is enough. 2. **The delivery-success metrics** for the destination, which fall away from the ingest rate when deliveries start failing. 3. **`DeliveryToS3.DataFreshness`**, the age of the oldest undelivered record, which climbs steadily rather than sawtoothing when a stream is stuck retrying. ## The answer in one shape Retry for a bounded window; park what still fails in S3 with typed error prefixes; alarm on that prefix because nothing else will tell you; own the replay job yourself, because Firehose has no redrive. A candidate who says all four has clearly operated one of these.
- Why does Firehose require an S3 backup bucket for destinations that are not S3?Because S3 is the fallback when the real destination refuses writes. Firehose stores nothing itself, so after the retry duration expires it needs somewhere durable to put records rather than dropping them. That bucket, under the error output prefix, is the only surviving copy of anything the destination rejected.
- How would you choose the retry duration for a delivery stream?By how long a plausible transient failure lasts against how stale the data may become. A long window lets blips self-heal but delays discovery and degrades freshness; a short one gets failures visible fast at the cost of more backfill work. Zero is defensible when freshness matters more than completeness.
- What makes this class of failure easy to miss in production?It is invisible from the producer side — `PutRecord` keeps succeeding and ingest metrics look normal while the data quietly diverts to S3. Only destination-side signals reveal it: the delivery-success metrics, `DeliveryToS3.DataFreshness` climbing steadily, and an alarm on objects landing under the error output prefix.
saying these in an interview costs you the question
- Assumes Firehose drops records the destination rejects
- Expects automatic redrive once the destination recovers
- Thinks failing deliveries surface as producer PutRecord errors
- Treats the error output prefix as a dead-letter queue with retry
- Blames the domain outage and ignores mapping-level rejections