A Kinesis Data Streams stream retains 24 hours of records by default. How do you decide between extending that retention and archiving the records elsewhere so you can replay them later?
answer
- retention is a recovery window, not storage
- replay speed is capped per shard
- one code path versus one bill
- archive is the system of record
basics
~20 sExtended retention keeps replay on the same code path and preserves per-shard ordering, but you pay stream prices for cold data and replay is capped by shard read throughput. Archiving to object storage is far cheaper and unbounded, at the price of a second read path.
solid answer
~60 sRetention on a stream is adjustable — 24 hours by default, raisable to seven days, and up to a year with long-term retention — via `IncreaseStreamRetentionPeriod`. The question is what you are buying. Extended retention gives replay through the exact same consumer code, with per-shard ordering intact, which is worth a lot during an incident: point the consumer at `AT_TIMESTAMP` and go. What it costs is money per shard for data nobody is reading, and time: replay is still bounded by 2 MB/s per shard, so reprocessing a week can take longer than the week's ingest did unless you add enhanced fan-out or shards. Archiving to S3 — via Firehose or a dedicated consumer — is an order of magnitude cheaper per GB, keeps history indefinitely, and lets you reprocess with parallelism S3 allows rather than shard count. The price is a second, differently shaped read path that must be built and tested. My rule of thumb: retention long enough to cover operational recovery — a bad deploy, a consumer outage over a weekend — and archive for anything older, because beyond a few days the reason for replay is analytical, not operational.
code
bash · 14 lines# Size the operational rewind window deliberately (hours).
aws kinesis increase-stream-retention-period \
--stream-name events \
--retention-period-hours 168
aws kinesis describe-stream-summary --stream-name events \
--query 'StreamDescriptionSummary.RetentionPeriodHours'
# Replay a consumer from a chosen point in time rather than the tip.
aws kinesis get-shard-iterator \
--stream-name events \
--shard-id shardId-000000000000 \
--shard-iterator-type AT_TIMESTAMP \
--timestamp 2025-01-14T02:30:00Zgo deeper
Know that a Kinesis stream keeps records for 24 hours by default, that reading does not delete them, and that the retention period can be increased so consumers can replay.
Explain the retention tiers and the APIs that change them, and know that a consumer replays by starting at TRIM_HORIZON or AT_TIMESTAMP rather than LATEST.
Reason about replay as an operation with a cost: throughput per shard bounds how fast it finishes, a backfill can starve production consumers, and lowering retention destroys data immediately.
Set the boundary and defend it: retention sized to a named operational failure mode, an archive as the system of record beyond it, and a rehearsal schedule that keeps the archive replay path real rather than theoretical.
## What the knob actually is A stream's retention period is the window during which a written record can still be read. It defaults to 24 hours; `IncreaseStreamRetentionPeriod` raises it to as much as seven days as ordinary extended retention, and long-term retention extends it further, up to a year. `DecreaseStreamRetentionPeriod` lowers it — and lowering immediately makes older data unreadable, so it is a one-way door for anything already past the new window. Billing has layers. Extended retention beyond the default carries an additional charge tied to your shards; long-term retention beyond seven days carries a separate storage charge plus a retrieval charge on data read from that older tier. The important structural point is that you are paying **stream** prices to store data that, past a day or two, is behaving like an archive. ## What extended retention genuinely buys **One code path.** This is the underrated benefit. Replaying from the stream means the same consumer, same deserialization, same idempotency logic, same checkpointing. Starting position becomes `AT_TIMESTAMP` or `TRIM_HORIZON` and you are done. An archive-based replay is a different program: read objects, parse a different container format, reconstruct order, feed the processor. That program is only correct if you have exercised it — and teams that have never rehearsed a replay from archive usually discover during the incident that they cannot. **Ordering survives.** Records replayed from a shard come back in the shard's original order with their sequence numbers. Reconstructing per-key order from files in object storage requires you to have preserved a key and a sequence in the payload and to sort on read. **Recovery from operational failure.** A consumer that crashed on Friday evening, a bad deploy that wrote garbage downstream, a checkpoint accidentally advanced past unprocessed records — these are 24-to-72-hour problems, and retention that covers a long weekend converts a data-loss incident into a rewind. ## What it does not buy **Speed.** Replay is still governed by shard read throughput: 2 MB/s per shard for standard consumers, or 2 MB/s per shard per registered enhanced fan-out consumer. Replaying seven days of a stream that ingests near its 1 MB/s per shard ceiling takes days at 2 MB/s unless you add read capacity. The mitigations are real but must be planned: register a dedicated EFO consumer for the backfill so it does not starve production, or split the replay by shard and parallelise the workers. **Cheap history.** Past a few days, per-GB stream storage is simply an expensive way to hold cold data compared with object storage, especially with lifecycle transitions to colder classes. ## What archiving buys and costs Writing every record to S3 — through Firehose, or through a consumer application you own — gives you unbounded history at object-storage prices, durable independently of the stream's lifecycle, and reprocessable with whatever parallelism you can afford rather than whatever shard count you have. It is also the substrate analytics wants anyway: partitioned, columnar, queryable. The cost is the second read path and its correctness: batching boundaries, deduplication (the archive writer is itself at-least-once), ordering reconstruction, schema evolution across months of files, and a reprocessing job that must be tested on a schedule or it will not work when needed. ## Framing the decision I would answer with a boundary rather than a number, and justify the boundary: 1. **How far back does an *operational* rewind ever need to go?** Take the worst realistic detection-plus-repair time — a consumer bug found Monday morning that started Friday night — and set retention to cover it with margin. That is usually somewhere between three and seven days, and it is the value with a defensible reason behind it. 2. **How far back does *analysis* need to go?** Months or years. That is the archive, and it is not a retention setting. 3. **How often will you actually replay?** Rarely from the stream and never from archive means the archive path is untested and effectively does not exist; budget a rehearsal or accept longer retention as insurance. 4. **Can the pipeline absorb the replay?** If reprocessing at full speed would overwhelm a downstream database or double-charge a payment provider, replay speed is not the binding constraint — throttling and idempotency are, and that argues against paying for retention you cannot exploit. 5. **Is there a compliance or contractual floor?** Some data must be retained for a stated period and some must be deleted within one; both are constraints on the archive, not on the stream, and holding regulated data in a stream you cannot easily query or expunge is usually the wrong place for it. ## The position worth defending The two options are not alternatives — the mature answer is both, with a clear boundary between them. Retention is an **operational recovery window** sized to the failure modes you can name. The archive is the **system of record** for anything older, and its replay job is treated as production code with a rehearsal schedule. Arguing for a year of stream retention because "replay is easier" is paying stream prices to avoid writing and testing a job — a trade that looks cheap right up to the point where the stream is wide.
- You extend retention to seven days and then need to reprocess all of it. What actually limits how long that takes?Read throughput per shard. Standard consumers share 2 MB/s per shard, so a stream ingesting near 1 MB/s per shard replays at best at twice real time — and slower if production consumers are sharing that budget. Register the backfill as its own enhanced fan-out consumer to give it a dedicated 2 MB/s per shard, and check that the downstream sink can absorb the resulting write rate.
- What is the risk of lowering a stream's retention period?`DecreaseStreamRetentionPeriod` takes effect immediately, and anything already older than the new window becomes unreadable at once. If a consumer is lagging or a replay is in flight, that data is simply gone. Verify the maximum `GetRecords.IteratorAgeMilliseconds` across all shards before lowering it, and lower in steps rather than in one jump.
- If everything is archived to S3 anyway, why keep more than the default 24 hours on the stream?Because the archive replay path is a different program, and one that is rarely exercised is rarely correct — batching boundaries, duplicate handling, ordering reconstruction and schema drift all bite during the incident. Extended retention buys a rewind through the code that already runs in production every day. Whether that insurance is worth its price depends on shard count and on how recently you rehearsed the archive path.
saying these in an interview costs you the question
- Assumes replay from a stream is instantaneous
- Treats stream retention as a cheap long-term archive
- Forgets that lowering retention deletes older data immediately
- Never rehearses the archive replay path and assumes it works
- Runs the backfill on the shared read budget and starves production