skip to content

What do EventBridge archives and replay give you, and what would you be careful about before replaying events into a production bus?

level: seniorimportance: should knowfreq 38%

answer

  1. a bus that forgets versus one that remembers
  2. filtered by a pattern, kept for a period
  3. nothing before it existed is captured
  4. restrict the destination rules
  5. replayed events are tagged

basics

~20 s

An archive durably retains the events on a bus, optionally filtered by a pattern and kept for a chosen retention period. Replay re-publishes a time range of archived events to that bus's rules, which lets you recover from a broken consumer or backfill a new one.

solid answer

~50 s

An archive is attached to a bus, can filter with an event pattern, and holds matching events for a retention period you choose (or indefinitely). Replay then re-publishes events from a time window back to the bus, either to all its rules or to a named subset via the rule filter. That is what makes an event bus recoverable: a consumer that crashed or dropped events for two hours can be re-fed, and a brand-new consumer can be backfilled with history. The cautions are real. An archive only captures events that arrived *after* it was created — there is no retroactive capture, so you create archives on day one. Replay is not ordered, replayed events reach every rule you did not exclude, and they carry a `replay-name` field so consumers can distinguish them. Because targets are already at-least-once, replaying into live rules means duplicate side effects — emails resent, charges retried — unless the consumers are idempotent or you narrow the replay to specific rules.

code

bash · 8 lines
bash
aws events create-archive --archive-name orders-archive \
  --event-source-arn arn:aws:events:eu-west-1:111122223333:event-bus/orders-bus \
  --event-pattern '{"source":["com.acme.orders"]}' --retention-days 30

aws events start-replay --replay-name backfill-projector \
  --event-source-arn arn:aws:events:eu-west-1:111122223333:archive/orders-archive \
  --event-start-time 2025-03-04T02:00:00Z --event-end-time 2025-03-04T04:00:00Z \
  --destination '{"Arn":"arn:aws:events:eu-west-1:111122223333:event-bus/orders-bus","FilterArns":["arn:aws:events:eu-west-1:111122223333:rule/orders-bus/projector"]}'

go deeper

for a junior

Know that an archive stores a bus's events for a retention period and that replay can send a time range of them back through the bus's rules.

for a middle

Explain that archives take an event pattern and are not retroactive, and that a replay is bounded by a time window and can be restricted to named rules on the destination bus.

for a senior

Show the operational caution: replay is unordered, reaches every rule you did not exclude, floods downstreams, and duplicates side effects unless consumers branch on replay-name and are idempotent.

for a principal

Own the policy: which buses get archives on day one, what retention and cost the business actually needs, and the organisation-wide convention that consumers suppress external side effects during a replay.

## What an archive is By default an event bus is amnesiac. An event arrives, rules evaluate it, targets are invoked, and nothing remains. That is fine until a consumer has a bug, and then the events it mishandled are simply gone. An **archive** fixes that. You create it on a specific bus with: - an optional **event pattern**, using exactly the same syntax as a rule, so you can archive only what matters rather than every event on the bus; - a **retention period** in days, or zero for indefinite retention. From that moment, matching events are stored as they arrive. Archiving is passive: it does not intercept, delay or divert anything, and rules fire exactly as they would without it. The single most important property is that an archive is **not retroactive**. It captures nothing that arrived before it existed. This is why archives belong in the initial setup of any bus that carries business-meaningful events — by the time you want history, it is too late to start. ## What replay does `StartReplay` takes a named replay, the archive to read from, an `EventStartTime` and `EventEndTime` bounding the window, and a destination consisting of the event bus plus an optional list of rule ARNs to restrict delivery to. EventBridge then re-publishes the archived events in that window, and the destination's rules evaluate them like any other event. The uses are exactly what you would expect: - **Recovery.** A consumer was down or buggy between 02:00 and 04:00; fix it, replay that window to just that consumer's rule. - **Backfill.** A new service joins the system and needs the last month of events to build its state. - **Debugging with real traffic.** Point a replay at a rule whose target is a test consumer and watch it handle production-shaped events. Several constraints shape how you use it. Events are replayed to the **same bus** the archive was made from, so the rule filter is how you control who sees them. **Ordering is not guaranteed**, so a consumer that assumes chronological arrival will misbehave on a replay even if it looked fine live. Replayed events carry a **`replay-name`** field identifying the replay, which is the hook consumers use to branch — for example, to skip sending customer-facing notifications. And a replay's progress is visible as a state you poll; it is not instantaneous for a large window. ## The care that matters before replaying into production This is the part interviewers push on, because a careless replay is a self-inflicted incident. 1. **Everything not excluded will fire.** Unless you pass the rule filter, every rule on the destination bus reprocesses those events. If one of them sends email, charges a card, or posts to a partner API, you have just done it all again. The default posture should be to name the rules explicitly. 2. **Idempotency is your only real protection.** Target delivery is already at-least-once, so consumers should be idempotent anyway — but replay converts "rare duplicate" into "thousands of duplicates in minutes", which is where weak deduplication breaks (an in-memory cache, a TTL shorter than the replay window). 3. **Downstream load.** A two-hour window replays as fast as EventBridge can push it. Third-party APIs get rate-limited, a database gets a write spike, a Lambda hits its concurrency ceiling. Consider replaying a narrow window first and watching. 4. **Stale side effects.** Some events are only meaningful when fresh. Replaying a week-old "price changed" event into a rule that pushes a notification is worse than doing nothing. 5. **Nothing scheduled comes back.** Replay reproduces archived bus events, not schedule-driven invocations that never existed as archived events. A safe default pattern is: consumers check for `replay-name` and suppress externally visible side effects when it is present, while still updating internal state. That single convention makes replay a routine operation rather than a change-managed event. ## Cost and retention judgment Archives are charged on storage and replay is charged per event processed, so infinite retention of a chatty bus is a real bill. The useful discipline is to scope the archive's pattern to the events you would actually replay — the domain events with business meaning — and set retention to the window in which replaying is still plausible, which for most systems is weeks, not years. Long-term analytical history is a data-lake problem, not an archive one.

  • How does a consumer tell a replayed event apart from a live one, and why would it care?
    Replayed events carry a `replay-name` field. Consumers use it to branch: update internal state as normal, but suppress externally visible side effects such as emails, payments or partner API calls. Adopting that convention across consumers turns replay from a risky change-managed operation into something you can run during an incident without a meeting.
  • A team asks you to replay last quarter's events, but the archive was created a month ago. What do you tell them?
    That data does not exist. Archives are not retroactive — they capture only events that arrived after creation, so there is nothing to replay from before that date. The lesson is to create the archive when the bus is created. If a longer history is genuinely needed, that is a data-lake concern: land events in S3 continuously and rebuild from there.
  • What operational risk comes from replaying a large time window at once?
    Replay pushes at EventBridge's pace, not your downstream's. A wide window becomes a write spike, a Lambda concurrency ceiling, or a rate-limited third-party API — and target retries then amplify the failures. Replay a narrow window first, watch the target's error and throttle metrics, and scope the destination to specific rules so only the consumer you mean is loaded.
  • Can you assume events arrive in their original order during a replay?
    No. EventBridge does not guarantee ordering live and does not guarantee it during replay either, so a consumer that reconstructs state must tolerate arbitrary order — typically by carrying a version or business timestamp in the payload and discarding stale updates. A consumer that worked fine live can still corrupt state under replay for exactly this reason.

saying these in an interview costs you the question

  • Assuming an archive captures events from before it was created
  • Believing replay preserves the original event order
  • Forgetting that replay reaches every rule you did not filter out
  • Treating an archive as a general long-term analytics store
  • Thinking replayed events are indistinguishable from live ones

context