skip to content

A rule's target intermittently rejects deliveries. How does EventBridge retry a target invocation, and what do you configure so failed events are not silently lost?

level: seniorimportance: should knowfreq 56%

answer

  1. asynchronous, so nobody is waiting
  2. backoff plus an age bound
  3. attempts or age, whichever first
  4. without a DLQ the event is gone
  5. per target, and it must be a queue

basics

~20 s

EventBridge retries a failing target with exponential backoff for up to 24 hours by default. You bound that with a retry policy setting maximum attempts and maximum event age, and you attach a dead-letter queue — an SQS queue, configured per target — so exhausted or non-retryable events are captured instead of dropped.

solid answer

~40 s

Delivery to a target is asynchronous and at-least-once. When an invocation fails, EventBridge retries with exponential backoff, by default for up to 24 hours. You bound that per target with a `RetryPolicy` carrying `MaximumRetryAttempts` and `MaximumEventAgeInSeconds` — a short age is what you want for anything time-sensitive, because a stale event delivered hours later can be worse than no event. Critically, if retries are exhausted and there is no dead-letter queue, **the event is discarded with no trace**. So attach a `DeadLetterConfig` pointing at an SQS queue, on each target, not on the rule. The DLQ also catches non-retryable failures — a deleted target, a missing permission — which go straight there without retrying. Then alarm on the rule's `FailedInvocations` and on `DeadLetterInvocations`, and make the target idempotent, because retries mean duplicates.

code

bash · 6 lines
bash
aws events put-targets --rule order-alerts --event-bus-name orders-bus --targets '[{
  "Id": "processor",
  "Arn": "arn:aws:lambda:eu-west-1:111122223333:function:order-processor",
  "RetryPolicy": { "MaximumRetryAttempts": 10, "MaximumEventAgeInSeconds": 900 },
  "DeadLetterConfig": { "Arn": "arn:aws:sqs:eu-west-1:111122223333:order-target-dlq" }
}]'

go deeper

for a junior

Know that EventBridge retries a target that fails, and that you attach a dead-letter queue so events that keep failing are kept rather than dropped.

for a middle

Explain the two retry-policy bounds — maximum attempts and maximum event age — that the DLQ is per target and must be SQS, and that non-retryable errors skip retries and go straight there.

for a senior

Show the operating picture: alarm on FailedInvocations, DeadLetterInvocations and InvocationsFailedToBeSentToDlq, own the redrive path, and make targets idempotent because delivery is at-least-once.

for a principal

Own the failure architecture: how long an event stays meaningful, whether retries amplify pressure on a fragile downstream, and when a queue belongs between the rule and the service so backpressure exists at all.

## Asynchronous, at-least-once delivery Once a rule matches, invoking its targets is EventBridge's problem, not the publisher's. The publisher already got its `PutEvents` acknowledgement and is gone. That makes the delivery path the part you have to design for failure, and it has three knobs: retries, a dead-letter queue, and the metrics that tell you either is happening. The contract is **at-least-once**. A target can see the same event more than once — a retry after a timeout where the target actually succeeded is the everyday case. Idempotency is therefore not optional: dedupe on the event's `id`, or make the operation naturally idempotent (upsert rather than insert, set-state rather than increment). ## Retries By default EventBridge retries a failed target invocation with exponential backoff for up to 24 hours. That default is generous and often wrong. A per-target `RetryPolicy` bounds it: - **`MaximumRetryAttempts`** — the attempt cap (up to 185 as of 2025). - **`MaximumEventAgeInSeconds`** — how old an event may get before EventBridge gives up (60 to 86,400 seconds as of 2025). Whichever is reached first ends the retries. The age bound is usually the one that matters, because it encodes *freshness*: an "order placed" notification arriving twenty hours late may cause more damage than one never arriving. Set the age to the window in which the event is still meaningful, and let the DLQ hold the rest. Not every failure is retried. Errors EventBridge knows are permanent — the target no longer exists, the invocation is denied by permissions, a KMS key cannot be used — are not worth backing off against, so they go **straight to the dead-letter queue**. This is a useful diagnostic: a DLQ filling instantly rather than gradually usually means a policy or ARN problem, not a flaky downstream. ## The dead-letter queue A few properties trip people up: - It is configured **per target**, in that target's `DeadLetterConfig`. A rule with three targets can have three different DLQs, or only one of them protected. - It must be an **SQS queue**. Not SNS, not S3, not another bus. - The queue's own policy must allow EventBridge to send to it, and if the queue is encrypted with a customer managed KMS key, the key policy has to permit it too — a very common reason a DLQ stays mysteriously empty while events vanish. - The dead-lettered message carries the original event as its body, with **message attributes describing the failure** — the rule ARN, the target ARN, an error code and error message, and the condition that exhausted retries. That metadata is what turns a redrive into a five-minute job instead of an archaeology exercise. And the default position is the dangerous one: **no DLQ means the event is gone.** There is no automatic archive of failures, no notification to the publisher, and nothing in the target's logs, because the target never ran. ## Making it observable The metrics that matter, per rule: - **`TriggeredRules`** — the pattern matched. Zero here means a matching problem, not a delivery problem. - **`Invocations`** and **`FailedInvocations`** — invocation attempts and the ones that failed permanently. - **`DeadLetterInvocations`** — events written to the DLQ. Alarm on this being greater than zero over a short window; it is the cleanest "we are losing work" signal. - **`InvocationsFailedToBeSentToDlq`** — the second-order failure, usually a queue policy or KMS problem. Alarm on it too; without it, a broken DLQ looks exactly like a healthy one. Also alarm on the DLQ's own `ApproximateNumberOfMessagesVisible`, and treat that queue as work with an owner rather than a graveyard. The recovery path is to fix the target, then redrive the queue back through processing. ## Design implications Two consequences are worth stating out loud in an interview. First, retries plus at-least-once means **every target is a distributed-systems participant**: it must tolerate duplicates and out-of-order arrival. Second, retries are a form of backpressure amplification — a downstream that is failing because it is overloaded will receive escalating retry traffic from every event in flight. When the target is fragile or slow, the right shape is often to make the target an SQS queue and let a consumer drain it at its own pace, rather than pointing the rule directly at the fragile service.

  • Why bound MaximumEventAgeInSeconds instead of leaving the 24-hour default?
    Because freshness is part of correctness. An event that triggers a notification, a cache invalidation or a time-limited offer is worthless — sometimes harmful — hours later, and a long window also keeps a large backlog in flight that all lands at once when the target recovers. Set the age to the period the event still means something, and let the DLQ hold whatever misses it.
  • The DLQ is configured but stays empty while events are clearly being lost. What do you check?
    Look at the `InvocationsFailedToBeSentToDlq` metric — it exists exactly for this. The usual causes are the SQS queue policy not allowing EventBridge to send, or a customer managed KMS key on the queue whose key policy does not permit the service. Both make a broken DLQ look identical to a healthy one from the rule's side.
  • The target is a slow, fragile internal service. Is pointing the rule straight at it a good design?
    Usually not. Direct invocation means retry traffic scales with the failure, so an overloaded service gets hit harder as it degrades. Make an SQS queue the target instead and let a consumer drain it at a rate the service tolerates. You gain backpressure, a natural buffer, and per-message retry with its own DLQ.
  • How do you keep duplicate deliveries from causing damage?
    Design the target to be idempotent. Either dedupe on the event's `id` in a store with a TTL, or make the operation naturally repeatable — upsert rather than insert, assign a state rather than increment a counter. Never rely on "it hardly ever happens": at-least-once delivery means duplicates are a normal event, not an incident.

saying these in an interview costs you the question

  • Assuming failed events are stored somewhere by default
  • Configuring the dead-letter queue on the rule instead of the target
  • Thinking the DLQ can be an SNS topic or an S3 bucket
  • Believing delivery is exactly-once so targets need no idempotency
  • Leaving the 24-hour retry window for time-sensitive events

context