For asynchronous AWS Lambda invocations, what is the difference between an on-failure destination and the older DeadLetterConfig dead-letter queue, and which should you configure on a new function?
answer
- both catch the event nobody is holding
- one carries the input, one carries the story
- the reason it failed is a field
- four target types versus two
- success has a slot too
basics
~20 sA dead-letter queue receives only the original event payload and can target SQS or SNS. An on-failure destination receives a richer record with the invocation context and the error response, and can also target EventBridge or another Lambda function. Prefer destinations on new functions.
solid answer
~50 sBoth exist to catch an asynchronous event whose invocations have all failed, so the payload is not silently dropped. The `DeadLetterConfig` on the function is the older mechanism: it takes a single SQS queue or SNS topic ARN and forwards the original event, with the request ID and error text tucked into message attributes. Destinations, configured through the function's event-invoke config, are richer in three ways. They emit a structured record containing the request context — including a `condition` field telling you *why* it failed, such as `RetriesExhausted` or `EventAgeExceeded` — plus both the original payload and the function's error response, so you can triage without correlating logs. They support SQS, SNS, EventBridge and Lambda as targets. And they have an on-success counterpart, which the DLQ has no equivalent for. Configure destinations on anything new; if a function somehow has both, the destination takes precedence.
code
bash · 4 linesaws lambda put-function-event-invoke-config \
--function-name my-function \
--maximum-retry-attempts 2 \
--destination-config '{"OnFailure":{"Destination":"arn:aws:sqs:eu-west-1:111122223333:failed-events"}}'go deeper
Know that an asynchronous event whose retries are exhausted is thrown away unless you configure somewhere for it to go, and that destinations are the modern way to do that.
Contrast the two precisely: DLQ takes SQS or SNS and forwards the raw event; destinations take four target types, carry request and response context plus a condition field, and add an on-success slot.
Bring the operational failure modes — execution-role permissions on the target, the DestinationDeliveryFailures metric, alarming on failure-queue depth, and keeping replay safe because the handler is idempotent.
Set the standard: every asynchronous function ships with a failure destination and a named owner for draining it, and decide whether failures route to a per-function queue or a central bus that triage tooling subscribes to.
## Why either one exists An asynchronously invoked function has no response channel. When Lambda exhausts its retries or the event exceeds its maximum age, the event is discarded and the original caller — long gone, holding a 202 — never learns anything. Both mechanisms in this question exist to give that dying event somewhere to land. ## DeadLetterConfig: the original `DeadLetterConfig` is a property of the function itself, holding a single `TargetArn` that must be an SQS queue or an SNS topic. When the retries run out, Lambda forwards the **original event payload** to that target. Diagnostic information travels as message attributes rather than in the body: the request ID, an error code and an error message. ```bash aws lambda update-function-configuration \ --function-name my-function \ --dead-letter-config TargetArn=arn:aws:sqs:eu-west-1:111122223333:my-dlq ``` It works, and plenty of production functions still use it. Its limits are that you get the input but not the function's actual error response, you cannot tell from the body why the event died, and there is no equivalent for successful invocations. ## Destinations: the current mechanism Destinations are set on the function's asynchronous invocation configuration through `PutFunctionEventInvokeConfig`, in a `DestinationConfig` with two independent slots: `OnFailure` and `OnSuccess`. ```bash aws lambda put-function-event-invoke-config \ --function-name my-function \ --destination-config '{"OnFailure":{"Destination":"arn:aws:sqs:eu-west-1:111122223333:failed-events"}}' ``` Four target types are supported: an SQS queue, an SNS topic, an EventBridge event bus, and another Lambda function. The EventBridge and Lambda targets are what make destinations feel like a routing primitive rather than a rubbish bin — you can fan a failure out to several handlers by rule, or hand it to a compensating function directly. What arrives is not the bare event but a JSON record wrapping it, carrying: - `requestContext` — including `requestId`, `functionArn`, `approximateInvokeCount`, and a `condition` field naming the reason the event ended up here: `RetriesExhausted`, `EventAgeExceeded`, or `ZeroReservedConcurrency`. - `requestPayload` — the original event, so replay is possible. - `responseContext` and `responsePayload` — the function's own error output. That `condition` field is the practical difference. "The dependency was down and we exhausted retries" and "this event sat in the queue for six hours" and "the function had zero reserved concurrency" are three completely different incidents, and with a DLQ you would be reconstructing which one happened from logs. ## Things that catch people out **Permissions.** The delivery is performed using the function's **execution role**, so that role needs `sqs:SendMessage`, `sns:Publish`, `events:PutEvents` or `lambda:InvokeFunction` on the target as appropriate. Miss it and the capture silently fails — the whole safety net does nothing. Watch the `DestinationDeliveryFailures` metric (and `DeadLetterErrors` for the older mechanism); an alarm on either is not optional if you rely on the capture. **Precedence.** If a function has both a `DeadLetterConfig` and an `OnFailure` destination, the destination is used. Do not configure both and expect two copies. **Scope.** Destinations described here apply to asynchronous invocation. Poller-based event sources have their own failure-handling configuration, which is a different mechanism with different fields. **Encryption.** If the target queue or topic uses a customer managed KMS key, the execution role needs key permissions too — a common cause of a capture path that tests fine against an unencrypted queue and fails in production. ## Which to choose For a new function: an on-failure destination, essentially always. Use `OnSuccess` more sparingly — it is genuinely useful for chaining a downstream step off a successful asynchronous invocation without the function knowing about its successor, but it also doubles your event volume into whatever consumes it. Keep the older DLQ only where it already exists and works, and treat migrating it as low-risk cleanup rather than urgent. Whichever you pick, the capture is only half the job. Someone has to look at the failure target, and a queue of dead events nobody drains is a compliance artefact, not a recovery mechanism. Alarm on its depth, and make replay safe by keeping the handler idempotent.
- Which permissions does the delivery to a failure destination use, and what happens if they are missing?Lambda uses the function's own execution role, so it needs sqs:SendMessage, sns:Publish, events:PutEvents or lambda:InvokeFunction on the target — plus KMS key permissions if the target is encrypted with a customer managed key. Without them the delivery fails silently and only the DestinationDeliveryFailures metric reveals it.
- What does the condition field on a destination record tell you?Why the event ended up at the failure destination: RetriesExhausted means the handler kept failing, EventAgeExceeded means the event outlived MaximumEventAgeInSeconds while queued, and ZeroReservedConcurrency means the function had no capacity allocated at all. Three different incidents that a bare dead-letter payload would leave you to infer from logs.
- When is an OnSuccess destination worth configuring?When you want to chain a follow-on step off a successful asynchronous invocation without the function knowing about its successor — routing the result to an EventBridge bus, for example, so consumers subscribe by rule. Use it deliberately: it emits a record for every successful invocation, which can be a very large volume.
saying these in an interview costs you the question
- Thinks a dead-letter queue is configured by default
- Believes the DLQ receives the function's error response
- Configures both a DLQ and a destination expecting two copies
- Forgets the execution role needs write access to the target
- Treats a filling failure queue as a solved problem