skip to content

You are designing a workflow backed by AWS Lambda and must decide whether the caller retries a failed synchronous invoke, or the work is handed to Lambda's asynchronous invocation path so Lambda retries it. How do you make that call, and what does each choice cost you?

level: principalimportance: should knowfreq 38%

answer

  1. start from who is waiting for the answer
  2. the visible error versus the durable event
  3. minutes of delay is the fixed price
  4. retries at two layers multiply
  5. one layer owns the policy

basics

~20 s

Decide by who must know the outcome and how long they can wait. Caller-side retry keeps the error visible and the policy yours, but multiplies in-flight load during an outage. Lambda-side asynchronous retry buys durability and backpressure at the cost of a coarse, untunable policy and no response channel.

solid answer

~50 s

The question is really where the *authority* over failure should sit. If a human or an upstream service is waiting on the answer, retry has to be synchronous and caller-side: only the caller knows its own latency budget, and only it can decide between retrying, degrading, and failing fast. The cost is that retries multiply concurrent in-flight work exactly when the system is already struggling, so they need capped attempts, jittered backoff and a circuit breaker, or they become the outage. If nobody is waiting, hand the event to the asynchronous path: Lambda takes durable custody, absorbs throttles as queue depth instead of errors, and retries for you. The costs are real too — the policy is coarse (at most two retries, minute-scale delays, a six-hour age ceiling), the caller learns nothing, and failure visibility must be engineered through destinations and metrics. What you must never do is layer both without thinking, because retry counts multiply.

go deeper

for a junior

Know the basic split: a synchronous caller sees the error and decides what to do, while asynchronous invocation hands the retry to Lambda and returns immediately.

for a middle

Explain the concrete constraints on each side — Lambda's asynchronous policy is capped at two retries with minute-scale delays, while caller-side retry needs capped attempts, backoff and jitter.

for a senior

Show you have operated this: describe retry amplification during an outage, shared account concurrency as the resource that gets starved, and how you make asynchronous failures visible with destinations and alarms.

for a principal

Own the retry budget as an explicit architectural contract — one nominated owner per path, fail-fast everywhere else, idempotency as a precondition, and a clear rule for when the built-in policy has been outgrown.

## Reframe the question "Who retries" sounds like a configuration detail. It is actually a decision about where the authority over failure lives, and it determines the contract your system offers: does the caller get an answer, or a promise? ## The deciding questions **Is anyone waiting on the result?** If a user or an upstream service needs the outcome to proceed, you are synchronous by necessity. If the work only has to happen eventually, asynchronous is available and usually better. **What is the latency budget?** Lambda's asynchronous retries are spaced in minutes. If the work must complete inside a few seconds, that policy is unusable regardless of how appealing the durability is. **What happens if it never completes?** Work with financial or legal consequence needs durable custody and a capture path. Best-effort work — a cache warm, an enrichment — can be dropped. **Is the handler idempotent?** If it is not, every retry mechanism is dangerous and the design work is idempotency, not retry placement. ## Caller-side retry: control, and a loaded gun Retrying a synchronous invoke keeps everything visible. The caller sees the error immediately, applies its own policy, can distinguish a retryable dependency failure from a permanent validation error, and can degrade gracefully — serve stale data, queue for later, tell the user honestly. The danger is amplification. Each retry is another concurrent execution. Three attempts per caller against a slow dependency means three times the in-flight requests and three times the concurrency consumed, at precisely the moment the system is least able to afford it — and Lambda's concurrency is a shared account-level resource, so an amplifying caller starves functions that have nothing to do with the incident. That is the metastable failure pattern: the retries keep the system down after the original trigger has passed. So caller-side retry is only safe with the full discipline: a small attempt cap, exponential backoff **with jitter**, retrying only errors that can plausibly succeed on a second try, a circuit breaker that stops trying when the dependency is clearly down, and a deadline that ends the whole thing. Modern AWS SDKs give you part of this — configurable retry modes and a maximum attempt count — but their policy covers throttling and service errors, not your application's semantics, so you still own the decision about *your* failures. ## Lambda-side asynchronous retry: durability, and a policy you do not control Invoking asynchronously moves custody to AWS. The caller gets a 202 in milliseconds and is finished. Lambda holds the event durably, retries handler failures a bounded number of times, and — importantly — keeps re-attempting throttled invocations until the event's age limit, turning a concurrency shortage into queue depth instead of a cascade of errors. That backpressure property is the strongest argument for this path. The costs, stated honestly: - **The policy is coarse and nearly untunable.** At most two retries, minute-scale delays you do not choose, no per-error-class behaviour, no jitter control. - **The caller is blind.** No result, no error, ever. Failure observability has to be built: an on-failure destination so payloads survive, alarms on `Errors` and on delivery failures to the target, and attention to `AsyncEventAge` as the backlog signal. - **Ordering and timing are unspecified.** Events may arrive out of order and may be delivered more than once. - **A poison event is expensive.** It burns three invocations before dying, and a flood of them triples volume and cost. ## The third option, and why to name it When neither policy fits — you need ten attempts, or hours of backoff, or explicit per-message failure handling — the honest answer is that you want an explicit queue between the caller and the function, so the retry policy becomes yours to write rather than AWS's to impose. Say so in an interview: recognising that the built-in asynchronous policy is a convenience with fixed parameters, and knowing when your requirements have outgrown it, is exactly the judgment being probed. ## The rule about layering Retries compose multiplicatively. A caller with three attempts, invoking asynchronously into a path with three attempts, against a handler that itself retries a dependency three times, produces up to twenty-seven executions of one logical request. Pick **one** layer to own retry, make every other layer fail fast and propagate, and write that down where the next team can find it. Retry budgets that are implicit in three separate codebases are how a small dependency wobble becomes an all-hands incident. ## What a strong answer sounds like Start from who is waiting; pick the model that matches; state the cost you are accepting; name the single layer that owns retry; require idempotency in the handler regardless; and describe how failures become visible in each case — a returned error for the synchronous path, a destination plus alarms for the asynchronous one.

  • Why is retrying failed synchronous invokes aggressively dangerous during a partial outage?
    Every retry is another concurrent execution, so attempts multiply in-flight load exactly when the system is weakest, and Lambda concurrency is shared account-wide — an amplifying caller starves unrelated functions. This is the metastable pattern where retries sustain the outage after its trigger has passed. Cap attempts, jitter backoff, and break the circuit.
  • What is the case for putting an explicit queue between the caller and the function instead of using asynchronous invocation?
    Control over the retry policy. The built-in asynchronous path gives at most two retries at delays you do not choose, with no per-error behaviour. When you need many attempts, long backoff, explicit per-message failure handling or your own dead-letter semantics, you have outgrown the convenience and want the policy to be yours.
  • How do you stop retry counts multiplying across layers?
    Nominate one layer as the retry owner and make every other layer fail fast and propagate. A caller retrying three times into an asynchronous path retrying three times, over a handler retrying its dependency three times, is up to twenty-seven executions of one request. Write the chosen owner down; implicit budgets across codebases are how incidents scale.
  • Does the choice change if the handler is not idempotent?
    It makes idempotency the prerequisite work rather than a nicety. Every retry mechanism — caller-side, Lambda's asynchronous path, an explicit queue — can execute the same logical request twice, and so can a client timeout on a call that actually succeeded. Until duplicate execution is harmless, no retry placement is safe.

saying these in an interview costs you the question

  • Treats asynchronous invocation as free durability with no downsides
  • Adds retries at every layer without a budget
  • Retries non-retryable errors such as validation failures
  • Uses fixed-interval retries with no jitter or circuit breaker
  • Assumes Lambda's async retry delays can be tuned per error class

context