skip to content

In an AWS Step Functions Task state, how do the Retry and Catch fields interact, and which retrier fields control the wait between attempts?

level: middleimportance: must knowfreq 72%

answer

  1. two arrays, one runs before the other
  2. exhaust attempts, then hand it over
  3. interval times rate to the power n
  4. zero attempts is a deliberate setting
  5. the whole state runs again

basics

~20 s

Step Functions matches a failure against the Retry array first and re-runs the state until that retrier's MaxAttempts is used up. Only then is the error offered to Catch. IntervalSeconds, BackoffRate and MaxAttempts define the wait between attempts.

solid answer

~40 s

`Retry` and `Catch` are two ordered arrays on the same state, and they run in sequence rather than in parallel. When the state fails, Step Functions scans `Retry` top to bottom for the first retrier whose `ErrorEquals` matches, waits, and re-runs the whole state. The wait for attempt *n* is `IntervalSeconds * BackoffRate^(n-1)` — the defaults are `IntervalSeconds` 1, `BackoffRate` 2.0 and `MaxAttempts` 3, so a bare retrier gives you roughly 1s, 2s, 4s. `MaxDelaySeconds` caps that growth and `JitterStrategy: "FULL"` randomises each wait. Once that retrier's attempts are exhausted, the error moves to `Catch`, and the first catcher whose `ErrorEquals` matches sends the workflow to its `Next` state instead of failing. If nothing matches in either array, the state fails and the failure propagates up.

code

json · 28 lines
json
{
  "Comment": "Retriers are tried first; only survivors reach Catch",
  "StartAt": "ChargeCard",
  "States": {
    "ChargeCard": {
      "Type": "Task",
      "Resource": "arn:aws:states:::lambda:invoke",
      "Parameters": { "FunctionName": "charge-card", "Payload.$": "$" },
      "Retry": [
        { "ErrorEquals": ["CardDeclined"], "MaxAttempts": 0 },
        {
          "ErrorEquals": ["States.ALL"],
          "IntervalSeconds": 2,
          "BackoffRate": 2,
          "MaxAttempts": 4,
          "MaxDelaySeconds": 30,
          "JitterStrategy": "FULL"
        }
      ],
      "Catch": [
        { "ErrorEquals": ["States.ALL"], "Next": "RecordFailure" }
      ],
      "Next": "ShipOrder"
    },
    "ShipOrder": { "Type": "Succeed" },
    "RecordFailure": { "Type": "Fail", "Error": "ChargeFailed" }
  }
}

go deeper

for a junior

Be able to say that Retry re-runs the state and Catch routes to another state, that retries happen first, and that IntervalSeconds, BackoffRate and MaxAttempts control the waiting.

for a middle

Explain the ordering precisely: first matching retrier only, its own attempt counter, the interval times backoff-rate formula, and the hand-off to Catch once attempts are exhausted. Know the defaults of 1 second, rate 2 and 3 attempts.

for a senior

Show you decide what deserves a retry at all. Separate transient service errors from deterministic ones, pin the latter at MaxAttempts 0, add jitter when many executions fail together, and raise idempotency before anyone asks.

for a principal

Own the retry budget across layers — SDK retries inside the task, the state machine's retriers, and any outer re-invocation — so a brief dependency outage is not amplified into a much longer one, and set the house convention for how failure branches are structured.

## Two arrays, one order of operations A `Task`, `Parallel` or `Map` state in Amazon States Language can carry two error-handling fields, both arrays of objects: - **`Retry`** — a list of *retriers*. A retrier says "for these error names, run this state again, this many times, waiting like this." - **`Catch`** — a list of *catchers*. A catcher says "for these error names, do not fail; go to this other state instead." They are not alternatives, and they do not race. The order is fixed and worth stating plainly in an interview: **retry first, catch second**. A failure is offered to `Retry`; only a failure that survives retrying reaches `Catch`. ## How a retrier is chosen and what it does When the state reports an error, Step Functions scans the `Retry` array from top to bottom and picks the **first** retrier whose `ErrorEquals` contains the error name. Only that retrier is used. A later retrier that also matches the name never takes over when the first one is exhausted — the error goes to `Catch` instead. Each retrier keeps its own attempt counter, scoped to that state and that execution. A retry re-runs the **entire state**, not "the failed part of it". The state's input is unchanged, so the invoked work must tolerate being called again — this is where idempotency enters the conversation. The fields: | Field | Default | Meaning | | --- | --- | --- | | `ErrorEquals` | (required) | Error names this retrier handles | | `IntervalSeconds` | 1 | Wait before the first retry | | `MaxAttempts` | 3 | How many retries (0 = never retry) | | `BackoffRate` | 2.0 | Multiplier applied to each successive wait | | `MaxDelaySeconds` | — | Ceiling on the computed wait | | `JitterStrategy` | `NONE` | `FULL` randomises the wait | The wait before retry *n* is `IntervalSeconds * BackoffRate^(n-1)`. With `IntervalSeconds: 2, BackoffRate: 2, MaxAttempts: 3` the waits are 2s, 4s, 8s. `MaxDelaySeconds` clamps that so a long `MaxAttempts` on an aggressive `BackoffRate` cannot drift into hours; `JitterStrategy: "FULL"` spreads the retries of many concurrent executions instead of having them all hammer the recovering dependency at the same instant, which matters when hundreds of executions failed on the same downstream blip. Note `MaxAttempts: 0`. It is legal and useful: it names an error as explicitly non-retriable so that a later wildcard retrier does not pick it up, sending it straight to `Catch`. ## What Catch does with the survivor Once retries are exhausted (or no retrier matched at all), the `Catch` array is scanned in order and the first matching catcher wins. Its `Next` names the state to run — typically a cleanup branch, a notification, or a state that records the failure and ends the workflow deliberately rather than crashing it. The catcher can also reshape what that state receives, via `ResultPath`, which is a separate topic worth knowing well because the default discards the state's original input. If no catcher matches, the state fails. Inside a `Parallel` or `Map` branch, that failure surfaces to the parent as `States.BranchFailed`, which the parent state can retry or catch in turn — so error handling nests. ```json "ChargeCard": { "Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "Parameters": { "FunctionName": "charge-card", "Payload.$": "$" }, "Retry": [ { "ErrorEquals": ["CardDeclined"], "MaxAttempts": 0 }, { "ErrorEquals": ["States.ALL"], "IntervalSeconds": 2, "BackoffRate": 2, "MaxAttempts": 4, "MaxDelaySeconds": 30, "JitterStrategy": "FULL" } ], "Catch": [ { "ErrorEquals": ["States.ALL"], "Next": "CompensateOrder" } ], "Next": "ShipOrder" } ``` ## The judgment interviewers are listening for The mechanics are easy; the discrimination is not. Retry is for failures that are plausibly **transient** — throttling, a connection reset, a service exception. Retrying a deterministic failure such as a validation error or a declined card just burns state transitions and delays the inevitable, which is why the example pins `CardDeclined` at `MaxAttempts: 0`. Catch is for failures that need a **different path**: cleanup, a manual-review branch, a recorded outcome. The second thing to say out loud is that a retried task runs again with the same input, so if the first attempt had already reached the downstream system before failing, the effect happens twice. Retries are a correctness decision, not just a resilience knob.

  • What do MaxDelaySeconds and JitterStrategy add to a retrier?
    `MaxDelaySeconds` caps the computed wait, so an aggressive `BackoffRate` with many attempts cannot grow into an absurd delay. `JitterStrategy: "FULL"` randomises each wait between zero and the computed interval. Jitter matters when a downstream outage failed hundreds of executions at once: without it they all retry in lockstep and re-overload the dependency the moment it recovers.
  • If two retriers both match an error and the first one runs out of attempts, does the second take over?
    No. Step Functions picks the first retrier whose `ErrorEquals` matches and uses only that one. When its `MaxAttempts` is exhausted the error goes to `Catch`, not to a later matching retrier. That is why ordering is a design decision: put the narrow, well-tuned retriers first and treat a `States.ALL` retrier as the fallback for everything not named above it.
  • Your Task state retries a non-idempotent API call. What can go wrong?
    A retry re-runs the whole state with the same input, and Step Functions cannot tell "the call failed" from "the call succeeded but the response was lost". So a timeout on a request that actually landed produces a duplicate side effect — a double charge, a duplicate order. The fix is on the downstream side: pass a stable idempotency key, for example one derived from the execution name available on the context object.

saying these in an interview costs you the question

  • Thinking Catch fires immediately, before retries are attempted
  • Assuming a retry re-runs only the failed call, not the whole state
  • Treating MaxAttempts as total attempts rather than retries
  • Retrying deterministic failures like validation errors or declined payments
  • Believing a second matching retrier takes over when the first is exhausted
  • Ignoring that retries duplicate side effects on non-idempotent calls

context