skip to content

Step Functions

You will learn Step Functions as managed orchestration: a JSON state machine that calls AWS services directly, keeps the workflow state for you, and makes long-running or human-in-the-loop steps durable. Interviewers ask it as the orchestration half of the choreography-versus-orchestration question.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

11

In an AWS Step Functions Task state, how do the Retry and Catch fields interact, and which retrier fields control the wait between attempts?

level: middleimportance: must knowfreq 72%

answer

  1. two arrays, one runs before the other
  2. exhaust attempts, then hand it over
  3. interval times rate to the power n
  4. zero attempts is a deliberate setting
  5. the whole state runs again

basics

~20 s

Step Functions matches a failure against the Retry array first and re-runs the state until that retrier's MaxAttempts is used up. Only then is the error offered to Catch. IntervalSeconds, BackoffRate and MaxAttempts define the wait between attempts.

solid answer

~40 s

`Retry` and `Catch` are two ordered arrays on the same state, and they run in sequence rather than in parallel. When the state fails, Step Functions scans `Retry` top to bottom for the first retrier whose `ErrorEquals` matches, waits, and re-runs the whole state. The wait for attempt *n* is `IntervalSeconds * BackoffRate^(n-1)` — the defaults are `IntervalSeconds` 1, `BackoffRate` 2.0 and `MaxAttempts` 3, so a bare retrier gives you roughly 1s, 2s, 4s. `MaxDelaySeconds` caps that growth and `JitterStrategy: "FULL"` randomises each wait. Once that retrier's attempts are exhausted, the error moves to `Catch`, and the first catcher whose `ErrorEquals` matches sends the workflow to its `Next` state instead of failing. If nothing matches in either array, the state fails and the failure propagates up.

code

json · 28 lines
json
{
  "Comment": "Retriers are tried first; only survivors reach Catch",
  "StartAt": "ChargeCard",
  "States": {
    "ChargeCard": {
      "Type": "Task",
      "Resource": "arn:aws:states:::lambda:invoke",
      "Parameters": { "FunctionName": "charge-card", "Payload.$": "$" },
      "Retry": [
        { "ErrorEquals": ["CardDeclined"], "MaxAttempts": 0 },
        {
          "ErrorEquals": ["States.ALL"],
          "IntervalSeconds": 2,
          "BackoffRate": 2,
          "MaxAttempts": 4,
          "MaxDelaySeconds": 30,
          "JitterStrategy": "FULL"
        }
      ],
      "Catch": [
        { "ErrorEquals": ["States.ALL"], "Next": "RecordFailure" }
      ],
      "Next": "ShipOrder"
    },
    "ShipOrder": { "Type": "Succeed" },
    "RecordFailure": { "Type": "Fail", "Error": "ChargeFailed" }
  }
}

go deeper

for a junior

Be able to say that Retry re-runs the state and Catch routes to another state, that retries happen first, and that IntervalSeconds, BackoffRate and MaxAttempts control the waiting.

for a middle

Explain the ordering precisely: first matching retrier only, its own attempt counter, the interval times backoff-rate formula, and the hand-off to Catch once attempts are exhausted. Know the defaults of 1 second, rate 2 and 3 attempts.

for a senior

Show you decide what deserves a retry at all. Separate transient service errors from deterministic ones, pin the latter at MaxAttempts 0, add jitter when many executions fail together, and raise idempotency before anyone asks.

for a principal

Own the retry budget across layers — SDK retries inside the task, the state machine's retriers, and any outer re-invocation — so a brief dependency outage is not amplified into a much longer one, and set the house convention for how failure branches are structured.

## Two arrays, one order of operations A `Task`, `Parallel` or `Map` state in Amazon States Language can carry two error-handling fields, both arrays of objects: - **`Retry`** — a list of *retriers*. A retrier says "for these error names, run this state again, this many times, waiting like this." - **`Catch`** — a list of *catchers*. A catcher says "for these error names, do not fail; go to this other state instead." They are not alternatives, and they do not race. The order is fixed and worth stating plainly in an interview: **retry first, catch second**. A failure is offered to `Retry`; only a failure that survives retrying reaches `Catch`. ## How a retrier is chosen and what it does When the state reports an error, Step Functions scans the `Retry` array from top to bottom and picks the **first** retrier whose `ErrorEquals` contains the error name. Only that retrier is used. A later retrier that also matches the name never takes over when the first one is exhausted — the error goes to `Catch` instead. Each retrier keeps its own attempt counter, scoped to that state and that execution. A retry re-runs the **entire state**, not "the failed part of it". The state's input is unchanged, so the invoked work must tolerate being called again — this is where idempotency enters the conversation. The fields: | Field | Default | Meaning | | --- | --- | --- | | `ErrorEquals` | (required) | Error names this retrier handles | | `IntervalSeconds` | 1 | Wait before the first retry | | `MaxAttempts` | 3 | How many retries (0 = never retry) | | `BackoffRate` | 2.0 | Multiplier applied to each successive wait | | `MaxDelaySeconds` | — | Ceiling on the computed wait | | `JitterStrategy` | `NONE` | `FULL` randomises the wait | The wait before retry *n* is `IntervalSeconds * BackoffRate^(n-1)`. With `IntervalSeconds: 2, BackoffRate: 2, MaxAttempts: 3` the waits are 2s, 4s, 8s. `MaxDelaySeconds` clamps that so a long `MaxAttempts` on an aggressive `BackoffRate` cannot drift into hours; `JitterStrategy: "FULL"` spreads the retries of many concurrent executions instead of having them all hammer the recovering dependency at the same instant, which matters when hundreds of executions failed on the same downstream blip. Note `MaxAttempts: 0`. It is legal and useful: it names an error as explicitly non-retriable so that a later wildcard retrier does not pick it up, sending it straight to `Catch`. ## What Catch does with the survivor Once retries are exhausted (or no retrier matched at all), the `Catch` array is scanned in order and the first matching catcher wins. Its `Next` names the state to run — typically a cleanup branch, a notification, or a state that records the failure and ends the workflow deliberately rather than crashing it. The catcher can also reshape what that state receives, via `ResultPath`, which is a separate topic worth knowing well because the default discards the state's original input. If no catcher matches, the state fails. Inside a `Parallel` or `Map` branch, that failure surfaces to the parent as `States.BranchFailed`, which the parent state can retry or catch in turn — so error handling nests. ```json "ChargeCard": { "Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "Parameters": { "FunctionName": "charge-card", "Payload.$": "$" }, "Retry": [ { "ErrorEquals": ["CardDeclined"], "MaxAttempts": 0 }, { "ErrorEquals": ["States.ALL"], "IntervalSeconds": 2, "BackoffRate": 2, "MaxAttempts": 4, "MaxDelaySeconds": 30, "JitterStrategy": "FULL" } ], "Catch": [ { "ErrorEquals": ["States.ALL"], "Next": "CompensateOrder" } ], "Next": "ShipOrder" } ``` ## The judgment interviewers are listening for The mechanics are easy; the discrimination is not. Retry is for failures that are plausibly **transient** — throttling, a connection reset, a service exception. Retrying a deterministic failure such as a validation error or a declined card just burns state transitions and delays the inevitable, which is why the example pins `CardDeclined` at `MaxAttempts: 0`. Catch is for failures that need a **different path**: cleanup, a manual-review branch, a recorded outcome. The second thing to say out loud is that a retried task runs again with the same input, so if the first attempt had already reached the downstream system before failing, the effect happens twice. Retries are a correctness decision, not just a resilience knob.

  • What do MaxDelaySeconds and JitterStrategy add to a retrier?
    `MaxDelaySeconds` caps the computed wait, so an aggressive `BackoffRate` with many attempts cannot grow into an absurd delay. `JitterStrategy: "FULL"` randomises each wait between zero and the computed interval. Jitter matters when a downstream outage failed hundreds of executions at once: without it they all retry in lockstep and re-overload the dependency the moment it recovers.
  • If two retriers both match an error and the first one runs out of attempts, does the second take over?
    No. Step Functions picks the first retrier whose `ErrorEquals` matches and uses only that one. When its `MaxAttempts` is exhausted the error goes to `Catch`, not to a later matching retrier. That is why ordering is a design decision: put the narrow, well-tuned retriers first and treat a `States.ALL` retrier as the fallback for everything not named above it.
  • Your Task state retries a non-idempotent API call. What can go wrong?
    A retry re-runs the whole state with the same input, and Step Functions cannot tell "the call failed" from "the call succeeded but the response was lost". So a timeout on a request that actually landed produces a duplicate side effect — a double charge, a duplicate order. The fix is on the downstream side: pass a stable idempotency key, for example one derived from the execution name available on the context object.

saying these in an interview costs you the question

  • Thinking Catch fires immediately, before retries are attempted
  • Assuming a retry re-runs only the failed call, not the whole state
  • Treating MaxAttempts as total attempts rather than retries
  • Retrying deterministic failures like validation errors or declined payments
  • Believing a second matching retrier takes over when the first is exhausted
  • Ignoring that retries duplicate side effects on non-idempotent calls

context

open as a page

In an AWS Step Functions Task state, what is the difference between the default request-response integration, the .sync pattern, and .waitForTaskToken?

level: middleimportance: must knowfreq 66%

basics

~20 s

Request-response calls the API and moves on as soon as it returns. The .sync suffix makes Step Functions wait until the underlying job reaches a terminal state and returns its result. The .waitForTaskToken suffix pauses the execution until something calls SendTaskSuccess or SendTaskFailure with the injected token.

open as a page

In AWS Step Functions, how do Standard and Express workflows differ, and how would you choose between them?

level: middleimportance: must knowfreq 78%

basics

~20 s

Standard workflows are durable and auditable: they run up to a year, execute exactly once, keep a queryable execution history, and bill per state transition. Express workflows run up to five minutes, are at-least-once, log to CloudWatch, and bill by requests plus duration.

open as a page

In AWS Step Functions, what does the error name States.ALL match inside a Retry or Catch ErrorEquals array, and where is it allowed to appear?

level: juniorimportance: should knowfreq 55%

basics

~20 s

States.ALL is the Step Functions wildcard error name that matches nearly every error a state can raise. It must be the only entry in its ErrorEquals array and must sit in the final retrier or catcher, because they are evaluated top to bottom.

open as a page

In AWS Step Functions, what are the required top-level fields of an Amazon States Language definition, and how do StartAt, Next and End control the flow?

level: juniorimportance: should knowfreq 52%

basics

~20 s

An Amazon States Language definition requires two top-level fields: StartAt and States. StartAt names the first state to run, each state's Next names the state that follows it, and End set to true finishes that branch.

open as a page

In AWS Step Functions, what JSON does a Catch block pass to the target state, and what changes if the catcher sets ResultPath to "$.error"?

level: middleimportance: should knowfreq 45%

basics

~20 s

A Step Functions catcher passes an object with two fields, Error and Cause. By default that object replaces the failed state's input entirely; setting ResultPath to "$.error" instead nests it inside the original input, so the recovery state still sees the business data it needs.

open as a page

An AWS Step Functions Task state calls a worker that sometimes dies without reporting back, leaving executions stuck for hours. How do TimeoutSeconds and HeartbeatSeconds address this, and how do they differ?

level: seniorimportance: should knowfreq 50%

basics

~20 s

TimeoutSeconds caps a Step Functions task's total wall-clock duration and raises States.Timeout when exceeded. HeartbeatSeconds caps the gap between SendTaskHeartbeat calls from the worker and raises States.HeartbeatTimeout, detecting a dead worker long before a generous overall timeout would.

open as a page

A Step Functions workflow must process several million JSON objects stored in an S3 bucket. Why would you use a Distributed Map state instead of an inline Map?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An inline Map runs every iteration inside the parent execution, so the items must fit in the state payload and every iteration's events land in one execution history, which is bounded. Distributed Map reads items straight from S3 and runs them as separate child executions at far higher concurrency.

open as a page

Standard Step Functions workflows bill per state transition. When would you keep multi-step coordination inside application code or a plain event-driven chain on AWS instead of building a state machine?

level: principalimportance: should knowfreq 40%

basics

~20 s

Skip the state machine when the coordination is cheap, fast and stateless: very high-volume per-event work where per-transition billing dominates, latency-critical paths, and pure computation that belongs inside one function. Reach for it when you need durable state, waits, or an auditable trail per execution.

open as a page

An AWS Step Functions Standard execution failed near the end of a long workflow after a downstream outage. What does redriving that execution do, and when is an execution not eligible for it?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Redrive restarts a failed Step Functions Standard execution from its first unsuccessful state, keeping the same execution ARN and history and skipping everything that already succeeded. Only failed, aborted or timed-out Standard executions are eligible, and only for a limited window after they ended.

open as a page

AWS Step Functions caps the data passed into and out of a state. What is that limit, and how do you design a workflow whose steps produce large results?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

State input and output are capped at 256 KiB as of 2025, and exceeding it fails the execution with a data-limit error. Keep large data in S3 and pass object keys through the workflow, trimming each state's output so payloads never accumulate.

open as a page