skip to content

An AWS Step Functions Standard execution failed near the end of a long workflow after a downstream outage. What does redriving that execution do, and when is an execution not eligible for it?

level: seniorimportance: nice to knowfreq 32%

answer

  1. resume, do not restart
  2. completed steps stay completed
  3. same identity, longer history
  4. one workflow type only
  5. the window does not stay open

basics

~20 s

Redrive restarts a failed Step Functions Standard execution from its first unsuccessful state, keeping the same execution ARN and history and skipping everything that already succeeded. Only failed, aborted or timed-out Standard executions are eligible, and only for a limited window after they ended.

solid answer

~50 s

Redrive, available through the `RedriveExecution` API or the console, resumes a Standard execution from the point of failure rather than from the beginning. States that already succeeded are not re-run, so their side effects are not duplicated and you do not pay for those state transitions again; the execution keeps its original ARN and name and the new events are appended to the same history, with `DescribeExecution` reporting a redrive count. Eligibility is narrow: Standard workflows only — Express executions cannot be redriven — and the execution must have failed, been aborted, or timed out, within the redrive window after it ended (14 days as of 2025). A successful execution is never redrivable. Redrive is for re-running the same work once the *downstream* cause is fixed; if the workflow logic itself was wrong, start a fresh execution instead.

go deeper

for a junior

Know that a failed Step Functions execution can be redriven, and that redrive resumes from the failed state rather than starting the workflow over from the first state.

for a middle

Explain the eligibility rules — Standard workflows only, only failed, aborted or timed-out executions, and only within a limited window — and that successful states are not re-run.

for a senior

Show the operational pattern: catchers record enough context and fail deliberately, an alarm brings in a human, the external cause is fixed, then affected executions are redriven in bulk safely because identity and completed work are preserved.

for a principal

Frame it as a workflow-type decision. Needing mid-workflow recovery is an argument for Standard over Express, and the recovery runbook — who redrives, on what evidence, within what window — belongs in the design, not in an incident channel.

## The problem redrive solves Before redrive existed, a Standard execution that failed at step twelve of fifteen left you with two bad options. Start a new execution from the beginning and re-run the eleven expensive, already-successful steps — duplicating their side effects, paying for their state transitions again, and possibly re-charging a customer. Or hand-craft a second state machine that begins at the failure point and feed it a reconstructed payload, which is exactly the kind of one-off operational script that rots. Redrive replaces both. Introduced in November 2023, it restarts a failed execution **from its first unsuccessful state**, with the input that state originally received. ## What it actually does - **Skips successful work.** Every state that already succeeded stays succeeded. Only the failed state and everything after it runs. - **Keeps identity.** The redriven run reuses the same state machine ARN, execution ARN and execution name. This matters more than it first appears: any idempotency key derived from the execution — the context object exposes `$$.Execution.Name` — is stable across a redrive, so downstream deduplication keeps working. - **Appends to the same history.** The new events land in the existing execution history rather than a separate one, so the whole story of the attempt stays in one place. `DescribeExecution` reports how many times the execution has been redriven and when it last was. - **Works on partial failures too.** A `Map` run whose child executions partially failed can be redriven so only the failed items re-run, rather than the whole batch. You trigger it with the `RedriveExecution` API, from the CLI, or with the Redrive button in the console on a failed execution. ## When it is not available This is the half interviewers actually probe, because it constrains architecture: 1. **Standard workflows only.** Express workflows are not redrivable. Express executions do not retain the durable, per-state execution history that redrive resumes from — one of the concrete reasons the Standard-versus-Express choice is not purely about cost and duration. 2. **The execution must have ended unsuccessfully** — failed, aborted, or timed out. A successful execution is not redrivable; there is nothing to resume. 3. **There is a time window.** Redrive is available for a bounded period after the execution ended — 14 days as of 2025. Beyond that the execution is no longer eligible, and `DescribeExecution` reports its redrive status accordingly. Check that status programmatically rather than assuming, because it also tells you when an execution can only be redriven through its parent Map run. ## What redrive is not Redrive re-runs the *same work* on the assumption that the reason it failed has been removed — a dependency that was down is back, a permission that was missing has been granted, a quota has been raised. It is an operational recovery tool, not a deployment mechanism: if the workflow's own logic was wrong, the honest move is to fix the definition and start a fresh execution, because the failed one was doing the wrong thing, not the right thing at the wrong time. Nor is it a substitute for designing failure paths. Retriers still handle transient errors inside a single execution, catchers still route to compensating branches, and both operate automatically at machine speed. Redrive is the manual, human-in-the-loop layer above them, for the case where automated handling correctly gave up because the problem was outside the workflow's power to solve. ## Operating with it The practical pattern is: catchers route irrecoverable failures to a state that records enough context to diagnose them and then fails the execution deliberately; an alarm on failed executions brings a human in; the human fixes the external cause and redrives the affected executions, possibly in bulk from a list of failed execution ARNs. Because redrive preserves identity and skips completed work, that bulk recovery is safe to run even against executions that got most of the way through — which is precisely what makes it worth designing for rather than discovering during an incident.

  • Why can't Express workflow executions be redriven?
    Express executions do not retain the durable per-state execution history that redrive resumes from — their history goes to CloudWatch Logs rather than being kept as replayable execution state. If mid-workflow recovery matters for a use case, that is an argument for Standard, alongside the usual duration and exactly-once considerations, rather than something you can add later.
  • Does redriving an execution change the idempotency keys downstream systems have already seen?
    No — a redriven run reuses the same execution ARN and execution name, so a key derived from `$$.Execution.Name` on the context object is identical to the one the first attempt used. That is deliberate: downstream deduplication keeps working, and a state that partially completed before failing will not be treated as new work by a system that already recorded that key.
  • When should you start a fresh execution instead of redriving?
    When the cause was inside the workflow rather than outside it. Redrive re-runs the same work on the assumption the external problem is fixed — a dependency restored, a permission granted, a quota raised. If the definition or the input was wrong, the failed execution was doing the wrong thing, and resuming it just does the wrong thing from a later state.

saying these in an interview costs you the question

  • Thinking redrive restarts the execution from the beginning
  • Expecting redrive to work on Express workflow executions
  • Assuming a successful execution can be redriven to re-run it
  • Believing redrive picks up a state machine fix you just deployed
  • Treating redrive as a replacement for retriers and catchers
  • Assuming failed executions stay redrivable indefinitely

context