In AWS Step Functions, how do Standard and Express workflows differ, and how would you choose between them?
answer
- two execution modes, fixed at creation
- one runs for a year, one for minutes
- one guarantee is exactly-once, the other is not
- history API versus CloudWatch Logs only
- per-transition billing versus duration billing
basics
~20 sStandard workflows are durable and auditable: they run up to a year, execute exactly once, keep a queryable execution history, and bill per state transition. Express workflows run up to five minutes, are at-least-once, log to CloudWatch, and bill by requests plus duration.
solid answer
~50 sThey are two execution modes of the same language, chosen when you create the state machine and not changeable afterwards. **Standard** is the durable, auditable mode: executions can run up to a year, the workflow executes exactly once, every state transition is recorded in a history you can query with `GetExecutionHistory` and see in the console for 90 days, and it supports the `.sync` and `.waitForTaskToken` integration patterns. It bills per state transition, so a chatty workflow at high volume gets expensive. **Express** is the high-volume mode: a five-minute ceiling, no execution-history API — you must ship logs to CloudWatch Logs to see anything — at-least-once semantics for asynchronous starts, and billing by number of requests plus duration and memory. Choose Standard for long-running or human-in-the-loop business processes; choose Express for short, high-throughput event processing where steps are idempotent. A common hybrid is a Standard parent that runs Express children.
code
bash · 3 linesaws stepfunctions start-sync-execution \
--state-machine-arn arn:aws:states:us-east-1:123456789012:stateMachine:price-quote-express \
--input '{"sku":"A-1001","qty":3}'go deeper
Know that the two types exist and that Standard is the durable long-running one while Express is for short, high-volume work. Naming the five-minute ceiling is enough at this level.
Explain the execution semantics — exactly-once for Standard, at-least-once for asynchronous Express — and why that forces idempotent steps. Be ready to describe the two billing shapes without quoting prices.
Demonstrate the operational consequences: Express gives you nothing to debug unless logging was configured in advance, and the missing .sync/callback support rules it out for anything that waits.
Own the crossover analysis for a real workload — executions per second times states per execution against duration-based billing — and the Standard-parent/Express-child split that keeps the audit trail where it is worth paying for.
## One language, two runtimes Standard and Express workflows share Amazon States Language almost entirely — the difference is the runtime underneath, and it is chosen at creation time via the state machine's type. You cannot flip an existing state machine from one to the other; you create a new one (which is why the type shows up in reviews of a design, not in a later tuning pass). ## Duration and throughput As of 2025, a Standard execution may run for up to **one year**; an Express execution is capped at **five minutes** and fails when it exceeds that. That single number decides many designs on its own: anything with a human approval, a multi-hour batch job, or a wait-for-callback is Standard by construction. In the other direction, Express is built for volume. Standard's `StartExecution` rate is an account quota you can hit with a busy event source; Express is designed for very high start rates and short lifetimes, which is why it is the mode behind API-request handling and per-event stream processing. ## Execution semantics — the part candidates miss - **Standard: exactly-once workflow execution.** The service will not run your workflow twice off one start, and it persists state between transitions. - **Express, asynchronous (`StartExecution`): at-least-once.** The workflow may be executed more than once. Every side effect must therefore be idempotent — writes keyed on an id, conditional puts, dedup keys. - **Express, synchronous (`StartSyncExecution`): at-most-once.** The caller blocks and receives the result inline, and there is no retry behind your back — so the caller owns retries, and *that* retry may duplicate work. This is the single most consequential difference. "Express is just cheaper Standard" is wrong; you are trading a durability guarantee for throughput. ## Observability Standard records every state transition into an execution history that you read with `DescribeExecution` and `GetExecutionHistory`, and the console replays it visually; history is retained for 90 days. Express records nothing in that API. If you do not enable logging to CloudWatch Logs on an Express state machine, a failed execution leaves you with essentially no evidence. On a Standard machine you can debug after the fact; on an Express machine you have to have decided to log beforehand. ## Integration patterns Express workflows do not support the job-run (`.sync`) or callback (`.waitForTaskToken`) service-integration patterns, and they do not support Activities. Both of those mean "pause and wait", and a runtime with a five-minute ceiling and no durable checkpointing has nowhere to park. If your workflow must wait for an ECS task, a Glue job, or a human, that state machine is Standard. ## Cost model The billing shapes are qualitatively different, and it is the shape, not the price, that matters in an interview: - Standard bills **per state transition**. A workflow with twenty states run a million times a day bills twenty million transitions a day. Adding a Pass state costs real money at that volume. - Express bills **per request plus duration × memory**, like Lambda. A twenty-state workflow that finishes in 300 ms costs roughly what a 300 ms workflow costs regardless of how many states it walked through. So the crossover is about chattiness and volume: few executions of a long, stateful workflow → Standard; many executions of a short workflow → Express, often by a large factor. ## The hybrid that shows judgment The pattern interviewers like to hear: a **Standard parent** that owns the durable, auditable outer process, invoking **Express children** for the high-volume inner work via `arn:aws:states:::states:startExecution.sync:2`. You keep the audit trail and the long-running waits where they matter, and you stop paying per transition for the hot inner loop. Distributed Map does the same thing structurally — its child executions are commonly Express. ```json "RunBatch": { "Type": "Task", "Resource": "arn:aws:states:::states:startExecution.sync:2", "Parameters": { "StateMachineArn": "arn:aws:states:us-east-1:123456789012:stateMachine:process-batch-express", "Input.$": "$.batch" }, "End": true } ``` ## How to answer the choice question Walk the decision, do not recite the table: how long does the process run, does it ever wait for something external, can steps be made idempotent, how many executions per second, and do I need to be able to reconstruct what happened to a single execution three weeks later? The first "yes" to long/waiting/auditable pins it to Standard; otherwise volume and cost push you to Express.
- Why can't an Express workflow use the .waitForTaskToken callback pattern?Both callbacks and `.sync` mean "suspend this execution durably until something external reports back", and Express has no durable checkpointing plus a five-minute ceiling — there is nowhere to park a paused execution. If you need a callback, the state machine holding it must be Standard, even if it delegates the fast work to Express children.
- You inherit an Express workflow that fails intermittently in production and there is nothing to look at. What went wrong?Express executions are not recorded in the execution-history API, so `DescribeExecution` and the console replay give you nothing. Logging must be enabled on the state machine to ship execution data to CloudWatch Logs, and it has to be enabled before the failure. Turn it on, reproduce, then query the log group.
- Can you convert a Standard state machine to Express to cut costs?Not in place — the type is fixed at creation, so you create a new state machine and cut traffic over. And it is not a pure cost lever: you would give up exactly-once execution, the history API, waits longer than five minutes, and the `.sync`/callback patterns. Check those first, cost second.
saying these in an interview costs you the question
- Saying Express is simply a cheaper Standard workflow
- Assuming asynchronous Express executions run exactly once
- Expecting the console to replay an Express execution
- Thinking the workflow type can be switched later
- Believing Express supports waits longer than a few minutes