skip to content

Standard Step Functions workflows bill per state transition. When would you keep multi-step coordination inside application code or a plain event-driven chain on AWS instead of building a state machine?

level: principalimportance: should knowfreq 40%

answer

  1. ask what the service is actually selling you
  2. volume multiplies by the number of steps
  3. coordination language, not a programming language
  4. a state should be worth resuming from
  5. the alternative has costs you now own

basics

~20 s

Skip the state machine when the coordination is cheap, fast and stateless: very high-volume per-event work where per-transition billing dominates, latency-critical paths, and pure computation that belongs inside one function. Reach for it when you need durable state, waits, or an auditable trail per execution.

solid answer

~50 s

The honest test is what the workflow buys you that you would otherwise have to build. Step Functions sells durable state between steps, built-in retry and error routing, waits measured in days, callbacks, and an execution history you can replay per request — all real engineering you would otherwise own. If a process needs none of those, the state machine is overhead: Standard bills **per state transition**, so a twenty-state workflow at high volume multiplies cost by twenty, and each transition adds latency you cannot tune away. For chatty, short, high-volume work the answer is usually Express rather than "no workflow", but for a tight computation loop or two calls that always succeed together, an in-process sequence, an SQS hop, or an EventBridge rule is simpler and cheaper. The other side matters too: hand-rolled orchestration means you now own persistence, retries, idempotency and visibility — usually badly.

go deeper

for a junior

Know that Standard workflows are charged per step transition, so a workflow with many small steps run very often costs more than doing the same work inside one function.

for a middle

Compare the two billing shapes concretely and explain what you would have to build yourself — retries, idempotency, a record of in-flight work — if you drop the state machine.

for a senior

Show the threshold reasoning on a real workload: executions per second times states, the latency budget of a synchronous path, and whether anything in the process waits on something outside AWS.

for a principal

Own the standard for the organisation: which processes must be state machines because their audit trail is an operational requirement, where Express is the default, and what teams are forbidden from hand-rolling.

## Frame it as "what am I buying" A state machine is not free structure; it is a purchase. What you get is specific: - **Durable state between steps.** The service remembers where an execution is and what it was carrying, for up to a year. - **Retry and error routing without writing it.** Backoff and failure paths are configuration, not code you maintain. - **Long waits and callbacks.** A process can pause for two days waiting for a person, consuming nothing. - **Per-execution auditability.** You can open one execution from three weeks ago and see exactly which step failed with what input. - **Control flow visible outside the code.** A reviewer diffs the definition and sees the process change. Every one of those is expensive to build yourself and easy to build badly. The decision is whether this particular process needs them. ## When the answer is no **Very high volume with trivial steps.** Standard bills per state transition. A ten-state workflow run 50 million times a month bills 500 million transitions. If the work is "parse the event, write to DynamoDB, publish to SNS", a Lambda triggered off the source does it in one billing unit. The first move here is usually Express rather than abandoning workflows — Express bills by request plus duration and memory, which reprices chattiness to nearly nothing — but if there is no branching, no waiting and no retry policy beyond the event source's own, a plain consumer is simpler. **Latency-sensitive synchronous paths.** Every transition costs some milliseconds. On a user-facing request path with a tight budget, a workflow of eight states spends real time on coordination. Synchronous Express narrows the gap, but a single function doing three calls will still beat it. **Tight loops and pure computation.** ASL is a coordination language, not a programming language. Iterating over an in-memory list, doing arithmetic, or branching a dozen ways on a field is code — and expressing it as Choice-state sprawl produces a definition nobody can review and that costs a transition per iteration. The heuristic: a state should be a *unit of work worth resuming from*. If a step could not meaningfully be retried on its own, it does not deserve to be a state. **Steps that always succeed or fail together.** Two calls with no independent failure handling and no need to resume between them are one function, not two states. ## When the answer is yes even though it looks expensive **Anything that waits on a human or a partner.** There is no cheap way to hand-roll "pause for three days, then resume exactly where you were" — you end up building a scheduler, a store and a resume path. **Long-running, multi-system business processes.** Order fulfilment, tenant provisioning, media pipelines, data-migration runs. The per-transition cost is negligible against the volume, and the audit trail is the operational product. **Processes where "which step failed for customer X" is asked routinely.** Support and compliance load is a real cost. Reconstructing that from logs across five services is far more expensive than the transitions. **Fan-out with partial-failure semantics.** Distributed Map with a tolerated-failure threshold and results written to S3 is a batch engine you would not want to rebuild. ## The cost of the alternative The weak version of this answer stops at "workflows are expensive". The strong version prices the alternative. Orchestration in application code means you own: a durable record of where each in-flight process is (a table, and its schema migrations); retry with backoff and a dead-letter path; idempotency for every step, because your retries will duplicate; timeouts and stuck-process detection; and enough structured logging that someone can answer the "which step failed" question. Teams routinely under-price that and ship a chain of Lambdas invoking Lambdas with no durable state, then discover the gap during their first partial outage. ## Testability and change velocity — a real, secondary factor ASL is not unit-testable the way a function is, though Step Functions provides a `TestState` API to exercise a single state's input and output processing, and a local test runner exists for definitions. Even so, a team whose entire toolchain is built around code tests will move slower in a definition, and definitions in source control are reviewed less carefully than code. That is an argument for keeping *business logic* inside tasks and using the state machine only for *coordination* — not an argument against the service. ## The shape of a strong answer Name the axes (volume × states, latency budget, does anything wait, is per-execution audit a requirement, who owns durability if not the service), give a threshold you would actually compute, and reject the false binary: the common right answer is Express workflows for the hot path and a Standard workflow around the durable outer process, not "workflow" versus "no workflow".

  • How do you decide what deserves to be its own state versus code inside a task?
    A state should be a unit of work worth resuming from. If a step has its own failure mode, its own retry policy, or a meaningful checkpoint after it, make it a state. If it is arithmetic, parsing, or a call that always succeeds or fails alongside its neighbour, it is code inside a task. That keeps definitions reviewable and transitions justified.
  • A team proposes replacing a state machine with a chain of Lambdas invoking each other to cut cost. What do you push back on?
    They have to rebuild what they are giving up: a durable record of where each in-flight process sits, retry with backoff and a dead-letter path, idempotency for the duplicates those retries create, stuck-process detection, and enough correlation to answer which step failed for a given customer. That is usually more expensive than the transitions, and it fails first during a partial outage.
  • Is 'we cannot unit-test ASL' a good reason to avoid Step Functions?
    It is a real friction but a weak reason. The `TestState` API exercises a single state's input and output processing, and a local runner can execute definitions. The better response is architectural: keep business logic inside tasks where your normal tests apply, and let the state machine own only coordination — which is what should be reviewed in the definition anyway.

saying these in an interview costs you the question

  • Treating a state machine as free structure with no cost axis
  • Costing the workflow but never costing the hand-rolled alternative
  • Expressing tight loops and arithmetic as Choice-state sprawl
  • Framing it as workflow versus no workflow, ignoring Express
  • Assuming chained Lambda invocations give durable in-flight state

context