A team implements a multi-step order-fulfillment process (validate payment, reserve inventory, ship, notify) as a single long-running Lambda function that keeps intermediate results in local variables across the steps, retrying the whole function from the top on any failure. What's wrong with this design, and how would a workflow orchestrator like AWS Step Functions change it?
answer
- execution timeout caps single-function workflows
- retry-from-top duplicates non-idempotent steps
- orchestrator, not function, holds workflow position
- each state = its own stateless invocation + persisted output
- per-step retry/catch policies + execution history
basics
~20 sKeeping progress in a running function's local variables means a crash or timeout loses everything and forces starting over, possibly repeating things like a payment charge; a workflow orchestrator instead saves each step's result externally so the process can resume exactly where it stopped.
solid answer
~50 sA single long-running function that holds intermediate step results in local variables ties correctness to that one process instance staying alive for the entire multi-step flow — any timeout, crash, or cold-start recycling loses all progress, and 'retry from the top' means non-idempotent steps like charging a payment or reserving inventory can run twice. Step Functions inverts this: it's an external state machine that is itself the source of truth for which step you're on, and each step is invoked as its own separate, stateless function call, with the orchestrator persisting the step's output before invoking the next step. On failure, it can retry just the failed step, with per-step retry/backoff/catch policies, and it enforces idempotency at the step boundary rather than the whole-flow boundary. This decouples individual function execution limits from overall workflow duration, and gives a visual/audit trail of exactly which step failed, very hard to reconstruct from logs of a monolithic function.
go deeper
Should recognize that a crash mid-way through the single function loses all progress made so far.
Should identify that retrying from the top can re-run steps that already succeeded, and know Step Functions exists to coordinate multi-step flows.
Should explain how the orchestrator externally persists step position/output, enabling per-step retry, and connect this to idempotency requirements at each step boundary.
Should weigh Standard vs Express workflow trade-offs, payload-size externalization patterns, and design org-wide conventions for when a process belongs in an orchestrated workflow versus a single function, including cost and observability implications at scale.
## What breaks in the single-function design The single-function design described sits directly on Lambda's execution-time and statelessness constraints and violates both. 1. Lambda functions have a **maximum execution duration**, and every invocation runs in one of the ephemeral, isolated execution environments — if that environment crashes, is throttled, hits its timeout, or is recycled mid-execution for any platform-internal reason, every local variable holding intermediate progress is lost instantly, because that state existed only in the process memory of one instance that no longer exists. 2. The 'retry the whole function from the top' fallback compounds this: if the function successfully validated payment and reserved inventory, then failed during the shipping step, a naive full retry re-runs payment validation and inventory reservation again. 3. If those operations aren't carefully designed to be **idempotent**, the customer gets double-charged, or two units of inventory get reserved for one order. ## How an orchestrator changes it AWS Step Functions (and equivalents like Temporal or Azure Durable Functions) solve this by moving the 'what step are we on and what did each prior step produce' bookkeeping out of any single function's memory and into an externally-persisted **state machine** definition executed by the orchestration service itself. Concretely: - you define the workflow as a state machine — a sequence/graph of states, each typically an invocation of one small, stateless Lambda function that performs exactly one step; - the orchestrator, not any one function, holds the current position in that graph and the accumulated output of each completed state; this is stored durably by the AWS-managed service, entirely independent from any Lambda execution environment; - when a state's Lambda invocation completes, the orchestrator persists its output and only then triggers the next state as an entirely fresh, independent Lambda invocation; - each state can carry its own configured retry policy and catch/fallback behavior, defined declaratively rather than hand-coded try/catch logic strung through one giant function. ## The core trade-off The core trade-off is added architectural complexity in exchange for durability and observability. Instead of one function you can step through in a debugger, you now have N small functions plus a state machine definition to author and reason about, and cross-step data has to flow through the orchestrator's input/output plumbing rather than shared local variables, which can hit payload-size limits on what can be passed between states, encouraging the same 'pass a reference, not the blob' pattern used with S3 externalization elsewhere. There's also per-state-transition cost with Step Functions' **Standard** workflow pricing, though **Express** workflows exist specifically for high-volume, short-duration flows. ## The reliability payoff The reliability payoff is substantial. - Because the orchestrator, not any function instance, is the durable source of truth for workflow position, a crashed or timed-out Lambda invocation during the 'ship' step causes Step Functions to retry only the 'ship' state, not re-run payment validation or inventory reservation — provided each state is itself designed to be idempotent, since Step Functions documents **at-least-once** delivery semantics for state transitions. - Failure modes are also far easier to diagnose: Step Functions provides an execution history and visual graph showing exactly which state ran, its input, its output or error, and how long it took, versus reconstructing that from application logs scattered across however many invocations of a monolithic function ran during a failed multi-step attempt. - Long-running or human-in-the-loop steps, such as waiting days for a manager's approval, become trivial to express as a wait/task-token pattern, something essentially impossible to do safely inside one Lambda invocation given its hard execution-time limit. ## Where this shows up This exact evolution — from a monolithic Lambda doing everything to Step Functions coordinating small, focused functions — is a well-documented real-world pattern; AWS describes order-processing and media-processing pipelines as canonical Step Functions use cases, precisely because those workflows have multiple external side effects that must not silently duplicate on retry, and multi-step durations that can exceed what's safe to keep alive in a single function's local state.
- Does making each step idempotent alone solve the problem, without needing an orchestrator?Idempotency is necessary but not sufficient — it prevents duplicate side effects on retry, but without an orchestrator you still need something to durably track which steps have completed so you know what to retry versus skip, plus something to invoke the remaining steps after a crash; an orchestrator provides that tracking and re-invocation mechanism rather than requiring a hand-rolled checkpoint table.
- What happens to data that needs to pass between steps if it's too large for Step Functions' state input/output payload limits?The same externalization pattern used elsewhere applies: write the large payload to S3 from the producing step, and pass only the S3 object reference through the state machine's input/output to the consuming step, keeping the orchestrator's payload small and fast.
- Why can Step Functions safely retry a single failed state instead of the whole workflow, when the monolithic-function design couldn't?Because the orchestrator persisted the output of every prior state before advancing, it can resume execution at exactly the failed state with the correct upstream inputs already available, whereas the monolithic function had no durable record of prior progress once its process instance was gone, forcing a full restart.
It's the difference between one person trying to carry an entire relay race baton solo across a marathon course (if they collapse, the whole race restarts from zero) versus a real relay team where a race official at each handoff point records exactly which runner finished which leg, so a stumble only means re-running that one leg, not the whole race.
saying these in an interview costs you the question
- Treats a single Lambda invocation as suitable for an arbitrarily long multi-step business process
- Doesn't recognize that 'retry the whole function' duplicates non-idempotent side effects like payments
- Believes Step Functions removes the need for idempotent step design
- Passes large payloads directly through state machine input/output instead of externalizing to S3
- Assumes local variables in one Lambda invocation can survive a timeout or crash