skip to content

State Machines & Service Integrations

You will learn Amazon States Language and the workflow shapes it gives you — sequential tasks, branching, fan-out with Map and Parallel, and waits — plus how a state calls another AWS service synchronously, asynchronously, or with a callback token. Interviewers ask Standard vs Express because pricing and semantics diverge.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

In an AWS Step Functions Task state, what is the difference between the default request-response integration, the .sync pattern, and .waitForTaskToken?

level: middleimportance: must knowfreq 66%

answer

  1. the ARN suffix decides when the state finishes
  2. default returns as soon as the API answers
  3. one suffix waits for the job itself
  4. one suffix waits for an external actor
  5. the token must be put into the payload

basics

~20 s

Request-response calls the API and moves on as soon as it returns. The .sync suffix makes Step Functions wait until the underlying job reaches a terminal state and returns its result. The .waitForTaskToken suffix pauses the execution until something calls SendTaskSuccess or SendTaskFailure with the injected token.

solid answer

~40 s

All three are selected by the Task state's `Resource` ARN, and they differ in *when the state is considered finished*. **Request-response** is the default: `arn:aws:states:::ecs:runTask` calls the API, gets an HTTP response, and transitions immediately — for an async API that means you moved on before the work happened. **Job-run** appends `.sync`: `arn:aws:states:::ecs:runTask.sync` keeps the state open until the task actually reaches a terminal state, and returns the job's result, so you get a synchronous step out of an asynchronous service. **Callback** appends `.waitForTaskToken`: Step Functions injects `$$.Task.Token` into the payload you send, then suspends the state indefinitely until some worker calls `SendTaskSuccess` or `SendTaskFailure` with that token. Use the callback for human approval or a third-party system with no `.sync` support. Not every integration supports every suffix, and Express workflows support only request-response.

code

json · 24 lines
json
{
  "Comment": "Pause until an approver reports back with the task token",
  "StartAt": "RequestApproval",
  "States": {
    "RequestApproval": {
      "Type": "Task",
      "Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken",
      "Parameters": {
        "QueueUrl": "https://sqs.us-east-1.amazonaws.com/123456789012/approvals",
        "MessageBody": {
          "orderId.$": "$.orderId",
          "TaskToken.$": "$$.Task.Token"
        }
      },
      "Next": "Ship"
    },
    "Ship": {
      "Type": "Task",
      "Resource": "arn:aws:states:::lambda:invoke",
      "Parameters": { "FunctionName": "ship-order" },
      "End": true
    }
  }
}

go deeper

for a junior

Recall that the Resource ARN can carry a suffix and that without one the state finishes as soon as the API responds. Be able to point at .sync as the "wait for the job" form.

for a middle

Explain all three patterns and what each returns as the state's output, and show how $$.Task.Token gets from the context object into the message a worker receives.

for a senior

Diagnose the failure mode: a workflow advancing past an asynchronous job, or a callback that never resumes because the token was dropped. Talk about the extra IAM permissions .sync needs to track the job.

for a principal

Own the integration strategy — when SDK integrations remove Lambda glue entirely, when a callback is the right seam to a partner or a human, and how that choice pins the workflow to the Standard type.

## The suffix is the contract A Task state's `Resource` field says both *what* to call and *when the state is done*. Three patterns exist, and they are distinguished by a suffix on the ARN. ``` arn:aws:states:::ecs:runTask # request-response arn:aws:states:::ecs:runTask.sync # job-run: wait for the task to finish arn:aws:states:::ecs:runTask.waitForTaskToken # callback: wait for someone to report back ``` Getting this wrong is one of the classic Step Functions bugs, so it makes a good interview question: a workflow that "works" but whose downstream step reads results that do not exist yet is almost always a request-response state where the author assumed `.sync`. ## Request-response (the default) Step Functions makes the API call, waits only for the HTTP response, records that response as the state's output, and transitions. For a synchronous API such as `lambda:invoke` (a RequestResponse invocation) that is exactly what you want — the function ran, and its return value is your output. For an *asynchronous* API it is a trap. `ecs:runTask` returns as soon as the task is accepted; `glue:startJobRun` returns a job run id. The state succeeds in a second, and the next state runs against a job that has not started producing output. Nothing errors — the workflow is simply wrong. ## Job-run: `.sync` The `.sync` suffix asks Step Functions to keep the state open until the underlying job reaches a terminal state, and to fail the state if the job failed. Under the hood the service tracks the job (via events and polling) on your behalf, which is why the execution role needs more than just the start permission — it typically also needs describe/stop permissions and, for event-driven tracking, permission for the managed EventBridge rule the service creates. Commonly used ones: ``` arn:aws:states:::ecs:runTask.sync arn:aws:states:::batch:submitJob.sync arn:aws:states:::glue:startJobRun.sync arn:aws:states:::states:startExecution.sync:2 ``` That last one — a state machine invoking another state machine and waiting — is how you compose workflows, and the `:2` variant returns the child's output already parsed as JSON rather than as a string. Not every service offers `.sync`; the supported list is fixed by AWS, and if the one you need is absent, the callback pattern is the fallback. ## Callback: `.waitForTaskToken` This is the pattern that makes Step Functions useful for processes involving things AWS cannot see. When the state runs, Step Functions generates a task token, exposes it through the context object as `$$.Task.Token`, and — crucially — **you must pass it into the payload yourself**, because the token is only useful if the worker receives it: ```json "Parameters": { "QueueUrl": "https://sqs.us-east-1.amazonaws.com/123456789012/approvals", "MessageBody": { "orderId.$": "$.orderId", "TaskToken.$": "$$.Task.Token" } } ``` The state then suspends. It does not consume compute, and the workflow's persisted state is exactly where it was. It resumes only when something calls `SendTaskSuccess` (with a JSON output that becomes the state's result) or `SendTaskFailure` with that token. That something can be a Lambda function, an on-premises worker, a support engineer clicking a link in an email, or a partner's webhook handler — anything with credentials for `states:SendTaskSuccess`. The worker needs to be given the token and needs to keep it; losing it means the execution sits until the state's configured timeout, which is why you always configure one on a callback state rather than trusting the caller. ## Optimized versus SDK integrations Alongside the three patterns there is a second axis. **Optimized integrations** are the ones AWS wrote special support for — `lambda:invoke`, `sqs:sendMessage`, `dynamodb:putItem`, `ecs:runTask`, `sns:publish` and friends. They have tailored parameters and they are the ones that offer `.sync` and callbacks. **AWS SDK integrations** are the generic escape hatch, addressed as `arn:aws:states:::aws-sdk:<service>:<apiAction>`, for example `arn:aws:states:::aws-sdk:s3:listObjectsV2`. They expose a very large share of the AWS API surface directly from a Task state, with parameter names taken straight from the API model, and they let you drop a Lambda function whose only job was to call one API. They are request-response by nature, though many also accept `.waitForTaskToken`. ## Choosing Ask what "done" means for this step. If the API's return *is* the result, request-response. If the API starts something whose completion you care about, `.sync` when it exists. If completion is reported by something outside AWS's view — a person, a partner, a long poll-based worker — the callback. And remember Express workflows support only the request-response form, so a state machine that needs `.sync` or a callback is Standard.

  • What is the difference between an optimized service integration and an AWS SDK integration?
    Optimized integrations, like `arn:aws:states:::ecs:runTask`, are ones AWS built explicit support for — tailored parameters, and the `.sync` and callback patterns where they make sense. SDK integrations, addressed as `arn:aws:states:::aws-sdk:s3:listObjectsV2`, generically expose the AWS API surface with the model's own parameter names, and mostly only in request-response form. They exist so you stop writing Lambda functions that call one API.
  • How does the worker actually get the task token?
    You put it there. Step Functions exposes the token in the context object as `$$.Task.Token`, and you reference it inside the state's `Parameters` — dropped into an SQS message body, an SNS payload, or a Lambda event. If you forget, the worker has nothing to call `SendTaskSuccess` with and the execution waits until the state times out.
  • Your workflow starts a Glue job and the next state reads its output, but the output is always missing. What is the likely cause?
    The Task almost certainly uses `arn:aws:states:::glue:startJobRun` rather than the `.sync` variant. The plain form returns the moment the API accepts the job, so the workflow advances while the job is still running. Switching to `.sync` makes the state stay open until the job reaches a terminal state and returns its result.
  • Does a suspended callback state consume any compute or concurrency?
    No. The execution's state is persisted by the service and nothing of yours is running — no Lambda invocation is held open, no container idles. That is precisely why a callback beats polling in a loop: waiting is free apart from the state transitions on either side of it, and it can last far longer than any single compute invocation could.

saying these in an interview costs you the question

  • Assuming every Task state waits for the work to finish
  • Expecting .sync to be available on every service
  • Forgetting to pass $$.Task.Token into the payload
  • Thinking a paused callback state burns compute or concurrency
  • Believing Express workflows can use callbacks

context

open as a page

In AWS Step Functions, how do Standard and Express workflows differ, and how would you choose between them?

level: middleimportance: must knowfreq 78%

basics

~20 s

Standard workflows are durable and auditable: they run up to a year, execute exactly once, keep a queryable execution history, and bill per state transition. Express workflows run up to five minutes, are at-least-once, log to CloudWatch, and bill by requests plus duration.

open as a page

In AWS Step Functions, what are the required top-level fields of an Amazon States Language definition, and how do StartAt, Next and End control the flow?

level: juniorimportance: should knowfreq 52%

basics

~20 s

An Amazon States Language definition requires two top-level fields: StartAt and States. StartAt names the first state to run, each state's Next names the state that follows it, and End set to true finishes that branch.

open as a page

A Step Functions workflow must process several million JSON objects stored in an S3 bucket. Why would you use a Distributed Map state instead of an inline Map?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An inline Map runs every iteration inside the parent execution, so the items must fit in the state payload and every iteration's events land in one execution history, which is bounded. Distributed Map reads items straight from S3 and runs them as separate child executions at far higher concurrency.

open as a page

Standard Step Functions workflows bill per state transition. When would you keep multi-step coordination inside application code or a plain event-driven chain on AWS instead of building a state machine?

level: principalimportance: should knowfreq 40%

basics

~20 s

Skip the state machine when the coordination is cheap, fast and stateless: very high-volume per-event work where per-transition billing dominates, latency-critical paths, and pure computation that belongs inside one function. Reach for it when you need durable state, waits, or an auditable trail per execution.

open as a page

AWS Step Functions caps the data passed into and out of a state. What is that limit, and how do you design a workflow whose steps produce large results?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

State input and output are capped at 256 KiB as of 2025, and exceeding it fails the execution with a data-limit error. Keep large data in S3 and pass object keys through the workflow, trimming each state's output so payloads never accumulate.

open as a page