skip to content

An AWS Step Functions Task state calls a worker that sometimes dies without reporting back, leaving executions stuck for hours. How do TimeoutSeconds and HeartbeatSeconds address this, and how do they differ?

level: seniorimportance: should knowfreq 50%

answer

  1. two clocks, not one setting
  2. total budget versus longest silence
  3. the worker has to call in
  4. the smaller one must fit inside
  5. expiry stops waiting, not working

basics

~20 s

TimeoutSeconds caps a Step Functions task's total wall-clock duration and raises States.Timeout when exceeded. HeartbeatSeconds caps the gap between SendTaskHeartbeat calls from the worker and raises States.HeartbeatTimeout, detecting a dead worker long before a generous overall timeout would.

solid answer

~40 s

They are two different clocks. `TimeoutSeconds` is the total budget for the task: exceed it and Step Functions stops waiting and raises `States.Timeout`. `HeartbeatSeconds` is the maximum silence allowed between liveness signals — the worker calls `SendTaskHeartbeat` with its task token, and if none arrives in time the task fails with `States.HeartbeatTimeout`. Heartbeats only apply where a worker holds a token, so activities and `.waitForTaskToken` integrations; `HeartbeatSeconds` must be smaller than `TimeoutSeconds`. The point of having both is that a legitimately long job needs a generous overall timeout, and a generous timeout means a crashed worker is undetected for hours. A short heartbeat catches the crash in seconds while the long timeout still bounds the honest case. Both can be driven from execution data with `TimeoutSecondsPath` and `HeartbeatSecondsPath`.

go deeper

for a junior

Know that TimeoutSeconds bounds how long a task may take and that exceeding it raises States.Timeout, so an unbounded task can otherwise hang for a very long time.

for a middle

Explain the second clock: HeartbeatSeconds bounds the silence between SendTaskHeartbeat calls, applies only where a worker holds a task token, must be smaller than the timeout, and raises States.HeartbeatTimeout.

for a senior

Show why both exist — a long job needs a long budget, and a long budget hides a dead worker — and that expiry does not cancel the work, so downstream effects need an idempotency key such as the execution name.

for a principal

Own the policy: which classes of task must declare timeouts at all, whether a state-machine-level timeout is mandatory, and how detection latency for stuck work is held to a standard rather than left to each workflow author.

## The failure being solved A Step Functions `Task` state waits for something to report back. If that something never does — the container was terminated, the process segfaulted, the machine lost its network — the state waits. With no timeout configured, it waits until the execution itself expires, and a Standard workflow's maximum execution duration is measured in months. The observable symptom is an execution sitting on one state indefinitely, with nothing failing, nothing alarming, and nothing retrying, because from Step Functions' point of view the task simply has not finished yet. Both mechanisms turn that silence into an error name, which is what makes the rest of the error-handling machinery — retriers, catchers, a compensating branch — able to act. ## TimeoutSeconds: the total budget `TimeoutSeconds` is a field on the `Task` state giving the maximum wall-clock time from when the task starts to when it reports success. Exceed it and the state fails with `States.Timeout`, which you can name in `ErrorEquals` like any other error. ```json "ConvertVideo": { "Type": "Task", "Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken", "TimeoutSeconds": 3600, "HeartbeatSeconds": 60, "Retry": [{ "ErrorEquals": ["States.HeartbeatTimeout"], "MaxAttempts": 2 }], "Next": "Publish" } ``` There is also a top-level `TimeoutSeconds` on the state machine definition itself, bounding a whole execution. Setting it is cheap insurance against a workflow that gets stuck in a loop or on a state you forgot to bound. ## HeartbeatSeconds: the liveness clock `HeartbeatSeconds` sets the maximum interval Step Functions will tolerate between liveness signals. The worker is responsible for sending them: it calls the `SendTaskHeartbeat` API with the task token it was handed, on whatever cadence it chooses, as long as the gap stays under `HeartbeatSeconds`. Miss the window and the task fails with `States.HeartbeatTimeout`. This only makes sense where a worker holds a task token and can call back — activities polled with `GetActivityTask`, and service integrations using the `.waitForTaskToken` pattern. There is nothing to heartbeat in a plain request/response task that returns in one call. `HeartbeatSeconds` must be less than `TimeoutSeconds`; a heartbeat window as long as the whole budget detects nothing. ## Why you want both The two clocks answer different questions, and the gap between them is the entire value: - `TimeoutSeconds` asks *"is this job taking unreasonably long?"* - `HeartbeatSeconds` asks *"is anybody still working on it?"* A video transcode may legitimately need an hour, so the budget has to be an hour. If the transcoding worker crashes ninety seconds in, a one-hour timeout means fifty-eight minutes of an execution that is already dead. A sixty-second heartbeat surfaces it almost immediately, so the retrier can hand the work to a healthy worker while the customer is still on the page. Notice in the snippet above that the retrier names `States.HeartbeatTimeout` specifically — a crashed worker is worth retrying, whereas a genuine `States.Timeout` may mean the input is simply too large and retrying will only time out again. ## The fact people miss: the work does not stop When the clock expires, Step Functions stops waiting. It does **not** reach out and kill the worker. An external activity worker or an already-running function invocation carries on, unaware, and may complete its side effects after the state has already failed and the workflow has moved to a compensating branch. If a retrier then hands the same input to a second worker, both are doing the job. This has two consequences worth volunteering in an interview. First, downstream operations under a timed-out task need to be idempotent, keyed on something stable — the execution name is available on the context object as `$$.Execution.Name` and makes a natural idempotency key that survives across retries of the same execution. Second, a worker that finishes late and calls `SendTaskSuccess` for a task that already timed out will get an error rather than resurrecting the execution, so it must handle that path rather than assuming success. ## Choosing the numbers Set `TimeoutSeconds` from the honest p99 of the work plus headroom, not from the average, or you will manufacture failures on the tail. Set `HeartbeatSeconds` from how quickly you need to detect a dead worker, and make sure the worker's heartbeat cadence has room for a couple of missed sends inside that window. Where the right budget depends on the payload — a small file versus a large one — `TimeoutSecondsPath` and `HeartbeatSecondsPath` read the values from the state's input instead of hard-coding them.

  • Does Step Functions stop the underlying work when TimeoutSeconds expires?
    No. It stops waiting and fails the state, but it has no way to terminate an external activity worker or an already-running function invocation. The work may complete afterwards, so its side effects must be idempotent — and a worker that later calls `SendTaskSuccess` for a task that already timed out will get an error and must handle that path instead of assuming it succeeded.
  • Which service integration patterns can actually use HeartbeatSeconds?
    Only the ones where a worker holds a task token and can call the API back: activities polled with `GetActivityTask`, and integrations using the `.waitForTaskToken` pattern. A plain request/response task has no token and nothing to report during the call, so a heartbeat setting would have nothing to check.
  • How would you make the timeout depend on the size of the work item?
    Use `TimeoutSecondsPath` instead of `TimeoutSeconds` — it reads the value from the state's input at a JSONPath you name, so an upstream state can compute a budget from the payload. `HeartbeatSecondsPath` does the same for heartbeats. That avoids the usual compromise of one hard-coded timeout that is too tight for large items and useless for small ones.

saying these in an interview costs you the question

  • Believing a task cannot hang because Lambda has its own timeout
  • Assuming Step Functions kills the worker when the task times out
  • Setting HeartbeatSeconds larger than or equal to TimeoutSeconds
  • Expecting heartbeats to work without the worker calling SendTaskHeartbeat
  • Sizing the timeout from the average duration rather than the tail

context