A Prefect task calling a flaky API fails intermittently. How do you configure retries and timeouts?
answer
- Two knobs, and one of them is forgotten
- A hang is not a failure
- Spread the herd before it stampedes
- Some errors will fail identically every time
- A second attempt repeats the side effects
basics
~20 sSet retries and retry_delay_seconds on the @task. retry_delay_seconds accepts a list of per-attempt delays or exponential_backoff(), with retry_jitter_factor to spread a thundering herd, and timeout_seconds bounds a hung call so it fails instead of blocking the flow forever.
solid answer
~40 sConfigure it on the task, because the task run is the retry boundary: `@task(retries=3, retry_delay_seconds=exponential_backoff(backoff_factor=10), retry_jitter_factor=1, timeout_seconds=60)`. `retry_delay_seconds` also takes a single number or an explicit list like `[1, 10, 60]` for per-attempt delays. Jitter matters when a whole fan-out hits the same API and would otherwise retry in lockstep. `timeout_seconds` is the part people forget: without it a hung socket occupies a worker slot indefinitely and no retry ever fires, because the run never fails. Use `retry_condition_fn` when only some failures are worth retrying — a 503 yes, a 400 no. And retries only help transient faults: if the call has side effects, the retried attempt repeats them, so the work must be idempotent or the retry turns one flaky request into duplicate writes.
code
python · 26 linesfrom prefect import flow, task
from prefect.tasks import exponential_backoff
class PermanentError(Exception):
pass
@task(
retries=3,
retry_delay_seconds=exponential_backoff(backoff_factor=10),
retry_jitter_factor=1,
timeout_seconds=60,
)
def fetch_page(page: int, idempotency_key: str) -> dict:
# client-side timeout as well: fail near the cause
resp = call_api(page, key=idempotency_key, timeout=20)
if 400 <= resp.status < 500 and resp.status != 429:
raise PermanentError(resp.status) # not worth retrying
resp.raise_for_status()
return resp.json()
@flow(timeout_seconds=1800)
def ingest(pages: list[int]) -> list[dict]:
return fetch_page.map(pages, "run-2026-08-21").result()go deeper
Know the two basic options on a task — retries and retry_delay_seconds — and that retries exist for transient failures such as a network blip, not for bugs.
Explain the delay shapes: a number, a per-attempt list, or exponential backoff with jitter, plus what timeout_seconds does and why it lives beside the retry settings.
Diagnose the real incidents: a hang that no retry can fix, a lockstep fan-out retry storm, deterministic errors burning the backoff budget, and duplicate writes from a non-idempotent retry.
Own the policy — which classes of failure retry at all, how retry budgets interact with a shared dependency's capacity, and how retry rates get monitored so quietly degrading pipelines surface before they break.
## Where retries are configured The task run is the retry boundary, so the knobs live on the decorator: ```python from prefect import task from prefect.tasks import exponential_backoff @task( retries=3, retry_delay_seconds=exponential_backoff(backoff_factor=10), retry_jitter_factor=1, timeout_seconds=60, ) def fetch_page(page: int) -> dict: ... ``` `retries=3` means up to three additional attempts after the first, four executions worst case. Between attempts the run sits in a retrying state and the same task run is reused, so its identity and logs are continuous rather than a fresh run per attempt. ## Shaping the delay `retry_delay_seconds` accepts three shapes and choosing between them is a real answer: - **A single number** — `retry_delay_seconds=5`. Fine for a service that recovers in seconds. - **A list** — `retry_delay_seconds=[1, 10, 60]`. One entry per attempt, which is the most explicit form: fast first retry for a blip, long final retry for a deploy-shaped outage. The list length should match `retries`. - **`exponential_backoff(backoff_factor=n)`** — generates a growing sequence, the standard choice against a rate-limited or overloaded dependency. Add `retry_jitter_factor` on top. It randomises each delay, and it matters precisely in the fan-out case this tree cares about: 200 mapped task runs all hit the same API, all get throttled at the same instant, and all retry at the same instant unless the delays are spread. Lockstep retries turn a small overload into a self-sustaining one. ## Not every failure deserves a retry Retrying a deterministic failure is pure waste: a 400 from a malformed payload, a schema mismatch, a missing permission or an assertion error will fail identically four times and simply delay the alert by the sum of the backoff. Distinguishing transient from deterministic is the senior half of this question. `retry_condition_fn` lets you encode it — it receives the task, the task run and the state, and returns whether to retry — so you can retry a 429 or 503 and fail fast on a 4xx that will never change. If you cannot use it, raise a distinct exception type for permanent failures and keep the retriable path narrow. ## Timeouts `timeout_seconds` is available on tasks and on flows, and it is the knob most often missing in a real incident. The failure mode is specific: a request with no client-side timeout hangs on a half-open socket. The task run never fails, so `retries` never fires — retries are triggered by failure, and a hang is not a failure. The run occupies a worker slot forever, the flow run never completes, and if the flow is scheduled you accumulate overlapping runs. A timeout converts a hang into a failure, which then interacts correctly with everything else: it becomes visible, it can be retried, and the slot is released. Set it deliberately — comfortably above the observed p99 of the operation, not at the mean, or you will manufacture failures on ordinary slow days. Belt and braces: also set the client library's own timeout, because that fails faster and closer to the cause. Flow-level `timeout_seconds` is a different guarantee: it bounds the whole flow run, catching the case where no single task hangs but the run as a whole drags past its useful window. ## Task retries versus flow retries Both exist, and confusing them is a red flag. `@task(retries=...)` re-runs one unit of work. `@flow(retries=...)` re-runs the flow function — in general re-executing its steps, which is far heavier and only appropriate when the whole stage is safely repeatable. The default instinct should be task-level retries for transient dependency faults, and flow-level retries reserved for a stage that is genuinely idempotent end to end. ## The idempotency condition A retry is a second execution of side effects. If `fetch_page` only reads, retrying is free. If the task charges a card, sends an email, or appends rows to a table, the retried attempt repeats it, and a flaky network on the *response* path — where the write succeeded but the acknowledgement was lost — makes duplicates especially likely. Make the write safe to repeat: an idempotency key on the API call, a merge or upsert keyed on a natural identifier, writing to a deterministic path that overwrites, or a check-then-act guarded by a unique constraint. "Add retries" without asking this question is how a retry policy quietly becomes a data-quality incident. ## Observability Retries hide instability. A pipeline whose tasks quietly retry twice every night is failing constantly and nobody notices until the third attempt starts failing too. Watch attempt counts, not only final states, and treat a rising retry rate as a signal about the dependency. ## How to answer Name the knobs precisely, explain backoff plus jitter and why jitter matters for a fan-out, insist on `timeout_seconds` because retries cannot rescue a hang, separate transient from deterministic failures, and finish on idempotency. That ordering is the difference between reciting parameters and demonstrating you have run this in production.
- Why can a task with retries=3 still block a flow run forever?Because retries fire on failure, and a hang is not a failure. Without timeout_seconds, a request stuck on a half-open socket keeps the task run in a running state indefinitely, holding a worker slot and preventing the flow run from completing. The timeout converts the hang into a failure, which then becomes retriable.
- When is a flow-level retry the right choice rather than a task-level one?When the whole stage is idempotent and a partial rerun is meaningless — for instance a flow that rebuilds a temporary dataset from scratch. Flow retries re-execute the steps, so they are far heavier and unsafe if any step writes non-repeatably. For transient dependency faults, retry the individual task instead.
- Why does retry_jitter_factor matter for a mapped fan-out?Two hundred mapped task runs hitting a throttled API get rejected at almost the same moment and, with identical delays, retry in lockstep — recreating the overload on schedule. Jitter randomises each delay so the attempts spread out, which is what actually lets the dependency recover.
- How do you avoid retrying failures that will never succeed?Use retry_condition_fn to inspect the failed state and decide — retry a 429 or 503, fail fast on a 400 or a permissions error. Alternatively narrow the retriable surface by raising a distinct exception type for permanent faults, so the backoff is not spent delaying an alert you needed immediately.
saying these in an interview costs you the question
- Adding retries without asking whether the task is idempotent
- Believing retries can rescue a hung task with no timeout
- Retrying deterministic errors like a 400 or schema mismatch
- Setting the timeout at the average duration rather than above p99
- Confusing flow-level retries with task-level retries