skip to content

In an orchestrated data pipeline, what happens when a task with three configured retries fails?

level: juniorimportance: should knowfreq 60%

answer

  1. it starts over, it does not carry on
  2. a failing task is not yet a failed task
  3. count the attempts against the deadline
  4. whatever the first attempt already did, happens again

basics

~20 s

The orchestrator marks that attempt failed, waits the configured delay, then re-runs the whole task body from the start. Only when the retry budget is exhausted does the task count as failed and downstream work stop.

solid answer

~50 s

Each attempt is a separate execution of the same task with the same run parameters. When attempt one fails, the orchestrator puts the task into a retrying state rather than a failed state, waits the configured delay, and starts a fresh attempt — commonly on a different worker. Crucially it **restarts the task, it does not resume it**: everything the first attempt did before it died happens again, so a task that sent an email or appended rows will repeat that side effect unless it was written to be safe to run twice. Downstream tasks stay in a waiting state throughout; they are not failed and not skipped. Only after the last attempt in the budget fails does the task become failed, the failure propagate downstream, and the failure alert fire. Worst-case wall clock is roughly attempts × (runtime + delay), which is what eats an SLA.

code

text · 6 lines
text
task=load_orders  attempt=1/4  state=failed     err=ConnectionResetError
  -- waiting 300s --
task=load_orders  attempt=2/4  state=failed     err=ConnectionResetError
  -- waiting 300s --
task=load_orders  attempt=3/4  state=success    rows=812_004
downstream task=build_daily_mart  state=queued -> running

go deeper

for a junior

Be ready to say plainly that a retry re-runs the entire task from the beginning with the same parameters, and that the task only counts as failed once its attempts run out.

for a middle

Explain the state transitions — attempt failed, retrying, terminal failure — and do the wall-clock arithmetic showing how attempts and delays consume the delivery window.

for a senior

Show the operational consequences: duplicated side effects, retries masking a real defect, and a failure-only alerting policy staying silent through a long retry chain.

for a principal

Argue for retry budgets as a platform default that follows from each dataset's deadline, and for tracking retry rate as a leading indicator rather than treating retries as free.

## What a retry is In any orchestrator, a task's execution is modelled as a series of **attempts**. The scheduler records, per attempt, a start time, an end time, and an outcome. A retry policy is two numbers plus a rule: how many extra attempts are allowed after the first one fails, how long to wait between them, and (sometimes) whether that wait grows. When an attempt ends in failure, the orchestrator does not immediately declare the task failed — it moves it to a *retrying* state, which is a distinct, non-terminal state that most schedulers surface separately in the UI precisely so that a transient blip does not look like an outage. ## The loop, step by step 1. Attempt 1 runs. The process exits non-zero, raises an uncaught exception, or the orchestrator loses its heartbeat. The attempt is recorded as failed. 2. If attempts remain in the budget, the task enters the retrying state. Downstream tasks that depend on it remain *waiting* — they are neither failed nor skipped, because the task has not reached a terminal state. 3. The orchestrator waits the retry delay, then queues attempt 2. It typically lands on whichever worker is free, which may not be the one that ran attempt 1. 4. This repeats until an attempt succeeds (task succeeds, downstream unblocks) or the budget is exhausted (task fails, and downstream reacts according to the dependency rules the pipeline declares). ## Retries restart; they do not resume This is the single most misunderstood part. There is no checkpoint. The retry re-executes the task body from its first line, with the *same* logical parameters — the same target window, the same partition, the same input paths. Reusing the parameters is deliberate: the retry is another go at the *same* unit of work, not a new unit of work. But it means every effect the failed attempt produced before dying happens again: - rows appended to a table are appended a second time, - a notification already sent is sent again, - a metered API call is billed again, - a temp file written to the previous worker's local disk is simply gone, because attempt 2 may be on a different machine. That is why "tasks must be safe to run twice" is the precondition for enabling retries at all. The usual shape is: write to a scratch location and swap or overwrite the target atomically at the end, so a half-finished attempt leaves nothing behind for the next one to duplicate. ## Delay, backoff and where the wall clock goes A fixed delay (say five minutes) is the common default; a growing delay reduces pressure on a dependency that is already struggling. Either way, do the arithmetic before you set the numbers: a task that normally runs 20 minutes, with 3 retries and a 10-minute delay, can occupy up to roughly 4 × 20 + 3 × 10 = 110 minutes before anyone is told anything is wrong. If the dataset it produces is promised for 06:00, the retry budget has quietly consumed most of the margin. Retry budgets are an SLA decision, not a cosmetic default. ## Retries and alerting Most teams alert on the terminal failure, not on individual attempt failures, otherwise every transient blip pages someone. The cost of that choice is that a task silently burning through retries looks healthy on a failure-only dashboard. Two habits fix it: track the retry *rate* per task as a metric (a task that suddenly needs two attempts every night is telling you something before it starts failing outright), and keep a separate deadline monitor that fires when the output is late regardless of what the run's state says. ## What retries are not for A retry is not a wait. If a task fails because the upstream file has not landed yet, a retry chain is a crude poll: it works by accident, it hides the real condition (lateness) behind a failure count, and it produces a misleading error in the log. Express waiting as waiting, with its own deadline. A retry is also not a timeout. Retries only trigger on a *failure*; a task that hangs forever on an unanswered socket never fails, so it never retries — it just holds its worker slot. Retries and execution timeouts are complementary settings and you generally want both. Finally, a retry is not a fix for a deterministic error. If the cause of the failure is the same on every attempt — a missing column, a bad credential, a bug — the retries are guaranteed to fail identically, and their only effect is to delay the alert and spend compute.

  • While a task is between retry attempts, what state are its downstream tasks in?
    They are still waiting. A retrying task has not reached a terminal state, so dependencies are neither satisfied nor broken, and downstream tasks simply sit queued. That is why a long retry chain shows up as a stalled pipeline rather than a failed one, and why failure-only alerting stays quiet the whole time.
  • How do you size a retry budget so it does not eat the delivery deadline?
    Work backwards from the promised time. Multiply the task's typical duration plus the retry delay by the number of attempts, add that to the rest of the critical path, and check it still lands before the deadline with margin. If it does not, shorten the delay, cut attempts, or move the schedule earlier.

A retry is like restarting a recipe from the top after burning the sauce, not picking up from step six — anything you already put in the oven is still in the oven.

saying these in an interview costs you the question

  • Thinking a retry resumes the task where it left off
  • Assuming the orchestrator rolls back a failed attempt's writes
  • Believing a task in a retrying state has already alerted someone
  • Setting large retry counts on every task as a default
  • Expecting retries to fire for a task that hangs instead of failing

context