In a data pipeline, which task failures are worth retrying and which will just fail again?
answer
- ask whether anything could differ next time
- the machine's fault versus the data's fault
- identical input, identical crash
- every wasted attempt is deadline you spent
- classify in the task, not in the default
basics
~20 sRetry failures whose cause may differ next attempt: timeouts, throttling, preempted workers, transient locks. Deterministic causes — bad input data, a schema mismatch, a missing permission, a code bug — fail identically every time, so retries only delay the alert and burn compute.
solid answer
~50 sSplit failures by whether anything about the world could plausibly change before the next attempt. **Transient**: connection resets, read timeouts, HTTP 429/503, a preempted or evicted worker, a lock or concurrency-slot conflict, a warehouse queue rejection. Those are worth a bounded retry chain. **Deterministic**: a column that does not exist, a type or constraint violation from the input file, a permission denied, a malformed credential, an unhandled exception in the transformation code. Every attempt reproduces them exactly, so the retry budget just adds attempts × delay to the time before a human hears about it — and pays for the failed compute each time. The practical move is to make the pipeline itself classify: catch the known-permanent conditions and re-raise them as a non-retryable error, and set validation-style tasks to zero retries so a bad file fails fast and loud.
code
python · 11 linesclass NonRetryable(Exception):
pass
def load_orders(window_start, window_end):
try:
rows = source.fetch(window_start, window_end)
except (ConnectionResetError, TimeoutError, RateLimited):
raise # transient: let the orchestrator spend an attempt
except (SchemaMismatch, PermissionDenied) as e:
raise NonRetryable(f"will fail identically on every attempt: {e}")
target.overwrite_window(window_start, window_end, rows)go deeper
Be able to sort common errors into two buckets — connection timeouts and rate limits are worth another attempt, missing columns and permission errors are not — and say why.
Explain the test you apply (could anything relevant differ next attempt?), and name the three costs of retrying a deterministic failure: delayed alert, wasted compute, misleading diagnosis.
Show how you encode the classification in production: non-retryable error types, zero retries on validation gates, generous budgets only at flaky external boundaries, and failure metrics labelled by category.
Own the platform default. Argue what the fleet-wide retry policy should be, how teams opt into a bigger budget, and how you keep retry spend and alert latency visible to the people setting deadlines.
## The question a retry policy is really asking A retry is a bet that the failure was about *conditions*, not about *inputs or code*. So the classification test is one sentence: **could anything relevant plausibly be different when the next attempt runs?** If yes, retry. If the same code will read the same bytes and hit the same rule, the retry is arithmetic — the same computation, the same result, one budget later. ## Failures that are worth retrying - **Network and connection errors** — resets, DNS blips, TLS handshake failures, read timeouts talking to a source API or a warehouse endpoint. - **Throttling and back-pressure** — rate-limit rejections, a warehouse rejecting a query because the queue or concurrency slots are full, a service returning "try again later". - **Infrastructure churn** — a spot or preemptible worker reclaimed mid-task, a pod evicted under memory pressure, a node drained during a rolling upgrade. The task did nothing wrong; the machine disappeared. - **Concurrency conflicts** — a lock held by another writer, a serialization failure, a metadata operation that clashed with a concurrent one. - **Eventual-consistency artefacts** — a file listed but not yet fully visible, a table's metadata not yet propagated. These share a shape: the cause is external, stateful and short-lived. ## Failures where retrying is pointless - **Schema and contract violations** — a referenced column is absent, a type changed, a NOT NULL column contains nulls, a JSON payload does not parse. - **Bad input data** — a malformed row, an unparseable date, a value outside the allowed domain. - **Authorization and configuration** — permission denied, an expired or wrong credential, a missing bucket or table, a misspelled path. - **Code defects** — a null dereference, an index error, a division by zero in the transformation. - **Resource sizing that is deterministically wrong** — a job that runs out of memory at the same point of the same input every single time. (Note the subtlety: an out-of-memory error is transient if the box was shared and busy, deterministic if the input simply does not fit. That is exactly why classification is a judgement, not a lookup table.) ## Why retrying the deterministic ones actively hurts Three costs, and interviewers want all three: 1. **It delays detection.** A task with 4 attempts and a 10-minute delay pushes the failure alert out by the better part of an hour. If the pipeline had a 06:00 deadline, you have spent the recovery window discovering something the first attempt already knew. 2. **It costs money.** Each attempt re-runs whatever ran before the failure point — often a full scan or a large read — and cloud warehouses and clusters bill for the failed work. 3. **It obscures the diagnosis.** Four identical stack traces in the log make the failure look like an infrastructure problem when it is a data problem, and on-call reaches for the wrong runbook. ## The third category: not a failure at all There is a case that looks deterministic on attempt one and transient on attempt three — the upstream file simply has not landed yet. A retry chain will eventually "work" here, and that is a trap. Lateness is a different condition from failure: it deserves an explicit wait with its own deadline, so the log says "still waiting for the 02:00 extract at 04:30" instead of "FileNotFoundError, attempt 3 of 5". Conflating the two is how teams end up with retry budgets sized to cover an upstream's tardiness, and no signal at all when the upstream stops arriving entirely. ## Making the pipeline classify for you Defaults cannot tell a 429 from a missing column, so encode the distinction in the task: - Catch the conditions you know to be permanent and re-raise them as a distinct non-retryable error type; where the orchestrator supports it, that error type ends the task immediately without consuming the budget. - Give schema checks, contract checks and input validation zero retries, and run them as an early, cheap gate — a bad file should fail in the first minute, not after the expensive load has been attempted four times. - Give the genuinely flaky boundaries — the external API, the shared warehouse — their own generous budget with a growing delay. - Emit the failure category as a label on your failure metric, so "we retried 400 times this week, 380 of them for the same throttling problem" is a visible fact rather than folklore. The payoff is that the retry budget starts meaning something: attempts are spent only where an attempt can help, and the first sign of a data or code defect is an alert rather than a stalled pipeline.
- How would you make the pipeline itself distinguish retryable from non-retryable failures?Catch the known-permanent conditions inside the task and re-raise them as a distinct non-retryable exception type that the orchestrator is configured to fail immediately, and set validation-style tasks to zero retries. Everything unclassified keeps a small default budget. Label the failure metric with the category so the split is visible in aggregate.
- An out-of-memory failure — transient or deterministic?It depends on the cause, which is why it deserves a look rather than a default. On a shared or oversubscribed worker it is transient and a retry elsewhere may succeed. If the same input always exceeds the allocation at the same point, it is deterministic and retries just repeat an expensive crash; the fix is more memory, a smaller partition, or streaming the input.
- A task fails because the upstream file has not landed yet. Should it retry?No — that is lateness, not failure. Model it as an explicit wait with a deadline, so the pipeline reports "still waiting past the expected arrival time" instead of a stack trace. Retry chains that cover an upstream's tardiness hide the difference between late and never coming, and you lose the signal when the source stops entirely.
saying these in an interview costs you the question
- Setting the same generous retry budget on every task regardless of error
- Claiming retries make a pipeline resilient to bad input data
- Treating a permission or credential error as something backoff can cure
- Using a long retry chain as a substitute for waiting on late data
- Never inspecting which errors are actually consuming the retry budget