skip to content

Why should a supervised Python service leave restart backoff to the supervisor instead of retrying internally?

level: seniorimportance: should knowfreq 35%

answer

  1. A living process hides a dead service
  2. Restart counters are the visible signal
  3. Retry scope: one call versus the whole process
  4. Classify: skip, retry, or exit
  5. Backoff is a deployment-time decision

basics

~20 s

A process that catches everything and sleeps stays alive while doing nothing, so the supervisor sees a healthy service and restart-based alerting never fires. Exiting non-zero makes the failure visible and lets the supervisor apply its own backoff.

solid answer

~50 s

The supervisor already owns restart policy: how fast to retry, how much to back off, how many failures in a window before it gives up. A `while True:` around the main loop with a broad `except` and a `time.sleep` re-implements that badly and, worse, hides it — the process stays up, restart counters never move, and the dashboards keyed to them stay green while the service does no work. The rule is about *scope*, not about retries in general: bounded, jittered retries around a single unit of work against a flaky dependency are good engineering. What must not live in the program is the whole-process restart policy. So classify the failure: skip and count a bad record, retry a transient call, and for an unrecoverable condition log the reason to `sys.stderr` and return a non-zero code from `main()` so the supervisor can back off and someone can see it.

code

python · 14 lines
python
import time


def sync_batch() -> None:
    raise RuntimeError("peer record dated 15 minutes in the future")


# Anti-pattern: the process stays alive, so nothing outside it ever learns.
for _ in range(3):
    try:
        sync_batch()
    except Exception as exc:
        print(f"retrying after: {exc}")
        time.sleep(0.1)

go deeper

for a junior

Know that a supervised program is allowed to fail: exiting with a non-zero status is how it reports a problem, and something outside it decides when to try again. Catching everything and continuing is not robustness.

for a middle

Explain the mechanics of the anti-pattern — a catch-all loop keeps the process alive, so restart counters and exit codes never change — and describe the three responses to an error: skip the item, retry the call, or exit with a code.

for a senior

Show the judgement: classifying failures into per-item, transient and unrecoverable; keeping restart budgets in mind so a dependency blip does not exhaust them; and diagnosing a service that looked healthy for days while doing no work.

for a principal

Own the operating contract: which failures are allowed to exit, what the exit codes mean, how restart policy and alerting consume them, and why no service re-implements backoff in application code. Be ready to defend that against a team that wants its own retry supervisor.

## A concrete shape of the bug A team of four runs an inventory sync between two systems. The peer's clock drifts, and records start arriving stamped several minutes in the future; a freshness assertion in the sync loop raises. The service was written defensively — `while True:` around the loop, `except Exception:`, log the error, `time.sleep(5)`, continue. The process therefore **never exits**. The supervisor sees an uninterrupted uptime; the restart counter that alerting is keyed to stays at zero; the dashboard is green. Records stop flowing, and the team hears about it days later from the people on the other side of the sync. Had the process exited non-zero, the supervisor would have restarted it, backed off as the failure repeated, and the crash-loop would have been visible within minutes. ## Why the supervisor is the right owner Restart policy is a **deployment-time decision, not a code-time one**: - how soon to retry, - how the delay grows, - how many failures in a window before the supervisor stops trying, - whether to alert or page. The supervisor can change those without a release, applies them uniformly to every service on the machine, and — crucially — records them. Counters, restart timestamps and the exit codes it observed are the operational trail. A retry loop inside the process produces none of that; its state is a local variable. There is also a **correctness argument**. A restarted process starts from a known state. A process that swallowed an exception and continued is in whatever state the half-finished operation left behind: a partially applied batch, an open transaction, a connection the driver considers poisoned, a cache that no longer matches the source. Restarts are a cheap and reliable way to get back to a state you can reason about, and they are cheap precisely because the process is supervised. ## What this rule does not say It does not say never retry. Inside one unit of work, retrying a flaky call with bounded attempts, jittered delays and an overall deadline is exactly right, because the failure is expected, local and recoverable, and re-running the whole process to redo one call would be absurd. Nor does it say every error is fatal: one malformed record in a batch of thousands should be logged, counted and skipped, not turned into a process exit that stalls the other records forever. The distinction is between an error you can **classify and continue past**, and a condition under which **continuing is meaningless**. Unrecoverable usually means one of a few things: - configuration that cannot work no matter how long you wait, - a credential that is missing or rejected, - corrupted local state, - or a bug that has produced an invariant violation you cannot reason about. For those, the correct behaviour is to write a clear reason to the standard error stream and return a non-zero code from `main()`. ## Distinguish the codes Exiting non-zero is the minimum; exiting with a code that carries meaning is better. A **permanent misconfiguration** and a **transient dependency failure** both stop the process, but only the second gets better on its own, and a supervisor whose policy can distinguish codes — or an operator reading the logs — will handle them differently. If the supervisor gives up after N failures in a window, that is usually the behaviour you want for the permanent class and the behaviour you must design around for the transient class: crashing on every dependency blip can exhaust the restart budget and take the service down for a problem that resolved itself, so transient failures generally belong inside the request path with timeouts and retries, not in a process exit. ## Anti-patterns to name - The bare `except:` around the main loop, which also swallows the deliberate `SystemExit` a shutdown path raised. - The `time.sleep` before exiting, added to `slow down the crash loop` — it delays the supervisor's own accounting and makes the failure look slower than it is; the supervisor's backoff already exists for this. - And the process that re-executes itself with `os.execv` to `restart cleanly`: it replaces the image while keeping the pid, so from the supervisor's point of view nothing ever failed, and every diagnostic it would have collected is lost. ## The shape that works 1. Let unexpected exceptions propagate out of `main()` — the traceback lands on the standard error stream and the interpreter exits 1. 2. Catch the exceptions you can genuinely classify, and for each one decide explicitly: skip the item, retry the call, or return a code. 3. Log the decision so the reason survives in whatever collected the process's output. 4. Then leave the timing of the next attempt entirely to the thing that started you.

  • Where is an internal retry still the right call?
    Around a single unit of work against a dependency that fails transiently: bounded attempts, jittered delays and an overall deadline, so the operation either succeeds or gives up predictably. That is local recovery of an expected failure. What must not be internal is the whole-process restart policy, because that is the supervisor's job and the only place the failure becomes visible.
  • If the supervisor gives up after several restarts in a window, how do you keep a dependency outage from taking the service down permanently?
    Do not turn every dependency failure into a process exit. Handle transient downstream problems inside the request or batch path with timeouts, bounded retries and a degraded response, and reserve exits for conditions that will not fix themselves. Distinct exit codes help too: an operator or a policy can then tell a permanent misconfiguration from a transient failure.
  • What is wrong with re-executing the program with os.execv instead of exiting?
    It replaces the process image while keeping the same pid, so from the supervisor's point of view nothing ever failed: no restart is recorded, no backoff applies, no alert fires, and any diagnostics the supervisor would have captured are lost. It also carries over inherited state such as environment and open descriptors, which is the opposite of the clean slate a restart is meant to provide.

A smoke alarm that quietly resets itself every five seconds is not a resilient smoke alarm; it is a silent one.

saying these in an interview costs you the question

  • Wraps the main loop in while True with a bare except
  • Sleeps and continues after an unrecoverable error
  • Assumes a running process means a working service
  • Re-executes itself instead of exiting on failure
  • Treats one malformed record as fatal to the process
  • Implements its own exponential backoff before exiting

context