skip to content

When should a Go service supervise a crashed worker in-process, and when should it exit and let the platform restart it?

level: principalimportance: nice to knowfreq 30%

answer

  1. two restart machines, one boundary
  2. local fault versus unprovable damage
  3. what does a process restart cost you
  4. residue accumulates across in-process restarts
  5. an infinite retry loop hides the fault from on-call

basics

~10 s

Supervise in-process when the failure is scoped to one worker and the process holds state worth keeping. Exit when the damage is process-wide or unprovable, or the restart budget is spent.

solid answer

~50 s

In-process supervision keeps the other workers serving, which matters when a process holds thousands of live sessions and a full restart means a fleet-wide reconnect storm. It is the right default while the failure is provably local: one device refusing connections, one bad payload. Exit instead when you cannot bound the damage — goroutines and sockets leaked on every relaunch, corrupted shared state, configuration only read at startup — because a fresh process is the only cheap way to reclaim any of that. The judgment is where the handoff sits: bound in-process restarts with a backoff cap, export a restart-rate counter, and once the cap is spent stop retrying and exit non-zero so the platform's restart policy takes over. That cap is a shared number, because the platform team owns process-level restart and paging policy and can overrule it.

go deeper

for a junior

Know that a crashed worker can be restarted inside the process or by letting the process exit, and that the two have very different costs for everything else running alongside it.

for a middle

Be able to compare the two: in-process restarts are cheap and touch one worker, a process restart is slow but is the only thing that reclaims leaked goroutines, sockets and stale startup configuration.

for a senior

Argue the boundary from evidence — resource counts rising with the restart counter, faults you cannot prove are local — and show the bounded design: capped backoff, a failure cap, and an exported restart rate.

for a principal

Own the posture end to end: where in-process retry stops and the platform's restart policy begins, why an unbounded loop is a governance failure rather than resilience, and how that handoff is written down and agreed with the team that carries the pager.

## Two restart machines, one service Every long-running Go service that supervises workers has two restart mechanisms available, and the interesting decision is where the boundary between them sits. **In-process supervision.** A supervisor goroutine relaunches the worker function. Cost: microseconds. Blast radius: one worker. Everything else in the process — the other device sessions, warm connection pools, in-memory routing state, the admin surface — keeps running untouched. **Process restart.** The process exits and whatever started it starts it again. Cost: seconds, plus whatever it takes to rebuild every session. Blast radius: everything. But it is the only mechanism that reclaims anything the process has irreversibly damaged. Neither is a default you can apply everywhere. Choosing per failure class is the job. ## When in-process wins The failure is **local and identified**: this device is refusing connections, this stream sent a frame the parser rejects. Nothing about it says the rest of the process is unsound. The process holds **expensive state**. A session manager with several thousand live sockets is the clearest case: restarting the process means every peer reconnects at once, and that reconnect storm can be worse for the fleet than the original fault — especially if the peers themselves retry aggressively. Keeping the process alive is what protects the ninety-nine percent that were fine. The restart is **cheap and complete**: relaunching the worker function genuinely re-establishes it, with no residue. That is a claim you should be able to defend, not assume. ## When exiting wins The **residue accumulates**. Each crashed worker leaves something behind — a goroutine blocked on a socket that was never closed, a file descriptor, an entry in a map keyed by session. Restarting in-process a thousand times leaves a thousand copies. If the goroutine count climbs in lockstep with the restart counter, in-process restarting is converting a fault into a slow leak, and only a fresh process reclaims it. The **damage is not provably local**. Shared mutable state that a worker may have left half-updated, a global cache the failure could have poisoned, an initialisation the code cannot redo idempotently. A process restart is a very cheap way to get back to a state you can reason about; "restart the goroutine and hope" is not. The fault is **environmental and only fixed by starting over**: stale configuration read at startup, a rotated credential, a dependency whose address changed. Retrying inside the same process cannot pick up any of those. The **budget is spent**. Past a cap of consecutive fast failures, in-process retry has stopped being a recovery strategy and become a way of hiding a fault. ## The organisational half, which is what makes this a judgment call A supervisor that retries forever is a decision to make a fault invisible. From the outside the process is up; its health endpoint answers; the platform's restart count is zero; no page fires. Degradation quietly becomes the steady state and nobody is accountable for it. That is why an unbounded retry loop is a governance problem, not just a technical one. The mirror image is equally real: a process that exits on the first error hands every fault to a restart policy tuned by a different team, with a different backoff, and a single flaky peer can drive a restart cycle that drops all the healthy sessions repeatedly. So the posture I argue for, and the one that survives being overruled: 1. **Bound in-process restarts.** A cap on consecutive fast failures per worker, with capped backoff between attempts. Named constants, not literals buried in a loop. 2. **Always export the restart rate.** Per worker, alerted on rate rather than total, alongside current-run uptime and the last exit error. If a crash loop can run without a page, the design is not finished. 3. **Escalate deliberately when the cap is spent.** Either mark that worker dead and keep serving the rest — an explicit, visible, alerted state — or return the error up so `main` exits non-zero. Both are defensible; drifting into infinite retry is not. 4. **Write the handoff down.** The platform team owns process-level restart pacing and the paging policy; the service team owns per-worker restart pacing. The cap is where those two meet, and it belongs in a runbook with the alert threshold beside it, because the on-call engineer needs to know which machine is currently restarting things and how to stop it. ## How I would present the tradeoff Framing it as a comparison of two numbers usually settles the argument fast: the cost of one process restart, measured in reconnects and seconds of unavailability across every session, against the cost of tolerating one degraded worker for the time it takes a human to respond. When the first number is large and the fault is local, supervise. When the first number is small, or the fault could be anywhere, exit — the fresh process is more trustworthy than any recovery you write by hand.

  • What evidence would change your mind from in-process restarting to letting the process exit?
    Resource counts that climb with the restart counter — goroutines, file descriptors, heap — because that says each restart leaves residue the process cannot reclaim. Also any fault whose scope I cannot bound: half-updated shared state, a poisoned cache, configuration only read at startup. Once recovery is unprovable, a fresh process is the cheaper and more honest fix.
  • The platform team wants your service to exit on any worker failure so their restart policy handles everything. How do you argue the case?
    With numbers rather than principle: this process holds several thousand live sessions, so every exit is a fleet-wide reconnect storm, while the fault in question is one unreachable peer. I would offer the bounded compromise — capped in-process restarts, an exported restart rate they can alert on, and a deliberate non-zero exit once the cap is spent — so their policy still governs the cases that matter.
  • What goes in the runbook for a service that supervises workers in-process?
    Which worker restarts are normal and which rate is an alert; where the restart counter and last-exit-error are exposed; the configured backoff ceiling and failure cap and how to change them; what the process does when the cap is spent; and who owns each layer, so on-call knows whether the service or the platform is currently deciding to restart things.

saying these in an interview costs you the question

  • Retries forever in-process with no cap or metric
  • Exits the whole process on any single worker failure
  • Assumes relaunching a goroutine reclaims its leaked resources
  • Never measures what a full process restart actually costs
  • Leaves the restart policy undocumented between the two teams