When should a Go service supervise a crashed worker in-process, and when should it exit and let the platform restart it?
answer
- two restart machines, one boundary
- local fault versus unprovable damage
- what does a process restart cost you
- residue accumulates across in-process restarts
- an infinite retry loop hides the fault from on-call
basics
~10 sSupervise in-process when the failure is scoped to one worker and the process holds state worth keeping. Exit when the damage is process-wide or unprovable, or the restart budget is spent.
solid answer
~50 sIn-process supervision keeps the other workers serving, which matters when a process holds thousands of live sessions and a full restart means a fleet-wide reconnect storm. It is the right default while the failure is provably local: one device refusing connections, one bad payload. Exit instead when you cannot bound the damage — goroutines and sockets leaked on every relaunch, corrupted shared state, configuration only read at startup — because a fresh process is the only cheap way to reclaim any of that. The judgment is where the handoff sits: bound in-process restarts with a backoff cap, export a restart-rate counter, and once the cap is spent stop retrying and exit non-zero so the platform's restart policy takes over. That cap is a shared number, because the platform team owns process-level restart and paging policy and can overrule it.
go deeper
Know that a crashed worker can be restarted inside the process or by letting the process exit, and that the two have very different costs for everything else running alongside it.
Be able to compare the two: in-process restarts are cheap and touch one worker, a process restart is slow but is the only thing that reclaims leaked goroutines, sockets and stale startup configuration.
Argue the boundary from evidence — resource counts rising with the restart counter, faults you cannot prove are local — and show the bounded design: capped backoff, a failure cap, and an exported restart rate.
Own the posture end to end: where in-process retry stops and the platform's restart policy begins, why an unbounded loop is a governance failure rather than resilience, and how that handoff is written down and agreed with the team that carries the pager.
## Two restart machines, one service Every long-running Go service that supervises workers has two restart mechanisms available, and the interesting decision is where the boundary between them sits. **In-process supervision.** A supervisor goroutine relaunches the worker function. Cost: microseconds. Blast radius: one worker. Everything else in the process — the other device sessions, warm connection pools, in-memory routing state, the admin surface — keeps running untouched. **Process restart.** The process exits and whatever started it starts it again. Cost: seconds, plus whatever it takes to rebuild every session. Blast radius: everything. But it is the only mechanism that reclaims anything the process has irreversibly damaged. Neither is a default you can apply everywhere. Choosing per failure class is the job. ## When in-process wins The failure is **local and identified**: this device is refusing connections, this stream sent a frame the parser rejects. Nothing about it says the rest of the process is unsound. The process holds **expensive state**. A session manager with several thousand live sockets is the clearest case: restarting the process means every peer reconnects at once, and that reconnect storm can be worse for the fleet than the original fault — especially if the peers themselves retry aggressively. Keeping the process alive is what protects the ninety-nine percent that were fine. The restart is **cheap and complete**: relaunching the worker function genuinely re-establishes it, with no residue. That is a claim you should be able to defend, not assume. ## When exiting wins The **residue accumulates**. Each crashed worker leaves something behind — a goroutine blocked on a socket that was never closed, a file descriptor, an entry in a map keyed by session. Restarting in-process a thousand times leaves a thousand copies. If the goroutine count climbs in lockstep with the restart counter, in-process restarting is converting a fault into a slow leak, and only a fresh process reclaims it. The **damage is not provably local**. Shared mutable state that a worker may have left half-updated, a global cache the failure could have poisoned, an initialisation the code cannot redo idempotently. A process restart is a very cheap way to get back to a state you can reason about; "restart the goroutine and hope" is not. The fault is **environmental and only fixed by starting over**: stale configuration read at startup, a rotated credential, a dependency whose address changed. Retrying inside the same process cannot pick up any of those. The **budget is spent**. Past a cap of consecutive fast failures, in-process retry has stopped being a recovery strategy and become a way of hiding a fault. ## The organisational half, which is what makes this a judgment call A supervisor that retries forever is a decision to make a fault invisible. From the outside the process is up; its health endpoint answers; the platform's restart count is zero; no page fires. Degradation quietly becomes the steady state and nobody is accountable for it. That is why an unbounded retry loop is a governance problem, not just a technical one. The mirror image is equally real: a process that exits on the first error hands every fault to a restart policy tuned by a different team, with a different backoff, and a single flaky peer can drive a restart cycle that drops all the healthy sessions repeatedly. So the posture I argue for, and the one that survives being overruled: 1. **Bound in-process restarts.** A cap on consecutive fast failures per worker, with capped backoff between attempts. Named constants, not literals buried in a loop. 2. **Always export the restart rate.** Per worker, alerted on rate rather than total, alongside current-run uptime and the last exit error. If a crash loop can run without a page, the design is not finished. 3. **Escalate deliberately when the cap is spent.** Either mark that worker dead and keep serving the rest — an explicit, visible, alerted state — or return the error up so `main` exits non-zero. Both are defensible; drifting into infinite retry is not. 4. **Write the handoff down.** The platform team owns process-level restart pacing and the paging policy; the service team owns per-worker restart pacing. The cap is where those two meet, and it belongs in a runbook with the alert threshold beside it, because the on-call engineer needs to know which machine is currently restarting things and how to stop it. ## How I would present the tradeoff Framing it as a comparison of two numbers usually settles the argument fast: the cost of one process restart, measured in reconnects and seconds of unavailability across every session, against the cost of tolerating one degraded worker for the time it takes a human to respond. When the first number is large and the fault is local, supervise. When the first number is small, or the fault could be anywhere, exit — the fresh process is more trustworthy than any recovery you write by hand.
- What evidence would change your mind from in-process restarting to letting the process exit?Resource counts that climb with the restart counter — goroutines, file descriptors, heap — because that says each restart leaves residue the process cannot reclaim. Also any fault whose scope I cannot bound: half-updated shared state, a poisoned cache, configuration only read at startup. Once recovery is unprovable, a fresh process is the cheaper and more honest fix.
- The platform team wants your service to exit on any worker failure so their restart policy handles everything. How do you argue the case?With numbers rather than principle: this process holds several thousand live sessions, so every exit is a fleet-wide reconnect storm, while the fault in question is one unreachable peer. I would offer the bounded compromise — capped in-process restarts, an exported restart rate they can alert on, and a deliberate non-zero exit once the cap is spent — so their policy still governs the cases that matter.
- What goes in the runbook for a service that supervises workers in-process?Which worker restarts are normal and which rate is an alert; where the restart counter and last-exit-error are exposed; the configured backoff ceiling and failure cap and how to change them; what the process does when the cap is spent; and who owns each layer, so on-call knows whether the service or the platform is currently deciding to restart things.
saying these in an interview costs you the question
- Retries forever in-process with no cap or metric
- Exits the whole process on any single worker failure
- Assumes relaunching a goroutine reclaims its leaked resources
- Never measures what a full process restart actually costs
- Leaves the restart policy undocumented between the two teams