skip to content

A supervised TCP session worker dies at startup and is relaunched instantly, pinning a core. How do you pace and cap the restarts?

level: seniorimportance: should knowfreq 42%

answer

  1. a restart is a cost, so budget it
  2. grow the delay, but cap it
  3. wait in a select, never a bare sleep
  4. only a long healthy run clears the backoff
  5. the process stays up, so export a restart rate

basics

~20 s

Put a growing, capped delay between relaunches and wait it out in a select against the context. Reset the delay only after a run stayed up a while, and stop after N consecutive fast failures.

solid answer

~50 s

Relaunching on the exit signal with no delay turns an instant failure — a refused connection, a bad certificate, a config typo — into a hot loop that burns a core and floods the logs. Add exponential backoff: start around 200ms, double after each failed run, cap it near 30 seconds, and wait it out with `select { case <-ctx.Done(): ...; case <-time.After(delay): }` so shutdown does not sit through the full delay. Reset to the base only after a run stayed up past a health threshold; otherwise a worker that ran fine for hours reconnects at the ceiling after its first blip. Cap consecutive fast failures too — after N in a row, return the wrapped error and let the layer above decide. And export a per-worker restart counter, because a process in a crash loop still looks up.

code

go · 34 lines
go
const (
	baseDelay   = 200 * time.Millisecond
	maxDelay    = 30 * time.Second
	healthyRun  = time.Minute // a run this long clears the backoff
	maxRestarts = 8
)

func superviseSession(ctx context.Context, dial func(context.Context) error) error {
	delay, fails := baseDelay, 0
	for {
		start := time.Now()
		err := dial(ctx)
		if ctx.Err() != nil {
			return nil
		}
		restarts.Add(1) // exported counter: alert on its rate
		if time.Since(start) >= healthyRun {
			delay, fails = baseDelay, 0
			continue
		}
		fails++
		if fails >= maxRestarts {
			return fmt.Errorf("session failed %d times in a row: %w", fails, err)
		}
		select {
		case <-ctx.Done():
			return nil
		case <-time.After(delay):
		}
		if delay *= 2; delay > maxDelay {
			delay = maxDelay
		}
	}
}

go deeper

for a junior

Remember that a supervisor must not relaunch instantly: put a growing delay between attempts and stop after a set number of failures in a row.

for a middle

Explain the four moving parts — base delay, doubling, ceiling, and what resets it — and why the wait is written as a select over the context rather than a sleep.

for a senior

Diagnose the live incident: name the resource a hot restart loop burns, show the reset-on-healthy-run rule, cap consecutive failures, and say what you export so the loop is visible while the process still looks healthy.

for a principal

Own the numbers as policy: the ceiling is a statement about tolerated downtime, the cap is a statement about when a worker is declared dead, and both belong in a runbook with the alert that watches the restart rate.

## The symptom One device session is misconfigured — wrong port, expired certificate, DNS that resolves to nothing. The worker's `dial` returns in microseconds. The supervisor sees the exit, relaunches, and the whole cycle repeats thousands of times a second. What you observe: one core pinned at 100%, gigabytes of identical log lines, connection attempts hammering a remote endpoint that may itself be struggling, and, if each attempt allocates buffers or leaves a socket in TIME_WAIT, resource growth that eventually takes out the *healthy* sessions on the same process. The supervisor is doing exactly what it was told. The missing piece is that a restart is a cost, and costs need pacing and a budget. ## Pacing: exponential backoff, capped Start at a delay short enough that a genuine transient blip barely shows — 100 to 500 milliseconds — and double it after each failed run. Cap it: without a ceiling the delay reaches minutes and then hours, and a device that comes back is left offline long after it was fixed. Thirty seconds is a common ceiling for a reconnect loop; the right number is "how long am I willing to be down after the peer recovers?". Do not use `time.Sleep` for the wait. A sleeping supervisor cannot notice that the service is shutting down, so a 30-second backoff becomes 30 seconds of shutdown latency. Write the wait as a `select` over the context and a timer channel: ```go select { case <-ctx.Done(): return nil case <-time.After(delay): } ``` Now cancellation cuts the wait short, which is the difference between a service that stops in milliseconds and one that hangs long enough for the platform to kill it. ## Resetting: the mistake almost everyone makes first If the delay only ever grows, it is not a backoff — it is a ratchet. A session that ran healthily for six hours and then hit a one-second network blip will reconnect at the 30-second ceiling, because the counter still remembers failures from this morning. The fix is to reset on evidence of health, not on any successful start. Record the time before running the session; when it returns, if it lasted longer than a health threshold (a minute is a reasonable default for a reconnect loop), set the delay back to the base and clear the failure count. Anything shorter counts as a fast failure and advances the backoff. "It started" is not evidence — a worker in a crash loop starts every time. ## Budget: capping the restarts Backoff paces the damage; it does not stop it. A device that will never connect again — decommissioned, permanently misconfigured — deserves a decision, not an eternal retry at the ceiling. Keep a count of *consecutive* fast failures and, past a limit, return a wrapped error instead of looping. What happens next belongs to the layer above: another supervisor tier, the group joining all the device supervisors, or the process exiting so the platform's own restart policy takes over. The cap is what turns "this worker is broken" from a log line into an event something can act on. ## Seeing it: a process in a crash loop is still "up" This is the part that gets skipped, and it is why crash loops survive for weeks. From the outside the process is running, its health endpoint answers, its memory is flat. Nothing restarts, so the platform's restart count stays at zero and no alert fires. The signal has to come from inside: - a **restart counter per worker**, incremented on every relaunch — alert on the *rate*, for example more than a few restarts a minute sustained, not the total; - the **time the current run has been up**, which distinguishes flapping from stable; - the **last exit error**, kept per worker and exposed, so the reason does not require digging through log volume that the loop itself is generating; - a **rate limit on the log line**, or log only on backoff transitions, so the loop cannot drown everything else. One more corroborating signal: a crash loop that leaks — a socket not closed, a goroutine that outlives its session — shows up as a steadily rising goroutine count in the goroutine profile or `runtime/metrics`. If restarts are climbing and goroutines climb with them, the worker is not cleaning up after itself, and the backoff is only buying time. ## Where the numbers come from Base delay: below the noise floor of a transient network error. Ceiling: your tolerated downtime after the peer recovers. Health threshold: comfortably longer than the time a doomed session takes to fail. Failure cap: small enough that a permanently broken worker is escalated within minutes, large enough to ride out a dependency's own restart. Write them as named constants — they are policy, and someone on call will want to change them.

  • Why reset the backoff only after a long run rather than on every successful start?
    Because a worker in a crash loop starts successfully every time — starting is not evidence of health. If you reset on start, the delay never actually grows and the loop stays hot. Timing the run and resetting only when it lasted past a threshold, say a minute, separates a worker that is genuinely serving from one that dies immediately after connecting.
  • Why not just call time.Sleep for the backoff delay?
    A sleeping supervisor cannot observe cancellation, so shutdown has to wait out the full delay — with a 30-second ceiling that is 30 seconds of hang, long enough for a platform to kill the process mid-drain. Waiting inside a `select` over `ctx.Done()` and a timer channel makes the delay interruptible while behaving identically in the normal case.
  • The process is up, memory is flat, and every session worker is restarting twice a second. What fires an alert?
    Only something the process exports itself. A per-worker restart counter with an alert on its rate, plus the current run's uptime and the last exit error, are what make a crash loop visible; external restart counts stay at zero because the process never exits. A goroutine count climbing in step with restarts additionally says the worker is leaking on the way out.
  • What should the supervisor do once the consecutive-failure cap is hit?
    Stop looping and return the wrapped last error, so the decision moves up a layer rather than being retried forever in place. Above it, the choice is to keep the rest of the service running with that one worker marked dead, or to treat it as fatal and let the process exit so the platform's restart policy takes over — a policy call that should be explicit rather than implied by an infinite loop.

It is the difference between a doorbell and someone leaning on the button. Backoff makes each retry a knock; the cap is when you stop knocking and go find a key.

saying these in an interview costs you the question

  • Relaunches immediately with no delay between attempts
  • Grows the delay but never resets it after a healthy run
  • Resets the backoff on a successful start rather than a long run
  • Uses time.Sleep so shutdown waits out the full delay
  • Retries forever with no cap and no restart metric