skip to content

Your controller recovers every reconcile panic and keeps running; the on-call SRE wants it to crash instead. How do you decide?

level: principalimportance: should knowfreq 30%

answer

  1. ask what the panic can damage
  2. one bad item versus shared state
  3. a released lock is not a restored invariant
  4. restart only helps if state is rebuildable
  5. clean up, record, then decide

basics

~20 s

Decide by what a panic can damage. If the failure is confined to one item and shared state is untouched, recover and fail that item. If shared in-memory state can be left half-updated, crashing and restarting clean is safer.

solid answer

~50 s

This is a blast-radius question, not a style argument. Recovering per reconcile is right when a panic can only damage the one item being processed: its entry in the desired-state cache is written whole or not at all, no lock was left protecting a half-updated structure, and the item can be marked failed and requeued — one poisonous object should not stop a thousand healthy ones. Crashing is right when a panic can leave shared state torn, because the process then keeps serving data nobody validated and the damage spreads silently. The SRE's case is strongest when restart is cheap: state rebuilds from the source of truth, a supervisor exists, no restart storm follows. My usual answer is both — recover at the boundary to log the stack and count panics per key, quarantine a repeat offender, and re-panic after cleanup when the rate says the problem is systemic.

code

go · 14 lines
go
func (c *Controller) runOne(k Key) (err error) {
	defer func() {
		r := recover()
		if r == nil {
			return
		}
		c.logger.Printf("reconcile %v panicked: %v\n%s", k, r, debug.Stack())
		if c.stateTorn(k) {
			panic(r)
		}
		err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r)
	}()
	return c.reconcile(k)
}

go deeper

for a junior

Know the two options exist and what separates them: containing a failure to one item versus ending the process so it restarts with clean state. You are not expected to own this call yet.

for a middle

Be able to argue both sides on mechanics: what a recovered panic may have left behind in shared memory, and what a restart actually resets.

for a senior

Show the operational reasoning — blast radius, restart cost, crash-loop risk — and the hybrid where the handler records evidence before deciding to continue or re-panic.

for a principal

Own the posture as policy: who decides, what evidence settles it, how escalation thresholds are encoded, and how you disagree with an on-call team without turning it into a configuration flag.

## Framing the decision honestly The disagreement looks like "availability versus safety", but the useful question is narrower: **what can a panic in a reconcile leave behind?** Everything else follows from the answer. Go makes this decision unusually sharp for two reasons. First, an unrecovered panic in any goroutine takes down the **whole process**, so the default with no boundary is maximum blast radius. Second, a recovery boundary is trivially cheap to install, so the temptation is to wrap everything and never think about it again — which is how a service ends up running for days on a cache that a half-finished update corrupted. ## When recovering per reconcile is right Recover and continue when all of the following hold: - **The unit of work is genuinely isolated.** The panic can only damage the item being reconciled. Its entry in the desired-state cache, keyed by namespace and name, is replaced atomically at the end of a successful reconcile, so a crash mid-way leaves the previous value intact. - **No shared invariant can be half-established.** Nothing was mutating a shared structure with a lock held. Note that `defer mu.Unlock()` still runs during unwinding, which is exactly the danger: the lock is released and other goroutines proceed over a structure that was left mid-update. Releasing the lock is not the same as restoring the invariant. - **The failure is reportable and requeueable.** The item can be marked failed with the captured stack, retried later, or quarantined. Under those conditions crashing is actively bad: one malformed object becomes a full outage, and if the object is still there after restart you get a crash loop that never converges. ## When crashing is right Crash — or recover, clean up and re-panic — when a panic can put the process into a state you would not trust: - The panic happened while a shared cache, index or queue was being restructured. - A lock's invariant was broken mid-update and the deferred unlock published the broken state. - The panic indicates an assumption failure so fundamental (a corrupt decoded structure, an impossible internal state) that continuing means guessing. Here the restart is not a failure, it is the recovery mechanism: the process comes back with empty in-memory state and rebuilds it from the source of truth. The precondition for that argument is operational, and it is what the SRE is really asserting: there **is** a supervisor, startup is fast, the cache can be rebuilt, and a mass restart will not stampede the API it re-lists from. If those are false, "just crash" is a much weaker position and you should say so. ## The shape that usually settles the argument Rarely is the answer purely one or the other. The boundary can recover *and* decide: ```go defer func() { r := recover() if r == nil { return } c.logger.Printf("reconcile %v panicked: %v\n%s", k, r, debug.Stack()) if c.stateTorn(k) { panic(r) // shared state suspect: die and restart clean } err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r) }() ``` Recovering first buys the things a bare crash loses: the stack captured while the frames still exist, the identity of the item, a metric incremented, a flush of buffered telemetry. Re-panicking afterwards keeps the process-death outcome the SRE wants, and the runtime still reports the panic. That combination — clean up, record, then die — is usually the settlement, because it concedes the safety argument while keeping the evidence. The second half of the settlement is **escalation**. A single recovered panic on one key is an item-level failure; the same key panicking five times is a poison record that should be quarantined and alerted on, not retried forever; a panic rate across many keys is a systemic bug, and continuing to serve is no longer a kindness. Encoding those thresholds turns the posture from a belief into behaviour you can review. ## Who owns it, and how to disagree well The service owner writes the policy, but this is a decision the SRE can legitimately overrule, because it determines whether a process is allowed to keep serving on state nobody has validated — and they are the ones who carry the consequence at 3am. What makes the conversation productive is bringing evidence rather than preference: what the panics actually were over the last quarter, whether any of them touched shared state, how long a restart takes, and what a crash loop would do to the systems downstream. Write the outcome down as an operational contract — what the process survives, what it dies on, what pages — rather than leaving it implicit in a `recover` somebody added in a hotfix. One thing to resist: making it a configuration flag to avoid deciding. A flag doubles the number of behaviours to reason about and the one production uses will be whatever it was set to during an incident. ## Summary Contain what is isolated, die on what is corrupt, and prove which is which. Recover to capture evidence and clean up even when the answer is to crash; escalate on repetition; agree the posture with the people who get paged, and record it as policy rather than as a line of code nobody revisits.

  • A deferred unlock runs during unwinding, so the mutex is released. Why is that not reassuring?
    Releasing the lock restores availability, not correctness. If the panic happened halfway through updating the structure the lock protects, the unlock publishes a half-updated view to every goroutine that was waiting. The invariant the lock existed to protect is gone, and the process now serves from it happily — which is the strongest argument for crashing rather than recovering.
  • If you agree to crash, why recover at all before re-panicking?
    To keep the evidence and finish the cleanup. Inside the handler you can capture the stack while the frames still exist, name the item that failed, increment a metric and flush buffered telemetry. Re-panicking afterwards still ends the process, so the SRE gets the outcome they asked for and the next person gets a diagnosis rather than a bare traceback.
  • What would make you reject the SRE's crash-and-restart position outright?
    If restart is not actually cheap: slow startup, a cache that takes minutes to rebuild, no supervisor, or a re-list that stampedes the upstream API when many instances restart together. Crash-only is a design that borrows against fast, safe restarts; where that credit does not exist, crashing converts a contained bug into an outage.
  • How do you keep the posture from drifting once it is agreed?
    Write it down as an operational contract — what the process survives, what it dies on, what pages — and encode the escalation in code: panic counts per key and per minute, quarantine after repeats, deliberate abort above a rate. Then a recovery added in a hotfix is a visible change to policy rather than an invisible one.

saying these in an interview costs you the question

  • Recovering everywhere so the process can never die, with no other reasoning
  • Treating crash-only as always correct regardless of restart cost
  • Assuming a deferred unlock leaves the protected structure consistent
  • Making it a configuration flag instead of deciding a default
  • No escalation: the same key panicking forever with no quarantine or alert