skip to content

Your Go migration tool exits 0 even when the migration failed. Which termination paths produce that status?

level: seniorimportance: should knowfreq 35%

answer

  1. green means nobody set the status
  2. the error stopped travelling somewhere
  3. logged is not the same as reported
  4. a goroutine failure main never collected
  5. only os.Exit changes the number

basics

~20 s

A Go process reports 0 whenever main returns, so an error that was only logged, a discarded return value, or a failure in a goroutine main never waited for all end green. Only a non-zero os.Exit makes failure visible.

solid answer

~50 s

Every way out of `main` except `os.Exit` with a non-zero code, or an unrecovered panic, gives status 0 - so I look for the point where the error stopped travelling. The usual causes: the failure was logged and then ignored; the returned error was assigned to `_`; the work happened in a goroutine whose result `main` never collected before returning; or a deferred `recover` in a worker turned a panic into a log line. The fix is structural rather than a patch at the failure site: the work returns an error to a single `run() error`, `main` prints it to `os.Stderr` and calls `os.Exit(1)`, and anything concurrent reports back through a channel or an error group that `run` waits on. Then I add a check that runs the tool against a deliberately broken migration and asserts the status is non-zero, because that number is the only thing CI reads.

code

go · 6 lines
go
func main() {
	if err := applyAll(); err != nil {
		log.Printf("migration failed: %v", err)
	}
	// main returns here, so the process exits 0 and CI reports success
}

go deeper

for a junior

Know the trap: an error you only log does not change the exit status. If a failure has to fail the build, it must reach main and become an os.Exit with a non-zero code.

for a middle

Be able to enumerate how a process ends up at 0 - main returned, the error was logged or assigned to the blank identifier, a goroutine's failure was never collected - and how each one is closed off.

for a senior

Diagnose it as a plumbing problem: find the frame where the error stops travelling, restore a single reporting path into main, and prove it with a check that asserts a non-zero status on a known-bad input.

for a principal

Decide what the status must mean across a fleet of tools so pipelines can trust it: partial success, retryable versus fatal, and who reviews a change that alters the code a tool returns.

## Why this is the most expensive exit-status bug A CI job, a deployment step or a scheduled runner judges a program by one integer. If the tool prints a stack of red text and then exits 0, everything downstream proceeds as though the work succeeded. Nobody looks at the log of a green step. The failure surfaces later, somewhere else, as data that does not match the schema. ## Every path that lands on 0 **main returned.** This is the default and the root of all the cases below. The runtime exits with 0 the moment `main` returns, regardless of what happened. **The error was logged and dropped.** `if err != nil { log.Printf(...) }` with no `return`, no `os.Exit`, no propagation. The most common form, and it reads as handled code. **The error was assigned to the blank identifier**, or a function that returns `(T, error)` was called for its first result only in a context where the compiler does not object - for example a deferred `Close` whose error is never examined, or a wrapper that returns only `T`. **A goroutine's failure was never collected.** `main` starts work concurrently, the work fails, and either nothing sends the error back or nobody receives it. When `main` returns, the process ends and the surviving goroutines are simply gone - unread errors included. **A recover swallowed a panic.** A worker defers a function that recovers and logs. The crash that would have exited with 2 becomes a log line, and `main` returns normally. **os.Exit(0) on a path that is not success.** Rare but real: a "nothing to do" branch that turns out to also cover "could not determine what to do". ## Diagnosing it Trace the error backwards from the log line that shows the failure. At each frame ask: does this function return the error, and does its caller do anything with it? The break is usually one frame - a helper that logs instead of returning, or a call site that ignores the second result. `go vet` catches some ignored results and shadowed error variables, and a review pass over every `if err != nil` block that contains no `return` finds most of the rest. For the concurrent case, the question is different: which function waits for the work, and what does it do with what the work reported? If nothing waits, the failure never had a route to `main` at all. ## Closing it off - One `run() error` that all the work reports into, and one `os.Exit(1)` in `main`. - Concurrent work reports through a channel or an error group that `run` waits on, so a worker's failure becomes `run`'s return value. - A `recover` that only logs is a bug unless the recovering function also turns the failure into a returned error. - Where a failure must be non-fatal - one of two hundred files could not be processed - count the failures and let the count decide the status, so "mostly worked" is still an explicit decision rather than an accident. ## Proving it Two levels of test. Unit-test `run` and assert it returns an error for a broken input. Then one end-to-end check that builds the binary, runs it against a known-bad migration and asserts a non-zero status. Only the second catches a `main` that forgets to call `os.Exit`, and that is exactly the bug that produced the green run in the first place. ## The one failure that cannot hide An unrecovered panic takes the whole process down with status 2 and a dump of the goroutines, from any goroutine, whether or not `main` was waiting. It is the only failure mode here that reports itself without cooperation, which is a good reason to be suspicious of a blanket recover that only writes a log line.

  • A worker goroutine panics, but each worker defers a function that recovers and logs. What status does the tool report?
    0, provided main then returns normally. recover turns the crash that would have exited 2 into an ordinary log line, so the failure never reaches the status. A recovering function has to convert the failure into an error that travels back to main, or the panic has been deleted rather than handled.
  • How do you test that the tool really exits non-zero?
    At two levels. Unit-test run() and assert it returns an error for a broken input. Then add one end-to-end check that builds the binary, runs it against a deliberately broken migration and asserts the exit status is non-zero. Only the second catches a main that forgot os.Exit.
  • Does an unrecovered panic in a non-main goroutine still fail the run?
    Yes. A panic nothing recovers takes the whole process down: the runtime prints the panic value and a dump of every goroutine to standard error and exits with status 2. It is one of the few failures that cannot pass silently, which is why a blanket recover that only logs is a hazard.
  • Half the migration steps succeeded and half failed. What status should the tool report?
    Non-zero, unless partial application is an explicitly supported outcome. Count the failures, report them, and decide the status from the count rather than from whichever step ran last. If callers need to distinguish 'nothing applied' from 'partially applied', that is a documented second code, not an inference from the log.

saying these in an interview costs you the question

  • Says the run failed because the log shows an error
  • Assumes a goroutine's error reaches main by itself
  • Thinks a recover that logs preserves the failure
  • Believes writing to os.Stderr sets a non-zero status
  • Sprinkles os.Exit at every failure site as the fix