skip to content

What rule would you set for os.Exit and log.Fatal in shared Go packages, and how would you enforce it?

level: principalimportance: nice to knowfreq 28%

answer

  1. Who owns the process makes this call
  2. A library cannot know its caller
  3. Fatal is an exit, not a log line
  4. Think about what a test binary does
  5. A rule nobody checks decays

basics

~20 s

Confine process exit to package main. A shared package calling os.Exit or log.Fatal takes a whole-program decision away from every caller and skips their cleanup, so libraries return errors and one place in main flushes, then exits with a status.

solid answer

~50 s

Ending the process is a decision that belongs to whoever owns the process, not to a package that happens to be three frames down. My rule is that `os.Exit` and `log.Fatal` appear only in `package main`, and ideally at one call site there, reached after everything that must be flushed has been flushed. The reasons are concrete: a package that exits cannot be reused by a long-running service, cannot be tested (it kills the test binary and takes the report of which test failed with it), and destroys the caller's deferred cleanup and buffered output. Enforcement is a CI check over non-main packages plus a short allowlist for the exceptions you consciously accept, because a rule nobody can check decays within a quarter. Where I would be flexible is a small internal one-shot tool; where I would not is anything another team imports.

go deeper

for a junior

Know the safe default: your package returns errors and lets main decide whether the program stops. Recognise log.Fatal as an exit call rather than a log line.

for a middle

Be able to explain the concrete damage: skipped defers, unflushed writers, a test binary killed mid-run, and a package that cannot be reused inside a long-running service.

for a senior

Show you can apply this in review and migration: find the existing exits, prioritise the paths that produce output, and propose a check rather than a convention nobody verifies.

for a principal

Own the rule and its exceptions. Price the inconvenience honestly, say which boundaries are non-negotiable and which are not, and be ready to be overruled on a tool where an immediate hard stop is genuinely safer.

## The decision being made Calling `os.Exit` is not an error-handling choice; it is a statement about the whole program. It says: whatever else this process was doing, whatever cleanup was pending, whatever output was buffered, none of it matters. That is a legitimate decision — but only the owner of the process can make it, because only they know what else is running. A shared package does not know. It might be linked into a CLI where exiting is fine, a long-running service where exiting is an outage, a test binary where exiting destroys the run, or another team's batch job where exiting truncates a report. Which is why the rule is about *where* the call may live rather than *whether* exiting is ever right. ## The rule **`os.Exit` and `log.Fatal`/`Fatalf`/`Fatalln` may appear only in `package main`.** Everything else returns an error. Inside `main`, aim for a single exit site, positioned after buffered writers have been flushed and their errors checked, so the exit is the last thing that happens and not a shortcut taken halfway through. Two corollaries worth writing down explicitly, because they are the ones people get wrong: - `log.Fatal` is an exit call. Teams that ban `os.Exit` in libraries and leave `log.Fatal` alone have banned nothing; the Fatal family prints and then calls `os.Exit(1)`. - A `defer` does not save you inside `main` either. `defer w.Flush()` in `main` does not run if `main` itself ends with `os.Exit`. The flush must precede the exit on every path. ## What the rule buys - **Testability.** A package that exits cannot be tested for that path at all: the test binary dies, usually without telling you which test was running. The equivalent in tests is `t.Fatal`, which ends only the test's goroutine (via `runtime.Goexit`) and lets cleanup run — exactly the scoped behaviour library code should be aiming for. - **Reusability.** The utility written for a cron job gets imported by a service six months later. If it exits on bad input, the service now has a crash path it cannot intercept. - **Diagnosability.** An error that travels up the stack can be wrapped with context at each layer. An exit produces one line and a status code. - **Data safety.** Every caller's deferred flush, close and unlock is skipped. In a tool whose output *is* the deliverable, that is the difference between a usable report and a truncated one. ## What it costs, and where I would bend The cost is real and worth naming rather than waving away. `log.Fatalf("bad config: %v", err)` is one line; threading an error up through four callers is a signature change at each of them. In a fifty-line internal tool with one author, insisting on the ceremony is a poor trade, and I would accept exits in a `main` package's own helpers there. Where I would not bend: anything with an import path other teams use, anything that might end up inside a server, and anything whose output must be complete to be worth anything. Those are the boundaries you cannot renegotiate cheaply once people depend on them. There is also a genuine counter-argument to hear out: for a condition where continuing would corrupt data, exiting immediately can be safer than returning an error a caller might swallow. My answer is usually `panic` rather than `os.Exit` — it runs the defers, prints a stack that names the cause, and still leaves a caller the option to recover at a boundary — but if someone owns a tool where an immediate hard stop really is the safer failure mode, that is a decision to document and scope, not to spread. ## Enforcement A rule without a check is a preference. In order of what actually holds: 1. **A CI check** that fails when `os.Exit` or the `log` Fatal family appears outside `package main`, with a small, reviewed allowlist. It runs on every change, so new violations cannot arrive quietly. 2. **An allowlist that is visible and short.** Each entry names who accepted the exception and why. Growth in that file is the signal that the rule needs revisiting. 3. **Making the right path the easy one.** If every tool in the repo has the same shape — helpers return errors, `main` flushes and exits — the correct code is the code people copy. 4. **Review as the last line, not the first.** Reviewers miss a `log.Fatal` in a diff of two hundred lines; the check does not. Migrating an existing codebase, I would gate new packages immediately, fix existing violations on the paths that produce output first (those are the ones losing data), and let the rest come with normal churn rather than opening a large mechanical change nobody can review. ## What an interviewer listens for That you treat exiting as an ownership question rather than a style question, that you can price the inconvenience honestly instead of declaring the rule obviously correct, that you know `log.Fatal` is the same call, and that you propose something checkable rather than a paragraph in a wiki.

  • An engineer argues log.Fatal is fine because their package is internal. How do you respond?
    Internal packages get imported by other internal programs, including test binaries and services, and the exit is invisible at the call site — the caller sees a function that returns nothing unusual, right up until the process is gone. The convenience saved is one line; the cost is a crash path no caller can intercept. I would rather grant a named, reviewed exception than let the general habit spread.
  • A package hits a condition where continuing would corrupt data. What should it do?
    Return the error if the caller can plausibly act on it. If the condition really is a broken invariant, `panic` is the better hard stop than `os.Exit`: it runs deferred calls, prints a stack that names the cause, and a server can still recover at its request boundary. `os.Exit` gives up all three and hands the caller a bare status code.
  • How would you actually enforce this without slowing every review down?
    A CI check over the repository that fails on os.Exit or the log Fatal family outside package main, with a short allowlist recording who accepted each exception. Automation catches the routine cases; review then only has to argue about the exceptions, which is where judgment is worth spending.
  • How does this rule interact with tests?
    Directly. A `log.Fatal` reached during a test kills the test binary, so you lose the remaining tests and often the report of which one failed. The in-test equivalent is `t.Fatal`, which ends only the test's goroutine through `runtime.Goexit` and still runs deferred cleanup — the scoped stop that library code should be modelling.

A library calling os.Exit is a subcontractor deciding to evacuate the whole building because their own task went wrong.

saying these in an interview costs you the question

  • Treats log.Fatal as logging rather than a process exit
  • Lets any package exit because the error is fatal anyway
  • Allows library exits behind a global disable flag
  • Ignores that exiting skips every caller's cleanup
  • States a rule with no way to check it in CI
  • Bans exits everywhere, including main, with no alternative