Where in a Go service do you allow slog.Error, and where do you require returning the error instead?
answer
- who is allowed to say a failure is reported?
- the rule is one line; the exceptions need writing down
- structural beats a review checklist
- a library imposing output on its importers
- on-call and the storage bill pull opposite ways
basics
~20 sDraw the line at ownership: a package that returns an error must not report it; only layers where errors stop travelling may log. Write the exceptions down and enforce the boundary structurally, not by review habit.
solid answer
~50 sThe convention I set is that a failure gets exactly one report, written where it stops propagating - `main` after `run()` returns, a request boundary, or the top of a worker goroutine - and every package below that returns wrapped errors and never imports a logger. I write the exceptions down rather than leaving them to taste: goroutines with no caller, best-effort work, retries that later succeeded, and errors deliberately dropped in a deferred call. Enforcement has to be structural: keeping the logger out of the lower packages makes a violation an import a check can catch, where a review habit decays. The part I hold loosely is what the single record must contain: if on-call cannot diagnose from it, the answer is richer wrapping or typed errors carrying attributes, not permission to log lower down.
go deeper
Know the rule you will be asked to follow: if your function returns the error, do not log it as well, and never call log.Fatal outside main.
Be able to point at the report site in the codebase you work on and say why it sits there, and to move a log line into a wrap message without losing detail.
Show that you can apply the rule under pressure during an incident and afterwards choose the right repair - better wrapping or typed errors - rather than adding records back at every layer.
Own the convention: its written exceptions, its structural enforcement, the higher bar for shared libraries, and the tradeoff between one readable record and what on-call and the log bill each demand.
## The decision, and who can overrule it "Log at the top, return everywhere else" sounds like a style rule, but it is really an allocation of authority: it decides who is allowed to declare a failure reported, and therefore what the people reading the logs will see. Two constituencies can legitimately push back. The on-call engineer who cannot diagnose an incident from the single record wants more logging lower down. The person paying for log storage, or tuning alerts that count error records, wants less. A lead owns the balance, and should expect to defend it. ## The rule State it in one line so it can be repeated in review: **a function that returns an error does not report it.** Reporting is done once, at the layer where the error stops travelling: - a command or job binary: `main` receives the error from `run() error`, writes one record, exits non-zero - and because `run` returned normally, its deferred cleanup already finished; - a server: the boundary that converts the failure into a response; - a background worker: the top of the goroutine's loop, which has no caller above it. Everything below wraps with `fmt.Errorf("operation: %w", err)`, adding the operation and the identifiers only that layer knows. ## Write the exceptions down A rule with unwritten exceptions is a rule people quietly stop following. Mine are the cases where nothing is returned, so the log call *is* the single report: - a goroutine started with no one waiting on its result; - best-effort work whose failure does not change the outcome; - a retry loop recording the attempts that failed before one succeeded; - an error deliberately discarded in a deferred call. And one hard prohibition that is not negotiable: **no `log.Fatal` outside `main`.** It is `os.Exit(1)` in disguise, it runs no deferred cleanup, and a helper that calls it has decided for every caller, present and future, that the failure is fatal. ## Enforcement that survives turnover Rank the mechanisms by how much human attention they need: 1. **Structural** - confine the logger to the top-level packages. A lower package that wants to log must add an import, and an automated import check can fail the build on it. This is the only mechanism that keeps working when the people who agreed to the rule have moved on. 2. **Automated review signal** - a check for `log.Fatal` and `os.Exit` outside the main package. Cheap, unambiguous, and it catches the expensive failure mode rather than the merely noisy one. 3. **Review checklist** - useful for the judgment calls the tools cannot make, such as whether a wrap message actually says something. 4. **Documentation alone** - the weakest, and worth writing only because it explains the *why* the tools cannot. ## For a library other teams import The bar is higher, and it is worth saying explicitly in the package's own docs: a library that logs is imposing its output on every importer, including one whose process writes structured records to a different destination and one running inside a test. So a shared library returns errors and logs nothing. Where a library genuinely has diagnostics to emit - a retry it recovered from, a cache it rebuilt - the way to expose them is a hook the importer supplies, not a logger the library chose. That is an API decision you cannot walk back once a dozen teams depend on it. ## When on-call pushes back The complaint - "I got one line and it did not tell me enough" - is usually correct as an observation and wrong as a diagnosis. The fixes, in order of preference: wrap at more layers so the chain names each operation; carry structured detail in typed errors the report site unpacks with `errors.As` and attaches to its record; and only then consider whether some part of the system genuinely handles failures internally and therefore deserves its own report site. Reopening the rule to allow logging on the return path costs you the property that makes the log readable - one incident, one record. ## When to bend the rule deliberately There are honest cases for temporarily allowing more. A subsystem being migrated, where the errors it returns are still poorly worded and an incident is more expensive than the noise. A path so rare that the extra records cost nothing and the confidence is worth it. Handle these as time-boxed exceptions with an owner rather than as amendments to the rule, because the failure mode of a convention is not one violation - it is the second team that cites the first as precedent.
- A team argues their shared library must log because callers keep ignoring its errors. How do you respond?That is a caller bug the library cannot fix by logging, and logging imposes the library's output on every importer, including tests. Return the error, make the message good, and if the library really has diagnostics to emit, expose a hook the importer supplies rather than choosing a logger for everyone.
- On-call says the single record is not enough to diagnose an incident. What do you change first?Wrapping, before anything else - each layer should name the operation and its identifiers so the chain reads as a cause. Next, typed errors carrying structured detail the report site unpacks with `errors.As`. Reopening the rule to permit logging on the return path is the last resort, because it costs the one-incident-one-record property.
- Which single enforcement mechanism would you keep if you could only have one?The structural one: keep the logger out of the lower packages so a violation is an import an automated check can fail on. Checklists and documentation depend on continuous human attention and decay with turnover; an import boundary keeps holding after everyone who agreed to it has left.
- How do you handle a subsystem where the convention is genuinely painful right now?Grant a time-boxed exception with a named owner and a reason recorded, rather than amending the rule. The real risk to a convention is not one violation but the second team citing the first as precedent, and an explicit temporary exception is much easier to withdraw than a quiet one.
saying these in an interview costs you the question
- Let each team decide, it is only style
- Documenting the convention is enough enforcement
- A shared library should log so callers cannot ignore failures
- Allow log.Fatal below main when the error is clearly unrecoverable
- Answer on-call complaints by permitting logs on the return path