In a Go job runner, one failure produces five slog.Error lines with five different wordings. How do you fix it?
answer
- count the records, count the layers
- if you return it, do not report it
- each layer contributes a phrase, not a record
- one site where the error stops travelling
- main logs once, then exits non-zero
basics
~20 sDelete the log call from every layer that returns the error, and wrap instead with the operation that layer attempted. One site - where the error stops propagating - logs the whole chain once and exits non-zero.
solid answer
~50 sFive wordings means five layers each reported the same error on its way up. The repair is mechanical: in every function that still returns the error, remove the `slog.Error` call and replace it with `fmt.Errorf("acquiring lock: %w", err)` so the detail that justified the log line lives in the message instead. Then pick the one place that stops propagation - for a job runner that is `main`, which calls a `run() error`, logs once with `slog.Error` and exits non-zero - and make sure the deferred cleanup inside `run` still gets to finish before the exit happens. Afterwards a failure is one record whose message reads as a causal chain, `run: loading config: acquiring lock: permission denied`, and `errors.Is` still matches the original at the top. The exceptions worth keeping are the ones that report nothing upward: a goroutine with no caller, and errors deliberately discarded in a deferred call.
code
go · 16 linesfunc main() {
if err := run(); err != nil {
slog.Error("job run failed", "err", err)
os.Exit(1)
}
}
func run() error {
lock, err := os.Create("/var/run/reindex.lock")
if err != nil {
return fmt.Errorf("acquiring lock: %w", err)
}
defer os.Remove(lock.Name())
defer lock.Close()
return reindex() // returns wrapped errors, logs nothing
}go deeper
Recognise the symptom: many records for one failure means several layers logged on the way up. The lower ones should return instead.
Walk through the mechanical repair - remove the log call, add a wrap that names the operation and the identifiers, keep %w so errors.Is still matches at the top.
Show the whole cleanup including the report site and its exceptions, and explain how structured detail survives as typed errors the top unpacks rather than as extra records.
Frame it as a cost decision: record count tracking call depth inflates alerting and storage, and the fix has to be enforced structurally or it comes back.
## What five wordings tells you Five different messages for one failure is the signature of log-and-return applied at every layer. Each function reported what it knew and then handed the error on, so the number of records is the depth of the call chain rather than the number of things that went wrong. The operational damage is concrete: - the on-call engineer cannot tell whether one job failed or five subsystems did; - searching for any single phrasing finds only one slice of the story; - alert rules that count error records fire proportionally to stack depth; - log volume, and its bill, scales with how deeply nested the code is. ## The repair, layer by layer Work from the bottom up and apply one rule: **if the function still returns the error, it does not log it.** Replace each removed log line with a wrap that carries the same detail: ```go if err != nil { return fmt.Errorf("acquiring lock %s: %w", path, err) } ``` The wrap text should name the operation this layer was attempting and the identifiers only this layer knows - the path, the environment variable name, the record id. `%w` keeps the original error underneath, so `errors.Is(err, os.ErrPermission)` and `errors.As` still work at the top no matter how many wraps were added on the way. A useful discipline for the text itself: it will be concatenated into a chain with colons, so each fragment should be a short lowercase phrase with no trailing punctuation and no "failed to" prefix, because the surrounding sentence already says something failed. ## The single report site Every error needs exactly one place where it stops travelling. For a scheduled job runner the shape is: ```go func main() { if err := run(); err != nil { slog.Error("job run failed", "err", err) os.Exit(1) } } ``` `run` owns the whole job, uses `defer` freely for its lock file and open handles, and returns an error. Because the process only ends after `run` has returned, all of that deferred cleanup actually runs - which is exactly what a `log.Fatal` somewhere in the middle would have destroyed. `main` writes one record and one exit status. Other shapes have a different, equally single site: a request boundary for a server, the top of the loop for a worker goroutine, the supervising function for a fan-out of tasks. What they share is that the error's journey ends there. ## What you lose, and how to keep it Two objections come up, and both have answers. **"The lower layer knew things the message cannot carry."** Anything structured - a record id, a retry count, a duration - can either go into the wrap text or into a typed error that the report site inspects with `errors.As` and turns into attributes on the single record. Nothing forces you to flatten to a string. **"Now I cannot tell where it happened."** The chain names each operation in order, which is usually more useful than a frame list, because it says what the program was trying to do rather than which functions were on the stack. If you genuinely need more, wrap at more layers - wrapping is cheap, a second record is not. ## The exceptions that survive The rule is one report per failure, not zero logs below `main`. Keep a log call where nothing is returned: - a goroutine started without anyone waiting on its result has no caller, so it must report its own failures; - an error discarded on purpose in a deferred call - the log line is the only trace; - best-effort work whose failure does not change the outcome, like a metrics push; - a retry loop recording the attempts that failed before one succeeded, where the caller sees only success. Each of these ends the error's journey, so each is a report site in its own right. ## Making it stick After the cleanup, the cheap enforcement is structural rather than aspirational: keep the logger out of the packages that are supposed to return errors, so importing it is the thing review or an automated import check notices, rather than asking reviewers to spot a log line next to a `return err` every time.
- The lower layers were logging structured attributes. How do those survive the cleanup?Put them in a typed error the layer returns, and let the report site pull them out with `errors.As` and attach them to its single record. Values that read well inline - a path, a variable name - can simply go in the wrap text. Nothing has to be flattened into prose.
- How do you keep the rule from eroding as the codebase grows?Make it structural. Confine the logger to the top-level packages so that a lower package importing it is visible in review or fails an automated import check, rather than relying on reviewers to notice a log call sitting next to a return.
- Why is putting the exit inside a run() error function rather than in main a regression?Because ending the process from inside `run` skips `run`'s own deferred cleanup - the lock file removal, the flush, the temporary directory. Returning the error lets every defer finish first, and only then does `main` write its one record and exit non-zero.
saying these in an interview costs you the question
- Keep all five lines, more logs is more information
- Wrapping with %w captures a stack trace
- Log at each layer because the top message is vague
- Exit from the deepest layer that detected the failure
- Deduplicate the five records in the log pipeline instead