skip to content

A crawler returns errors.Join over one failure per module and a heap profile shows it retaining hundreds of MB. Why?

level: seniorimportance: nice to knowfreq 26%

answer

  1. the error is holding the evidence
  2. retention grows with the failure rate
  3. worst exactly when everything is failing
  4. keep a few, count the rest
  5. the causes count failures, not attempts

basics

~20 s

A joined error holds a live reference to every cause, and each cause keeps alive whatever it captured — response bodies, buffers, request structs. With one cause per failed module and no cap, the error itself becomes the leak. Cap what you retain and summarise the rest.

solid answer

~50 s

`errors.Join` keeps a slice of every cause you gave it, so the joined value transitively retains everything those errors reference. If a per-module failure wraps the fetched body, a `*http.Response`, or a decoded manifest, none of it can be collected while the error is alive, and a crawl over a hundred thousand modules turns the error into the largest object in the process — a heap profile's `inuse_space` view will show the joined value's slice at the root of the retention. The fix is to shrink the causes and cap the count: wrap with the module path and a short message rather than the payload, keep the first N distinct failures plus a counted summary of the remainder, and drop or log the rest as you go. Also stop inferring the failure count from the joined error: nils are dropped, so count what you appended, or read the causes back through `Unwrap() []error`.

code

go · 8 lines
go
const maxKept = 20

if len(errs) > maxKept {
	extra := fmt.Errorf("and %d more modules failed", len(errs)-maxKept)
	errs = append(errs[:maxKept:maxKept], extra)
}

return errors.Join(errs...)

go deeper

for a junior

Take away the basic rule: an error you keep around keeps alive everything it references, so put an identity and a reason in it, never a payload.

for a middle

Be able to explain the retention chain from the joined value's slice to each cause to whatever that cause captured, and why the growth tracks the failure count rather than the item count.

for a senior

Show the whole loop: read inuse_space in a heap profile, identify the joined error's slice as the retaining root, cap or group the causes, and verify the retained bytes stop scaling with failures.

for a principal

Own the reporting contract for long-running jobs: what a batch's error value is allowed to carry, what belongs in logs and metrics instead, and how a bounded summary still lets callers match the conditions they depend on.

## Why the error is the leak `errors.Join` allocates a slice sized to the non-nil arguments and stores them. Nothing about that is unusual — but an error value in Go is an interface holding a concrete value, and that concrete value may hold anything. `fmt.Errorf("decode %s: %w", path, err)` retains the formatted string and the wrapped cause. A custom error carrying `Body []byte` retains the body. An error that holds a `*http.Response` retains the response and, if the body is not drained and closed, potentially the connection state as well. Each of those is small on its own; multiplied by the number of failed items in a long crawl, and held for the whole crawl because the joined error accumulates until the function returns, they become the process's dominant allocation. The giveaway in a heap profile taken with `inuse_space` is that the retained bytes are not in the fetch path where you would expect them, but rooted at the slice inside the joined error, still reachable from a live local variable in the crawl function. Nothing has leaked in the C sense; the collector is doing exactly what it should, because your program still holds the reference. This is worse than an equivalent leak elsewhere for two reasons. First, it grows with the *failure* rate, so it appears during precisely the incident you are trying to diagnose — the registry is down, every fetch fails, and the crawler that normally holds twelve errors now holds a hundred thousand. Second, it is invisible in review: the accumulate-then-join shape is the recommended idiom, and nothing in it says "bounded". ## Shrinking the causes The first fix is to make each retained error small and self-describing. An error that a human will read needs an identity and a reason, not evidence: - Wrap with the module path and the operation, not the payload: `fmt.Errorf("fetch %s: %w", path, err)`. - Never capture a large `[]byte`, a decoded document, or a response value in a long-lived error. If you need a sample for debugging, truncate it to a bounded prefix at the point of failure. - Make sure the response body is drained and closed on the failure path, so the error does not keep transport resources alive alongside the message. ## Capping the count The second fix is a ceiling. Keep the first N causes — N in the tens, not thousands — and replace the tail with a single counted summary: ```go const maxKept = 20 if len(errs) > maxKept { extra := fmt.Errorf("and %d more modules failed", len(errs)-maxKept) errs = append(errs[:maxKept:maxKept], extra) } return errors.Join(errs...) ``` If the causes cluster — and during an incident they nearly always do — grouping is better than truncating: count occurrences per classified reason and emit one cause per class, `"142 modules: connection refused"`. That preserves the information a reader actually wants while retaining a constant number of values. Everything you drop should still be *observed*: log or count each failure as it happens, so the error is a summary and the log is the record. An error is a return value, not a storage medium. ## The count that lies A second, quieter defect lives in the same code. Because Join drops nil arguments, a joined error's causes count the failures, not the attempts, and there is no way to recover the attempts from the value. Two anti-patterns follow: - **Splitting the message on newlines to count failures.** A cause whose own message contains a newline inflates the count, and a single cause is textually identical to the bare error. Read the causes properly instead, by asserting for `interface{ Unwrap() []error }`. - **Reporting "n failures" from a value you have capped.** Once you truncate or group, the number of causes is no longer the number of failures, so carry the real counters yourself — attempted, succeeded, failed — and let the error carry only the human-readable summary. ## Doing it without changing behaviour If the crawler's contract is that callers can `errors.Is` the result against specific sentinels, capping must preserve that. Two guards: make sure at least one cause of each *distinct* class survives the cap, so no matchable condition disappears; and keep the wrap chain intact on the causes you retain, using `%w` rather than `%v`, so the sentinels remain reachable. Then confirm with a benchmark run with `-benchmem` over a synthetic all-failing input, and a heap profile of the same run, that the retained bytes are now flat in the number of failures rather than linear in it.

  • How do you cap the causes without breaking callers that match against the result with errors.Is?
    Keep at least one cause per distinct failure class rather than the first N arbitrary ones, and retain the `%w` wrapping on everything you keep so the sentinels stay reachable. Then the set of conditions a caller can match is unchanged even though the number of causes is bounded.
  • How would you confirm from a profile that the joined error is what is retaining the memory?
    Take a heap profile and look at inuse_space rather than alloc_space: live bytes rooted at the crawl function's error slice, not at the fetch path, point at retention rather than allocation rate. Re-run after capping and check the retained bytes stop growing with the failure count.
  • Why can you not report the number of failed modules by counting lines in the joined error's text?
    Nil arguments were dropped before the value existed, a single cause's own message may contain newlines, and any cap or grouping you applied changes the count again. Track attempted, succeeded and failed as ordinary counters and let the error carry only the summary.

saying these in an interview costs you the question

  • Blames the garbage collector rather than the live reference
  • Captures response bodies or decoded payloads inside long-lived errors
  • Accumulates one cause per item with no ceiling on the count
  • Counts failures by splitting the joined message on newlines
  • Truncates causes arbitrarily and loses whole failure classes