skip to content

Why does logging only err.Error() from http.Client.Do leave a postmortem guessing?

level: seniorimportance: nice to knowfreq 32%

answer

  1. a message is not a measurement
  2. the verdict lives in the type
  3. one whole failure class never errors at all
  4. addresses and syscall wording are not stable
  5. classify at the call site, log fields

basics

~20 s

Because the classification lives in the error's type, not its text. Once flattened to a string you can no longer ask Timeout() or run errors.Is and errors.As - and the most common failure of all, a 500 response, never produces an error to log.

solid answer

~50 s

Calling `err.Error()` throws away everything a postmortem needs. The verdict lives in the type: `errors.As` to `*url.Error` gives you `Timeout()`, `errors.Is` against `context.Canceled` and `context.DeadlineExceeded` separates a cancellation from a deadline, and `errors.As` to `*net.DNSError` or `*net.OpError` separates a name failure from a refused dial. A string supports none of that. Worse, a log that records only non-nil errors misses the single most common production failure, because a `500` from the server leaves `err` nil - the status has to be recorded on its own. And the text itself is unstable: it embeds a resolved address, a port and a platform-specific syscall message, so any dashboard grouping by message fragment quietly breaks across a Go upgrade. Classify at the call site while the type still exists, log the verdict as fields, and wrap with `%w` so callers upstream can classify too.

code

go · 14 lines
go
resp, err := client.Do(req)
if err != nil {
	var urlErr *url.Error
	isTimeout := errors.As(err, &urlErr) && urlErr.Timeout()

	slog.Error("partner call failed",
		"method", req.Method,
		"url", req.URL.String(),
		"timeout", isTimeout,
		"canceled", errors.Is(err, context.Canceled),
		"deadline", errors.Is(err, context.DeadlineExceeded),
		"err", err)
	return fmt.Errorf("reconcile partner orders: %w", err)
}

go deeper

for a junior

Be ready to say that an error's text is for people while its type is for code, and that logging the string alone means nothing later can branch on what kind of failure it was.

for a middle

An interviewer expects you to name the tests you lose - Timeout(), errors.Is against the context sentinels, errors.As to *url.Error or *net.DNSError - and to know that %v wrapping causes the same loss as stringifying.

for a senior

Show the incident angle: the classification fields you emit at the call site, the separate recording of status codes because a 500 produces no error, and why grouping dashboards by message substrings breaks quietly across a Go upgrade.

for a principal

Own the standard: which failure classes every service reports, where the classification happens so it is not reimplemented per call site, and the cardinality rules that keep error labels from turning telemetry into an unbounded pile of series.

## Stringifying is a one-way door `err.Error()` produces a message for a human to read. Everything that made the error queryable - its concrete type, its methods, its wrapped chain - is gone the moment you take that string. And the client's errors are built precisely so that the type carries the diagnosis: - `errors.As(err, &urlErr)` with `var urlErr *url.Error` yields `Op`, `URL`, and `Timeout()`. - `errors.Is(err, context.DeadlineExceeded)` says a deadline you attached to the request expired. - `errors.Is(err, context.Canceled)` says something cancelled the call deliberately. - `errors.Is(err, io.EOF)` says the connection closed before any response arrived. - `errors.As(err, &dnsErr)` with `var dnsErr *net.DNSError` isolates name resolution, and its `IsNotFound` field separates a missing record from a resolver outage. - `errors.As(err, &opErr)` with `var opErr *net.OpError` isolates dial and connection-level failures. That is six distinct incident classes with six different owners and six different responses. Flattened to a string, they are one line each in a log, distinguishable only by eyeballing English text. ## The failure that is not in the log at all There is a second, larger hole. A job that logs on `err != nil` records nothing when the partner answers `500`, because that exchange succeeded and `err` is nil. In a real incident the pattern is bleak: the partner served error pages for four hours, the job decoded them into an empty result and reconciled against nothing, and the error log for the whole window is empty. So the status code must be recorded on its own path, not folded into error handling. If you only ever ask 'was there an error', you have instrumented the least likely failure and ignored the most likely one. ## The text is not a contract A typical message is `Get "https://api.example.com/v1/orders": dial tcp 10.0.0.7:443: connect: connection refused`. Inside it are a resolved IP address, a port, and a syscall message whose exact wording depends on the operating system. None of that is stable. Two consequences follow: - **Dashboards that group by message substring break silently.** They do not error; they just stop matching, and the graph goes flat while the incidents continue. - **Metric labels built from `err.Error()` have unbounded cardinality**, because every distinct address and port produces a new series. The fix for both is the same: label with the classification, a handful of stable values such as `timeout`, `canceled`, `dns`, `refused`, `eof`, `http_5xx`, and leave the raw text in the log message where a human will read it. ## What to record Classify once, at the call site, while the type is still intact, and emit fields: ```go var urlErr *url.Error isTimeout := errors.As(err, &urlErr) && urlErr.Timeout() slog.Error("partner call failed", "method", req.Method, "url", req.URL.String(), "timeout", isTimeout, "canceled", errors.Is(err, context.Canceled), "deadline", errors.Is(err, context.DeadlineExceeded), "err", err) ``` Note that `url.Error` already carries the method and URL, so you are not inventing context - you are surfacing what the error was holding all along. ## Keep the chain alive above the call site The same mistake has a subtler form: wrapping with `%v` instead of `%w`. `fmt.Errorf("reconcile partner orders: %v", err)` produces a brand-new error whose entire content is text, which is exactly the loss described above, just moved one frame up. `%w` keeps the chain, so a caller three levels away can still run `errors.As` for `*url.Error` or `errors.Is` for `context.DeadlineExceeded` and reach its own conclusion. ## The postmortem test A useful bar to hold your client code to: after an incident, can you answer 'how many of these were deadlines, how many were refused dials, how many were 500s, and did the mix change over the four hours' from what was recorded, without re-reading raw text? If not, the instrumentation is wrong, and adding more log lines will not fix it - classifying earlier will.

  • Which classification fields are actually worth recording?
    A small fixed set the postmortem can pivot on: the method and URL, which *url.Error already carries as Op and URL; a timeout boolean from Timeout(); canceled and deadline booleans from errors.Is against context.Canceled and context.DeadlineExceeded; and, separately on the success path, the HTTP status code. Everything else can stay as free text in the message.
  • How does the wrapper preserve the ability to classify further up the stack?
    By wrapping with fmt.Errorf("...: %w", err) rather than %v. %w keeps the chain intact, so a caller several frames up can still run errors.As for *url.Error or errors.Is for context.DeadlineExceeded. %v builds a new error whose only content is text, which is the same loss as logging err.Error() - just moved higher.
  • Does the same reasoning apply to metrics, not only to logs?
    More sharply, yes. A counter labelled with err.Error() has unbounded cardinality, because the text embeds resolved addresses and ports and a new series appears for each. Label with the classification instead - a handful of stable values such as timeout, dns, refused, eof and http_5xx - and leave the raw text in the log line.

It is the difference between a spreadsheet and a screenshot of a spreadsheet. Everything is visible in both; only one of them can be counted, sorted or filtered.

saying these in an interview costs you the question

  • Logs err.Error() and calls that structured logging
  • Alerts only on non-nil errors, missing every 500
  • Wraps with %v and destroys the chain
  • Groups dashboards by raw error message substrings
  • Assumes Go's error text is stable across releases