skip to content

Your Go RPC transport reports an error on every successful call, logged as <nil>; how do you diagnose it?

level: seniorimportance: should knowfreq 46%

answer

  1. the peer says success, you say failure
  2. an error with no message at all
  3. print more than the value
  4. one verb names what is inside
  5. fix the function that returned it

basics

~20 s

Log the dynamic type alongside the value with %T. If it names a concrete pointer type while %v prints nothing, an interface is holding a nil concrete pointer, so the nil check can never be false. Fix the function that produced the error, not the caller.

solid answer

~50 s

The signature of this incident is a failure branch taken on requests that actually succeeded, with a log line carrying no cause. Add the dynamic type to the log: `log.Printf("send failed: %v (%T)", err, err)`. When `%v` prints `<nil>` but `%T` names something like `*rpc.TransportError`, the error interface is holding a nil concrete pointer - the type half is filled, so `err != nil` is true forever, and `fmt` prints `<nil>` because the Error method panicked on the nil receiver and formatting recovered it. Then walk back to the function whose result type is `error` and find the success path returning a concrete-typed variable. Fix it there: return the literal nil, or hold the value in a variable declared as `error`. Add a test asserting the success path returns a nil error so it cannot regress.

code

go · 4 lines
go
log.Printf("send failed: %v (%T)", err, err)
// %v can print <nil> when Error() panics on a nil receiver and fmt recovers it;
// %T still names the dynamic type, e.g. *rpc.TransportError, proving the
// interface is non-nil while the pointer inside it is nil.

go deeper

for a junior

Know that an error printing as nothing is a signal in itself, and that printing the dynamic type with %T is the next thing to try before changing any code.

for a middle

Be able to explain the mechanism you found: an interface with its type half filled and its value half nil, produced by an implicit conversion at a return statement in the calling chain.

for a senior

Show the whole loop - the contradictory signals, the one-line diagnostic, the producer-side fix, the regression test, and the review rule - and reject the tempting fixes that only remove the symptom.

for a principal

Own the aftermath: what convention goes into the codebase, who enforces it at package boundaries other teams consume, and how much retry amplification the service should tolerate before an unexplained error stops traffic.

### Reading the incident At 3am the signals are contradictory: the downstream service reports normal traffic and success responses, while your service's error rate is pinned at 100% and the retry loop is re-sending work the peer already accepted. The log line says something like `send failed: <nil>`. A failure with no cause, on a call that worked, is the fingerprint of an interface value holding a nil concrete pointer. ### Step one: print the dynamic type The single most useful thing you can add is the `%T` verb next to the `%v`: ```go log.Printf("send failed: %v (%T)", err, err) ``` There are exactly three outcomes and each points somewhere different: - `%T` prints `<nil>` - the error interface is genuinely empty, and the bug is that something else set the failure flag. Look at the control flow, not the error. - `%T` names a concrete type and `%v` prints a real message - this is an ordinary error and you should read the message. - `%T` names a concrete pointer type while `%v` prints `<nil>` - this is the typed-nil case. The interface's type half is filled with, say, `*rpc.TransportError`, and its value half is a nil pointer. `err != nil` is therefore true on every call. `%v` shows `<nil>` because the formatter called the Error method, that method dereferenced the nil receiver and panicked, and formatting recovered the panic and printed `<nil>` for the nil pointer. That one log line separates the three cases in a single request, which matters when you are being paged and cannot reproduce locally. ### Step two: find the producer The caller is innocent - `if err != nil` is the correct idiom and there is no caller-side check that would be better. Walk back to the function whose declared result type is `error` and look at what it returns on the success path. The bug is always the same shape: a variable declared with a concrete type, returned directly. ```go func send(req Request) error { var e *TransportError // nil unless the failure branch runs ... return e // implicit conversion fills the dynamic type half } ``` Also check helpers whose own result type is concrete. A helper may legitimately return `*TransportError` so that code inside the package can read its fields, but if its result is plumbed straight out through an `error` return, the same conversion happens one level up and is even easier to miss in review. ### Step three: fix it at the boundary Three equivalent repairs, in order of preference: 1. Return the untyped `nil` literal on the success path. 2. Declare the accumulating variable as `var err error` so no conversion happens at the return. 3. Where a concrete-returning helper must be adapted, convert explicitly behind a nil check: `if e := roundTrip(req); e != nil { return e }; return nil`. Then write the regression test. It is one line of assertion - the success path must return a nil error - and it is exactly the test nobody wrote, because everyone tests the failure path of an error-returning function and assumes the success path is trivial. ### Fixes to reject in review - **Making the Error method nil-tolerant.** It removes the `<nil>` in the log and leaves the incident in place: the interface stays non-nil, the failure branch still fires, the retries still hammer. Worse, the next person has no symptom to follow. - **Reflection at every call site.** Checking whether the value inside a non-nil error is a nil pointer works, but spreading that over every caller institutionalises the defect and is easy to forget in the one place that matters. - **Suppressing the retry.** Treating an unexplained error as success papers over a real failure mode the next time the transport genuinely breaks. ### Hardening afterwards The durable outcome of this incident is a rule, not a patch. Write it down: a function whose result type is an interface never returns a concrete-typed variable on a success path. Pair it with two habits - log `%T` alongside `%v` for errors crossing a service boundary, and assert nil errors on success paths in tests for anything on the request path. Both cost nothing and turn a class of silent 3am pages into a compile-and-test-time concern. ### What a strong answer sounds like Start from the contradiction between your error rate and the peer's success rate, name the one-line diagnostic, state the mechanism in terms of the two halves of an interface value, fix the producer rather than the caller, and finish with the test and the review rule. Candidates who have lived it usually mention the retry amplification without being asked, because that is what actually wakes people up.

  • Why does the log print <nil> instead of a formatting panic taking down the handler?
    The fmt package calls the Error method to format the value and guards that call. When the method panics and the argument is a nil pointer, formatting recovers and prints `<nil>` rather than propagating. So the process survives and you are left with an error that has no message - which is the diagnostic clue, once you know to read it that way.
  • The team proposes checking with reflection at every call site instead. What do you say?
    It works and it is the wrong place. Every caller now carries a workaround for a defect one function upstream, and the one caller that forgets reintroduces the incident. Fix the producing function so its success path returns an empty interface, keep the callers on the plain nil check, and add a test on the success path so the fix cannot silently regress.
  • What would you add to prevent a repeat rather than just closing the incident?
    Two things. A review rule that a function with an interface result never returns a concrete-typed variable on a success path, with the conversion done explicitly at package boundaries. And a test convention that every error-returning function on the request path has a case asserting the success path returns a nil error, which is the assertion nobody writes by default.

saying these in an interview costs you the question

  • Treats it as a logging bug and changes the format verb only
  • Adds a nil-receiver guard to Error and closes the incident
  • Suppresses the retry instead of finding the cause
  • Blames the downstream service for lying about success
  • Pushes reflection-based nil checks onto every caller