skip to content

In Go, should a library's error value declare its own retryability, or should the caller classify it?

level: principalimportance: nice to knowfreq 28%

answer

  1. half of it is fact, half is policy
  2. who pays for a duplicate side effect?
  3. a verdict in a method is a contract
  4. changing it breaks nobody's build
  5. the on-call lever cannot be a release

basics

~20 s

Export facts, not verdicts. A library knows what happened — whether the request was sent, what the peer said — but not the caller's cost of a duplicate effect. Declare retryability only where the library alone can know it.

solid answer

~50 s

Exporting `Retryable() bool` or an `ErrTemporary` sentinel makes retry policy part of your package's compatibility surface: consumers' loops depend on your judgment, and narrowing it later changes their behaviour with no compile error. The library is also the wrong place for the judgment, because retryability is half fact and half policy. The library owns the fact — did the request leave the process, did the peer send a wait hint, was the write never applied — and the caller owns the policy, because only it knows whether a duplicate effect is a rounding error or a double charge. So I export a typed error carrying those facts and let each consumer keep one `classify` function. The exception is where the library is the only informed party, and then I document the guarantee, because someone on call for the downstream will eventually want that classification changed.

code

go · 9 lines
go
type CallError struct {
	Op         string
	Sent       bool          // did the request leave this process?
	RetryAfter time.Duration // if the peer told us to wait
	Err        error
}

func (e *CallError) Error() string { return e.Op + ": " + e.Err.Error() }
func (e *CallError) Unwrap() error { return e.Err }

go deeper

for a junior

Know that the error value is where retryability travels, and that a library and its caller may reasonably disagree about whether a failure is worth another attempt.

for a middle

Be able to explain that a library can observe what happened but not what a repeat would cost, and that a typed error carrying facts leaves the caller room to decide.

for a senior

Show that an exported classification is a behavioural contract: changing it alters consumers' retry loops with no compile error, so it needs a documented promise and a migration path.

for a principal

Own the split explicitly — facts in the library, policy in the service, the lever in configuration — and name who can overrule you: the downstream's on-call for amplification, and the team paying for duplicate side effects.

## Why this is a decision and not a detail Retryability looks like a property of a failure. It is not. It is a joint function of three things: what actually happened, what a repeat would cost, and what the dependency can absorb. The library that produced the error knows only the first. That split is what makes this an ownership question. If the library exports a verdict, the library has taken a decision on behalf of consumers it has never met, in a place they cannot override without wrapping every call. ## What exporting a verdict actually commits you to Once `Retryable() bool` or `ErrTemporary` is in your public surface, consumers build loops on it. Then: - **Narrowing it silently changes their behaviour.** If a later release stops classifying a case as retryable, nothing fails to compile. Their loop simply stops retrying, and they discover it as an availability regression attributed to their own service. - **Widening it is worse.** Adding a case to the retryable set turns into extra load on someone else's dependency, generated by a fleet you do not run, at a time you do not choose. - **You cannot deprecate it cleanly.** There is no compiler warning for a behavioural contract expressed as a method's return value. The standard library has already been through this: an older "temporary" classification on its network error interface was deprecated because the word never acquired an agreed meaning, and implementations set it inconsistently enough that callers could not build on it. The lesson is not "never classify" — it is that an undefined verdict is worse than no verdict at all. ## The line that holds: export facts, own policy A typed error is the better shape because it can carry what the caller needs to decide: ```go type CallError struct { Op string Sent bool // did the request leave this process? RetryAfter time.Duration // if the peer told us Err error } ``` `Sent` is the fact nobody else can supply — the caller cannot see whether the write reached the wire — and it is precisely the fact that decides whether re-running a non-idempotent operation is safe. `RetryAfter` is a peer instruction, not the library's opinion. With these, each consumer writes one `classify` function embodying its own policy: a read-only reporting job may retry aggressively, a payment worker may refuse to retry anything with `Sent == true`. The cases where the library should declare the verdict are the ones where it is the only informed party: a driver that knows a transaction was rolled back with no effects, or a client that knows the peer explicitly signalled overload. Even then, document the promise in terms of effects — "this error guarantees the operation was not applied" — rather than in terms of the caller's action. ## Who can overrule you This is the part that makes it an organisational decision rather than a taste one. The service owner writes the worker's retry policy, but at least two other parties have standing to overrule it. The **on-call engineer for the downstream** owns the load your retries create. When your worker's amplification is what is keeping their incident alive, they need a lever, and "we will ship a release" is not one. That argues for policy living in your service behind configuration — the retryable set and the attempt cap adjustable without a deploy — rather than inside a dependency's method. The **team that pays for duplicate side effects** owns the other direction. If a retry can double-charge or double-notify, they can legitimately say that a class you consider transient must not be retried at all unless the error proves the operation was not applied. Their veto is only actionable if the error carries `Sent`-style facts. So the design follows the escalation path: facts in the library because only it can observe them; policy in the service because that is where both objections land; the lever in configuration because the objections arrive at 3am. ## When the pragmatic answer is different Inside one team's own module boundary, with all consumers in the same repository and deployable together, exporting a retryability method is fine and saves duplication — the compatibility argument is about consumers you cannot redeploy. The moment the package is imported by teams on their own release cadence, the calculus flips, and you should be exporting facts. A middle path that works: export the facts as a typed error, and additionally export a *helper* — `IsTransient(err) error` living in the same package — that consumers may use as a default policy or ignore. The default is available, the override is a one-line change, and you have not made policy into a method signature. ## What an interviewer is listening for That you distinguish fact from policy and can say which side each piece belongs on; that you can name the failure modes of exporting a verdict, including that it changes consumers' behaviour without a compile error; that you name who can overrule the policy and what lever they need; and that you have a defensible exception rather than an absolute rule.

  • Your library already exports Retryable() bool and consumers depend on it. How do you back out?
    Not by removing it, since nothing would fail to compile. Add the facts first — a typed error carrying whether the request was sent and any peer wait hint — and document the method as a default policy rather than a guarantee. Then move consumers onto their own classify functions one at a time, and only narrow the method's behaviour once no loop you know of still relies on it.
  • Which single fact matters most for deciding whether a non-idempotent operation may be retried?
    Whether the request actually left the process and reached the peer. A failure before the send is safe to repeat; a failure after it may mean the operation was applied and only the response was lost. The caller cannot observe this, so a library that hides it forces every consumer to be pessimistic or reckless.
  • Why should the retryable set and attempt cap be configuration rather than constants?
    Because the people who need them changed are on call for a dependency you do not own, at a time when shipping a release is not an option. Amplification from your worker can keep their incident alive, and the only fast mitigation is narrowing what you retry or capping attempts. A constant makes that a deploy; a knob makes it a minute.

saying these in an interview costs you the question

  • Exporting Retryable() with no documented meaning
  • Assuming the library knows the caller's cost of a duplicate
  • Hard-coding the retryable set with no operational lever
  • Treating a behavioural contract as safe to narrow
  • Hiding whether the request was actually sent