A Go queue worker retries every non-nil error and is overloading a failing downstream. How do you fix its classification?
answer
- err != nil is not a retry signal
- one classifier, not scattered checks
- three outcomes, not two
- what should an unknown error default to?
- break failures out by class and attempt
basics
~20 sReplace the retry-everything default with one classify function mapping an error to retryable, terminal or abandoned via errors.Is and errors.As, cap attempts with the envelope's delivery count, and count failures by error class and attempt number.
solid answer
~50 sThe bug is using `err != nil` as the retry signal, so malformed payloads and cancelled work are re-attempted alongside genuine transients — and a downstream that is already failing gets each of those multiplied by the retry budget. I put the decision in one place: a `classify(err)` returning retryable, terminal or abandoned, built from `errors.Is` against the sentinels the dependencies export and `errors.As` into a typed cause carrying `Retryable() bool` or a status. Unknown errors default to terminal, so a new failure mode cannot silently become a storm. Terminal goes to the dead-letter queue with the error attached; abandoned goes back for redelivery without consuming an attempt. Then I make it visible: failures counted by error class and attempt number, so at 3am you can tell real traffic from the same hundred messages on attempt seven.
code
go · 13 linesfunc classify(err error) disposition {
var r interface{ Retryable() bool }
switch {
case errors.Is(err, context.Canceled):
return abandoned
case errors.Is(err, ErrBadMessage):
return terminal
case errors.As(err, &r) && r.Retryable():
return retryable
default:
return terminal
}
}go deeper
Know that not every non-nil error deserves another attempt, and that a malformed message will fail identically every time. Be able to say what a dead-letter queue is for.
Explain how retrying everything multiplies load on a dependency that is already failing, and write the classifier that separates retryable from terminal using errors.Is and errors.As rather than message text.
Show the operational judgment: the safe default for unknown errors, the attempt cap that survives a wrong classification, and the two dashboard labels that let you tell a storm from an incident at 3am.
Own the amplification budget across the fleet: what retry factor the downstream can absorb, who is told when the retryable set changes, and how the mitigation lever is exposed so on-call can pull it without a deploy.
## How the storm forms The worker's handler returns an error and the loop does the obvious thing: ```go if err := handle(ctx, msg); err != nil { return msg.Nack() // try again later } ``` This is correct for exactly one class of failure and wrong for every other. When a downstream starts failing, each message now produces N attempts instead of one, so the offered load on the sick dependency rises by the retry factor at precisely the moment it needs less. Messages that can never succeed — a payload that does not unmarshal, a record that was deleted — are re-attempted forever, occupying worker slots that healthy messages need. And if the failure was cancellation from a rolling restart, the worker re-attempts work nobody wants. All three effects compound, which is why a retry-everything worker turns a downstream blip into a self-sustaining outage that survives the original cause. ## The fix: one classifier, three dispositions The first move is to stop making the decision at call sites. Scattered `if strings.Contains(...)` and ad-hoc `errors.Is` checks drift apart; the retry policy should be readable in one function. ```go type disposition int const ( terminal disposition = iota retryable abandoned ) func classify(err error) disposition { var r interface{ Retryable() bool } switch { case errors.Is(err, context.Canceled): return abandoned case errors.Is(err, ErrBadMessage): return terminal case errors.As(err, &r) && r.Retryable(): return retryable default: return terminal } } ``` Three things about this shape matter more than the specific cases. **It reads the unwrapped chain.** `errors.Is` and `errors.As` both walk `Unwrap`, so the classifier keeps working when a handler adds context with `%w`. Any check that inspects only the top-level value breaks the first time someone improves an error message. **The default is terminal, not retryable.** This is the deliberate inversion of the buggy worker. An unrecognised error is one you have not reasoned about; sending it to the dead-letter queue costs you one message and a page, while retrying it costs you an amplification factor against a dependency you do not understand. Dead-lettered messages are recoverable — you can replay them once you have classified the new failure mode. A storm is not recoverable in the same sense. **Abandoned is its own outcome.** Cancellation means the attempt never reached a verdict, so it should neither consume the retry budget nor mark the message bad. Leave it unacknowledged for redelivery. ## The budget you keep even when the classifier is wrong Classification will be wrong sometimes — a new dependency version returns a failure shape you have never seen. So the envelope's delivery count is the backstop: after N deliveries the message is dead-lettered regardless of how it was classified. That converts "retried forever" into "retried N times", which is a bounded, survivable mistake. Keep the cap in the worker, not in each handler, and record the attempt number on the dead-lettered message so the replay has context. The cap also protects against the subtler version of the bug: an error that is genuinely transient in principle but not in practice, such as a downstream that will be down for six hours. Classification says retryable and it is not lying; the budget is what stops you from proving it. ## Making it visible from the 3am chair A single `errors_total` counter cannot answer the question you actually have at 3am, which is *what kind* of failure and *which attempt*. Break the counter out by two labels: the class the classifier assigned (or a short symbol for the matched sentinel), and the attempt number from the envelope. That gives you three readings instantly: - Volume concentrated at attempt 1 with low classes of retryable: a genuine downstream incident, the classifier is behaving. - Volume concentrated at attempts 5-10 with the same class: a storm — the same small set of messages is cycling, and you should stop retrying that class. - A rising count in the terminal bucket after a deploy: a new failure shape hitting the default branch, which is your signal to classify it properly rather than to widen the retryable set. Without the attempt-number dimension, a storm and real traffic look identical on the dashboard, which is exactly why these incidents run long. ## The immediate mitigation Before the code change ships, the lever is the classifier's inputs: drop the retryable set to only the classes you are sure about, or set the attempt cap to 1, so the worker drains rather than amplifies. Then classify the new failure mode, replay the dead-letter queue, and add the class to the retryable set deliberately. ## What an interviewer is listening for That you can name the amplification mechanism, that you centralise the decision and read the unwrapped chain rather than error text, that your unknown-error default is the safe one, and that you know what you would have to see on a dashboard to distinguish a storm from an incident.
- Why should an unrecognised error default to terminal rather than retryable?Because the two mistakes are not symmetric. Wrongly dead-lettering costs you one message, and it is recoverable: you classify the new shape and replay it. Wrongly retrying costs you an amplification factor against a dependency you do not understand, and it can keep an outage alive after its original cause is gone. Default to the recoverable mistake.
- The classification is right but the downstream is down for hours. What still stops the worker?The delivery count in the message envelope. Classification only says an attempt could succeed in principle; the budget bounds how many times you are willing to find out. After N deliveries the message is dead-lettered with the attempt number and error attached, so it can be replayed once the dependency is back.
- What would you put on the dashboard so a storm is distinguishable from a real incident?Failure counts labelled by assigned class and by attempt number. Real traffic concentrates at attempt 1 across many distinct messages; a storm concentrates at high attempt numbers on one class, with a flat count of distinct messages. A single total-errors counter shows the same spike for both, which is why these incidents run long.
saying these in an interview costs you the question
- Treating any non-nil error as retryable
- Retry checks scattered across every call site
- Defaulting unknown errors to retryable
- No cap on delivery attempts
- Classifying on error text so wrapping changes behaviour
- One errors_total counter with no class or attempt labels