skip to content

In a Competing Consumers setup, a single message keeps causing a processing exception no matter which consumer handles it, and it keeps getting redelivered — starving the queue for other work. How do you design around this 'poison message' problem?

level: seniorimportance: should knowfreq 60%

answer

  1. deterministic failure, not transient
  2. max receive count -> DLQ
  3. protects throughput + prevents head-of-line blocking
  4. preserve, don't drop, for inspection
  5. classify transient vs permanent failures

basics

~20 s

One bad message that always fails can get retried forever and clog things up. The fix is to count how many times it's failed, and after a limit, move it out of the main queue into a separate 'failed' queue so a human can look at it later, instead of retrying forever.

solid answer

~50 s

A poison message is one that a consumer can never successfully process — a malformed payload, a bug triggered only by that input, a reference to data that no longer exists — so every delivery attempt ends in the same exception and the message becomes eligible for redelivery again. Left unchecked, this wastes consumer capacity on an unwinnable retry loop and, on a strict FIFO queue, can block everything behind it. The standard fix is a redrive/maximum-receive-count policy: the broker (or consumer) tracks how many times a message has been delivered, and after N failed attempts routes it to a separate dead-letter queue instead of redelivering it to the main queue. That isolates the queue's healthy throughput from the poison message, while preserving the message (rather than dropping it) for inspection, alerting, and either a manual fix-and-replay or automated triage.

go deeper

for a junior

Should recognize that a message which always fails shouldn't be retried forever, and that there needs to be some limit.

for a middle

Should know the basic mechanism: a max-redelivery-count threshold routes the message to a separate dead-letter queue instead of endless retry.

for a senior

Should design the threshold trade-off (too low vs too high), know how it interacts with ordered/FIFO queues (head-of-line blocking), and set up monitoring/alerting on the DLQ.

for a principal

Should push for consumers to classify transient vs. permanent failures explicitly rather than treating all exceptions identically, and own the operational lifecycle of DLQ triage and redrive as a first-class part of the system's reliability posture.

## What makes a message poison A **poison message** is any message that a consumer will never be able to successfully process, no matter how many times or which specific instance retries it — as opposed to a **transient failure** (a downstream API being briefly unavailable, a database connection blip) where a retry has a real chance of succeeding. The distinguishing property is that the failure is deterministic with respect to the message's content or the code path it triggers. Each of these will fail identically on every attempt, on every instance, forever: - a malformed JSON payload that fails to parse - a reference to a foreign-key ID that was since deleted - an edge case that triggers a genuine bug in the consumer Because the visibility-timeout/redelivery mechanism that makes Competing Consumers resilient to transient failures cannot distinguish "this will succeed if I just try again" from "this can never succeed," it treats every unacknowledged message the same way — it goes back on the queue for redelivery. Without an additional mechanism, a poison message enters an infinite retry loop. ## Why it matters This matters for two concrete reasons. 1. **First, it wastes consumer capacity.** Every redelivery consumes a worker's attention for the duration of the doomed processing attempt (including whatever time it takes to fail — parsing, partial work, the exception itself), capacity that could have gone to genuinely processable messages. Under load, a handful of poison messages cycling through the pool can measurably reduce effective throughput. 2. **Second, and more severely**, on a strict FIFO or single-partition queue where order matters, a poison message sitting at the head of the line can block every message behind it from ever being reached, since the consumer(s) assigned to that partition keep retrying the blocking message instead of moving on — this is sometimes called **head-of-line blocking**, and it turns one bad message into a total outage for that ordered stream, not just a wasted retry. ## The standard fix: a dead-letter queue The standard, broadly adopted fix is a **maximum-receive-count** (or maximum-redelivery-count) policy paired with a **dead-letter queue (DLQ)**. The broker (SQS, RabbitMQ with a policy, Azure Service Bus, etc.) tracks how many times each message has been delivered without being acknowledged. Once that count crosses a configured threshold — say, 5 attempts — instead of making the message visible again on the main queue, the broker automatically moves it to a separate DLQ. This achieves two things simultaneously: - the main queue's healthy throughput is no longer affected by that message (it stops competing for consumer attention and, on ordered queues, stops blocking subsequent messages); - the message itself is preserved rather than silently dropped, sitting in the DLQ for inspection. From there, typical practice is to alert on DLQ depth (a growing DLQ usually signals either a data-quality problem in producers or a real bug in the consumer), inspect the failed payloads to diagnose the root cause, and — once fixed — either manually replay the messages back onto the main queue or build an automated redrive pipeline for the common case. ## Choosing the threshold There's a real trade-off in how the max-receive-count threshold is chosen. | Threshold choice | What it costs you | |---|---| | Set it too low | legitimate transient failures (a downstream dependency that's flaky for a minute) get prematurely dead-lettered before a retry would have actually succeeded, turning a recoverable blip into manual intervention work | | Set it too high | a genuine poison message wastes more retries — and, on an ordered queue, blocks the line longer — before it's finally isolated | A common refinement is to distinguish failure types explicitly in the consumer: catch known-transient exceptions (a timeout, a 503) and let those retry normally up to the threshold, but catch known-permanent exceptions (a deserialization failure, a validation error on the payload itself) and route straight to the DLQ (or a "rejected" queue) on the first attempt, since retrying a guaranteed failure is pure waste — this requires the consumer to actively classify failures rather than treating every exception identically. ## Where it shows up A concrete, well-documented real-world instance: AWS SQS's redrive policy lets you configure a maximum receive count on a source queue together with a target DLQ, and this is the officially recommended pattern in AWS's own architecture guidance specifically to prevent poison messages from looping indefinitely; monitoring alarms on DLQ message count are the standard operational signal that something in the pipeline needs human attention. RabbitMQ achieves the equivalent via a queue's dead-letter-exchange policy combined with an explicit negative-acknowledgment (reject without requeue), routing repeatedly-rejected messages to a designated dead-letter exchange rather than looping them forever.

  • How is a poison message different from a message that fails due to a temporary downstream outage?
    A poison message fails deterministically because of something intrinsic to the message or a permanent bug — retrying it, no matter how many times or on what instance, produces the same failure. A transient failure (the downstream API is down for two minutes) will actually succeed on retry once the underlying condition clears, so redelivery is the correct, working recovery mechanism for it, not a symptom of a problem.
  • What operational practices should surround a dead-letter queue once it's in place?
    At minimum, an alert on DLQ depth or age (a DLQ that's silently growing is a live incident, not a passive archive), a defined process for triaging messages that land there (is it a bad producer, a consumer bug, bad data), and either a manual or automated redrive path once the root cause is fixed — a DLQ with no one watching it just becomes a silent data-loss sink with extra steps.
  • Could you avoid poison messages entirely by validating payloads before enqueuing them?
    Producer-side validation catches malformed payloads before they ever hit the queue, which helps, but it can't catch every case — a message might be perfectly valid when produced but become unprocessable later (the entity it references gets deleted), or trigger a consumer-side bug that validation can't anticipate. So a DLQ/max-receive-count safety net is still needed even with strong producer-side validation.

Like a jammed envelope stuck in a mail-sorting machine that keeps getting fed back through, jamming again every time, while every other envelope behind it waits. The fix is a rule: after it jams a few times, an attendant pulls it out to a side bin for manual handling, so the rest of the mail keeps moving.

saying these in an interview costs you the question

  • Treats every failed message as retryable indefinitely with no cap
  • Doesn't distinguish transient failures from deterministic/permanent ones
  • Proposes dropping/deleting the poison message with no record of it
  • Unaware that a poison message can block an entire ordered queue behind it
  • No mention of monitoring or alerting on the dead-letter queue itself

context