skip to content

A team asks to convert a high-volume standard SQS queue to FIFO because production shows duplicate and out-of-order messages. How do you evaluate that request?

level: principalimportance: should knowfreq 38%

answer

  1. two complaints, only one needs FIFO
  2. duplicates are answered by idempotency
  3. which swaps produce a wrong state?
  4. type is fixed at creation
  5. order only the slice that needs it

basics

~20 s

Separate the two complaints first: duplicates are inherent to at-least-once delivery and are answered by idempotent consumers, not by FIFO. Only genuine order dependence justifies FIFO, and it costs a new queue, a throughput ceiling, and head-of-line blocking per group.

solid answer

~50 s

I would push back until I know which symptom actually breaks something. Duplicates on a standard queue are expected behaviour, and FIFO does not remove them end to end — a consumer that dies before deleting still triggers redelivery — so if duplicates are the complaint, the fix is idempotent handlers regardless of queue type. Out-of-order is the only symptom FIFO addresses, and I would ask which specific pairs of messages produce a wrong end state when swapped. If that set is narrow, the better answer is often to route only that subset to a FIFO queue keyed on the entity id and leave the bulk stream on standard. If we do go FIFO, I would scope the work honestly: the type cannot be changed in place, so it is a new `.fifo` queue with a cutover, plus a throughput quota, higher per-request pricing, and a poison message now stalling a whole group.

go deeper

for a junior

Know that the queue type is chosen at creation, so switching to FIFO means creating a new queue whose name ends in .fifo and moving producers and consumers over to it.

for a middle

Explain why duplicates and reordering are different problems: at-least-once delivery causes the first and is handled by idempotent consumers, while only genuine order dependence justifies the FIFO throughput tradeoff.

for a senior

Lead the cutover: confirm the peak message rate against the FIFO quota, pick a narrow group key, plan the drain and overlap window, and account for one bad message now blocking its group.

for a principal

Own the framing that ordering is a purchased constraint. Push to isolate the order-sensitive slice onto its own queue so the bulk stream keeps unbounded throughput, and treat the group key as a boundary that is expensive to revisit.

## Split the request into two questions The request bundles two very different complaints, and conflating them is the mistake to catch. **"We see duplicates."** A standard queue guarantees at-least-once delivery; duplicates are the contract, not a fault. More importantly, FIFO does not eliminate them. Its deduplication protects the *send* path within a five-minute window; it does nothing when a consumer processes a message and dies before deleting it, which is the duplicate that actually corrupts state. So a team that moves to FIFO for this reason pays the full cost and still needs idempotent handlers. This complaint is answered by making the effect converge on a repeat, keyed on a business identifier — and once that exists, the duplicates stop mattering. **"We see reordering."** This is the only symptom FIFO genuinely addresses. The question to ask is narrow and concrete: which pairs of messages, if applied in the wrong order, leave the system in a wrong end state? Very often the honest answer is "only events for the same account" or "only create-then-update for the same record", and sometimes it is "none — the handler is order-insensitive and someone just found the log confusing". ## Consider the cheaper designs first Before the queue type changes, three alternatives usually deserve a hearing. *Idempotency plus tolerance.* If handlers are order-insensitive and convergent, neither symptom is a defect. *Version or sequence numbers.* Producers stamp a monotonically increasing version per entity, and the consumer discards anything older than what it has already applied. This gives correct end state under reordering on a standard queue, with no throughput ceiling. It is the pattern to reach for when the domain is last-write-wins. *Split the stream.* Route only order-sensitive events to a FIFO queue keyed on the entity, and keep the high-volume remainder on the standard queue. This is frequently the best outcome: the ordering constraint applies to the traffic that needs it, and the bulk volume never touches the FIFO quota. ## If FIFO is the right answer, scope it honestly Four costs should be on the table before anyone agrees. **It is a migration, not a setting.** Queue type is fixed at creation and a FIFO queue's name must end in `.fifo`, so the queue URL changes. The realistic plan: create the FIFO queue, deploy consumers that read both, switch producers, drain the old queue, then delete it. Note that during the switch, in-flight messages exist in two queues with no ordering between them — for a genuinely order-sensitive stream that may require a brief producer pause rather than a rolling cutover. **Throughput.** A default FIFO queue is quota-limited per API action, unlike the standard queue the team is currently enjoying. "High volume" in the request is a warning sign; you need the current peak rate before agreeing, and you need to know whether high throughput mode plus batching covers it. **Blast radius.** With ordering comes head-of-line blocking: a message that cannot be processed halts every later message in its group. On the standard queue a poison message ties up one consumer; afterwards it freezes a customer's entire stream. The group key decides how much is frozen, so the key must be as narrow as correctness allows. **Feature and cost differences.** FIFO requests are priced higher per million. FIFO queues also do not support per-message delay values — you can set a delay on the queue, but not on an individual send — which occasionally breaks an existing scheduling trick. ## The group key is the real design decision If the migration proceeds, the durable decision is not "FIFO" but the group key, because it is embedded in producer code and re-keying later means another drain window. Choose the narrowest identity for which ordering is required — the entity being mutated — so the queue behaves as thousands of small ordered lanes rather than one pipe. ## How to answer in an interview State the position: ordering is a constraint you buy, and its price is throughput, blast radius and migration cost. Show that you separate the duplicate complaint from the ordering complaint, that idempotency is required either way, and that you would look for the smallest slice of traffic that truly needs order before converting the whole queue.

  • How would you sequence the cutover from the standard queue to the new FIFO queue?
    Create the `.fifo` queue, deploy consumers that poll both, then switch producers and let the old queue drain before retiring it. The gap is that during the overlap the two queues have no ordering between them, so for a strictly ordered stream I would prefer a short producer pause and a full drain over a rolling switch.
  • What alternative gives correct end state under reordering without FIFO?
    A monotonic version or sequence number per entity, stamped by the producer. The consumer applies a message only if its version is newer than what it has already stored and drops anything older. That yields last-write-wins correctness on a standard queue with no throughput ceiling and no head-of-line blocking, and it composes well with idempotent handlers.
  • What would make you refuse the FIFO request outright?
    If the team cannot name a concrete pair of messages whose swapped order produces a wrong result. Without that, they are buying a quota, higher per-request cost and group-level blocking to fix a log that merely looks untidy. I would ask for the failing scenario first and, absent one, invest in idempotency and observability instead.

saying these in an interview costs you the question

  • Moving to FIFO to eliminate duplicate processing
  • Assuming the queue type can be flipped with an attribute
  • Ordering an entire stream when one entity needs it
  • Ignoring the FIFO throughput quota on a high-volume queue
  • Forgetting that one poison message stalls a whole group

context