In an SQS FIFO queue, what does MessageGroupId control, and how does the choice of group key affect consumer parallelism?
answer
- ordering scope, not queue scope
- one in flight per group
- cardinality sets your parallelism
- coarse key equals serial queue
- a stuck head blocks its lane
basics
~20 sMessageGroupId is the ordering scope of a FIFO queue: SQS delivers messages sharing a group id strictly in order, one in flight at a time, while different groups proceed independently. The number of distinct group ids therefore sets the maximum consumer parallelism.
solid answer
~50 sEvery send to an SQS FIFO queue carries a `MessageGroupId`, and that value is the unit of ordering. Within one group, SQS delivers messages in the order it accepted them and will not hand out the next one until the current message is deleted or its visibility expires. Across groups there is no ordering relationship, and SQS hands different groups to different consumers concurrently. That makes the group key a capacity decision as much as a correctness one: pick the narrowest identity that actually needs ordering — an account id, an order id, a device id — and you get many independent lanes. Pick something coarse like a constant or a tenant name and you have collapsed the queue into one or a few serial streams, so adding consumers buys you nothing and one slow message stalls everything behind it.
code
bash · 4 linesaws sqs send-message --queue-url "$QUEUE_URL" \
--message-body '{"event":"AddressChanged","customerId":"cust-42"}' \
--message-group-id "cust-42" \
--message-deduplication-id "evt-9f3c1a"go deeper
Know that every send to a FIFO queue must include a MessageGroupId, and that ordering is guaranteed only between messages that share the same group value.
Explain that a group allows one in-flight message at a time, so the count of distinct group ids caps how many messages can be processed concurrently, and show how you would derive the key from an entity id.
Demonstrate that you have felt head-of-line blocking: describe the failure where one bad message freezes a customer's stream, and how key granularity determines whether that stalls one entity or the whole queue.
Own the group key as an architectural boundary that is expensive to change later, since it lives in producer code and a re-key needs a drain window. Frame it as the tradeoff between ordering scope and achievable throughput.
## The group is the ordering unit A FIFO queue does not order the queue. It orders each `MessageGroupId` independently. The parameter is required on every `SendMessage` to a FIFO queue — there is no default — and it is an opaque string you choose. The guarantee within one group is strong and worth stating precisely: SQS delivers the messages of a group in the order it accepted them, and it will not make the next message of that group available to any consumer until the current one is deleted or becomes visible again. That is what makes ordering survive multiple consumers — the queue itself refuses to let two messages of the same group be in flight simultaneously. Across groups there is no guarantee at all. Group `cust-7` and group `cust-8` have no defined relative order, and SQS will happily give them to two different consumers at the same instant. ## Group cardinality is your parallelism ceiling Because a group admits one in-flight message at a time, the number of distinct groups with pending work is the maximum number of messages that can be processed concurrently from that queue. This is the fact interviewers are usually probing. - One constant group id for everything: the queue is a single serial stream. A hundred consumers will not go faster than one, because ninety-nine of them get nothing while the head message is in flight. - A group per tenant, with one whale tenant: that tenant's lane is a bottleneck no matter how many consumers you run, while small tenants stream through. - A group per entity — order id, account id, device id — with thousands of live values: full parallelism, with order preserved exactly where it means something. The design rule is to pick the **narrowest** key that still makes ordering correct. Ask: which pairs of messages would produce a wrong result if swapped? Those pairs must share a group; everything else should not. ```bash aws sqs send-message --queue-url "$QUEUE_URL" \ --message-body '{"event":"AddressChanged","customerId":"cust-42"}' \ --message-group-id "cust-42" \ --message-deduplication-id "evt-9f3c1a" ``` ## Head-of-line blocking The flip side of the in-flight rule is that a message which cannot be processed blocks its group. If a handler fails and the message returns to the queue, it is retried from the head of that group, and every later message for that entity waits. Nothing behind it can pass. This is not a defect; it is the guarantee working. If order matters, skipping the failed message would produce exactly the wrong state. But it changes your failure blast radius: on a standard queue one poison message occupies one consumer slot, while on a FIFO queue it freezes an entire logical stream. Two consequences follow. First, a bad group key concentrates that damage — a whole tenant, or the whole queue, rather than one customer. Second, you need an explicit escape route, so that a persistently failing message eventually leaves the queue instead of blocking its group indefinitely. ## Receiving from groups A single `ReceiveMessage` call on a FIFO queue can return several messages, and when it returns more than one from the same group they come back in order; requesting a batch tends to draw from a single group first before moving on. The practical implication for consumer code is that you must process and delete a group's messages in the order received, not fan a received batch out to a thread pool. Parallelising inside a batch silently destroys the ordering you are paying for. ## Getting the key wrong, and fixing it A group key is embedded in the producer, so changing it is a code deploy, and during the transition messages for the same entity can exist under two keys with no ordering between them. That makes this a decision worth getting right up front. If you discover the key is too coarse, the usual remediation is to move to a finer key and accept a drain window: stop producing under the old key, let the old groups empty, then start producing under the new one. The useful mental model is that a FIFO queue is not one ordered pipe. It is a large, dynamic set of tiny ordered pipes, and the group id decides how many pipes you have.
- What happens to later messages of a group when the message at its head keeps failing?They wait. SQS will not release the next message of a group until the current one is deleted or becomes visible again, so a message that keeps failing is retried from the head and the rest of that group's backlog stalls behind it. That is the ordering guarantee working as designed, which is why a persistently failing message needs a route out of the queue.
- Is it safe to hand a received FIFO batch to a thread pool for speed?No, if the batch contains multiple messages from one group. Those arrive in order precisely so you can apply them in order; processing them concurrently reintroduces the reordering you chose FIFO to avoid. Process each group's messages sequentially and gain concurrency across groups instead, which is where FIFO actually offers it.
- How do you choose the group key for a multi-tenant workload?Go finer than the tenant unless ordering is genuinely tenant-wide. Tenant-level keys make every large customer a serial bottleneck and put all of their traffic behind one bad message. Keying on the entity being mutated — order, account, document — keeps ordering correct where it matters and gives each tenant many parallel lanes.
saying these in an interview costs you the question
- Thinks MessageGroupId affects deduplication rather than ordering
- Uses one constant group id and expects consumer scaling to help
- Assumes messages in different groups keep their relative order
- Believes more consumers can bypass a blocked group
- Fans a received FIFO batch out to worker threads