skip to content

What is a consumer group in Spring Cloud Stream, and what changes when you set the `group` property on a binding?

level: juniorimportance: must knowfreq 70%

answer

  1. group = competing consumers, one instance per message
  2. no group = anonymous, non-durable, everyone gets a copy
  3. different groups = pub-sub, each its own copy
  4. named group = Kafka committed offsets = durable
  5. DLQ topic = error.<dest>.<group>

basics

~20 s

A consumer group is a named set of app instances that share the workload: each message is delivered to only one instance in the group (competing consumers). Without a group, every instance gets its own copy.

solid answer

~40 s

Setting `spring.cloud.stream.bindings.<name>.group=myGroup` makes all instances with that group form competing consumers over the destination — each message is processed by exactly one instance, so you scale horizontally without duplicate processing. Instances in different groups each receive their own copy (publish-subscribe between groups). A named group also gives durability: with the Kafka binder it creates a real consumer group with committed offsets, so messages published while the app is down are still delivered on restart. Omitting `group` yields an anonymous, auto-generated group — non-durable, typically starting from the latest offset, and every instance gets every message. For any production consumer you almost always set an explicit `group`. It is also a prerequisite for features like DLQ routing.

code

java · 18 lines
java
// application.yml equivalent, expressed for clarity:
// spring.cloud.stream.bindings.process-in-0.destination=orders
// spring.cloud.stream.bindings.process-in-0.group=order-workers   // <-- durable, competing consumers

import java.util.function.Consumer;
import org.springframework.context.annotation.Bean;
import org.springframework.stereotype.Component;

@Component
public class OrderProcessor {

    // Bound to destination 'orders', group 'order-workers'.
    // Run N replicas: each order is handled by exactly ONE of them.
    @Bean
    public Consumer<String> process() {
        return order -> System.out.println("Handling: " + order);
    }
}

go deeper

for a junior

Know the one-line definition: a group = competing consumers, each message handled once; no group = everyone gets a copy.

for a middle

Explain durability (committed offsets) and cross-group pub-sub, and that DLQ needs a group.

for a senior

Tie group parallelism to partition count and reason about rebalancing and idle instances.

for a principal

Discuss group naming as an operational contract, offset-reset implications of renaming, and standby capacity planning.

**Spring Cloud Stream** is an abstraction over messaging middleware (Kafka, RabbitMQ) where your code is just functional beans — `Supplier`, `Function`, `Consumer` — and the framework binds them to *destinations* (Kafka topics / Rabbit exchanges) via *binders*. A **binding** is one such connection, named like `<function>-in-0` (input) or `-out-0` (output). **Consumer group** = a logical name shared by app instances that should *cooperate* on consuming a destination. Configure it with `spring.cloud.stream.bindings.<bindingName>.group=<groupName>`. Two delivery semantics matter: 1. **Within a group — competing consumers.** Each message goes to exactly one instance in the group. Run 5 replicas of your service in group `orders`, and the load is spread across them; no message is processed twice. This is how you scale throughput and get high availability. 2. **Across groups — publish-subscribe.** Group `orders` and group `analytics` both subscribed to the same destination each receive their *own* independent copy of every message. This lets multiple independent consumers react to the same stream. **Durability.** A named group is *durable*. With the Kafka binder it maps to a native Kafka consumer group that **commits offsets**, so if the app is offline, messages accumulate and are delivered when it comes back. If you omit `group`, Spring Cloud Stream generates an **anonymous group** (a random UUID). Anonymous groups are **non-durable**: every instance gets every message (pub-sub-like), and they typically start reading from the *latest* offset, so anything published while disconnected is missed. Anonymous groups are meant for things like a live dashboard tailing events, not durable work queues. **Why it's foundational.** Many error-handling features depend on a group. The **dead-letter queue** default topic name is `error.<destination>.<group>` — you need a group to route poison messages. Consumer-side retry and offset commit semantics also assume a durable group. **Gotchas.** - Forgetting `group` in production → duplicate processing across replicas and lost messages during downtime. - Under Kafka, the group's parallelism is bounded by the number of **partitions**: more instances than partitions means some instances sit idle. - The group name is part of your operational contract — renaming it creates a *new* group that resets to its start offset. **When to use.** Always set an explicit, stable `group` for durable work-queue semantics; rely on anonymous groups only for ephemeral, tail-the-latest scenarios.

  • You run 6 instances in one group against a Kafka topic with 3 partitions — what happens?
    Kafka assigns each partition to at most one consumer in the group, so only 3 instances get partitions and actively consume; the other 3 stay idle as hot standbys, ready to take over on rebalance. Parallelism is capped by partition count.
  • Why might messages published while your app was down never arrive after restart?
    You likely omitted `group`, so you got an anonymous non-durable group that starts from the latest offset and doesn't commit — nothing published during the gap is retained for you. A named group with committed offsets fixes this.

saying these in an interview costs you the question

  • Thinking every instance in a group receives every message (that's cross-group pub-sub, not within-group)
  • Believing an anonymous (no-group) consumer is durable and replays missed messages
  • Claiming you can scale consumer parallelism beyond the number of Kafka partitions

context