How do concurrency settings and prefetch (prefetchCount) affect throughput, ordering, and fairness in a @RabbitListener? How do they interact?
answer
- concurrency = parallel consumer threads
- prefetch = unacked messages broker pushes (basic.qos)
- in-flight ≈ consumers × prefetch
- high prefetch = throughput but unfair + more redelivery
- prefetch 1 = fair dispatch; ordering only single-consumer
basics
~20 sConcurrency sets how many consumer threads process in parallel; prefetch (basic.qos) sets how many unacked messages the broker sends each consumer before waiting for acks. More concurrency + higher prefetch = higher throughput but weaker ordering and worse load fairness; low prefetch spreads work evenly.
solid answer
~40 sTwo knobs control consumption. Concurrency — on SMLC concurrentConsumers/maxConcurrentConsumers (thread pool of consumers), or on the annotation via concurrency = "min-max"; on DMLC consumersPerQueue plus the shared executor. Prefetch (prefetchCount, the AMQP basic.qos) caps how many unacknowledged messages the broker will push to each consumer at once. High prefetch keeps consumers busy (throughput) but a slow consumer can hoard a batch, hurting fairness, and increases redelivery risk on crash. Low prefetch (e.g. 1) gives near-fair round-robin at the cost of round-trip latency. They multiply: effective in-flight messages ≈ consumers × prefetch. Ordering is only preserved within a single consumer/queue; adding concurrency parallelizes and breaks global order. Tune prefetch to roughly cover processing time and set concurrency to your CPU/IO profile; use MANUAL ack or txSize carefully as they change how prefetch is replenished.
code
java · 13 lines@Bean
SimpleRabbitListenerContainerFactory rabbitListenerContainerFactory(ConnectionFactory cf) {
var f = new SimpleRabbitListenerContainerFactory();
f.setConnectionFactory(cf);
f.setConcurrentConsumers(3); // start with 3 consumer threads
f.setMaxConcurrentConsumers(8); // auto-scale up to 8 under load
f.setPrefetchCount(20); // basic.qos: <=20 unacked per consumer
return f;
}
// Or per-listener, overriding the factory's concurrency:
// @RabbitListener(queues = "work.q", concurrency = "3-8")
// public void handle(Task t) { ... }go deeper
Know prefetch limits how many messages are in flight and concurrency adds parallel consumers.
Explain the throughput-vs-fairness tradeoff and prefetch=1 fair dispatch.
Reason about in-flight ≈ consumers×prefetch, ordering loss, and ack interplay.
Design tuning strategy across memory, redelivery risk, ordering guarantees, and per-key queue partitioning.
## The two knobs ### 1. Concurrency — how many things run in parallel - **SimpleMessageListenerContainer**: `concurrentConsumers` (initial) and `maxConcurrentConsumers` (auto-scale ceiling). Each consumer is a thread with its own channel. On the annotation you can write `@RabbitListener(concurrency = "3-8")` meaning min 3, max 8. - **DirectMessageListenerContainer**: `consumersPerQueue` sets how many consumers subscribe per queue; actual parallel *processing* is bounded by the shared `TaskExecutor` size. ### 2. Prefetch — how many unacked messages the broker pushes `prefetchCount` maps to AMQP **basic.qos**. It's the maximum number of **delivered-but-not-yet-acknowledged** messages the broker will send to a consumer before it must receive acks. In Spring AMQP set it via the container/factory `prefetchCount` (Boot: `spring.rabbitmq.listener.simple.prefetch` / `...direct.prefetch`). Historically SMLC's default related to `txSize`; modern Spring uses a sensible default (250) unless you change it. ## How they interact **Effective in-flight work ≈ number of consumers × prefetchCount.** Two consumers with prefetch 10 means up to 20 unacked messages out at once. This total governs memory pressure on the consumer and how many messages are lost-to-redelivery if the app crashes (all unacked get requeued). ## Throughput vs. fairness vs. ordering - **Throughput**: higher prefetch avoids the consumer sitting idle waiting for the next message (no per-message round trip). Too low (like 1) and each ack→next-message round trip caps throughput. - **Fairness / load balancing**: with high prefetch, RabbitMQ front-loads a batch to whichever consumer is ready; a **slow consumer can grab a big batch and starve** faster ones. Setting prefetch low (e.g. 1) gives near round-robin, fair dispatch — the classic 'fair dispatch' pattern for uneven work. - **Ordering**: RabbitMQ preserves order **only within one queue delivered to one consumer**. The moment you add concurrency (multiple consumers on the same queue) or requeue-on-failure, strict ordering is gone. If you need per-key ordering, route to per-key queues with a single consumer, or use one consumer + prefetch tuning. ## Prefetch and acknowledgment Prefetch is about **unacknowledged** messages, so it's tightly coupled to ack mode. In AUTO ack mode Spring acks after the method returns, freeing a prefetch slot. In MANUAL mode the slot isn't freed until *you* call `basicAck`; forget to ack and the consumer's window fills, then the broker stops delivering to it (a common 'consumer stops receiving' bug). With `txSize`/batch on SMLC, acks happen per batch, which changes when slots free up. ## Practical tuning - Start with prefetch ≈ how many messages you can process during one network round trip; a common range is small (1–50) for slow/uneven work, larger for fast, uniform work. - Set concurrency to match your workload (CPU-bound: near core count; IO-bound: higher). - Watch for **memory**: consumers × prefetch × message size is buffered client-side. - For strict ordering, prefer single consumer + prefetch 1 (or per-entity queues). ## Gotchas - Prefetch 0 in raw AMQP means unlimited — Spring doesn't use 0 by default; be careful setting it, as unlimited prefetch can OOM a consumer. - Auto-scaling consumers (SMLC) plus high prefetch can create bursty imbalance. - global vs per-consumer qos: Spring applies prefetch per consumer channel.
- Your slow consumers keep hogging messages while others idle. What single setting fixes fairness?Lower prefetchCount toward 1 ('fair dispatch'). With prefetch 1 the broker only sends a consumer one message at a time and waits for the ack before sending the next, so work spreads evenly instead of one consumer grabbing a large batch.
- Why can a MANUAL-ack listener suddenly stop receiving messages under high prefetch?Prefetch counts unacknowledged messages. If the code forgets to basicAck (or nack), those deliveries stay unacked and fill the prefetch window; once it's full the broker stops delivering to that consumer until acks free slots.
saying these in an interview costs you the question
- Believing higher prefetch always increases throughput with no downside (it hurts fairness and increases redelivery/memory).
- Assuming message ordering is preserved when using multiple concurrent consumers.
- Confusing concurrency (threads) with prefetch (unacked window) — they are different and multiply.