How does kafka-consumer-perf-test.sh measure consumer performance, and how do you interpret its output?
answer
- reads --messages records then stops
- columns: MB.sec, nMsg.sec, fetch.* , rebalance.time.ms
- fetch.* excludes rebalance; overall includes it
- page cache → consume often faster than produce
- --show-detailed-stats exposes warm-up
basics
~20 sIt runs a consumer that reads a fixed number of messages from a topic and reports how much data and how many records it consumed per second. You read its MB/sec and nMsg/sec columns to judge consume throughput.
solid answer
~40 skafka-consumer-perf-test.sh drives a consumer (or consumer group) against a topic and measures read throughput. Key flags: --bootstrap-server, --topic, --messages (how many records to consume before stopping), --threads, --group, --fetch-size, and --timeout. It outputs CSV-style columns: start.time, end.time, data.consumed.in.MB, MB.sec, data.consumed.in.nMsg, nMsg.sec, plus rebalance and fetch timing. The headline numbers are MB.sec (megabytes/second) and nMsg.sec (records/second). To get a clean read you first preload the topic with the producer perf tool, then consume. Interpretation caveats: the run includes consumer-group rebalance time and the cost of fetching from the broker page cache vs disk, and with --show-detailed-stats you see per-interval rates that reveal warm-up. Because consumers usually read from the broker page cache, consume benchmarks can look faster than produce.
go deeper
Know it consumes a set number of messages and reports MB.sec and nMsg.sec.
Read all output columns, preload the topic first, and explain rebalance vs fetch timing.
Reason about page-cache vs disk reads and why consume can outrun produce; isolate steady-state.
Design consume benchmarks that defeat page cache for honest capacity numbers and integrate group/partition parallelism.
## Purpose `kafka-consumer-perf-test.sh` (wrapping `org.apache.kafka.tools.ConsumerPerformance`) measures how fast a consumer can pull records from a topic. Whereas the producer perf tool tests the write path, this tool tests the read/fetch path: broker → network → consumer deserialization. ## Key flags - **`--bootstrap-server`**: broker connection string. - **`--topic`**: topic to read from (it must already contain data — typically you preload it with `kafka-producer-perf-test.sh`). - **`--messages`**: total number of records to consume before the test stops. This bounds the run. - **`--threads`**: number of consumer threads. - **`--group`**: consumer group id (affects rebalancing and offset semantics). - **`--fetch-size`** / **`--consumer.config`**: fetch sizing and arbitrary consumer properties. - **`--show-detailed-stats`** + **`--reporting-interval`**: emit per-interval rows instead of just a final summary, which exposes warm-up and variance. ## Output columns It prints a header then data row(s): `start.time, end.time, data.consumed.in.MB, MB.sec, data.consumed.in.nMsg, nMsg.sec, rebalance.time.ms, fetch.time.ms, fetch.MB.sec, fetch.nMsg.sec`. - **MB.sec / nMsg.sec**: overall throughput including rebalance time — the end-to-end view. - **fetch.MB.sec / fetch.nMsg.sec**: throughput counting *only* the fetching window (excluding rebalance), so it isolates raw fetch speed. - **rebalance.time.ms**: time spent joining the group/assigning partitions before data flowed. ## Interpretation 1. **Page cache effect**: consumers usually read recently produced data straight from the broker OS page cache, so consume MB/s can exceed produce MB/s. To benchmark cold/disk reads you must consume data old enough to have been evicted from cache. 2. **Rebalance overhead**: on a fresh group the first rebalance shows up in `rebalance.time.ms` and drags the overall MB.sec below fetch.MB.sec — that gap is normal startup cost. 3. **Warm-up**: with `--show-detailed-stats` the first interval is often slower (JIT warm-up, cold connections, cache fill); steady-state intervals are the number to quote. 4. **Partitions/threads**: a single consumer thread reads partitions serially; parallelism comes from more threads or group members across partitions. ## Gotchas - Consuming faster than realistic because of page cache leads to overstated capacity planning. - Forgetting it stops at `--messages`, not at end-of-topic. - Comparing consumer numbers to producer numbers directly without noting the page-cache asymmetry.
- Why might consumer-perf MB/sec be higher than producer-perf MB/sec on the same cluster?Because consumers typically read freshly written data from the broker OS page cache (RAM) rather than disk, so the read path avoids disk I/O the write path had to pay for fsync/replication on.
- What is the difference between MB.sec and fetch.MB.sec in the output?MB.sec is end-to-end including consumer-group rebalance time, while fetch.MB.sec counts only the active fetching window, isolating raw fetch throughput from join/assignment overhead.
saying these in an interview costs you the question
- Claiming the tool reads to end-of-topic by default — it stops at --messages.
- Treating cache-served consume numbers as cold-disk capacity.
- Confusing fetch.MB.sec (fetch-only) with MB.sec (includes rebalance).
- Forgetting the topic must be preloaded first.