skip to content

Benchmarking with Perf Test Tools

Benchmarking with the bundled perf-test tools and reading the throughput-versus-latency result honestly. Interviewers ask how you would prove a configuration change actually helped.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is kafka-producer-perf-test.sh, and what do its --num-records, --record-size, and --throughput flags control?

level: juniorimportance: must knowfreq 62%

answer

  1. num-records = total count, stops the run
  2. record-size = bytes per message → drives MB/s
  3. throughput = records/sec cap; -1 = flat out
  4. must add --topic + --producer-props bootstrap.servers
  5. reports records/sec, MB/sec, p99 latency

basics

~10 s

It is Kafka's built-in load generator for producers. --num-records sets how many messages to send, --record-size sets each message's size in bytes, and --throughput caps messages per second (-1 means unthrottled, full speed).

solid answer

~40 s

kafka-producer-perf-test.sh is a CLI tool shipped with Kafka that generates synthetic producer load and reports throughput and latency. --num-records is the total count of records to publish (the test stops after that many). --record-size is the per-record payload size in bytes, so total data sent ≈ num-records × record-size. --throughput is the target rate in records/second; set it to a positive number to throttle to a fixed rate (useful for latency tests at a controlled load) or to -1 to send as fast as possible (max-throughput tests). You also pass --topic and --producer-props (e.g. bootstrap.servers, acks, batch.size, linger.ms, compression.type). Output reports records/sec, MB/sec, and latency percentiles (avg, p50, p95, p99, max).

go deeper

for a junior

Know the three flags by name and that throughput is records/sec with -1 meaning unthrottled.

for a middle

Connect record-size to MB/s and know the required --topic/--producer-props and what the summary reports.

for a senior

Discuss payload compressibility, parallel producer instances, and choosing throttle vs flat-out per measurement goal.

for a principal

Frame the tool's role in a repeatable benchmark methodology and its limits as a synthetic single-JVM generator.

## What the tool is `kafka-producer-perf-test.sh` is a command-line benchmarking utility bundled in every Kafka distribution's `bin/` directory (the wrapper around the `org.apache.kafka.tools.ProducerPerformance` class). Its job is to push a controlled stream of synthetic records into a topic and measure how fast the producer can go and how long each send takes, without you having to write any code. ## The three core flags - **`--num-records`**: the total number of records the test will send before stopping. This bounds the run. Bigger values give more statistically stable numbers and let the system reach steady state, but take longer. - **`--record-size`**: the size, in bytes, of each record's value payload. The tool fills records with random bytes of this size. Total bytes pushed ≈ `num-records × record-size`. This matters because Kafka throughput is often network/disk-bandwidth bound, so MB/s depends directly on record size. (Alternatively `--payload-file` feeds real payloads.) - **`--throughput`**: the target send rate in **records per second**. A positive value throttles the producer to that rate using a rate limiter — essential when you want to measure *latency at a fixed offered load*. Setting it to `-1` removes the throttle so the producer runs flat-out, which is how you measure *maximum sustainable throughput*. ## Required companions You must also pass `--topic <name>` and `--producer-props key=value ...` (at minimum `bootstrap.servers`). The producer-props let you vary the knobs that actually move throughput/latency: `acks`, `batch.size`, `linger.ms`, `compression.type`, `buffer.memory`. There is also `--print-metrics` to dump the full client metric set at the end. ## Output The tool prints periodic progress lines and a final summary like: `1000000 records sent, 250000.0 records/sec (23.84 MB/sec), 5.20 ms avg latency, 120.00 ms max latency, ... 99th 45 ms`. Throughput is reported in both **records/sec** and **MB/sec**; latency is reported as average, max, and percentiles (p50/p95/p99). ## Edge cases / gotchas - Random payloads compress poorly, so if you enable `compression.type` your MB/s on the wire won't reflect real (compressible) data — use `--payload-file` with representative data. - A single producer process is single-JVM; one instance may not saturate a large cluster, so you may run several in parallel. - The reported MB/sec is uncompressed payload throughput from the client's perspective.

  • How do you compute the total data volume a run will push?
    Approximately num-records × record-size bytes of payload (e.g. 1,000,000 records × 1000 bytes ≈ 1 GB), before any compression and excluding Kafka record/batch overhead.
  • When would you set --throughput to -1 versus a fixed number?
    Use -1 to find maximum sustainable throughput (unthrottled). Use a fixed positive value to hold a controlled offered load while measuring latency, which is how you build a throughput-vs-latency curve.

saying these in an interview costs you the question

  • Saying --throughput is in MB/s — it is records/second.
  • Thinking the tool needs custom code; it's a ready CLI in bin/.
  • Ignoring that random payloads make compression results meaningless.
  • Believing one producer instance always saturates the cluster.

context

open as a page

How do you use the perf-test tools to build a throughput-vs-latency curve, and how do you interpret its 'knee'?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Run kafka-producer-perf-test.sh repeatedly at increasing fixed --throughput values and record p99 latency each time. Plot offered rate (x) against latency (y). Latency stays flat then sharply rises at the 'knee' — the point where you hit saturation.

open as a page

How does kafka-consumer-perf-test.sh measure consumer performance, and how do you interpret its output?

level: middleimportance: should knowfreq 45%

basics

~20 s

It runs a consumer that reads a fixed number of messages from a topic and reports how much data and how many records it consumed per second. You read its MB/sec and nMsg/sec columns to judge consume throughput.

open as a page

When benchmarking Kafka, what does it mean to measure 'sustained MB/s' versus 'records/s', and why can they tell different stories?

level: middleimportance: should knowfreq 30%

basics

~20 s

MB/s is data-volume throughput (bytes per second); records/s is message-count throughput (messages per second). They differ because record size links them: small records can give high records/s but low MB/s, and vice versa. 'Sustained' means the rate held over a long, steady run, not a peak burst.

open as a page

Why must you account for warm-up effects when benchmarking Kafka, and how do you isolate steady-state throughput?

level: middleimportance: should knowfreq 33%

basics

~20 s

Early in a run the JVM JIT, OS page cache, TCP connections, and producer batching haven't stabilized, so the first numbers are misleadingly slow. Run long enough and discard the warm-up interval, quoting only the steady-state rate.

open as a page

What is Trogdor and when would you use it instead of the kafka-*-perf-test.sh scripts?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

Trogdor is Kafka's distributed test framework: a coordinator plus agents that run coordinated workloads and fault injections across many nodes. Use it for large-scale, multi-node, repeatable load and chaos tests; use the perf-test scripts for quick single-box benchmarks.

open as a page