What is kafka-producer-perf-test.sh, and what do its --num-records, --record-size, and --throughput flags control?
answer
- num-records = total count, stops the run
- record-size = bytes per message → drives MB/s
- throughput = records/sec cap; -1 = flat out
- must add --topic + --producer-props bootstrap.servers
- reports records/sec, MB/sec, p99 latency
basics
~10 sIt is Kafka's built-in load generator for producers. --num-records sets how many messages to send, --record-size sets each message's size in bytes, and --throughput caps messages per second (-1 means unthrottled, full speed).
solid answer
~40 skafka-producer-perf-test.sh is a CLI tool shipped with Kafka that generates synthetic producer load and reports throughput and latency. --num-records is the total count of records to publish (the test stops after that many). --record-size is the per-record payload size in bytes, so total data sent ≈ num-records × record-size. --throughput is the target rate in records/second; set it to a positive number to throttle to a fixed rate (useful for latency tests at a controlled load) or to -1 to send as fast as possible (max-throughput tests). You also pass --topic and --producer-props (e.g. bootstrap.servers, acks, batch.size, linger.ms, compression.type). Output reports records/sec, MB/sec, and latency percentiles (avg, p50, p95, p99, max).
go deeper
Know the three flags by name and that throughput is records/sec with -1 meaning unthrottled.
Connect record-size to MB/s and know the required --topic/--producer-props and what the summary reports.
Discuss payload compressibility, parallel producer instances, and choosing throttle vs flat-out per measurement goal.
Frame the tool's role in a repeatable benchmark methodology and its limits as a synthetic single-JVM generator.
## What the tool is `kafka-producer-perf-test.sh` is a command-line benchmarking utility bundled in every Kafka distribution's `bin/` directory (the wrapper around the `org.apache.kafka.tools.ProducerPerformance` class). Its job is to push a controlled stream of synthetic records into a topic and measure how fast the producer can go and how long each send takes, without you having to write any code. ## The three core flags - **`--num-records`**: the total number of records the test will send before stopping. This bounds the run. Bigger values give more statistically stable numbers and let the system reach steady state, but take longer. - **`--record-size`**: the size, in bytes, of each record's value payload. The tool fills records with random bytes of this size. Total bytes pushed ≈ `num-records × record-size`. This matters because Kafka throughput is often network/disk-bandwidth bound, so MB/s depends directly on record size. (Alternatively `--payload-file` feeds real payloads.) - **`--throughput`**: the target send rate in **records per second**. A positive value throttles the producer to that rate using a rate limiter — essential when you want to measure *latency at a fixed offered load*. Setting it to `-1` removes the throttle so the producer runs flat-out, which is how you measure *maximum sustainable throughput*. ## Required companions You must also pass `--topic <name>` and `--producer-props key=value ...` (at minimum `bootstrap.servers`). The producer-props let you vary the knobs that actually move throughput/latency: `acks`, `batch.size`, `linger.ms`, `compression.type`, `buffer.memory`. There is also `--print-metrics` to dump the full client metric set at the end. ## Output The tool prints periodic progress lines and a final summary like: `1000000 records sent, 250000.0 records/sec (23.84 MB/sec), 5.20 ms avg latency, 120.00 ms max latency, ... 99th 45 ms`. Throughput is reported in both **records/sec** and **MB/sec**; latency is reported as average, max, and percentiles (p50/p95/p99). ## Edge cases / gotchas - Random payloads compress poorly, so if you enable `compression.type` your MB/s on the wire won't reflect real (compressible) data — use `--payload-file` with representative data. - A single producer process is single-JVM; one instance may not saturate a large cluster, so you may run several in parallel. - The reported MB/sec is uncompressed payload throughput from the client's perspective.
- How do you compute the total data volume a run will push?Approximately num-records × record-size bytes of payload (e.g. 1,000,000 records × 1000 bytes ≈ 1 GB), before any compression and excluding Kafka record/batch overhead.
- When would you set --throughput to -1 versus a fixed number?Use -1 to find maximum sustainable throughput (unthrottled). Use a fixed positive value to hold a controlled offered load while measuring latency, which is how you build a throughput-vs-latency curve.
saying these in an interview costs you the question
- Saying --throughput is in MB/s — it is records/second.
- Thinking the tool needs custom code; it's a ready CLI in bin/.
- Ignoring that random payloads make compression results meaningless.
- Believing one producer instance always saturates the cluster.