When benchmarking Kafka, what does it mean to measure 'sustained MB/s' versus 'records/s', and why can they tell different stories?
answer
- MB/s ≈ records/s × record-size
- per-record cost (CPU/request) vs per-byte cost (net/disk)
- small records → records/s bound; large → MB/s bound
- sustained = held long-run, not a burst peak
- always report both + record size + duration
basics
~20 sMB/s is data-volume throughput (bytes per second); records/s is message-count throughput (messages per second). They differ because record size links them: small records can give high records/s but low MB/s, and vice versa. 'Sustained' means the rate held over a long, steady run, not a peak burst.
solid answer
~50 sMB/s measures byte volume per second; records/s measures message count per second. They are related by record size: MB/s ≈ records/s × record-size. They tell different stories because Kafka has both per-record overheads (request handling, indexing, CPU per message) and per-byte costs (network, disk bandwidth, replication). With tiny records you often hit a records/s ceiling (CPU/request-bound) long before saturating MB/s; with large records you hit the MB/s ceiling (network/disk-bound) at modest records/s. 'Sustained' throughput is the rate the cluster holds continuously over a long run after warm-up — distinct from a momentary peak — accounting for GC pauses, segment rolls, replication, and flush. So you should report both metrics plus the record size, and state that the number is sustained, because a headline 'X MB/s' is meaningless without knowing record size and whether it held steady.
go deeper
Know MB/s is bytes/sec and records/s is messages/sec, linked by record size.
Explain per-record vs per-byte bottlenecks and why small vs large records bind different metrics.
Define sustained vs peak and insist on reporting both metrics plus record size and durability.
Set benchmark reporting standards and map record-size regimes to the subsystem each saturates for capacity planning.
## Two throughput dimensions Kafka throughput can be expressed two ways, and the perf-test tools report both: - **records/s (nMsg.sec)**: how many discrete messages move per second. - **MB/s**: how many megabytes of payload move per second. They are tied together by **record size**: `MB/s ≈ records/s × record-size-in-bytes`. Fix any two and the third is determined. ## Why they tell different stories Kafka spends effort in two ways: 1. **Per-record costs**: each message incurs request processing, offset assignment, index entries, CRC, and CPU for (de)serialization and compression framing. These scale with *message count*. 2. **Per-byte costs**: network transfer, disk write/read bandwidth, page-cache pressure, and replication traffic scale with *data volume*. Consequences: - **Small records** (e.g. 100 bytes): you may saturate the **records/s** ceiling — CPU and request-handling bound — while MB/s is still low. The bottleneck is message rate, so batching (`linger.ms`, `batch.size`) helps enormously by amortizing per-record overhead. - **Large records** (e.g. 100 KB): you saturate the **MB/s** ceiling — network or disk bandwidth bound — at a modest records/s. Here batching helps less; you're bandwidth-limited. So a system that does 1,000,000 records/s of 100-byte messages (~100 MB/s) and one that does 10,000 records/s of 100-KB messages (~1000 MB/s) stress completely different subsystems despite both being 'fast.' ## 'Sustained' vs peak **Sustained** throughput is the rate the cluster maintains continuously over a long, steady-state run — after warm-up and across the periodic costs of JVM GC pauses, log segment rolls, follower replication, and page-cache flushes. A short burst can momentarily exceed the sustainable rate because buffers (producer `buffer.memory`, OS page cache) absorb it; once those fill, the rate falls back to what disk/replication can actually drain. Quoting a peak burst as capacity is a classic mistake. ## Reporting discipline A credible benchmark result always states: (1) records/s, (2) MB/s, (3) the record size, (4) that it is sustained over a stated duration, and (5) the relevant config (acks, compression, partitions, replication). 'We do 500 MB/s' alone is unfalsifiable — at what record size, sustained how long, with what durability?
- Why can a small-record workload be records/s-bound while a large-record one is MB/s-bound?Small records pay heavy per-message overhead (request handling, indexing, CPU) so they hit a message-rate ceiling first; large records pay mostly per-byte (network/disk bandwidth) so they hit a data-volume ceiling first.
- Why is 'we sustain 500 MB/s' an incomplete benchmark claim?It omits record size (which determines records/s and which subsystem is stressed), the run duration / whether it's truly sustained vs a buffered burst, and durability settings like acks and replication that bound real throughput.
saying these in an interview costs you the question
- Quoting MB/s without the record size.
- Reporting a momentary burst peak as sustained capacity.
- Assuming records/s and MB/s always rise together.
- Thinking batching helps large-record (bandwidth-bound) workloads as much as small-record ones.