skip to content

When benchmarking Kafka, what does it mean to measure 'sustained MB/s' versus 'records/s', and why can they tell different stories?

level: middleimportance: should knowfreq 30%

answer

  1. MB/s ≈ records/s × record-size
  2. per-record cost (CPU/request) vs per-byte cost (net/disk)
  3. small records → records/s bound; large → MB/s bound
  4. sustained = held long-run, not a burst peak
  5. always report both + record size + duration

basics

~20 s

MB/s is data-volume throughput (bytes per second); records/s is message-count throughput (messages per second). They differ because record size links them: small records can give high records/s but low MB/s, and vice versa. 'Sustained' means the rate held over a long, steady run, not a peak burst.

solid answer

~50 s

MB/s measures byte volume per second; records/s measures message count per second. They are related by record size: MB/s ≈ records/s × record-size. They tell different stories because Kafka has both per-record overheads (request handling, indexing, CPU per message) and per-byte costs (network, disk bandwidth, replication). With tiny records you often hit a records/s ceiling (CPU/request-bound) long before saturating MB/s; with large records you hit the MB/s ceiling (network/disk-bound) at modest records/s. 'Sustained' throughput is the rate the cluster holds continuously over a long run after warm-up — distinct from a momentary peak — accounting for GC pauses, segment rolls, replication, and flush. So you should report both metrics plus the record size, and state that the number is sustained, because a headline 'X MB/s' is meaningless without knowing record size and whether it held steady.

go deeper

for a junior

Know MB/s is bytes/sec and records/s is messages/sec, linked by record size.

for a middle

Explain per-record vs per-byte bottlenecks and why small vs large records bind different metrics.

for a senior

Define sustained vs peak and insist on reporting both metrics plus record size and durability.

for a principal

Set benchmark reporting standards and map record-size regimes to the subsystem each saturates for capacity planning.

## Two throughput dimensions Kafka throughput can be expressed two ways, and the perf-test tools report both: - **records/s (nMsg.sec)**: how many discrete messages move per second. - **MB/s**: how many megabytes of payload move per second. They are tied together by **record size**: `MB/s ≈ records/s × record-size-in-bytes`. Fix any two and the third is determined. ## Why they tell different stories Kafka spends effort in two ways: 1. **Per-record costs**: each message incurs request processing, offset assignment, index entries, CRC, and CPU for (de)serialization and compression framing. These scale with *message count*. 2. **Per-byte costs**: network transfer, disk write/read bandwidth, page-cache pressure, and replication traffic scale with *data volume*. Consequences: - **Small records** (e.g. 100 bytes): you may saturate the **records/s** ceiling — CPU and request-handling bound — while MB/s is still low. The bottleneck is message rate, so batching (`linger.ms`, `batch.size`) helps enormously by amortizing per-record overhead. - **Large records** (e.g. 100 KB): you saturate the **MB/s** ceiling — network or disk bandwidth bound — at a modest records/s. Here batching helps less; you're bandwidth-limited. So a system that does 1,000,000 records/s of 100-byte messages (~100 MB/s) and one that does 10,000 records/s of 100-KB messages (~1000 MB/s) stress completely different subsystems despite both being 'fast.' ## 'Sustained' vs peak **Sustained** throughput is the rate the cluster maintains continuously over a long, steady-state run — after warm-up and across the periodic costs of JVM GC pauses, log segment rolls, follower replication, and page-cache flushes. A short burst can momentarily exceed the sustainable rate because buffers (producer `buffer.memory`, OS page cache) absorb it; once those fill, the rate falls back to what disk/replication can actually drain. Quoting a peak burst as capacity is a classic mistake. ## Reporting discipline A credible benchmark result always states: (1) records/s, (2) MB/s, (3) the record size, (4) that it is sustained over a stated duration, and (5) the relevant config (acks, compression, partitions, replication). 'We do 500 MB/s' alone is unfalsifiable — at what record size, sustained how long, with what durability?

  • Why can a small-record workload be records/s-bound while a large-record one is MB/s-bound?
    Small records pay heavy per-message overhead (request handling, indexing, CPU) so they hit a message-rate ceiling first; large records pay mostly per-byte (network/disk bandwidth) so they hit a data-volume ceiling first.
  • Why is 'we sustain 500 MB/s' an incomplete benchmark claim?
    It omits record size (which determines records/s and which subsystem is stressed), the run duration / whether it's truly sustained vs a buffered burst, and durability settings like acks and replication that bound real throughput.

saying these in an interview costs you the question

  • Quoting MB/s without the record size.
  • Reporting a momentary burst peak as sustained capacity.
  • Assuming records/s and MB/s always rise together.
  • Thinking batching helps large-record (bandwidth-bound) workloads as much as small-record ones.

context