skip to content

How do you use the perf-test tools to build a throughput-vs-latency curve, and how do you interpret its 'knee'?

level: seniorimportance: must knowfreq 40%

answer

  1. sweep fixed --throughput, never -1, for latency
  2. x=offered throughput, y=p99 latency
  3. flat → knee (elbow) → saturation
  4. operate below the knee for headroom
  5. watch coordinated omission: delivered≈offered?

basics

~20 s

Run kafka-producer-perf-test.sh repeatedly at increasing fixed --throughput values and record p99 latency each time. Plot offered rate (x) against latency (y). Latency stays flat then sharply rises at the 'knee' — the point where you hit saturation.

solid answer

~50 s

You measure latency under controlled load by holding --throughput at a series of fixed records/sec values (not -1) and recording p99/p999 latency at each step, keeping record-size and producer config constant. Plotting offered throughput on the x-axis and latency on the y-axis gives the classic curve: latency is roughly flat while the system is under-utilized, then bends upward sharply at the 'knee' as queues (producer buffer.memory, broker request queue, replication) start filling. Past the knee, increasing offered load yields little extra delivered throughput but exploding latency — that's saturation. You operate below the knee for headroom. The maximum sustainable throughput is found separately with --throughput -1, but that flat-out number trades away latency, so the knee, not the peak, is your real capacity target. Watch for coordinated-omission style underreporting when the throttler can't keep up.

go deeper

for a junior

Understand that pushing more load eventually makes latency spike; that's the basic trade-off.

for a middle

Run multiple fixed-throughput sweeps and plot offered rate vs p99 latency.

for a senior

Locate and interpret the knee, target below it, and explain queueing causes of the inflection.

for a principal

Compare whole curves across configs (acks/linger/batch), account for coordinated omission, and turn the knee into capacity-planning policy with headroom.

## The fundamental tension In any queueing system, throughput and latency trade off. At low load, a request is served almost immediately. As offered load approaches the system's service capacity, requests start waiting in queues, so latency climbs. The relationship is non-linear: it follows roughly a hockey-stick shape predicted by queueing theory (latency ∝ 1/(1−utilization)). ## Building the curve with the tools 1. **Fix everything except offered load**: same `--record-size`, same `--producer-props` (acks, batch.size, linger.ms, compression.type), same topic/partition/replication layout. 2. **Sweep `--throughput`**: run the producer perf test at, say, 50k, 100k, 150k, 200k, 250k records/sec — each a separate run with a *positive* `--throughput`, never `-1`. 3. **Record latency at each point**: capture the reported avg/p50/p95/p99 (and ideally p999) latency, plus the *delivered* records/sec to confirm the producer actually hit the offered rate. 4. **Plot**: x = offered (delivered) throughput, y = a chosen latency percentile. ## Reading the knee - **Flat region**: at low utilization latency is dominated by fixed costs (network round-trip, linger.ms batching wait) and barely moves as load rises. - **The knee/elbow**: the inflection where latency starts rising steeply. It marks the onset of saturation — queues (producer accumulator, broker network/request handler queues, the replication pipeline) no longer drain as fast as they fill. - **Past the knee**: offered load can't be delivered; delivered throughput plateaus while latency (and timeouts/retries) explode. The producer's `buffer.memory` fills and sends block. ## Why the knee matters more than the peak The absolute max throughput (from `--throughput -1`) sits at or beyond the knee with terrible tail latency. Production systems should run **below the knee** to keep headroom for spikes, GC pauses, partition leader moves, and rebalances. So capacity planning targets the knee, not the flat-out peak. ## Measurement pitfalls - **Coordinated omission**: if the load generator is itself stalled waiting on slow responses, it under-issues requests and under-reports tail latency. Watch that delivered rate ≈ offered rate; if it falls short, you're past saturation and the latency numbers are optimistic. - **Tuning shifts the curve**: raising `linger.ms`/`batch.size` raises baseline latency but pushes the knee right (more throughput); `acks=all` lowers the knee (replication cost) but improves durability. So you compare *curves*, not single points. - **Tail vs average**: always reason on p99/p999 — averages hide the knee.

  • Why should you sweep fixed --throughput values rather than just running --throughput -1?
    -1 runs flat-out at the saturation point, giving only the max-throughput-with-worst-latency data point. To map latency across the operating range you need controlled offered loads below saturation, which only positive --throughput values provide.
  • How would raising linger.ms and batch.size change the curve?
    Larger batching raises the baseline (flat-region) latency because records wait longer to fill a batch, but it increases efficiency and shifts the knee to a higher throughput, raising overall capacity. It is a latency-for-throughput trade.
  • What is coordinated omission and how does it corrupt these results?
    When the generator blocks waiting on slow responses it stops issuing new requests on schedule, so the requests that would have seen the worst latency never get measured, understating the tail. Detect it by checking delivered rate matches the offered --throughput.

saying these in an interview costs you the question

  • Using --throughput -1 to build a latency curve — it only gives the saturation point.
  • Reporting average latency instead of p99/p999 to find the knee.
  • Treating the flat-out peak as the production capacity target instead of the knee.
  • Ignoring that delivered < offered means you're past saturation and latency is understated.

context