Design producer configs for two workloads: (a) lowest p99 latency for small synchronous events, and (b) maximum throughput for a high-volume bulk ingest. Justify each setting.
answer
- one curve: fuller batch = throughput up, latency up
- low p99: linger=0, small batch, none/lz4, acks=1, fail-fast
- max tput: linger 20-100, big batch, zstd/lz4, acks=all+idem, big buffer
- pipelining hides acks=all latency
- partitions cap throughput; measure, don't guess
basics
~10 sLow latency: linger.ms=0, small batches, compression none/lz4, acks=1, modest buffer. High throughput: linger.ms=20-100, large batch.size, zstd/lz4 compression, acks=all with pipelining, big buffer.memory. The two sit at opposite ends of the batching/latency curve.
solid answer
~50 sThe batching/latency curve says fuller batches raise throughput but add up to linger.ms of latency, so the two workloads tune opposite directions. (a) Low p99 latency: linger.ms=0 so nothing is held back; smaller batch.size; compression none or lz4 (avoid heavy zstd/gzip CPU on the hot path); acks=1 (or acks=all only if durability demands it, accepting the replication round-trip); modest buffer.memory with low max.block.ms to fail fast rather than queue; keep max.in.flight high so single records still pipeline. (b) Max throughput: linger.ms=20-100 and large batch.size (256KB-1MB) to fill batches; zstd or lz4 to shrink bytes on the wire and disk; acks=all with idempotence and max.in.flight=5 so replication latency is hidden by pipelining; large buffer.memory (128-256MB) to absorb bursts; ensure enough partitions to parallelize. Always validate against real payloads and watch record-queue-time, request-latency, and batch-size-avg, since the optimum is empirical.
go deeper
Recognize the two extremes: linger.ms=0 for latency, big batches for throughput.
Map each knob (linger, batch.size, compression, acks) to its effect on the curve.
Justify combined configs and explain pipelining + idempotence for high-throughput durability.
Drive empirical tuning, account for partitions/broker capacity and tail-latency sources, and set per-workload policy with measurement loops.
## The underlying curve Producer tuning is a single tradeoff curve. As you let batches grow fuller (via `linger.ms` and `batch.size`), you send **fewer, larger requests** → higher throughput and better compression, but each record may wait up to `linger.ms` longer → higher latency. The two workloads below sit at opposite ends. ## (a) Lowest p99 latency — small synchronous events Goal: minimize time from `send()` to broker ack, especially the tail. - **linger.ms=0**: never hold a record back artificially; ship as soon as a sender thread is free. This is the dominant latency lever. - **batch.size**: keep modest (default 16KB is fine); you are not trying to fill batches, and oversized batches just sit waiting under low volume. - **compression.type=none or lz4**: avoid gzip/zstd CPU on the latency-critical path; lz4 is acceptable if bytes matter, but compression always adds CPU time per batch. - **acks**: prefer **acks=1** for lowest latency; use acks=all only when the data cannot be lost, accepting the extra replication round-trip in p99. - **buffer.memory**: modest; pair with a **low max.block.ms** (or fail-fast) so a stalled cluster surfaces as an error fast instead of silently queueing and inflating tail latency. - **max.in.flight**: keep at 5 — even single records benefit from pipelining so the sender never idles. - **Beware**: tail latency is often dominated by GC pauses, leader elections, and acks=all replication outliers — not by linger; monitor request-latency p99 and time-in-buffer. ## (b) Maximum throughput — high-volume bulk ingest Goal: maximize sustained records/bytes per second; per-record latency is secondary. - **linger.ms=20-100**: deliberately wait so batches fill, cutting request count and improving compression. Diminishing returns past ~50-100ms. - **batch.size large (256KB-1MB)**: lets each request carry many records; raise `max.request.size` and broker `message.max.bytes` accordingly if needed. - **compression.type=zstd (or lz4)**: zstd when network/storage is the bottleneck and CPU is available — fewer bytes on the wire, on disk, and in replication; lz4 when CPU is tighter. - **acks=all + enable.idempotence=true + max.in.flight=5**: durability with no-loss, and pipelining hides the replication latency so throughput stays high. Idempotence keeps ordering and avoids duplicates. - **buffer.memory large (128-256MB)**: absorb bursts so `send()` rarely blocks; ensure it exceeds batch.size x active partitions. - **Partitions**: throughput ultimately scales with partition count and broker capacity — config tuning cannot exceed the cluster's parallelism. ## Why measurement is mandatory The optimal point on the curve depends on payload size, compressibility, network RTT, broker load, and partition count. Tune empirically: watch `batch-size-avg` (are batches filling?), `record-queue-time-avg` (backpressure?), `request-latency-avg/p99`, `compression-rate-avg`, and broker request-handler idle ratio. Change one knob at a time. ## Common architectural mistakes - Cranking linger.ms very high (e.g., 1000ms) for throughput — diminishing returns and large latency hit; batches usually fill far sooner. - Using acks=0 'for speed' on data that must not be lost. - Assuming acks=all roughly halves throughput — pipelining makes the real cost small. - Forgetting that too few partitions caps throughput regardless of producer config.
- For the high-throughput config, why not push linger.ms to 1000ms to fill batches even more?Under high volume, batches reach batch.size and ship well before 1000ms, so the extra linger adds latency without filling batches further — diminishing returns. A few tens of ms typically captures nearly all the batching benefit; beyond that you pay latency for almost no throughput gain.
- What single cluster-side factor can make producer tuning irrelevant for throughput?Partition count (and broker capacity). Throughput parallelizes across partitions; if a topic has too few partitions or the brokers are saturated, no producer config can exceed that ceiling. You must scale partitions/brokers first.
- Why might you still choose acks=all in the low-latency profile?If the data cannot tolerate loss (financial events, etc.), durability outweighs the extra p99 from the replication round-trip. You accept a slightly higher tail in exchange for the no-loss guarantee, possibly mitigating with placement to keep ISR replicas close (same AZ).
saying these in an interview costs you the question
- Recommending acks=0 for any data that must not be lost
- Pushing linger.ms into hundreds/thousands of ms expecting linear throughput gains
- Ignoring partition count as the real throughput ceiling
- Putting heavy zstd/gzip on a latency-critical path
- Assuming one config is universally 'best' rather than measuring per workload