skip to content

Producers

Everything on the write path: how send() batches records, what acks buys you, idempotence, partitioning, and retry semantics. Interviewers focus here because most data-loss and duplicate stories start with a producer setting.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What do the producer acks settings 0, 1, and all (-1) mean, and how do they trade durability against latency?

level: juniorimportance: must knowfreq 85%

answer

  1. 0 = fire-and-forget
  2. 1 = leader only
  3. all/-1 = full ISR
  4. durability up, latency up
  5. all = committed, 1 = written at leader

basics

~20 s

acks=0 means the producer never waits for acknowledgement (fastest, can lose data). acks=1 waits only for the leader to write the record. acks=all waits for the leader plus all in-sync replicas, giving the strongest durability but the highest latency.

solid answer

~40 s

The producer config `acks` controls how many broker acknowledgements a producer waits for before considering a send successful. `acks=0`: fire-and-forget, the producer doesn't wait at all and assumes success the moment the record leaves the client buffer — lowest latency, no delivery guarantee, records lost on any failure. `acks=1`: the leader replica writes the record to its log and acknowledges immediately, without waiting for followers — moderate latency, but data is lost if the leader fails before a follower replicates it. `acks=all` (alias `acks=-1`): the leader waits until all replicas in the in-sync replica set (ISR) have replicated the record before acknowledging — highest latency, strongest durability. With idempotence enabled (default in modern clients), `acks=all` is the recommended setting for no data loss.

go deeper

for a junior

Memorize the three values and the durability/latency direction: 0 fastest/least safe, all slowest/safest.

for a middle

Explain the leader-crash data-loss window under acks=1 and that all means the ISR, not all replicas.

for a senior

Connect acks=all to committed vs leader-written semantics and the high watermark, and to idempotence requiring acks=all.

for a principal

Frame acks as one knob in a durability budget alongside min.insync.replicas, replication.factor, and unclean.leader.election for an org-wide reliability standard.

## What `acks` is In Apache Kafka, a producer sends records to a **partition**, which is replicated across several **brokers**. One broker is the **leader** for that partition (handles all reads/writes); the others are **followers** that copy the leader's log. The set of replicas currently caught up to the leader is the **in-sync replica set (ISR)**. The producer config `acks` (set on `ProducerConfig.ACKS_CONFIG`) decides how many acknowledgements the producer waits for before its `send()` future completes successfully. ### `acks=0` — fire and forget The producer does **not** wait for any response from the broker. It considers the record sent as soon as it is written to the client's socket buffer. Lowest latency and highest throughput, but **any** failure (network drop, broker down, leader change) silently loses the record. Retries are meaningless because the producer never learns of failures. Use only for tolerable-loss telemetry/metrics. ### `acks=1` — leader acknowledgement The **leader** writes the record to its local log and immediately acknowledges, **without** waiting for followers. If the leader crashes after acknowledging but before a follower replicated the record, that record is **lost** — a classic window of data loss. Good latency, partial durability. ### `acks=all` (a.k.a. `acks=-1`) — full ISR acknowledgement The leader waits until **every replica in the ISR** has fetched and appended the record before acknowledging. Combined with `min.insync.replicas` this gives a tunable durability floor (see follow-ups). This is the only setting that guarantees no acknowledged record is lost as long as at least one ISR member survives. It costs an extra replication round-trip of latency. ### Important nuance: committed vs written A record acknowledged under `acks=all` is **committed** — visible to consumers and durable across the ISR. Under `acks=1` it is merely **written at the leader** and only becomes committed once followers catch up. Consumers only ever read up to the **high watermark** (the offset replicated to all ISR members), so they never see uncommitted records regardless of `acks`. ### Edge cases - With idempotent producers (`enable.idempotence=true`, default since 3.0), Kafka **requires** `acks=all`; setting `acks=0/1` with idempotence throws a config error. - `acks=all` does **not** mean 'all replicas' — it means all **in-sync** replicas. If the ISR has shrunk to just the leader, `acks=all` behaves like `acks=1` unless `min.insync.replicas` forbids it.

  • Why is acks=1 still vulnerable to data loss?
    The leader acknowledges before followers replicate. If the leader crashes during that window and a follower without the record becomes the new leader, the record is permanently lost.
  • Does acks=all wait for all replicas or all in-sync replicas?
    All in-sync replicas (the ISR), not all assigned replicas. A replica that has fallen behind is removed from the ISR and isn't waited on.

saying these in an interview costs you the question

  • Saying acks=all waits for every assigned replica regardless of ISR membership.
  • Claiming acks=0 supports retries or delivery guarantees.
  • Confusing acks=2 as a valid value — only 0, 1, and all/-1 exist.
  • Saying acks controls how many partitions are written, not replicas.

context

open as a page

What do batch.size and linger.ms control in a Kafka producer, and how do they trade latency for throughput?

level: juniorimportance: must knowfreq 78%

basics

~20 s

batch.size caps how many bytes a producer collects per partition before sending; linger.ms tells it to wait up to that many milliseconds for more records to fill a batch. Bigger batches and more lingering raise throughput but add latency.

open as a page

What is the difference between a retriable and a fatal (non-retriable) exception in the Kafka producer, and how does each affect a send?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Retriable errors (like a leader change or timeout) are transient, so the producer automatically resends. Fatal errors (like an unknown topic, message too large, or auth failure) cannot be fixed by resending, so the send fails immediately.

open as a page

What is Kafka's idempotent producer, and how do you turn it on?

level: juniorimportance: must knowfreq 78%

basics

~10 s

An idempotent producer guarantees that producer retries don't create duplicate messages on a partition. You enable it by setting enable.idempotence=true (the default since Kafka 3.0).

open as a page

How does a Kafka producer decide which partition a record goes to when the record has a non-null key?

level: juniorimportance: must knowfreq 78%

basics

~20 s

For a keyed record, Kafka hashes the key and maps it to a partition. The same key always lands on the same partition (as long as the partition count is unchanged), which keeps records with that key in order.

open as a page

Walk through what happens when you call KafkaProducer.send(record). What are the stages between the call returning and the record actually reaching the broker?

level: juniorimportance: must knowfreq 78%

basics

~20 s

send() is asynchronous. The producer serializes the key/value, picks a partition, and appends the record to an in-memory buffer (RecordAccumulator). A background Sender thread later batches and transmits it. send() returns a Future immediately, before the broker has the record.

open as a page

What is a Kafka producer Serializer, and how do you configure the key and value serializers?

level: juniorimportance: must knowfreq 80%

basics

~10 s

A Serializer turns your key/value objects into the byte arrays Kafka stores. You set them with key.serializer and value.serializer producer configs, e.g. StringSerializer for text.

open as a page

How do acks=all, min.insync.replicas, and replication.factor work together to define a durability guarantee?

level: middleimportance: must knowfreq 80%

basics

~20 s

replication.factor sets how many copies of each partition exist. min.insync.replicas sets how many in-sync copies must acknowledge a write under acks=all. Only acks=all enforces min.insync.replicas; together they define how many failures the system tolerates without losing acknowledged data.

open as a page

How does compression.type work in a Kafka producer, and how do gzip, snappy, lz4, and zstd compare?

level: middleimportance: must knowfreq 70%

basics

~20 s

compression.type tells the producer to compress each batch as a whole using a codec: gzip (best ratio, slow), snappy (fast, modest ratio), lz4 (fast), or zstd (great ratio and good speed). Compression is per-batch, so bigger batches compress better.

open as a page

How does the broker actually detect and drop duplicate batches from an idempotent producer?

level: middleimportance: must knowfreq 70%

basics

~20 s

Each producer gets a producer ID (PID). Each batch carries a per-partition sequence number that increases by one. The broker remembers the last sequence written per (PID, partition) and discards any batch it has already seen.

open as a page

What happens to partition selection when a record has a null key, and why was the sticky partitioner (KIP-480) introduced?

level: middleimportance: must knowfreq 70%

basics

~20 s

With a null key there is nothing to hash, so Kafka picks a partition without a key. The old DefaultPartitioner round-robined per record; KIP-480's sticky partitioner instead fills one partition's batch, then switches, producing fuller batches and lower latency.

open as a page

Describe the RecordAccumulator: how does it organize buffered records, and how do batch.size and linger.ms govern when a batch is sent?

level: middleimportance: must knowfreq 70%

basics

~20 s

The RecordAccumulator is the producer's in-memory buffer. It keeps one deque of ProducerBatches per topic-partition. A record is appended to the partition's current batch. A batch becomes sendable when it fills to batch.size or when linger.ms elapses, whichever comes first.

open as a page

How do max.request.size (producer), message.max.bytes (broker), and replica.fetch.max.bytes interact, and what breaks if they're misaligned?

level: seniorimportance: must knowfreq 55%

basics

~20 s

max.request.size limits how big a single producer request/record can be (default ~1 MB). message.max.bytes is the broker's per-batch ceiling, and max.message.bytes is its per-topic version. If the producer allows bigger messages than the broker accepts, the broker rejects them; replicas also need replica.fetch.max.bytes large enough or replication stalls.

open as a page

Explain delivery.timeout.ms (KIP-91) and how it relates to retries, request.timeout.ms, and linger.ms in bounding the total time a send can take.

level: seniorimportance: must knowfreq 60%

basics

~20 s

delivery.timeout.ms is the single upper bound on the total time from send() to success or failure, covering batching, all retries, and inflight requests. It must be >= linger.ms + request.timeout.ms. When it expires, the record fails with TimeoutException regardless of remaining retries.

open as a page

How can max.in.flight.requests.per.connection cause message reordering with retries, and how do you prevent it while keeping throughput?

level: seniorimportance: must knowfreq 65%

basics

~20 s

If more than one request is in flight per connection and an earlier batch is retried while a later one already succeeded, the retried batch lands after it, reordering messages within a partition. Enabling idempotence (enable.idempotence=true) prevents reordering even with up to 5 in-flight requests.

open as a page

What does idempotence guarantee versus what it does NOT, and when do you need transactions instead?

level: seniorimportance: must knowfreq 64%

basics

~10 s

Idempotence guarantees exactly-once writes to a single partition within one producer session. It does NOT give cross-partition atomicity, cross-session deduplication, or atomic consume-process-produce. Those require transactions (transactional.id, initTransactions, begin/commit).

open as a page

How does buffer.memory and max.block.ms create back-pressure on a producer, and what does the application experience when the buffer fills?

level: seniorimportance: must knowfreq 66%

basics

~20 s

buffer.memory caps the total bytes the accumulator can hold. If the app produces faster than the Sender can drain, the buffer fills. Then send() blocks waiting for free space for up to max.block.ms; if it can't get memory in time, send() throws a TimeoutException.

open as a page

How does the Confluent KafkaAvroSerializer work with Schema Registry, including the wire format and subject compatibility?

level: seniorimportance: must knowfreq 65%

basics

~20 s

KafkaAvroSerializer registers the record's Avro schema in Schema Registry, gets a schema ID, and writes a magic byte + 4-byte schema ID + Avro-encoded payload. The registry enforces compatibility per subject before allowing new schema versions.

open as a page

What do retry.backoff.ms and request.timeout.ms control, and how do they interact during a failed send?

level: middleimportance: should knowfreq 45%

basics

~20 s

request.timeout.ms is how long the producer waits for a single request's ack before giving up on that attempt. retry.backoff.ms is the pause before retrying a failed attempt, so the producer doesn't hammer a struggling broker. Both repeat until delivery.timeout.ms expires.

open as a page

Which producer configs does enable.idempotence require, and what happens if you set a conflicting value?

level: middleimportance: should knowfreq 58%

basics

~10 s

Idempotence requires acks=all, retries>0, and max.in.flight.requests.per.connection<=5. If you explicitly set a conflicting value (e.g., acks=1), the producer fails fast with a ConfigException.

open as a page

What is the difference between DefaultPartitioner and UniformStickyPartitioner, and when would you choose one over the other?

level: middleimportance: should knowfreq 48%

basics

~10 s

Both use sticky batching for null keys. DefaultPartitioner hashes non-null keys with murmur2 to preserve per-key ordering; UniformStickyPartitioner ignores the key entirely and sticks for every record, so it never gives key-based ordering.

open as a page

Compare the two ways to get the result of a send: the returned Future versus a Callback. When is each appropriate, and what are the pitfalls?

level: middleimportance: should knowfreq 58%

basics

~20 s

send() returns a Future<RecordMetadata>; calling future.get() blocks until the broker responds, turning async into sync. Alternatively pass a Callback to send(record, callback) that the producer invokes asynchronously on completion. Use Future.get() for sync confirmation, Callback for non-blocking handling.

open as a page

In what order do interceptors, serializers, and the partitioner execute in the producer send path, and what does each operate on?

level: middleimportance: should knowfreq 30%

basics

~10 s

Order inside send(): interceptor.onSend (sees objects) -> key/value serializers (produce bytes) -> partitioner (uses serialized key) -> accumulator/batching -> network. onAcknowledgement fires later on ack/failure.

open as a page

When would you use StringSerializer vs ByteArraySerializer, and what are the trade-offs of raw byte[] keys/values?

level: middleimportance: should knowfreq 50%

basics

~10 s

StringSerializer encodes text to UTF-8 bytes; ByteArraySerializer passes byte[] through unchanged. Use String for human-readable text/JSON-as-string; use ByteArray when you've already encoded bytes yourself.

open as a page

When is a produced record considered 'committed', and how does that differ from being written at the leader? What role does the high watermark play?

level: seniorimportance: should knowfreq 60%

basics

~20 s

A record is 'written at the leader' once the leader appends it to its log. It becomes 'committed' only once all in-sync replicas have replicated it — at which point the high watermark advances. Consumers only read up to the high watermark, so they never see uncommitted records.

open as a page

Explain the NotEnoughReplicas and NotEnoughReplicasAfterAppend errors: when does each occur, are they retriable, and how should a producer handle them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Both occur with acks=all when the in-sync replica count is below min.insync.replicas. NotEnoughReplicas is raised before the leader appends the record; NotEnoughReplicasAfterAppend is raised after appending but before full replication. Both are retriable, so the producer retries until enough replicas rejoin the ISR.

open as a page

When does a Kafka broker recompress (or decompress) batches instead of storing them as-is, and why does that matter?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Normally a broker stores the producer's compressed batch untouched (zero-copy friendly). But if the topic's compression.type forces a different codec, or older message-format conversion / timestamp validation / offset assignment requires it, the broker must decompress and recompress, costing CPU and breaking the zero-copy path.

open as a page

Why does idempotence require acks=all, and how do durability settings (min.insync.replicas) interact with the exactly-once-per-partition guarantee?

level: seniorimportance: should knowfreq 42%

basics

~20 s

acks=all means a write is acknowledged only after all in-sync replicas store it, so an acked record can't be silently lost on leader failover. Idempotence dedups retries, but only acks=all makes the underlying write durable enough for the guarantee to mean exactly-once.

open as a page

How do you implement a custom Partitioner in Kafka, and what is a legitimate use case for one?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Implement org.apache.kafka.clients.producer.Partitioner (the partition() method returns an int partition), then set partitioner.class to your class. A common reason is routing hot or special keys to dedicated partitions to control skew or isolate VIP traffic.

open as a page

Your keyed topic shows severe partition skew (a few partitions hold most of the load). What are your options, and what is the fundamental tension you must navigate?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Skew comes from uneven key distribution (hot keys), not the hash. Options: salt or compound the key, route hot keys with a custom partitioner, or increase partitions. The tension: spreading a key for balance destroys the per-key ordering the key was meant to guarantee.

open as a page

showing 1–30 of 36