skip to content

Explain the NotEnoughReplicas and NotEnoughReplicasAfterAppend errors: when does each occur, are they retriable, and how should a producer handle them?

level: seniorimportance: should knowfreq 45%

answer

  1. code 19 = before append (rejected)
  2. code 20 = after append (uncommitted)
  3. both retriable
  4. AfterAppend → duplicate risk → need idempotence
  5. only happens with acks=all

basics

~20 s

Both occur with acks=all when the in-sync replica count is below min.insync.replicas. NotEnoughReplicas is raised before the leader appends the record; NotEnoughReplicasAfterAppend is raised after appending but before full replication. Both are retriable, so the producer retries until enough replicas rejoin the ISR.

solid answer

~50 s

These two broker errors enforce `min.insync.replicas` under `acks=all`. **`NotEnoughReplicasException`** (error code 19) is returned when, at the time of the produce request, the current ISR size is **below** `min.insync.replicas` — the leader rejects the write **before** appending it, so the record never enters the log. **`NotEnoughReplicasAfterAppendException`** (error code 20) occurs when the ISR was sufficient at append time but shrank below the minimum **before** the record could be fully replicated/committed — the record is in the leader's log but not committed and won't advance the high watermark. Both are **retriable** errors; the Java producer's default retry logic (and idempotence) will keep retrying until enough replicas rejoin the ISR, at which point the write succeeds. The `AfterAppend` variant is the trickier one because a non-idempotent producer retrying can create duplicates (the record may already be in the log) — which is exactly why idempotent producers are recommended with acks=all.

go deeper

for a junior

Know these errors mean too few in-sync replicas to satisfy acks=all and that the producer will retry.

for a middle

Distinguish the before-append vs after-append timing and that both are retriable.

for a senior

Explain the uncommitted-record state of AfterAppend, the duplicate hazard, and why idempotence makes retries safe; tie to delivery.timeout.ms.

for a principal

Treat frequent NotEnoughReplicas as a fleet-health signal, set retry/timeout policy, and resist the anti-pattern of relaxing min.isr to clear the alert.

## The contract being enforced With `acks=all`, the broker promises that an acknowledged record lives on at least `min.insync.replicas` brokers. To keep that promise it must refuse writes when too few replicas are in-sync. Two distinct errors implement this. ## `NotEnoughReplicasException` (NOT_ENOUGH_REPLICAS, code 19) - **When**: the leader receives a produce request, checks `|ISR| < min.insync.replicas`, and rejects it **before appending** anything. - **State**: the record is **not** in any log. No side effect. - **Cause**: brokers down, network partition, or followers lagging beyond `replica.lag.time.max.ms` so they dropped out of the ISR. ## `NotEnoughReplicasAfterAppendException` (NOT_ENOUGH_REPLICAS_AFTER_APPEND, code 20) - **When**: the ISR was adequate when the leader appended the record, but **before the record was committed** (replicated to all ISR members) the ISR shrank below `min.insync.replicas`. - **State**: the record **is** in the leader's log (LEO advanced) but the high watermark did **not** advance — it is uncommitted and invisible to consumers. - **Why a separate error**: it signals a partially-applied write. If a leader change happens now, this record may be truncated. ## Are they retriable? Yes — both are classified **retriable** in the Java client. The producer (with default `retries`/`delivery.timeout.ms`) backs off and retries. When enough followers re-enter the ISR, the retry commits successfully. If `delivery.timeout.ms` elapses first, the send fails permanently and the application sees the exception. ## The duplicate hazard with AfterAppend Because `NotEnoughReplicasAfterAppend` leaves the record in the leader's log, a **non-idempotent** producer that retries can append the same record again → **duplicate**. With `enable.idempotence=true` (default since Kafka 3.0), the producer attaches a producer ID + sequence number, and the broker deduplicates retries, so the retry is safe. This is a core reason idempotence is mandatory-recommended alongside `acks=all`. ## How to handle them as an operator/developer 1. **Don't catch-and-drop** — let the producer's retry machinery work; these are transient. 2. **Enable idempotence** so AfterAppend retries don't duplicate. 3. **Alert on them** — frequent NotEnoughReplicas means brokers are unhealthy or the ISR is chronically shrunk; investigate broker availability, GC pauses, or disk/network saturation. 4. **Don't 'fix' by lowering min.insync.replicas** under load — that silently weakens durability. 5. Tune `delivery.timeout.ms` so transient ISR shrinkage (e.g. a rolling restart) doesn't fail sends prematurely. ## Relationship to acks These errors are **specific to acks=all**. With acks=0/1 the broker doesn't consult `min.insync.replicas`, so these errors never appear — but neither does the durability guarantee.

  • Why is NotEnoughReplicasAfterAppend more dangerous for a non-idempotent producer than NotEnoughReplicas?
    AfterAppend means the record is already in the leader's log. A non-idempotent retry appends it a second time, creating a duplicate. NotEnoughReplicas rejects before appending, so a retry can't duplicate. Idempotence (producer ID + sequence) eliminates the AfterAppend duplicate risk.
  • These errors are retriable — what stops the producer from retrying forever?
    delivery.timeout.ms bounds the total time for a send (including retries and backoff). When it expires, the send fails permanently and the application gets the exception, even though the error itself was retriable.

saying these in an interview costs you the question

  • Saying these errors are fatal/non-retriable.
  • Recommending lowering min.insync.replicas as the fix (it weakens durability).
  • Claiming the record is never in the log for AfterAppend (it is in the leader's log, just uncommitted).
  • Expecting these errors with acks=1 (they only occur under acks=all).

context