skip to content

In-Sync Replicas and Replica Lag

What puts a replica in the in-sync set and what makes it drop out, driven by replica.lag.time.max.ms. Central to every durability question, since acks=all only ever waits for the ISR.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the In-Sync Replicas (ISR) set in Apache Kafka, and how does it relate to the assigned replicas of a partition?

level: juniorimportance: must knowfreq 78%

answer

  1. ISR ⊆ AR (assigned replicas)
  2. Leader always in its own ISR
  3. Caught-up subset within a time bound
  4. acks=all waits for all ISR
  5. Only ISR members electable (by default)

basics

~20 s

The ISR is the subset of a partition's replicas (leader plus followers) that are fully caught up with the leader's log. Assigned replicas (AR) are all replicas; ISR is the healthy, up-to-date ones eligible to serve and be elected leader.

solid answer

~40 s

Every Kafka partition has a list of assigned replicas (AR) spread across brokers: one leader and the rest followers. The In-Sync Replicas (ISR) set is the subset of those replicas that are currently caught up with the leader within a configured time bound. The leader is always in the ISR. Followers stay in the ISR by continuously fetching and keeping pace with the leader's log; if a follower falls behind or stops fetching, the leader removes it from the ISR. The ISR matters for durability and availability: with acks=all, a produce request is only acknowledged once all ISR members have replicated it, and by default only an ISR member can be elected the new leader. So ISR ⊆ AR, and ISR is the set Kafka trusts at any moment.

go deeper

for a junior

Know that ISR = caught-up replicas, a subset of all assigned replicas, and the leader is always in it.

for a middle

Connect ISR to acks=all durability and to leader-election eligibility.

for a senior

Explain how high watermark advances with ISR and the role of min.insync.replicas in rejecting writes.

for a principal

Reason about ISR as the consistency boundary and the durability/availability tradeoffs it encodes across the cluster.

## The problem ISR solves Kafka replicates each **partition** (an ordered log of messages) across multiple brokers for fault tolerance. The full list of brokers hosting a partition is the **assigned replicas (AR)** — sometimes called the replica set. One of them is the **leader**; the rest are **followers**. All reads and writes go through the leader; followers continuously **fetch** new records from the leader to stay current. But replication is asynchronous: a follower might be slow, GC-pausing, network-partitioned, or simply rebooting. Kafka needs a way to know *which* replicas are trustworthy right now. That set is the **In-Sync Replicas (ISR)**. ## Definition The **ISR** is the subset of the AR that is currently caught up to the leader within a time bound (`replica.lag.time.max.ms`, default 30000 ms / 30s). The leader is always a member of its own ISR. A follower is in the ISR if it has fetched up to the leader's log end offset recently enough. So always: **ISR ⊆ AR**. In a healthy cluster with replication factor 3, AR has 3 members and ISR also has 3. If one follower dies, ISR shrinks to 2 while AR stays 3. ## Why ISR matters 1. **Durability (acks=all):** A producer using `acks=all` only gets its write acknowledged after *every member of the current ISR* has appended the record. The **high watermark** — the offset up to which consumers can read — advances only when all ISR members have the data. 2. **Leader election:** By default (`unclean.leader.election.enable=false`), only a replica that is in the ISR can be elected as the new leader if the current leader fails. This guarantees no committed data is lost. 3. **min.insync.replicas:** A topic/broker config that sets the minimum ISR size required to accept an `acks=all` write. If ISR shrinks below it, producers get `NotEnoughReplicasException`. ## Key terms - **AR (assigned replicas):** all brokers that host this partition. - **ISR:** the caught-up subset, tracked and published by the leader. - **High watermark (HW):** highest offset replicated to all ISR members; the consumer-visible boundary. - **Log end offset (LEO):** the next offset to be written on a given replica. A follower is in-sync when its LEO has caught up to the leader's LEO recently enough (within `replica.lag.time.max.ms`).

  • Is the leader always part of the ISR?
    Yes. The leader is by definition in-sync with itself, so it is always a member of its partition's ISR. Followers may join and leave, but the leader stays until it fails and a new leader is elected.
  • If replication factor is 3 and one follower crashes, what are AR and ISR?
    AR stays 3 — the assignment doesn't change just because a broker is down. ISR shrinks to 2 (leader plus the one healthy follower) once the crashed follower exceeds replica.lag.time.max.ms.

saying these in an interview costs you the question

  • Saying ISR and AR are the same thing — ISR is a dynamic subset of AR.
  • Claiming the leader can be excluded from the ISR.
  • Thinking ISR includes replicas that are merely assigned but lagging or down.

context

open as a page

How does replica.lag.time.max.ms control ISR shrink and expansion, and why did Kafka switch from a message-count-based lag check?

level: middleimportance: must knowfreq 70%

basics

~20 s

A follower stays in the ISR if it fetched up to the leader's latest offset within replica.lag.time.max.ms (default 30s). Miss that window and the leader removes it; catch back up and it rejoins. The old count-based check (lag in messages) misjudged bursty traffic, so it was replaced by time.

open as a page

Your monitoring shows UnderReplicatedPartitions spiking and IsrShrinksPerSec elevated, but brokers are all up. How do you diagnose and act?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Under-replicated partitions with all brokers up means followers can't keep up, not that brokers are down. Check follower I/O (disk, network, GC), look at IsrShrinks/ExpandsPerSec for flapping, inspect which brokers are dropping out, and either fix the bottleneck or tune replica.lag.time.max.ms / replica fetcher threads.

open as a page

Walk through what happens, step by step, when a follower falls out of the ISR and later rejoins. Who decides, and how is the change propagated?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The leader detects a follower exceeding replica.lag.time.max.ms, shrinks the ISR, and persists the new ISR via the controller (AlterPartition in KRaft, ZooKeeper in older versions). When the follower catches back up to the leader's LEO, the leader expands the ISR and re-propagates it.

open as a page

How do ISR shrink dynamics interact with min.insync.replicas, acks=all, and unclean.leader.election.enable to shape the durability vs. availability tradeoff?

level: principalimportance: should knowfreq 48%

basics

~20 s

When the ISR shrinks, fewer replicas hold committed data. min.insync.replicas sets the floor: below it, acks=all writes are rejected (availability lost to protect durability). unclean.leader.election lets an out-of-sync replica become leader if ISR is empty — restoring availability but risking data loss.

open as a page