Explain follower fetching (KIP-392) in a multi-AZ Kafka cluster: what problem it solves, how to enable it, and its consistency implications.
answer
- KIP-392, Kafka 2.4
- client.rack + replica.selector.class
- RackAwareReplicaSelector
- reads only up to high watermark
- cuts cross-AZ $ but adds staleness; producers still hit leader
basics
~20 sFollower fetching lets a consumer read from a nearby replica in its own AZ instead of always from the leader. This cuts cross-AZ network traffic and cost. You enable it with a broker replica.selector.class and a consumer client.rack matching its zone.
solid answer
~50 sNormally all consumer reads go to the partition leader, which may sit in a different AZ — generating expensive cross-AZ traffic. KIP-392 (Kafka 2.4) added 'fetch from closest replica': the consumer sends its location via client.rack, and the broker uses a pluggable ReplicaSelector (set with replica.selector.class, typically RackAwareReplicaSelector) to redirect the fetch to an in-sync replica in the same rack/AZ. This slashes inter-AZ bandwidth charges and can lower latency. Consistency is preserved because followers only serve data up to the high watermark (the offset all ISR have replicated), so consumers never read uncommitted records. The trade-off is increased read latency tail: a follower lags the leader slightly, so a consumer may see a record a beat later than reading the leader would. Writes/produces are unaffected — they still go to the leader. It requires followers to be in ISR to be selectable.
go deeper
Know it exists: consumers can read from a nearby copy instead of the leader to save cross-zone cost.
Name the two configs (replica.selector.class + client.rack) and that producers still write to the leader.
Explain the high-watermark consistency bound and the cost-vs-staleness trade-off in detail.
Weigh follower-fetch savings against latency SLAs, design rack labelling, and consider custom ReplicaSelector for non-AZ topologies.
## The problem By default every consumer fetch goes to the **partition leader**. In a stretch cluster the leader for a given partition lives in one AZ, but consumers can be in any AZ. A consumer in AZ-b reading a leader in AZ-a generates **cross-AZ network traffic**, which cloud providers bill per GB and which adds round-trip latency. At scale this inter-AZ data transfer is often a dominant Kafka cost. ## What KIP-392 does Introduced in **Kafka 2.4**, *fetch from closest replica* lets a consumer read from a **follower** replica near it instead of the leader. Two pieces: 1. **Broker side:** `replica.selector.class`. Out of the box Kafka ships `org.apache.kafka.common.replica.RackAwareReplicaSelector`. Set it on the brokers to turn on rack-aware fetch routing. (Default is null → always leader.) 2. **Consumer side:** `client.rack` — the consumer declares its rack/AZ, e.g. `client.rack=us-east-1b`. At fetch time the broker passes the consumer's `client.rack` and the replica list to the selector, which returns the **preferred read replica** — ideally an in-sync follower in the same rack. The consumer is told to redirect subsequent fetches there. If no same-rack in-sync replica exists, it falls back to the leader. ## Consistency guarantees The key safety property: **followers only serve up to the high watermark (HW)** — the offset that *all* ISR members have replicated. Records above the HW are not yet committed and are never returned. So follower reads are **consistent with leader reads in terms of committed data** — you cannot read an uncommitted or about-to-be-truncated record. What you *do* trade: - **Latency / staleness:** a follower is always slightly behind the leader (replication lag), so a record visible on the leader may appear on the follower a few milliseconds later. End-to-end consumer latency can rise even as cost drops. - **Read-your-writes across replicas** is not guaranteed to be instantaneous from a follower. - The follower must be **in ISR** to be eligible; a lagging out-of-sync follower won't be selected. ## Operational notes - Only **consumer fetches** are redirected; **producers still write to the leader**. - Offsets are global per partition, so switching read replica does not change offsets the consumer sees. - Works hand-in-hand with `broker.rack` (the brokers must be rack-labelled for the rack-aware selector to match consumer `client.rack`). - You can implement a custom `ReplicaSelector` for more exotic topologies. - This is the main lever for cutting cross-AZ egress cost in a stretch deployment without changing replication semantics.
- Does follower fetching ever let a consumer read uncommitted data?No. Followers serve reads only up to the high watermark — the offset replicated by all in-sync replicas. Records above the HW are uncommitted and never returned, so follower reads see exactly the same committed data as leader reads, just possibly a few ms later.
- What are the two configs needed to turn on rack-aware follower fetching?On the broker: replica.selector.class=org.apache.kafka.common.replica.RackAwareReplicaSelector. On the consumer: client.rack set to its AZ. The brokers must also have broker.rack set so the selector can match the consumer's rack to a co-located replica.
saying these in an interview costs you the question
- Claiming follower fetching redirects producer writes (only consumer reads move; producers still hit the leader)
- Saying it can return uncommitted/dirty data (it is bounded by the high watermark)
- Forgetting client.rack on the consumer and assuming broker.rack alone enables it
- Thinking it improves latency (it primarily cuts cost and may add slight staleness)