Explain fetch-from-follower (KIP-392): how does rack-aware consumer fetching work, what configs enable it, and why would you use it?
answer
- replica.selector.class = RackAwareReplicaSelector (broker)
- client.rack = AZ (consumer)
- consumers only; producers still hit leader
- read up to follower high-watermark
- driver = cross-AZ network cost
basics
~20 sKIP-392 lets a consumer read from a follower replica in its own rack/AZ instead of always from the leader. You set replica.selector.class on the broker and client.rack on the consumer. It cuts cross-AZ network cost and latency.
solid answer
~40 sBefore KIP-392, consumers always fetched from the partition leader, so a consumer in a different AZ paid cross-AZ network charges and latency. KIP-392 (fetch-from-follower) lets a consumer read from an in-sync follower in its own rack. You enable it broker-side with replica.selector.class=org.apache.kafka.common.replica.RackAwareReplicaSelector, and each consumer sets client.rack=<its-AZ>. On fetch, the broker's selector matches the consumer's client.rack against replica racks and returns a same-rack in-sync replica as the preferred read replica; the consumer then fetches from it. Only consumers benefit — producers still write to the leader, and replication still flows leader→follower. The follower may lag slightly, so reads see data up to the follower's high-watermark. The main driver is cost: in the cloud, cross-AZ traffic is billed, and follower reads keep consumer traffic intra-AZ. Default selector (LeaderSelector) preserves classic leader-only reads.
go deeper
Know consumers can read from a nearby follower instead of the leader to save cross-AZ cost.
Name both configs (replica.selector.class + client.rack) and that only consumers, not producers, are affected.
Explain the fetch redirect flow, high-watermark bound, ISR-only eligibility, and the cost-vs-freshness tradeoff.
Reason about custom ReplicaSelector strategies, truncation/epoch correctness under follower reads, and quantifying cross-AZ cost savings at scale.
## The problem KIP-392 solves Historically a Kafka consumer **always read from the partition leader**. The leader for a given partition lives on exactly one broker in one rack/AZ. If your consumer runs in a different AZ, every fetched byte crosses an AZ boundary — which in cloud providers (AWS/GCP/Azure) is **metered, billed network traffic** and adds round-trip latency. For high-throughput pipelines this cross-AZ cost can dominate the bill. ## What KIP-392 (fetch-from-follower) does It lets a consumer read from a **follower replica** — a non-leader in-sync copy — chosen to be in the **same rack** as the consumer, instead of the leader. Replication is unchanged (leader→followers); only the *consumer read path* gains a choice. ## The two configs 1. **Broker side:** `replica.selector.class`. Set it to `org.apache.kafka.common.replica.RackAwareReplicaSelector`. This plugs in a `ReplicaSelector` that, on each fetch, picks a preferred read replica. The default is `LeaderSelector` (always the leader = old behavior). You can implement a custom `ReplicaSelector` for other strategies. 2. **Client side:** `client.rack=<consumer's AZ>` on the consumer. This tells the broker which rack the consumer is in so the selector can match. ## The fetch flow 1. Consumer sends a Fetch request including its `client.rack`. 2. The leader broker runs the `ReplicaSelector`, which compares `client.rack` to the racks of the partition's in-sync replicas and returns a **preferred read replica** in the same rack (falling back to the leader if none matches). 3. The leader's fetch response includes that preferred replica id; the consumer then **redirects subsequent fetches to that follower**. 4. The consumer reads up to the **follower's high-watermark** (HW) — the offset the follower has confirmed replicated. Because a follower can lag the leader slightly, a follower read may be a touch behind, but never reads uncommitted data. ## Correctness and edge cases - **Only in-sync replicas** are eligible as read replicas; an out-of-sync follower won't be chosen. - The consumer tracks **offset/lag against the follower's HW**; the broker can send it back to the leader if the chosen replica becomes unavailable or falls out of ISR. - **OffsetForLeaderEpoch** handling and HW propagation were extended so a consumer reading from a follower still detects truncation/leadership changes correctly. - **Producers are unaffected** — writes always go to the leader; min.insync.replicas/acks semantics are unchanged. - Follower reads can show **slightly higher end-to-end latency to the freshest record** because of replication lag, a deliberate cost/freshness tradeoff. ## Why use it - **Cost:** keep consumer traffic intra-AZ → eliminate cross-AZ data transfer charges for reads. - **Latency/locality:** read from a nearby broker. - **Prereq:** brokers must have `broker.rack` set and consumers `client.rack` set; without rack info the selector can't match and falls back to the leader.
- Does fetch-from-follower change durability or producer acknowledgment semantics?No. Producers still write to the leader and acks/min.insync.replicas are unchanged. Only the consumer read path changes, and consumers still only read committed data (up to the follower's high-watermark).
- What happens if a consumer sets client.rack but no follower exists in that rack?The RackAwareReplicaSelector finds no same-rack in-sync replica and falls back to returning the leader, so the consumer simply reads from the leader as before — correct, just without the cost saving.
- Why might a follower read return slightly staler data than a leader read?A follower only exposes data up to its own high-watermark, which can trail the leader by replication lag. The read is still consistent (never uncommitted), just potentially a few records behind.
saying these in an interview costs you the question
- Claiming producers can also write to followers — only the read/consumer path is affected.
- Saying follower reads can return uncommitted data — they're bounded by the follower's high-watermark.
- Forgetting that both broker (replica.selector.class) AND client (client.rack) configs are required.
- Confusing broker.rack with client.rack — broker.rack labels brokers, client.rack labels consumers.