skip to content

How does broker.rack enable rack awareness, and what does it (and doesn't it) guarantee for replica placement and reads?

level: seniorimportance: should knowfreq 45%

answer

  1. broker.rack = failure domain (AZ)
  2. spread replicas across racks
  3. placement only, not leadership
  4. not retroactive
  5. KIP-392 + client.rack for local reads

basics

~20 s

Setting broker.rack tags each broker with a failure-domain label (e.g. an AZ). Kafka then spreads a partition's replicas across distinct racks so one rack failure doesn't take all replicas down. It improves availability but doesn't change which replica is leader.

solid answer

~40 s

broker.rack assigns each broker a rack identifier representing a failure domain, typically a cloud availability zone. When a topic's replicas are assigned, Kafka's rack-aware placement spreads each partition's replicas across as many distinct racks as possible, so losing one rack (AZ) loses at most one replica per partition and the partition stays available. It applies at topic creation and partition reassignment, not retroactively. Important limits: it doesn't guarantee perfectly even placement if rack sizes are unequal, and it doesn't change the read path by itself, clients still fetch from the leader, which may be in a remote rack incurring cross-AZ cost. To read from a same-rack follower you additionally need follower fetching (KIP-392): set the broker replica selector (replica.selector.class=...RackAwareReplicaSelector) and the consumer's client.rack. So broker.rack handles durability/availability placement; KIP-392 handles locality reads.

go deeper

for a junior

Know broker.rack tags a broker's zone so replicas spread across zones.

for a middle

Explain placement at topic creation and that it isn't retroactive.

for a senior

Separate placement (broker.rack) from local reads (KIP-392, client.rack) and cross-AZ cost.

for a principal

Design multi-AZ topologies balancing durability, cost, and read locality, including uneven-rack tradeoffs.

## The problem rack awareness solves Kafka replicates each partition across several brokers for fault tolerance. But if all replicas of a partition happen to sit in the same **failure domain** (a server rack, or in the cloud an **availability zone / AZ**), one zone outage can take down every replica and make the partition unavailable or cause data loss. Rack awareness prevents this. ## broker.rack `broker.rack` is a per-broker string tagging the broker with its failure domain, e.g. `broker.rack=us-east-1a`. The cluster groups brokers by this label. When Kafka assigns replicas (at **topic creation** or during a **partition reassignment**), the rack-aware algorithm tries to place a partition's N replicas across N **distinct racks**. With 3 racks and replication factor 3, each replica lands in a different AZ; losing one AZ loses exactly one replica per partition, leaving the others (and the ISR/quorum) intact. This is purely about **placement**. ## What it does NOT do 1. **Not retroactive**: existing partitions aren't rebalanced just because you set broker.rack; you must reassign. 2. **No perfect balance with uneven racks**: if racks hold different numbers of brokers, placement skews; with fewer racks than the replication factor, some racks necessarily hold multiple replicas. 3. **Does not change leadership**: the leader can be in any rack. Clients by default fetch from the **leader**, so a consumer in AZ-a may pull from a leader in AZ-b, incurring **cross-AZ network cost** and latency. ## Reading locally: KIP-392 (fetch from follower) To actually read from a **same-rack replica** and save cross-AZ traffic, broker.rack alone is insufficient. You enable **follower fetching** (KIP-392): - On brokers: `replica.selector.class=org.apache.kafka.common.replica.RackAwareReplicaSelector` (default is the leader-only selector). - On consumers: set `client.rack` to the consumer's rack/AZ. Then the broker steers each consumer to an in-sync replica in its own rack when one exists, while writes still go to the leader. broker.rack provides the rack labels this selector relies on. ## Putting it together - **broker.rack** = durability/availability via cross-rack replica spread. - **KIP-392 (replica.selector.class + client.rack)** = read locality / cost savings. ## Edge cases - A follower must be **in-sync** to serve a fetch; a lagging follower won't, so the consumer falls back to the leader. - Producers always write to the leader regardless of rack. - Mislabeling racks (e.g. two AZs sharing a label) silently defeats the placement guarantee.

  • A consumer in AZ-a still reads across AZs despite broker.rack being set. Why, and how do you fix it?
    broker.rack only spreads replicas; reads still go to the leader by default. Enable KIP-392: set replica.selector.class to RackAwareReplicaSelector on brokers and client.rack on the consumer to fetch from a same-rack in-sync follower.
  • Does setting broker.rack rebalance existing partitions?
    No. It affects new topics and explicit partition reassignments; existing partitions must be reassigned to honor rack placement.

saying these in an interview costs you the question

  • Claiming broker.rack makes consumers read from the nearest replica automatically
  • Saying it guarantees perfectly even replica distribution
  • Thinking it changes which replica is leader
  • Believing it retroactively moves existing replicas
  • Saying producers can write to a follower in their rack

context