skip to content

Rack Awareness and Replica Placement

Spreading replicas across racks or availability zones with broker.rack, plus rack-aware follower fetching. Comes up in any multi-AZ design question about surviving a zone failure.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the broker.rack configuration in Kafka, and what does setting it actually do?

level: juniorimportance: must knowfreq 60%

answer

  1. per-broker string label = rack/AZ
  2. spreads replicas across racks
  3. survive one rack/AZ failure
  4. just a tag, Kafka doesn't validate
  5. set on ALL brokers

basics

~20 s

broker.rack is a per-broker string tagging which rack or availability zone a broker is in. Kafka uses these tags to spread a partition's replicas across different racks, so one rack failing doesn't take down all copies.

solid answer

~40 s

broker.rack is a per-broker config (e.g. broker.rack=us-east-1a) that labels the broker's physical rack or cloud availability zone. When set on all brokers, Kafka's rack-aware replica assignment tries to place each partition's replicas in as many distinct racks as possible. The goal is fault isolation: if an entire rack or AZ goes down, every partition should still have at least one surviving replica, so leadership can move and no data is lost. It takes effect during automatic topic creation, partition reassignment, and the kafka-topics CLI. It's just a label — Kafka doesn't validate it against real topology; you must set it correctly to match your physical/cloud layout.

go deeper

for a junior

Know it's a per-broker label for rack/AZ that makes Kafka spread replicas across racks to survive a rack failure.

for a middle

Know it's all-or-nothing across brokers, only affects new placements, and is just a label Kafka doesn't validate.

for a senior

Tie it to AZ mapping in cloud, RF vs rack count math, and that it's the prerequisite for both fault isolation and KIP-392 follower reads.

for a principal

Reason about topology design, mixed-config fallback behavior, and operational migration of existing topics into rack-aware layout.

## What broker.rack is `broker.rack` is a string configuration property you set in each broker's `server.properties` (or via dynamic config). It names the failure domain the broker lives in — typically a physical server rack in a data center, or a cloud **Availability Zone (AZ)** like `us-east-1a`. Example: `broker.rack=us-east-1a`. A **replica** is one copy of a partition's data. Every Kafka partition has a configurable **replication factor** (RF) — e.g. RF=3 means three brokers each hold a full copy: one **leader** (handles reads/writes) and two **followers** (continuously replicate from the leader). ## What it does When `broker.rack` is set on all brokers, Kafka switches to **rack-aware replica assignment**. Instead of just round-robining replicas across brokers, the assignment algorithm tries to spread a partition's RF replicas across as many *distinct racks* as possible. With RF=3 and 3 racks, each replica lands in a different rack. This is **fault isolation**: if one whole rack/AZ fails, every partition still has ≥1 surviving replica, leadership fails over to a survivor, and you lose no data and (usually) no availability. ## When it takes effect Rack-aware placement runs at: automatic topic creation, the `kafka-topics.sh --create` CLI, and partition reassignment (`kafka-reassign-partitions.sh`) when you let Kafka generate the assignment. It does **not** retroactively move replicas of topics that already exist — you must reassign them. ## Important caveats - It's **just a label**. Kafka does not verify it matches real topology. If you mislabel brokers, you get a false sense of safety. - **All-or-nothing**: if *some* brokers set `broker.rack` and others don't, by default Kafka refuses rack-aware assignment and falls back to non-rack-aware (controlled by whether you let it tolerate mixed config). Set it on every broker. - It enables, but is independent from, **fetch-from-follower** (KIP-392), which uses rack info on the *client* side to reduce cross-AZ network cost. ## Why it matters in the cloud Cross-AZ traffic costs money and adds latency, and AZ-wide outages happen. Mapping `broker.rack` to AZs gives you both an availability guarantee (survive an AZ loss) and the foundation for cost-saving follower reads.

  • Does setting broker.rack rebalance replicas of topics that already exist?
    No. It only affects new placement decisions — topic creation and reassignment. Existing topics keep their current layout until you explicitly run a partition reassignment.
  • What happens if only some brokers have broker.rack set?
    By default Kafka treats the cluster as having incomplete rack info and falls back to non-rack-aware assignment (it won't silently place some replicas rack-aware and others not). Set it on every broker for consistent behavior.

saying these in an interview costs you the question

  • Saying broker.rack physically moves data or validates real topology — it's just a label.
  • Claiming it automatically rebalances existing topics.
  • Confusing broker.rack (server-side placement) with client rack config for follower fetching.

context

open as a page

Explain fetch-from-follower (KIP-392): how does rack-aware consumer fetching work, what configs enable it, and why would you use it?

level: seniorimportance: must knowfreq 50%

basics

~20 s

KIP-392 lets a consumer read from a follower replica in its own rack/AZ instead of always from the leader. You set replica.selector.class on the broker and client.rack on the consumer. It cuts cross-AZ network cost and latency.

open as a page

Walk through what happens to a Kafka topic (RF=3 across 3 AZs, min.insync.replicas=2, acks=all) when one entire AZ fails. What survives, and what would break this guarantee?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Each partition keeps 2 of its 3 replicas because they're in different AZs. Leaders that were in the dead AZ fail over to surviving replicas, ISR drops to 2 which still meets min.insync.replicas=2, so acks=all producers and consumers keep working with no data loss.

open as a page

How does Kafka's rack-aware replica assignment algorithm decide where to place replicas, and what is the relationship between replication factor and the number of racks?

level: middleimportance: should knowfreq 45%

basics

~20 s

Kafka assigns each partition's replicas by cycling through racks so no two replicas share a rack until it runs out of racks. With replication factor 3 and 3 racks, each replica lands in a different rack; with fewer racks than RF, some replicas double up.

open as a page

As a platform architect, how would you design rack/AZ topology, replication, and fetch strategy for a multi-AZ Kafka cluster to balance durability, availability, and cross-AZ cost? What are the principal tradeoffs?

level: principalimportance: should knowfreq 30%

basics

~20 s

Map broker.rack to AZs, use RF=3 across 3 AZs with min.insync.replicas=2 and unclean.leader.election=false for durability. Enable KIP-392 fetch-from-follower (RackAwareReplicaSelector + client.rack) to cut cross-AZ read costs. The main tradeoffs are cost vs freshness and balanced AZ sizing.

open as a page