skip to content

What is broker.rack rack awareness, and how should it influence how you add or remove brokers in a cloud deployment across availability zones?

level: seniorimportance: should knowfreq 45%

answer

  1. broker.rack = fault domain = AZ
  2. spreads replicas across distinct racks
  3. RF>=3 across 3 AZs, min.insync=2
  4. add/remove brokers symmetrically per AZ
  5. manual reassignment JSON bypasses rack checks
  6. KIP-392 fetch-from-follower cuts cross-AZ cost

basics

~20 s

broker.rack tags each broker with a fault domain (e.g., an availability zone). Kafka then spreads a partition's replicas across different racks, so losing one rack/AZ doesn't take all copies. When scaling or decommissioning, keep replicas balanced across racks.

solid answer

~50 s

Setting broker.rack=<az-id> on each broker enables rack-aware replica placement: when assigning a partition's replicas (at topic creation or via reassignment with --generate), Kafka tries to put each replica in a different rack so no single rack failure loses a quorum of copies. In cloud, racks map to availability zones, giving AZ-fault tolerance. Implications: (1) Replication factor should be >= number of racks you want to survive (often RF=3 across 3 AZs). (2) When adding brokers, add them rack-tagged and regenerate balanced assignments so new partitions still span AZs. (3) When decommissioning, ensure the drain doesn't collapse two replicas of a partition into the same rack or below your AZ-survival target. (4) Hand-written reassignment JSON bypasses rack-aware logic — you own correctness. (5) Cross-AZ replication costs network/latency; rack-aware consumer fetch (fetch-from-follower, KIP-392, replica.selector.class) lets consumers read from a same-AZ follower to cut egress cost.

go deeper

for a junior

Know broker.rack tags an AZ/fault domain so replicas spread across racks.

for a middle

Explain that placement is best-effort and disabled if any broker is untagged; tie RF to rack count.

for a senior

Reason about AZ-balanced scaling/decommission and manual-JSON pitfalls.

for a principal

Design multi-AZ topology: RF/min.insync defaults, fetch-from-follower for cost, and rack-aware balancing automation.

**Rack awareness defined:** A **rack** in Kafka is an abstract **fault domain** — a group of brokers that can fail together. You declare it per broker with the config `broker.rack=<id>`. In on-prem deployments a rack is a literal server rack or power/network domain; in the cloud it maps to an **availability zone (AZ)**. **What it changes:** When Kafka decides where to place a partition's replicas — at **topic creation**, partition increase, or during `kafka-reassign-partitions.sh --generate` — the rack-aware assignment algorithm tries to **spread the replicas of each partition across as many distinct racks as possible**. Goal: for a partition with replication factor 3, put the three copies in three different AZs so that an entire-AZ outage still leaves >=2 copies (and an ISR) alive. It also balances leadership across racks. If `broker.rack` is unset on *any* broker, Kafka falls back to rack-unaware (round-robin) placement for safety. **Cloud sizing implications:** - **Replication factor vs racks:** To survive one AZ failure with no loss, you want RF >= 2 with replicas in distinct AZs (RF=3 across 3 AZs is the common, durable choice). `min.insync.replicas=2` then keeps `acks=all` producers writing as long as 2 AZs are up. - **Adding brokers:** Always set `broker.rack` *before* the broker joins, and add capacity **symmetrically across AZs** when possible. Then regenerate reassignment plans so newly balanced partitions still satisfy rack spread. Adding all new brokers into one AZ skews placement and undermines AZ fault tolerance. - **Removing/decommissioning brokers:** A naive drain can move a partition's replica onto a broker that's in the *same* rack as another of its replicas, silently collapsing two copies into one AZ. Use rack-aware tooling (Cruise Control with rack-awareness goal, or `--generate`) rather than hand-editing JSON, and verify each partition still spans the intended number of racks afterward. - **Manual JSON caveat:** `--execute` with a hand-written reassignment file does **not** re-run the rack-aware checker — it does exactly what you wrote. You are responsible for rack correctness. **Cost/perf angle — fetch from follower (KIP-392):** Cross-AZ traffic costs money and adds latency. By configuring `replica.selector.class=org.apache.kafka.common.replica.RackAwareReplicaSelector` on brokers and setting `client.rack` on consumers, a consumer can **fetch from a same-AZ follower replica** instead of always the leader, dramatically cutting cross-AZ egress. This relies on the same `broker.rack` tags. **Edge cases:** Uneven broker counts per rack make perfect spread impossible; Kafka does best-effort. If you have more racks than the replication factor, not every rack hosts a copy of a given partition — that's fine; the guarantee is 'replicas in distinct racks up to RF'. Mixing rack-tagged and untagged brokers disables rack-aware placement entirely.

  • If you hand-write a reassignment JSON instead of using --generate, is rack awareness still enforced?
    No. --execute applies exactly the replica lists you provide; the rack-aware assignment algorithm only runs during --generate or automatic placement. With manual JSON you can accidentally put two replicas of a partition in the same rack/AZ, so you must verify rack spread yourself.
  • How can rack awareness reduce cloud networking cost beyond fault tolerance?
    Via fetch-from-follower (KIP-392): set replica.selector.class=RackAwareReplicaSelector on brokers and client.rack on consumers so consumers read from a same-AZ follower instead of the cross-AZ leader, cutting inter-AZ egress charges and latency.

saying these in an interview costs you the question

  • Thinking rack awareness guarantees replicas in distinct racks even if some brokers lack broker.rack (any untagged broker disables it).
  • Believing manual reassignment JSON is checked for rack correctness.
  • Adding all new brokers into a single AZ and assuming AZ fault tolerance is preserved.
  • Confusing rack awareness with consumer fetch-from-follower (related but separate configs).

context