skip to content

Stretch Clusters and Rack-Aware Placement

Running one logical cluster across availability zones with rack-aware placement and synchronous cross-AZ replication. Interviewers ask when the design goal is zero RPO rather than asynchronous replication.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the broker.rack configuration in Kafka, and why does it matter for a cluster spread across availability zones?

level: juniorimportance: must knowfreq 70%

answer

  1. per-broker AZ label
  2. spread replicas across racks
  3. survive an AZ outage
  4. best-effort, not enforced
  5. no auto-move of existing replicas

basics

~20 s

broker.rack labels each broker with its location (e.g. its availability zone). Kafka uses these labels to spread the copies of each partition across different zones, so losing one zone does not lose all copies of the data.

solid answer

~40 s

broker.rack is a per-broker string config that tags a broker with a failure domain, typically the availability zone (AZ) it runs in, e.g. broker.rack=us-east-1a. When a topic is created or partitions are reassigned, Kafka's rack-aware replica placement tries to put each partition's replicas on brokers in different racks. The goal is that the replicas of any one partition span multiple AZs, so an entire AZ outage does not take all replicas (and thus the partition) offline. It also influences rack-aware consumer/follower fetching when configured. Without broker.rack set, Kafka treats all brokers as one rack and may pile all replicas of a partition into a single AZ, defeating the point of a multi-AZ deployment.

go deeper

for a junior

Know it is a per-broker label (usually the AZ) that makes Kafka spread partition copies across zones to survive a zone outage.

for a middle

Explain best-effort placement, that it must be set per broker, and that it does not retroactively move replicas.

for a senior

Connect it to reassignment tooling, rack count vs replication factor trade-offs, and its role in follower-fetch reads.

for a principal

Reason about it as part of a multi-AZ availability design: rack count, RF, min.insync.replicas, and failure-domain modeling together.

## What broker.rack is Kafka stores each topic partition as multiple **replicas** (copies) on different brokers. The number of copies is the **replication factor**. One replica is the **leader** (handles reads/writes); the others are **followers** that copy the leader's data. A **rack** in Kafka terminology is a *failure domain* — a group of brokers that can fail together. In the cloud, the natural failure domain is an **availability zone (AZ)**: an isolated datacenter within a region. `broker.rack` is a string config set on each broker (in `server.properties`) that names which failure domain that broker belongs to, e.g.: ``` broker.rack=us-east-1a ``` ## Why it matters When Kafka assigns replicas to brokers (at topic creation, partition increase, or reassignment), the **rack-aware replica assignment** algorithm tries to place the replicas of each partition on brokers in *as many distinct racks as possible*. So with replication factor 3 and three AZs, you ideally get one replica in each AZ. The payoff: if an entire AZ goes down, every partition still has surviving replicas in the other AZs, so the cluster keeps serving. Without `broker.rack`, Kafka considers every broker to be in the same (unnamed) rack and may place all 3 replicas of a partition on brokers that happen to be in the same AZ. Then one AZ outage takes that partition completely offline — data is unavailable until the AZ returns. ## Edge cases / details - `broker.rack` is **not** automatically derived; an operator (or the cloud provisioning/Helm chart) must set it correctly per broker. - Rack-awareness is **best effort**: if there are fewer racks than the replication factor, some racks will hold more than one replica. - Changing `broker.rack` on an existing broker does **not** move existing replicas; you must run a partition reassignment to rebalance. - It also feeds **rack-aware fetching** (KIP-392 follower fetching) so consumers can read from a same-AZ replica to cut cross-AZ network cost/latency. - Rack info is exposed in metadata and used by tools like `kafka-reassign-partitions` (with `--enable-rack-aware`, which is on by default when racks are present).

  • If you add broker.rack to brokers that already host topics, are existing partitions rebalanced automatically?
    No. broker.rack only affects new placement decisions. Existing replica assignments are untouched until you run a partition reassignment (e.g. kafka-reassign-partitions), which can take rack info into account.
  • What happens if you have replication factor 3 but only 2 racks defined?
    Rack-awareness is best-effort, so it spreads as widely as it can: two replicas land in distinct racks and the third doubles up in one of them. You no longer survive an arbitrary single-rack loss for every partition without care, so 3 racks for RF=3 is preferred.

saying these in an interview costs you the question

  • Saying broker.rack physically moves data when changed (it only affects future placement)
  • Claiming rack-awareness guarantees one replica per AZ (it is best-effort and depends on rack/RF counts)
  • Confusing broker.rack with consumer rack.id used for follower fetching

context

open as a page

In a stretch cluster across 3 AZs, how should replication.factor, min.insync.replicas, and acks be set so the cluster survives the loss of one AZ without data loss?

level: middleimportance: must knowfreq 65%

basics

~20 s

Use replication factor 3 (one replica per AZ), min.insync.replicas=2, and producers with acks=all. Then a write is only acknowledged after two zones have it, so losing one AZ still leaves a committed copy and writes keep working.

open as a page

How does synchronous cross-AZ (or cross-DC) replication latency affect producer throughput and tail latency in a stretch cluster, and what levers tune that trade-off?

level: seniorimportance: should knowfreq 35%

basics

~20 s

With acks=all the leader waits for in-sync replicas in other zones before acknowledging, so each write pays a cross-AZ round trip. That raises per-record latency. You hide it with batching (linger.ms, batch.size) and concurrency (in-flight requests), and by keeping zones close.

open as a page

Explain follower fetching (KIP-392) in a multi-AZ Kafka cluster: what problem it solves, how to enable it, and its consistency implications.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Follower fetching lets a consumer read from a nearby replica in its own AZ instead of always from the leader. This cuts cross-AZ network traffic and cost. You enable it with a broker replica.selector.class and a consumer client.rack matching its zone.

open as a page

What is a 2.5-DC stretch cluster design, and how do observers / asymmetric quorums (e.g. ZooKeeper or KRaft witnesses) help avoid split-brain across two main datacenters?

level: principalimportance: should knowfreq 30%

basics

~20 s

A 2.5-DC design runs Kafka across two full datacenters plus a tiny third site holding just a tie-breaker (a witness/observer). The third site holds no data but lets the cluster keep a majority quorum if one of the two main DCs fails, avoiding split-brain.

open as a page