skip to content

What is a rolling upgrade of a Kafka cluster, and why do you restart brokers one at a time instead of all at once?

level: juniorimportance: must knowfreq 70%

answer

  1. one broker down at a time
  2. wait UnderReplicatedPartitions == 0
  3. leadership fails over to ISR
  4. replication = the reason it works
  5. RF=3, min.insync.replicas=2

basics

~20 s

A rolling upgrade restarts brokers one at a time, swapping in the new version on each, so the cluster keeps serving while every broker is replaced. One broker is down at a time; the rest keep handling reads and writes.

solid answer

~40 s

A rolling upgrade upgrades a multi-broker Kafka cluster with zero downtime by taking down and restarting only one broker at a time. You stop a broker, install the new binaries, restart it, wait until it has fully rejoined and its partitions are back in sync (all replicas in the ISR, under-replicated partitions back to zero), then move to the next broker. Because Kafka replicates each partition across multiple brokers, the remaining brokers keep serving produce and consume requests for partitions whose leaders were on the downed broker — leadership simply fails over to an in-sync replica. Restarting all brokers at once would take the whole cluster offline and risk data unavailability. The one-at-a-time discipline preserves availability and gives you a checkpoint to abort if a broker comes back unhealthy.

go deeper

for a junior

Know the definition: restart one broker at a time so the cluster stays up; replication makes it safe.

for a middle

Know the health gate (UnderReplicatedPartitions=0), controlled shutdown, and the RF=3/min.insync.replicas=2 combo.

for a senior

Reason about availability math: which acks/min.insync settings keep producers alive with one broker down, and when leadership failover stalls.

for a principal

Design the upgrade runbook and automation, define health gates, and teach why redundancy is consumed and restored on each step.

## What is Kafka and a broker? Apache Kafka is a distributed log/messaging system. A *broker* is a single server process in a Kafka cluster. Topics are split into *partitions*; each partition is replicated to several brokers for fault tolerance. - One replica is the *leader* (handles all reads/writes for that partition) and the others are *followers* that copy the leader's data. - The set of replicas that are caught up is the *ISR* (in-sync replica set). ## What is a rolling upgrade? It is the procedure for moving a running cluster from version A to version B (or applying config changes) without taking the whole cluster offline. You upgrade ***one broker at a time***: 1. stop broker 1, 2. replace its software, 3. start it, 4. wait for it to be fully healthy, 5. then do broker 2, and so on. ## Why one at a time? Kafka tolerates the loss of a single broker because each partition has replicas elsewhere. When you stop a broker, every partition it was leading fails over: the controller elects a new leader from the ISR on a surviving broker. Clients transparently retry against the new leader. If you instead restarted *all* brokers at once, every partition would lose all replicas simultaneously — total outage and potential data unavailability until they all come back. One-at-a-time keeps `min.insync.replicas` satisfiable so producers using `acks=all` keep succeeding. ## The health checkpoint between steps After restarting a broker you must wait until the cluster is fully healthy before touching the next one. The key signal is the **`UnderReplicatedPartitions`** JMX metric (kafka.server:type=ReplicaManager) returning to **0** — meaning every partition is fully replicated again. You also watch that: - the restarted broker has rejoined the ISR for its partitions and - that there are no offline partitions. Only then is it safe to take down the next broker, because the cluster has recovered the redundancy you just spent. ## Edge cases 1. If a topic has replication factor 1, the partitions on the downed broker are simply unavailable during its restart — rolling upgrades only give zero downtime for replicated topics. 2. `min.insync.replicas` must be low enough that the cluster still meets it with one broker down (RF=3 + min.insync.replicas=2 is the canonical safe combo). 3. Controlled shutdown (`controlled.shutdown.enable=true`, default) makes the broker proactively move leadership off itself before stopping, shortening the unavailability window. ## Beyond binaries A rolling restart is also how you apply many broker config changes and, importantly, how you bump protocol/format versions during a *two-phase* version upgrade (covered separately).

  • What metric tells you it is safe to move on to the next broker?
    UnderReplicatedPartitions (kafka.server:type=ReplicaManager) back to 0, meaning every partition is fully replicated again; also confirm no offline partitions and the broker has rejoined its ISRs.
  • Why does a rolling upgrade not work for a topic with replication factor 1?
    There is only one copy of each partition, so when its broker goes down there is no in-sync replica to fail over to — those partitions are unavailable for the duration of that broker's restart.

saying these in an interview costs you the question

  • Saying you can restart all brokers simultaneously with no downtime.
  • Claiming rolling upgrades give zero downtime even for RF=1 topics.
  • Not waiting for under-replicated partitions to clear before moving to the next broker.
  • Confusing a rolling upgrade (one broker at a time) with simply restarting the cluster.

context