skip to content

Rolling Upgrades and Version Compatibility

Upgrading a cluster one broker at a time, with protocol and message-format version bumps as a deliberate second pass. Interviewers ask because getting that two-phase order wrong destroys your rollback option.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a rolling upgrade of a Kafka cluster, and why do you restart brokers one at a time instead of all at once?

level: juniorimportance: must knowfreq 70%

answer

  1. one broker down at a time
  2. wait UnderReplicatedPartitions == 0
  3. leadership fails over to ISR
  4. replication = the reason it works
  5. RF=3, min.insync.replicas=2

basics

~20 s

A rolling upgrade restarts brokers one at a time, swapping in the new version on each, so the cluster keeps serving while every broker is replaced. One broker is down at a time; the rest keep handling reads and writes.

solid answer

~40 s

A rolling upgrade upgrades a multi-broker Kafka cluster with zero downtime by taking down and restarting only one broker at a time. You stop a broker, install the new binaries, restart it, wait until it has fully rejoined and its partitions are back in sync (all replicas in the ISR, under-replicated partitions back to zero), then move to the next broker. Because Kafka replicates each partition across multiple brokers, the remaining brokers keep serving produce and consume requests for partitions whose leaders were on the downed broker — leadership simply fails over to an in-sync replica. Restarting all brokers at once would take the whole cluster offline and risk data unavailability. The one-at-a-time discipline preserves availability and gives you a checkpoint to abort if a broker comes back unhealthy.

go deeper

for a junior

Know the definition: restart one broker at a time so the cluster stays up; replication makes it safe.

for a middle

Know the health gate (UnderReplicatedPartitions=0), controlled shutdown, and the RF=3/min.insync.replicas=2 combo.

for a senior

Reason about availability math: which acks/min.insync settings keep producers alive with one broker down, and when leadership failover stalls.

for a principal

Design the upgrade runbook and automation, define health gates, and teach why redundancy is consumed and restored on each step.

## What is Kafka and a broker? Apache Kafka is a distributed log/messaging system. A *broker* is a single server process in a Kafka cluster. Topics are split into *partitions*; each partition is replicated to several brokers for fault tolerance. - One replica is the *leader* (handles all reads/writes for that partition) and the others are *followers* that copy the leader's data. - The set of replicas that are caught up is the *ISR* (in-sync replica set). ## What is a rolling upgrade? It is the procedure for moving a running cluster from version A to version B (or applying config changes) without taking the whole cluster offline. You upgrade ***one broker at a time***: 1. stop broker 1, 2. replace its software, 3. start it, 4. wait for it to be fully healthy, 5. then do broker 2, and so on. ## Why one at a time? Kafka tolerates the loss of a single broker because each partition has replicas elsewhere. When you stop a broker, every partition it was leading fails over: the controller elects a new leader from the ISR on a surviving broker. Clients transparently retry against the new leader. If you instead restarted *all* brokers at once, every partition would lose all replicas simultaneously — total outage and potential data unavailability until they all come back. One-at-a-time keeps `min.insync.replicas` satisfiable so producers using `acks=all` keep succeeding. ## The health checkpoint between steps After restarting a broker you must wait until the cluster is fully healthy before touching the next one. The key signal is the **`UnderReplicatedPartitions`** JMX metric (kafka.server:type=ReplicaManager) returning to **0** — meaning every partition is fully replicated again. You also watch that: - the restarted broker has rejoined the ISR for its partitions and - that there are no offline partitions. Only then is it safe to take down the next broker, because the cluster has recovered the redundancy you just spent. ## Edge cases 1. If a topic has replication factor 1, the partitions on the downed broker are simply unavailable during its restart — rolling upgrades only give zero downtime for replicated topics. 2. `min.insync.replicas` must be low enough that the cluster still meets it with one broker down (RF=3 + min.insync.replicas=2 is the canonical safe combo). 3. Controlled shutdown (`controlled.shutdown.enable=true`, default) makes the broker proactively move leadership off itself before stopping, shortening the unavailability window. ## Beyond binaries A rolling restart is also how you apply many broker config changes and, importantly, how you bump protocol/format versions during a *two-phase* version upgrade (covered separately).

  • What metric tells you it is safe to move on to the next broker?
    UnderReplicatedPartitions (kafka.server:type=ReplicaManager) back to 0, meaning every partition is fully replicated again; also confirm no offline partitions and the broker has rejoined its ISRs.
  • Why does a rolling upgrade not work for a topic with replication factor 1?
    There is only one copy of each partition, so when its broker goes down there is no in-sync replica to fail over to — those partitions are unavailable for the duration of that broker's restart.

saying these in an interview costs you the question

  • Saying you can restart all brokers simultaneously with no downtime.
  • Claiming rolling upgrades give zero downtime even for RF=1 topics.
  • Not waiting for under-replicated partitions to clear before moving to the next broker.
  • Confusing a rolling upgrade (one broker at a time) with simply restarting the cluster.

context

open as a page

Explain the classic ZooKeeper-era two-phase Kafka upgrade using inter.broker.protocol.version and log.message.format.version. Why must the binary upgrade and the protocol bump be separate steps?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Phase 1: roll new binaries while pinning inter.broker.protocol.version (and log.message.format.version) to the OLD version, so new and old brokers still speak the old protocol. Phase 2: after all brokers are upgraded, roll again removing/raising the pins. Splitting it keeps a mixed-version cluster compatible and makes phase 1 downgradable.

open as a page

How does Kafka's bidirectional client/broker compatibility work, and what does it mean for upgrading clients vs. brokers?

level: middleimportance: should knowfreq 55%

basics

~20 s

Each Kafka API (request type) is versioned. On connect, the client asks the broker which versions it supports (ApiVersions) and uses the highest both understand. Since ~0.10.2 this works both ways: newer clients talk to older brokers and older clients talk to newer brokers. Upgrade brokers and clients independently.

open as a page

What operational safeguards (controlled shutdown, health gating, acks/min.insync.replicas) make a rolling restart safe, and how do you pace it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Enable controlled shutdown so a broker moves leadership off itself before stopping. Between brokers, wait until UnderReplicatedPartitions is 0 and there are no offline partitions. Use replication factor 3 with min.insync.replicas 2 and acks=all so producers survive one broker being down.

open as a page

In a KRaft cluster, how do feature flags and metadata.version replace inter.broker.protocol.version, and how do you upgrade metadata.version during a rolling upgrade?

level: seniorimportance: should knowfreq 45%

basics

~20 s

KRaft has no ZooKeeper, so there is no inter.broker.protocol.version to set in server.properties. Instead the cluster has a feature flag named metadata.version that gates new behavior. You roll new binaries first (metadata.version stays put), then explicitly raise it with kafka-features.sh upgrade once all nodes are on the new version.

open as a page