What operational safeguards (controlled shutdown, health gating, acks/min.insync.replicas) make a rolling restart safe, and how do you pace it?
answer
- controlled.shutdown.enable=true (default)
- gate: UnderReplicatedPartitions=0, no offline
- RF=3 + min.insync.replicas=2 + acks=all
- one broker at a time, let handoff finish
- throttle replication catch-up on rejoin
basics
~20 sEnable controlled shutdown so a broker moves leadership off itself before stopping. Between brokers, wait until UnderReplicatedPartitions is 0 and there are no offline partitions. Use replication factor 3 with min.insync.replicas 2 and acks=all so producers survive one broker being down.
solid answer
~50 sThree safeguards make a rolling restart safe. (1) **Controlled shutdown** (`controlled.shutdown.enable=true`, the default): before a broker stops it asks the controller to migrate its partition leaderships to in-sync replicas, so clients fail over cleanly instead of timing out on a dead leader. (2) **Health gating between steps**: after each restart, wait until the cluster is whole again — `UnderReplicatedPartitions=0`, no offline partitions, the broker rejoined its ISRs — before taking down the next broker. (3) **Durability config**: replication factor ≥3 with `min.insync.replicas=2` and producers using `acks=all`, so with exactly one broker down every partition still has 2 in-sync replicas and `acks=all` writes still succeed. Pacing: do strictly one broker at a time, allow controlled shutdown to complete (watch `controlled.shutdown.max.retries`), restart, gate on health, repeat. Optionally throttle replication catch-up to avoid saturating the network when the broker rejoins.
go deeper
Know to restart one broker at a time and wait until partitions are fully replicated before continuing.
Know controlled shutdown, the health-gate metrics, and the RF=3/min.insync.replicas=2/acks=all combo.
Reason about the availability math, replication throttling, unclean leader election, and rebalance/coordinator effects.
Codify the safeguards into automated, health-gated upgrade orchestration with throttling and rack-aware pacing.
**The risk a rolling restart must manage.** Each broker you stop is leading some partitions and is a replica for others. If you stop it carelessly, (a) its lead partitions briefly have no leader (clients error until failover), and (b) the cluster temporarily loses one copy of every partition it held. The safeguards below shrink the unavailability window and prevent stopping the *next* broker before redundancy is restored. **1. Controlled shutdown.** With `controlled.shutdown.enable=true` (the **default**), a broker about to stop sends a request to the **controller** asking it to move leadership for the broker's lead-partitions to other in-sync replicas *first*. Only after leadership has migrated does the broker exit. This turns a hard leader failure (clients time out, then rediscover) into a graceful handoff (clients are redirected almost immediately). `controlled.shutdown.max.retries` and `controlled.shutdown.retry.backoff.ms` tune how hard it tries. Always confirm controlled shutdown actually completed in the broker log; if it failed, leadership failover still happens but less gracefully. **2. Health gating between brokers.** The cardinal rule: never stop broker N+1 until the cluster has fully recovered from stopping broker N. Concretely, gate on: - **`UnderReplicatedPartitions` == 0** (kafka.server:type=ReplicaManager) — every partition is fully replicated. - **No offline partitions** (`OfflinePartitionsCount` == 0) — nothing is leaderless. - The restarted broker has **rejoined the ISR** for its partitions (it finished catching up). Automations often poll these JMX metrics or `kafka-topics.sh --describe --under-replicated-partitions` and block until clean. **3. Durability configuration so producers survive.** The availability math: with replication factor **3** and **`min.insync.replicas=2`**, a partition needs at least 2 in-sync replicas to accept `acks=all` writes. With exactly one broker down, every partition still has 2 in-sync replicas, so `acks=all` producers keep succeeding. If you instead ran RF=2 with min.insync.replicas=2, taking one broker down would drop a partition to 1 ISR and **block all acks=all writes** to it — a self-inflicted outage during the upgrade. So RF=3 + min.insync.replicas=2 + acks=all is the canonical safe combination. **Pacing the roll.** - Strictly **one broker at a time**. - Let **controlled shutdown** finish before the process exits. - Restart, then **gate on health metrics** before continuing. - When the broker rejoins it must **catch up** its replicas; this generates replication traffic. On large clusters you can cap it with **replication throttling** (`leader.replication.throttled.rate` / `follower.replication.throttled.rate` set via `kafka-configs.sh`) so catch-up doesn't saturate the NICs and starve client traffic. - For the **controller/KRaft controllers**, upgrade them following the same one-at-a-time + health-gate discipline; ensure the controller quorum stays available. **Edge cases.** (1) Unclean leader election should stay **disabled** (`unclean.leader.election.enable=false`) so you never promote an out-of-sync replica and lose data during the churn. (2) Watch consumer group rebalances — restarting brokers that host group coordinators triggers rebalances; pacing avoids a rebalance storm. (3) Rack awareness: if replicas are rack-spread, ensure you're not down-ing the only copy in a rack for partitions that matter. **Summary mental model.** Controlled shutdown shrinks the *unavailability window*; health gating ensures *redundancy is restored before you spend it again*; RF=3/min.insync.replicas=2/acks=all keeps *durability and producer availability intact* throughout.
- Why is RF=2 with min.insync.replicas=2 dangerous during a rolling restart?Taking one broker down drops each affected partition to a single in-sync replica, which is below min.insync.replicas=2, so all acks=all producers to those partitions get blocked (NotEnoughReplicas) — a self-inflicted write outage. RF=3 keeps 2 ISRs with one broker down.
- What does controlled shutdown actually do that a plain kill -9 doesn't?It asks the controller to migrate the broker's partition leaderships to in-sync replicas before the process exits, so clients are redirected to new leaders gracefully instead of timing out on a dead leader and then rediscovering.
saying these in an interview costs you the question
- Stopping the next broker before under-replicated partitions return to zero.
- Enabling unclean leader election to 'speed up' failover during the roll.
- Running RF=2 with min.insync.replicas=2 and expecting acks=all to keep working with a broker down.
- Using kill -9 instead of a graceful controlled shutdown.