Walk through how you add a new controller voter to a running dynamic KRaft cluster and how you remove one.
answer
- format --no-initial-controllers ⇒ observer
- catch up on __cluster_metadata first
- add-controller ⇒ AddRaftVoter ⇒ VotersRecord commit
- remove-controller, THEN stop node
- one voter change at a time
basics
~20 sFormat the new controller with --no-initial-controllers so it joins as an observer, start it, let it catch up on the metadata log, then run kafka-metadata-quorum add-controller (AddRaftVoter) to promote it to a voter. To remove one, run remove-controller (RemoveRaftVoter), then shut the node down.
solid answer
~40 sAdding: provision the new node, run kafka-storage format --no-initial-controllers (and --cluster-id) so it starts as a non-voting observer, point it at controller.quorum.bootstrap.servers, and start it. It replicates the __cluster_metadata log to catch up. Once caught up, run kafka-metadata-quorum --bootstrap-controller ... add-controller, which issues AddRaftVoter with the node's id, directory.id, and endpoints; the leader appends a VotersRecord adding it, and after that record commits the node is a voter. Removing: run remove-controller (RemoveRaftVoter) with the target's id and directory.id; the leader commits a VotersRecord dropping it, after which you stop the node. Both operations change membership by exactly one voter at a time, each via a committed joint-consensus transition, so the quorum stays available throughout. You must keep a majority live during the change.
go deeper
Know that you add a controller by starting it then running add-controller, and remove with remove-controller.
Describe the observer-then-promote sequence and the AddRaftVoter/RemoveRaftVoter RPCs.
Detail --no-initial-controllers formatting, catch-up, commit semantics, and correct remove ordering.
Reason about availability math across the transition, one-change-at-a-time joint consensus, and runbook safety.
## Goal Grow or shrink the controller quorum without downtime and without editing config on every node. Dynamic reconfiguration (KIP-853) does this by writing membership changes into the metadata log itself. ## Adding a controller — step by step 1. **Provision** the host and pick a unique `node.id`. 2. **Format storage as an observer:** ``` kafka-storage format --cluster-id <id> \ --config server.properties \ --no-initial-controllers ``` `--no-initial-controllers` means the node does *not* assume it is part of the initial voter set; it boots as a **non-voting observer**. Formatting also generates its `directory.id` into meta.properties. 3. **Configure** `controller.quorum.bootstrap.servers` to point at existing controllers, set `process.roles=controller` (or `broker,controller`), and start the node. 4. **Catch up:** as an observer it replicates `__cluster_metadata` from the leader until its log nearly matches. You can watch lag with `kafka-metadata-quorum describe --status` (LastFetchTimestamp / log end offset). 5. **Promote to voter:** ``` kafka-metadata-quorum --bootstrap-controller host:9093 \ add-controller ``` This sends **AddRaftVoter** with the node's id, directory.id, and endpoints. The leader appends a **VotersRecord** that includes the new voter. Once that record is **committed** (replicated to a majority of the *new* configuration), the node is a full voter. ## Removing a controller 1. Run: ``` kafka-metadata-quorum --bootstrap-controller host:9093 \ remove-controller --controller-id <id> \ --controller-directory-id <dirId> ``` This sends **RemoveRaftVoter**. The leader commits a VotersRecord dropping that voter. 2. After the change commits, **shut down** the node. (Shutting it down first, while still a voter, shrinks your live majority and is the wrong order.) ## One change at a time Membership changes are applied **single-server at a time** (add one or remove one). Each change is a **joint-consensus** transition that requires agreement from both the old and new voter sets before committing, which guarantees no two disjoint majorities can form. You cannot batch a +2/-2 swap as one atomic step; do them sequentially. ## Availability math Keep a majority alive throughout. Example: going 3→5, after adding the 4th voter a majority is 3 of 4; after the 5th it is 3 of 5. When removing, remove before stopping so you never drop below majority of the *current* voter set. ## Common pitfalls - Forgetting `--no-initial-controllers` (the node tries to bootstrap its own quorum). - Promoting before the observer has caught up (a far-behind voter weakens availability until it catches up). - Stopping a node before remove-controller commits. - Passing the wrong directory.id to remove-controller (it must match the voter set entry).
- Why must a new controller catch up as an observer before being promoted?A far-behind voter still counts toward the majority needed to commit, so promoting it early temporarily weakens availability and slows commits until it catches up.
- What is the correct order when removing a voter, and why?Run remove-controller and let it commit first, then shut the node down. Stopping it while still a voter shrinks the live majority and can stall the quorum.
- Can you swap two controllers in a single atomic reconfiguration?No — changes are applied one voter at a time. Add the replacement, let it commit, then remove the old one (or vice versa).
saying these in an interview costs you the question
- Editing controller.quorum.voters / bootstrap.servers on every node to change membership — that's the static model.
- Adding and removing multiple voters in one step.
- Stopping a controller before issuing remove-controller.
- Promoting an observer to voter before it has replicated the metadata log.