Explain the dual-write phase and how brokers are rolled from ZooKeeper mode to KRaft mode during migration.
answer
- dual-write = controller writes KRaft + ZK
- initial snapshot copy then ongoing mirror
- mixed mode: ZK + KRaft brokers coexist
- roll brokers one at a time, drop zookeeper.connect
- ZK mirror = rollback safety
basics
~20 sIn dual-write, the active KRaft controller writes every metadata change to both KRaft and ZooKeeper, keeping them in sync. Then brokers are restarted one at a time into KRaft mode (process.roles=broker, no zookeeper.connect) until all are migrated.
solid answer
~50 sOnce the migration controllers are up and the brokers have registered, the KRaft controller becomes the active controller and the cluster enters the **dual-write phase**. The controller copies the snapshot of existing ZK metadata into the KRaft __cluster_metadata log, then for every subsequent metadata change (topic create, ISR shrink, config update, ACL change) it writes to *both* KRaft and ZooKeeper. ZK stays a faithful mirror, which is what preserves the rollback window. While in dual-write, brokers still in ZK mode keep working normally. You then perform a **rolling restart of the brokers**: for each broker you remove zookeeper.connect and the migration flag, set process.roles=broker and the KRaft controller.quorum.voters, and restart it as a pure KRaft broker. The cluster runs in a mixed state (some ZK-mode, some KRaft-mode brokers) until every broker is rolled. Only after all brokers are KRaft brokers do you finalize.
go deeper
Know that dual-write means writing to both KRaft and ZooKeeper.
Describe the snapshot copy plus the one-at-a-time broker rolling restart.
Explain mixed-mode operation, why ZK mirroring preserves rollback, and the exact broker config flip.
Reason about availability during the roll (ISR/min.insync.replicas), ZK-write health monitoring, and abort criteria.
## The dual-write phase — what it actually does After the migration controllers start and the (still-ZK) brokers register with them, the KRaft quorum elects an **active controller** and the migration transitions into **dual-write**. Two things happen: 1. **Initial copy.** The active controller reads the full metadata snapshot from ZooKeeper (all topics, partition assignments, ISRs, configs, ACLs, client quotas, delegation tokens) and writes it into the KRaft `__cluster_metadata` log. After this, KRaft holds a complete copy. 2. **Ongoing dual-write.** From now on, *every* metadata mutation flows through the KRaft controller, which commits it to the KRaft log **and** writes it back into ZooKeeper. ZooKeeper is therefore kept as a live, up-to-date mirror of KRaft. Why mirror ZK? Because the brokers haven't all moved yet, and because keeping ZK current is precisely what makes **rollback** safe. If you abort during dual-write, ZK still reflects reality and you can revert brokers and controllers to ZK mode. ## Mixed-mode operation During dual-write the cluster is intentionally heterogeneous: - **ZK-mode brokers** still read some state from ZK but now receive controller RPCs from the KRaft controller. - **KRaft-mode brokers** (once rolled) consume the `__cluster_metadata` log directly. The KRaft controller speaks to both, which is the whole point of dual-write — it bridges the two worlds so the rolling restart can proceed gradually with no downtime. ## Rolling the brokers For each broker, one at a time (respecting replication so partitions stay available): 1. Stop the broker. 2. Change its config: - `process.roles=broker` - `controller.quorum.voters=<the KRaft controllers>` and `controller.listener.names` - **Remove** `zookeeper.connect` and **remove** `zookeeper.metadata.migration.enable`. - Keep the same `node.id` (it carries over) and the same data directories. 3. Start the broker — it now operates as a pure KRaft broker consuming the metadata log. Because you roll one broker at a time and wait for ISRs to recover between restarts, producers and consumers see no outage (assuming replication factor >= 2/3 and `min.insync.replicas` is satisfied). ## Edge cases and gotchas - **Don't finalize early.** If you flip the controllers out of migration mode before *every* broker is in KRaft mode, the remaining ZK-mode brokers are orphaned. - **Watch under-replicated partitions** between restarts; rolling too fast can drop a partition below `min.insync.replicas`. - **Dual-write lag/errors** to ZooKeeper should be monitored (the controller exposes ZK-write metrics); persistent ZK write failures mean rollback is no longer clean. - **node.id continuity** matters — the broker keeps its identity and log directories across the mode switch.
- Why does the controller keep writing to ZooKeeper during dual-write instead of stopping immediately?To keep ZK a faithful mirror so rollback stays safe and so any still-ZK-mode brokers continue to function until they are rolled into KRaft mode.
- What config changes flip a broker from ZK mode to KRaft mode during the roll?Set process.roles=broker and controller.quorum.voters/controller.listener.names; remove zookeeper.connect and zookeeper.metadata.migration.enable. Keep the same node.id and data dirs.
saying these in an interview costs you the question
- Saying ZooKeeper stops being written to as soon as dual-write starts — it is actively mirrored throughout.
- Restarting all brokers at once instead of a controlled rolling restart.
- Finalizing the migration before every broker is in KRaft mode.
- Generating new node.ids for brokers during the roll instead of preserving identity.