Should KRaft controllers run on the same nodes as brokers, and why do production deployments isolate them?
answer
- process.roles: controller | broker | broker,controller
- combined = dev/small only
- isolate so broker load can't trip elections
- co-location couples failure domains
- 3 or 5 small controllers, fast disk, AZ-spread
basics
~20 sIn production, run controllers as dedicated nodes (process.roles=controller), separate from brokers (process.roles=broker). Combined-mode nodes are fine for dev but mean broker load can starve the metadata quorum and a failure takes out both roles at once.
solid answer
~50 sA KRaft node's `process.roles` can be `controller`, `broker`, or `broker,controller` (combined mode). Combined mode is convenient for development and small clusters but is discouraged for production. Dedicated controllers isolate the metadata quorum from data-plane load: heavy produce/consume traffic, large fetch requests, GC pauses, and page-cache pressure on a broker can delay the controller's Raft heartbeats/fetches and trigger spurious elections or slow metadata commits. Co-locating also couples failure domains — losing one node removes both a broker and a voter, eroding quorum tolerance faster. Dedicated controllers are typically small (modest CPU/RAM, fast disk for the metadata log) and you run 3 or 5 of them, ideally across availability zones, while scaling brokers independently. Combined mode also has operational limits: you generally can't cleanly migrate roles, and rolling restarts are riskier because each restart removes a voter and a broker simultaneously.
go deeper
Know process.roles can be broker, controller, or both; production uses dedicated controllers.
Explain resource interference and that combined nodes couple a broker and a voter failure.
Discuss spurious elections from data-plane load, independent scaling, and AZ placement of 3/5 dedicated controllers.
Reason about control-plane/data-plane separation as an architecture principle, capacity-plan the quorum vs broker fleet, and the operational/upgrade implications.
## The three role configurations Every KRaft node sets `process.roles`: - `controller` — a dedicated voter: participates in the metadata Raft quorum, does not serve client produce/consume traffic. - `broker` — a dedicated data node: stores partitions, serves clients, fetches metadata from the controllers but does not vote. - `broker,controller` — **combined mode**: one process does both. ## Why combined mode exists Combined mode lets you stand up a working KRaft cluster on as few as one or three machines — ideal for laptops, CI, demos, and small non-critical clusters. It is fully functional, just operationally constrained. ## Why production isolates controllers ### 1. Resource isolation / interference Brokers do heavy I/O: large sequential writes, fetch requests serving consumers, page-cache churn, and occasionally long GC pauses under load. The controller's job — keeping the Raft metadata log healthy and sending/serving timely replication — is latency-sensitive. If a combined node is busy serving a traffic spike, the controller side can miss its `controller.quorum.fetch.timeout.ms` window, causing **spurious leader elections** or delayed metadata commits that ripple across the cluster. Dedicated controllers keep the control plane quiet and predictable. ### 2. Failure-domain separation With combined nodes, one machine failure removes a broker **and** a voter at once. In a 3-node combined cluster, losing one node simultaneously drops you to a 2-voter majority-of-2 (no remaining tolerance) and removes a data node. Separating roles means a broker loss doesn't touch quorum headroom and vice versa. You can also place the (few) controllers carefully across AZs without that constraint dictating broker placement. ### 3. Independent scaling Quorums don't scale with cluster size — you want 3 or 5 voters regardless of whether you have 5 brokers or 500. Dedicated controllers let you fix the quorum at 3/5 small nodes and scale brokers freely. In combined mode, every broker you add would (if also a controller) bloat the quorum, and if not, you'd have an awkward mix. ### 4. Operational safety Rolling restarts and upgrades: restarting a dedicated controller affects only quorum headroom briefly; restarting a combined node disrupts both planes. Role migration (e.g., separating roles later) is also cleaner when they started separate. Kafka's own guidance recommends dedicated controllers for production. ## Sizing the dedicated controllers Controllers need: modest CPU and memory, but a **fast, reliable disk** for the metadata log (it's on the commit path of every metadata change), and low-latency networking to each other and to brokers. They don't need the large disks brokers use for partition data. Run 3 (tolerate 1) or 5 (tolerate 2), spread across failure domains. ## Key takeaway Combined mode = dev/small. Production = dedicated controllers (3 or 5, isolated, AZ-spread) so data-plane load and failures can't destabilize the metadata quorum, and so you can scale and operate the two planes independently.
- What's the risk of broker GC pauses or traffic spikes on a combined controller+broker node?The data-plane work can delay the controller's Raft fetch/heartbeat past controller.quorum.fetch.timeout.ms, triggering spurious leader elections and slowing metadata commits — destabilizing the control plane under exactly the load you most need it stable.
- Why don't controllers scale with the number of brokers?Quorum availability is a function of voter count, not cluster size — you want 3 or 5 voters whether you have 5 or 500 brokers. So you fix a small dedicated quorum and scale brokers independently.
saying these in an interview costs you the question
- Recommending combined mode for large production clusters.
- Saying you should add a controller every time you add a broker (the quorum stays 3/5).
- Claiming controllers need big disks like brokers (they need a fast metadata-log disk, not bulk storage).
- Assuming co-location has no downside because 'it's the same Raft either way' — it couples load and failure domains.