skip to content

What is LinkedIn Cruise Control, and how do its goals and self-healing automate cluster balancing?

level: seniorimportance: should knowfreq 45%

answer

  1. load model from metrics reporter
  2. hard goals must hold; soft goals optimized by priority
  3. RackAware/Capacity = hard; Distribution = soft
  4. anomaly detector → self-healing
  5. add/remove/demote broker via REST

basics

~20 s

Cruise Control is a tool that continuously monitors a Kafka cluster's load and automatically generates and executes reassignment plans to keep it balanced. It optimizes against a prioritized list of 'goals' (hard and soft constraints) and can self-heal by reacting to broker failures or anomalies.

solid answer

~40 s

kafka-reassign-partitions is manual and naive; Cruise Control (open-sourced by LinkedIn) automates rebalancing. It collects per-partition/broker load metrics, builds a workload model, and runs an optimizer over a prioritized list of goals — hard goals (e.g. RackAwareGoal, ReplicaCapacityGoal, Disk/Network/CpuCapacityGoal) that must be satisfied, and soft goals (e.g. ReplicaDistributionGoal, LeaderReplicaDistributionGoal, resource-distribution goals) optimized in priority order. It produces an optimization proposal (a reassignment plan) and can execute it with built-in throttling and concurrency limits. Self-healing watches for anomalies — broker failures, goal violations, metric anomalies, disk failures — and automatically triggers remediation (e.g. move replicas off a dead broker). It also supports add/remove-broker and rebalance operations via REST/UI, making capacity changes and balancing hands-off, unlike the all-manual CLI workflow.

go deeper

for a junior

Know Cruise Control automatically keeps a Kafka cluster balanced so you don't run the CLI by hand.

for a middle

Explain goals (balance replicas/leaders/resources) and that it can react to broker failures.

for a senior

Distinguish hard vs soft goals with concrete names and describe the load model + self-healing anomaly types.

for a principal

Design goal priority policy and self-healing scope for a large multi-rack cluster, weighing safety vs feasibility.

## Why Cruise Control exists The built-in `kafka-reassign-partitions` tool is **manual** and **dumb about load**: it round-robins replicas without knowing actual CPU, disk, or network usage, and it does nothing continuously. On large clusters you need *ongoing* balancing and *automatic* reaction to failures. **Cruise Control** (open-sourced by LinkedIn) provides this. ## Architecture in brief - A **Metrics Reporter** runs on each broker and emits load metrics to a Kafka topic. - Cruise Control's **Load Monitor** aggregates these into a **cluster workload model** (per-partition byte rates, per-broker resource utilization). - The **Analyzer** runs an optimization over **goals** to produce an **optimization proposal**. - The **Executor** applies the proposal as a series of throttled, concurrency-limited reassignments. - The **Anomaly Detector** drives **self-healing**. ## Goals: hard vs soft A **goal** is a constraint/objective. Goals are evaluated in a **priority order** (configurable). They split into: **Hard goals** — must be satisfied or the proposal is rejected. Examples: - `RackAwareGoal` — replicas of a partition spread across racks. - `ReplicaCapacityGoal`, `DiskCapacityGoal`, `NetworkInboundCapacityGoal`, `NetworkOutboundCapacityGoal`, `CpuCapacityGoal` — no broker exceeds capacity. **Soft goals** — optimized as well as possible, in priority order, without violating hard goals. Examples: - `ReplicaDistributionGoal` — even replica counts per broker. - `LeaderReplicaDistributionGoal` / `LeaderBytesInDistributionGoal` — even leadership/leader load. - `DiskUsageDistributionGoal`, `NetworkInbound/OutboundUsageDistributionGoal`, `CpuUsageDistributionGoal` — even resource utilization. The optimizer satisfies all hard goals first, then improves soft goals by priority, each goal respecting the ones above it. ## Self-healing The **Anomaly Detector** continuously checks for: - **Broker failures** (a broker leaves the cluster) → move its replicas elsewhere. - **Goal violations** (cluster drifts out of balance) → rebalance. - **Metric anomalies** (sudden load spikes). - **Disk failures** (JBOD volume dies). With self-healing enabled, detected anomalies automatically trigger remediation proposals + execution — no human in the loop. You can enable it per anomaly type. ## Operations exposed Via REST/UI: `rebalance`, `add_broker`, `remove_broker` (drain), `demote_broker` (move leadership off without removing replicas), `fix_offline_replicas`. All use throttling and `--concurrent`/concurrency caps to bound impact. ## Edge cases / tradeoffs - Goal **priority ordering** is policy: e.g. putting rack-awareness as hard prevents unsafe moves, but a tight capacity goal can make a balanced solution infeasible. - Self-healing can mask a failing broker — pair with alerting. - Cruise Control still executes ordinary Kafka reassignments under the hood, so the same throttle dynamics apply.

  • What is the difference between a hard goal and a soft goal in Cruise Control?
    A hard goal must be satisfied or the optimization proposal is rejected (e.g. rack-awareness, capacity limits). Soft goals are optimized as much as possible in priority order without violating any hard goal (e.g. even replica/leader/resource distribution).
  • How does Cruise Control's self-healing react to a broker failure, and what is the risk?
    The anomaly detector notices the broker left and automatically generates and executes a reassignment moving its replicas to surviving brokers. The risk is masking a recurring hardware problem, so it should be paired with alerting rather than relied on silently.

saying these in an interview costs you the question

  • Describing Cruise Control as just a wrapper that runs the same naive plan as the CLI — it builds a real load model and optimizes against goals.
  • Saying all goals are equal — they are prioritized, and hard vs soft is a fundamental distinction.
  • Claiming self-healing is always on by default for everything — it is configured per anomaly type.

context