One partition's leader broker shows much higher p99 than its peers even though cluster-wide throughput looks balanced. How would you diagnose and confirm a hot partition?
answer
- averages hide single-partition skew
- one leader handles all of a partition's traffic
- murmur2(key) % N -> skewed key = hot partition
- compare per-broker BytesIn/Out, then per-partition
- fix key / add partitions / reassign leadership
basics
~20 sA hot partition gets a disproportionate share of traffic, so its leader broker works harder and its tail latency rises while cluster averages look fine. Confirm by comparing per-partition/per-broker byte and message rates and checking the producer's partitioning key for skew.
solid answer
~40 sCluster-average balance can hide a single overloaded partition. A hot partition concentrates produce/fetch load on one leader broker, so that broker's request-handler threads, disk, and network do more work and its p99 climbs while the cluster mean stays flat. To diagnose: compare per-broker BytesInPerSec / BytesOutPerSec / MessagesInPerSec and find the outlier broker, then drill to per-topic/per-partition rates to find the hot partition. Confirm the cause is keyed-producer skew — a low-cardinality or skewed partition key (e.g. everything keyed by one tenant) routes most records to one partition via the default murmur2 partitioner. Cross-check that broker's RequestHandlerAvgIdlePercent (lower than peers), LocalTimeMs, and disk/network. Fixes: improve key distribution, increase partition count and rebalance, use a custom/sticky partitioner, or move the hot partition's leadership to a less-loaded broker.
go deeper
Know a hot partition gets too much traffic, that its single leader broker bears the load, and that key skew is the usual cause.
Diagnose by comparing per-broker then per-partition throughput, tie skew to murmur2(key) % N, and name the standard fixes.
Reason about ordering/partition-remap tradeoffs of each fix, distinguish produce vs consumer hotspots and hot-partition vs slow-broker, and confirm with local broker pressure metrics.
Set partitioning and key-design standards, balance leadership across brokers, and design for skew-resilience and capacity headroom.
## Why averages lie Monitoring dashboards often show **cluster-wide** or **per-broker average** throughput, which can look perfectly balanced while one *partition* is overloaded. Because every partition has exactly one **leader** broker that handles all of its produce and consumer-fetch traffic, a single hot partition concentrates load on one broker — and tail latency is a *local*, per-broker phenomenon. ## What makes a partition hot Kafka routes a produced record to a partition by: - **Keyed records:** `partition = murmur2(key) % numPartitions` (default partitioner). If keys are skewed or low-cardinality (e.g. all events keyed by a single big customer, or a null-ish constant key), most records land on one partition. - **Consumer side:** a partition with far more data, or a single slow/expensive consumer assignment, can be a read hotspot. Hot partitions can also arise from **uneven leadership distribution** (many partition leaders on one broker) even without key skew — but the classic case is producer key skew. ## Diagnosis workflow 1. **Find the outlier broker.** Compare `kafka.server:type=BrokerTopicMetrics,name=BytesInPerSec / BytesOutPerSec / MessagesInPerSec` across brokers. The hot broker stands out. 2. **Confirm local pressure on that broker.** Its `RequestHandlerAvgIdlePercent` is lower than peers, `LocalTimeMs` and possibly `RequestQueueTimeMs` higher, disk/network busier. 3. **Drill to the partition.** Per-topic-partition metrics (or `kafka-topics --describe` plus per-partition byte rates / log size growth) reveal which partition on that broker is taking the load. A partition whose log grows far faster than its siblings is the hot one. 4. **Confirm the cause.** Inspect the producer's keying: low-cardinality or skewed keys produce skew. A quick check is the distribution of message counts across partitions of the topic — a hot partition has a far larger count. ## Fixes - **Fix the key**: choose a higher-cardinality / better-distributed key, or hash differently. - **Increase partitions** and rebalance — but note adding partitions changes `murmur2(key) % N`, remapping keys (and breaking per-key ordering across the change). - **Custom or sticky partitioner** for null-key records (the modern sticky partitioner batches to one partition per batch to improve throughput, but for *balance* you may want round-robin/uniform). - **Reassign leadership** so the hot partition's leader sits on a less-loaded broker, or spread leadership with preferred-leader election. ## Edge cases - A hot partition can be *consumer-driven* (one partition repeatedly re-read or backfilled) rather than produce-driven; check BytesOut vs BytesIn. - Adding partitions to relieve a hot key only helps if the key cardinality exceeds the new partition count; one dominant key still maps to one partition. - Per-key ordering guarantees constrain your fix: you can't freely re-key data that consumers depend on being ordered per key. - Distinguish hot partition (load skew) from a slow broker (the broker is degraded but load is even) — the per-broker throughput comparison separates them.
- Why doesn't simply increasing the partition count always fix a hot partition?If one key dominates the traffic, murmur2(key) % N still maps that single key to exactly one partition regardless of N. More partitions only help when the skew comes from many keys colliding, i.e. when key cardinality exceeds the partition count.
- Which broker metric, compared across brokers, most directly reveals a hot partition?Per-broker BytesInPerSec / BytesOutPerSec / MessagesInPerSec (BrokerTopicMetrics). An outlier broker carrying far more traffic than peers, despite balanced cluster averages, points to a hot partition whose leader lives there.
saying these in an interview costs you the question
- Trusting cluster-average throughput and concluding the cluster is balanced
- Claiming adding partitions always fixes a hot key (a single dominant key still maps to one partition)
- Forgetting that rebalancing/adding partitions remaps murmur2 keys and can break per-key ordering
- Confusing a hot partition (skewed load) with a degraded broker (even load, sick node)