skip to content

A three-broker Kafka cluster on Kubernetes rejects acks=all writes during routine node drains, although a PodDisruptionBudget allows only one unavailable broker. How can that happen, and what do you change?

level: seniorimportance: should knowfreq 41%

answer

  1. Ready is not in sync
  2. budget counts pods, not replicas
  3. second eviction too soon
  4. RWO detach wait stretches window
  5. wait for zero under-replicated partitions

basics

~20 s

A PodDisruptionBudget counts Ready pods, not in-sync replicas. A restarted broker can be Ready while still catching up, so a second eviction is allowed. Slow volume reattachment stretches that window. Gate restarts on replica health and serialise drains.

solid answer

~50 s

With replication factor 3 and `min.insync.replicas` 2, `acks=all` producers survive one broker out of the in-sync set, not two. A PodDisruptionBudget with `maxUnavailable: 1` only counts **Ready pods**. After the first drain, the replacement broker turns Ready once its process answers, which can be long before it has rejoined the in-sync set. The budget then allows the next drain to evict a second broker, and two replicas are now out of sync. The **ReadWriteOnce** volume adds time: the new pod cannot start until the attach/detach controller has detached the volume from the old node, and if that node dies before unmounting, the controller waits up to six minutes before force-detaching. The fixes: let restarts be decided on replica health (Strimzi's operator-driven restarts do this), drain one node at a time and wait for zero under-replicated partitions, and spread brokers across nodes and zones.

code

yaml · 20 lines
yaml
apiVersion: v1
kind: Pod
metadata:
  name: ocr-kafka-broker-1
  labels:
    app: ocr-kafka
spec:
  topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: ocr-kafka
  containers:
  - name: kafka
    image: registry.example.com/ocr/kafka:3.9
    resources:
      requests:
        cpu: 350m

go deeper

for a junior

Remember that Kubernetes counts Ready pods, while Kafka counts in-sync replicas, and the two can disagree.

for a middle

Walk through the timeline where a second eviction is allowed too early, and do the arithmetic for replication factor 3 and min.insync.replicas 2.

for a senior

Diagnose from signals: under-replicated partitions, FailedAttachVolume events, drain timing. Change the maintenance process to wait on replica health rather than only the budget.

for a principal

Decide how maintenance automation should learn a data service's real health, and whether the node fleet and storage topology make these services worth running in-cluster.

## The setup The OCR pipeline on a 3-control-plane, 27-worker self-managed cluster publishes page events to a Kafka topic with **replication factor 3** and **`min.insync.replicas` 2**. Producers use `acks=all`, so a write succeeds only when at least two replicas in the **in-sync replica set (ISR)** have it. Losing one broker is fine. Two replicas out of the ISR at once means `acks=all` writes fail. The platform team has a PodDisruptionBudget allowing one unavailable broker and drains nodes in a maintenance pipeline. Writes still fail during some maintenance windows. This question is about how Kubernetes' view of "available" differs from the data service's view. The budget's fields and the drain procedure belong elsewhere. ## Why the budget is not enough A **PodDisruptionBudget** limits **voluntary evictions** made through the Eviction API, which `kubectl drain` uses. It decides based on how many selected pods are currently **healthy**, meaning Ready. Kafka's safety depends on something else: how many replicas are **in sync**. The gap opens like this: 1. The drain evicts `broker-1`. The budget now has no room. 2. The replacement `broker-1` pod starts, its readiness probe (a TCP check on the listener, say) passes, and the budget has room again. 3. `broker-1` is still copying data from the partition leaders. It is **Ready but not in sync**. 4. The pipeline drains the next node and evicts `broker-2`. The budget allows it. 5. For partitions whose replicas include `broker-1` and `broker-2`, only one in-sync replica is left, below `min.insync.replicas`. `acks=all` writes to those partitions fail. | Signal | What it measures | Protects Kafka writes? | |---|---|---| | Pod Ready | The process answers its probe | No | | PodDisruptionBudget | Count of Ready pods | Only if Ready means in sync | | Under-replicated partitions = 0 | All replicas caught up | Yes | | Operator's restart check | ISR against `min.insync.replicas` | Yes, for restarts it performs | ## How ReadWriteOnce storage stretches the window Each broker's volume is typically a **ReadWriteOnce** block device, which can be attached to one node at a time. When the pod moves: - The new pod cannot start until the **attach/detach controller** in kube-controller-manager has detached the volume from the old node and attached it to the new one. A `FailedAttachVolume` event on the pod shows it is waiting. - If the old node goes away **before** its kubelet has unmounted the volume, recovery is slower. The old pod may not finish terminating on its own, and the controller waits up to **six minutes** for an unmount before force-detaching from an unhealthy node, and does not force-detach at all if `--disable-force-detach-on-timeout` is set. - A cloud block volume is often **zonal**, so the pod can only be scheduled into the same zone. If that zone has no room, the pod stays Pending. Every minute here is a minute with one fewer in-sync replica, and more data to copy once the broker is back. ## What to change 1. **Let the data service decide.** Use operator-driven restarts (Strimzi checks `min.insync.replicas` before rolling a broker) instead of bare evictions. A pod-deleting script bypasses those checks. 2. **Serialise and gate drains.** Drain one node, then wait until under-replicated partitions are zero before starting the next. A budget alone does not wait for that. 3. **Make readiness mean more,** where the software supports it, so a broker is not Ready until it can serve its role. Be careful: a probe that fails during a long catch-up can also block other rollouts. 4. **Spread the members.** Use topology spread constraints or anti-affinity across nodes and zones so no single node or zone holds two of the three brokers. 5. **Handle dead nodes explicitly.** Apply the `node.kubernetes.io/out-of-service` taint once a node is confirmed down, so volumes are detached and pods moved without waiting out a timeout. 6. **Consider more headroom.** A larger replication factor raises tolerance, at a cost in storage and latency. The same reasoning applies to PostgreSQL or etcd-style quorums: find the database's real availability signal and make the disruption process wait on it.

  • Why not just set the Kafka PodDisruptionBudget to `maxUnavailable: 0`?
    Then no eviction is ever allowed, so every drain blocks until someone intervenes. That only works if something else performs the restart safely. That pattern exists: a drain helper intercepts the eviction and asks the operator to roll the broker with its own replica checks. Without such a helper, a zero budget just turns maintenance into a manual process.
  • A node running a broker lost power. The replacement pod has been waiting for its volume for several minutes. What do you do?
    Once the node is confirmed down and will not return by itself, apply the `node.kubernetes.io/out-of-service` taint. Kubernetes then force-deletes the pods on it and detaches their volumes, so the new pod can attach the volume elsewhere. Do not do this while the node might still be writing, because two writers on one block device can corrupt it.

saying these in an interview costs you the question

  • A PodDisruptionBudget guarantees Kafka keeps enough in-sync replicas.
  • A Ready broker pod has fully caught up with its partitions.
  • A ReadWriteOnce volume moves to the new node instantly.
  • Deleting pods directly is equivalent to draining with a budget.
  • Replication factor 3 always survives two brokers down for acks=all writes.