skip to content

What is an ephemeral znode, and how did Kafka use ephemeral znodes and ZooKeeper sessions to detect a dead broker?

level: middleimportance: should knowfreq 55%

answer

  1. ephemeral = dies with the session
  2. /brokers/ids/<id> beacon
  3. zookeeper.session.timeout.ms ~18s
  4. controller watches, re-elects from ISR
  5. GC pause can false-expire

basics

~20 s

An ephemeral znode is a ZooKeeper node that exists only while the client's session is alive; it auto-deletes when the session ends. Each broker created one under /brokers/ids. When the broker's session timed out, the znode vanished, signaling the broker was dead.

solid answer

~50 s

A znode is a node in ZooKeeper's tree. An **ephemeral** znode is tied to the lifetime of the ZK client **session** that created it — when that session ends (clean disconnect or timeout), ZK automatically deletes the node. Each Kafka broker, on startup, created an ephemeral znode at `/brokers/ids/<broker.id>` holding its endpoints. The broker kept its ZK session alive with periodic heartbeats; if it stopped heartbeating for longer than the negotiated session timeout (driven by `zookeeper.session.timeout.ms`, default 18s in Kafka), ZK expired the session and deleted the ephemeral znode. The controller had a **watch** on `/brokers/ids`, so the deletion fired a notification; the controller then treated the broker as failed and triggered leader re-election for partitions it led. The same ephemeral mechanism backed `/controller`: when the controller died, its znode vanished and another broker won the election.

go deeper

for a junior

Know ephemeral znode = auto-deleted when the broker's session ends, signaling it's dead.

for a middle

Explain session timeout, heartbeats, the /brokers/ids beacon, and controller watch-driven re-election.

for a senior

Discuss GC-pause false positives and the timeout tradeoff, plus why ISR tracking complements ZK liveness.

for a principal

Reason about failure-detector semantics (server-side expiry, partition scenarios) and how this design influenced KRaft's broker heartbeat model.

## znodes and their flavors ZooKeeper stores data in a tree of **znodes** (like files/dirs). Each znode can be: - **persistent** — stays until explicitly deleted. - **ephemeral** — bound to the creating client's **session**; ZK deletes it automatically when that session ends. Ephemeral znodes cannot have children. - (and **sequential**, which appends a monotonic counter — orthogonal to ephemeral.) ## ZooKeeper sessions A ZK client opens a **session** with the ZK ensemble. The session has a negotiated **timeout**. The client must send heartbeats (pings) within that window. If ZK hears nothing for the timeout, it declares the session **expired** and removes everything that session owned — including its ephemeral znodes. Importantly, session expiry is decided by the **server**, not the client; a partitioned client may think it's connected while ZK has already expired it. ## How Kafka used this for failure detection 1. On startup, a broker connects to ZK and creates an **ephemeral** znode `/brokers/ids/<id>` with its listener endpoints. This is the broker's 'I'm alive' beacon. 2. The broker maintains its ZK session via heartbeats. `zookeeper.session.timeout.ms` (Kafka default 18000ms) bounds how long a silent broker stays 'alive'. 3. The **controller** registers a **watch** on the children of `/brokers/ids`. 4. If the broker crashes, GC-pauses, or is network-partitioned long enough, ZK expires its session and deletes the ephemeral znode. 5. That deletion fires the controller's watch. The controller marks the broker offline and re-elects leaders for every partition that broker led (choosing a new leader from the ISR), then propagates the new metadata to other brokers. ## Controller election uses the same trick `/controller` is itself an ephemeral znode. Brokers race to create it; the winner is the controller. If the controller dies, its session expires, `/controller` is deleted, and the surviving brokers race again. This is a classic ZK leader-election pattern. ## Edge cases / gotchas - **GC pauses / long STW**: a long stop-the-world pause can blow the session timeout and falsely mark a healthy broker dead, triggering needless leader churn. Tuning the timeout trades faster failure detection against false positives. - **Watches are one-shot**: a watch fires once and must be re-registered, so Kafka re-set watches after each notification. - **Session vs connection**: a brief TCP disconnect is recoverable if the client reconnects before the timeout; only full session expiry deletes ephemerals. - **Soft vs hard failure**: ZK gives liveness via session, but the broker could be alive-but-slow; that's why Kafka also tracks ISR via replica fetch progress, not just ZK liveness.

  • Why can a long GC pause cause a healthy broker to be marked dead?
    Heartbeats stop during a stop-the-world pause. If the pause exceeds zookeeper.session.timeout.ms, ZK expires the session and deletes the ephemeral znode, so the controller treats the broker as failed and re-elects leaders — even though the broker recovers moments later.
  • What's the difference between a ZK session expiring and just a TCP disconnect?
    A TCP disconnect is recoverable: if the client reconnects to any ensemble member before the session timeout, the session and its ephemeral znodes survive. Only true session expiry (server-decided, after the timeout) deletes ephemeral znodes.

saying these in an interview costs you the question

  • Saying the broker deletes its own znode on crash — a crashed broker can't act; ZK auto-deletes it on session expiry.
  • Treating watches as persistent subscriptions — they are one-shot and must be re-registered.
  • Confusing connection loss with session expiry; short disconnects don't drop ephemeral znodes.

context