skip to content

How does a Kafka broker know on startup whether it was shut down cleanly, and why does that distinction matter?

level: juniorimportance: should knowfreq 35%

answer

  1. .kafka_cleanshutdown marker, per log dir
  2. written on graceful stop, deleted on startup
  3. present = skip recovery (fast)
  4. absent = unclean = run recovery
  5. controlled.shutdown.enable + SIGTERM

basics

~20 s

On a clean shutdown the broker flushes all data, updates the checkpoint files, and writes a clean-shutdown marker file in each log dir. On startup, if that marker is present it skips log recovery; if it is missing (a crash), it runs recovery.

solid answer

~40 s

During an orderly shutdown the broker flushes every log, advances the `recovery-point-offset-checkpoint` to the log end, and writes a marker file named `.kafka_cleanshutdown` into each log directory. On the next startup the broker checks for this marker in each directory: if present, it deletes the marker and trusts the checkpoints, so startup is fast and no log re-validation is needed. If the marker is absent — because the process was SIGKILLed, OOM-killed, or the host crashed/lost power — the broker assumes an unclean shutdown and performs recovery: it scans each affected partition from its recovery point to the log end, re-validating records and rebuilding indexes. This is the slow path, parallelized by `num.recovery.threads.per.data.dir`. The marker is therefore the per-directory signal that distinguishes a quick startup from a potentially long, CPU/disk-intensive recovery.

go deeper

for a junior

Know that a clean shutdown writes a marker file so the next startup can skip the slow recovery scan.

for a middle

Describe the .kafka_cleanshutdown marker, the per-dir check, and the clean vs unclean startup paths.

for a senior

Connect the marker to recovery-point checkpoints, controlled.shutdown.enable, and the under-replication window on unclean restart.

for a principal

Set operational standards (graceful drain, controlled shutdown, kill-signal handling) and reason about per-dir markers in JBOD and crash-interruption edge cases.

## The problem When a broker restarts, it needs to know whether the data on disk is in a known-good, fully-flushed state, or whether it might be partially written because the process died mid-flight. Recovery (re-scanning and validating the tail of each log) is correct but slow, so Kafka wants to skip it whenever it safely can. ## The clean-shutdown marker Kafka solves this with a **marker file written per log directory** during an orderly shutdown, named **`.kafka_cleanshutdown`** (a hidden file in each `log.dirs` entry). The shutdown sequence is roughly: 1. Stop accepting new requests. 2. Flush all partition logs to disk (fsync). 3. Advance the `recovery-point-offset-checkpoint` (and write the other checkpoints) so the recovery point equals the log end offset. 4. Write the `.kafka_cleanshutdown` marker into each directory. ## On startup For each log directory the broker checks for the marker: - **Marker present →** the previous shutdown was clean. The broker deletes the marker (so a later crash without re-writing it will be detected) and loads logs using the trusted checkpoints. No tail re-validation is needed → **fast startup**. - **Marker absent →** the previous stop was unclean (kill -9, OOM, crash, power loss). The broker performs **log recovery**: it re-scans each affected partition from its recovery point to the end, validating record batches and rebuilding/ truncating indexes as needed. This is the slow path, parallelized by `num.recovery.threads.per.data.dir`. ## Why the distinction matters - **Startup time / availability:** clean shutdown → seconds; unclean shutdown on a many-partition broker → potentially many minutes of recovery, during which those partitions are unavailable on this broker and under-replicated. - **Operational practice:** always stop brokers with a graceful shutdown (SIGTERM and let it drain) rather than SIGKILL, and set `controlled.shutdown.enable=true` so leadership migrates off the broker before it stops — this both avoids recovery and avoids client-visible disruption. ## Edge cases - The marker is **per directory**, so in a JBOD setup one disk could be clean while another (that failed) is treated as unclean. - If a shutdown is interrupted *after* flushing but *before* writing the marker, the broker conservatively runs recovery — slower but safe. - Recovery is bounded by the recovery point, so even an unclean restart only re-scans the unflushed tail, not the entire log. ## Bottom line The `.kafka_cleanshutdown` marker, written per log dir at orderly shutdown and deleted at startup, is the signal that lets a broker skip recovery. Its presence means 'trust the checkpoints, start fast'; its absence means 'I might have crashed — recover the log tails.'

  • What is controlled.shutdown.enable and how does it relate to clean shutdown?
    When true (the default), a broker being stopped first migrates partition leadership to other in-sync replicas before exiting, minimizing client disruption. Combined with a graceful SIGTERM it allows the flush + clean-shutdown-marker sequence to complete, so the next startup skips recovery.
  • If you kill -9 a broker, what happens on the next startup?
    No clean-shutdown marker was written, so the broker treats it as unclean and runs log recovery — re-scanning each partition from its recovery point to the log end, which can be slow on many-partition brokers.

saying these in an interview costs you the question

  • Saying the broker always runs full recovery on every startup — it skips recovery when the clean-shutdown marker is present.
  • Claiming the marker is a single broker-wide file — it is per log directory.
  • Thinking SIGKILL and SIGTERM are equivalent for Kafka — only a graceful stop produces a clean shutdown.

context