How does a Kafka broker know on startup whether it was shut down cleanly, and why does that distinction matter?
answer
- .kafka_cleanshutdown marker, per log dir
- written on graceful stop, deleted on startup
- present = skip recovery (fast)
- absent = unclean = run recovery
- controlled.shutdown.enable + SIGTERM
basics
~20 sOn a clean shutdown the broker flushes all data, updates the checkpoint files, and writes a clean-shutdown marker file in each log dir. On startup, if that marker is present it skips log recovery; if it is missing (a crash), it runs recovery.
solid answer
~40 sDuring an orderly shutdown the broker flushes every log, advances the `recovery-point-offset-checkpoint` to the log end, and writes a marker file named `.kafka_cleanshutdown` into each log directory. On the next startup the broker checks for this marker in each directory: if present, it deletes the marker and trusts the checkpoints, so startup is fast and no log re-validation is needed. If the marker is absent — because the process was SIGKILLed, OOM-killed, or the host crashed/lost power — the broker assumes an unclean shutdown and performs recovery: it scans each affected partition from its recovery point to the log end, re-validating records and rebuilding indexes. This is the slow path, parallelized by `num.recovery.threads.per.data.dir`. The marker is therefore the per-directory signal that distinguishes a quick startup from a potentially long, CPU/disk-intensive recovery.
go deeper
Know that a clean shutdown writes a marker file so the next startup can skip the slow recovery scan.
Describe the .kafka_cleanshutdown marker, the per-dir check, and the clean vs unclean startup paths.
Connect the marker to recovery-point checkpoints, controlled.shutdown.enable, and the under-replication window on unclean restart.
Set operational standards (graceful drain, controlled shutdown, kill-signal handling) and reason about per-dir markers in JBOD and crash-interruption edge cases.
## The problem When a broker restarts, it needs to know whether the data on disk is in a known-good, fully-flushed state, or whether it might be partially written because the process died mid-flight. Recovery (re-scanning and validating the tail of each log) is correct but slow, so Kafka wants to skip it whenever it safely can. ## The clean-shutdown marker Kafka solves this with a **marker file written per log directory** during an orderly shutdown, named **`.kafka_cleanshutdown`** (a hidden file in each `log.dirs` entry). The shutdown sequence is roughly: 1. Stop accepting new requests. 2. Flush all partition logs to disk (fsync). 3. Advance the `recovery-point-offset-checkpoint` (and write the other checkpoints) so the recovery point equals the log end offset. 4. Write the `.kafka_cleanshutdown` marker into each directory. ## On startup For each log directory the broker checks for the marker: - **Marker present →** the previous shutdown was clean. The broker deletes the marker (so a later crash without re-writing it will be detected) and loads logs using the trusted checkpoints. No tail re-validation is needed → **fast startup**. - **Marker absent →** the previous stop was unclean (kill -9, OOM, crash, power loss). The broker performs **log recovery**: it re-scans each affected partition from its recovery point to the end, validating record batches and rebuilding/ truncating indexes as needed. This is the slow path, parallelized by `num.recovery.threads.per.data.dir`. ## Why the distinction matters - **Startup time / availability:** clean shutdown → seconds; unclean shutdown on a many-partition broker → potentially many minutes of recovery, during which those partitions are unavailable on this broker and under-replicated. - **Operational practice:** always stop brokers with a graceful shutdown (SIGTERM and let it drain) rather than SIGKILL, and set `controlled.shutdown.enable=true` so leadership migrates off the broker before it stops — this both avoids recovery and avoids client-visible disruption. ## Edge cases - The marker is **per directory**, so in a JBOD setup one disk could be clean while another (that failed) is treated as unclean. - If a shutdown is interrupted *after* flushing but *before* writing the marker, the broker conservatively runs recovery — slower but safe. - Recovery is bounded by the recovery point, so even an unclean restart only re-scans the unflushed tail, not the entire log. ## Bottom line The `.kafka_cleanshutdown` marker, written per log dir at orderly shutdown and deleted at startup, is the signal that lets a broker skip recovery. Its presence means 'trust the checkpoints, start fast'; its absence means 'I might have crashed — recover the log tails.'
- What is controlled.shutdown.enable and how does it relate to clean shutdown?When true (the default), a broker being stopped first migrates partition leadership to other in-sync replicas before exiting, minimizing client disruption. Combined with a graceful SIGTERM it allows the flush + clean-shutdown-marker sequence to complete, so the next startup skips recovery.
- If you kill -9 a broker, what happens on the next startup?No clean-shutdown marker was written, so the broker treats it as unclean and runs log recovery — re-scanning each partition from its recovery point to the log end, which can be slow on many-partition brokers.
saying these in an interview costs you the question
- Saying the broker always runs full recovery on every startup — it skips recovery when the clean-shutdown marker is present.
- Claiming the marker is a single broker-wide file — it is per log directory.
- Thinking SIGKILL and SIGTERM are equivalent for Kafka — only a graceful stop produces a clean shutdown.