What are the checkpoint files in a Kafka log directory (recovery-point-offset-checkpoint and replication-offset-checkpoint), and what does each track?
answer
- recovery-point = flushed-to-disk offset
- replication-offset = high watermark (HW)
- log-start-offset = retention floor
- cleaner-offset = compaction progress
- per log dir, plaintext, version + count + lines
basics
~20 sEach log directory holds small checkpoint files. recovery-point-offset-checkpoint records, per partition, the offset up to which data has been flushed to disk (so recovery can skip it). replication-offset-checkpoint records the high watermark, the offset up to which messages are replicated/committed.
solid answer
~50 sInside every `log.dirs` directory Kafka keeps a handful of plaintext checkpoint files, written periodically and on clean shutdown. The two central ones are: **`recovery-point-offset-checkpoint`** — for each partition, the offset up to which records are known to be flushed (fsync'd) to disk; on startup Kafka only needs to recover/re-validate records *after* this point, so it bounds recovery work. **`replication-offset-checkpoint`** — for each partition, the **high watermark (HW)**, the highest offset that is replicated to all in-sync replicas and therefore visible to consumers (committed). On restart this lets the broker restore the HW without re-deriving it. Other files in the dir include `log-start-offset-checkpoint` (the log start offset after retention/deletes or compaction) and `cleaner-offset-checkpoint` (per-partition progress of the log cleaner / compaction). Each file is per-data-directory and lists `topic partition offset` lines under a version and entry-count header. They are the durable metadata that makes recovery fast and the HW survive restarts.
go deeper
Know each log dir has small files recording per-partition offsets used at restart.
Distinguish recovery-point (flushed offset) from replication-offset (high watermark) and name the other checkpoint files.
Explain how recovery point bounds recovery, how HW restoration affects reads and truncation, and per-dir scope.
Reason about checkpoint durability, corruption handling, leader-epoch truncation, and the metadata's role in restart correctness across a fleet.
## Where these live Every directory listed in `log.dirs` is self-contained: alongside the partition subdirectories (e.g. `my-topic-0/`) sit several small text checkpoint files that record per-partition offsets for *that directory only*. They are rewritten periodically by background tasks and on clean shutdown. ## recovery-point-offset-checkpoint — the durability boundary Kafka does not fsync every message individually (that would be slow); it lets the OS page cache absorb writes and flushes periodically. The **recovery point** for a partition is the offset up to which data has actually been flushed to disk. This file stores `topic partition recoveryPointOffset` for every partition in the directory. - **On a clean shutdown**, Kafka flushes everything and advances each partition's recovery point to the log's end offset. - **On startup**, Kafka must only recover (re-scan/validate) records *after* the recovery point — because everything up to it is known-good on disk. This is exactly what bounds the recovery work discussed under `num.recovery.threads.per.data.dir`. - After an **unclean shutdown**, the recovery point lags the true log end, so Kafka re-validates the tail of each affected log. ## replication-offset-checkpoint — the high watermark The **high watermark (HW)** of a partition is the highest offset that has been replicated to *all* in-sync replicas (the ISR). Consumers can only read up to the HW — messages above it are not yet 'committed' and could be lost if the leader fails. This file stores the HW per partition so it survives a restart. - Without persisting the HW, a restarting broker couldn't immediately know how far reads are safe. - The HW is distinct from the **log end offset (LEO)**, which is the offset of the next record to be appended on this replica. LEO ≥ HW always. ## The other checkpoint files - **`log-start-offset-checkpoint`** — the **log start offset** per partition: the lowest offset still retained, which moves forward as retention deletes old segments or as records are explicitly deleted (DeleteRecords) or compacted. Restores the start offset across restarts. - **`cleaner-offset-checkpoint`** — for compacted topics, tracks how far the **log cleaner** has compacted each partition, so compaction resumes rather than restarting. ## File format Each file is plaintext: ``` 0 <- version 3 <- number of entries my-topic 0 1500 my-topic 1 1490 orders 0 9921 ``` Line 1 = version, line 2 = entry count, then one `topic partition offset` per partition. ## Why this matters operationally - These files are the **durable metadata** that make startup recovery bounded (recovery point) and reads correct after restart (HW). - They are written **per log directory**, so each disk in a JBOD setup carries the checkpoints for the partitions it hosts. Losing a disk loses only that dir's checkpoints, and the affected replicas re-sync from the leader. - Corruption or deletion of these files forces full recovery of the affected partitions (treated like an unclean state). ## Edge cases - A truncated/zero-length checkpoint after a crash mid-write is tolerated: Kafka falls back to full recovery for those partitions. - The recovery point can equal the LEO (fully flushed) or lag it (pending flush). - HW restoration interacts with leader election: a new leader may truncate followers to the HW (or use leader epochs to truncate precisely). ## Bottom line `recovery-point-offset-checkpoint` = how far data is flushed (bounds recovery). `replication-offset-checkpoint` = the high watermark (bounds safe reads). Plus `log-start-offset-checkpoint` and `cleaner-offset-checkpoint` for retention and compaction progress. All per-directory, all small text files, all critical to fast, correct restarts.
- What is the difference between the high watermark and the log end offset?The log end offset (LEO) is the offset of the next record to append on a replica. The high watermark (HW) is the highest offset replicated to all in-sync replicas and thus visible to consumers. LEO ≥ HW; the gap is data appended on the leader but not yet fully replicated.
- Why does the recovery-point checkpoint make startup faster?It records the offset up to which each partition is flushed to disk. On restart Kafka only re-validates records after that point, so it skips the bulk of the log instead of scanning everything from the start.
- What happens if these checkpoint files are missing or corrupt after a crash?Kafka treats the affected partitions as unflushed and performs full recovery (re-scanning their logs), and re-derives/restores the HW. It is tolerant of truncated checkpoints written during a mid-crash flush.
saying these in an interview costs you the question
- Confusing the recovery point (flush boundary) with the high watermark (replication/consumer-visibility boundary).
- Saying replication-offset-checkpoint stores consumer group offsets — those live in __consumer_offsets, not here.
- Thinking the checkpoints are global to the broker — they are per log directory.