What controls when KRaft generates a new metadata snapshot, and what does metadata.log.max.record.bytes.between.snapshots do?
answer
- bytes-between-snapshots default 20 MB
- size trigger bounds replay tail
- max.snapshot.interval.ms = time trigger (default 1h)
- lower threshold -> faster failover, more IO
- snapshot then truncate log prefix
basics
~10 sSnapshots are triggered by thresholds. metadata.log.max.record.bytes.between.snapshots sets how many bytes of new log records may accumulate since the last snapshot before a new one is generated. A time-based setting also forces periodic snapshots.
solid answer
~50 sKRaft generates snapshots on thresholds rather than continuously. `metadata.log.max.record.bytes.between.snapshots` (default 20 MB) bounds how many bytes of metadata records may be appended since the last snapshot before a new snapshot is taken — a size trigger that keeps the replay tail and recovery time bounded relative to write volume. `metadata.log.max.snapshot.interval.ms` adds a time trigger so that even a low-traffic cluster still snapshots periodically (important so the log can be truncated and a long-running node's tail doesn't grow stale). When either threshold is crossed, the node materializes the current image into a new snapshot file named by its end offset/epoch, then becomes eligible to delete log segments fully covered by that snapshot (subject to retention of records still needed by lagging quorum members). Lower thresholds mean smaller replay tails and faster failover but more frequent snapshotting IO/CPU; higher thresholds reduce snapshot overhead at the cost of longer catch-up. These apply to both controllers and brokers, which each maintain their own snapshots of the metadata log.
go deeper
Know a config limits how much log piles up before a new snapshot is taken.
Name the byte threshold and that there is also a time-based trigger; default ~20 MB.
Explain the size vs time triggers, the failover-vs-IO trade-off, and the truncation interaction.
Tune thresholds against recovery-time objectives and metadata size; reason about lagging-member retention vs snapshot install.
## Why thresholds, not continuous Writing a snapshot serializes the entire metadata image to disk, costing CPU and IO proportional to state size. Doing it on every record would be wasteful. So KRaft snapshots when accumulated change crosses a configured threshold, trading **snapshot frequency** against **replay-tail length**. ## `metadata.log.max.record.bytes.between.snapshots` This is the ***size* trigger**. It is the maximum number of bytes of metadata log records that may accumulate after the latest snapshot before the node generates a new snapshot. Default is 20971520 bytes (20 MB). Interpretation: it bounds the size of the log tail that any node would have to replay on top of the latest snapshot. - A smaller value -> shorter replay tail -> faster catch-up/failover, but more frequent (costlier) snapshotting. - A larger value -> rarer snapshots, less overhead, but longer recovery. ## `metadata.log.max.snapshot.interval.ms` This is the ***time* trigger** (default 1 hour). Even if a cluster is nearly idle and never crosses the byte threshold, this forces a snapshot after the configured elapsed time. Why it matters: without a time trigger, a low-write cluster might never snapshot, so the metadata log could never be truncated and an old node could face a long replay; periodic snapshots keep the log truncatable and recovery bounded. Setting it to 0 disables the time-based trigger. ## What 'generate a snapshot' does The node freezes the current materialized `MetadataImage` (or controller state) and writes it as a checkpoint file named with the offset and epoch it covers, e.g. `0000000000000123456-0000000000000007.checkpoint`. The write must be **fully durable** before the corresponding log prefix can be deleted. ## Log truncation interaction After a durable snapshot at offset X, segments whose records are all <= X are eligible for deletion. But the node must not delete records still needed by a lagging quorum member that hasn't caught up — otherwise that member could never catch up without a snapshot install. KRaft balances this: it can serve far-behind followers the snapshot directly rather than retaining the whole log. ## Who has these configs Both controllers and brokers maintain snapshots of the metadata log (brokers as observers), so these settings govern snapshotting on all KRaft nodes. ## Tuning guidance - For tight recovery-time objectives (fast failover), lower the byte threshold so the replay tail stays small. - For very large metadata (many topics/partitions), be mindful that frequent full-image snapshots are expensive; balance against IO budget. - Most clusters are fine on defaults.
- Why is a time-based snapshot trigger needed in addition to the byte threshold?On a low-traffic cluster the byte threshold may rarely be crossed, so the log would never be truncated and a recovering node could face a stale/long tail. The time trigger guarantees periodic snapshots regardless of write volume.
- What is the trade-off in lowering metadata.log.max.record.bytes.between.snapshots?Faster catch-up and failover because the replay tail stays small, at the cost of more frequent snapshot generation, which consumes CPU and disk IO proportional to metadata size.
saying these in an interview costs you the question
- Thinking snapshots are taken on a fixed schedule only (size threshold also triggers them)
- Claiming the byte threshold counts user/topic data rather than metadata records
- Saying only controllers snapshot (brokers as observers also maintain snapshots)
- Believing the log prefix can be deleted before the snapshot is durable