What is a checkpoint in a database storage engine, and what two problems would a system have if it never took one?
answer
- gap: log knows changes the data files don't
- flush dirty pages + write checkpoint record
- recovery starts at last checkpoint, not at the beginning
- log reclaim needs the guarantee a checkpoint gives
- frequent = fast recovery, more write amplification
basics
~20 sA checkpoint flushes modified in-memory pages to disk and records a position in the transaction log. Without it, crash recovery would have to replay the log from the beginning, and the log could never be truncated, so it would grow forever.
solid answer
~50 sA checkpoint is a periodic action that makes the on-disk data files catch up with the changes already recorded in the write-ahead log. It flushes dirty pages from the buffer pool and writes a checkpoint record naming a log position from which recovery may safely start. Without checkpoints, two things break: 1. **Recovery time is unbounded.** After a crash the engine must replay log records to reconstruct changes not yet in the data files. With no checkpoint, the earliest such record could be from any time since the database was created, so recovery time grows with the age of the system rather than with recent activity. 2. **The log can never be reclaimed.** Log segments can only be recycled or archived once their changes are known to be safely in the data files. Without a checkpoint no such guarantee exists, so the log grows without bound until the disk fills. Checkpoint frequency is therefore a tuning knob trading recovery time against steady write cost.
go deeper
Be able to state it plainly: dirty pages get flushed and a marker is written, so recovery starts from there and old log can be dropped.
Add the tradeoff - shorter interval means faster recovery and more repeated writes of the same hot pages - and separate checkpoints from commit durability.
Connect it to operations: recovery-time objectives, log growth incidents and what pins the reclaim point, and how checkpoint pacing shows up in latency.
Frame it as an RTO-versus-throughput budget across the fleet, including replica catch-up, archiving pipelines, and the failure modes when the reclaim point stalls.
## The setup A relational engine does not write changed pages to their data files at commit time. It writes a compact **log record** describing each change, flushes the log at commit, and leaves the modified page **dirty** in memory. This is what makes commits fast: one sequential log write rather than scattered random page writes, and many updates to the same page collapse into a single eventual write. The consequence is a permanent gap. The log knows about changes the data files have never seen. Something must periodically close that gap, and that something is the **checkpoint**. ## What a checkpoint does Conceptually a checkpoint has three parts: 1. **Write out dirty pages** from the buffer pool - at minimum, all pages dirtied before some cut-off point in the log. 2. **Force the log** up to that point so log and data are consistent with each other. 3. **Record the checkpoint**: write a checkpoint record into the log and, crucially, store a pointer to it in a well-known place (a control file or header) that recovery reads first. After it completes, the engine can assert: every change logged before position X is present in the data files on disk. ## Problem 1: bounding recovery time When an engine restarts after a crash, it must find every change that was logged but might not have reached the data files, and replay it. The only way to know where to start is the last checkpoint - the checkpoint record says 'everything before here is already on disk', so recovery begins there and scans forward to the end of the log. The length of that scan is the dominant term in recovery time. Checkpoint every 30 seconds and recovery reads roughly 30 seconds' worth of log; checkpoint every hour and it reads an hour's worth, which on a write-heavy system can be many gigabytes and many minutes. Because the checkpoint interval sets an upper bound on how much log recovery must process, this is the knob operators tune when a recovery-time objective exists: 'we must be back within 60 seconds' translates directly into a checkpoint target. Note what checkpoints do *not* do: they do not make committed data durable. Durability comes from flushing the log at commit. A checkpoint is about *how long it takes to reconstruct* state after a crash, not about whether the state survives. ## Problem 2: reclaiming log space Log files are finite. A segment may only be recycled or archived once nothing will ever need to read it again - which means once every change it contains is guaranteed present in the data files. That guarantee is exactly what a checkpoint establishes. With no checkpoints, no segment can ever be released and the log grows until the filesystem is full, at which point the database typically stops accepting writes. This failure mode appears in real systems whenever something *blocks* checkpoint progress or holds the reclaim point back - a stuck long-running transaction, a replication slot no consumer is reading, an archiving command that keeps failing. The symptom is a log directory growing without limit, and the diagnosis is always 'what is preventing the reclaim point from advancing?' ## The cost side Checkpoints are not free. Flushing dirty pages is random write I/O, and doing it in a burst can visibly slow every concurrent query. Checkpointing more often also means the same hot page is written to disk more times - a page updated a thousand times between two checkpoints costs one write, but with ten times more frequent checkpoints it may cost ten. This is **write amplification**, and it is the direct price of shorter recovery. Hence the tradeoff every engine exposes in some form: frequent checkpoints give fast recovery and small logs at the cost of more steady-state write I/O; infrequent checkpoints give cheap steady state at the cost of slow recovery and large logs. ## Related mechanisms worth naming - **Fuzzy checkpoints**: production engines do not stop the world to take a checkpoint; they let transactions continue while pages are flushed, and record enough information for recovery to cope with the resulting inconsistency. - **Spreading the I/O**: rather than dumping all dirty pages at once, engines pace the writes across the interval so the burst becomes a hum. - **Shutdown checkpoints**: a clean shutdown takes a final checkpoint so that startup has essentially nothing to replay - which is why a clean restart is fast and a crash restart is not. ## The one-line answer A checkpoint synchronizes the data files with the log so that recovery can start from a recent, known-good position and old log can be thrown away; its frequency trades recovery time against write I/O.
- Does a checkpoint make committed transactions durable?No. Durability is established when the write-ahead log records for a transaction are flushed to stable storage at commit; that is what allows the engine to acknowledge the commit. A checkpoint only moves the data files forward so recovery has less log to replay and old log can be reclaimed. A committed transaction that never appears in any checkpoint is still fully recoverable.
- A production system's transaction log directory keeps growing and never shrinks. How does that relate to checkpoints?Log segments are released only once their changes are guaranteed present in the data files, which is what a checkpoint establishes, and only if nothing else still needs to read them. Growth means the reclaim point is not advancing - checkpoints are failing or being blocked, or a consumer such as a replication slot, an archive process, or a very old open transaction is pinning an old log position. The fix is to find what holds the reclaim point back, not to enlarge the disk.
Working from a running notebook of edits (the log) instead of rewriting the book each time. A checkpoint is the moment you actually apply the pending edits to the book and draw a line in the notebook - after that you can tear out the earlier pages, and if you are interrupted you only redo edits after the line.
saying these in an interview costs you the question
- Saying a checkpoint is what makes a commit durable
- Claiming the log is truncated on every commit
- Believing recovery must always replay the entire log
- Thinking a checkpoint writes only the log and no data pages
- Treating checkpoints as free and recommending 'checkpoint constantly'