Explain the difference between a 'sharp' (consistent) checkpoint and a 'fuzzy' checkpoint in a storage engine, and why production databases use the fuzzy variety.
answer
- sharp = stop the world, consistent image, unusable in prod
- fuzzy = flush while writes continue, image is a smear
- record active transactions + oldest dirty page position
- per-page LSN makes redo idempotent
- redo everything, then undo the uncommitted
basics
~20 sA sharp checkpoint quiesces the system so the flushed image is transaction-consistent. A fuzzy checkpoint lets writes continue while pages are flushed, so the image is inconsistent, and records enough state - such as the oldest dirty page's log position and the active-transaction list - for recovery to fix it up.
solid answer
~50 sA **sharp** (consistent) checkpoint stops new work, waits for in-flight transactions, flushes every dirty page, and writes the marker. Recovery is then trivial - the data files are a consistent snapshot - but the database is unavailable for as long as the flush takes, which on a large buffer pool is far too long to accept. A **fuzzy** checkpoint runs concurrently with the workload. Pages are written while other transactions are still modifying pages, so the on-disk image mixes states from different points in time and may contain changes from uncommitted transactions. To make that recoverable, the checkpoint record captures the bookkeeping recovery needs: the list of transactions active at checkpoint time, and the earliest log position of any page still dirty in the pool (the dirty-page table / recovery start point). Recovery therefore starts from that earliest-dirty-page position, redoes changes not present on a page (comparing each page's stored log sequence number), and then undoes transactions that never committed. Every production engine works this way.
go deeper
Know the distinction in one sentence: sharp stops the world for a clean image, fuzzy runs alongside the workload and leaves recovery to sort it out.
Name the state a fuzzy checkpoint records - active transactions and the oldest dirty page - and explain why idempotent redo makes an inconsistent image safe.
Connect it to operations: why the recovery start point can lag the checkpoint, why long-dirty pages matter, and why clean restarts are fast while crash restarts are not.
Discuss it as an availability tradeoff baked into the storage engine - no stop-the-world window in exchange for a more complex recovery contract - and what that costs in engine complexity and in recovery predictability.
## Sharp checkpoints The textbook version is simple to reason about: 1. Stop accepting new transactions. 2. Wait for all in-flight transactions to finish. 3. Flush every dirty page in the buffer pool. 4. Flush the log and write the checkpoint record. 5. Resume. The result is a **transaction-consistent** image on disk: no partial transactions, nothing missing. Recovery from such a checkpoint would need to do almost nothing - start at the marker and replay forward. The problem is step 1 through 3. On a server with tens of gigabytes of dirty pages, flushing them is minutes of random write I/O, and the whole time no transaction may start or commit. The system's availability would be punctuated by regular multi-minute stalls. Nobody accepts that, so sharp checkpoints survive only in special situations - notably a **clean shutdown**, where stopping the workload is the point anyway, which is why a cleanly shut down database restarts almost instantly. ## Fuzzy checkpoints A fuzzy checkpoint refuses to quiesce. It begins, notes some state, flushes pages over a period of time while the workload continues, and ends. Because transactions keep running throughout: - A page flushed early in the checkpoint may be modified again before the checkpoint finishes, so what is on disk is already stale. - A page flushed late may contain changes from a transaction that has not committed and may still abort. - Different pages on disk reflect different moments in time - the image is *fuzzy*, not a snapshot. This is fine as long as recovery is given enough information to repair it. Two pieces of state are recorded: - **The active transaction table**: which transactions were in flight, and where their log records begin. Recovery needs this to know what to undo. - **The dirty page table** - or at minimum the oldest log position among all currently dirty pages. This is the true **recovery start position**: any change logged before it is guaranteed to be on disk, because no dirty page is older than it. Crucially the recovery start point is *not* the checkpoint record itself. If a page has been sitting dirty in the pool for ten minutes, its changes are not on disk, so recovery must begin from that page's oldest change, which precedes the checkpoint. A page that stays dirty for a long time therefore holds the recovery start point back and lengthens recovery - which is why engines flush long-dirty pages preferentially and why the checkpoint interval alone does not fully determine recovery time. ## How recovery copes The repair machinery is the ARIES-style redo/undo pattern, and only two ideas are needed to see why a fuzzy image is safe: 1. **Per-page log sequence numbers.** Every page stores the identifier of the last log record applied to it. During redo, recovery reads a log record for a page, compares the record's position with the page's stored value, and applies it only if the page is behind. So replaying a change that is already present is harmless - redo is **idempotent**, which is exactly what makes 'some pages newer than others' tolerable. 2. **Redo everything, then undo the losers.** Recovery first replays *all* logged changes from the start position - including changes made by transactions that never committed - so that the pages reach the exact state memory was in at the crash. Then it uses the active-transaction information to roll back those transactions that had not committed. Undoing uncommitted work after redoing it is simpler and more robust than trying to filter during replay. With those two properties, an inconsistent on-disk image is a non-problem: redo brings every page forward to the crash point regardless of when it happened to be flushed, and undo removes the uncommitted parts. ## Why this shapes engine design Several familiar behaviours follow directly from fuzzy checkpointing: - **Checkpoints can be spread over time.** Since there is no stop-the-world window, the flush can be paced deliberately over most of the interval to avoid an I/O spike. A sharp checkpoint has no such freedom. - **A checkpoint is a process, not an instant.** It has a start and an end, and it may still be running when the next one is due - a signal that the system is dirtying pages faster than it can flush them. - **Recovery time depends on the oldest dirty page, not just on interval.** Operators tune both the interval and how aggressively long-dirty pages are written. - **A crash restart is slower than a clean restart**, because a clean shutdown ends with a sharp checkpoint and a crash leaves a fuzzy one plus a tail of log. ## The interview answer Sharp = quiesce, consistent image, unacceptable stall. Fuzzy = concurrent, inconsistent image, plus recorded active transactions and oldest-dirty-page position. Idempotent redo keyed on per-page log sequence numbers plus undo of uncommitted transactions makes the inconsistency safe.
- If a fuzzy checkpoint writes pages containing uncommitted changes, how does the database avoid leaving that uncommitted data in the files?It does not avoid writing it - a steal policy explicitly permits flushing pages dirtied by uncommitted transactions. Safety comes from having logged the undo information before the page was written, so recovery can roll the change back. After redoing everything up to the crash point, recovery consults the list of transactions that were active and had not committed, and undoes their changes using those log records.
- Why can recovery have to start from a log position earlier than the last checkpoint record?The safe start point is the oldest change among pages still dirty in the buffer pool, not the moment the checkpoint was recorded. A page dirtied ten minutes ago and not yet flushed has changes that are absent from disk, so replay must begin at that page's earliest change. This is why engines track the oldest dirty page and preferentially flush long-dirty pages to keep the recovery start point moving forward.
saying these in an interview costs you the question
- Claiming a fuzzy checkpoint produces a consistent on-disk snapshot
- Saying recovery always starts exactly at the checkpoint record
- Believing pages with uncommitted changes are never written to disk
- Not knowing that redo is made idempotent by comparing per-page log sequence numbers
- Describing a sharp checkpoint as the normal production behaviour