skip to content

A database engine has both a background writer and a checkpointer, and both write dirty pages out of the buffer pool. Why does it need both, and what different goal does each one serve?

level: middleimportance: should knowfreq 32%

answer

  1. Bgwriter = clean buffers available, no deadline
  2. Checkpointer = consistent point, hard deadline
  3. No checkpoint → unbounded recovery and log growth
  4. Spread checkpoint writes or get a latency sawtooth
  5. Metric: pages written by backends vs by writer

basics

~20 s

The background writer keeps clean, reusable buffers available so a query needing a free buffer does not have to flush one itself — it targets buffer availability. The checkpointer flushes everything dirty as of a point in time to bound crash-recovery work and let old log be recycled.

solid answer

~60 s

They write the same kind of thing for different reasons. **Background writer — availability.** It scans for dirty pages that are likely eviction candidates and writes them continuously, so when a backend needs a free buffer it finds a clean one. Its success metric is how rarely a *foreground* session has to write a page itself. It makes no durability guarantee and has no deadline. **Checkpointer — bounded recovery.** On a schedule, or after a volume of log, it flushes every page dirty as of a point in time, then records that point. Everything before it is known to be in the data files, so older log can be recycled and recovery starts there. It has a hard completion deadline, which is why engines spread its writes across the interval instead of dumping them at once. So: one smooths steady-state latency, the other bounds restart time and log retention. Neither substitutes for the other — a perfectly working background writer would still leave unbounded recovery, and checkpoints alone would leave backends flushing pages between checkpoints.

go deeper

for a junior

State the two goals in one line each: reusable clean buffers versus a bounded recovery point.

for a middle

Explain what breaks if either is removed, and why the checkpointer has a deadline the background writer does not.

for a senior

Discuss pacing, the durability phase, log-triggered checkpoints, and the backend-writes metric as the leading indicator.

for a principal

Present checkpoint cadence as an explicit dial between steady-state write amplification and recovery-time objective, with the background writer's tuning as the smoothing term around it.

## Same action, different contracts Both processes take dirty pages out of the shared buffer pool and write them to the data files. What differs is the *promise* each one makes. ## Background writer: keep clean buffers ahead of demand When a backend needs to read a page that is not cached, it must claim a buffer. If the buffer chosen for reuse is dirty, that backend has to write it out first — a synchronous, random write in the middle of a user query, and exactly the latency spike you are trying to avoid. The background writer exists to make that rare. It walks the pool, looks at the pages nearest to eviction, and writes some of them, leaving them clean but still cached. If the page is used again it is still there; if it is evicted, eviction is free. It is deliberately conservative. Writing too aggressively wastes I/O on pages that will be dirtied again in a moment; writing too little pushes the work onto backends. Engines self-tune it against recent demand, and the diagnostic worth watching is the ratio of pages written *by backends* versus by the writer — a high backend share means the pool is too small or the writer is too timid. Crucially, it offers no guarantee about *which* pages are on disk at any moment. It cannot bound recovery, because it never promises to have finished anything. ## Checkpointer: establish a recovery starting point Recovery after a crash must replay log from a point where the data files are known consistent. Producing that point is the checkpoint's job: flush everything dirty as of an instant, wait for it to be durable, then record "recovery may start here". Two direct payoffs. Restart time is bounded by how much log accumulated since the last checkpoint. And log segments older than that point are no longer needed for crash recovery, so they can be recycled or archived — without checkpoints, log storage grows without limit. The checkpointer therefore has an obligation the background writer does not: a set of pages that must all be written, by a deadline. Do it naively and you get a burst of random writes that saturates the device, stalls every query, and produces the classic sawtooth latency graph. So engines *spread* checkpoint writes across a fraction of the interval, pacing them against elapsed time and log consumed. The second half of a checkpoint — forcing the file system to make those writes durable — is often the more painful part, because the storage layer has been accumulating them. ## Why both, concretely Imagine dropping the background writer. Between checkpoints, dirty pages accumulate; every backend that needs a buffer finds a dirty one and writes it itself. Query latency becomes hostage to eviction luck. Now imagine dropping the checkpointer. Pages still get written, eventually, in no particular order and with no completion guarantee. Nothing establishes a consistent starting point, so recovery would have to replay from the beginning of time and log could never be discarded. They are complementary: one is a smoothing mechanism with no deadline, the other a deadline-driven barrier. ## Tuning tension Checkpoint frequency is a two-sided dial. Frequent checkpoints mean short recovery and small log retention, but the same hot page gets written repeatedly — more total I/O. Infrequent checkpoints coalesce more modifications per write but stretch recovery time and swell log storage, and each checkpoint moves more pages when it does run. Meanwhile the background writer's aggressiveness trades wasted writes against foreground stalls, and a well-tuned writer also flattens checkpoints by having already cleaned many pages before the checkpoint starts. ## What to look at in production Count how many page writes were done by backends versus by the background writer versus by the checkpointer. Watch for checkpoints triggered by log volume rather than by time — that means the write rate is outrunning the configured pacing. Watch the duration of the durability phase of each checkpoint. Rising values in any of these mean background flushing is losing the race, and the shortfall always shows up as foreground latency.

  • What happens if checkpoints are made much less frequent to reduce write amplification?
    Steady-state I/O drops because a hot page is modified many more times per write, but two costs rise. Crash recovery must replay much more log, so restart takes longer, and log storage must hold everything since the last checkpoint. Each checkpoint also has more dirty pages to move when it finally runs, so the burst is larger unless pacing is adjusted.
  • Which single metric best tells you the background writer is not keeping up?
    The proportion of buffer writes performed by ordinary backends rather than by the background writer or checkpointer. Every such write is a user query stalling to flush a page before it can claim a buffer. A rising share points to a buffer pool that is too small for the working set, a writer configured too conservatively, or a write rate the storage cannot absorb.

saying these in an interview costs you the question

  • Saying the background writer guarantees data is durable
  • Thinking a checkpoint is just the background writer running harder
  • Assuming checkpoints can be dropped because pages get written anyway
  • Believing more frequent checkpoints are always better
  • Ignoring that the durability phase of a checkpoint, not the write phase, often causes the stall

context