skip to content

What is a torn page, why can a crash while the engine is flushing an 8 KB or 16 KB page produce one, and how do storage engines protect against it - for example with full-page images in the log or a doublewrite area?

level: middleimportance: should knowfreq 42%

answer

  1. page 8-16 KB vs 4 KB atomic write unit
  2. half old, half new - not any valid version
  3. redo needs a valid earlier version; header lies
  4. first touch after a checkpoint -> full page image in log
  5. doublewrite: write twice, restore the intact copy

basics

~20 s

A page is larger than the sector the device writes atomically, so a crash mid-write can leave a page half old, half new - internally corrupt and unusable by redo. Engines fix it by first writing a whole clean copy of the page: a full-page image in the log, or a doublewrite area, from which recovery restores it.

solid answer

~1 min

A data page is typically 8 KB or 16 KB, but storage only guarantees atomicity for a much smaller unit - historically a 512-byte sector, 4 KB on modern devices. Writing one page is therefore several device-level writes, and a crash or power loss in the middle can leave the on-disk page as a mixture of old and new fragments: a **torn page**. That breaks recovery's core assumption. Redo applies a change to a page assuming the page is a valid earlier version; a torn page is not any version at all - its header, item pointers, and rows can disagree - so replaying onto it produces corruption rather than repair. The defence is to have a known-good full copy of the page somewhere before the risky write: - **Full-page writes in the log**: the first time a page is modified after a checkpoint, the entire page image is written into the log, so recovery can restore it wholesale and then replay from there. - **Doublewrite area**: pages are first written to a dedicated contiguous region, then to their real location; if the second write tears, recovery copies the good version back. Both trade extra write volume for safety.

go deeper

for a junior

Know that a page is bigger than what the disk writes atomically, so a crash can leave it half-written, and that engines keep a spare good copy to repair it.

for a middle

Explain why redo cannot fix a torn page, and describe both full-page images and the doublewrite area with their costs.

for a senior

Discuss the tuning consequences - full-image volume versus checkpoint interval, write amplification, and the conditions under which the protection can be safely disabled.

for a principal

Frame it as an end-to-end durability contract across device, RAID, virtualization, and filesystem, and weigh the write-amplification budget against a silent-corruption risk that surfaces long after the incident.

## Where the tear comes from Databases work in pages of 4, 8, or 16 KB. Storage devices guarantee atomicity only at a much smaller granularity: classically a 512-byte sector, 4 KB on modern drives, and the guarantee stops at the device anyway - a filesystem, a RAID layer, or a virtualized disk can split or reorder pieces. Writing one 8 KB page is therefore a multi-part operation with no all-or-nothing promise. If the machine loses power or panics partway through, the page on disk becomes a hybrid: some sectors carry the new contents, some the old. This is a **torn page** (also called a partial page write or fractured block). ## Why a torn page is worse than an old page An *old* page is harmless. Recovery expects that - it compares the page's stored log sequence number with the log records and replays the missing changes. The whole redo design assumes the page is some valid version of itself. A torn page is not a valid version. A page has internal structure - a header with a free-space pointer and a log sequence number, an array of item pointers, and the tuples themselves - and those parts can now come from different generations. The header may claim a layout the body does not have; an item pointer can address a region containing something else. The engine may even read a plausible-looking log sequence number from a stale header and conclude the page is up to date, skipping the redo that would have rewritten the damaged half. Silent corruption, discovered days later, is the realistic outcome. Checksums help detect this but do not repair it. ## Why checkpoints are the natural place to discuss it Tearing happens when a page is written to its home location, and the checkpoint mechanism is what schedules the bulk of those writes. More importantly, the standard protection is anchored on checkpoints, as the next section shows: 'the first modification after a checkpoint' is the trigger, which makes checkpoint frequency directly control the protection's cost. ## Protection 1: full-page images in the log The idea: instead of relying on the home-location write being atomic, put a complete, known-good copy of the page somewhere that *is* safe - the log, which is written sequentially and protected by its own checksums. The rule is: **the first time a page is modified after a checkpoint, write its entire image into the log** rather than just the delta. Subsequent modifications of the same page before the next checkpoint log only deltas. Recovery then does something stronger than replay: when it encounters a full-page image, it *overwrites* the on-disk page with that image, no questions asked, without consulting the possibly-garbage header. Whatever tearing had occurred is obliterated, and replay continues from a known state. The cost is log volume. The first touch of a page after each checkpoint costs a full page in the log instead of a few dozen bytes. On a workload that scatters small updates across many pages, this can multiply log write volume several-fold - and it is why **checkpointing more frequently makes the log bigger**, which is counterintuitive until you see the mechanism: each new checkpoint resets every page to 'first touch', so every hot page pays the full image again. ## Protection 2: doublewrite area An alternative: keep a dedicated, contiguous region on disk (the doublewrite buffer). Flushing works in two stages: 1. Write the batch of pages sequentially into the doublewrite area and flush it. 2. Write each page to its real home location. If a crash tears a home-location write, recovery finds the page's checksum invalid and copies the intact copy from the doublewrite area. If the crash happens during stage 1, nothing is lost - the home pages are still their previous valid selves. The cost is that every page is written twice. It is less painful than it sounds, because the doublewrite writes are sequential and batched, but it is real and it is why the mechanism is sometimes disabled on storage that guarantees atomic page-sized writes. ## When protection can be turned off Both mechanisms are insurance against non-atomic page writes. If the storage stack genuinely provides atomic writes at the page size - some enterprise SSDs, some filesystems with copy-on-write semantics or explicit atomic-write support - the protection is redundant and can be disabled for a substantial reduction in write volume. The judgement required is real: the claim must hold for the *entire* stack, including any virtualization, RAID layer, and the filesystem, not just the drive's datasheet. Disabling it on an unverified stack converts a rare crash into silent data corruption, so the default answer in an interview is 'keep it on unless you can prove atomicity end to end'. ## What to say in an interview Define the tear (page > atomic write unit), explain why it defeats redo (the page is not a valid version and its header cannot be trusted), name both protection styles and the cost each pays, and note the tuning consequence that more frequent checkpoints increase full-page-image volume.

  • Why does checkpointing more frequently increase the size of the transaction log rather than decrease it?
    With full-page-image protection, the first modification of a page after each checkpoint logs the entire page instead of a small delta. Shortening the interval resets every page to 'first touch' more often, so hot pages pay the full-image cost repeatedly. The log-space saving from earlier reclamation is often outweighed by this extra volume, which is a classic surprise when someone tightens the checkpoint target.
  • Under what conditions is it defensible to disable torn-page protection?
    Only when the entire storage stack guarantees atomic writes at the database's page size - the device, any RAID or virtualization layer, and the filesystem. Some enterprise SSDs and copy-on-write filesystems provide this. The verification burden is high and the failure mode is silent corruption discovered long after the crash, so the safe default is to leave it enabled and treat disabling it as a deliberate, tested decision.

saying these in an interview costs you the question

  • Believing a page write is atomic because the filesystem or the drive 'handles it'
  • Thinking a torn page is equivalent to an out-of-date page that redo can simply fix
  • Claiming checksums prevent torn pages rather than merely detecting them
  • Assuming more frequent checkpoints always shrink the log
  • Confusing a doublewrite area with a replica or a second copy of the whole database

context