skip to content

You take an LVM snapshot of a mounted, actively written filesystem on a Linux host. What consistency does the resulting snapshot actually give you, and what does running `fsfreeze -f` on the mount point immediately before the snapshot add?

level: seniorimportance: nice to knowfreq 33%

answer

  1. block layer only, nothing above it
  2. like pulling the power cable
  3. journal replays, recent writes may not be there
  4. flush and block writers, then snapshot
  5. freeze window must be tiny, always thaw

basics

~20 s

The snapshot is crash-consistent: it looks exactly like the volume would after a power cut, so a journalling filesystem mounts it after replaying its journal, but anything still buffered above the block layer is missing. fsfreeze -f flushes and quiesces the filesystem first, giving a cleanly-shut-down image instead.

solid answer

~60 s

An LVM snapshot is a block-level operation. Creating it suspends the origin device briefly so outstanding block I/O is flushed and the image is atomic at the block layer, but LVM knows nothing about the filesystem or the applications above it — dirty page-cache data that has not been written back, and application state that was never written at all, simply are not there. That is *crash consistency*: mounting the snapshot behaves like mounting after a power cut, so ext4 or XFS replays its journal and comes up, usually intact but missing recent work. `fsfreeze -f /mnt/data` upgrades that: the kernel flushes the filesystem's dirty data and metadata to disk and blocks new writers until `fsfreeze -u`, so the snapshot taken in between is a cleanly quiesced filesystem image needing no journal replay. Two operational caveats: keep the frozen window down to the snapshot creation itself, because every writer blocks meanwhile, and remember the freeze still does not flush *application* buffers — anything holding state in memory has to be told to checkpoint before you freeze.

code

bash · 14 lines
bash
#!/bin/bash
set -euo pipefail
MNT=/var/lib/app

# always thaw, even if the snapshot fails
trap 'fsfreeze -u "$MNT" || true' EXIT

fsfreeze -f "$MNT"
lvcreate -s -L 10G -n app-snap vg0/app
fsfreeze -u "$MNT"
trap - EXIT

# the copy happens afterwards, with the filesystem writable again
mount -o ro /dev/vg0/app-snap /mnt/snap

go deeper

for a junior

Know that a snapshot of a mounted filesystem is like a sudden power loss: it usually mounts because the journal replays, but very recent writes may be missing.

for a middle

Explain the layering — LVM captures the block device, the page cache and application memory are above it — and what fsfreeze -f changes by flushing the filesystem and blocking writers until the thaw.

for a senior

Show the operational discipline: script freeze, snapshot and thaw as one sequence with an unconditional thaw, keep the frozen window to a single command, judge when blocking writers is worse than accepting crash consistency, and prove the choice with a restore test.

for a principal

Own the backup contract across the estate: which workloads need application-level quiescing versus crash consistency, what recovery point that implies, and how the organisation verifies restorability instead of assuming a journal replay will be enough.

## Three layers, three kinds of state When a snapshot is taken, data relevant to "is this consistent?" lives in three places: 1. **The block device** — data already written to the logical volume. The snapshot captures this. 2. **The kernel page cache and filesystem journal** — data the filesystem has accepted but not yet written back. The snapshot captures only what happened to have been flushed. 3. **Application memory** — records, buffers and in-flight transactions the application has not handed to the filesystem at all. The snapshot captures none of this. LVM operates at layer 1. Creating the snapshot suspends the origin's device-mapper device long enough to flush in-flight block I/O and install the new mapping atomically, which is why you never see a torn half-old/half-new image at the block layer. It does not reach up into layers 2 and 3. ## Crash consistency, and why it is usually survivable The resulting image is exactly what the volume would look like if the power cable were pulled at that instant. Journalling filesystems are built for precisely this: ext4 and XFS write metadata changes to a journal before applying them, so mounting the snapshot replays the journal and yields a structurally sound filesystem. That is why crash-consistent snapshots "just work" most of the time. What survives the replay is a weaker statement than people assume. Structural integrity is restored; recent *content* may not be. A file that the application wrote and did not fsync may be absent, truncated, or present with stale contents. Two files that must be updated together may be captured in different states unless the application ordered its writes with fsync and expected a crash at any point. In short: crash consistency gives you what a crash-tolerant application would have survived, and nothing more. ## What `fsfreeze -f` adds `fsfreeze` (from util-linux) drives the kernel's filesystem freeze operation on a mount point: ``` fsfreeze -f /var/lib/app # flush and quiesce lvcreate -s -L 10G -n app-snap vg0/app fsfreeze -u /var/lib/app # thaw ``` Between the freeze and the thaw the kernel has written back the filesystem's dirty data and metadata and put the filesystem into a consistent, quiescent state; new write attempts block rather than proceeding. The snapshot taken in that window is therefore a *cleanly unmounted-looking* filesystem image: it mounts without journal replay, and it does not contain a half-finished metadata operation. ext4 and XFS both support this, and taking the snapshot is fast, so the window is milliseconds to a second or two. The operational rules around it are strict, and this is where an interviewer is listening: - **Freeze for as short a time as possible.** Every writer to that filesystem blocks while it is frozen. A frozen root filesystem can wedge the machine, including the very shell you would use to thaw it — always run freeze/snapshot/thaw as one scripted sequence, never as three interactive commands you might get distracted between. - **Always pair the thaw**, including on failure. Put the `fsfreeze -u` in a trap or an unconditional cleanup path so a failing `lvcreate` does not leave the filesystem frozen. - **Do not freeze what you cannot afford to block.** For a latency-sensitive service, even a second of blocked writes may be more disruptive than a journal replay on the restore side. ## What freeze still does not do Freezing quiesces the *filesystem*, not the *applications*. Anything holding state in its own memory — buffered records, an open write it intends to complete, a cache it plans to flush later — is not represented in the snapshot. Making a snapshot usable at the application level means telling the application to flush and reach a consistent point before you freeze, and only that application can define what such a point is. That coordination is the application's own problem and its own procedure; from the storage side, your obligation is to give it a window and to freeze after it says it is ready. ## Putting it together A defensible snapshot-backup sequence looks like: ask the application to quiesce or checkpoint → `fsfreeze -f` → `lvcreate -s` → `fsfreeze -u` (unconditionally) → release the application → mount the snapshot read-only → stream it off-host → verify → `lvremove`. The freeze window contains exactly one cheap command. And the honest fallback, when a freeze is impossible, is to accept crash consistency and to prove it works by actually restoring from it — a restore test is the only evidence that "the journal will replay" is true for your workload.

  • Why is a crash-consistent snapshot of an ext4 or XFS filesystem usually still mountable?
    Because both are journalling filesystems: metadata changes are recorded in a journal before being applied, so an abrupt stop leaves either the old state or a replayable record of the new one. Mounting the snapshot replays that journal and produces a structurally sound filesystem. Structural soundness is the guarantee; recent unflushed file contents are not.
  • What is the danger of running `fsfreeze -f` interactively on a busy filesystem?
    Every writer to that filesystem blocks until you thaw it, so a pause between your commands turns into an outage — and if you freeze a filesystem the system itself needs, you can block the tooling you would use to recover. Script freeze, snapshot and thaw as one unit with the thaw in an unconditional cleanup path.
  • Does freezing the filesystem make the snapshot usable by the application that owns the data?
    Not by itself. The freeze flushes what the filesystem holds; it cannot reach state the application never wrote. The application has to reach a checkpoint of its own choosing first, and only it can define what that means. The storage side's job is to provide a short window after the application says it is ready.

saying these in an interview costs you the question

  • Claims LVM flushes application buffers for you
  • Thinks a mountable snapshot means the data is complete
  • Leaves the filesystem frozen after a failed snapshot
  • Freezes for the whole duration of the backup copy
  • Believes crash-consistent and application-consistent are the same

context