skip to content

Snapshots & Thin Provisioning

An LVM snapshot freezes a point-in-time view of a volume so you can back it up or roll an upgrade back, and thin pools let volumes over-commit shared space. Interviewers probe the copy-on-write mechanics and the classic "the snapshot filled up and got invalidated" failure.

on this pageshow

questions

7

You run `lvcreate -s -L 5G -n db-snap vg0/db` on a Linux host. What has LVM created, and what happens on disk the first time a block of the origin volume is overwritten afterwards?

level: middleimportance: must knowfreq 62%

answer

  1. nothing is copied at creation
  2. exception table maps origin chunk to COW chunk
  3. old data is preserved, not new
  4. extra read plus extra write on first touch
  5. grows with change, not with volume size

basics

~20 s

LVM allocated a 5 GiB copy-on-write area and a snapshot device; no data was copied. The first overwrite of a chunk on the origin makes device-mapper read the old chunk, write it into the copy-on-write area, record the exception, then apply the new write.

solid answer

~50 s

`lvcreate -s` allocates 5 GiB of free extents in `vg0` as a copy-on-write exception area and creates a device-mapper snapshot device at `/dev/vg0/db-snap`. Nothing is copied at creation, which is why snapshotting a multi-terabyte volume is instant. The snapshot's contents are defined as "the origin, as of now". When a chunk of the origin is written for the first time after that, device-mapper preserves the *old* version: it reads the chunk, writes it into the COW area, updates the exception table, and only then lets the new data land on the origin. Reads through the snapshot check the exception table — modified chunks come from the COW area, everything else is redirected to the origin. So the origin pays roughly an extra read plus an extra write on each chunk's first modification, and the snapshot grows in proportion to how much of the origin changes, not to how big the origin is.

code

bash · 12 lines
bash
# create a 5 GiB copy-on-write snapshot of vg0/db (returns immediately)
lvcreate -s -L 5G -n db-snap vg0/db

# the snapshot is an ordinary block device presenting the origin's size
mount -o ro /dev/vg0/db-snap /mnt/snap

# watch how much of the copy-on-write area has been consumed
lvs -o lv_name,lv_size,origin,data_percent vg0

# tear it down when the copy is finished
umount /mnt/snap
lvremove -f vg0/db-snap

go deeper

for a junior

Know that creating a snapshot copies nothing and is instant, and that the snapshot only fills up as the origin changes. Say that the space you pass with -L is for preserved old data.

for a middle

Walk the write path precisely: read the old chunk, write it to the copy-on-write area, commit the exception, then apply the new write — and note that the second write to the same chunk is free. Explain the read path's redirect to the origin.

for a senior

Bring the operational consequences: the write penalty during a snapshot window, why it multiplies with several classic snapshots, and why long-lived snapshots on a busy origin are a capacity and latency decision you should be sizing and monitoring deliberately.

for a principal

Own the choice of mechanism: classic copy-on-write snapshots versus thin snapshots versus array- or filesystem-level snapshots, argued from write amplification, retention needs and blast radius rather than from familiarity with one tool.

## What exists after the command returns Three things now exist in `vg0`: - the origin logical volume `vg0/db`, unchanged in content; - a hidden copy-on-write logical volume of 5 GiB, carved from the volume group's free extents; - a new visible device, `/dev/vg0/db-snap`, which you can mount (normally read-only) and read as if it were a frozen clone of the origin. `lvs` will now report `db` with the origin/snapshot relationship and show `db-snap` with an `Origin` of `db` and a `Data%` that starts near zero: ``` $ lvs -o lv_name,lv_size,origin,data_percent vg0 LV LSize Origin Data% db 500.00g db-snap 5.00g db 0.01 ``` No bulk copy happened. That is the first thing to say out loud, because a lot of candidates assume snapshot creation duplicates data and are then confused about why it is instantaneous. ## The copy-on-write write path The origin and the snapshot are both fronted by device-mapper targets. The unit of work is a *chunk* — a fixed block size chosen at creation and adjustable with `--chunksize`. Device-mapper maintains an **exception table**: a map from "chunk N of the origin" to "chunk M of the COW area", holding the pre-snapshot contents of any chunk that has since been modified. A write to a chunk of the origin that has no exception yet triggers, in order: 1. read the current (old) contents of that chunk from the origin; 2. write those old contents into a free slot in the COW area; 3. commit the new entry in the exception table (metadata is also stored in the COW LV); 4. write the new data to the origin. A write to a chunk that already has an exception skips all of that and goes straight through — the old version is already preserved, so it only ever gets copied once. This is why a workload that repeatedly rewrites the same hot region pays the penalty briefly and then runs at close to normal speed, while a workload that touches fresh regions across the volume keeps paying. The cost is real: at worst one extra read and one extra write per chunk first-touched, plus metadata commits. Order-of-magnitude write slowdowns on a busy volume during a snapshot window are unremarkable, and the effect multiplies with the number of classic snapshots — each snapshot has its own COW area and needs its own copy of the old chunk, so N snapshots means up to N copies before the write can proceed. ## The read path Reading `/dev/vg0/db-snap` consults the exception table. If the chunk has an exception, the read is served from the COW area (the old, preserved contents). If it does not, the read is redirected to the origin, because that chunk has not changed since the snapshot was created and the origin still holds the correct point-in-time data. This redirect is exactly why the snapshot cannot outlive the origin. Reading the *origin* is unaffected — it is served directly, with no exception lookup on the read path. ## Chunk size, and what the 5G means The `-L 5G` sizes the COW area, not the visible size of the snapshot. `/dev/vg0/db-snap` presents the same logical size as the origin (500 GiB in the listing above); the 5 GiB is the budget for preserved old chunks. Sizing it is therefore a question about the origin's *write* volume during the snapshot's lifetime, never about the origin's capacity. Chunk size matters at the margins: larger chunks mean fewer, larger copies and a smaller exception table but more write amplification when writes are small and scattered (touching one 4 KiB block still copies a whole chunk). Smaller chunks reduce the copied volume but grow the metadata. You can set it explicitly with `lvcreate -s --chunksize`. ## The consequences worth naming - **Creation is O(1)**, deletion is O(1); the cost is spread over the origin's writes while the snapshot lives. - **The snapshot grows with change, not with data**, so a quiet volume can be snapshotted with a tiny COW area and a write-heavy one cannot. - **A full COW area is fatal to the snapshot** — device-mapper marks it invalid and reads from it fail, which is the classic production incident on this feature. - **Long-lived classic snapshots are a performance decision**, not a free one; this is the main reason thin snapshots, which redirect writes to newly allocated pool blocks instead of copying old ones, replaced them for anything long-running.

  • Why does the origin get slower while a snapshot exists, and does the penalty stay constant?
    Each first write to a chunk becomes read-old-chunk, write-it-to-the-COW-area, commit metadata, then do the real write. Chunks that already have an exception are written straight through, so the penalty is heaviest early and on workloads that keep touching fresh regions. With several classic snapshots the old chunk must be copied into each one, so the cost scales with snapshot count.
  • Does the visible size of `/dev/vg0/db-snap` equal the 5G passed to `-L`?
    No. The snapshot device presents the same logical size as the origin — mount it and you see the origin's filesystem at its full size. The `-L 5G` sizes only the copy-on-write area that holds preserved old chunks, which is a budget for how much of the origin may change before the snapshot becomes invalid.
  • What does `--chunksize` trade off?
    Larger chunks mean a smaller exception table and fewer, bigger copy operations, but more write amplification: modifying a few kilobytes still copies a whole chunk into the COW area. Smaller chunks copy less data per write but grow the metadata and the lookup cost. Scattered small writes favour smaller chunks; large sequential rewrites tolerate bigger ones.

saying these in an interview costs you the question

  • Claims the snapshot is copied out at creation time
  • Says new writes go into the snapshot area
  • Thinks the -L size is how much data the snapshot shows
  • Assumes reads from the snapshot never touch the origin
  • Believes snapshots are free for the origin's performance

context

open as a page

A nightly job snapshots a busy Linux volume so a backup can stream from it. This morning `lvs` shows the snapshot at 100% Data% with an `I` in its attribute string, and the backup that ran from it is unusable. What happened, and how do you size and monitor the snapshot so it does not recur?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The origin changed more than the snapshot's copy-on-write area could hold, so LVM marked the snapshot invalid and reads from it now fail. Size the COW area for the origin's write volume during the snapshot's lifetime, monitor Data%, and enable snapshot autoextend with lvm2-monitor running.

open as a page

A colleague says the nightly LVM snapshot of the server's data volume is the backup. Why is an LVM snapshot sitting in the same volume group not a backup, and what is it actually good for?

level: juniorimportance: should knowfreq 48%

basics

~20 s

An LVM snapshot stores only the blocks that changed since it was taken and reads everything else from the origin volume, on the same disks in the same volume group. Lose the origin or the disks and the snapshot is worthless. It is a stable point-in-time source to copy from, not a copy.

open as a page

How does a thin snapshot of an LVM thin logical volume differ from a classic `lvcreate -s -L` snapshot in the way it stores data, and why can you keep dozens of thin snapshots but not dozens of classic ones?

level: middleimportance: should knowfreq 45%

basics

~20 s

A classic LVM snapshot owns a fixed copy-on-write area and preserves old chunks by copying them into it. A thin snapshot just shares block references in the thin pool: a write allocates a new pool block, nothing is copied, and the cost does not grow with the number of snapshots.

open as a page

Before a risky package upgrade you snapshot the Linux server's root-adjacent data volume with `lvcreate -s`. The upgrade goes badly and you want to roll back. What does `lvconvert --merge vg0/snap` do, and why might the rollback not take effect the moment the command returns?

level: seniorimportance: should knowfreq 40%

basics

~20 s

lvconvert --merge schedules the snapshot's preserved contents to be written back over the origin, reverting it to the snapshot's point in time and deleting the snapshot when finished. If the origin is open — mounted or in use — the merge is deferred until the volume is next deactivated and reactivated.

open as a page

An LVM thin pool created with `lvcreate -L 200G -T vg0/pool` backs thin volumes whose virtual sizes total 800G. What happens to those volumes when the pool's data space — or its metadata space — actually runs out, and how do you keep that from happening?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Once the pool has no free blocks, any write needing a new allocation fails or hangs, filesystems on the thin volumes go read-only or corrupt, and metadata exhaustion is worse — the pool goes read-only and needs offline repair. Prevent it with Data%/Meta% alerting, autoextend, and volume group headroom.

open as a page

You take an LVM snapshot of a mounted, actively written filesystem on a Linux host. What consistency does the resulting snapshot actually give you, and what does running `fsfreeze -f` on the mount point immediately before the snapshot add?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The snapshot is crash-consistent: it looks exactly like the volume would after a power cut, so a journalling filesystem mounts it after replaying its journal, but anything still buffered above the block layer is missing. fsfreeze -f flushes and quiesces the filesystem first, giving a cleanly-shut-down image instead.

open as a page