skip to content

You run `lvcreate -s -L 5G -n db-snap vg0/db` on a Linux host. What has LVM created, and what happens on disk the first time a block of the origin volume is overwritten afterwards?

level: middleimportance: must knowfreq 62%

answer

  1. nothing is copied at creation
  2. exception table maps origin chunk to COW chunk
  3. old data is preserved, not new
  4. extra read plus extra write on first touch
  5. grows with change, not with volume size

basics

~20 s

LVM allocated a 5 GiB copy-on-write area and a snapshot device; no data was copied. The first overwrite of a chunk on the origin makes device-mapper read the old chunk, write it into the copy-on-write area, record the exception, then apply the new write.

solid answer

~50 s

`lvcreate -s` allocates 5 GiB of free extents in `vg0` as a copy-on-write exception area and creates a device-mapper snapshot device at `/dev/vg0/db-snap`. Nothing is copied at creation, which is why snapshotting a multi-terabyte volume is instant. The snapshot's contents are defined as "the origin, as of now". When a chunk of the origin is written for the first time after that, device-mapper preserves the *old* version: it reads the chunk, writes it into the COW area, updates the exception table, and only then lets the new data land on the origin. Reads through the snapshot check the exception table — modified chunks come from the COW area, everything else is redirected to the origin. So the origin pays roughly an extra read plus an extra write on each chunk's first modification, and the snapshot grows in proportion to how much of the origin changes, not to how big the origin is.

code

bash · 12 lines
bash
# create a 5 GiB copy-on-write snapshot of vg0/db (returns immediately)
lvcreate -s -L 5G -n db-snap vg0/db

# the snapshot is an ordinary block device presenting the origin's size
mount -o ro /dev/vg0/db-snap /mnt/snap

# watch how much of the copy-on-write area has been consumed
lvs -o lv_name,lv_size,origin,data_percent vg0

# tear it down when the copy is finished
umount /mnt/snap
lvremove -f vg0/db-snap

go deeper

for a junior

Know that creating a snapshot copies nothing and is instant, and that the snapshot only fills up as the origin changes. Say that the space you pass with -L is for preserved old data.

for a middle

Walk the write path precisely: read the old chunk, write it to the copy-on-write area, commit the exception, then apply the new write — and note that the second write to the same chunk is free. Explain the read path's redirect to the origin.

for a senior

Bring the operational consequences: the write penalty during a snapshot window, why it multiplies with several classic snapshots, and why long-lived snapshots on a busy origin are a capacity and latency decision you should be sizing and monitoring deliberately.

for a principal

Own the choice of mechanism: classic copy-on-write snapshots versus thin snapshots versus array- or filesystem-level snapshots, argued from write amplification, retention needs and blast radius rather than from familiarity with one tool.

## What exists after the command returns Three things now exist in `vg0`: - the origin logical volume `vg0/db`, unchanged in content; - a hidden copy-on-write logical volume of 5 GiB, carved from the volume group's free extents; - a new visible device, `/dev/vg0/db-snap`, which you can mount (normally read-only) and read as if it were a frozen clone of the origin. `lvs` will now report `db` with the origin/snapshot relationship and show `db-snap` with an `Origin` of `db` and a `Data%` that starts near zero: ``` $ lvs -o lv_name,lv_size,origin,data_percent vg0 LV LSize Origin Data% db 500.00g db-snap 5.00g db 0.01 ``` No bulk copy happened. That is the first thing to say out loud, because a lot of candidates assume snapshot creation duplicates data and are then confused about why it is instantaneous. ## The copy-on-write write path The origin and the snapshot are both fronted by device-mapper targets. The unit of work is a *chunk* — a fixed block size chosen at creation and adjustable with `--chunksize`. Device-mapper maintains an **exception table**: a map from "chunk N of the origin" to "chunk M of the COW area", holding the pre-snapshot contents of any chunk that has since been modified. A write to a chunk of the origin that has no exception yet triggers, in order: 1. read the current (old) contents of that chunk from the origin; 2. write those old contents into a free slot in the COW area; 3. commit the new entry in the exception table (metadata is also stored in the COW LV); 4. write the new data to the origin. A write to a chunk that already has an exception skips all of that and goes straight through — the old version is already preserved, so it only ever gets copied once. This is why a workload that repeatedly rewrites the same hot region pays the penalty briefly and then runs at close to normal speed, while a workload that touches fresh regions across the volume keeps paying. The cost is real: at worst one extra read and one extra write per chunk first-touched, plus metadata commits. Order-of-magnitude write slowdowns on a busy volume during a snapshot window are unremarkable, and the effect multiplies with the number of classic snapshots — each snapshot has its own COW area and needs its own copy of the old chunk, so N snapshots means up to N copies before the write can proceed. ## The read path Reading `/dev/vg0/db-snap` consults the exception table. If the chunk has an exception, the read is served from the COW area (the old, preserved contents). If it does not, the read is redirected to the origin, because that chunk has not changed since the snapshot was created and the origin still holds the correct point-in-time data. This redirect is exactly why the snapshot cannot outlive the origin. Reading the *origin* is unaffected — it is served directly, with no exception lookup on the read path. ## Chunk size, and what the 5G means The `-L 5G` sizes the COW area, not the visible size of the snapshot. `/dev/vg0/db-snap` presents the same logical size as the origin (500 GiB in the listing above); the 5 GiB is the budget for preserved old chunks. Sizing it is therefore a question about the origin's *write* volume during the snapshot's lifetime, never about the origin's capacity. Chunk size matters at the margins: larger chunks mean fewer, larger copies and a smaller exception table but more write amplification when writes are small and scattered (touching one 4 KiB block still copies a whole chunk). Smaller chunks reduce the copied volume but grow the metadata. You can set it explicitly with `lvcreate -s --chunksize`. ## The consequences worth naming - **Creation is O(1)**, deletion is O(1); the cost is spread over the origin's writes while the snapshot lives. - **The snapshot grows with change, not with data**, so a quiet volume can be snapshotted with a tiny COW area and a write-heavy one cannot. - **A full COW area is fatal to the snapshot** — device-mapper marks it invalid and reads from it fail, which is the classic production incident on this feature. - **Long-lived classic snapshots are a performance decision**, not a free one; this is the main reason thin snapshots, which redirect writes to newly allocated pool blocks instead of copying old ones, replaced them for anything long-running.

  • Why does the origin get slower while a snapshot exists, and does the penalty stay constant?
    Each first write to a chunk becomes read-old-chunk, write-it-to-the-COW-area, commit metadata, then do the real write. Chunks that already have an exception are written straight through, so the penalty is heaviest early and on workloads that keep touching fresh regions. With several classic snapshots the old chunk must be copied into each one, so the cost scales with snapshot count.
  • Does the visible size of `/dev/vg0/db-snap` equal the 5G passed to `-L`?
    No. The snapshot device presents the same logical size as the origin — mount it and you see the origin's filesystem at its full size. The `-L 5G` sizes only the copy-on-write area that holds preserved old chunks, which is a budget for how much of the origin may change before the snapshot becomes invalid.
  • What does `--chunksize` trade off?
    Larger chunks mean a smaller exception table and fewer, bigger copy operations, but more write amplification: modifying a few kilobytes still copies a whole chunk into the COW area. Smaller chunks copy less data per write but grow the metadata and the lookup cost. Scattered small writes favour smaller chunks; large sequential rewrites tolerate bigger ones.

saying these in an interview costs you the question

  • Claims the snapshot is copied out at creation time
  • Says new writes go into the snapshot area
  • Thinks the -L size is how much data the snapshot shows
  • Assumes reads from the snapshot never touch the origin
  • Believes snapshots are free for the origin's performance

context