How does a thin snapshot of an LVM thin logical volume differ from a classic `lvcreate -s -L` snapshot in the way it stores data, and why can you keep dozens of thin snapshots but not dozens of classic ones?
answer
- fixed COW area versus shared pool
- copy the old block versus redirect the new one
- cost per snapshot versus cost per write
- many cheap snapshots, one shared risk
- no -L when snapshotting a thin volume
basics
~20 sA classic LVM snapshot owns a fixed copy-on-write area and preserves old chunks by copying them into it. A thin snapshot just shares block references in the thin pool: a write allocates a new pool block, nothing is copied, and the cost does not grow with the number of snapshots.
solid answer
~60 sA classic snapshot is created with its own sized copy-on-write area, and every first write to an origin chunk copies the old data into that area before the write proceeds. With several classic snapshots, that old chunk must be copied into each one's area, so origin writes get slower and space usage grows roughly linearly with snapshot count — and each snapshot can independently fill up and be invalidated. A thin snapshot of a thin LV is created with no size at all: `lvcreate -s -n snap vg0/thinlv` just clones the volume's block-mapping tree inside the thin pool, so origin and snapshot initially point at the same pool blocks. A write to a shared block allocates a *new* block from the pool and repoints the writer at it — a redirect, not a copy — which costs the same whether one snapshot or thirty are referencing the block. That is why thin snapshots support real retention chains and snapshots of snapshots, at the price of a single shared risk: if the pool itself runs out of data or metadata space, everything using it is affected at once.
code
bash · 12 lines# a thin pool, a 400G thin volume inside a 200G pool, and a snapshot of it
lvcreate -L 200G -T vg0/pool
lvcreate -V 400G -T vg0/pool -n app
# no size argument: the snapshot shares the origin's pool blocks
lvcreate -s -n app-snap vg0/app
# thin snapshots carry the activation-skip flag; -K overrides it
lvchange -ay -K vg0/app-snap
# pool usage is the number that matters, not per-snapshot usage
lvs -o lv_name,lv_size,data_percent,metadata_percent vg0go deeper
Know that a classic snapshot is created with a fixed size for preserved data while a thin snapshot takes no size and draws from a shared pool, and that thin snapshots are cheap enough to keep several of.
Contrast the mechanics precisely: copying the old chunk into a per-snapshot area versus allocating a new pool block for the writer and leaving the shared block alone, and derive from that why classic cost scales with snapshot count and thin cost does not.
Make the risk argument: thin trades many individually-sized budgets for one shared pool whose data and metadata usage must be monitored, and explain what you would put in place — thresholds, autoextend, VG headroom, discard handling — before relying on it.
Decide the storage standard and defend it: whether the fleet's retention and rollback needs justify thin provisioning's concentrated failure mode, who owns pool capacity, and how over-commit ratios are governed rather than left to whoever creates the next volume.
## Two different data structures **Classic (dm-snapshot).** `lvcreate -s -L 5G -n snap vg0/lv` allocates a separate copy-on-write logical volume from the volume group. Device-mapper keeps an exception table mapping origin chunks to preserved copies in that COW LV. The preservation is a genuine *copy-on-write*: before an origin chunk is first modified, its old contents are read and written into the COW area, then the new write proceeds. **Thin (dm-thin).** A thin pool created with `lvcreate -L 200G -T vg0/pool` owns a data area and a metadata area. Every thin volume in the pool is a *mapping tree* from its logical blocks to blocks in the pool's data area. `lvcreate -s -n snap vg0/app` — note: no `-L` — creates a new thin volume whose mapping tree starts as a clone of `app`'s, so both volumes reference exactly the same pool blocks and the snapshot costs almost nothing but a little metadata. When a block referenced by more than one thin volume is written, dm-thin does not copy the old block anywhere. It **allocates a fresh block from the pool, writes the new data there, and updates the writing volume's mapping tree to point at it**. The other volumes keep pointing at the original block, which still holds the old contents. This is redirect-on-write rather than copy-on-write, and it is the whole difference. ## Why the counts diverge With classic snapshots, N snapshots of the same origin means the pre-write contents of a chunk must be preserved into N separate exception stores. Origin write latency grows with N, and space consumption grows with N. Three or four concurrent classic snapshots of a busy volume is already an operational problem. With thin snapshots, a first write after a snapshot costs one pool allocation and one metadata update regardless of how many snapshots reference the block. Space is charged once, to the pool, for the newly written data — not once per snapshot. So hourly snapshots with a rolling retention, or a chain of snapshots-of-snapshots for testing, are practical. ``` # a pool, a thin volume, then a snapshot of it (no size argument) lvcreate -L 200G -T vg0/pool lvcreate -V 400G -T vg0/pool -n app lvcreate -s -n app-snap vg0/app ``` A practical wrinkle: thin snapshots are created with the activation-skip flag set, so they are not activated automatically. Use `lvchange -ay -K vg0/app-snap` (the `-K` overrides the skip) before you can open the device. ## Where each one fails - **Classic:** the failure is *per-snapshot*. Fill the copy-on-write area and that snapshot is marked invalid; the origin keeps running. Blast radius is small, but the snapshot has an individual budget you must size correctly, and long-lived snapshots on a churning origin are a losing bet. - **Thin:** the failure is *poolwide*. Because thin volumes over-commit shared space, running the pool out of data blocks — or out of metadata — affects every volume and snapshot in it at once, typically with I/O errors and filesystems flipping read-only. Nothing is individually invalidated; instead the whole pool needs headroom management, `Data%` and `Meta%` monitoring, and autoextend policy. So the tradeoff is not "thin is better" but "thin trades many small, individually-sized budgets for one large shared one". Concentrating risk is a good deal when you monitor the pool and keep free extents in the volume group; it is a bad deal when nobody watches it. ## Other practical differences - **Rollback.** Both support merging a snapshot back over its origin, but with different commands: `lvconvert --merge` for classic snapshots, `lvconvert --mergethin` for thin ones. - **Writability.** Both can be writable, but writable classic snapshots consume their COW area from both sides. Thin snapshots are ordinary thin volumes once created, so a writable clone for a test environment is a natural use. - **Reclaim.** Deleting files inside a thin volume does not return blocks to the pool by itself; the filesystem has to issue discards (mounted with discard support, or a periodic `fstrim`) and the pool has to pass them down. Classic snapshots have no equivalent — their space is returned when the snapshot is removed. - **Origin requirement.** A thin snapshot requires the origin to already be a thin volume in a pool; you cannot take a thin snapshot of an ordinary LV. Choosing thin provisioning is a decision made when the volume is created, not at snapshot time.
- Why does taking a thin snapshot complete almost instantly no matter how large the volume is?Because it only clones the volume's block-mapping tree inside the pool's metadata. Origin and snapshot start out referencing exactly the same data blocks, so no data is read or written. Divergence happens lazily, one block at a time, as either volume is written and dm-thin allocates a fresh pool block for the writer.
- If a thin snapshot never copies old blocks, where does the space it consumes come from?From the shared thin pool, charged once for each newly written block rather than once per snapshot. The pool's Data% is the number that matters, and it rises with total change across all thin volumes in the pool. Deleted files only return space if discards reach the pool, via a discard-enabled mount or a periodic fstrim.
- You have a fresh thin snapshot but its device node will not open. What is missing?Thin snapshots are created with the activation-skip flag set, so LVM deliberately does not activate them with the rest of the volume group. Activate it explicitly with `lvchange -ay -K vg/snap` — the `-K` overrides the skip flag — and then the device can be mounted or read.
saying these in an interview costs you the question
- Says thin snapshots also copy the old block somewhere
- Passes -L when snapshotting a thin volume
- Thinks each thin snapshot has its own size budget
- Assumes deleting files inside a thin volume frees pool space
- Believes thin snapshots remove the risk of running out of space