A nightly job snapshots a busy Linux volume so a backup can stream from it. This morning `lvs` shows the snapshot at 100% Data% with an `I` in its attribute string, and the backup that ran from it is unusable. What happened, and how do you size and monitor the snapshot so it does not recur?
answer
- COW area sized for change, not capacity
- the I in lvs Attr means invalid
- origin survives, snapshot is sacrificed
- longer backup window, larger exposure
- autoextend needs dmeventd and free extents
basics
~20 sThe origin changed more than the snapshot's copy-on-write area could hold, so LVM marked the snapshot invalid and reads from it now fail. Size the COW area for the origin's write volume during the snapshot's lifetime, monitor Data%, and enable snapshot autoextend with lvm2-monitor running.
solid answer
~60 sA classic LVM snapshot can only preserve as much changed data as its copy-on-write area holds. When the origin's writes exhaust it, device-mapper cannot preserve any more old chunks, so it drops the snapshot: `lvs` reports 100% `Data%` and an `I` (invalid) in the attributes, and every read through the snapshot device errors out. The origin itself is fine and keeps serving traffic — only the snapshot and anything reading it are lost, which is why the backup is garbage. The fix has three parts. Size the COW area against the volume of writes the origin takes during the snapshot's lifetime, not against the origin's capacity — measure it rather than guessing. Shorten that lifetime, so the backup streams and the snapshot is removed in minutes. And set `snapshot_autoextend_threshold` and `snapshot_autoextend_percent` in `lvm.conf` with `lvm2-monitor` running so dmeventd grows the COW area before it fills, keeping spare free extents in the volume group for it to grow into. Alert on `Data%` too — an invalid snapshot should never be discovered by a failed restore.
code
bash · 12 lines# how full is each snapshot's copy-on-write area, and is any invalid?
lvs -o lv_name,origin,lv_attr,data_percent vg0
# read back the effective autoextend policy (not just the file you edited)
lvmconfig activation/snapshot_autoextend_threshold \
activation/snapshot_autoextend_percent
# autoextend only happens while dmeventd monitoring is running
systemctl status lvm2-monitor
# and only if the volume group still has free extents to grow into
vgs -o vg_name,vg_size,vg_free vg0go deeper
Know that a snapshot has a limited copy-on-write area, that it is consumed by writes to the origin, and that when it fills the snapshot becomes unusable while the origin keeps working.
Explain why consumption tracks the origin's changed chunks rather than its size, read the Data% and the invalid state flag out of lvs, and describe the autoextend settings and the dmeventd monitoring they depend on.
Show the production judgment: measure the write volume over the real window, shorten the snapshot's lifetime as the first lever, keep free extents in the volume group, alert on Data%, and make the backup verify rather than trust its exit code.
Argue the mechanism choice for the fleet — classic snapshots versus thin pools versus array-level snapshots — and set the capacity policy that keeps volume groups from being allocated to the last extent, since that reserve is what makes the safety net work at all.
## What the symptoms mean Two fields in `lvs` tell the whole story: ``` $ lvs -o lv_name,origin,lv_attr,data_percent vg0 LV Origin Attr Data% db owi-aos--- db-snap db swi-I-s--- 100.00 ``` `Data%` is the fraction of the copy-on-write area consumed by preserved old chunks. The fifth character of `Attr` is the volume's *state*, and `I` means **invalid snapshot**. Once that flag appears, the snapshot device returns errors for reads; there is no repair, only `lvremove`. Anything that was mid-stream from it — your backup — is truncated or corrupt. The important reassurance: the **origin is untouched**. Device-mapper stops preserving old chunks because it has nowhere to put them, and the snapshot is sacrificed so the origin can carry on. A production database on the origin does not go read-only because of this. The damage is confined to the snapshot and to whatever trusted it. ## Why it filled The COW area consumes space at the rate the *origin changes*, once per chunk. A 5 GiB COW area does not mean "5 GiB of the volume are protected"; it means "the origin may have up to roughly 5 GiB worth of distinct chunks modified before I die". Three things push a job over that line: - **Write volume.** A busy volume can rewrite tens of gigabytes in an hour. - **Snapshot lifetime.** A backup that used to take 20 minutes and now takes six hours has multiplied the exposure without anyone changing a setting. This is the most common cause of a job that worked for a year and then did not. - **Write scatter and chunk size.** Because preservation happens per chunk, small writes spread across the volume consume far more COW space than the same byte count written sequentially. A nightly `fstrim`, a defragmentation pass, a compaction job or an index rebuild inside the snapshot window can touch a huge fraction of the volume. ## Sizing it honestly Size from measurement, not intuition: 1. Determine the write volume over the window. Per-device write counters — for example the sectors-written figures exposed under `/sys/block/<dev>/stat`, or a period of `iostat -x` over a representative night — give you gigabytes-written for the period. 2. Inflate for chunk-level amplification and for the fact that repeated writes to the same chunk cost nothing but scattered ones cost a full chunk each. A common rule of thumb is to allow a healthy multiple of the measured figure. 3. Add headroom for the abnormal night — the migration, the bulk import, the reindex. 4. Cross-check with reality: run the job and record the peak `Data%` for a week. If peaks sit at 70–90%, you are one busy night from an incident. The cheapest lever is usually not a bigger COW area but a **shorter window**. If the backup can read from the snapshot at full speed and finish in ten minutes, the exposure shrinks by an order of magnitude. ## Autoextend and monitoring LVM can grow the COW area on its own, driven by dmeventd: ``` # in /etc/lvm/lvm.conf, activation section snapshot_autoextend_threshold = 70 snapshot_autoextend_percent = 20 ``` With `snapshot_autoextend_threshold` below 100 and the monitoring service active, crossing the threshold triggers an extension by `snapshot_autoextend_percent`. Two preconditions bite in practice: the monitoring daemon must actually be running (`systemctl status lvm2-monitor`), and the **volume group must have free extents** to grow into. A VG allocated to 100% turns autoextend into a no-op, and the snapshot dies anyway. Keeping deliberate free space in the VG is part of the design, not slack. Check the effective settings with `lvmconfig` rather than trusting the file you edited, since distributions ship layered configuration under `lvm.conf.d`-style directories. Monitoring is the other half. Emit `Data%` per snapshot from `lvs` and alert at, say, 80%. The failure you are protecting against is silent: nothing pages you when a snapshot is invalidated, and the backup job may exit zero having written a corrupt stream. Whatever consumes the snapshot should also **verify** — a restore test, or at minimum a check that the snapshot's `Attr` has no `I` at the end of the run. ## The structural fix If the workload genuinely churns, stop using classic snapshots for it. A **thin snapshot** in an LVM thin pool has no fixed COW area: writes allocate new blocks from the shared pool, so the snapshot does not have an individual budget to blow through, and keeping several of them is affordable. That moves the exhaustion risk to the pool as a whole — a different, poolwide problem with its own monitoring — but it removes the specific failure where one long-running backup silently destroys its own source.
- Can you recover an LVM snapshot after it has been marked invalid?No. Once the copy-on-write area is exhausted, device-mapper stops preserving old chunks and the point-in-time image is no longer reconstructable. Extending the logical volume afterwards does not undo it. The only action is `lvremove` and a fresh snapshot. That irreversibility is exactly why the alert threshold matters more than the recovery procedure.
- The volume group is fully allocated. What does that do to snapshot autoextend?It disables it in practice. Autoextend can only grow the copy-on-write logical volume if the volume group has free extents; with none, dmeventd has nothing to allocate and the snapshot fills and is invalidated exactly as if autoextend were off. Reserving free space in the VG is a prerequisite of the policy, not spare capacity to be consumed.
- Why did a job that ran fine for a year suddenly start failing without any configuration change?Usually the snapshot's lifetime grew or the origin's write rate did. A backup that slowed from twenty minutes to several hours multiplies the volume of origin writes the copy-on-write area must absorb. A new nightly compaction, reindex or bulk import inside the same window has the same effect by touching far more distinct chunks.
- How would you catch this failure before a restore does?Alert on the snapshot's `Data%` from `lvs` at around 80%, and have the backup job check the snapshot's attributes for the invalid flag before it reports success. Beyond that, test restores on a schedule — the only proof that a backup chain works is having restored from it recently.
saying these in an interview costs you the question
- Thinks the origin volume is damaged or goes read-only
- Sizes the COW area as a fraction of the origin's capacity
- Believes an invalid snapshot can be repaired by extending it
- Enables autoextend in a volume group with no free extents
- Assumes a zero exit code means the backup is valid