skip to content

An LVM thin pool created with `lvcreate -L 200G -T vg0/pool` backs thin volumes whose virtual sizes total 800G. What happens to those volumes when the pool's data space — or its metadata space — actually runs out, and how do you keep that from happening?

level: seniorimportance: should knowfreq 42%

answer

  1. virtual size promised, real blocks finite
  2. reads fine, new allocations fail
  3. metadata full is worse than data full
  4. autoextend needs free extents in the VG
  5. deleted files need discard to return blocks

basics

~20 s

Once the pool has no free blocks, any write needing a new allocation fails or hangs, filesystems on the thin volumes go read-only or corrupt, and metadata exhaustion is worse — the pool goes read-only and needs offline repair. Prevent it with Data%/Meta% alerting, autoextend, and volume group headroom.

solid answer

~60 s

Over-provisioning means the pool promises 800G it does not have. Reads and rewrites of already-allocated blocks keep working, but the moment a write needs a fresh block and the pool has none, that I/O cannot be satisfied: depending on the pool's `--errorwhenfull` setting it either queues for a while and then errors, or errors immediately, and the filesystems above it react the way they react to any I/O error — XFS shuts down, ext4 remounts read-only or worse. Every thin volume and snapshot in the pool is hit at once, because the shortage is shared. Running out of *metadata* is the nastier variant: the pool itself goes read-only, and recovering typically means taking it offline and running `thin_check`/`thin_repair` against a copy of the metadata. Prevention is the whole job: monitor `Data%` and `Meta%` from `lvs` with alerts well below full, set `thin_pool_autoextend_threshold` and `thin_pool_autoextend_percent` with `lvm2-monitor` running, keep genuinely free extents in the volume group for autoextend to consume, size the metadata area generously, and make sure discards reach the pool — via a discard-capable mount or a scheduled `fstrim` — so deleted data actually returns blocks.

code

bash · 12 lines
bash
# both numbers matter: data blocks and mapping metadata
lvs -o lv_name,lv_size,data_percent,metadata_percent vg0

# emergency capacity, from free extents in the volume group
lvextend -L +100G vg0/pool
lvextend --poolmetadatasize +1G vg0/pool

# fail loudly instead of queueing when the pool is full
lvchange --errorwhenfull y vg0/pool

# return deleted blocks to the pool from the filesystem above it
fstrim -v /var/lib/app

go deeper

for a junior

Know that thin volumes promise more space than the pool has, and that once the pool is full writes needing new space fail and the filesystems above it stop accepting writes.

for a middle

Explain which I/O still succeeds when the pool is full and which cannot, read Data% and Meta% from lvs, and describe the autoextend settings and the monitoring daemon that has to be running for them to fire.

for a senior

Demonstrate that you have run this: alert thresholds well below full, volume group headroom reserved for autoextend, discards reaching the pool via fstrim, bounded snapshot retention, and a rehearsed extend-or-drop-a-snapshot runbook for the page at 99%.

for a principal

Own the over-commit policy across the fleet — what ratio is acceptable given the monitoring actually in place, who is accountable for pool capacity, and whether concentrating many workloads' failure mode into one shared pool is a trade you want at all.

## What over-provisioning actually promises A thin pool has a real data area (200 GiB here) and a real metadata area. Thin volumes created with `lvcreate -V 400G -T vg0/pool -n app` present a *virtual* size; blocks are allocated from the pool only when written. Summing virtual sizes to 800G against a 200G pool is legitimate and is the point of the feature — but it is a bet that the volumes will not all fill. Usage is visible in two columns: ``` $ lvs -o lv_name,lv_size,data_percent,metadata_percent vg0 LV LSize Data% Meta% pool 200.00g 71.40 18.20 app 400.00g 34.10 ``` `Data%` is the pool's data area consumption; `Meta%` is its metadata area consumption. Both matter, and people watch only the first. ## Data exhaustion When `Data%` reaches 100: - **Reads still work**, and so do overwrites of blocks that are already mapped — no new allocation is needed. - **Any write requiring a new block cannot be satisfied.** That includes all writes to a region never written before, and *every* first write to a block shared with a snapshot, because redirect-on-write must allocate. - The pool's behaviour on that condition is configurable: `lvchange --errorwhenfull y|n vg0/pool`. With errors enabled, I/O fails immediately and loudly. With the default queueing behaviour, I/O is held for a grace period in the hope that space appears, and then fails — which presents as a hang first and errors second, and is often the more confusing symptom on a pager. - Filesystems then do what they do with I/O errors: XFS shuts the filesystem down; ext4 follows its `errors=` mount option, typically remounting read-only. Applications see write failures. Data written into the page cache but never successfully flushed is lost. Because the pool is shared, this hits every thin volume and snapshot in it simultaneously. That is the structural difference from a classic snapshot, where filling the copy-on-write area invalidates exactly one snapshot and leaves the origin alone. ## Metadata exhaustion The metadata area holds the mapping trees. It is small and easy to ignore, and it grows with the *number of mappings*, which is driven by snapshot count, small chunk sizes, and how much the volumes have diverged. If `Meta%` reaches 100, dm-thin puts the pool into read-only mode to protect the metadata it still has. That is not something you resolve by adding data space; recovery generally means deactivating the pool, working on a copy of the metadata with `thin_check` and `thin_repair` from the device-mapper persistent-data tools, and restoring the repaired metadata. Plan the metadata area up front (`lvcreate --poolmetadatasize`, extendable later with `lvextend --poolmetadatasize`) rather than discovering it during an incident. ## Prevention, in the order it pays off 1. **Alert on both numbers.** Export `Data%` and `Meta%` from `lvs` and page below the cliff — 75–80% gives room to act. This is the single highest-value control. 2. **Enable autoextend and make it possible.** `thin_pool_autoextend_threshold` and `thin_pool_autoextend_percent` in `lvm.conf`, with `lvm2-monitor` running so dmeventd can act. Autoextend allocates from the volume group, so the VG must retain free extents. A VG allocated to the last extent silently disables the safety net. 3. **Make discards reach the pool.** Deleting a file does not return blocks by itself. Either mount the filesystem with discard support or run `fstrim` on a schedule, and keep the pool's discard handling at the default that passes them down. Without this, a pool that has churned through data climbs monotonically even though the filesystems look half empty. 4. **Bound snapshot retention.** Every retained snapshot pins the blocks it references, so old snapshots are the usual reason a pool grows while the live data does not. 5. **Govern the over-commit ratio.** 4:1 on a pool nobody watches is reckless; 4:1 on a pool with alerting, autoextend and VG headroom is a considered decision. Write down which one you are running. 6. **Have the emergency runbook.** `lvextend -L +100G vg0/pool` for data, `lvextend --poolmetadatasize +1G vg0/pool` for metadata, plus removing an expendable snapshot to release blocks immediately. Rehearse it, because the situation where you need it is exactly the one where filesystems are already going read-only. ## The interview point Thin provisioning does not create capacity; it defers the moment you must have it, in exchange for a shared failure mode that is worse than the one it replaced. The engineer who has operated it talks about `Meta%`, discards, free extents in the volume group, and the alert threshold. The one who has only read about it talks about how much space it saves.

  • Why can a thin pool keep filling even though the filesystems on top of it look half empty?
    Because deleting a file only frees filesystem blocks, not pool blocks. Unless discards travel down — a discard-enabled mount or a scheduled `fstrim`, with the pool passing them through — the pool keeps every block it ever allocated. Retained snapshots compound it by pinning blocks the live filesystem has already released.
  • How is running out of pool metadata different from running out of pool data?
    Data exhaustion fails the writes that need new blocks, and adding data space resumes normal service. Metadata exhaustion puts the pool into read-only mode to protect the mapping trees, and recovery usually means deactivating it and running thin_check and thin_repair against the metadata rather than simply extending it. Size and monitor the metadata area separately.
  • What does `lvchange --errorwhenfull y` change about a full pool's behaviour?
    It makes I/O that needs an allocation fail immediately instead of being queued for a grace period first. That trades a confusing hang for a clear, fast error, which usually makes the incident easier to diagnose and lets applications and filesystems react promptly rather than piling up blocked writers.
  • You are paged with the pool at 99% and climbing. What are the immediate levers?
    Extend the pool from free volume group extents with `lvextend`, and if the VG has none, add a physical volume first. Delete an expendable snapshot to release the blocks it pins. Run `fstrim` on the thin volumes so already-deleted data returns blocks. Then stop or throttle whatever is writing hardest while capacity is arranged.

saying these in an interview costs you the question

  • Watches Data% but never Meta%
  • Assumes the pool grows itself with no free extents in the VG
  • Thinks deleting files automatically returns pool blocks
  • Believes only the volume being written is affected
  • Treats over-provisioning as free capacity rather than deferred capacity

context