On a Linux server, /dev/sdb is one of three physical volumes in an LVM volume group serving live databases, and it has started logging SMART errors and I/O timeouts. How do you get the data off that disk and remove it from the volume group without downtime?
answer
- can the remaining disks even hold it
- one command empties a physical volume
- online, slow, and restartable
- membership first, then the label
basics
~20 sConfirm the remaining physical volumes hold enough free extents, then run pvmove /dev/sdb to relocate its extents onto them while the volumes stay online. Finish with vgreduce to drop the disk from the group and pvremove to clear its LVM label.
solid answer
~60 sFirst find out what is actually on it and whether there is anywhere to put it: `pvs` shows how many extents on `/dev/sdb` are allocated and how much free space the other physical volumes have, and `lvs -o +devices` shows which logical volumes have extents there. If the remaining members have room, `pvmove /dev/sdb` relocates every allocated extent off it, online, with the volumes mounted and in use — LVM builds a temporary mirror per extent, syncs it, then switches the mapping. It is slow and it competes with production I/O, so run it in a window and use `-i` to watch progress, or name a destination like `pvmove /dev/sdb /dev/sdd` to control where things land. It is restartable: re-running bare `pvmove` resumes an interrupted move, and `pvmove --abort` rolls it back. When the disk shows zero allocated extents, `vgreduce vg0 /dev/sdb` removes it from the group and `pvremove /dev/sdb` wipes its LVM label. The catch is that pvmove must read every extent from a disk that is already failing — if reads are erroring, restore from backup instead.
code
bash · 14 lines# 1. can the healthy PVs hold what is on the failing one, and who is exposed?
pvs -o +pv_used
lvs -o +devices vg0
# 2. evacuate online, with progress, optionally onto a named destination
pvmove -i 10 /dev/sdb /dev/sdd
# resume after an interruption, or roll the move back
pvmove
pvmove --abort
# 3. once /dev/sdb shows zero used extents
vgreduce vg0 /dev/sdb
pvremove /dev/sdbgo deeper
Know that LVM can relocate a logical volume's extents off one disk onto another with pvmove, and that a disk leaves the group with vgreduce. You are not expected to plan the operation yet.
Explain that pvmove works online by mirroring each segment and swapping the mapping, and give the finishing sequence — pvmove, then vgreduce for group membership, then pvremove for the label.
Show the assessment first: per-device free extents, which logical volumes are exposed via lvs -o +devices, and the I/O cost of the copy against live traffic. Name the pivot to a restore when the disk cannot be read reliably.
Decide the policy: at what SMART threshold a disk is evacuated, whether redundancy lives below LVM so this is a rebuild rather than a migration, and how much spare capacity every group carries so an evacuation always has somewhere to go.
## Assess before you move Two facts decide whether this is a routine evacuation or an incident: ``` pvs -o +pv_used # allocated versus free per physical volume vgs # VFree for the group as a whole lvs -o +devices vg0 # which logical volumes have extents on /dev/sdb ``` You need free extents on the *remaining* physical volumes at least equal to the allocated extents on `/dev/sdb`. Group-wide `VFree` is not sufficient on its own — free space that lives on the dying disk cannot receive its own data. If there is not enough room, add a replacement disk to the group first (`pvcreate` then `vgextend`) and evacuate onto it. That is usually the cleaner move anyway, since it keeps the group's total capacity constant. `lvs -o +devices` also tells you the blast radius: which services are exposed if the disk dies mid-evacuation. ## What pvmove actually does `pvmove /dev/sdb` relocates every allocated extent on that physical volume to free extents elsewhere in the group. Under the hood LVM converts each affected segment into a temporary mirror, synchronises the copy onto the destination extents, then atomically switches the logical volume's mapping to the new location and drops the old leg. Because the switch happens in the device-mapper table, the logical volume never goes away and the filesystem on it never notices — this is a genuinely online operation on mounted, busy volumes. Useful forms: ``` pvmove /dev/sdb # everything, LVM picks destinations pvmove /dev/sdb /dev/sdd # everything, onto a named destination pvmove -n data /dev/sdb # only the extents belonging to LV 'data' pvmove -i 10 /dev/sdb # report progress every 10 seconds ``` Progress is recorded in LVM metadata, not just in the running process, which is what makes it restartable. If the terminal dies, the machine reboots, or you stop it deliberately, running `pvmove` with no arguments resumes the outstanding move; `pvmove --abort` cancels it and leaves the extents where they started. Nothing is left half-mapped either way. ## The cost, which is real Every allocated extent is read from the failing disk and written elsewhere. On a multi-terabyte volume that is hours, and it is sequential-ish bulk I/O contending with production traffic for the same devices and the same queues. Watch service latency while it runs, and schedule it like the bulk copy it is. `-n` lets you evacuate the most critical logical volume first, so the exposure window for the data you care about most is short even if the whole disk takes all night. ## The failure case: reads that error pvmove's fundamental requirement is that it can *read* every extent on the disk it is emptying. A disk with reallocated sectors and rising timeouts may still satisfy that; a disk returning hard read errors will not, and the move will fail partway. That is the moment to stop optimising for elegance: - If the volumes are mirrored or RAID-backed below LVM, the redundancy — not pvmove — supplies the good copy, and the repair happens at that layer. - Otherwise, restore the affected logical volumes from backup onto healthy extents. And if the disk is already gone, evacuation is no longer on the table. LVM then shows the physical volume as missing, and `vgreduce --removemissing` (with `--force` when volumes still reference it) drops it from the group — at the cost of the logical volumes whose extents lived there, which then have to be recreated and restored. Knowing that this command sacrifices data is the difference between a controlled recovery and a second outage. ## Finishing the job Once `pvs` shows zero used extents on the disk: ``` vgreduce vg0 /dev/sdb # remove it from the volume group pvremove /dev/sdb # wipe the LVM label so it is no longer a PV ``` The order matters and is easy to reason about: `vgreduce` is group membership, `pvremove` is the label on the device. Trying `pvremove` on a device that is still a group member is refused. After that the disk can be pulled, and it is worth confirming with `pvs`/`vgs` that the group is healthy and that no logical volume lost extents. ## Say the whole arc in an interview The answer that lands is the ordered one: assess capacity and exposure, ensure a destination exists, evacuate online with pvmove while acknowledging its runtime cost, verify the disk is empty, remove it from the group, clear its label — and name the condition under which the plan changes to a restore, because a failing disk is not guaranteed to be a readable one.
- The server reboots halfway through a multi-hour pvmove. What state is the migration in?A consistent one. The in-flight move is recorded in LVM metadata, not only in the running process, so no logical volume is left half-mapped. Running `pvmove` with no arguments resumes the outstanding move from where it stopped, and LVM's polling daemon may pick it up automatically on activation. `pvmove --abort` cancels it instead and leaves the extents in place.
- The volume group as a whole shows plenty of free space. Is that enough to plan the evacuation?No — you need free extents on the *other* physical volumes. Group-wide VFree may include free space sitting on the very disk you are emptying, which cannot receive its own data. Check per-device free space with `pvs`, and if the healthy members are short, add a replacement disk with `vgextend` before starting.
- What if the disk fails completely before you finish?Evacuation is no longer possible; there is nothing left to read. LVM marks the physical volume missing, and `vgreduce --removemissing --force` removes it from the group while sacrificing the logical volumes that had extents on it. Those must be recreated and restored from backup, which is exactly why the evacuation should start on the first SMART warning rather than the last.
- Why does pvmove not require unmounting the affected filesystems?Because the relocation happens below the filesystem, in device-mapper. LVM mirrors each segment to its new extents, waits for the sync, then swaps the logical volume's mapping table entry. The logical volume's identity, size and device node never change, so the mounted filesystem sees a continuously available block device throughout.
saying these in an interview costs you the question
- Thinks the filesystems must be unmounted for pvmove
- Counts group-wide free space instead of per-device free space
- Assumes an interrupted pvmove leaves volumes corrupted
- Runs pvremove before removing the disk from the group
- Uses vgreduce --removemissing without realising it discards data