skip to content

In LVM on Linux, what is a physical extent (PE), what extent size does vgcreate use by default, and what does that choice actually affect?

level: middleimportance: should knowfreq 52%

answer

  1. allocation unit, not an I/O unit
  2. fixed per group, chosen at creation
  3. 4 by default, in mebibytes
  4. sizes round up, never down
  5. granularity versus metadata volume

basics

~20 s

A physical extent is LVM's fixed-size unit of allocation — 4 MiB by default in LVM2, set per volume group at vgcreate time and shared by every PV in it. Logical volume sizes are always rounded up to a whole number of extents.

solid answer

~40 s

When `vgcreate` builds a volume group it divides every member PV into equal-sized chunks called physical extents; the default is 4 MiB and you override it with `vgcreate -s 16M`. All PVs in one VG use the same extent size — it is a property of the group, not of the disk. Everything above is extent bookkeeping: a logical volume is just a list of extents, which is why `lvcreate -L 10M` on a 4 MiB-extent VG prints "Rounding up size to full physical extent" and gives you 12 MiB. `-l` lets you ask in extents or percentages directly. Extent size is mostly a granularity-versus-metadata tradeoff: a small extent wastes less on rounding, a large one keeps the extent count and the per-operation bookkeeping down on very large groups.

code

bash · 4 lines
bash
vgcreate -s 16M vg0 /dev/sdb
vgdisplay vg0 | grep -E 'PE Size|Total PE'
lvcreate -L 10M -n test vg0
lvs -o lv_name,lv_size,seg_count vg0

go deeper

for a junior

Know that LVM allocates space in fixed-size extents, that the default is 4 MiB, and that a size you ask for gets rounded up to the next whole extent.

for a middle

Explain that the extent size is fixed for the whole volume group at vgcreate time via -s, that logical extents map one-to-one onto physical extents, and why -l 100%FREE avoids rounding surprises.

for a senior

Show the tradeoff judgment: when a very large group justifies a coarser extent, why granularity suffers, and how to read the segment map from lvdisplay -m to see which disk actually backs a volume.

for a principal

Decide the standard once for the fleet, since the value is unchangeable in practice after creation: pick an extent size that suits the largest groups you provision and make automation set it explicitly rather than inheriting a default.

## The unit LVM actually thinks in LVM does not track byte ranges. When a volume group is created, each of its physical volumes is carved into equal-sized **physical extents** (PEs), and every allocation decision after that is "which extents, from which PV". A logical volume's on-disk definition is essentially a mapping from its own logical extents (LEs) to physical extents on named PVs. Logical extents are the same size as physical extents in the group, so the two counts line up one-to-one for a simple linear volume. That design is why LVM operations are cheap. Growing a volume is appending extents to a list and reloading a device-mapper table; nothing is copied. ## The default, and how to change it LVM2 uses **4 MiB** extents by default. You set a different size at group-creation time: ```bash vgcreate -s 16M vg0 /dev/sdb /dev/sdc vgdisplay vg0 | grep 'PE Size' ``` The size must be a power of two and at least 1 KiB. It is a property of the **volume group**: every PV that joins the group is divided using the same size, so you cannot mix 4 MiB and 16 MiB extents inside one VG. ## What the choice actually affects Three things, in descending order of how often they matter: **Granularity.** Extent size is the quantum of every size you can ask for. With 4 MiB extents an LV's size is always a multiple of 4 MiB; with 128 MiB extents you cannot create a 100 MiB volume at all — you get 128 MiB. On a group hosting many small volumes, a coarse extent size silently inflates every one of them. **Metadata and operation cost.** Each extent is an entry LVM has to account for. A 100 TiB group with 4 MiB extents is roughly 26 million extents; the metadata gets large and allocation and reporting commands get slower. Bumping the extent size to 32 MiB or 64 MiB on very large groups is the standard reason to deviate from the default. **Nothing about performance.** This is the common misconception. Extent size is an allocation granularity, not an I/O size. It does not set a stripe width, a chunk size or a readahead value; a filesystem doing 4 KiB writes does 4 KiB writes regardless of whether the extents beneath it are 4 MiB or 64 MiB. (The one place a size *does* affect I/O layout is the stripe size chosen when creating a striped LV, which is a separate parameter entirely.) ## The rounding you will actually see ```bash # 4 MiB extents; 10 MiB is 2.5 extents, so LVM takes 3 lvcreate -L 10M -n test vg0 # Rounding up size to full physical extent 12.00 MiB ``` LVM rounds **up**, never down, and it tells you when it does. This trips people who script size arithmetic and then assert on an exact byte count afterwards. When you want to be exact, ask in extents rather than bytes. `-l` accepts a raw count or a percentage expression: ```bash lvcreate -l 2560 -n test vg0 # exactly 2560 extents lvcreate -l 100%FREE -n data vg0 # every unallocated extent in the VG lvcreate -l 50%VG -n half vg0 # half the group's total size ``` `100%FREE` is the idiomatic way to consume the remainder of a group without doing the division yourself. ## Reading the extent numbers `vgdisplay` is the clearest view: it prints `PE Size`, `Total PE`, `Alloc PE / Size` and `Free PE / Size`. The compact commands expose the same figures as columns — `vgs -o vg_name,vg_extent_size,vg_extent_count,vg_free_count`, and `pvs -o pv_name,pv_pe_count,pv_pe_alloc_count` per disk. `lvdisplay -m` goes one level further and prints the segment map: for each range of logical extents, which PV and which physical extent range backs it. That segment map is the practical payoff of understanding extents. It is how you answer questions like "is this volume actually spread across both disks, or did it all land on the first one" — which is not visible from `lvs` alone. ## A note on capacity you cannot use A PV whose size is not a whole multiple of the extent size loses the remainder: the leftover tail is smaller than one extent, so nothing can be allocated from it. With 4 MiB extents that is a rounding error. With a deliberately huge extent size on many small disks it can add up, which is one more reason not to raise the default without a concrete reason.

  • Why might you deliberately choose a 32 MiB extent size instead of the 4 MiB default?
    On a very large volume group — hundreds of terabytes — 4 MiB extents produce tens of millions of entries, which bloats the metadata and slows allocation and reporting. A coarser extent cuts the count by an order of magnitude. The cost is granularity: every logical volume's size becomes a multiple of 32 MiB, which is irrelevant for multi-terabyte volumes and wasteful for a group of small ones.
  • Can two physical volumes in the same volume group use different extent sizes?
    No. Extent size is a property of the volume group, fixed when `vgcreate` runs, and every PV that joins is divided using it. That uniformity is what lets LVM treat the group as one flat pool of interchangeable extents rather than tracking per-disk geometry.
  • Does raising the extent size improve I/O throughput?
    No — that is the standard misconception. An extent is an allocation quantum, not a transfer size. The block layer and the filesystem still issue whatever request sizes they were going to issue. The parameter that genuinely shapes I/O layout is the stripe size on a striped logical volume, which is chosen separately at lvcreate time.

saying these in an interview costs you the question

  • Thinks bigger extents mean faster I/O
  • Believes extent size can differ per disk
  • Assumes lvcreate -L gives an exact byte size
  • Confuses the extent size with a stripe or chunk size
  • Thinks extent size is set on the PV by pvcreate

context