skip to content

What does the fill factor (page fill percentage) of a B+tree index mean, why do engines let you leave free space in a leaf page on purpose, and how does a page's fill decay under different insert and update patterns?

level: middleimportance: should knowfreq 40%

answer

  1. Fill factor = packing target at build time
  2. Free space = landing zone for in-range writes
  3. Append-only keys: reserve nothing
  4. Random churn: settles near ~2/3 density
  5. Every reserved point is a read cost forever

basics

~20 s

Fill factor is how full a leaf page is packed when the index is built. Leaving free space lets later entries land in place instead of forcing new pages. Random churn drives density down; append-only keys keep it high.

solid answer

~60 s

**Fill factor** is the target percentage of each leaf page that a build or rebuild fills, leaving the rest as reserved free space. Packing pages 100% full gives the smallest, densest index — best for read-only or append-only data. But if later inserts and updates need to place entries into an already-full page, the engine must allocate a new page and move entries, which costs write I/O and leaves both pages partly empty. So fill factor is a trade: density and scan efficiency now, versus absorbing future in-range writes without disturbing the structure. A mostly-static reporting table wants a high fill; a table with heavy in-place churn on indexed columns benefits from reserving space. Decay depends on workload shape. Keys inserted in ever-increasing order fill pages once, densely, and never revisit them. Random keys and updates that move keys spread writes across the whole index, and the long-run average density for a randomly churned B+tree settles somewhere around two-thirds. Deletes push it lower still, because emptied slots in a page are only reused by keys that fall in that page's range.

code

text · 6 lines
text
page size 8 KB, ~400 entries per page at 100% fill

build at 100% fill, append-only workload : 2,500 leaf pages, stays ~2,500
build at  70% fill, append-only workload : 3,572 leaf pages, stays ~3,572  <- pure loss
build at 100% fill, random-churn workload: 2,500 pages -> drifts to ~3,700
build at  70% fill, random-churn workload: 3,572 pages -> drifts to ~3,700 too

go deeper

for a junior

Know that fill factor is how tightly pages are packed at build time, and that leaving room helps later inserts land without creating new pages.

for a middle

Explain the trade in both directions — read density versus structural write work — and that the right value depends on whether writes land at the edge or in the middle of the key range.

for a senior

Reason about steady-state density per workload shape (append-only, random, update-heavy, range-delete) and insist on measurement before changing a default.

for a principal

Treat it as one lever among several: key design, partitioning, and index count usually move density far more than a fill-factor setting does.

## What fill factor is A B+tree index is made of fixed-size pages. When an index is built — or rebuilt — from existing data, the engine walks the keys in sorted order and packs them into leaf pages. **Fill factor** (some engines call it PCTFREE from the other direction) is the target percentage of each page it fills before starting the next one. A fill factor of 100 means pack the page completely; 70 means fill 70% and reserve 30% as free space for future entries whose keys belong on that page. The reserved space is not wasted in the sense of being unusable — it is a deliberate landing zone. ## Why anyone would deliberately leave a page part-empty A fully packed page has no room. The moment a new key must go into it, the engine has to create room by allocating another page and redistributing entries, which costs write I/O, dirties extra pages, generates extra write-ahead log, and leaves the resulting pages only partly filled anyway. Reserving space up front converts that structural operation into a cheap in-page insert. There is a second, quieter reason in MVCC engines: an update to a row may need a new index entry while the old one is still needed by open snapshots, so both live in the index simultaneously. Free space absorbs that transient doubling. ## The trade, stated plainly - **High fill (95–100%)** — smallest index, fewest pages per scan, best cache density, fastest bulk reads. Right for static or append-only data: history tables, warehouse-style tables loaded once, indexes on ever-increasing keys. - **Lower fill (70–85%)** — bigger index and slightly more pages per scan, but future writes in the middle of the key range settle in place instead of restructuring. Right for tables with sustained in-place churn on indexed columns. Over-reserving is a real cost, not a free insurance policy: every point of fill factor you give up is a point of scan efficiency and cache density you pay on every read, forever. ## How fill decays over time Fill factor describes the state at build time. What matters in production is the *steady-state density* that the workload drives the index toward. **Append-only / monotonically increasing keys.** New keys always exceed every existing key, so they always land at the right-hand edge. Pages behind the edge are written once and never revisited: density stays essentially at whatever the insert path produced, and it stays high. Reserving free space in these pages is pure waste — nothing will ever be inserted into them. **Random keys.** Inserts land uniformly across the whole index. Pages fill, run out of room, and are split into two partly-filled pages. Averaged over a long random-insert workload, the classic result is a steady-state density around 2/3 of capacity. That is the baseline an actively churned index tends to; you cannot beat it without maintenance. **Update-heavy on indexed columns.** An update that changes an indexed value is logically a delete plus an insert in a different part of the key range. The old entry stays dead until snapshots release it, and the new entry consumes space somewhere else. This is the fastest route to low density, because it simultaneously fills new pages and hollows old ones. **Delete-heavy, especially range deletes.** Deleting all of last quarter's rows hollows out a contiguous stretch of the index. Those pages hold few live entries, and unless the workload later inserts keys in exactly that range — which, for a timestamp or sequence key, it never will — they stay hollow forever. Completely emptied pages can be recycled onto the index's free list; partly-empty ones cannot. ## Why density matters for reads Everything an index costs at read time is denominated in pages: pages read from storage, pages resident in the buffer cache, pages prefetched during an ordered scan. Density is the exchange rate between live entries and pages. Halving density roughly doubles the pages a range scan touches and roughly halves how much of the index the same cache memory can cover. ## Practical stance Default fill is usually right. Deviate deliberately: raise it toward full for indexes on append-only or read-mostly data; lower it only for indexes you have measured splitting and re-splitting under mid-range churn. Then re-measure — the point of tuning it is to reduce structural write work without paying too much read cost, and only observation tells you which side you are on.

  • For an index on an auto-incrementing id column, what fill factor makes sense and why?
    Essentially full. New keys are always larger than every existing key, so they land at the right-hand edge and never need to be inserted into older pages. Any reserved space in those older pages will never be used, so it is permanent read overhead — more pages per scan and less of the index resident in cache — bought for no benefit.
  • Is a lower fill factor a way to prevent index bloat?
    No — it changes where the empty space comes from, not whether it exists. A lower fill factor pre-pays the space in exchange for fewer structural write operations later. Bloat from dead entries is caused by delete and update churn plus deferred reclamation, and reserved free space does nothing to reclaim those entries.

Shelving books with a gap left at the end of each shelf: the gaps let new titles slide in without reshuffling the whole library, but they also mean you walk past more shelves to read the same collection.

saying these in an interview costs you the question

  • "Lower fill factor always means better write performance" — it only helps when writes land in the middle of the key range; for append-only keys it is pure overhead
  • "Fill factor is enforced continuously" — it is a target for build/rebuild; ordinary inserts fill pages beyond it
  • "100% fill is always wrong" — it is the right answer for read-mostly and append-only indexes
  • "Reserving free space prevents bloat" — bloat comes from dead entries and deferred cleanup, which free space does not address
  • "Density doesn't matter, it's all in cache anyway" — density is exactly what determines how much of the index fits in that cache

context