skip to content

What are write, read and space amplification in a log-structured engine, and why can tuning never minimise all three at once?

level: seniorimportance: should knowfreq 38%

answer

  1. bytes written per byte
  2. files touched per read
  3. disk used per live byte
  4. merging more helps two, costs one

basics

~20 s

Write amplification is disk bytes written per byte the application writes; read amplification is work per read; space amplification is disk used per byte of live data. Merging files more eagerly cuts read and space costs but raises writes.

solid answer

~50 s

- **Write amplification**: total bytes written to storage divided by bytes the application wrote. The log, the flush and every later compaction that rewrites the same data all count. - **Read amplification**: how much work a read does — files or blocks consulted — relative to the one answer it needs. - **Space amplification**: bytes on disk divided by bytes of live data, inflated by obsolete versions, delete markers and temporary compaction space. The knob is **how eagerly files are merged**. Merging often keeps few overlapping files, so reads and space are lean, but each byte is rewritten many times: high write amplification. Merging rarely saves writes but leaves many overlapping files and stale data: high read and space amplification. Compaction policy chooses a point on this trade-off; workload shape (append-only, update-heavy, time-ordered) decides which point is cheapest. Measure all three before tuning.

go deeper

for a junior

Know that an LSM engine rewrites data during compaction and that this costs extra disk writes.

for a middle

Define write, read and space amplification and name the main source of each.

for a senior

Explain the eager-versus-lazy merge trade-off, relate it to workload shapes, and measure all three before tuning.

for a principal

Be ready to set storage and hardware budgets from amplification estimates and defend which cost a system should pay.

## Three ratios that describe an LSM engine A log-structured engine turns random writes into sequential ones, then pays for it later by merging. Three **amplification** ratios capture where the cost goes. | ratio | definition | main sources | |---|---|---| | **write amplification** | bytes written to storage ÷ bytes written by the application | commit log, flush, every compaction that rewrites the data | | **read amplification** | work per read ÷ the minimum needed (files, blocks, seeks) | overlapping files, versions and delete markers to merge | | **space amplification** | bytes on disk ÷ bytes of live data | superseded versions, delete markers, temporary space during compaction | ## Why they pull against each other The central decision is **how often and how aggressively to merge files**. - **Merge eagerly** (keep data in few, non-overlapping files): - reads touch few files — **low read amplification**; - obsolete data is dropped quickly — **low space amplification**; - but each byte is rewritten each time it moves into a larger file — **high write amplification**. - **Merge lazily** (let files accumulate, merge in big batches): - each byte is rewritten fewer times — **low write amplification**; - but reads must check many overlapping files — **high read amplification**; - and stale versions linger — **high space amplification**, with large temporary space needed when a big merge runs. Research on storage engines has formalised this: improving any two of read, update and memory or space overhead tends to worsen the third. No setting minimises all three at once. ## How workload shape moves the balance 1. **Append-only, never updated** (logs, events): little stale data exists, so lazy merging costs little space; write amplification dominates. 2. **Update- and delete-heavy**: stale versions accumulate, so lazy merging hurts reads and space badly; eager merging pays off. 3. **Time-ordered with expiry**: grouping files by time lets whole files expire without rewriting, cutting all three. 4. **Read-latency critical**: keep read amplification low even at write cost. The concrete compaction strategies that sit at different points of this trade-off, and how deletes and expiry interact with them, are a subject of their own. ## Measuring before tuning - **Write amplification**: compare disk bytes written (from the operating system or engine statistics) with bytes ingested. - **Read amplification**: engine metrics for files or blocks read per query, and read latency percentiles. - **Space amplification**: disk used versus an estimate of live data, plus headroom needed for the largest compaction. Watch them together: a change that improves read latency may quietly double disk writes, which shortens SSD life and eats I/O budget. ## Interview angle Define all three precisely, explain the eager-versus-lazy merge trade-off, tie it to workload shapes, and say how you would measure before changing anything.

  • Why does write amplification matter more on SSDs?
    Flash cells wear out after a limited number of writes, so a higher device write volume shortens drive life. It also consumes device bandwidth that foreground writes and reads need.
  • Why can lazy merging need a lot of free disk space at once?
    When many files are finally merged into one large output, the inputs stay on disk until the output is complete. That temporary peak can approach the size of the data being merged.

saying these in an interview costs you the question

  • Believing a compaction setting exists that minimises write, read and space costs together
  • Counting only the application's writes when estimating disk write volume
  • Ignoring temporary space needed during large compactions
  • Tuning for read latency without measuring the change in disk writes