skip to content

What is write amplification in a lakehouse table, and what raises it?

level: middleimportance: should knowfreq 50%

answer

  1. bytes written divided by bytes actually changed
  2. one row changed, one whole file rewritten
  3. the rewrite unit is the file
  4. maintenance passes write data that has not changed
  5. batch the edits so one rewrite covers many

basics

~20 s

Write amplification is the ratio of bytes physically written to bytes logically changed. Immutable files force whole-file rewrites, so changing one row in a 512 MB file writes 512 MB. Larger files, frequent small updates, and repeated compaction of the same data all raise it.

solid answer

~50 s

Because data files are immutable, a logical change of a few bytes becomes a physical rewrite of everything sharing the file. Update one row in a 512 MB file and you write 512 MB — an amplification factor of millions. Three things push it up: **file size**, since the rewrite unit is the file; **change frequency and scatter**, since updates spread across many files rewrite all of them; and **compaction itself**, since every pass rewrites data that was already written, and data compacted repeatedly is written several times over its life. The counter-measures are deferring the rewrite (record deletions or new versions separately and merge at read time, then batch the physical rewrite), batching changes so one rewrite absorbs many logical edits, and clustering so changes land in fewer files. It is a genuine tradeoff against read cost, not a bug to eliminate.

go deeper

for a junior

Recall that files cannot be edited in place, so changing one row means writing a whole new file — that ratio is write amplification.

for a middle

Explain the three drivers — file size, how scattered the changes are, and repeated maintenance passes — and describe batching and deferred rewrites as the standard mitigations with their read-side cost.

for a senior

Be ready to budget bytes written per day across ingest, mutation and maintenance for a real table, spot a compaction policy that re-touches already-good files, and argue a target size against the mutation rate.

for a principal

Own the economics: when a platform pays repeated rewrite cost across many tables, decide where deferred-rewrite semantics become the default, what maintenance is allowed to touch, and how that spend is measured and capped.

## Definition **Write amplification** = bytes physically written to storage ÷ bytes of data that logically changed. A single-row update of 200 bytes that causes a 512 MB file to be rewritten has an amplification factor around 2.7 million. That number sounds alarming and is entirely normal for immutable-file table formats. The question is never "is there amplification" but "is it bounded and is it worth what it buys." ## Why it exists at all Data files in these formats are immutable, and that immutability is load-bearing: it is what lets readers hold a consistent view without locking, lets snapshots and time travel work by simply retaining old files, and lets writers commit by swapping a pointer rather than mutating shared state. The price is that you cannot edit a file in place. Any change to any row in a file means producing a new file. Columnar encoding compounds this. A Parquet-style file is compressed, dictionary-encoded and organized by column; you cannot surgically patch a value inside it without decoding and re-encoding the containing structure. ## The three main drivers **1. File size.** The rewrite unit is the file, so amplification is roughly proportional to file size for scattered single-row changes. This is exactly why "just make files enormous" is the wrong answer to the small-files problem — you trade read overhead for write amplification. The target-size band exists because both ends cost something. **2. Change scatter.** Ten updates that all fall in one file cost one file rewrite. Ten updates scattered across ten files cost ten. Scatter is a function of data layout: a table clustered on the key you update by concentrates changes; an unclustered table spreads them everywhere. This is a second, less-obvious benefit of clustering — it reduces write amplification for mutations, not just scan cost for reads. **3. Repeated compaction.** Compaction is itself amplification. Every byte it processes is read and written again with no logical change whatsoever. Worse, a poorly designed schedule can rewrite the same bytes many times: a file compacted from 4 MB to 128 MB, then later merged into a 512 MB file, then re-clustered, has been written three times beyond its original write. Good compaction policy avoids re-touching already-target-sized files — only files meaningfully below target are candidates — and treats cold partitions as done. ## The main mitigation: defer the physical rewrite Rather than rewriting a file when a row changes, a table format can record the change *beside* the file and reconcile at read time: - Record which rows are logically deleted, so readers filter them out while the original file stays untouched. - Write new versions of changed rows into new small files, and have readers merge them over the base data by key. Both approaches make the write cheap — proportional to what actually changed — and push cost to the read path, which must now consult extra information for every scan. Amplification does not disappear; it is deferred until a later maintenance job materializes the changes into new base files. The tradeoff is explicit: **cheap writes and more expensive reads, versus expensive writes and clean reads.** Choosing between them is a workload decision — heavy mutation with tolerant read latency favors deferral; append-mostly data with strict query SLAs favors rewriting immediately. Note the second-order effect: deferring creates small files (the delete records or new-version files), so the mitigation for write amplification *feeds* the small-files problem. The two are coupled, which is why one maintenance policy has to address both. ## The other mitigation: batch Amplification is per rewrite, not per logical change. If you accumulate an hour of updates and apply them in one pass, every affected file is rewritten once rather than once per update. Micro-batching is the single most effective lever available to a pipeline author, and it also directly reduces the small-file count. Row-at-a-time modification of a columnar lakehouse table is the pathological case for both problems simultaneously. ## Budgeting it A useful way to reason about a table: estimate the bytes written per day from (a) ingestion, (b) mutation rewrites, and (c) maintenance passes. Object storage bills for writes and for storage of the versions you retain; compute bills for the rewriting. Teams are routinely surprised that maintenance dominates ingest on a heavily updated table — the table writes several times its own daily volume without gaining a byte of new information. When that ratio is uncomfortable, the levers in order of impact are: batch more coarsely, defer physical rewrites via delete/merge-at-read, cluster so changes concentrate, lower target file size for mutation-heavy tables, and stop compacting partitions that no longer receive writes.

  • How does deferring the rewrite make writes cheap, and what does it cost instead?
    Instead of rewriting the containing file, the writer records the change separately — which rows are gone, or new versions of them — so write cost is proportional to what changed. Readers must then reconcile base data with those records on every scan, so read cost rises and small files accumulate until a later job materializes them into new base files.
  • Why does clustering reduce write amplification, not just read cost?
    Amplification scales with how many distinct files a change set touches. If the table is ordered by the key you update by, a batch of updates concentrates into a few files instead of smearing across all of them, so far fewer files are rewritten for the same logical change.
  • When is high write amplification acceptable?
    When the table is append-mostly and queried constantly. Occasional rewrites for compaction buy clean, well-sized, well-clustered files that every subsequent query benefits from. Amplification only becomes a problem when mutation is frequent or maintenance rewrites the same bytes repeatedly.
  • How can a compaction policy itself become the dominant source of amplification?
    By re-touching data that is already fine. If every run considers all files rather than only those below target, or re-clusters cold partitions that no longer receive writes, the same bytes are read and rewritten pass after pass with no benefit. Restrict candidates to under-sized files and to partitions that actually changed.

Reprinting an entire book because one sentence changed — and doing it again next week for a different sentence.

saying these in an interview costs you the question

  • Believing a single-row update rewrites only that row's bytes
  • Claiming larger target files always reduce total write cost
  • Ignoring compaction and clustering passes when counting bytes written
  • Thinking delete-and-merge-at-read removes amplification rather than deferring it
  • Treating write amplification as a defect instead of a deliberate tradeoff

context