skip to content

With deletion vectors enabled on a Delta Lake table, what does a DELETE write?

level: seniorimportance: nice to knowfreq 36%

answer

  1. rewriting a whole file to drop three rows
  2. mark the rows instead of rebuilding the file
  3. a bitmap of positions inside one file
  4. the reader must subtract while scanning
  5. old readers are refused, not fooled

basics

~20 s

With deletion vectors enabled, a Delta DELETE leaves the Parquet file untouched and writes a bitmap of deleted row positions to a side file; the commit re-registers the same file path carrying a deletionVector reference instead of rewriting data.

solid answer

~40 s

Deletion vectors turn Delta's copy-on-write deletes into merge-on-read. Instead of rewriting a whole Parquet file to drop a few rows, the writer emits a compressed bitmap of the deleted **row positions** within that file and commits an `add` action for the *same* data-file `path` carrying a `deletionVector` object — `storageType`, `pathOrInlineDv`, `offset`, `sizeInBytes`, `cardinality` — plus a `remove` for the file's previous state. Readers must subtract those positions while scanning, which is why `deletionVectors` is a **reader-writer** table feature: an unaware reader is refused the table rather than allowed to return deleted rows. Deletes become fast and cheap, at the cost of read-side work and accumulating vectors. `OPTIMIZE` materialises them by rewriting the files, and `REORG TABLE ... APPLY (PURGE)` forces that rewrite when the rows must physically go.

code

json · 3 lines
json
{"commitInfo":{"timestamp":1700000200000,"operation":"DELETE","operationMetrics":{"numDeletionVectorsAdded":"1","numRemovedFiles":"0","numDeletedRows":"3"}}}
{"remove":{"path":"part-00003-a1b2c3d4.snappy.parquet","deletionTimestamp":1700000200000,"dataChange":true}}
{"add":{"path":"part-00003-a1b2c3d4.snappy.parquet","partitionValues":{},"size":268435456,"modificationTime":1699000000000,"dataChange":true,"deletionVector":{"storageType":"u","pathOrInlineDv":"r7Kq2mXd9","offset":1,"sizeInBytes":38,"cardinality":3}}}

go deeper

for a junior

Know the idea: instead of rewriting a Parquet file to drop a few rows, Delta can record which row positions are deleted and let readers skip them.

for a middle

Explain the commit shape — the same file path re-added with a deletionVector reference — and why the reader must apply the bitmap while scanning rather than trusting the file wholesale.

for a senior

Weigh the tradeoff in production: fast deletes and merges against read-side cost, accumulating vectors, and the need for regular compaction. Know that deleted bytes persist until a purge and cleanup actually run.

for a principal

Own the enablement decision. It is a permanent protocol upgrade on a shared table that can lock out older engines, so treat it as a breaking contract change weighed against the delete workload it relieves.

## The problem Without deletion vectors, a Delta `DELETE`, `UPDATE` or `MERGE` is copy-on-write: any data file containing an affected row is rewritten in full, minus or with the change. Deleting three rows from a 256 MB Parquet file rewrites 256 MB and copies a quarter of a million untouched rows — visible in the commit's `numCopiedRows` metric. For GDPR-style scattered deletes, or for a `MERGE` touching a few rows in many files, this write amplification dominates the job. ## The mechanism Enabling the feature (`delta.enableDeletionVectors = true`) switches those operations to merge-on-read. The writer: 1. determines which **row positions** inside each affected data file are deleted — ordinal positions in the file, not primary keys; 2. serialises those positions as a compressed bitmap (a roaring-bitmap encoding) into a side file in the table directory, or inline in the log when tiny; 3. commits a `remove` action for the file's previous state and an `add` action for the **same `path`**, now carrying a `deletionVector` object. The `deletionVector` object names `storageType` (whether the bitmap lives in a separate file, at an absolute path, or inline), `pathOrInlineDv`, `offset` and `sizeInBytes` locating the bitmap, and `cardinality` — how many rows it marks deleted. That cardinality is what lets the engine report accurate row counts without opening the bitmap. No Parquet data is written. A delete becomes proportional to the number of *files touched*, not to the bytes they contain. ## What readers must do A scan of a file that has a deletion vector must fetch the bitmap and skip the marked row positions. That is a genuine change to how an `add` action is interpreted: previously, a live `add` meant every row in that file is part of the table. This is why `deletionVectors` is a **reader-writer** table feature, appearing in both `readerFeatures` and `writerFeatures`, and why the table's protocol jumps to reader version 3 / writer version 7. An engine that does not implement deletion vectors must refuse the table outright — the alternative would be silently returning rows the user deleted, which is a correctness and compliance failure, not a performance one. Enabling the feature on a table read by an older Spark, an older connector or a third-party tool will break those consumers, so inventory them first. ## Costs and accumulation Merge-on-read moves work rather than removing it: - **Read cost.** Every scan of a covered file loads and applies its bitmap. Cheap per file, but non-zero, and it constrains vectorised reading. - **Skew between physical and logical size.** A file may be 256 MB on disk while only a fraction of its rows are live. Enough deletes and the table is mostly reading rows it then throws away. - **Small side files.** Each round of deletes adds bitmap files. The remedy is compaction. `OPTIMIZE` rewrites affected files and, in doing so, materialises the deletions — the rewritten files contain only live rows and carry no vector. Where the physical bytes must genuinely go (a right-to-erasure request, for example), `REORG TABLE ... APPLY (PURGE)` forces the rewrite of files carrying deletion vectors, and the retention-based cleanup then removes the tombstoned originals once the retention window has passed. Until both steps run, the deleted rows still exist in the old Parquet files even though no query returns them, which is exactly the distinction a compliance reviewer will probe. ## Contrast with the neighbouring formats Every table format has some answer to "delete without rewriting", and mixing them up is the classic interview error. Delta's answer is the deletion vector: a positional bitmap attached to one data file. Iceberg's answer is delete files — positional deletes and equality deletes — which are separate data files tracked in its own metadata. Hudi's answer is a Merge-on-Read table type where updates land in log files that are later compacted into base files. Same problem, three different mechanisms, and the vocabulary does not transfer. ## Where it also helps Because a deletion vector rewrites no data file, two writers modifying different rows in the same file are less likely to be found in conflict; recent Databricks runtimes build row-level concurrency on top of deletion vectors and row tracking specifically to reduce the concurrent-modification failures that plague `MERGE` on unpartitioned tables. That is a runtime-level capability layered on the feature, not a property of the bitmap itself. ## When to leave it off Append-mostly tables with rare deletes gain little and still pay the protocol upgrade, which may lock out consumers. Tables read by a heterogeneous fleet of engines are the risky case. Delete-heavy or `MERGE`-heavy tables read by an up-to-date engine are where the feature pays for itself — provided compaction runs regularly, because merge-on-read without compaction degrades exactly the way it does in every other format.

  • Why must deletionVectors be a reader-writer feature rather than writer-only?
    Because it changes how existing data is interpreted. Before, a live `add` meant every row in that file belongs to the table; now some positions are logically gone. A reader that ignored the bitmap would return deleted rows with no error, so the protocol forces such readers to refuse the table instead.
  • After deleting personal data on a table with deletion vectors, are the bytes gone?
    No. The rows are marked deleted and no query returns them, but they still sit in the Parquet file. Materialise the deletion with `OPTIMIZE` or `REORG TABLE ... APPLY (PURGE)`, then let the retention-based cleanup remove the tombstoned originals once the window has passed.
  • How is a Delta deletion vector different from Iceberg's delete files?
    They solve the same problem with different machinery. A Delta deletion vector is a positional bitmap attached to one data file and referenced from that file's `add` action. Iceberg tracks separate delete files in its manifests, and supports equality deletes keyed on column values as well as positional ones. The vocabulary does not transfer between the formats.
  • What happens to accumulated deletion vectors when OPTIMIZE runs?
    The rewritten files contain only live rows, so the vectors are materialised and dropped — the new `add` actions carry no `deletionVector`. That restores full-speed scans and removes the physical/logical size skew, which is why a merge-on-read delete strategy needs regular compaction to stay healthy.

saying these in an interview costs you the question

  • Saying deleted rows are removed from the Parquet file immediately
  • Assuming any Delta reader can read a table with deletion vectors
  • Calling a deletion vector a list of primary keys of deleted rows
  • Confusing deletion vectors with Iceberg delete files or Hudi log files
  • Expecting deletion vectors to keep tables healthy with no compaction

context