skip to content

In a Delta Lake table, what actions does the _delta_log commit for a row-level DELETE contain?

level: middleimportance: must knowfreq 66%

answer

  1. Parquet files cannot be edited in place
  2. the file set changes, not the file
  3. one line says out, another says in
  4. look at numCopiedRows in the metrics
  5. the old file lingers for time travel

basics

~20 s

A row-level DELETE on a Delta table without deletion vectors rewrites the affected Parquet file: the commit holds a remove tombstone for the old file path, an add for the freshly written file minus those rows, and a commitInfo entry describing the operation.

solid answer

~50 s

Parquet files are immutable, so on a table without deletion vectors Delta deletes rows by **rewriting whole files**. It first finds which files can contain matching rows — using the partition values and the `stats` recorded on each `add` action — then rewrites only those, dropping the matched rows. The commit contains a `remove` action naming the old file `path` with a `deletionTimestamp` and `dataChange: true`, an `add` action for the new file with its own size and statistics, and a `commitInfo` line whose `operationMetrics` report `numRemovedFiles`, `numAddedFiles` and `numDeletedRows`. The old Parquet file stays on storage — the tombstone is logical, so earlier versions remain readable until retention-based cleanup removes the bytes. Deleting one row from a 1 GB file costs a 1 GB rewrite, which is exactly the write amplification deletion vectors were introduced to avoid.

code

json · 3 lines
json
{"commitInfo":{"timestamp":1700000100000,"operation":"DELETE","operationParameters":{"predicate":"[\"(user_id = 42)\"]"},"readVersion":11,"isBlindAppend":false,"operationMetrics":{"numRemovedFiles":"1","numAddedFiles":"1","numDeletedRows":"3","numCopiedRows":"249997"}}}
{"remove":{"path":"part-00003-a1b2c3d4.snappy.parquet","deletionTimestamp":1700000100000,"dataChange":true,"extendedFileMetadata":true,"partitionValues":{},"size":268435456}}
{"add":{"path":"part-00007-e5f6a7b8.snappy.parquet","partitionValues":{},"size":268402000,"modificationTime":1700000100000,"dataChange":true,"stats":"{\"numRecords\":249997}"}}

go deeper

for a junior

Recall that Parquet files are immutable, so a Delta delete writes a new file and marks the old one removed in the next commit. Know that the old file is not erased right away.

for a middle

Walk the commit line by line: remove with a tombstone timestamp, add with new size and stats, commitInfo with the predicate and metrics. Explain how partition values and stats limit which files get rewritten.

for a senior

Diagnose the cost. Be able to read numCopiedRows off the history, connect a slow nightly delete job to file layout, and argue for clustering, partition alignment or deletion vectors with numbers behind it.

for a principal

Frame the retention policy this creates: every mutation leaves tombstoned files, so storage growth, time-travel depth and compliance deletion deadlines are one linked decision you set for the platform, not per table by accident.

## Why a delete is an add plus a remove Delta Lake stores rows in Parquet files, and Parquet files are immutable — no engine edits bytes inside one. So every row-level change is expressed as a change to the *set* of files that make up the table, recorded in the next numbered JSON commit under `_delta_log`. For a `DELETE`, that means: write a new file containing the surviving rows, then in one commit say that the old file is out and the new file is in. This is called copy-on-write, and it is the default behaviour for a Delta table that does not have the `deletionVectors` table feature enabled. ## Step one: find the candidate files Delta does not rewrite the whole table. The planner narrows the work twice. First, partition pruning: each `add` action carries `partitionValues`, so a predicate on a partition column eliminates whole partitions without touching them. Second, file skipping on statistics: each `add` carries a `stats` string with `numRecords` and per-column `minValues`, `maxValues` and `nullCount`, so a file whose min/max range cannot contain the predicate value is skipped. Only the surviving files are read and rewritten. A `DELETE` whose predicate matches nothing in a file leaves that file untouched, and its `add` action is not restated. ## Step two: the commit The commit file contains, at minimum: - one **`remove`** per rewritten file — `path` (relative to the table root), `deletionTimestamp`, `dataChange: true`, and usually `size`, `partitionValues` and `extendedFileMetadata: true` so a checkpoint can carry the tombstone without re-reading the file; - one or more **`add`** actions for the newly written files — new `path`, `size`, `modificationTime`, `dataChange: true`, `partitionValues` and fresh `stats`; - a **`commitInfo`** line: `operation: "DELETE"`, the predicate in `operationParameters`, and `operationMetrics` such as `numRemovedFiles`, `numAddedFiles`, `numDeletedRows` and `numCopiedRows`. `readVersion` records the snapshot the delete was planned against, and `isBlindAppend` is `false` because the operation read data. A reader replaying the log reconciles by path: the old path was added at some earlier version and removed here, so it drops out; the new path is live from this version forward. ## numCopiedRows is the cost signal The metric worth staring at is `numCopiedRows` — rows that were read out of the old file and written into the new one unchanged, purely because they shared a file with a deleted row. Deleting a single row from a 1 GB file rewrites 1 GB and copies every surviving row. That write amplification is the defining cost of copy-on-write deletes, and it is why a delete-heavy table wants either partition or clustering layout aligned with the delete predicate, or deletion vectors. ## The old file does not disappear A `remove` action is a **tombstone**, not a file deletion. The Parquet bytes stay on storage, which is what keeps version N-1 readable for time travel and rollback. A separate retention-based cleanup deletes the physical files after the tombstone retention window has passed. Two consequences follow: the storage a Delta table occupies is always larger than the live data it exposes, and any tooling that estimates table size by listing the directory over-reports it. ## UPDATE and MERGE are the same shape An `UPDATE` is a delete plus an insert in one commit: the matched files are removed and rewritten with the updated values, so the commit again holds `remove` plus `add` actions and `commitInfo.operation` is `UPDATE`. A `MERGE` produces the same pair of action types with `operation: "MERGE"` and richer metrics — `numTargetRowsInserted`, `numTargetRowsUpdated`, `numTargetRowsDeleted`. All row-level mutation in Delta reduces to the same two actions; only the metrics differ. ## dataChange, and why compaction is different Both `add` and `remove` carry a `dataChange` boolean. A `DELETE` sets it to `true`: the logical contents of the table changed, and a streaming reader tailing the table must react. A compaction that only repacks rows into larger files sets `dataChange: false` on its adds and removes — the file set changed but no row was added or lost, so a streaming consumer can ignore that version instead of re-emitting every repacked row as new data. Being able to explain that flag is a reliable signal that you have actually read the log rather than a blog summary of it. ## With deletion vectors enabled If the table carries the `deletionVectors` table feature, the same `DELETE` takes a different shape: the data file is not rewritten, and the commit instead re-registers the same `path` with a `deletionVector` pointer to a bitmap of deleted row positions. That is a distinct mechanism with its own reader requirements, and it changes the metrics you will see in `commitInfo`.

  • How does Delta decide which files a DELETE has to rewrite?
    By pruning against metadata already in the log. `partitionValues` on each `add` eliminates whole partitions, and the `stats` string — `numRecords` plus per-column min, max and null counts — eliminates files whose value range cannot contain the predicate. Only surviving candidates are read and rewritten; everything else is left untouched.
  • Why does deleting 10 rows sometimes rewrite gigabytes?
    Because rewriting is per file, not per row. If the 10 rows are scattered across 50 large files, all 50 are rewritten and every surviving row in them is copied — visible as a huge `numCopiedRows`. Aligning partitioning or clustering with the delete predicate, or enabling deletion vectors, is the fix.
  • What does dataChange: false on an add action mean?
    That the commit moved rows between files without changing the table's logical contents — a compaction or file repack. Streaming readers tailing the table skip those versions instead of re-emitting the repacked rows as new data. A `DELETE`, `UPDATE` or `MERGE` always sets `dataChange: true`.

saying these in an interview costs you the question

  • Claiming Delta edits rows inside the existing Parquet file
  • Saying the DELETE physically frees storage immediately
  • Assuming a small DELETE only rewrites the matched rows
  • Thinking every DELETE rewrites the entire table
  • Confusing commitInfo metrics with the authoritative table state

context