What does a Delta Lake table's _delta_log directory contain?
answer
- the Parquet files alone are not the table
- one new file per successful write
- numbered JSON, one action per line
- add and remove, plus schema and protocol
- checkpoints so replay stays cheap
basics
~10 sA Delta table's _delta_log holds the transaction log: numbered JSON commit files listing actions such as add, remove, metaData and protocol, plus periodic Parquet checkpoints and a _last_checkpoint pointer to the newest one.
solid answer
~50 sA Delta Lake table is a directory of Parquet data files with a `_delta_log` subdirectory beside them, and the log is what makes it a table. Each successful write creates one commit file named for its version, zero-padded to twenty digits: `00000000000000000000.json`, `00000000000000000001.json`, and so on. A commit file is newline-delimited JSON, one *action* per line — `add` (a data file joins the table, with its size, partition values and column statistics), `remove` (a tombstone saying a file left the table), `metaData` (schema and partition columns), `protocol` (required reader/writer versions), `commitInfo` (provenance for `DESCRIBE HISTORY`). Every ten commits by default Delta also writes a `.checkpoint.parquet` summarising the whole state, and `_last_checkpoint` names the newest one. The current table is the replay of those actions, not whatever Parquet files happen to be lying in the directory.
code
text · 10 linesevents/
├── part-00000-3c1f9a2b.snappy.parquet
├── part-00001-77e0d4c5.snappy.parquet
└── _delta_log/
├── 00000000000000000000.json
├── 00000000000000000001.json
├── 00000000000000000002.json
├── 00000000000000000010.checkpoint.parquet
├── 00000000000000000010.json
└── _last_checkpointgo deeper
Be ready to say that a Delta table is Parquet files plus a _delta_log directory, that each write adds one numbered JSON commit file, and that add and remove name data files entering and leaving the table.
Explain the action types and how a snapshot is rebuilt: load the latest checkpoint, replay newer JSON commits, reconcile adds against removes by path. Know what stats on an add is used for.
Show you use the log operationally — reading commitInfo to explain a bad job, spotting a table whose log has grown to tens of thousands of tiny commits, knowing why directory size and table size diverge.
Own the consequences for the platform: which engines are allowed to write, why hand-placed files are never data, and how log retention interacts with the audit and time-travel guarantees you promise consumers.
## The directory, and why the log is the table A Delta Lake table on disk or object storage looks like this: some Parquet files, possibly nested in partition directories, and one subdirectory called `_delta_log`. The Parquet files hold rows. The log holds *the table*. That split is the whole idea — nothing in the data directory itself tells you which Parquet files are currently part of the table, so a reader that simply lists the directory and reads every `.parquet` file it finds will read files that have been logically deleted, files left behind by a job that crashed halfway, and old files that a compaction already replaced. Correct readers never do that: they read the log and open only the files the log says are live. ## Commit files Every successful write produces exactly one new file whose name is its version number, zero-padded to twenty digits with a `.json` suffix: `00000000000000000000.json` for the table's creation, then `...0001.json`, `...0002.json`. Versions are consecutive integers with no gaps. The file is newline-delimited JSON — one JSON object per line, each object a single **action**. Commit N is atomic because the writer is allowed to create `N.json` only if that file does not already exist. Either the file appears whole or the commit did not happen; there is no partially applied change. Readers therefore never see half a write. ## The action types - **`add`** — a data file becomes part of the table. Carries `path` (relative to the table root), `partitionValues`, `size`, `modificationTime`, `dataChange`, and `stats`: a JSON string with `numRecords` and per-column `minValues`, `maxValues` and `nullCount`. Those statistics are what lets the engine skip files without opening them. - **`remove`** — a tombstone. The named file is no longer part of the table as of this version, with a `deletionTimestamp`. The Parquet file itself still exists on storage; only a later cleanup deletes the bytes, which is what keeps time travel to earlier versions possible. - **`metaData`** — the schema as a JSON string, the partition columns, the format, and table properties. Rewritten whenever schema or properties change. - **`protocol`** — `minReaderVersion` and `minWriterVersion`, plus named `readerFeatures`/`writerFeatures` on tables using table features. A client below the required version must refuse the table rather than misread it. - **`txn`** — an application id and a monotonically increasing version, used to make streaming writes idempotent: if the same appId/version is already in the log, the batch is not applied twice. - **`commitInfo`** — provenance only: timestamp, operation name (`WRITE`, `DELETE`, `MERGE`, `OPTIMIZE`), operation parameters and metrics, `readVersion`, `isBlindAppend`. This is the row you see in `DESCRIBE HISTORY`. It is informational; state reconstruction does not depend on it. A useful subtlety is the `dataChange` flag on `add`/`remove`. A compaction that only reshuffles rows between files sets `dataChange: false`, so a streaming reader knows there is no new data to emit even though the file set changed. ## Checkpoints and _last_checkpoint Replaying thousands of JSON files on every query would be slow, so every tenth commit by default (`delta.checkpointInterval`) Delta writes a checkpoint: a Parquet file named for its version, `00000000000000000010.checkpoint.parquet`, containing the complete set of live actions at that version — all current `add`s, tombstones still within retention, and the latest `metaData`, `protocol` and `txn` entries. Alongside it sits `_last_checkpoint`, a tiny JSON file naming the newest checkpoint version and its size, so a reader can jump straight there instead of listing the directory. Very large tables may split a checkpoint across multiple part files, and newer Delta versions offer a v2 checkpoint layout with sidecar files. ## Reconstructing the current table A reader opens `_last_checkpoint`, loads that checkpoint, then replays only the JSON commits with higher version numbers. Adds and removes are reconciled by file path — a path that was added and later removed is gone; `metaData` and `protocol` are last-write-wins. What comes out is a *snapshot*: an exact list of data files, a schema, and a version number. Time travel is just replaying up to an earlier version instead of the latest. ## What the log is not It is not the data, and it is not a row-level redo log — Delta records file-level changes, never individual row images (Change Data Feed, when enabled, writes separate change files). It is also not infinite: `delta.logRetentionDuration`, 30 days by default, bounds how much history is kept, and expired commit files are cleaned up once a checkpoint covers them, which is why very old versions eventually stop being readable.
- If the log says a file was removed, why is the Parquet file still sitting in the directory?A `remove` action is a tombstone: it takes the file out of the current version only. The bytes stay so that time travel to earlier versions still works, and a separate retention-based cleanup deletes them later. That is also why listing the directory over-counts a Delta table's real size.
- What happens if someone drops an extra Parquet file into a Delta table directory by hand?Nothing — it is invisible. No `add` action references it, so no reader includes it, and it will eventually be treated as an orphan and cleaned up by the retention job. The only supported way to add data is a commit that writes an `add` action.
- Which action tells you who ran the last MERGE and how many rows it touched?`commitInfo`, the provenance line in each commit file, carries the operation name, its parameters, metrics such as rows updated and files added, and the engine identity. `DESCRIBE HISTORY` is just a rendering of those entries; it is informational and not used to rebuild table state.
The Parquet files are the warehouse shelves; _delta_log is the ledger. Only the ledger says which crates are currently part of the inventory — walking the aisles and counting what you see gives the wrong answer.
saying these in an interview costs you the question
- Saying the current table is whatever Parquet files exist in the directory
- Calling the log a row-level redo log of individual changed rows
- Believing a remove action physically deletes the Parquet file immediately
- Assuming _delta_log stores the data itself rather than metadata about files
- Thinking commit files can be edited in place after they are written