skip to content

Why does a LanceDB table keep growing on disk after deletes, and what fixes it?

level: seniorimportance: should knowfreq 40%

answer

  1. nothing is edited in place
  2. row count falls, bytes do not
  3. old versions pin old files
  4. compaction and pruning are two jobs
  5. retention buys space, costs history

basics

~20 s

Deletes commit a new version recording removed rows rather than rewriting data, and every earlier version still references the original files, so nothing is reclaimed. table.optimize() compacts fragments and prunes versions older than an age you pass, which is what actually returns the space.

solid answer

~50 s

Lance is copy-on-write: `delete()` writes deletion markers and commits a new version, and `update()` and `merge_insert()` rewrite affected fragments while the originals stay referenced by the versions before them. So a table under churn accumulates three things — superseded data files, deletion files, and version manifests — and its footprint grows even as `count_rows()` falls. Frequent small appends make it worse by leaving many small fragments, which also hurts scan performance. The fix is maintenance, not a different write pattern: `table.optimize()` compacts small and partially-deleted fragments into larger clean ones and prunes versions older than the threshold you give it. Pruning is what actually frees bytes, and it is destructive to history — after it, `checkout()` on a pruned version fails — so the retention window is a deliberate choice balancing rollback and reproducibility against storage cost.

go deeper

for a junior

Know that deleting rows in LanceDB does not immediately shrink the files, because the data is versioned and the old state is still readable until it is cleaned up.

for a middle

Explain copy-on-write: deletes write markers, updates rewrite fragments, and every earlier version keeps the old files alive; then name optimize as the maintenance that compacts and prunes.

for a senior

Diagnose it — row count versus directory size, deletion files versus small-fragment sprawl versus retained versions — and prescribe a compaction schedule plus a retention window, including its effect on query latency on object storage.

for a principal

Own the retention policy and its price: how far back rollback, reproducibility and audit must reach, what that costs in storage and requests at your scale, and where maintenance runs so it never races the ingest writer.

## The mechanism behind the symptom Nothing in a Lance dataset is edited in place. The dataset is a set of immutable data fragments plus metadata, and each commit produces a new manifest naming which fragments and deletion files constitute the table at that moment. Three write patterns therefore leave residue: - **Delete.** The rows are not removed from the fragment. A deletion file records which row positions in which fragment are gone, and a new version is committed. The original bytes are still on disk, still referenced by every earlier version. - **Update and merge_insert.** Modified rows cannot be patched in place, so the fragments containing them are rewritten with the new values and committed. The old fragments remain, referenced by prior versions. - **Small appends.** Each append writes its own fragment. Thousands of small appends leave thousands of small files, each with per-file overhead and per-file read cost. The user-visible symptom is a table whose row count drops while its directory grows, and whose queries get slower rather than faster after a cleanup. ## Why queries degrade too This is not only a bill problem. A fragment that is 90% deleted still has to be opened and its deletion file consulted; scans read structure they will immediately discard. Many small fragments multiply per-file metadata and, on object storage, multiply requests — the latency there is per request, so file count matters far more than it does on local disk. Recall and correctness are unaffected; throughput and cost are not. ## The maintenance operation `table.optimize()` is the answer, and it does two distinct jobs: 1. **Compaction.** Rewrite small and heavily-deleted fragments into fewer, larger, clean ones. Deleted rows are genuinely dropped in the rewritten data, and the file count collapses. Vector and scalar indexes need to keep up with the reorganised data, which optimize handles as part of the same maintenance pass. 2. **Version pruning.** Drop versions older than an age threshold you pass, along with the files only those versions referenced. This is the step that returns bytes to the filesystem or bucket, because until the last version referencing a fragment is gone, the fragment must stay. Compaction alone does not shrink the directory — it adds new files while the old ones remain pinned by history. People miss this and conclude compaction did nothing. ## Retention is a policy decision Pruning trades history for space. After a version is pruned, `checkout()` on it fails and any evaluation pinned to it can no longer be reproduced. So the threshold has to be chosen against real requirements: how far back must a rollback reach after a bad backfill; how long must an offline evaluation stay reproducible; is there an audit expectation. A pipeline that re-embeds nightly and evaluates weekly needs a window comfortably longer than a week; a scratch index for a demo needs almost none. ## Operating it Run optimize on a schedule rather than after every write — it rewrites data, so it costs I/O and, on object storage, requests. Sensible triggers are a nightly job, or a threshold on fragment count or deleted-row ratio if you track those. Because it is itself a write, it commits and it conflicts with concurrent writers, so it belongs in the same single-writer lane as your ingest rather than running beside it. On object storage, remember the bill is storage plus requests. Millions of tiny fragments cost more in GET requests during scans than the compacted equivalent costs in storage, and abandoned versions from a long-gone experiment can quietly dominate a bucket. ## Diagnosis in an interview Walk it as a diagnosis: compare `count_rows()` against the directory size, look at whether the growth is in data fragments, deletion files or version manifests, and match that to the write pattern — heavy deletes point at deletion files and pinned history, churny updates at superseded fragments, streaming ingest at small-file sprawl. Then prescribe compaction plus a retention window, and say out loud what the retention window costs you in lost rollback and reproducibility. The wrong answers are "delete and re-create the table", which throws away history and indexes to solve a maintenance problem, and "it will garbage collect itself", which it will not.

  • You ran compaction but the directory is the same size. What is missing?
    Version pruning. Compaction writes new consolidated fragments, but every older version still references the fragments it was built from, so nothing can be released while that history exists. Bytes come back only once versions older than your retention threshold are pruned and the files they alone referenced are deleted. Until then compaction has, if anything, added data.
  • How does a streaming ingest that appends every few seconds hurt this table?
    Each append writes its own fragment and commits its own version, so you accumulate thousands of small files and a long history. Scans then pay per-file overhead — badly on object storage, where cost and latency are per request — and pruning has far more manifests to walk. The fixes are batching writes into larger appends and running compaction on a schedule.
  • How do you choose the retention window for old versions?
    From what history is actually for. Rollback needs to reach past your worst realistic bad-backfill detection time; reproducible offline evaluations need to outlive the experiments pinned to them; audit needs whatever the policy says. Pick the longest of those, then price it — every retained version pins data files. Anything shorter silently breaks checkout on versions someone still depends on.

saying these in an interview costs you the question

  • Expecting deletes to free disk space immediately
  • Assuming old versions are garbage collected automatically
  • Thinking compaction alone shrinks the directory
  • Dropping and re-creating the table as routine maintenance
  • Running optimize alongside concurrent writers without care

context