What does CouchDB actually store when you DELETE a document, and what does compaction reclaim?
answer
- Deleting appends rather than erases
- Replicas must learn about the removal somehow
- Compaction targets superseded bodies
- One thing compaction never reclaims
- Only an explicit purge removes it
basics
~20 sA CouchDB delete writes a new revision flagged _deleted with an empty body — a tombstone that stays forever so replicas learn about the deletion. Compaction reclaims the bodies of superseded non-leaf revisions; it never removes tombstones.
solid answer
~40 s`DELETE /db/{docid}?rev=…` is not an erase. It appends one more revision to the document's revision tree, marked `_deleted: true`, with the fields stripped. That **tombstone** is load-bearing: replication propagates deletions by shipping the tombstone, so a replica that has never seen it would otherwise resurrect the document on the next sync. Compaction (`POST /db/_compact`) rewrites the database file and drops the *bodies* of superseded, non-leaf revisions, keeping only their identifiers so history comparisons still work — but it leaves tombstones in place, which is why a database of nothing but deleted documents still occupies space. `POST /db/_purge` is the only way to remove revisions completely; it is local to one node, breaks replication assumptions, and can make deleted documents reappear, so it is a last resort. This describes CouchDB 3.x.
code
bash · 6 linescurl -X DELETE 'http://localhost:5984/shop/order-42?rev=3-abc'
# {"ok":true,"id":"order-42","rev":"4-9f1d"}
curl 'http://localhost:5984/shop/order-42'
# HTTP/1.1 404 Not Found
# {"error":"not_found","reason":"deleted"}go deeper
Remember that deleting in CouchDB writes a new revision marked deleted rather than erasing anything, and that the document then returns 404 and vanishes from views.
Explain why replication forces tombstones to be permanent, and state precisely what compaction reclaims: the bodies of superseded non-leaf revisions, not tombstones and not current leaves.
Show you have run this: scheduling compaction off-peak, spotting a tombstone-dominated database, choosing database rotation over purge, and knowing that purge is node-local and can be undone by a peer.
Own the lifecycle policy: decide retention and rotation up front for ephemeral workloads, set revision-history limits against conflict risk, and treat regulatory erasure as a designed procedure rather than an ad-hoc purge.
## A delete is a write In CouchDB, deleting is just another edit. `DELETE /db/{docid}?rev=3-abc` appends a child revision `4-…` to the revision tree, sets `_deleted: true` on it, and discards the document's fields. The `_id` remains, the ancestry remains, and the document still exists as far as the storage engine is concerned — it simply reads as gone: a plain `GET` returns `404 Not Found` with reason `deleted`, and the document disappears from `_all_docs` and from view results. ## Why the tombstone has to survive CouchDB replicates by comparing revision trees. If a delete were an erase, a node would see the document present on the peer and absent locally, conclude it was missing, and copy it back — deletions would never converge, and any old replica reconnecting after a month would resurrect everything it remembered. The tombstone is the positive statement "this document was deleted at revision 4-…", and it is what actually travels. That means the cost of deletion in CouchDB is not zero and never reaches zero: every document you have ever deleted leaves a permanent, small record behind. ## What compaction does `POST /db/_compact` rewrites the database file, copying forward only what is still needed. It removes: - the **bodies of superseded revisions** — the old versions along each branch that are no longer leaves; - storage overhead from the append-only file layout, since CouchDB never updates in place. It keeps: - all current leaf revisions and their bodies; - the **revision identifiers** of the discarded ancestors, up to the retention limit, because replication needs to compare ancestry; - **every tombstone**. A practical consequence: after compaction you can usually no longer fetch an old revision by `?rev=`, even though its id still appears in `_revs_info`. This is why revision history must never be used as an audit trail — compaction is expected to run, and it deletes exactly the thing an audit trail would need. View indexes are compacted separately with `POST /db/{db}/_compact/{ddoc}`, and stale index files left behind by removed or edited design documents are removed by `POST /db/{db}/_view_cleanup`. ## Revision history retention Each database has a `_revs_limit` — how many revision identifiers to keep in a document's history — readable and writable at `GET`/`PUT /db/_revs_limit`, defaulting to 1000. Once history is trimmed past that point, two replicas can hold revisions whose common ancestor has been forgotten; they can then no longer prove one descends from the other and may register a conflict instead. Lowering the limit aggressively on a database with heavy edit churn saves space and buys conflicts. Raising it does the opposite. ## Purging `POST /db/_purge` takes a map of document ids to revisions and removes them from the database entirely, tombstones included. It exists mainly for regulatory erasure and for cleaning up accidentally written giant documents. Its caveats are severe: a purge applies to the node you sent it to and does not travel through normal replication the way a deletion does, so a peer that still holds the document can replicate it straight back; and purging a document that a replication checkpoint depends on can force replications to restart. Treat purge as a deliberate operational procedure with the topology in mind, never as routine cleanup. ## Deleting while keeping fields A `DELETE` throws the body away, which is inconvenient when replication is filtered — a filter function cannot route a deletion it cannot inspect. The idiom is to `PUT` the document yourself with `_deleted: true` **and** the few fields your filter or downstream consumers need, such as a tenant id or type. The document is just as deleted, but the tombstone still carries enough information to be examined. ## Operational shape CouchDB 3.x runs compaction automatically according to a configurable schedule, and administrators typically constrain it to off-peak windows because it rewrites the file and competes for disk I/O. Databases dominated by create-then-delete traffic — job queues, session stores, anything ephemeral — accumulate tombstones without bound; the usual remedy is not purging but rotation: use a fresh database per period and drop the whole database when it ages out, since deleting a database really does free everything. ## Interview framing Say it in one line: deletion writes a tombstone, compaction reclaims superseded bodies, and tombstones are forever because replication needs them. Then show the operational instinct — rotate databases rather than purge, and never plan around revision history surviving compaction.
- A database is used as a job queue: documents are created and deleted constantly. Why does its file keep growing even with compaction running?Because every completed job leaves a permanent tombstone that compaction will not reclaim. The usual fix is not purging but rotation — write into a new database per day or week and drop the old one outright, since dropping a database frees everything including tombstones.
- Why can lowering _revs_limit cause replication conflicts?Replication decides whether one revision descends from another by comparing revision histories. Trimming history too aggressively can leave two replicas with no shared ancestor in the retained ids, so neither branch can be proven to supersede the other and CouchDB records a conflict instead of a clean update.
- How would you delete a document while keeping filtered replication able to see it?Instead of the DELETE endpoint, PUT the document with _deleted set to true plus the handful of fields your filter needs, such as a type or tenant id. It is equally deleted, but the tombstone retains enough content for a filter function or a downstream consumer to inspect it.
saying these in an interview costs you the question
- Says DELETE removes the document from disk immediately
- Believes compaction reclaims tombstones
- Uses revision history as an audit trail
- Reaches for _purge as routine cleanup
- Assumes a deleted document cannot come back via replication