skip to content

Why must a CouchDB document update send the document's current _rev, and what happens if it is stale?

level: juniorimportance: must knowfreq 72%

answer

  1. The write carries something you read earlier
  2. A precondition the server checks
  3. Mismatch means a specific HTTP status
  4. 409 Conflict, not a lock wait
  5. Re-read, re-apply intent, retry

basics

~20 s

CouchDB uses _rev for optimistic concurrency. An update must carry the _rev the client last read; if that is no longer the document's current revision, CouchDB rejects the write with HTTP 409 Conflict instead of overwriting.

solid answer

~40 s

Every CouchDB document has an `_id` and a `_rev` such as `3-8a2b…` — a generation number plus a content hash. When you `PUT /db/{docid}` or `DELETE /db/{docid}?rev=…`, CouchDB compares the `_rev` you supplied against the document's current winning revision. If they match, it appends a new child revision with the generation incremented and returns the new `_rev`. If they don't, nobody is blocked and nothing is merged: you get **HTTP 409 Conflict** and your write is discarded. That is optimistic concurrency over a stateless HTTP API — there is no lock to hold and no open transaction. The correct response to a 409 is to re-read the document, re-apply your change on top of the fresh `_rev`, and write again, ideally in a bounded retry loop.

code

bash · 8 lines
bash
curl -X GET http://localhost:5984/shop/order-42
# {"_id":"order-42","_rev":"3-8a2b","status":"paid"}

curl -X PUT http://localhost:5984/shop/order-42 \
  -H 'Content-Type: application/json' \
  -d '{"_rev":"2-1f0c","status":"shipped"}'
# HTTP/1.1 409 Conflict
# {"error":"conflict","reason":"Document update conflict."}

go deeper

for a junior

Be able to state the loop from memory: GET the document, keep its _rev, send that _rev back on the update, and expect HTTP 409 if someone else wrote first.

for a middle

Explain why a stateless HTTP API pushes concurrency control to write time rather than read time, and why the retry must re-apply the change rather than re-post a stale snapshot.

for a senior

Show the production habits: bounded retries with jitter, inspecting every entry of the _bulk_docs response array, and designing hot documents so that conflicting writers are rare rather than retried constantly.

for a principal

Own the design consequence: optimistic concurrency shifts merge responsibility to the application, so decide up front which documents are single-writer, which need an append-only shape, and what the retry budget costs under load.

## What a revision identifier is Every document stored in CouchDB carries two reserved fields: `_id`, the document's permanent key, and `_rev`, the identifier of *this particular version* of it. A `_rev` looks like `3-8a2b1c0d…`. The part before the hyphen is the **generation number** — how many edits deep this revision sits in the document's revision tree. The part after the hyphen is a hash derived from that revision's content, which keeps the identifier unique even when two writers independently produce a third-generation edit of the same document. ## The update contract CouchDB's document API is plain HTTP over JSON: - `POST /db` with a body and no `_rev` creates a new document (generation 1). - `PUT /db/{docid}` with `_rev` in the body updates an existing one. - `DELETE /db/{docid}?rev=…` marks it deleted. On an update, the server compares the `_rev` you supplied with the document's current winning revision. Match: the write is accepted, a new revision is appended as a child of the one you named, and the response returns the new `_rev`, which you must use for your next edit. Mismatch: the server answers `409 Conflict` with an error body, and your bytes are thrown away. Nothing is partially applied. ## Why this is optimistic, not pessimistic There is no row lock, no `SELECT … FOR UPDATE`, no transaction you keep open across requests. Each HTTP request stands alone. That is the point: a stateless request/response protocol cannot hold a lock across a client's think-time without inviting abandoned locks, so CouchDB detects the collision at write time instead of preventing it at read time. The cost is that the *loser* of a race has to do work — retry — rather than simply waiting. ## Handling a 409 correctly The wrong reaction is to resend the same request; the `_rev` is still stale, so it will fail forever. The wrong reaction is also to blindly re-read the document and write your whole in-memory snapshot over it, because that silently discards whatever the other writer just changed. What you want is to re-apply your *intent*: 1. `GET` the document to obtain the current `_rev` and the current body. 2. Re-apply the field-level change you actually wanted (increment the counter, append to the array, set the one field). 3. `PUT` with the fresh `_rev`. 4. Retry a bounded number of times with a little jitter; log and surface a failure if the document is so hot that you cannot win. If you write batches through `POST /db/_bulk_docs`, a stale revision does not fail the batch — the response array carries a per-document entry with `"error": "conflict"` for exactly the documents that lost, and you retry those. Code that ignores the `_bulk_docs` response array silently drops writes; this is one of the most common real-world CouchDB bugs. ## What _rev is not - **Not a timestamp or a clock.** Generation counts edits along one branch of the revision tree; it says nothing about wall-clock time and nothing about which of two replicas edited first. - **Not a version history you can rely on.** You can `GET /db/{docid}?rev=2-…` while that revision's body still exists on disk, but database compaction discards the bodies of superseded revisions and keeps only their identifiers. If you need an audit trail, store it yourself as separate documents. - **Not a lock.** Holding a `_rev` grants you nothing; it only makes your next write's precondition checkable. - **Not a guarantee against conflicting versions existing.** A 409 happens on a *direct* client write. Replication is different: it inserts revisions together with their history rather than as new edits, so a divergent edit that arrives by replication is not rejected — it is stored as a second leaf of the revision tree and surfaces as a document conflict to be resolved later. ## Interview framing The short version an interviewer wants: `_rev` is a compare-and-set token. Read it, send it back, and the server tells you whether the world moved under you. Everything else about CouchDB's concurrency behaviour — conflicts, deterministic winners, tombstones — grows out of that one rule.

  • Does a higher generation number in a _rev mean the revision was written later in real time?
    No. The generation counts how many edits deep the revision sits along one branch of the revision tree, not when it happened. Two replicas that both edit a second-generation document independently each produce a third-generation revision, and neither is 'later' in any wall-clock sense. Never use _rev to order events.
  • If you send 200 documents through _bulk_docs and three have stale revisions, what does CouchDB return?
    HTTP 201 with a per-document result array. The 197 good writes succeed and return new revisions; the three stale ones appear in the array with an error of "conflict" and are not applied. There is no all-or-nothing batch, so client code must inspect every element of the response and retry the failures.
  • Can you retrieve an older revision of a document by its _rev?
    Sometimes, but never depend on it. A GET with ?rev= works only while that revision's body is still on disk; compaction removes the bodies of superseded revisions and keeps just their identifiers so replication can reason about history. Treat revisions as a concurrency mechanism, not as document versioning.

It works like editing a shared wiki page: you save with the version number you started from, and if someone else saved in the meantime the site refuses your save rather than clobbering theirs.

saying these in an interview costs you the question

  • Calls _rev a timestamp or a version clock
  • Says CouchDB locks the document during an update
  • Retries the same PUT unchanged after a 409
  • Treats the revision history as a durable audit log
  • Ignores the per-document error entries from _bulk_docs

context