skip to content

CouchDB

A document store built around an HTTP and JSON interface, MVCC revisions, and multi-master replication with explicit conflict handling. Interviewers ask about it for offline-first sync, where replicating to a device and merging later is the entire point.

on this pageshow

questions

12

Why must a CouchDB document update send the document's current _rev, and what happens if it is stale?

level: juniorimportance: must knowfreq 72%

answer

  1. The write carries something you read earlier
  2. A precondition the server checks
  3. Mismatch means a specific HTTP status
  4. 409 Conflict, not a lock wait
  5. Re-read, re-apply intent, retry

basics

~20 s

CouchDB uses _rev for optimistic concurrency. An update must carry the _rev the client last read; if that is no longer the document's current revision, CouchDB rejects the write with HTTP 409 Conflict instead of overwriting.

solid answer

~40 s

Every CouchDB document has an `_id` and a `_rev` such as `3-8a2b…` — a generation number plus a content hash. When you `PUT /db/{docid}` or `DELETE /db/{docid}?rev=…`, CouchDB compares the `_rev` you supplied against the document's current winning revision. If they match, it appends a new child revision with the generation incremented and returns the new `_rev`. If they don't, nobody is blocked and nothing is merged: you get **HTTP 409 Conflict** and your write is discarded. That is optimistic concurrency over a stateless HTTP API — there is no lock to hold and no open transaction. The correct response to a 409 is to re-read the document, re-apply your change on top of the fresh `_rev`, and write again, ideally in a bounded retry loop.

code

bash · 8 lines
bash
curl -X GET http://localhost:5984/shop/order-42
# {"_id":"order-42","_rev":"3-8a2b","status":"paid"}

curl -X PUT http://localhost:5984/shop/order-42 \
  -H 'Content-Type: application/json' \
  -d '{"_rev":"2-1f0c","status":"shipped"}'
# HTTP/1.1 409 Conflict
# {"error":"conflict","reason":"Document update conflict."}

go deeper

for a junior

Be able to state the loop from memory: GET the document, keep its _rev, send that _rev back on the update, and expect HTTP 409 if someone else wrote first.

for a middle

Explain why a stateless HTTP API pushes concurrency control to write time rather than read time, and why the retry must re-apply the change rather than re-post a stale snapshot.

for a senior

Show the production habits: bounded retries with jitter, inspecting every entry of the _bulk_docs response array, and designing hot documents so that conflicting writers are rare rather than retried constantly.

for a principal

Own the design consequence: optimistic concurrency shifts merge responsibility to the application, so decide up front which documents are single-writer, which need an append-only shape, and what the retry budget costs under load.

## What a revision identifier is Every document stored in CouchDB carries two reserved fields: `_id`, the document's permanent key, and `_rev`, the identifier of *this particular version* of it. A `_rev` looks like `3-8a2b1c0d…`. The part before the hyphen is the **generation number** — how many edits deep this revision sits in the document's revision tree. The part after the hyphen is a hash derived from that revision's content, which keeps the identifier unique even when two writers independently produce a third-generation edit of the same document. ## The update contract CouchDB's document API is plain HTTP over JSON: - `POST /db` with a body and no `_rev` creates a new document (generation 1). - `PUT /db/{docid}` with `_rev` in the body updates an existing one. - `DELETE /db/{docid}?rev=…` marks it deleted. On an update, the server compares the `_rev` you supplied with the document's current winning revision. Match: the write is accepted, a new revision is appended as a child of the one you named, and the response returns the new `_rev`, which you must use for your next edit. Mismatch: the server answers `409 Conflict` with an error body, and your bytes are thrown away. Nothing is partially applied. ## Why this is optimistic, not pessimistic There is no row lock, no `SELECT … FOR UPDATE`, no transaction you keep open across requests. Each HTTP request stands alone. That is the point: a stateless request/response protocol cannot hold a lock across a client's think-time without inviting abandoned locks, so CouchDB detects the collision at write time instead of preventing it at read time. The cost is that the *loser* of a race has to do work — retry — rather than simply waiting. ## Handling a 409 correctly The wrong reaction is to resend the same request; the `_rev` is still stale, so it will fail forever. The wrong reaction is also to blindly re-read the document and write your whole in-memory snapshot over it, because that silently discards whatever the other writer just changed. What you want is to re-apply your *intent*: 1. `GET` the document to obtain the current `_rev` and the current body. 2. Re-apply the field-level change you actually wanted (increment the counter, append to the array, set the one field). 3. `PUT` with the fresh `_rev`. 4. Retry a bounded number of times with a little jitter; log and surface a failure if the document is so hot that you cannot win. If you write batches through `POST /db/_bulk_docs`, a stale revision does not fail the batch — the response array carries a per-document entry with `"error": "conflict"` for exactly the documents that lost, and you retry those. Code that ignores the `_bulk_docs` response array silently drops writes; this is one of the most common real-world CouchDB bugs. ## What _rev is not - **Not a timestamp or a clock.** Generation counts edits along one branch of the revision tree; it says nothing about wall-clock time and nothing about which of two replicas edited first. - **Not a version history you can rely on.** You can `GET /db/{docid}?rev=2-…` while that revision's body still exists on disk, but database compaction discards the bodies of superseded revisions and keeps only their identifiers. If you need an audit trail, store it yourself as separate documents. - **Not a lock.** Holding a `_rev` grants you nothing; it only makes your next write's precondition checkable. - **Not a guarantee against conflicting versions existing.** A 409 happens on a *direct* client write. Replication is different: it inserts revisions together with their history rather than as new edits, so a divergent edit that arrives by replication is not rejected — it is stored as a second leaf of the revision tree and surfaces as a document conflict to be resolved later. ## Interview framing The short version an interviewer wants: `_rev` is a compare-and-set token. Read it, send it back, and the server tells you whether the world moved under you. Everything else about CouchDB's concurrency behaviour — conflicts, deterministic winners, tombstones — grows out of that one rule.

  • Does a higher generation number in a _rev mean the revision was written later in real time?
    No. The generation counts how many edits deep the revision sits along one branch of the revision tree, not when it happened. Two replicas that both edit a second-generation document independently each produce a third-generation revision, and neither is 'later' in any wall-clock sense. Never use _rev to order events.
  • If you send 200 documents through _bulk_docs and three have stale revisions, what does CouchDB return?
    HTTP 201 with a per-document result array. The 197 good writes succeed and return new revisions; the three stale ones appear in the array with an error of "conflict" and are not applied. There is no all-or-nothing batch, so client code must inspect every element of the response and retry the failures.
  • Can you retrieve an older revision of a document by its _rev?
    Sometimes, but never depend on it. A GET with ?rev= works only while that revision's body is still on disk; compaction removes the bodies of superseded revisions and keeps just their identifiers so replication can reason about history. Treat revisions as a concurrency mechanism, not as document versioning.

It works like editing a shared wiki page: you save with the version number you started from, and if someone else saved in the meantime the site refuses your save rather than clobbering theirs.

saying these in an interview costs you the question

  • Calls _rev a timestamp or a version clock
  • Says CouchDB locks the document during an update
  • Retries the same PUT unchanged after a 409
  • Treats the revision history as a durable audit log
  • Ignores the per-document error entries from _bulk_docs

context

open as a page

When a CouchDB document has conflicting revisions, how does every node deterministically pick the same winner?

level: middleimportance: must knowfreq 58%

basics

~20 s

CouchDB stores both branches and picks a winner by a fixed rule: a live leaf beats a deleted one, then the leaf with the highest generation number wins, and equal generations are broken by choosing the higher revision hash. Every node computes the same answer without coordinating.

open as a page

What does CouchDB's _changes feed return, and how do the normal, longpoll and continuous modes differ?

level: middleimportance: must knowfreq 70%

basics

~20 s

The _changes feed lists documents that changed since a given sequence, one row per document with its current revision. feed=normal returns immediately, feed=longpoll holds the request open until a change arrives, feed=continuous streams rows indefinitely.

open as a page

Which HTTP calls does a CouchDB replication make to move changes from source to target?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Replication reads the source's _changes feed, asks the target's _revs_diff which of those revisions it lacks, fetches the missing ones with open_revs, and posts them to the target's _bulk_docs with new_edits set to false. Progress is checkpointed in _local documents.

open as a page

What does PouchDB's db.sync(remote, {live: true, retry: true}) do in an offline-first app?

level: juniorimportance: should knowfreq 40%

basics

~20 s

It starts two continuous replications, local to remote and remote to local, that keep running as changes occur and reconnect automatically after failures. Writes land in the local database first, so the app works offline and catches up when connectivity returns.

open as a page

What does CouchDB actually store when you DELETE a document, and what does compaction reclaim?

level: middleimportance: should knowfreq 44%

basics

~20 s

A CouchDB delete writes a new revision flagged _deleted with an empty body — a tombstone that stays forever so replicas learn about the deletion. Compaction reclaims the bodies of superseded non-leaf revisions; it never removes tombstones.

open as a page

What does emit(key, value) in a CouchDB map function build, and when is that index updated?

level: middleimportance: should knowfreq 56%

basics

~20 s

Each emit call adds one row to a persistent B-tree sorted by key, so a view is a precomputed, ordered index you query by key or key range. CouchDB updates it incrementally, by default when the view is next queried.

open as a page

Why does a filtered CouchDB replication silently stop propagating deletions, and how do you fix the filter?

level: middleimportance: should knowfreq 45%

basics

~20 s

Deleting a CouchDB document leaves a tombstone containing only _id, _rev and _deleted, so a filter testing a business field such as doc.type no longer matches and the deletion is skipped. Fix it by returning true whenever doc._deleted is set.

open as a page

When does a CouchDB _find query fall back to a full database scan, and how do you detect it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

If no declared index covers the selector's fields, CouchDB answers the _find query by scanning every document through the built-in all-docs index and returns a warning saying no matching index was found. POST to _explain to see which index a selector will actually use.

open as a page

After bidirectional CouchDB sync, how do you find and resolve documents that ended up in conflict?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Replication stores divergent writes as extra leaf revisions rather than rejecting them, so an ordinary GET shows only the winner. Find conflicts with conflicts=true or a _changes feed using style=all_docs, then write a merged revision and delete every losing revision.

open as a page

Why do CouchDB offline-sync deployments use a database per user instead of one filtered replication?

level: principalimportance: should knowfreq 28%

basics

~20 s

CouchDB authorizes readers per database, not per document, so a replication filter only shapes what a sync transfers and never restricts what a client may read. Confidentiality between users therefore requires giving each user their own database.

open as a page

Why does a CouchDB reduce function receive a rereduce flag, and what must the function guarantee?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

CouchDB stores partial reduce results inside the view B-tree's inner nodes, so it must combine already-reduced values as well as raw emitted ones. The rereduce flag tells the function which it is being given, and the output must stay small no matter how much input it summarizes.

open as a page