skip to content

When a CouchDB document has conflicting revisions, how does every node deterministically pick the same winner?

level: middleimportance: must knowfreq 58%

answer

  1. Both branches are kept, one is shown
  2. No node asks any other node
  3. Length of the revision history matters
  4. Ties fall back to comparing the hashes
  5. Deterministic, but semantically meaningless

basics

~20 s

CouchDB stores both branches and picks a winner by a fixed rule: a live leaf beats a deleted one, then the leaf with the highest generation number wins, and equal generations are broken by choosing the higher revision hash. Every node computes the same answer without coordinating.

solid answer

~50 s

A conflicted CouchDB document has more than one **leaf** revision in its revision tree — typically because two replicas edited it independently and then replicated. CouchDB does not reject or merge those branches; it keeps them all and deterministically elects one to show. The ordering is: non-deleted leaves are preferred over deleted ones, then the leaf with the longest revision history (highest generation number), then, if still tied, the leaf whose revision hash sorts highest as a string. Because the rule depends only on data that every replica holds, every node picks the same winner without any coordination — that is what makes multi-master replication work. The winner is not "the correct one", just a stable one. Read the losers with `GET /db/{docid}?conflicts=true`, merge them in application code, and delete the losing leaves once you have.

code

bash · 5 lines
bash
curl 'http://localhost:5984/shop/order-42?conflicts=true'
# {"_id":"order-42","_rev":"3-bbb","status":"shipped",
#  "_conflicts":["3-aaa"]}

curl 'http://localhost:5984/shop/order-42?rev=3-aaa'

go deeper

for a junior

Know that two nodes can edit the same document independently, that CouchDB keeps both versions, and that a normal read shows only one of them.

for a middle

Be able to state the three-step ordering — deleted last, highest generation, highest revision hash — and explain why a rule computed purely from local data lets replicas agree without talking to each other.

for a senior

Demonstrate the operational side: a sweeper view to detect conflicts, a merge-and-delete routine in one bulk write, and document shapes chosen so concurrent writers rarely touch the same document.

for a principal

Own the guarantee boundary: deterministic winner selection buys convergence, not business correctness, so decide which data may be auto-merged, which needs human adjudication, and what the conflict backlog costs the product.

## Conflicts are stored, not rejected A direct write with a stale `_rev` fails with HTTP 409. Replication is different: it transfers revisions along with their ancestry, so a divergent edit made on another node does not look like a stale write — it looks like a second branch of the same revision tree. CouchDB accepts it. The document now has two (or more) **leaf** revisions: revisions with no children. That is what "the document is in conflict" means, and it is a normal, expected state in a multi-master system, not corruption. ## The winner-selection rule A plain `GET /db/{docid}` must return exactly one JSON body, so CouchDB elects one leaf as the winner. It sorts the leaves by three keys, in order: 1. **Deleted or not.** A leaf that is a deletion tombstone loses to a live leaf. Only if every leaf is deleted does the document read as deleted. 2. **Generation number.** The leaf whose revision history is longest — the higher number before the hyphen — wins. More edits deep beats fewer. 3. **Revision hash.** If two leaves sit at the same generation, CouchDB compares the hash strings and takes the higher one. Every input to that comparison is stored with the document itself, so any replica holding the same set of leaves computes the same winner. No node has to ask another node anything. That determinism is exactly why CouchDB can accept writes on every node while still converging on one visible document everywhere. ## The rule is arbitrary on purpose Step 3 is a coin flip dressed as an algorithm — the hash carries no meaning. Step 2 is nearly as arbitrary: a branch that received three trivial edits outranks a branch that received one important one. So "winning" carries no semantic weight whatsoever. If your application cares which version is right, it must resolve the conflict; if it just reads the winner and moves on, the other user's edit is invisible but still on disk, quietly replicating everywhere and consuming space. ## Finding the losers A plain read hides conflicts entirely, and `_all_docs` will not tell you either. To see them: - `GET /db/{docid}?conflicts=true` adds a `_conflicts` array listing the losing (non-deleted) leaf revision ids. - `GET /db/{docid}?deleted_conflicts=true` shows losing leaves that are deletions. - `GET /db/{docid}?revs_info=true` shows every known revision with its availability. - To find conflicts across a whole database, write a view whose map function checks for the field and emits the id, and query that view periodically. That is the standard "conflict sweeper" pattern, because otherwise nothing tells you conflicts exist. ## Resolving a conflict Resolution is application work, and it is always the same shape: 1. Read the document with `conflicts=true` to get the winner plus the losing revision ids. 2. Fetch each losing revision's body with `?rev=…`. 3. Merge them using domain rules — union the sets, take the max, prefer the later business timestamp you stored in the body, or ask a human. 4. Write the merged body as a new revision on top of the winner. 5. Delete each losing leaf by writing it with `_deleted: true` at its own revision. Steps 4 and 5 can go in one `_bulk_docs` call, which is the usual implementation. Deleting the losers matters: until you do, the conflict remains and will keep resurfacing on every replica. ## Designing so conflicts are rare or harmless Experienced teams do not lean on conflict resolution as a routine path. They shape documents so that concurrent edits do not collide: split a mutable hot field into its own document, model an activity stream as many immutable documents rather than one growing array, or keep per-device documents that are merged by a reader instead of a writer. Where conflicts are unavoidable — genuinely offline-first apps — the merge function should be commutative and idempotent so that running it on different nodes in different orders converges anyway. ## Interview framing The answer an interviewer is listening for has two halves. First, the mechanical rule: deleted-loses, then highest generation, then highest hash. Second, and more important, the judgment: the rule guarantees *convergence*, not *correctness*, so an application that never inspects `_conflicts` is silently losing user edits while believing replication worked.

  • After the winner is chosen, what happens to the losing revision?
    Nothing automatic. It stays on disk as another leaf of the revision tree, replicates to every other node, and keeps the document flagged as conflicted until the application explicitly deletes it by writing that revision with _deleted true. Unresolved conflicts accumulate and quietly consume space.
  • How would you find every conflicted document in a CouchDB database?
    Write a view whose map function requests the conflicts metadata and emits the document id when a conflict is present, then query that view on a schedule as a sweeper job. A plain GET hides conflicts and _all_docs does not report them, so without such a view conflicts stay invisible.
  • Why does CouchDB's tie-break use the revision hash rather than a timestamp?
    Because a timestamp would require trustworthy clocks on every node, and offline-first clients routinely have wrong ones. The hash is already present on every replica and yields the same ordering everywhere with zero coordination. It is arbitrary, but arbitrary and identical everywhere beats plausible and divergent.

saying these in an interview costs you the question

  • Says CouchDB merges the conflicting bodies automatically
  • Assumes the winner is the most recently written revision
  • Believes the loser is discarded once a winner is picked
  • Thinks a plain GET reveals that a document is conflicted
  • Calls a conflicted document corrupt or broken

context