skip to content

After bidirectional CouchDB sync, how do you find and resolve documents that ended up in conflict?

level: seniorimportance: should knowfreq 55%

answer

  1. Replication never returns 409
  2. Both versions kept, one wins
  3. A plain GET hides the losers
  4. conflicts=true reveals _conflicts
  5. Delete the losing revisions to finish

basics

~20 s

Replication stores divergent writes as extra leaf revisions rather than rejecting them, so an ordinary GET shows only the winner. Find conflicts with conflicts=true or a _changes feed using style=all_docs, then write a merged revision and delete every losing revision.

solid answer

~50 s

Replication never returns a 409. It writes with `new_edits: false`, so a revision that diverges from what the target holds becomes another **leaf** of that document's revision tree. Every replica ends up with the same set of leaves and independently picks the same winner, which means the cluster stays readable and consistent — and also means a plain `GET` hides the problem completely. To surface it, request `GET /db/{id}?conflicts=true` and read the `_conflicts` array of losing leaf revisions, follow `_changes?style=all_docs` which lists every leaf, or maintain a view that emits when `doc._conflicts` is present so you can sweep the whole database. Resolution is application work: fetch each leaf body, merge according to your domain rules, then in one `_bulk_docs` write the merged body as a new revision of the winner **and** a deleted revision for each loser. Skip that last part and the conflict persists and replicates onward.

code

bash · 5 lines
bash
# surface the losing leaf revisions for one document
curl 'http://host:5984/orders/order:912?conflicts=true'

# read a losing leaf's body
curl 'http://host:5984/orders/order:912?rev=4-a71f'

go deeper

for a junior

Know that two replicas can both accept a write to the same document, that CouchDB keeps both, and that a normal read shows only one of them without warning.

for a middle

Explain why replication stores a divergent write as an extra leaf instead of returning 409, and show how conflicts=true and style=all_docs expose the leaves a plain read hides.

for a senior

Demonstrate an operational strategy: a sweeper over conflicted documents, a merge plus loser-deletion in one _bulk_docs, and awareness that a half-resolved document replicates its unresolved leaves onward.

for a principal

Own the policy — who resolves, how merges stay deterministic across replicas, and which data should be modelled as single-writer or immutable so that conflicts cannot arise at all.

## Why sync creates conflicts at all CouchDB is multi-master by design. Two replicas — two servers, or a server and a phone that has been offline for a day — can each accept a write to the same document without talking to each other. When they later replicate, both versions have to survive: a replication that dropped one of them would be losing an acknowledged write. The mechanism is the write path. Replication uses `POST /_bulk_docs` with `new_edits: false`, which stores the incoming revision ids and ancestry exactly as supplied instead of validating a parent revision. So the divergent revision is not rejected with a 409, as an ordinary client write against a stale revision would be; it is attached as an additional **leaf** on the document's revision tree. ## Why the problem is invisible by default When a document has several leaves, CouchDB picks one as the winner using a rule that is deterministic and depends only on the tree, so every replica selects the same one without coordination. That is a genuinely useful property: reads never block, never fail, and never disagree between replicas. The cost is that a plain `GET /db/{id}` returns the winner with no hint that anything else exists. A conflicted database looks entirely healthy from the application's normal read path. Losing data here does not require a bug in CouchDB — only an application that never looks. ## Finding conflicts Three mechanisms, in increasing order of scope: **Per document.** `GET /db/{id}?conflicts=true` adds a `_conflicts` array listing the losing leaf revisions: ```json {"_id": "order:912", "_rev": "4-cf9b", "total": 40, "_conflicts": ["4-a71f"]} ``` **As they happen.** `_changes?style=all_docs` lists every leaf revision per row rather than only the winner, so a consumer watching the feed can spot a document that suddenly has two leaves and react immediately. Adding `include_docs=true&conflicts=true` inlines the bodies with their conflict markers. **Across the database.** A view in a design document that emits only when `doc._conflicts` is present gives you a queryable list of every conflicted document, which is what you want for a background sweeper or a monitoring check. Whichever mechanism you use, the important design decision is that *something* looks — a sweeper job, a change-feed consumer, or a check at read time in the application. ## Resolving a conflict Resolution is always application logic, because only the domain knows what merging two versions means. The mechanical steps are fixed: 1. Read the winner and the conflict list. 2. Fetch each losing leaf's body — `GET /db/{id}?rev=<losing-rev>`, or all leaves at once with `open_revs`. 3. Merge according to your rules: field-wise merge, pick by business timestamp, sum the sets, or escalate to a human. 4. Write the merged body as a new revision **on top of the winner**, and write a deleted revision for **each** losing leaf. Doing both in a single `_bulk_docs` keeps it atomic-ish from the application's point of view and avoids a window where the document looks resolved but is not. The step people skip is (4)'s second half. If you merge and PUT a new revision on the winner but leave the losing leaves alone, the document still has multiple leaves: it is still conflicted, `_conflicts` still reports it, and the unresolved leaf will replicate back out to every peer. Deleting the loser is what actually collapses the tree to one leaf. Note that a deleted losing revision is still a revision. It remains in the tree, occupying space, until database compaction removes what is no longer needed. ## Designing conflicts away Because resolution is real work, the strongest teams reduce how often it is needed: - **Make documents immutable.** Write a new document per event instead of updating one document, and compute the current state by reading them. Two devices writing two different documents cannot conflict. - **Partition by writer.** Give each device or user its own document (or its own database) for the data it owns, and merge on read. There is exactly one writer per document, so divergence is impossible. - **Keep documents small and single-purpose.** A document holding ten unrelated fields conflicts whenever any two of them are edited concurrently on different replicas; splitting it removes most false conflicts. - **Retry the write path.** An ordinary client write against a stale revision *does* get a 409, and retrying with the fresh revision avoids creating a conflict in the first place. Conflicts from replication are the ones you cannot avoid this way, which is precisely why they need a deliberate strategy. A CouchDB deployment that syncs to offline clients and has no conflict-handling code has not avoided conflicts; it has decided, implicitly, that the deterministic winner is always right and everything else can be discarded silently. Say that out loud in an interview and you have named the risk.

  • You merge the bodies and PUT a new revision on the winner, but leave the losing revisions in place. What is the state of the document?
    Still conflicted. The document now has your merged leaf plus every untouched losing leaf, so `_conflicts` keeps reporting it and those leaves replicate back out to every peer. Collapsing the tree requires writing a deleted revision for each loser, ideally in the same `_bulk_docs` as the merge, so the document never sits in a half-resolved state.
  • Two replicas independently resolve the same conflict at the same time. What happens?
    You get a new conflict: each resolver wrote its own merged revision on top of the winner, and those two merges are themselves divergent leaves. The usual defences are to make merges deterministic, so both replicas compute the identical body, or to designate a single resolver — a server-side sweeper rather than every client — so only one process ever writes resolutions.
  • How would you keep a conflicted document from being invisible to the application?
    Make something look on purpose. Run a background sweeper over a view that emits when `doc._conflicts` is present, or watch `_changes?style=all_docs` and react when a document gains a second leaf. Some applications also read with `conflicts=true` on the hot path for high-value documents. The failure mode to avoid is a system where only the winner is ever read.
  • Which design choices make conflicts structurally impossible for a piece of data?
    Single-writer ownership and immutability. If each device or user writes only its own documents and the application merges on read, no two writers ever touch the same document. Writing an immutable document per event instead of mutating a running total has the same effect. Both trade a read-time merge for the elimination of write-time divergence.

It is a shared document where two editors saved at once: the system shows you one version and quietly files the other in a drawer. Nothing is lost, but nothing is merged until someone opens the drawer.

saying these in an interview costs you the question

  • Thinks replication rejects the losing write with a 409
  • Believes CouchDB merges document bodies automatically
  • Assumes a plain GET signals that a conflict exists
  • Resolves by writing a merged revision without deleting losers
  • Says the deterministic winner is always the correct version

context