How do _seq_no and _primary_term give Elasticsearch optimistic concurrency control on document updates?
answer
- optimistic, not locking
- two numbers, not one, sent together
- one counter is per shard and monotonic
- the other changes on primary promotion
- a mismatch means HTTP 409
basics
~20 sEvery write stamps the document with the primary's sequence number and the primary term. A client re-sends both as if_seq_no and if_primary_term; the shard applies the write only if they still match the current document, otherwise it returns a 409 version conflict.
solid answer
~50 sWhen a primary shard indexes a document it assigns a monotonically increasing `_seq_no` from that shard's counter and stamps the current `_primary_term`, which increments each time a new primary is promoted for the shard. A read returns both. To do a compare-and-swap update, the client sends the values it read back as `if_seq_no` and `if_primary_term` — both are required together. The primary compares them against the document's current values and rejects the write with HTTP 409 and a `version_conflict_engine_exception` if anything changed in between, so the client re-reads and retries. The primary term is what makes this safe across failover: an operation carrying a stale term must have been produced against an old primary, so it can be rejected rather than silently overwriting. This replaced the older internal `_version` check. For read-modify-write via the `_update` API you can also let Elasticsearch retry for you with `retry_on_conflict`.
code
bash · 11 lines# Read the document and note its concurrency stamps
GET /products/_doc/1
# -> "_seq_no": 42, "_primary_term": 3
# Conditional write: applied only if nothing changed since that read
PUT /products/_doc/1?if_seq_no=42&if_primary_term=3
{
"name": "Widget",
"stock": 41
}
# -> 409 version_conflict_engine_exception if another writer got there firstgo deeper
Recall that Elasticsearch uses optimistic concurrency: you read a document's _seq_no and _primary_term and send both back to make the write conditional.
Explain the compare-and-swap mechanics, what each number is scoped to, and what a 409 version conflict obliges the client to do next.
Justify the primary term with the failover scenario, and pick correctly between surfacing conflicts, retry_on_conflict, and external versioning for a given write pattern.
Set the convention across services: which writes are conditional, how change events from the system of record are versioned so replays are idempotent, and how conflict rates are monitored as a design signal.
## The problem Elasticsearch has no transactions and no row locks. If two clients both read a document, modify a field and write it back, the second write wins wholesale and the first client's change disappears — the classic lost update. Optimistic concurrency control solves this without locking: attach the version of the document you read to the write, and let the server reject the write if the document has moved on. ## The two numbers **`_seq_no` (sequence number).** Every operation that a primary shard applies gets the next value from a per-shard counter. It is not global to the index — it is scoped to the shard — and it orders every indexing and delete operation that shard has processed. It is the backbone not only of concurrency control but also of replication and recovery, since a replica can be brought up to date by shipping the operations above its checkpoint. **`_primary_term`.** A counter, held in the cluster state, that increments every time a shard gets a new primary — that is, on every primary promotion after a failure. Every operation the primary applies is stamped with the term that was current when it applied it. Both are returned by the get API and, when requested, alongside search hits. ## Why one number is not enough Sequence numbers alone are not safe across failover. Consider a primary that becomes network-isolated but keeps serving; it may continue to assign sequence numbers while a replica has been promoted and is assigning the *same* numbers to different operations. A sequence number from the old primary could then look valid to the new one. The primary term breaks the tie: the new primary runs at a higher term, and any operation arriving with a lower term is known to originate from a superseded primary and is rejected. Sequence number plus term is a unique, totally-ordered identity for an operation on that shard. ## Using it Read the document and keep both fields, then send them back on the write: - `PUT /index/_doc/1?if_seq_no=42&if_primary_term=3` - The primary checks the stored `_seq_no` and `_primary_term` for that id. - Match: the write is applied and gets a fresh, higher `_seq_no`. - Mismatch: HTTP `409` with `version_conflict_engine_exception`, and nothing is written. Both parameters must be supplied together; sending only one is a request-validation error. On a 409 the correct client behaviour is to re-read the document, re-apply the business change to the new state, and retry — not to blindly retry the same body, which would either fail again or clobber whatever the other writer did. ## The _update API and retry_on_conflict The `_update` API performs the read-modify-write inside the shard, using a partial document or a script. It still hits version conflicts when two updates race. Setting `retry_on_conflict=N` tells Elasticsearch to redo the whole read-modify-write cycle up to N times internally before returning an error. This is the pragmatic option for commutative or idempotent updates — incrementing a counter, adding to a set — where any serialisation order is acceptable. It is *not* appropriate when the business logic needs the client to see the conflict and decide, because the retry silently applies your change to a state you never inspected. ## External versioning A different case is when the document's version is owned outside Elasticsearch — a row version or change-log offset in the system of record. There the client uses `version` together with `version_type=external`, and Elasticsearch accepts the write only if the supplied version is higher than the stored one. This makes out-of-order replays idempotent: a delayed event carrying an older version is dropped rather than resurrecting stale data. Note this is a different mechanism from `if_seq_no`/`if_primary_term`, which is Elasticsearch's own compare-and-swap. ## Bulk Each action line in a `_bulk` body can carry its own `if_seq_no` and `if_primary_term`. Because bulk returns HTTP 200 even when individual actions fail, the client must inspect the `errors` flag and the per-item `status` to spot the 409s and retry just those documents. ## What to say in an interview Name the numbers, say what each one is scoped to, explain the compare-and-swap protocol, and then produce the failover argument for why the term exists. Then mention the two escape hatches — `retry_on_conflict` for commutative updates, external versioning for a source-of-truth version — and the fact that 409 means re-read and re-apply rather than retry blindly.
- Why is _primary_term needed at all when _seq_no is already monotonic?Sequence numbers are only monotonic within a single primary. After a failover an isolated old primary could keep issuing numbers that collide with the new primary's, so a stale operation might pass a sequence-number-only check. The term increments on every promotion, so any operation carrying an older term is provably from a superseded primary and is rejected.
- When would you use retry_on_conflict instead of returning the 409 to the caller?When the update is commutative or idempotent — incrementing a counter, adding a tag to a set — so any serialisation order gives the same result. Then letting the shard redo the read-modify-write internally is simpler and cheaper. If the business logic must show the user the conflicting state or decide between versions, surface the 409 instead.
- How does external versioning differ from if_seq_no and if_primary_term?External versioning (`version` with `version_type=external`) lets the system of record own the version number: Elasticsearch accepts a write only if the supplied version is strictly higher than the stored one. It makes out-of-order or replayed change events idempotent. The seq_no/primary_term pair is Elasticsearch's own compare-and-swap against the value the client just read.
- How do you detect version conflicts inside a bulk request?A `_bulk` call returns HTTP 200 even when individual actions fail, so you must check the top-level `errors` flag and then each item's `status` and `error`. Conflicting items come back as 409 with `version_conflict_engine_exception`, and only those documents should be re-read and retried.
saying these in an interview costs you the question
- Says Elasticsearch locks the document during an update
- Sends only if_seq_no without if_primary_term
- Retries the same body unchanged after a 409
- Assumes _seq_no is unique across the whole index
- Ignores per-item errors because bulk returned HTTP 200