skip to content

A conditional write under a version token keeps being refused although no caller touched the field you changed - why, and what do you change?

level: middleimportance: nice to knowfreq 35%

answer

  1. the store versions what, exactly
  2. the entry, not the field
  3. a false conflict it cannot see
  4. granularity sets the conflict rate
  5. split the entry, or edit at the server

basics

~20 s

The token marks the whole entry, not the field. Any write to that entry moves it and refuses the next token holder, however unrelated the changes were. Fix the granularity, or move the edit to the server.

solid answer

~50 s

The store versions the **entry**, because the entry is the unit it writes. Two callers editing different parts of one value are, to the store, two writes to the same entry, and the second one's token no longer matches. That is a **false conflict** - the entry moved, but not in a way that invalidates the computation - and the store cannot distinguish it from a real one. Three responses exist. Where the server understands the value it can change one part in place, and the read-modify-write never crosses the network at all. Where it treats values as opaque bytes, that option does not exist, and the lever is key layout: put independently-written facts in separate entries, accepting that a reader now makes several calls. Or accept the refusals, which is fine exactly while they are rare.

go deeper

for a junior

Remember the unit: the token belongs to the entry, not to anything inside it. Any write to that entry - by anyone, to any part - is enough to refuse the next caller who read it beforehand.

for a middle

Explain that contention is created by key layout rather than by the data: two facts collide exactly when they share an entry. Then name the two structural fixes and say which one needs the server to understand the value.

for a senior

Argue the trade rather than the fix. Splitting removes conflicts and adds round trips and loses any single-entry invariant across the parts; living with refusals is right while they are rare, and the way to know is the refusal ratio on that entry.

for a principal

Treat entry shape as an interface decision, because it fixes both the conflict domain and the smallest unit anything can be enforced over. Once a value is split, no mechanism on this tier will make its parts move together off one node.

## The store versioned the entry, not the field A **version token** marks an entry as a whole. The store changes it on every write to that entry, whatever part of the value the write altered - because the entry is the unit the store writes. So when two callers read one entry and each rewrites a different part of it, the first accepted write moves the entry, and the second caller's token no longer matches. Refused, even though the two changes did not overlap in any way a person would call a conflict. This is not a defect in the mechanism; it is the resolution the mechanism has. A **false conflict** - the entry moved, but not in a way that invalidates the caller's computation - is indistinguishable at the store from a real one, and there is no portable way to tell it apart. On stores that hold values as **opaque bytes**, the entry is the only unit available at all: the server never parses the value, so it could not version a part of it even in principle. Notice what this makes contention a property of. Contention here is not a property of the data, or of how often a logical fact changes. It is a property of **how the facts were packed into entries**. Two facts written by different callers at different rates are in contention if and only if they share an entry. ## Three ways out, and what each gives up | approach | what it does | what it costs | |---|---|---| | move the edit to the server | a **server-side in-place update** changes one part of the value at the store, so no read-modify-write crosses the network and no token is needed for that part | exists only where the server can interpret the value, and only for edits the store can express | | narrow the entry | independently-written facts go in separate entries, so callers editing different facts never touch the same entry | a reader wanting all of them now makes several calls, and any invariant spanning them can no longer be checked in one place | | accept the refusals | retry under a bound and carry on | honest while conflicts are rare; the wasted work grows with the number of callers sharing the entry | The first two are structural and the third is a decision to live with the cost. It is a legitimate decision: the whole appeal of optimistic control is that it is nearly free when collisions are rare, and re-shaping a keyspace to eliminate a conflict that happens twice a day is not an improvement. ## The size of the entry is the size of the waste A refused attempt on a small entry costs a small read, a small computation and a small write. A refused attempt on a large value costs a full read of it across the network, a full recompute, and a full write - all discarded. Two things therefore scale together, and they scale against you: - **The conflict rate rises with the number of facts packed into one entry**, because more callers have a reason to write it. - **The cost of each conflict rises with the size of that entry**, because every discarded attempt moved the whole value twice. A large value rewritten whole by several callers is the worst case for this mechanism, and it is also the case where each attempt occupies the server and the connection longest. ## What not to conclude - **Do not conclude the store is broken.** It refused correctly; it told you the entry moved, which is exactly what it promised and all it can know. - **Do not reach for an expiring claim first.** A claim makes callers take turns, which is the right answer when they genuinely conflict. Here they do not - it would add waiting to a conflict that ought not to exist. - **Do not expect a longer pause between attempts to fix it.** Spacing the attempts lowers the collision rate a little; the granularity that produces the collisions is untouched. - **Do not assume a structured value always saves you.** Where the server can read and change part of a value, the in-place update is available and the token is not needed for that part. Where the server stores and returns bytes, the full read-modify-write is the only shape there is, and the remedies reduce to a token, a conditional create, or a claim. - **Do not split until you know the read pattern.** Separate entries mean separate calls; a design that turns one read into six has moved the cost from conflicts into round trips, which may be the worse trade.

  • Could the caller compare old and new values itself and decide the conflict was harmless?
    It can, and sometimes that is the right design - re-read, see that only fields it does not care about changed, and merge its own change in before writing again with the new token. But that is application-level reconciliation, not something the store provides, and it only works where the value's structure is known to the caller and the merge is genuinely commutative.
  • Does splitting one entry into several make the writes safe on its own?
    It makes each write conflict-free with respect to the others, which is the point. It does not give you any relationship between them: after the split there is no way to write two of them as one unit unless they are co-located and the store offers a group. If an invariant spans the parts, splitting has moved it out of reach.

saying these in an interview costs you the question

  • Expects the store to version individual fields of a value.
  • Thinks the store merges two callers' changes to one entry.
  • Blames the token instead of the entry's granularity.
  • Reaches for a claim before looking at key layout.
  • Assumes every server can edit part of a value in place.
  • Splits entries without considering the extra read calls.