You bound optimistic retries at five attempts on one contended entry - what contract does that give callers, and what happens to the one that exhausts the bound?
answer
- a bound changes the promise
- best effort, with a deadline attached
- who owns the invariant, and where
- what the exhausted caller is told
- degrades into latency, not errors
basics
~20 sThe bound trades an open-ended attempt for bounded latency and an explicit failure, turning a correctness mechanism into a best-effort one. The exhausted caller needs a real answer: fail outward, enforce downstream, serialise, or accept last-writer-wins.
solid answer
~50 sUnbounded retry promises nothing about completion - nothing reserves the entry, so a caller can lose every race for as long as the contention lasts. A bound replaces that with two guarantees worth having: bounded latency, and a definite outcome. What it does not do is decide what the outcome means. That is the design call, and there are four honest answers - fail the request outward and let the caller above decide; enforce the invariant in an engine that can open a transaction and keep this tier for the fast path; stop the racing by serialising writers on that entry; or decide the invariant does not need enforcing here and record that **last-writer-wins** is acceptable for this value. An undocumented last-writer-wins is a bug; a documented one is a design. Whichever you choose, exhaustion must be visible to operators, because this failure mode arrives as latency rather than as errors.
go deeper
Take away the simple part: a retry limit means a caller can run out of attempts, so some code has to say what happens then. Silence at that point is a bug someone will find later.
Be able to say what the bound bought and what it cost - predictable latency and a definite outcome, in exchange for the mechanism no longer guaranteeing that the write ever happens.
Show the four responses and pick one for a stated workload, then say what the loss of the entry does to the invariant you just enforced, and which signals make the degradation visible before a user reports it.
Answer the question under the question: should this tier be the coordinator at all, given no isolation level, no rollback, no cross-entry guarantee off one node and only the durability you arranged? Then write the contract down, including the tolerances you accepted.
## What the bound actually promises Unbounded optimistic retry has no completion guarantee. Nothing reserves the entry between attempts, so a caller re-enters the same race each time, and a caller that does more work between its read and its write loses more often than the others. There is no point at which the mechanism owes it a turn. A bound changes that into something a system can be built on: **a bounded worst-case latency, and a definite outcome**. That is a real improvement. But be clear about what it has done to the guarantee - the mechanism has stopped being a correctness device and become a best-effort one with an explicit failure path. Somebody now has to decide what that failure means, and if nobody decides, the code decides by accident. ## The four honest answers for the exhausted caller 1. **Fail the request outward.** Return a distinct, retryable failure and let the layer above decide - a user can be asked again, a job can be rescheduled. This is the right answer when the write is user-initiated and the invariant genuinely matters, and it is the only one of the four that never silently loses a change. 2. **Enforce the invariant downstream.** Put it in an engine that can open a transaction, hold a row and enforce the rule, and keep this tier for the fast, expendable copy. The contended write becomes a durable write somewhere that has the machinery, and the tier stops being asked to be a coordinator it was never built to be. 3. **Stop racing: serialise the writers for that entry.** An expiring claim, a single owning worker, or a queue in front of the entry all convert discarded work into orderly waiting. Throughput is bounded by one writer at a time either way; this makes that explicit and stops paying for it in wasted round trips. 4. **Decide the invariant does not need enforcing here.** For some values - a last-seen marker, a cached rendering, a best-effort counter used only for display - **last-writer-wins** is perfectly acceptable. Write it down. An undocumented last-writer-wins is a bug waiting to be found by someone downstream; a documented one is a design with a stated tolerance. ## What this mechanism can enforce at all Before choosing, be honest about the reach of what you are bounding: - It enforces **one entry**, or a declared set of entries on **one node's keyspace** where the store offers a read set and a group. Nothing wider. - Across nodes it enforces **nothing**. An invariant spanning entries that are not co-located has to live somewhere that can see all of them at once. - It enforces only **while the entry exists**. Eviction under memory pressure, an entry deadline passing, or the loss of the node holding the entry all remove the enforced state - and none of them notify the caller that once won the race for it. - It enforces only as **durably as the tier is**. Where replication is asynchronous, an accepted write is not yet a survived write; some stores in this class acknowledge only once a replica holds it, and some offer that per call. Stores differ here, so the durability of an atomic effect is a property of the deployment, not of the mechanism. That list is the real argument against making this tier the coordinator for anything whose loss requires an apology. It offers no isolation level to raise, no rollback, no cross-entry guarantee off one node, and no durability you did not separately arrange. ## What operators must be able to see This design does not fail loudly. It **degrades**: the accepted-write rate stops tracking demand, and the visible symptom is latency and load rather than errors. Unless the following are exported per entry, the incident will be diagnosed as a capacity problem and answered by adding callers, which makes it worse. - Refused attempts per accepted write, per entry. - Exhaustion count per entry, and what the system did on each one. - The attempt distribution, not its mean. - Latency of the whole retry sequence, since no single write looks slow. ## The contract to write down A usable contract names four things: which invariants this tier enforces and which it does not; what a caller is told when the bound is exhausted; what the loss of an entry does to the invariant; and how contention is visible before anyone files a ticket. Bounding the retries is how the tier behaves in the meantime. It is not an answer to any of the four.
- Is an unbounded retry ever the right choice?Occasionally - in a background job with no caller waiting, on an entry that is only intermittently contended, where finishing eventually is worth more than finishing predictably. Even then, bound the wall-clock time rather than the attempt count, and export the attempts, so an entry that has become permanently hot is discovered rather than absorbed.
- The invariant moves downstream to a durable engine. What is the tier still doing?Serving reads, absorbing the traffic the durable engine should never see, and holding state that is cheap to re-derive. The split to aim for is that this tier holds what can be rebuilt and the engine holds what must be defended. That keeps the fast path fast without asking the volatile tier to guarantee anything it cannot.
- Should the exhaustion failure look different from other failures to the caller above?Yes. Exhaustion means the write is safe to attempt again - nothing was applied, and the contention may have passed - which is quite different from a rejected or malformed write. Collapsing them into one failure produces callers that either retry what will never succeed or abandon what would have.
saying these in an interview costs you the question
- Raises the retry bound until the writes get through.
- Treats exhaustion as an error to log and swallow.
- Assumes this tier can enforce an invariant across nodes.
- Forgets that eviction or a deadline removes the enforced state.
- Calls an accepted write durable on an asynchronously replicated tier.
- Ships contention with no per-entry refusal or exhaustion signal.