One write in a run of two hundred sent to a volatile tier without waiting for replies fails — what happened to the others?
answer
- latency only
- nothing is rolled back
- the successful ones stay applied
- each reply carries its own outcome
basics
~20 sThe other one hundred and ninety-nine were executed and their effects stand. Sending a run of operations without waiting for replies buys latency and nothing else: no atomicity, no isolation, no rollback — each operation succeeds or fails on its own.
solid answer
~40 sNothing is undone. The server received two hundred independent operations and executed two hundred of them; one produced an error reply and the rest applied normally, including every one sent after the failure. The run is a network optimisation — it collapses two hundred round trips into one — and that is the whole of what it buys. It is not a group of steps run with no other caller interleaving, not a check-and-set retry and not a server-side program, so there is no all-or-nothing outcome to fall back on and no undo recorded anywhere. Practically this means two things: inspect every reply rather than the last one, and make each collapsed write safe to re-issue, because repairing the run means re-sending the failed operation yourself.
go deeper
Remember the headline: sending many operations without waiting for replies makes the request faster and changes nothing else. There is no undo, so a failure in the middle leaves everything else applied.
Explain why: the server saw independent operations and executed them independently, so it has nothing to roll back to. Then show the consequence in code — every reply is inspected, and each collapsed write is safe to re-issue.
Separate a per-operation failure from a connection-level one. The second is the dangerous case because the effects applied but the outcomes were lost, and it is what decides whether the collapsed writes have to be re-issuable at all.
Make it a rule other teams can follow: collapsing trips is a latency tool and may never appear in a design as the reason something is consistent. Where all-or-nothing behaviour is genuinely needed, it comes from a different mechanism, and the design should say which.
## What the mechanism actually is Sending a run of operations without waiting for each reply is a change to **when the caller reads**, and nothing else. Instead of write-read, write-read, write-read, the caller writes all two hundred operations onto the connection and then drains two hundred replies. The saving is real and large — two hundred round trips become one — and it is bought entirely on the network, not in the store. Nothing in that description changes how the server treats the operations. It received two hundred independent operations and executed two hundred independent operations. Reply 100 carries an error; replies 1-99 and 101-200 carry whatever those operations produced, and their effects are in the keyspace. ## The three things it does not buy 1. **No all-or-nothing outcome.** No undo is recorded for an operation that already ran, so there is nothing to roll back to. Whatever applied stays applied. 2. **No isolation.** The run does not reserve the keyspace, the connection or the server. Nothing prevents another caller's operations from executing between yours, and whether one connection's buffered operations happen to run consecutively is an implementation detail that varies between stores — it is not a property you may design against. 3. **No pre-validation.** There is no pass that inspects the run before executing it, so "the whole run was rejected because one operation was malformed" is not a behaviour of this mechanism. The positive way to say all of it: **collapsing trips buys latency, and buys latency only.** ## What a failure inside a run looks like Failures in a run come in two very different shapes, and candidates conflate them: - **A per-operation failure** — a malformed argument, an operation applied to a value whose shape does not support it, a conditional write whose condition did not hold. One reply slot carries the error; everything else is unaffected. This is the case in the question. - **A connection-level failure** — the connection drops, or the unread replies exceed what the tier will hold for that connection. Now the caller loses the **outcomes** of the whole run at once, while the operations the server already executed still applied. This is the genuinely nasty case: the caller does not know how far the run got. | | Per-operation failure | Connection-level failure | |---|---|---| | What failed | one operation | the delivery of replies | | Effects of the other operations | applied | applied, but unknown to the caller | | What the caller learns | exactly which one failed | nothing about how far the run got | | Repair | re-issue that one operation | re-issue everything that must be safe to re-issue | ## How to write code that survives this - **Read every reply.** A caller that sends two hundred operations and checks only the last one has published a silent-failure path. The per-operation outcome is the only thing the mechanism gives you; discarding it throws away its one guarantee. - **Make the collapsed writes re-issuable.** Setting a value to a known state, or an operation whose repetition is harmless, can simply be re-sent. An operation that folds the previous value into the new one — an increment, an append to a collection — cannot be blindly re-sent after a connection-level failure, because you do not know whether it already applied. - **Do not put a read-then-write dependency inside a run.** Every operation's arguments are fixed at the moment you write them onto the connection, so an operation in the run cannot use the result of an earlier one. If the logic needs that, it needs a different mechanism from the atomicity family, not a bigger run. - **Keep runs homogeneous where you can.** A run of writes that are each independently safe to re-issue is recoverable by re-sending; a run that mixes those with order-dependent work is not. ## Why interviewers ask this It separates a candidate who has read the word from one who has used the mechanism. The word used for it in most client libraries sits near words used for grouping steps together, and the two get merged into a single mental model in which "batching" implies "applied together". It does not, on any store of this class. The candidate who says "latency only — one round trip, two hundred independent outcomes, nothing rolled back, other callers may act in between" has the model right, and everything else about safe use follows from it.
- The connection drops halfway through draining the run's replies. What does the caller know?That some prefix of the operations ran and it cannot say which. The effects of everything the server executed are in the keyspace; what was lost is the caller's knowledge of them. This is why collapsed writes should be safe to re-issue: an operation that sets a value to a known state can simply be re-sent, while one that folds the previous value into the new one cannot be, because re-sending it may apply it twice.
- Can an operation later in the run use the value an earlier operation in the same run returned?No. Every operation's arguments are fixed when the caller writes it onto the connection, which happens before any reply is read, so nothing in the run can depend on another operation's result. A dependency like that needs a mechanism from the atomicity family — a check-and-set retry or a server-side program — and no amount of restructuring the run provides it.
- Does one operation taking many keys behave differently on partial failure?Yes, and not in the caller's favour. It is one operation with one outcome, so it either serves the request or fails it, and the reply does not tell you which key caused a failure. A run gives per-operation outcomes; a single many-key operation gives one. Neither undoes anything that already applied.
saying these in an interview costs you the question
- Believes the failed operation undoes the ones that succeeded
- Expects the store to validate the whole run before executing any of it
- Assumes no other caller can act while the run is in flight
- Thinks operations sent after the failing one are skipped
- Reads only the last reply and calls the run successful
- Uses a run to make a read-then-write sequence safe