An operation group can fail at two different moments on an in-memory store — which two, and what does each leave stored?
answer
- two moments, not one
- refused at assembly: nothing applied
- errored while applying: partial effects
- read the result of every step
basics
~20 sEither the store refuses the group as it is assembled — nothing applied, resubmit a corrected group — or it accepts the group and a step errors while applying, leaving the other steps applied and the repair to you.
solid answer
~50 sThe first moment is assembly: on stores that check each operation as it is queued, a malformed step can cause the whole group to be refused before anything ran, so the keyspace is untouched and a corrected group can be submitted safely. The second moment is application: the group was accepted, the steps ran in order, and one of them errored on its own terms — typically because the operation did not fit the value actually stored under its key. The steps before it stay applied, on most stores the steps after it are applied too, and the result comes back as one entry per step. There is a third, quieter outcome: the connection drops and the caller never learns which of the two happened, so any repair path must start by reading the current state.
go deeper
Hold on to the difference: a group refused before it ran changed nothing, while a group that errored part-way through has already changed things.
Describe both moments and what each leaves stored, and say that the per-step results are the only record of which effects landed.
Add the third case — the outcome you never learn — and show a repair path that begins with a read and tolerates being run twice.
Set the rule for the codebase: which operations are allowed inside a group at all, and what the service promises downstream when a group half-applies.
## Why one word for "failed" is not enough "The group failed, so I retry it" is the sentence that hides the whole subject. A group can fail before a single step touched the keyspace, or halfway through changing it, and those two outcomes demand opposite responses. A third possibility — the caller never finds out which happened — demands a third. An engineer who has run one of these stores names the moment before naming the remedy. ## Moment one: refused while the group is being assembled On stores that inspect each operation as it is queued, the check is shallow but real: is this an operation the server knows, does it carry a plausible number of arguments. A step that fails that check can cause the store to refuse the group as a whole, and **no step is applied**. The keyspace is exactly as it was before the group was submitted, so correcting the offending step and submitting a fresh group is safe. Two cautions belong with this. The check is about the *shape* of the request, not about whether the operation makes sense for the value stored under that key — that can only be known when the step is applied. And not every store performs the check at all; on some, a malformed step is discovered only during application. ## Moment two: errored while the group is being applied Here the group was accepted, the steps were applied in order, and one of them failed on its own terms: the operation did not fit the value actually present, an argument was out of range, or a limit was hit. Nothing is rolled back. The steps before it stay applied, and on most stores offering this construct the steps after it are applied as well rather than skipped. What comes back is **one result per step**, one of which is an error — and that list is the only record of which effects landed. This is the moment people mean when they say a group is not a transaction. The store did what it promised (nobody else interleaved) and did not do what you assumed (make the whole thing conditional). ## The third outcome: you never find out If the connection drops after the group was submitted, the caller cannot distinguish *never applied* from *applied, and I never saw the results*. No mechanism at this tier closes that gap for you. | | Refused while assembling | Errored while applying | Outcome unknown | |---|---|---|---| | What applied | Nothing | Everything except the failed step | Unknown | | What the caller sees | A refusal for the whole group | Per-step results, one an error | A connection error | | Safe to resubmit unchanged? | Yes | No — repair from what applied | Only if every step is idempotent | | Who repairs | Nobody, nothing to repair | The caller | Whoever reads the state next | ## What varies across stores - **Whether assembly-time checking happens at all**, and how deep it goes. Never rely on it to catch a step that is wrong for the data. - **Whether the remaining steps run after one errors.** The common behaviour is that they do; confirm it rather than assuming it. - **Whether the construct exists.** On a store with no grouping construct, neither moment exists and multi-step atomicity comes from a submitted program, a version token or a single operation instead. ## Handling it in code 1. **Always read the per-step results.** Treating one acknowledgement as "the group applied" is the defect underneath most half-applied states. 2. **Branch on the moment, not on the word "error".** A refusal at assembly is a bug in your request; an error during application is a state you now own. 3. **Never blind-retry a group that errored while applying.** Re-running steps that already succeeded can double an effect, and re-running the step that failed changes nothing unless you know why it failed. 4. **Make the repair path start with a read**, so that the unknown-outcome case collapses into the other two. 5. **Prefer steps that are idempotent** — writing a whole value rather than adjusting it, or creating an entry only if it is absent — so that the third outcome is survivable rather than a puzzle. The habit worth building is to write the sentence "after this fails, the store holds X" for each of the three columns above before the group ships. If you cannot write it, the group is doing work this tier should not be trusted with.
- The connection drops right after a group was submitted. How do you decide what to do?Do not guess. Read the entries the group would have changed and decide from what is actually stored, because a dropped connection cannot distinguish a group that never applied from one that applied with the results lost. If the steps were written to be idempotent, resubmitting is safe; if any step adjusts a value rather than setting it, resubmitting can double the effect.
- Why is it wrong to conclude from one error that nothing in the group applied?Because the error describes one step, not the group. The steps before it are applied and, on most stores offering this construct, the steps after it are applied too. Only the per-step results say which effects landed, and treating the single error as a verdict on the whole unit is how half-applied states go unnoticed.
saying these in an interview costs you the question
- Treats one acknowledgement as proof every step applied.
- Assumes a refused group and a half-applied group look the same.
- Thinks a step that errors stops the remaining steps everywhere.
- Believes every store validates each step when it is queued.
- Blindly resubmits a group that already applied half of its steps.