After a flush fails part way through sending its statements, what is the unit of work worth afterwards?
answer
- not an atomic act
- some sent, one failed, rest untried
- bookkeeping stopped mid-stream
- objects keep their in-memory values
- discard the set, reload, reapply inputs
basics
~20 sVery little. Some statements were sent, one failed, the rest never attempted, and the tracked set no longer describes any real state. Treat it as unusable, and never reuse its objects as if they were saved.
solid answer
~40 sA flush is a sequence of statements, not one atomic act. When one fails, earlier statements have already executed inside the transaction, the failing one has not, and the remainder were never sent — so the tracked set now claims a set of writes that only partly happened. The layer generally cannot tell you how far it got, and it cannot undo what it sent; only ending the transaction does that. Most layers therefore mark the unit of work unusable and expect it to be discarded. The subtle part is that **objects do not roll back**: fields you set stay set, and identifiers assigned during the partial flush stay attached, so carrying those instances into a fresh attempt is how stale data leaks in.
go deeper
Remember that a flush sends several statements, so a failure can leave some already executed. The unit of work should be abandoned, not patched up and flushed again.
Explain why the set is untrustworthy: bookkeeping stopped mid-sequence, a retry can double-write, and in-memory objects keep values and identifiers the database no longer backs.
Show the recovery you would actually write — abandon the set, end the transaction, rebuild from the original request — and where you validate so the error lands at the operation instead.
Own the blast radius. How wide a unit of work you allow decides how much work a single failed statement destroys and how well anyone can say afterwards what happened.
## What "part way" actually means A flush is a sequence: the unit of work derives its statements and sends them. If the fourth one is rejected — a constraint, a conversion, a conflict on a versioned row — then at the moment the error is raised: - statements one to three have **executed** inside the open transaction - statement four has **not** - statements five onward were **never attempted** - the transaction is still open, and still holds every lock those first statements took So the database holds a partial write and the tracked set holds the whole intended write. Neither describes the truth on its own. ## Why the unit of work is poisoned It is tempting to catch the error, fix the offending object and flush again. That is the trap, for reasons that stack: - **The layer usually cannot say what got through.** It knows the statement that failed; the mapping back from "statements sent" to "objects considered written" is not something it is obliged to keep accurate after an error. - **A second flush can double-write.** If the set still regards an already-inserted object as new, the retry sends the insert again. - **Internal bookkeeping may be inconsistent.** Assigned identifiers, version values and dirty marks were being updated as the statements went out, and that update stopped mid-stream. - **The transaction is already doomed in many engines.** After a failed statement, some engines refuse further work in that transaction until it is ended, so the retry cannot succeed regardless of what the layer thinks. This is why layers commonly mark the unit of work as unusable after a failed flush rather than letting you continue: the alternative is a set that writes plausible-looking nonsense. ## The part people get wrong: objects do not roll back Ending the transaction undoes the **rows**. It does nothing to the **objects**. | after a failed flush and a rolled-back transaction | state | |---|---| | rows written by the statements that ran | gone | | locks those statements held | released | | field values the code set in memory | still set | | identifiers assigned during the partial flush | still on the objects | | the tracked set's idea of what is new or dirty | stale and untrustworthy | An object carrying an identifier for a row that no longer exists is the single nastiest artefact here. Handed to another piece of work, it reads as an existing record, and the write that follows targets a row that was never committed. ## What to do instead 1. **Let the unit of work go.** Do not write through it again. Its statements are behind it and its bookkeeping is unreliable. 2. **End the transaction rather than continuing inside it.** The partial write must not be allowed anywhere near a commit. 3. **Rebuild from inputs, not from wreckage.** If the operation is to be attempted again, start a fresh unit of work, reload the rows, and reapply the values from the original request or command. The request is the durable description of intent; the half-written objects are not. 4. **Report against the cause, not the flush.** The failing statement identifies which object and which rule; that, rather than "flush failed", is what belongs in the log and the message. ## Designing so this hurts less - **Validate before the write, not at the flush.** Rules you can check in memory — required values, ranges, referenced-object presence — should fail at the operation, where the error is attributable and nothing has been sent. - **Keep the unit of work narrow.** A set spanning one logical operation loses one operation when it fails. A set spanning a whole batch of unrelated work loses everything, and can no longer tell you which part was at fault. - **Do not let half-written objects escape.** Returning entities from a failed operation, caching them, or holding them in a longer-lived structure spreads the damage past the failure. - **Distinguish the two questions.** "Are the rows safe?" is answered by the transaction. "Are these objects trustworthy?" is answered by the unit of work — and after a partial flush the answer is no, whatever the transaction did. ## The one-line version A failed flush leaves the database partly written, the objects unreverted, and the tracked set lying about both — so the unit of work is worth nothing but the error it raised.
- The transaction rolled back, so why can the objects still be dangerous?A rollback is a database operation. It restores rows and releases locks and touches nothing in memory. The objects keep the values the code set and any identifiers assigned while the statements were going out, so one can carry a key for a row that no longer exists — and downstream code will treat it as an existing record.
- How should a retry of the operation be built after a failed flush?As a genuinely fresh attempt: a new transaction, a new unit of work, rows reloaded from the database, values reapplied from the original request. Carrying the previous objects across is what turns a clean retry into a corruption. Whether retrying is appropriate at all depends on the failure, which is a separate decision from how to construct it.
- Does keeping the unit of work narrow really change the outcome of a failure?It changes what is lost and what can be said about it. A set covering one logical operation fails that operation, and the failing statement points at the object responsible. A set that accumulated many unrelated operations fails all of them together, and the error identifies one statement out of a sequence nobody can reconstruct.
It is a postal run where part of the bag was posted and the rest was not, and the bag does not record which. Sorting the remainder is guesswork; the honest move is to go back to the sender's list and start again.
saying these in an interview costs you the question
- Thinks a failed flush leaves the tracked set clean and reusable.
- Believes in-memory objects revert when the transaction rolls back.
- Assumes nothing was sent because the flush reported an error.
- Catches the error and calls flush again on the same set.
- Reuses objects from a failed unit of work as if they were saved.
- Thinks the layer can report exactly which statements got through.