Aborting halfway through a multi-step operation can leave the system in a partially modified state. How do you fail fast without leaving that inconsistency behind?
answer
- validate first, mutate second
- build new value, swap once
- transaction = all-or-nothing
- idempotent retry or compensating saga
- irreversible step last
basics
~20 sDo all the checking before you change anything, so a failure happens while nothing has been modified yet. When several changes must happen together, wrap them in a transaction, or make each step undoable or safely repeatable.
solid answer
~50 sFail fast and consistency are complementary if you separate decision from mutation. In-process, validate every precondition first and only then mutate; better still, build the new state on the side and swap it in with a single assignment, so a failure leaves the original untouched. Within one datastore, use a transaction so the abort rolls everything back. Across services, no shared transaction exists, so use one of: make each step idempotent and retry to completion with a durable record of intent (outbox, workflow, saga with compensating actions); or make the operation reservation-based, where the first step is a cheap reversible reservation and the irreversible step is last. Crash-only design takes it further: instead of trying to unwind in-process, fail fast by crashing, and rely on recovery from durable state on restart, which means recovery paths are exercised constantly instead of only in emergencies.
code
pseudocode · 13 lines// leaves the object half-updated if the second check fails
function rename(user, first, last) {
user.first = first
if (last == "") throw ArgumentError("last required") // first already changed
user.last = last
}
// validate-then-mutate, then build-and-swap (nothing changes unless all checks pass)
function rename(user, first, last) {
if (first == "") throw ArgumentError("first required")
if (last == "") throw ArgumentError("last required")
user.name = new Name(first, last) // single assignment
}go deeper
Say: do all the checks before changing anything, and use a database transaction when several changes must succeed together.
Add build-and-swap or immutable replacement, the strong-versus-basic exception guarantee, and why external side effects must not sit inside a transaction.
Cover cross-service work: durable intent plus idempotent roll-forward, sagas with compensating actions, the transactional outbox, and ordering irreversible steps last.
Discuss which invariants must be atomic versus eventually reconciled, the domain-modelling cost of exposing intermediate states, crash-only design as a way to have one well-exercised recovery path, and reconciliation with impossible-state alerting as the safety net.
## The tension Fail fast says stop as soon as something is wrong. But if you have already applied three of five changes, stopping leaves the system in a state that no design intended - the very thing fail fast was supposed to prevent. The resolution is not to abandon fail fast; it is to arrange the work so that the point of failure is a safe place to stop. ## Level 1: single object, in memory **Validate first, mutate second.** Do all argument and state checking at the top of the operation, before any field is assigned. This is why guards belong at the top of a function - not stylistic tidiness, but so an abort happens while nothing has changed. **Build then swap.** Where a change spans several fields, construct the new value entirely on the side and install it with a single assignment. If construction fails, the old state is intact and unchanged. This is the same reasoning as the copy-and-swap idiom and as immutable-update styles that replace a whole value rather than mutating parts of it. Immutability turns "partial update" into an impossibility. Useful vocabulary from exception-safety guarantees: - **No-throw** - the operation cannot fail. - **Strong** - if it fails, the state is exactly as before (commit-or-rollback). - **Basic** - if it fails, state remains valid but possibly changed. - **None** - anything may happen, including a broken invariant. Aim for strong on operations that guard important invariants; basic is often acceptable elsewhere; none is a bug. ## Level 2: one datastore A **transaction** provides exactly this: all-or-nothing, so failing fast anywhere inside it rolls back everything. The traps are practical rather than conceptual: - **Side effects inside a transaction that the transaction cannot roll back** - sending an email, calling a payment API, publishing a message. The database rolls back; the email is already sent. Move such effects after commit, or record intent in the same transaction and dispatch afterwards (the transactional outbox pattern). - **Failing after commit but before responding** - the client sees an error while the change is durable. Callers must therefore treat retries as possible duplicates, which is why idempotency keys matter. ## Level 3: multiple services, no shared transaction Across service boundaries there is no common rollback. Two workable shapes: 1. **Roll forward.** Persist the *intent* durably first, then drive each step to completion with retries. Every step must be **idempotent** - safe to apply repeatedly with the same effect - typically via a client-supplied idempotency key or a natural unique identifier. Failure becomes "not finished yet" rather than "half done". A workflow engine or an outbox plus a background processor implements this. 2. **Compensate.** A **saga** performs steps in order and, on failure, runs compensating actions to undo the completed ones. Compensation is business-level, not a true rollback: you cannot un-send an email, but you can send a cancellation. Sagas expose intermediate states to the outside world, which the domain must be able to express ("pending", "cancelled", "refunded"). **Ordering heuristic:** put cheap, reversible, and verifiable steps first, and the irreversible or externally visible step last. Reserve inventory (reversible) before charging a card; charge before dispatching. Then most failures happen while everything is still undoable. ## Level 4: crash-only thinking Crash-only design argues that trying to clean up carefully in-process is often more error-prone than simply failing fast by crashing and recovering from durable state on restart. Two benefits: there is only one recovery path instead of a rarely-run graceful-shutdown path plus a crash path, and that single path is exercised constantly, so it actually works. It requires that all important state is durable and that restart is cheap and idempotent. The prerequisite is precisely the transactional or roll-forward discipline above. ## What NOT to do - **Catch, log, and continue** past a failed step. That is the fail-silently pattern producing the inconsistency this question is about. - **Best-effort manual unwind in a catch block** without idempotency or durability - the unwind itself can fail or be interrupted, and now the state is unknown. - **Rely on "it will probably succeed"** for a step that has no compensation, executed first. ## Detecting what still slips through Some inconsistency will occur regardless. Design for it: reconciliation jobs that compare systems and repair or report drift, invariant checks over stored data, and alerts on states the model calls impossible. That converts an undetected inconsistency into a bounded, observable one - the same fail-fast philosophy applied after the fact.
- Why is publishing an event or sending an email inside a database transaction a problem?The external effect cannot be rolled back. If the transaction later aborts, the message has already gone and now describes something that did not happen. Record the intent in the same transaction (an outbox row) and dispatch after commit, with at-least-once delivery and idempotent consumers.
- When is a compensating saga preferable to retrying forward to completion?When a step is genuinely unable to complete - the resource is gone, the payment was declined, the business rule now forbids it - so retrying can never succeed. Roll forward is preferred when the obstacle is transient, because compensation adds business complexity and exposes intermediate states.
- What does idempotency actually require in practice?A stable key supplied by the caller or derived from the domain, storage of the outcome against that key, and returning the recorded outcome on a repeat rather than re-executing the effect. Without persisted deduplication, a retried step will apply twice, which is worse than the original failure.
A surgeon confirms blood type, consent and equipment before the first incision, not after. Once cutting has started the options are finish or repair - so all the checks that can abort the procedure happen while aborting is still free.
saying these in an interview costs you the question
- Catching an exception mid-operation, logging it, and continuing - guaranteeing exactly the inconsistent state fail fast is meant to prevent.
- Mutating fields one by one and only discovering a violated rule halfway through.
- Performing irreversible external effects (charging a card, sending mail) before cheap reversible steps.
- Publishing messages or calling external APIs inside a database transaction and assuming rollback undoes them.
- Writing ad-hoc cleanup in a catch block without idempotency or durable intent, so an interrupted cleanup leaves an unknown state.
- Assuming a saga's compensation is a true rollback rather than a business-level counter-action with its own failure modes.