skip to content

How do you decide whether a use case with several independent sub-operations runs under one boundary or one per sub-operation?

level: principalimportance: should knowfreq 44%

answer

  1. which intermediate states may exist
  2. invariants force atomicity, tidiness does not
  3. duration times contention is the cost
  4. a split creates a state to operate
  5. boring default, reviewed exceptions

basics

~20 s

Start from the invariant: only a state no reader may ever observe justifies one boundary. Otherwise weigh a long unit's lock time against the obligation a split creates - an intermediate state that must be detected, resumed and safely repeated.

solid answer

~50 s

The question is not "how many transactions" but "which states are allowed to exist". Write down what a reader or a crash must never see; if a partial result would violate a rule the system promises, those writes belong in one boundary and the cost is not negotiable. If partial progress is merely untidy, splitting is usually better: shorter units hold fewer locks, retry replays less work, and one failing sub-operation stops blocking the valuable one. What you take on in exchange is real - the intermediate state becomes a state the system can be found in, so you need idempotent re-runs, a way to detect a stalled one, and a path that finishes or abandons it. Middle grounds exist: a nested sub-scope for a tolerable failure, an independent unit for a record that must survive a rollback, and ordering so the irreversible step goes last.

go deeper

for a junior

Know that not every write in a use case has to be atomic with every other, and that the decision starts from what a partial result would mean rather than from the code.

for a middle

Explain the two costs side by side: a long boundary holds locks and replays everything on retry, while a split creates an intermediate state that has to be finished somehow.

for a senior

Show the design work a split actually demands - detection, idempotent completion, an alert on stalled items - and use nested or independent units for the single odd step instead of splitting wholesale.

for a principal

Argue for predictability: one default boundary shape, exceptions justified and documented with their recovery path, so failure semantics stay answerable across the whole codebase.

## Start from the invariant, not from the code The instinct is to ask how many transactions a use case should have. That question has no answer on its own. The answerable question is: **which intermediate states may exist?** Write the candidate bad state as a sentence a stakeholder would recognise - "an order exists whose stock was never reserved", "a balance was debited and never credited". If that sentence describes something the system genuinely promises never to happen, those writes share a boundary and the cost is simply the price of the promise. If the sentence is merely unpleasant - "the denormalised counter is briefly stale" - it is a candidate for splitting, and the burden of proof flips. Most use cases contain both kinds. The useful output of the analysis is not one number but a partition: this pair is atomic, that step is not. ## The cost of one big boundary - **Lock hold time.** Locks taken by the first write are held until the last commits. Duration multiplied by contention is the whole story of write throughput. - **All-or-nothing retry.** A transient failure in the last step replays everything, including the expensive parts that succeeded. - **Coupled availability.** A sub-operation that is slow or fragile drags the valuable one down with it; there is no "the important part landed". - **Longer tail latency.** The unit lasts as long as its slowest component, and connection demand tracks that duration. ## The cost of splitting Splitting does not remove the failure - it relocates it into a state the system can now be found in. That state has to be designed: 1. **Detectability.** Something must be able to recognise "step one done, step two not". Usually that is a status column or a pending record; if it does not exist, the split is not finished. 2. **Idempotent re-runs.** Whatever finishes the work must be safe to run twice, because it will be. Natural keys, conditional writes and version checks are the usual tools. 3. **A completion path.** A retry, a reconciliation pass or an operator action - named, owned and monitored. "It will probably work next time" is not a path. 4. **Visibility.** An alert on the age of the oldest unfinished item, because a split use case fails quietly by design. ## Middle grounds worth knowing | Technique | Use when | What it costs | |---|---|---| | One boundary | a real invariant spans the writes | lock time, coupled failure | | Nested sub-scope | an inner step may fail harmlessly | still one commit at the end | | Independent unit for one step | that step must survive a caller rollback | a second live transaction | | Split into separate units | partial progress is acceptable and recoverable | an intermediate state to operate | | Reorder the steps | one step is irreversible or expensive | nothing, and it is underused | Reordering deserves more attention than it gets. Putting the cheap, reversible, most-likely-to-fail step first, and the irreversible one last, shrinks the window in which a partial state can occur without changing the boundary count at all. ## A decision procedure 1. List the writes and name the states a partial run would leave. 2. Mark the states that violate a promise. Those writes are atomic together; stop arguing about them. 3. For the rest, estimate the boundary's duration and the contention on the rows it touches. If both are small, keep them in - simplicity wins by default. 4. Where a split is indicated, design the intermediate state first: how it is detected, resumed, and made safe to repeat. If you cannot design it, do not split. 5. Prefer a nested sub-scope or an independent unit for the single odd step before reaching for a full split; they are cheaper to reason about. 6. Write the decision and its recovery path down next to the code, because the next reader will otherwise "simplify" it back. ## The organisational angle The strongest position is usually a boring default - one boundary per use case - plus a small number of deliberate, reviewed exceptions. Per-case cleverness is expensive not because any single decision is wrong but because failure semantics become unpredictable across the codebase, and nobody can answer "what does the data look like after this fails?" without reading the implementation. Make the default the thing people do without thinking, and make the exceptions the thing they have to justify.

  • Where does reordering the steps help without changing the boundary count?
    Put the cheap, reversible, failure-prone step first and the irreversible or expensive one last. The window in which a partial state can exist shrinks to the gap after the last reversible step, and a failure before that point costs nothing but a rollback. It is the cheapest improvement available and routinely overlooked.
  • How do you argue against a team that wraps everything in one long boundary 'to be safe'?
    Ask which state they are preventing and whether anyone would notice it. Then show the price: locks held for the slowest component, retries replaying successful work, and one fragile step able to block the important one. Safety that nobody can name is usually contention nobody measured.

saying these in an interview costs you the question

  • Wraps everything in one boundary because it feels safer, without naming an invariant
  • Splits a use case without designing how the half-finished state is detected and finished
  • Assumes a retry makes a split safe even when the steps are not idempotent
  • Judges boundary size by the number of statements rather than duration and contention
  • Treats every write in a use case as equally atomic with every other