How do you decide where explicit flush points belong in a codebase, and what does forcing one early cost?
answer
- boundary owns it by default
- each exception needs a stated reason
- locks start at the flush
- grouping and ordering are given up
- never flush for durability
basics
~20 sDefault to flushing at the transaction boundary; treat an explicit flush as a rare exception with a stated reason. Forcing one early takes row locks sooner and holds them to commit, gives up grouping, and moves where errors surface.
solid answer
~40 sThe boundary should own the flush; everything else is an exception that has to justify itself. Legitimate reasons are narrow: reading a value that exists only once the statement has been sent, making a later read — especially one the layer does not issue — see pending work, or pulling a constraint error back to the operation that caused it so it is attributable. Everything else, and "flush to be safe" above all, is cargo cult: a flush provides no durability whatsoever. The cost is real. Locks are taken at the flush and held until commit, so the contention window widens; the layer loses the grouping and reordering it would have applied; and errors now surface in service code, so every caller needs an opinion about them.
go deeper
The rule to carry: do not add a flush because it seems safer. Flushing does not save anything; only a commit does, and the boundary already handles that.
Be able to justify a flush you add — a generated value, a read that must see pending work — and to state its cost in locks and round trips rather than treating it as free.
Judge existing flush points: find the ones inside shared code and loops, tie them to contention and round trips you can measure, and remove or relocate them with the reason recorded.
Set the policy and its enforcement. Decide who may force statements out, make it reviewable, and accept the trade you are making between lock footprint and error attribution across the whole codebase.
## The default: the boundary flushes The cheapest policy to reason about is one flush point per transaction, at the boundary that opened it. The unit of work accumulates for the whole operation, then writes once. That maximises what the layer can order, group and coalesce, keeps the lock window as short as the work allows, and puts every write error at one place — the boundary — where a single piece of handling can deal with it. Every explicit flush inside that span is a deviation. Deviations are sometimes right; they should be visible, local, and explained. ## When an explicit flush earns its place 1. **A value that exists only after the statement is sent.** A database-assigned identifier or a column default the code must use before the transaction ends. Note that layers can often obtain identifiers without a statement — pre-allocated key blocks, client-generated values — so check whether you need the flush before adding it. 2. **A later read that must count the pending work.** Especially one the layer does not issue itself, such as a hand-written statement or a call into a routine in the database. 3. **Error attribution.** A long operation touching many objects reports every violation at the boundary. A flush after the risky step names the step instead, which is worth real money in production diagnosis. 4. **A bounded set.** Writing out and clearing periodically so the tracked set does not grow without limit through a long operation. How those statements are grouped once sent is a separate concern. And the reason that is never valid: durability. Nothing a flush sends survives without a commit, so "flush so we don't lose it" protects nothing while paying every cost below. ## What an early flush costs - **Lock duration.** The statements take their row locks when they run and hold them until the transaction ends. Flushing halfway through a long operation can double how long a hot row is held, and lock-ordering differences between code paths turn that into deadlocks. - **Lost grouping.** The layer can only order and group what it holds. Splitting a flush into three sends three smaller sets, with more round trips and fewer opportunities to coalesce. - **Relocated failures.** Errors now surface at the flush inside service code. That is a benefit when deliberate (see attribution above) and a liability when accidental, because handling appears in places that had no error path before. - **Coupling to layer internals.** Code that depends on "the statement has been sent by now" is coupled to when the layer sends. It survives until someone moves the flush. ## Where flush points must not live - **Inside shared helpers and generic write methods.** A flush hidden behind a call every service makes imposes lock timing and error points that no call site can see. It is the most common way a codebase acquires flush-per-object behaviour that nobody chose. - **In a loop, per object, by convention.** That is the pathological case of the previous point: maximum round trips, maximum lock time, minimum grouping. - **In a read path.** If a read must flush to be correct, the write above it is in the wrong place; fix the ordering rather than adding writes to reads. - **As a workaround for a diagnosis nobody finished.** "It only works if we flush here" is a bug report, not a design. ## A policy a team can actually hold 1. The transaction boundary owns the commit, and by default owns the only flush. 2. An explicit flush requires a reason from the list above, written as a comment at the call site. 3. No flush inside shared or low-level code that callers cannot see. 4. Never flush for durability; that is what the commit is for. 5. Review write-heavy paths for lock duration, and treat a flush moved earlier as a change to the locking profile, not a formatting change. ## The tradeoff, stated plainly | flush late (at the boundary) | flush early (explicitly) | |---|---| | shortest lock window | locks held from the flush to commit | | best grouping and ordering | fewer statements per group | | every error at one place | error attributed to the step | | generated values unavailable mid-operation | values readable immediately | | reads through other paths see stale rows | pending work visible to them | Read that table as the real decision: you are trading contention and grouping for visibility and attribution. Do it where the second column names something you genuinely need, and nowhere else.
- How would you tell whether a codebase's flush points are actually hurting it?Look at where the write time and the contention are. Count the statements and round trips per operation on the hot paths, measure how long the busiest rows are locked within a transaction, and check whether deadlocks cluster on paths that flush midway. Then find the flushes: the ones inside shared helpers and loops are where the damage usually is.
- Is there a case for forbidding explicit flushes outright?It is a defensible default for teams that keep acquiring accidental ones, provided the boundary really can serve every read-after-write in the codebase. It breaks down as soon as you need a generated value mid-operation or mix hand-written statements with layer queries — so the workable version is not a ban but a review rule: allowed, rare, justified at the call site.
- Why is error attribution worth an early flush when it changes no data?Because an error at the boundary tells you an operation failed, while an error at the step tells you which object and which rule. In a long operation touching many rows, that difference is the whole diagnosis. You are buying a shorter path from alert to cause with a slightly longer lock window.
saying these in an interview costs you the question
- Flushes 'to be safe', believing it protects the work from loss.
- Puts a flush in a shared helper every service calls.
- Flushes once per object inside a write loop by convention.
- Thinks an early flush releases locks rather than taking them.
- Cannot name what an explicit flush costs, only what it fixes.
- Keeps a flush that 'makes it work' without diagnosing why.