skip to content

Flush Timing & Ordering

When accumulated changes become statements: an explicit flush, automatic before a query that would miss pending state, or flush at commit, and the order chosen. Asked because flush is not commit.

on this pageshow

questions

5

In a data-access layer that tracks loaded objects, what does a flush do, and how is it different from a commit?

level: juniorimportance: must knowfreq 82%

answer

  1. two events, not one
  2. in-memory changes become SQL
  3. transaction still open afterwards
  4. commit performs a final flush
  5. rollback erases flushed statements

basics

~20 s

A flush turns the unit of work's accumulated in-memory changes into SQL statements sent on the already-open transaction. A commit ends that transaction, making the changes durable and visible to others. Flushed-but-uncommitted work is still discarded by a rollback.

solid answer

~40 s

A tracking layer does not write on every setter. It records changes against its unit of work and waits. A **flush** walks that tracked set, decides which objects are new, changed or removed, and emits the matching `INSERT`, `UPDATE` and `DELETE` statements inside the transaction that is already open. A **commit** ends the transaction: only then is the work durable and visible to anyone else. Most layers perform a final flush as part of committing, so commit implies flush, but flush implies nothing about commit. Between the two, the rows exist for this transaction alone, the locks the statements needed are held, and a rollback still erases all of it.

go deeper

for a junior

Recall the two events: a flush writes accumulated changes as SQL on the open transaction, a commit ends the transaction. Know that a rollback still undoes everything a flush sent.

for a middle

Explain why the layer defers at all — ordering, coalescing, fewer round trips — that a commit performs a final flush, and that flushed rows stay invisible outside the transaction.

for a senior

Show that you place flush points deliberately, for a generated key or a query that must see pending work, and that you know an early flush takes row locks early and holds them to commit.

for a principal

Frame it as a contract: which layer of the code is allowed to force statements out, and what unrestricted flushing costs in lock footprint and in where errors get attributed.

## Two events, not one A data-access layer that tracks the objects it loaded does not send a statement every time you change a field. It records the change against its **unit of work** — the set of objects it is watching for this piece of work — and waits. Two distinct events then decide what the database sees. - **Flush** — the layer walks the tracked set, works out which objects are new, changed or removed, and emits the corresponding `INSERT`, `UPDATE` and `DELETE` statements on the connection, inside the transaction that is already open. - **Commit** — the transaction ends. Everything sent inside it becomes durable and becomes visible to other transactions. Flushing is a write to the *transaction*. Committing is a write to the *world*. ## What a flush leaves behind After a flush, and before a commit: - the statements have executed; the rows are there, but only inside this transaction's own view - the locks those statements needed are held, and stay held until the transaction ends - a rollback still erases every one of them — **flushed is not saved** - constraint checks the engine performs per statement have already run, so a violation is reported at the flush rather than at the line of code that changed the object - the tracked set is not emptied: the objects stay tracked, can be changed again, and a later flush will emit further statements for them ## What a commit adds A commit ends the transaction. Almost every tracking layer performs a **final flush as part of committing**, so pending changes are not silently dropped: the commit path asks the unit of work to write out whatever it still holds, then ends the transaction. That is why code that never calls flush explicitly still persists its work. | | flush | commit | |---|---|---| | turns tracked changes into SQL | yes | yes, through a final flush | | ends the transaction | no | yes | | makes the work durable | no | yes | | makes the work visible to others | no | yes | | undone by a rollback | yes | no | | may happen many times per transaction | yes | no | ## Why the separation exists Deferring statements buys the layer three things it cannot have if every setter writes immediately. 1. **Ordering.** It can send statements in an order the database will accept — parents before children, for example — instead of the order the code happened to run. 2. **Coalescing.** Ten changes to one object become a single `UPDATE`; an object created and then removed before the flush may produce no statement at all. 3. **Fewer round trips.** Statements accumulate and travel together rather than one at a time. The cost is that "the change is in my object" and "the change is in the database" become two different states, and the gap between them is where most surprises live. ## Where layers differ Layers are not uniform here. Some track loaded objects and defer everything to a flush; others — thin query layers and direct statement execution — hold no tracked set at all, send each statement as it is written, and have no flush concept, only a commit. Among the tracking ones, some flush automatically before a query whose result the pending changes could affect, while others flush only when asked, or only at commit. The honest first question on a new codebase is which of these you are working with. ## What this means in practice - Do not treat a flush as a save point. Nothing survives a rollback, and nothing outside the transaction can see it. - Do treat a flush as a **deadline**: it is where deferred errors arrive. A constraint violation surfacing far from the code that caused it is the deferral at work, not a mystery. - Reading a value the database produced — an assigned key, a column default — needs the insert to have been sent, so it needs a flush. And a flush pushes state out; it does not pull database-computed values back, so the object may still need a refresh. - Do not flush "to be safe". An early flush takes locks early and holds them until commit, and buys no durability whatsoever. ## Saying it in an interview "Flush turns the unit of work's accumulated changes into SQL on the open transaction; commit ends the transaction and makes them durable and visible. Commit implies a flush; a flush implies nothing about a commit. In between, the rows exist for me alone and a rollback still wipes them."

  • If a commit flushes anyway, why would anyone call flush explicitly?
    To read something that only exists once the statement has been sent — a database-assigned key or a column default — to make a later query see the pending work, or to have a constraint error reported at the operation that caused it instead of at the boundary. Each of those is a deliberate trade against holding locks for longer.
  • Does a flush end the unit of work or detach the objects it wrote?
    No. The objects stay tracked and stay editable; changing one again simply makes it dirty again, and the next flush emits another statement. Emptying or discarding the tracked set is a separate action from flushing it, and layers that offer both keep them separate for exactly that reason.
  • Can another transaction read rows that have been flushed but not committed?
    Not under the isolation levels real systems run at. The rows live inside the writing transaction's view until it commits; other transactions see the previous state, and may block on the locks the flush took. Only a read-uncommitted reader would see them, which is why that level is rarely used.

Flushing is handing your filled-in forms across the counter: the clerk now holds them and can reject one on the spot, but nothing is filed until the office closes the batch. That closing is the commit, and until it happens the whole pile can be handed straight back.

saying these in an interview costs you the question

  • Says flush and commit are the same operation under two names.
  • Thinks other transactions can read flushed rows before the commit.
  • Believes a flushed change can no longer be rolled back.
  • Assumes each setter sends its own UPDATE immediately.
  • Thinks a flush empties the tracked set and detaches the objects.
  • Expects constraint errors at the line that changed the object.
open as a page

Why does a mapper with a tracked set flush automatically before running a query, and when does that not help?

level: middleimportance: must knowfreq 66%

basics

~20 s

A query is evaluated by the database, which knows nothing about unwritten in-memory edits, so its filters ignore them. Many layers therefore flush pending changes before an affected query; that cannot help statements the layer never sees.

open as a page

Why does a flush fail on a unique constraint when the same unit of work deleted that row and then inserted a replacement?

level: seniorimportance: should knowfreq 54%

basics

~20 s

The unit of work does not replay statements in call order; it groups them by kind and dependency, sending inserts before deletes. So the replacement meets the original still holding the key. Flushing between them forces the order.

open as a page

After a flush fails part way through sending its statements, what is the unit of work worth afterwards?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Very little. Some statements were sent, one failed, the rest never attempted, and the tracked set no longer describes any real state. Treat it as unusable, and never reuse its objects as if they were saved.

open as a page

How do you decide where explicit flush points belong in a codebase, and what does forcing one early cost?

level: principalimportance: should knowfreq 42%

basics

~20 s

Default to flushing at the transaction boundary; treat an explicit flush as a rare exception with a stated reason. Forcing one early takes row locks sooner and holds them to commit, gives up grouping, and moves where errors surface.

open as a page