skip to content

In a data-access layer that tracks loaded objects, what does a flush do, and how is it different from a commit?

level: juniorimportance: must knowfreq 82%

answer

  1. two events, not one
  2. in-memory changes become SQL
  3. transaction still open afterwards
  4. commit performs a final flush
  5. rollback erases flushed statements

basics

~20 s

A flush turns the unit of work's accumulated in-memory changes into SQL statements sent on the already-open transaction. A commit ends that transaction, making the changes durable and visible to others. Flushed-but-uncommitted work is still discarded by a rollback.

solid answer

~40 s

A tracking layer does not write on every setter. It records changes against its unit of work and waits. A **flush** walks that tracked set, decides which objects are new, changed or removed, and emits the matching `INSERT`, `UPDATE` and `DELETE` statements inside the transaction that is already open. A **commit** ends the transaction: only then is the work durable and visible to anyone else. Most layers perform a final flush as part of committing, so commit implies flush, but flush implies nothing about commit. Between the two, the rows exist for this transaction alone, the locks the statements needed are held, and a rollback still erases all of it.

go deeper

for a junior

Recall the two events: a flush writes accumulated changes as SQL on the open transaction, a commit ends the transaction. Know that a rollback still undoes everything a flush sent.

for a middle

Explain why the layer defers at all — ordering, coalescing, fewer round trips — that a commit performs a final flush, and that flushed rows stay invisible outside the transaction.

for a senior

Show that you place flush points deliberately, for a generated key or a query that must see pending work, and that you know an early flush takes row locks early and holds them to commit.

for a principal

Frame it as a contract: which layer of the code is allowed to force statements out, and what unrestricted flushing costs in lock footprint and in where errors get attributed.

## Two events, not one A data-access layer that tracks the objects it loaded does not send a statement every time you change a field. It records the change against its **unit of work** — the set of objects it is watching for this piece of work — and waits. Two distinct events then decide what the database sees. - **Flush** — the layer walks the tracked set, works out which objects are new, changed or removed, and emits the corresponding `INSERT`, `UPDATE` and `DELETE` statements on the connection, inside the transaction that is already open. - **Commit** — the transaction ends. Everything sent inside it becomes durable and becomes visible to other transactions. Flushing is a write to the *transaction*. Committing is a write to the *world*. ## What a flush leaves behind After a flush, and before a commit: - the statements have executed; the rows are there, but only inside this transaction's own view - the locks those statements needed are held, and stay held until the transaction ends - a rollback still erases every one of them — **flushed is not saved** - constraint checks the engine performs per statement have already run, so a violation is reported at the flush rather than at the line of code that changed the object - the tracked set is not emptied: the objects stay tracked, can be changed again, and a later flush will emit further statements for them ## What a commit adds A commit ends the transaction. Almost every tracking layer performs a **final flush as part of committing**, so pending changes are not silently dropped: the commit path asks the unit of work to write out whatever it still holds, then ends the transaction. That is why code that never calls flush explicitly still persists its work. | | flush | commit | |---|---|---| | turns tracked changes into SQL | yes | yes, through a final flush | | ends the transaction | no | yes | | makes the work durable | no | yes | | makes the work visible to others | no | yes | | undone by a rollback | yes | no | | may happen many times per transaction | yes | no | ## Why the separation exists Deferring statements buys the layer three things it cannot have if every setter writes immediately. 1. **Ordering.** It can send statements in an order the database will accept — parents before children, for example — instead of the order the code happened to run. 2. **Coalescing.** Ten changes to one object become a single `UPDATE`; an object created and then removed before the flush may produce no statement at all. 3. **Fewer round trips.** Statements accumulate and travel together rather than one at a time. The cost is that "the change is in my object" and "the change is in the database" become two different states, and the gap between them is where most surprises live. ## Where layers differ Layers are not uniform here. Some track loaded objects and defer everything to a flush; others — thin query layers and direct statement execution — hold no tracked set at all, send each statement as it is written, and have no flush concept, only a commit. Among the tracking ones, some flush automatically before a query whose result the pending changes could affect, while others flush only when asked, or only at commit. The honest first question on a new codebase is which of these you are working with. ## What this means in practice - Do not treat a flush as a save point. Nothing survives a rollback, and nothing outside the transaction can see it. - Do treat a flush as a **deadline**: it is where deferred errors arrive. A constraint violation surfacing far from the code that caused it is the deferral at work, not a mystery. - Reading a value the database produced — an assigned key, a column default — needs the insert to have been sent, so it needs a flush. And a flush pushes state out; it does not pull database-computed values back, so the object may still need a refresh. - Do not flush "to be safe". An early flush takes locks early and holds them until commit, and buys no durability whatsoever. ## Saying it in an interview "Flush turns the unit of work's accumulated changes into SQL on the open transaction; commit ends the transaction and makes them durable and visible. Commit implies a flush; a flush implies nothing about a commit. In between, the rows exist for me alone and a rollback still wipes them."

  • If a commit flushes anyway, why would anyone call flush explicitly?
    To read something that only exists once the statement has been sent — a database-assigned key or a column default — to make a later query see the pending work, or to have a constraint error reported at the operation that caused it instead of at the boundary. Each of those is a deliberate trade against holding locks for longer.
  • Does a flush end the unit of work or detach the objects it wrote?
    No. The objects stay tracked and stay editable; changing one again simply makes it dirty again, and the next flush emits another statement. Emptying or discarding the tracked set is a separate action from flushing it, and layers that offer both keep them separate for exactly that reason.
  • Can another transaction read rows that have been flushed but not committed?
    Not under the isolation levels real systems run at. The rows live inside the writing transaction's view until it commits; other transactions see the previous state, and may block on the locks the flush took. Only a read-uncommitted reader would see them, which is why that level is rarely used.

Flushing is handing your filled-in forms across the counter: the clerk now holds them and can reject one on the spot, but nothing is filed until the office closes the batch. That closing is the commit, and until it happens the whole pile can be handed straight back.

saying these in an interview costs you the question

  • Says flush and commit are the same operation under two names.
  • Thinks other transactions can read flushed rows before the commit.
  • Believes a flushed change can no longer be rolled back.
  • Assumes each setter sends its own UPDATE immediately.
  • Thinks a flush empties the tracked set and detaches the objects.
  • Expects constraint errors at the line that changed the object.