skip to content

What is the Unit of Work pattern, and what problem does it solve when an application needs to save changes to a database?

level: juniorimportance: must knowfreq 70%

answer

  1. registerNew/Dirty/Deleted
  2. commit = one transaction
  3. dirty checking via snapshot diff
  4. collapses redundant writes
  5. Hibernate Session = UoW

basics

~20 s

It's an object that remembers every change you made to your data - new records, edits, deletions - while you work, then writes them all to the database together in one go, instead of saving each change separately.

solid answer

~50 s

The Unit of Work pattern is an in-memory object that tracks every business object touched during a logical operation - which ones are new, which were modified, which should be deleted - and coordinates persisting all of them as a single atomic transaction. Rather than each save() call hitting the database immediately, changes accumulate in the Unit of Work's internal lists, and a single commit() computes the minimal, correctly-ordered set of INSERT/UPDATE/DELETE statements and wraps them in one transaction. It solves two problems: consistency (the database never sees a half-applied set of changes because one logical operation triggered five independent, individually-committed writes) and efficiency (redundant writes to the same object collapse into one, and the framework can batch statements). It's the reason ORMs let you mutate several objects and call save() once at the end.

go deeper

for a junior

Should describe the shopping-cart-style batching intuition: changes accumulate, then get saved together. Doesn't need to explain dirty-checking internals or ORM specifics.

for a middle

Should name the three registration categories (new/dirty/deleted) and explain that commit wraps everything in one transaction; should be able to point to their ORM's session/context as the concrete implementation.

for a senior

Should explain how dirty checking actually works (snapshot diffing), why write ordering matters for foreign keys, and the operational risks of a Unit of Work living longer than one logical operation.

for a principal

Should be able to discuss when the pattern doesn't fit at all - e.g., bulk data operations, distributed transactions across services - and weigh implicit ORM magic against explicit, hand-written persistence code for a given system's needs.

## The bookkeeping object **Unit of Work** is fundamentally a piece of bookkeeping: a single object, created at the start of a logical operation (often a web request or a service-layer method), that offers three registration methods — `registerNew`, `registerDirty`, and `registerDeleted` — plus a `commit` method. As business code creates a new object, modifies an existing one, or marks one for removal, it notifies the Unit of Work instead of writing to the database directly. The Unit of Work simply appends the object (or a reference to it) to the appropriate internal list. **Nothing touches the database yet.** ## What commit() does When the operation is ready to finish, a single call to `commit()` walks those three lists and: 1. works out the correct SQL for each object; 2. figures out a safe ordering that respects foreign-key constraints (parents before children on insert, children before parents on delete); 3. executes everything inside one database transaction. | Registration | The object | Statement | |---|---|---| | `registerNew` | new objects | `INSERT` | | `registerDirty` | dirty ones | `UPDATE` | | `registerDeleted` | removed ones | `DELETE` | If any statement fails, the whole transaction rolls back and the database is left exactly as it was before the operation started. ## Why the pattern exists The pattern exists because the naive alternative — having every repository or DAO method open its own connection, run its own SQL, and commit immediately — breaks down in two ways. 1. **First, consistency.** A business operation that touches five objects but commits after each one leaves a window where the database holds a partially-applied change if something fails on object three, with no way to roll back objects one and two once they are already committed. 2. **Second, efficiency.** Without central bookkeeping, the same object might get saved twice because two different code paths both called `save()` on it, or an object created and then deleted within the same operation might still generate a wasted `INSERT` followed by a `DELETE`. Unit of Work collapses redundant work and guarantees that either everything from one logical operation lands in the database or none of it does. ## The trade-off The trade-off is complexity versus control. A hand-rolled Unit of Work, or the implicit one inside an ORM, hides a meaningful amount of machinery: - it has to decide what **'dirty'** means (usually by diffing an object's current state against a snapshot taken when it was loaded); - it has to order writes correctly; - it has to decide when an implicit flush happens versus when the developer explicitly commits. That hidden machinery is exactly what makes ORMs convenient — you mutate a few fields and call `save()` once — but it also means a developer who does not understand the flush/commit timing can be surprised by when SQL actually executes. Systems that skip the pattern and write directly with explicit SQL keep full control over exactly what happens when, at the cost of writing and maintaining that transaction-coordination code by hand for every operation. ## Failure modes in production In production, Unit of Work-related failures usually show up as one of three symptoms. 1. **The first is a forgotten or short-circuited commit:** an exception thrown after registering changes but before `commit()` silently discards work that looks like it succeeded from the caller's point of view, unless the surrounding code always commits-or-rolls-back in a `finally` block. 2. **The second is scope creep:** a Unit of Work kept alive far longer than one logical operation — for example, held for an entire user session — accumulates tracked objects indefinitely, becomes a memory leak, and starts committing changes the caller no longer expects to be related to the current action. 3. **The third is ordering bugs:** a custom Unit of Work that does not correctly sort inserts and deletes against foreign-key relationships will intermittently fail with constraint-violation errors that are hard to reproduce because they depend on which objects happen to be registered in which order. ## A concrete example A concrete, widely used example is Hibernate's `Session` (and, at the JPA specification level, the EntityManager's persistence context). When you load an entity through a Session, Hibernate keeps a reference to it and a snapshot of its original field values. Any field mutation afterward is invisible to the database until Hibernate flushes. At flush time — which happens automatically before a query that might need to see the change, and always at transaction commit — Hibernate: - diffs each managed entity against its snapshot; - generates `UPDATE` statements only for entities that actually changed; - orders all the generated SQL according to the entity's dependency graph. You never call an explicit `save()` for an update; the Unit of Work notices the dirty state on its own. That implicit dirty-checking is precisely what people mean when they say an ORM 'implements Unit of Work.'

  • Why is committing a Unit of Work once per operation better than saving each object as soon as it's changed?
    Saving immediately means a failure partway through a multi-object operation leaves the database in a partially updated, inconsistent state with no way to undo the already-committed writes. Batching everything into one transaction guarantees atomicity - either the whole operation's changes land or none do. It also lets the framework de-duplicate and order writes instead of firing them off one at a time as code happens to touch objects.
  • What happens if an exception is thrown after you register changes with the Unit of Work but before commit() is called?
    Nothing has touched the database, so nothing needs to roll back at the database level - but the caller must make sure commit() is never reached, typically by structuring the code so exceptions skip straight past commit. The real risk is the opposite bug: a caller that swallows the exception and calls commit() anyway, which would persist a half-built set of changes.
  • Does every ORM implement Unit of Work the same way?
    No - Hibernate/JPA and .NET's Entity Framework use a full Unit of Work with an identity map and automatic dirty checking via the persistence context. Rails' ActiveRecord is closer to the Active Record pattern: each object generally saves itself explicitly, without a shared session tracking every loaded object.

Like a shopping cart: you add and remove items as you browse, and nothing is charged to your card until you hit 'checkout' - one transaction, not one charge per item.

saying these in an interview costs you the question

  • Thinks Unit of Work means calling save() after every single change
  • Can't explain what 'dirty' means or how it's detected
  • Doesn't know a Unit of Work wraps writes in one transaction
  • Believes committing happens automatically with no explicit or implicit trigger
  • Confuses Unit of Work with just 'a database transaction'

context