skip to content

In what situations would you deliberately avoid relying on a Unit of Work / ORM-managed persistence context, and use explicit, hand-written persistence logic instead?

level: principalimportance: nice to knowfreq 35%

answer

  1. bulk/ETL -> set-based SQL beats per-row UoW
  2. cross-service writes -> Saga + compensating actions, not distributed 2PC via UoW
  3. hot write path -> bypass ORM tracking for raw batched inserts
  4. CQRS read side -> event-driven projection, no dirty-checking needed
  5. UoW fits bounded, single-DB, per-object correctness

basics

~20 s

When you're changing huge numbers of rows at once (bulk jobs), or writing across multiple independent services/databases where one shared transaction isn't possible, the automatic 'track everything then save' approach costs more than it helps - explicit, targeted SQL or separately coordinated writes work better.

solid answer

~50 s

Unit of Work shines for typical request-scoped, single-database business operations touching a modest object graph, but it becomes a liability in a few recognizable situations. Bulk/ETL-style operations that touch millions of rows are much cheaper as set-based SQL than as millions of individually loaded-and-tracked entities, because loading each row into the identity map and diffing it is pure overhead when you already know exactly what you want written. Cross-service or cross-database writes in a microservices architecture can't share one local ACID transaction at all, so a single in-process Unit of Work can't coordinate them - that requires a different pattern, like the Saga pattern with compensating actions. High-throughput write paths where every millisecond and every byte of per-request memory matters often bypass the ORM's tracking machinery in favor of direct, prepared-statement SQL. And CQRS-style read/write segregation typically keeps the write side using Unit of Work-style consistency, but does not use it for the read side, which is often a denormalized projection updated by very different, non-transactional means.

go deeper

for a junior

Not expected to have a firm answer here; a reasonable response is recognizing that 'saving one thing at a time' might sometimes be simpler, without needing to name alternative patterns.

for a middle

Should recognize that bulk operations on huge datasets are usually done differently from normal ORM saves, even if they can't fully explain why.

for a senior

Should be able to explain the per-row overhead argument for bulk operations and know that cross-service writes need a different coordination mechanism entirely.

for a principal

Should be able to name and reason about the alternative patterns directly - Saga with compensating actions, set-based SQL, CQRS read/write asymmetry - and make the call on which one fits a given system's constraints.

## When the default fits Unit of Work is a good default for the common case — a request-scoped or otherwise bounded business operation, working against a single database, touching a moderate number of related objects — but a principal engineer needs to recognize the several situations where its core assumptions stop holding and where reaching for it anyway actively hurts the system. ## Bulk and ETL-style data operations The first is bulk and ETL-style data operations. Unit of Work's value proposition is dirty-checking convenience: load an object, mutate a few fields, let the framework figure out the `UPDATE`. That convenience has a fixed per-row cost — identity map insertion, snapshot storage, diffing at flush — which is negligible for tens or hundreds of rows in a request but becomes the dominant cost when applied to millions of rows in a batch job. A set-based `UPDATE orders SET status = 'shipped' WHERE created_at < ?` does the same logical work as loading a million `Order` entities, mutating a status field on each, and flushing — but does it in one database round-trip with none of the per-row object/tracking overhead. Any operation naturally expressible as 'change this column for all rows matching this predicate' should almost always bypass the Unit of Work and go straight to set-based SQL or a bulk-loading tool, reserving the ORM's tracking machinery for genuinely per-object business logic that can't be expressed declaratively. ## Cross-service or cross-database writes The second, and architecturally the most important, is cross-service or cross-database writes. Unit of Work fundamentally coordinates writes into one local ACID transaction against one database connection. In a microservices architecture, a single business operation frequently needs to update state owned by two or more independently deployed services, each with its own database — and there is no single local transaction that can span both, because a shared database is precisely what microservices decomposition is meant to avoid. Reaching for a distributed transaction protocol (two-phase commit) to recreate Unit of Work-style atomicity across services is possible but notoriously operationally fragile: - it requires every participant to support the protocol; - it holds locks across a network round-trip; - a coordinator failure can leave participants blocked indefinitely. The pattern that actually fits this shape is the **Saga**: break the cross-service operation into a sequence of local transactions, each handled by an ordinary local Unit of Work, with each step publishing an event or triggering the next step, and an explicit compensating action for each step to undo it if a later step fails. Unit of Work still does useful work inside each individual service's local step; it's simply no longer the mechanism providing cross-service atomicity, because nothing at that scope can provide the same guarantee cheaply. ## Very high-throughput or latency-sensitive write paths The third is very high-throughput or latency-sensitive write paths, where the fixed overhead of identity-map bookkeeping and snapshot diffing is a cost the system can't afford at its target throughput or latency budget — for example, a hot event-ingestion pipeline writing tens of thousands of rows per second. These systems commonly bypass the ORM's Unit of Work layer entirely for the hot path and use direct prepared-statement batched inserts, accepting the loss of automatic dirty-checking convenience in exchange for predictable, minimal per-write overhead. ## CQRS-style architectures The fourth is CQRS-style architectures, which deliberately split the write model from the read model. - **The write side** typically still benefits from Unit of Work-style consistency, because enforcing an aggregate's invariants on write is exactly the bounded, single-transaction problem the pattern is built for. - **The read side**, however, is usually a denormalized projection built and refreshed by an entirely separate mechanism — often asynchronously, from the same domain events the write side publishes — and applying Unit of Work-style per-object tracking to that projection would be both unnecessary and actively wasteful. ## The unifying judgment call The unifying judgment call in all four cases is the same: Unit of Work earns its overhead when a bounded set of related objects needs correctness (all-or-nothing) and convenience within one local transaction. Whenever the operation's real shape is one of these, the pattern's assumptions no longer match the problem, and a principal engineer should reach for the tool built for that shape instead, respectively: | The operation's real shape | Reach for this instead | |---|---| | 'change many rows the same way' | set-based SQL | | 'coordinate across systems that can't share a transaction' | Sagas | | 'shave every possible microsecond off a hot write path' | direct batched writes | | 'rebuild a read-only projection from events' | event-driven projection rebuilding |

  • Why can't a Unit of Work simply be extended to coordinate a write across two microservices' separate databases?
    Unit of Work coordinates writes into one local ACID transaction on one connection; two independently deployed services with separate databases have no shared connection or transaction manager to coordinate through. Achieving atomicity across them requires either a distributed transaction protocol like two-phase commit - which is operationally fragile and rarely used in modern microservices - or restructuring the operation as a Saga of local transactions plus compensating actions.
  • If bulk set-based SQL updates a million rows directly, what does the application lose compared to doing it through the Unit of Work?
    It loses any per-row business logic encoded in the entity's setters, validation, or domain events that would normally fire during an ORM-tracked mutation - a raw UPDATE statement doesn't run application-layer code at all. That's an acceptable trade for pure data-shape changes, but risky if the update is supposed to also trigger domain-level side effects the ORM path would have handled.
  • In a CQRS system, does the read-side projection ever need anything like dirty checking?
    Generally no - the projection is typically rebuilt or upserted directly from a stream of events, so there's no need to diff an in-memory object against a snapshot to figure out what changed; the incoming event already says exactly what changed. Applying Unit of Work-style tracking there would add overhead for a guarantee the projection-update mechanism doesn't need.

Using Unit of Work for a million-row bulk update is like hand-addressing and stamping a million individual envelopes when a mail-merge bulk print job would do the same job in one pass - the per-item ceremony that's cheap at small scale becomes the whole cost at large scale.

saying these in an interview costs you the question

  • Believes Unit of Work can transparently span multiple databases/services
  • Suggests distributed two-phase commit as a routine, low-cost solution
  • Doesn't recognize bulk operations as a case where per-row tracking is wasteful
  • Can't name any alternative pattern (Saga, set-based SQL, event-driven projection)
  • Applies the same persistence strategy uniformly to both CQRS read and write sides

context