skip to content

In Domain-Driven Design, for a domain with a hard, non-negotiable invariant — such as a bank account balance that must never go negative under any concurrent transaction — how do you weigh keeping that invariant inside one aggregate against the throughput cost, and what alternatives exist if a single-aggregate design becomes a bottleneck?

level: principalimportance: nice to knowfreq 35%

answer

  1. hard invariant can't be eventually consistent
  2. per-account single-writer queue not global lock
  3. reservation/hold shrinks critical section
  4. DB CHECK constraint as backstop
  5. throughput ceiling is per-hot-aggregate, not whole system

basics

~20 s

When a rule can never be broken, even briefly, keep its data in one aggregate so one lock guards it. If that's too slow, speed up or queue that path — don't loosen the guarantee.

solid answer

~50 s

A true, hard invariant like 'balance never goes negative' forces the account balance and every operation that changes it into one consistency boundary, because DDD's invariant guarantee only holds within a single aggregate's atomic transaction — there's no clean way to enforce 'never negative' as an eventually-consistent rule without a window where it's temporarily violated. Accept the throughput cost of single-aggregate locking as the price of correctness, then attack the bottleneck without weakening the guarantee: batch/queue commands against the hot aggregate so contention becomes sequential processing instead of failed retries, use database-level constraints (a CHECK constraint or a computed balance) as a second line of defense, or, at genuine scale, move to patterns like reservation/hold (pre-authorize an amount, settle asynchronously) that shrink the critical section rather than removing it. Avoid silently loosening the invariant to relieve contention — for money that's a correctness regression, not a performance fix.

go deeper

for a junior

May know that money-related invariants are important but likely hasn't reasoned about the concurrency mechanics that make them fail under load.

for a middle

Understands that a hot aggregate causes lock contention but may default to 'just split it' without checking whether the invariant is genuinely separable.

for a senior

Can propose queuing or reservation-pattern mitigations and correctly distinguishes throughput problems from correctness problems.

for a principal

Makes and defends the call to keep a hard invariant's aggregate intact under load, names concrete scaling levers (per-account queue, reservation pattern, storage-layer backstop), and can push back on a team proposing a correctness-weakening 'fix.'

## When an honest boundary becomes the bottleneck At the principal level, the interesting version of the aggregate-sizing question isn't 'how do I apply the true-invariant heuristic' — it's **'what do I do when applying that heuristic honestly produces an aggregate that becomes a measurable bottleneck under real load, and the invariant genuinely cannot be relaxed.'** A bank account balance that must never go negative, even momentarily, even under concurrent withdrawal attempts, is the textbook case: any withdrawal must check the current balance and apply the debit as one atomic, isolated operation against every other concurrent withdrawal on the same account, or two withdrawals of $60 each against a $100 balance can both individually appear valid and jointly overdraw the account by $20. That requirement — **correctness that holds even under adversarial concurrency, not just 'usually true'** — is exactly what an aggregate's transaction boundary plus optimistic or pessimistic locking exists to provide, and it cannot be delivered by an eventually-consistent, cross-aggregate design, because 'eventually consistent' inherently accepts a window of temporary violation, and a temporarily-negative balance is not an acceptable temporary state for most financial systems (it can trigger real overdraft fees, real fraud, real regulatory exposure). ## Refusing the tempting fix So the first principal-level judgment call is refusing the tempting-but-wrong 'fix': splitting the account into smaller pieces or moving balance-checking to an asynchronous event handler to relieve contention, because that **changes the correctness guarantee, not just the performance profile**. The account aggregate (balance plus enough transaction history/state to validate an operation) stays as one aggregate, and the throughput problem is treated as exactly that — a throughput problem to be solved without touching the guarantee. ## The levers that stay available With the guarantee fixed, several real levers remain. 1. **The first is queuing/serializing writes against a hot aggregate** instead of letting concurrent requests race and retry: rather than N concurrent threads all attempting optimistic-locked writes and most failing and retrying (wasting work and adding latency variance), route all commands for a given account through a **single-writer queue** (per-account, not global) so operations against that one account process strictly sequentially, while different accounts still process fully in parallel — this converts lock-contention retries into queue wait time, which is usually more predictable and efficient at scale. 2. **The second lever is shrinking the critical section itself**: a **reservation/hold pattern** (as used by real payment processors) doesn't debit the full amount synchronously — it places a hold that decrements 'available balance' atomically and cheaply, then settles the actual transfer asynchronously, so the strictly-synchronous, contention-prone operation is reduced to the smallest possible check-and-decrement rather than a full transaction with side effects. 3. **The third lever is defense in depth at the storage layer**: a database-level `CHECK` constraint (`balance >= 0`) or a computed/derived balance from an **append-only ledger** of debits/credits (rather than a mutable balance field) can catch violations even if application-level locking has a bug, trading some write cost for a hard backstop that doesn't depend on the application getting concurrency control right every time. ## The ceiling worth naming out loud The trade-off to name explicitly in a design review is: **throughput on a single hot aggregate is fundamentally capped by how fast one sequence of operations can be validated and committed**, no matter how much you scale out the rest of the system — this is a real ceiling, not a tuning problem that disappears with more hardware, because the correctness requirement is inherently sequential for that one account. - What you can scale is the number of independent accounts processed in parallel (which is usually enormous and dwarfs any single account's transaction rate in practice), and the efficiency of the sequential path per account (batching, cheaper critical sections, faster storage). - Recognizing which of those two dials you're actually turning — and not confusing 'the whole system's throughput is fine' with 'this one popular account isn't a bottleneck' — is exactly the kind of distinction a principal engineer is expected to make when a team below them proposes 'let's just split the account aggregate to fix the slowness,' because that fix quietly breaks the guarantee the whole design existed to provide. ## How card networks already do this A concrete, real-world pattern illustrating this is how card payment networks and many banking cores handle **authorization holds**: the debit card authorization step is a small, fast, strictly serialized decrement of available balance (the hard invariant, protected tightly), while final settlement, merchant reconciliation, and statement generation happen asynchronously over the following hours or days — the hard invariant is kept small and fast, and everything that can tolerate delay is deliberately pushed out of the synchronous, contention-prone path.

  • A team suggests splitting one hot Account aggregate into per-currency or per-region sub-accounts to relieve write contention. When is that a legitimate fix versus a correctness regression?
    It's legitimate if the true invariant is genuinely scoped per sub-account (e.g., each currency wallet has its own independent non-negative-balance rule with no rule spanning currencies), because then each sub-account really is its own consistency boundary. It's a regression if the business actually requires a single combined balance or credit limit across those sub-accounts, since splitting would let two individually-valid sub-account operations jointly violate a rule that used to be enforced atomically.
  • How does a per-account single-writer queue differ from just retrying failed optimistic-lock writes more aggressively?
    A single-writer queue processes every command for a given account strictly one at a time by construction, so there's never a failed write to retry in the first place — contention becomes ordered wait time instead of wasted, discarded work. Aggressive retrying, by contrast, still lets concurrent writers race and fail against the version check, burning CPU and adding latency variance under load without changing the underlying serialization requirement.
  • Why isn't a database CHECK constraint like balance >= 0 sufficient on its own, without proper aggregate-level concurrency control?
    A CHECK constraint validates the final row value at commit time within one transaction, but it can't prevent two concurrent transactions from each independently computing a valid-looking debit against a stale read and both committing if isolation levels or locking aren't correctly enforced — it's a last-line backstop against bugs, not a substitute for correct concurrency control in the application/aggregate layer.

Like a single bank teller who must handle every transaction on one specific vault in strict order to make sure it's never overdrawn — you can hire more tellers for other vaults, but you can't parallelize that one vault's line without risking two withdrawals slipping through at once; instead you make each transaction at that teller faster.

saying these in an interview costs you the question

  • proposes relaxing a hard invariant to eventual consistency purely to fix a performance problem
  • conflates 'the whole system is slow' with 'this one hot aggregate is contended'
  • assumes splitting an aggregate always improves throughput without checking whether the invariant is truly separable
  • doesn't distinguish a per-account queue/lock from a global lock
  • treats a database CHECK constraint as sufficient replacement for correct aggregate-level concurrency control

context