skip to content

When an application loads an existing aggregate back from a database row or an event stream, why is this 'reconstitution' path usually kept separate from the factory method used to create a brand-new aggregate, even though both end up producing the same aggregate type?

level: middleimportance: must knowfreq 50%

answer

  1. create validates 'is this new instance legal'
  2. reconstitute trusts already-validated persisted state
  3. event sourcing: replay events vs handle first command
  4. ORMs bypass constructors via reflection - a real risk

basics

~20 s

Making something new has to check business rules like 'is this order allowed to exist yet,' but loading something that already existed and was already checked once shouldn't re-run those creation-time checks - it should just rebuild the object from stored data.

solid answer

~40 s

Creation-time factories enforce invariants that only make sense for something new: an order can't start already shipped, an account can't start with a negative balance from thin air. Reconstitution is rebuilding an aggregate whose state already passed those checks in the past and may now legitimately be in a state a 'create' factory would reject - a shipped order, for instance. If reconstitution reused the create-time factory, it would either wrongly reject valid historical states or the factory would need business-rule bypass flags, weakening the guarantee for real new creation. So teams add a separate reconstitution path (a 'rehydrate' factory method, an ORM mapper, or event-sourcing replay) that trusts the persisted state and skips create-only validation, while still assembling the aggregate through its intended internal structure rather than reflection hacks.

go deeper

for a junior

Should grasp that loading old data shouldn't re-check 'can this be born' rules; doesn't need ORM reflection details.

for a middle

Should describe a separate reconstitution path/method and give a concrete example like a shipped order failing create-time validation.

for a senior

Should discuss event-sourced replay vs command handling as a distinct mechanism and the performance/trust trade-off of skipping validation on load.

for a principal

Should reason about failure modes like reflection-based ORM hydration bypassing all invariant checks, and how to defend against corrupted persisted state despite that.

## What a creation factory actually asks A creation factory answers the question **'is it valid for a brand-new instance of this aggregate to come into existence with these values right now?'** That question only makes sense at the moment of birth. An `Order` factory might refuse to create an order with status 'shipped,' because a shipped order implies a sequence of prior events - placed, paid, packed, shipped - that a synthetic freshly-created object skipped entirely. But when the application restarts, or a request comes in for order #4821, the system needs to load an order that is, quite legitimately, already in the 'shipped' state, because it went through that whole lifecycle over the preceding weeks. If the only path back into an `Order` object were the same `Order.create(...)` factory used for new orders, that factory would either: - reject a perfectly valid stored order (because its creation-time invariants forbid a starting status of 'shipped') - or need special-case flags to bypass those checks - at which point every caller of 'create' has to reason about which flags are safe, quietly eroding the very guarantee the factory exists to provide ## The second path in The mechanism, concretely: most systems add a second path into the aggregate, often called **reconstitution**, **rehydration**, or (in ORMs) simply **'hydration.'** This path takes data already known to represent a once-valid, and presumably still-valid, aggregate state - a row fetched from the aggregate's own table(s), or the output of replaying a stream of domain events for an event-sourced aggregate - and assembles the aggregate object directly, without re-running the create-time business validation. It typically still goes through the aggregate's own controlled construction surface rather than a fully public constructor: - a package-private constructor - a static `reconstitute(...)` method - an ORM's ability to set private fields via reflection/bytecode manipulation The goal isn't 'anything goes' - it's 'trust the source of this data, but still assemble it through the type's real internal shape' so you don't end up with, say, a value object that skips its own construction-time normalization. ## Event sourcing makes the distinction sharpest Event-sourced aggregates make the distinction sharpest: reconstitution there isn't a database row mapping at all, it's replaying every domain event ever recorded for that aggregate's id, in order, applying each one to progressively rebuild current state (an `apply(OrderPlaced)`, `apply(OrderPaid)`, `apply(OrderShipped)` sequence). This is fundamentally not the same operation as creation, because creation-time factories in an event-sourced system typically just decide whether to emit the very first event (e.g., `OrderPlaced`) - they never see the full lifecycle at once, whereas reconstitution sees and replays the entire history to arrive at the current snapshot. ## The trade-off The trade-off of keeping these paths separate is a bit more code: - two distinct entry points into the aggregate to understand and maintain instead of one - a discipline requirement that any new construction-time invariant added to the 'create' factory be deliberately considered for whether reconstitution needs an equivalent (usually not, since reconstitution trusts already-validated data, but occasionally a defensive check is warranted if the store might have been corrupted or manually edited) The benefit is that 'create' stays strict and trustworthy - it's always safe to assume anything produced by the create factory is a legitimately new, rule-compliant instance - while reconstitution stays efficient and doesn't waste cycles re-validating rules on every single object load, which matters a lot at scale (loading an aggregate happens far more often than creating one). ## Failure modes Failure modes when this separation is missing or done carelessly are common in practice. 1. **Performance.** Teams that accidentally run full create-time validation (including things like external uniqueness checks against other aggregates) on every load slow down every read path unnecessarily. 2. **Data corruption tolerance.** Another failure mode: if reconstitution is implemented via raw reflection that bypasses the aggregate's constructor entirely (common with some ORMs, e.g., Hibernate's default no-arg constructor plus field injection), it's easy to end up with an aggregate whose invariants were never actually checked by anything, because both the create-time factory and the reconstitution path got skipped - a corrupted or manually-edited database row silently produces an aggregate object the domain model assumes is always valid. A well-known real-world pattern for exactly this is event-sourcing frameworks like Axon or EventStoreDB, which draw a hard architectural line between a command-handler-driven creation path that decides whether to emit a first event, and a separate event-replay/snapshot-loading path used purely to rebuild in-memory state - the two are never the same code path, by design.

  • What goes wrong if reconstitution is implemented by calling the same 'create' factory used for brand-new aggregates?
    The create factory's invariants are written for the moment of birth and will often legitimately reject valid historical states, like a shipped order being loaded with status 'shipped.' Teams then add bypass flags to the create factory to accommodate reloading, which weakens the guarantee for genuinely new creation and makes the factory's contract ambiguous.
  • In an event-sourced aggregate, what does reconstitution actually do differently from handling a creation command?
    Reconstitution replays the aggregate's entire recorded event history, in order, applying each event to rebuild current state, whereas a creation command handler only ever decides whether to emit the very first event. They operate on fundamentally different inputs - one on a full history, the other on a single incoming command.

Creating a new employee record goes through HR onboarding checks (background check, signed offer letter); pulling up an existing employee's file from the archive doesn't re-run the background check every time someone opens the folder - it trusts that the check already happened when they were hired.

saying these in an interview costs you the question

  • Thinks reconstitution and creation should always share one code path
  • Doesn't know ORMs commonly bypass constructors via reflection
  • Can't explain why re-running create-time validation on every load is wasteful
  • Confuses reconstitution with a database migration or backup/restore operation

context