skip to content

Why are Domain Services in DDD expected to be stateless, and what specifically breaks in production if a team adds mutable instance fields to one?

level: middleimportance: should knowfreq 55%

answer

  1. no mutable instance fields
  2. singleton + concurrency = race
  3. params/collaborators, not fields
  4. trivial to test, safe to scale
  5. corruption is sporadic, hard to repro

basics

~20 s

A Domain Service shouldn't remember anything between calls - each call gets everything it needs as input. If it stores data in its own fields instead, two concurrent calls can mix up or overwrite each other's data.

solid answer

~50 s

A Domain Service is meant to be a pure operation over the domain objects it's given, not an object with its own lifecycle or memory - all the data it needs should come in through method parameters or through read-only, injected collaborators like a repository interface, never through mutable fields the instance keeps between calls. This matters practically because domain services are almost always wired up as long-lived singletons in dependency-injection containers and shared across concurrent requests; if one instance holds mutable state, concurrent calls race on it, corrupting results for one or both callers, and the bug is nondeterministic and hard to reproduce. Statelessness also makes the service trivially safe to unit test (no setup/teardown of internal state between test cases), safe to scale horizontally (any instance can serve any call), and easy to reason about since correctness depends only on its inputs.

go deeper

for a junior

Should know the rule 'domain services don't hold mutable state' and recognize a mutable field on one as suspicious when pointed out.

for a middle

Should be able to explain the singleton-plus-concurrency mechanism that turns a mutable field into a race condition, not just cite the rule.

for a senior

Should recognize subtler variants (mutable fields on injected collaborators, accidental state via captured closures) and know how to fix them without breaking the service's API.

for a principal

Should set or review team conventions/lint rules that catch this class of bug before code review, and reason about its implications for horizontal scaling and multi-tenancy.

## What statelessness means here Statelessness for a Domain Service means the object holds **no mutable data between one method call and the next** - every fact the operation needs must arrive: - as a method parameter - or through an injected collaborator (like a repository) that is itself only used to read or persist data, not to accumulate state on the service instance This is a deliberate constraint, not an accident of style: it is what makes a domain service behave like a **pure function over domain objects** rather than like an entity with its own lifecycle. Contrast this with an entity, which is expected to hold state (an `Order` legitimately remembers its own line items and status) - a domain service is explicitly the opposite kind of object, and conflating the two by giving a service its own evolving fields undermines the reason it was introduced in the first place. ## Why it matters in a real deployment The practical reason this matters comes from how domain services are actually deployed in real applications. In almost every modern backend, dependency-injection containers (Spring, .NET's built-in DI, and equivalents) wire services up as **singletons by default** - one instance of the class serves every request for the lifetime of the application. If that single instance has a mutable field, say a `cachedExchangeRate` or a `runningTotal` set during one call, then every subsequent concurrent call sees and can overwrite that same field. Under load, two requests running on different threads can genuinely interleave: 1. request A sets the field 2. request B overwrites it before A reads it back 3. and A completes using B's data This produces a **race condition**, and race conditions are notoriously hard to catch in code review or in simple tests because they: - only manifest under concurrent load - disappear when you add a debugger breakpoint (which serializes execution) - and often show up in production as sporadic, unreproducible data corruption rather than a clean crash ## What the constraint costs and buys The trade-off is mostly one-sided in production code: statelessness costs a small amount of parameter-passing verbosity (you sometimes have to thread a value through a call rather than stash it on the object once), but buys **thread safety for free**, since a stateless object has nothing to corrupt no matter how many threads call it concurrently. - It also buys **simpler testing**: a stateless service test needs no setup/teardown of internal fields between test cases and can reuse a single instance across the whole suite without contaminating results; a stateful one risks test-order-dependent failures, where a test passes in isolation but fails when run after another test left the service in an unexpected internal state. - And it buys **straightforward horizontal scaling**: since no server-local memory holds business state, any instance behind a load balancer can serve any request, which matters once an application runs on more than one process or machine. ## Failure modes The failure modes when a team violates this are recognizable. 1. **The most direct is the race condition already described** - intermittent wrong answers or crossed data under concurrent load, often first noticed in production under peak traffic rather than in development. 2. **A subtler one is accidental cross-request memory retention**: a mutable field holding a reference to, say, a large object from one request can keep it alive longer than intended, contributing to memory bloat or, worse, letting data from one tenant or user leak into a response for a different one - a serious problem in multi-tenant systems. 3. **Another is 'my tests are flaky' bug reports** that turn out to be caused by shared mutable state on a domain service instance reused across test cases, wasting engineering time chasing a phantom test-infrastructure bug when the real defect is the service's own design. ## A concrete illustration A concrete illustration: imagine a `PricingService` with a method `calculateFinalPrice(Order order, Discount discount)` that, for a misguided 'optimization,' caches the last discount it computed in a field named `lastDiscount` so a subsequent debug log can print it. Under moderate concurrent traffic, two customers checking out within milliseconds of each other can end up with each other's discount applied to their order - a bug that would sail through single-threaded manual testing and appear only once concurrent load exists, exactly the kind of defect that's expensive to diagnose after the fact because the cause (a stray mutable field) is far from where the symptom (wrong price charged) shows up.

  • Is it ever safe for a domain service to have instance fields at all?
    Yes, if those fields are immutable and set once at construction time, such as an injected repository or another collaborator - the key distinction is immutability and stability, not whether the object literally has zero fields. Anything set or mutated inside a method call, and expected to persist to the next call, is what's unsafe.
  • How would caching fit into a stateless domain service if computing a value is expensive?
    Caching belongs in a dedicated caching layer or infrastructure component with its own explicit concurrency-safe design (e.g. a thread-safe cache with defined eviction), injected into the service as a collaborator - not as ad hoc mutable fields on the service itself. That keeps the caching concern separate from and doesn't compromise the service's own statelessness guarantee.
  • How would you catch this bug before it reaches production?
    Static analysis or a lint rule flagging mutable non-final fields on classes wired as singletons catches many cases automatically. Concurrency-focused tests that call the same service instance from multiple threads simultaneously and assert on correctness under load are the more reliable but heavier-weight way to catch it.

A stateless domain service is like a calculator that's wiped clean after every use - hand it fresh numbers each time and it always gives the right answer, no matter how many people are using identical calculators at once. A stateful one is like a shared whiteboard where two people scribbling numbers at the same time end up computing from each other's half-erased work.

saying these in an interview costs you the question

  • Adds a mutable field to a domain service 'just to remember the last result'
  • Doesn't know whether their DI container wires the service as a singleton or per-request
  • Dismisses concurrency concerns because 'it worked when I tested it manually'
  • Uses instance fields for caching instead of an explicit, thread-safe cache collaborator
  • Can't explain why a race condition in a service class would be intermittent rather than consistent

context