skip to content

When you decompose a monolith's single database into per-service data stores, how do you decide which service owns a given piece of data, and when is it appropriate to replicate that data into another service versus having that service call the owner synchronously on every read?

level: seniorimportance: should knowfreq 60%

answer

  1. one owner per data element, writes go there only
  2. replicate for read-heavy or availability-critical data
  3. call-through for low-frequency, must-be-fresh reads
  4. events drive replicated copies
  5. eventual consistency is the price of decoupling

basics

~20 s

Give each piece of data one owning service that's the source of truth for writes. Others call the owner for occasional fresh reads, or keep a replicated copy, updated via events, for frequent reads or to stay available if the owner is down.

solid answer

~50 s

Ownership follows the business capability: whichever service is responsible for creating and changing a piece of data owns it and is the single source of truth for writes; every other service treats it as read-only and gets it either by calling the owner's API or by keeping a locally replicated copy. The decision between synchronous call-through and local replication comes down to read frequency, latency needs, coupling tolerance, and availability requirements: call-through keeps one copy of the truth and is simple, but couples the caller's availability and latency to the owner's; replication via domain events (e.g., an owner publishing CustomerAddressChanged) decouples availability and latency at the cost of eventual consistency and the operational burden of a second copy that can drift or go stale. High-read, latency-sensitive, cross-boundary data, like a product's name and price shown on every order line, is usually replicated; low-frequency, must-be-current data, like checking current account balance before a large transfer, usually calls through.

go deeper

for a junior

Should understand that each piece of data should have one clear owner and that other services shouldn't just reach into that owner's database.

for a middle

Should describe both call-through and event-based replication as options and give one situation where each fits better.

for a senior

Should reason explicitly about the trade-off between coupling and simplicity versus latency, availability, and staleness, and design the event contract for a replicated read model, including idempotency.

for a principal

Should set organization-wide data-ownership conventions, such as who may publish which events and how consumers version their local schemas, and adjudicate disputes over shared or contested entities across teams.

## The ownership rule When a monolith's single shared database is broken apart during microservices decomposition, the first governing rule is that **every piece of data must have exactly one owning service** — the service responsible for the business capability that creates and changes that data becomes its single source of truth for writes, and every other service that needs it is a read-only consumer. This mirrors capability-based service boundaries: whichever service embodies the business function that legitimately mutates a piece of data, such as an Inventory service adjusting stock counts or a Catalog service renaming a product, is the only one allowed to write it, full stop — no other service should have direct write access to another's data store, whether through a shared table, a shared schema, or a back-door script. ## Two ways a non-owner reaches the data The harder decision is how a non-owning service gets access to data it needs but doesn't own, and this is a genuine trade-off with two live options rather than one obviously correct answer. - **Synchronous call-through.** The first option, synchronous call-through, has the consuming service call the owner's API each time it needs the data. This keeps exactly one copy of the truth in existence, avoids any staleness, and is conceptually simple — but it couples the caller's own latency and availability to the owner's: if the owner is slow or down, every caller that needs its data synchronously is slow or down too, and at high call volume this can turn a rarely-used internal API into a load-bearing dependency for many services at once. - **Replication via events.** The second option, replication via events, has the owning service publish domain events whenever the relevant data changes, for example a `ProductRenamed` or `PriceChanged` event, and consumers maintain their own local, purpose-built copy of just the fields they need, updated asynchronously as events arrive. This decouples the consumer's availability and latency entirely from the owner's — it can keep serving requests using its last-known copy even if the owner is offline — at the cost of eventual consistency: there is always some window, from milliseconds to minutes depending on the event pipeline, where the local copy can be stale relative to the source of truth. ## How to choose between them The decision between the two comes down to concretely naming, for a given piece of data, how often it's read, how fresh it must be, and how tolerant the caller is of the owner being temporarily unavailable. | The data profile | The fit | |---|---| | Data that's read very frequently, doesn't need to be instantaneously fresh, and would create unacceptable load or fragility if fetched synchronously every time — like a product's display name shown on every line of every order | a strong candidate for **replication** | | Data where staleness has real financial or correctness consequences and reads are comparatively infrequent — like an account balance checked immediately before authorizing a large transfer | a strong candidate for **synchronous call-through** to the authoritative source, accepting the coupling as the necessary cost of correctness | ## The cost of getting it wrong The cost of getting this wrong shows up differently depending on which mistake you make. - **Over-relying on synchronous call-through** for high-volume reads produces exactly the cascading-failure and availability-coupling symptoms of a distributed monolith: the 'owner' service becomes a single point of failure and a latency bottleneck for the whole system. - **Over-relying on replication** without the operational discipline to support it produces silent data drift: an event gets dropped, processed out of order, or a consumer's schema falls out of sync with what the owner now publishes, and nobody notices until a customer sees an obviously wrong price or a stale order status. Production teams typically guard against this with: - **idempotent event consumers**, safe to process the same event twice; - a **reconciliation job** that periodically diffs the replica against the source of truth and alerts or self-heals on mismatches; - and the ability to **replay an event log** to rebuild a consumer's local copy from scratch if drift is ever detected. ## Where it shows up A concrete, common instance of this pattern is an e-commerce Order service maintaining a lightweight, denormalized local copy of a product's name and image URL populated from Catalog's `ProductUpdated` events, so that rendering order history never depends on Catalog being up or fast — while the same Order service calls a Payments service synchronously, in real time, to authorize a charge, because approving a payment against stale account or fraud-risk data is not an acceptable trade to make for decoupling.

  • What's the risk of replicating data into multiple services via events instead of always calling the owning service?
    The replicated copies become eventually consistent, not immediately consistent — there's a window where a service's local copy is stale relative to the owner, which is fine for a product name on an order history page but dangerous for something like available inventory count at checkout. You also take on the operational cost of a second data store that can drift permanently if an event is ever missed or processed out of order, so you need idempotent consumers and some form of reconciliation or replay.
  • How would you handle a case where two services seem to need to jointly own the same entity, like both Sales and Fulfillment needing to update an 'Order' record?
    Split the entity conceptually rather than sharing one record: each service owns the fields and lifecycle stages it's actually responsible for, so Sales owns order-placed and pricing fields while Fulfillment owns shipment and delivery-status fields, and they're linked by a shared order ID, with each stage transition published as an event the other listens to. Trying to have two services write to the same row invites race conditions and violates the single-writer principle that keeps ownership unambiguous.
  • If a replicated copy of data drifts out of sync with its source of truth, how do you detect and recover from that in production?
    You typically add a reconciliation job that periodically compares checksums or key fields between the owner and the replica and flags or auto-corrects mismatches, plus dead-letter handling and alerting on the event pipeline so a dropped or failed event doesn't silently go unnoticed. Some teams also support event replay from an event log or outbox table so a consumer can rebuild its local copy from scratch if drift is detected.

Like a company's HR department owning the official personnel file while every other department, Payroll, IT, Facilities, keeps a small, purpose-built extract of just the fields it needs, such as salary grade or badge access level, that HR pushes updates to — nobody outside HR edits the master file, and each department chooses whether to keep its own copy or just call HR when it needs an answer, based on how often it asks and how badly it needs an instant answer.

saying these in an interview costs you the question

  • Suggests two services should both have write access to the same table or entity to 'stay in sync'
  • Treats every cross-service data need as requiring a synchronous call, ignoring replication as an option
  • Treats every cross-service data need as requiring full replication, ignoring that synchronous call-through is simpler and sometimes correct
  • Doesn't mention eventual consistency or staleness as a cost of replicating data

context