When Service A needs data that Service B owns (for example, a product's name and price displayed inside an order), what are the two main strategies for getting that data, and what does each trade off?
answer
- sync API call vs local denormalized copy
- events/CDC feed the copy
- availability coupling vs staleness window
- snapshot != cache
- distributed monolith risk from over-calling
basics
~20 sEither Service A asks Service B live every time it needs the data (simple but creates a runtime dependency), or Service A keeps its own local copy that's kept up to date via events (fast and independent, but the copy can be briefly out of date).
solid answer
~50 sThe two strategies are synchronous lookup (call Service B's API on demand) and data duplication (Service A stores a local, denormalized copy of the fields it needs, kept in sync via events or CDC from Service B). Synchronous lookup keeps a single source of truth and avoids staleness, but couples Service A's availability and latency to Service B — if B is slow or down, A degrades too, and a fan-out of many such calls creates a distributed-monolith. Duplication removes that runtime coupling and lets A serve requests even if B is unreachable, but introduces eventual consistency: A's copy can be stale for some window, and A must handle B's schema evolving via versioned events. The right choice depends on how fresh the data must be and how much A's availability should depend on B's.
go deeper
Can describe the two options in plain terms (ask live vs keep a copy) and give one pro/con of each.
Explains events/CDC as the mechanism for keeping a duplicate in sync, and names staleness vs availability-coupling as the core trade-off.
Distinguishes snapshot semantics from caching, discusses idempotency/ordering/versioning for event-driven duplication, and picks the right strategy per field based on business freshness requirements.
Sets org-wide guidance on when duplication is acceptable, designs the event contract/versioning strategy across many consumers, and reasons about the compounding cost of duplication sprawl across a large service graph.
## The two mechanisms Once each service owns its own database and cross-service SQL joins are off the table, a consuming service has two concrete mechanisms for getting data it doesn't own. 1. **The first is synchronous composition**: on each request, Service A calls Service B's API (typically over HTTP or gRPC), gets the current value, and uses it immediately without storing it. 2. **The second is duplication via asynchronous replication**: Service B publishes domain events whenever its relevant data changes (something like `ProductPriceChanged`), Service A subscribes to that stream and applies each event to its own local table, and from then on A reads its own copy for every request without contacting B at all. The same effect can be achieved with **change-data-capture (CDC)** tooling that reads B's database transaction log directly and turns row changes into a stream, without B's application code having to publish events explicitly. ## Why the choice exists This choice exists because, having already split the databases for independence, a consuming service still needs related data to satisfy its own requests, and it has to get that data from somewhere. - **Calling live every time** keeps a single source of truth and avoids ever serving stale data, but it means A's request can only succeed as fast, and as reliably, as B responds — A has taken on a runtime dependency on B for every request that needs that field. - **Keeping a local copy** removes that runtime dependency, so A keeps working even during a B outage, but it accepts that the copy can lag behind the true current value for some window of time, and it now owns the job of keeping that copy correct. ## What each one costs Both costs are real and specific. - **Synchronous lookup's cost is availability and latency coupling**: if B is slow, A gets slow; if B is down, A either fails the same request or has to build a fallback (cached last-known value, default, circuit breaker), and if A calls B once per item in a list, that N+1 pattern of remote calls compounds latency and failure probability badly at scale. - **Duplication's cost is everything that comes with eventual consistency**: the local copy needs storage and its own operational care, the events feeding it need to be processed in the right order (or carry a version so out-of-order delivery doesn't regress the copy to a stale value), the consumer needs to handle the owning service's event schema changing over time, and there's always some propagation delay between B's true state changing and A's copy catching up. ## The snapshot that is mistaken for a cache A specific and easy-to-hit failure mode with duplication is treating the local copy as a cache rather than being clear about what kind of copy it actually is. | Kind of copy | How it must behave | |---|---| | **A cache** | is supposed to track the current source-of-truth value and should be invalidated or refreshed whenever the source changes — staleness in a cache is a bug you try to minimize | | **A point-in-time snapshot** | is what some duplicated data is meant to be: it should never be refreshed to match the source, because an order's line items should keep the product's name and price as they were at the moment of purchase, not whatever the catalog says today | Confusing the two — 'refreshing' a snapshot as though it were a cache — silently rewrites history and produces bugs like a customer's past receipt changing price after the fact. Another common failure is **dual writes**: instead of deriving the local copy strictly from B's events, some part of A's code writes to the local copy directly, say to patch a bug quickly, which causes the copy to silently diverge from B's actual state the next time a real event arrives and overwrites the patch, or vice versa. ## Both strategies in one system A concrete example that shows both strategies coexisting sensibly: an e-commerce Order service duplicates a snapshot of each product's name and price at the moment an order is placed, because historical accuracy is a business requirement — orders must show what the customer actually paid, permanently, even after prices change. But the same Order service calls the Inventory service synchronously, live, right before confirming a purchase, because staleness there has a real cost: showing 'in stock' from a five-minute-old local copy risks confirming an order for an item that's actually sold out. The decision isn't generic — it follows from what each specific field needs: is staleness cheap or expensive, and does this data represent a fact frozen at a point in time or a live current value?
- If a service consumes events from another service to build its local copy, what's the risk of processing those events out of order, and how is it usually mitigated?Out-of-order processing can make the local copy regress to an older state — for example, applying an 'address changed to X' event after a later 'address changed to Y' event overwrites the newer value with the stale one. It's typically mitigated by including a version number or timestamp in each event and having the consumer discard any event older than what it already has, or by using a per-entity ordered stream, like a partitioned topic keyed by entity ID, so ordering is guaranteed per record.
- Why is duplicating a 'snapshot' of data, like the product price at order time, different from just caching another service's data?A cache is meant to mirror the current source-of-truth value and should be invalidated or refreshed when it changes, so staleness is a bug to minimize. A snapshot is intentionally a point-in-time copy that should NOT update when the source changes — the order needs the price as it was at purchase, not the current catalog price. Conflating the two leads to bugs, like a snapshot being 'refreshed' and silently rewriting history.
- When would synchronous lookup be clearly the better choice over duplicating data locally?When staleness carries real business or safety cost and the data changes frequently relative to how often it's read — for example, checking real-time inventory availability before confirming a sale, where showing a stale 'in stock' could cause overselling. In these cases the availability coupling to the owning service is an acceptable trade for correctness, sometimes combined with fallbacks or circuit breakers to limit blast radius.
Like choosing between calling a friend every time you need their phone number, always current, but you're stuck if they don't pick up, versus writing it down in your own address book, always available, but it goes stale if they change numbers and don't tell you.
saying these in an interview costs you the question
- Treats a duplicated snapshot as a cache that should always reflect the current source value
- Proposes writing directly to the local copy instead of deriving it from the owning service's events/API
- Ignores the staleness window entirely when picking duplication
- Suggests synchronous calls have no downside beyond a bit of latency
- Can't name a mitigation for out-of-order or duplicate event delivery