You're designing the integration between an order service and five downstream consumers (inventory, shipping, analytics, fraud, and notifications), each needing different amounts of order detail. Walk through how you would decide, per consumer, between event notification and event-carried state transfer, and when you'd introduce event sourcing instead of either.
answer
- decide per event type, not system-wide
- fan-out load favors carried state
- freshness-critical favors callback
- size/sensitivity favors keeping data behind API
- event sourcing = log IS the source of truth, bigger commitment
basics
~20 sPick the style per consumer based on how much data they need, how current it must be, and how many consumers are asking - big or many-and-varied needs favor a fat event, few-and-current needs favor callbacks. Event sourcing is a different, bigger decision: making events the actual system of record, not just a way to tell other services something happened.
solid answer
~1 minThe decision is per-event-type and driven by three questions: how many consumers need this data (many consumers calling back to fetch details is expensive - favor carried state), how fresh must the data be at the moment of use (fraud checks need current truth - favor notification/callback), and how large or sensitive is the data (large or sensitive data favors keeping it in the producer and using notification, to avoid duplicating it everywhere). In this scenario, fraud likely needs a synchronous check or thin notification since a stale carried snapshot could approve a bad transaction; inventory and shipping benefit from a moderately fat event carrying line items and address since they act immediately and can tolerate a few seconds of staleness; analytics and notifications can consume whatever is published, even thin. Event sourcing is a separate, more invasive decision: it means the event log itself becomes the durable system of record that current state is derived from by replay, rather than events being a byproduct of writes to a separate database. You'd reach for it when you need full audit history, temporal queries, or the ability to rebuild projections from scratch - not merely because you're choosing how to shape an integration event.
go deeper
Should understand that the decision can differ per event/consumer rather than being one fixed rule, even without producing a full decision framework.
Should name at least two of the three deciding factors (fan-out, freshness, payload size/sensitivity) and apply them to a simple example.
Should walk through a multi-consumer scenario like this one, making a distinct, justified call per consumer, and correctly separate event sourcing as a different kind of decision.
Should discuss the organizational cost of schema ownership as fan-out grows, propose governance (schema versioning, per-consumer event splitting) and clearly frame event sourcing's operational commitments (snapshotting, permanent replay compatibility, CQRS) as the real gate on adopting it.
## A per-event-type decision, not a system-wide one Choosing between event notification and event-carried state transfer is not a one-time, system-wide architectural decision - it's a **per-event-type design choice**, and a senior engineer should be able to walk through the concrete factors that tip it one way or the other for a given consumer relationship, as well as recognize when the right answer is neither of these two integration styles but the structurally different pattern of event sourcing. ## The three factors that tip the decision 1. **Consumer fan-out.** The first factor is consumer fan-out. If a single consumer needs order details, a callback is cheap and simple - one extra request, easy to reason about. If five, fifty, or five thousand consumer instances all need the same details after a thin notification, that fan-out multiplies into significant load on the producer's API, and it's usually more efficient to publish the data once in the event than to serve it N times over synchronous calls. This is exactly the load-multiplication problem that pushes high-fan-out systems toward carried state. 2. **Freshness sensitivity.** The second factor is freshness sensitivity. Some decisions are only correct if made against current truth: a fraud check evaluating account risk, an inventory reservation checking real stock levels at the exact moment of purchase, or a payment authorization checking account balance. For these, either a direct synchronous call or an event-notification-plus-callback pattern is safer, because embedding a snapshot in an event risks the consumer acting on data that's already been superseded. Other decisions tolerate looking slightly behind reality: an analytics dashboard, a search index, a recommendation feature, or a customer-facing 'your order was placed' notification email don't need millisecond-fresh data, so a carried-state event with a few seconds or minutes of propagation lag is perfectly fine and in fact preferable, since it doesn't block on the producer. 3. **Payload size, shape stability, and sensitivity.** The third factor is payload size, shape stability, and sensitivity. A field that's small, stable, and needed broadly (order total, currency) is cheap to embed. A field that's large (full item catalog details, images), volatile (frequently changing pricing rules), or sensitive (payment instrument details, PII beyond what's needed) is expensive or risky to duplicate into every event and every consumer's storage, and it argues for keeping it behind an API and using notification-plus-callback, or projecting only the minimal safe subset into the carried-state payload. ## Applying the factors to the five consumers Applying these factors to the five consumers in the scenario: - **Fraud detection** weighs heavily toward correctness over convenience, so it should either make a direct synchronous call or treat any embedded balance/limit data as advisory only, re-verifying anything security-critical at decision time. - **Inventory and shipping** act on the order almost immediately after it's placed and need concrete details (SKUs, quantities, address) to do their job, and they can tolerate the same order being corrected a few seconds later via a follow-up event, so a moderately rich event-carried state payload suits them well and avoids hammering the order service with synchronous calls during peak traffic. - **Analytics and notifications** are the least freshness-sensitive and can consume whatever's published - thin or fat doesn't materially change their correctness, so the deciding factor there is usually 'don't make analytics's needs bloat the shared event schema'; a separate, wider analytics-specific event is often cleaner than growing the shared `OrderPlaced` payload to satisfy a low-priority consumer. ## When the answer is event sourcing instead Event sourcing is a categorically different decision and shouldn't be reached for just because you're picking an integration style. In event sourcing, the append-only event log is the actual system of record - current state isn't stored directly at all, it's derived by replaying events (or replaying from the last snapshot forward). This buys you: - a complete audit trail for free, - the ability to answer 'what was true at any past point in time', - and the ability to rebuild any downstream projection from scratch if its logic changes or its data gets corrupted. But it's a much bigger commitment: you need a snapshotting strategy for performance, careful event schema versioning since old events must remain replayable forever, and typically a CQRS split so queries don't have to replay history on every read. You'd choose event sourcing when the audit/temporal-query requirement is a first-class product or compliance need - financial ledgers, e-commerce order lifecycle tracking for disputes, inventory movement history - not simply because you're already publishing events for integration purposes. Conflating 'we publish domain events to other services' with 'our events are the source of truth' is a common and costly architectural mistake; most systems doing straightforward service-to-service integration only need notification or carried-state events, with a conventional database as the real system of record underneath.
- Why shouldn't a team adopt event sourcing just because they're already publishing OrderPlaced-style integration events?Publishing integration events is a lightweight decision about how services talk to each other; event sourcing is a commitment to making the event log itself the durable system of record for a whole aggregate, which requires snapshotting, permanent schema-replay compatibility, and usually CQRS. Doing it without a real audit/temporal-query need adds significant operational complexity for no corresponding benefit.
- If the fraud service uses a synchronous callback for freshness, what happens to it during an order-service outage, and how would you mitigate that?The fraud check would block or fail, which could either wrongly block legitimate orders or force a fallback decision; mitigations include a circuit breaker with a conservative default (e.g., hold for manual review rather than auto-approve), a short-lived cache with an explicit staleness bound, or a dedicated low-latency read replica the fraud service can hit that's decoupled from the order service's primary write path.
- How would you keep the shared OrderPlaced event schema from growing indefinitely as more consumers request more fields?Establish an owner for the schema who pushes back on low-value additions, split off a separate, richer event for niche high-detail consumers like analytics instead of fattening the shared one, and use schema versioning with a clear deprecation policy so fields can be removed once unused.
It's like deciding, per department in a company, whether to send a detailed memo to everyone (carried state) or a short 'come ask me' note (notification) - finance needs the live numbers so they ask directly, but the newsletter team is happy with whatever's in last week's memo.
saying these in an interview costs you the question
- Picks one style for the whole system instead of reasoning per event type
- Recommends event sourcing as the default choice for any event-driven integration
- Ignores freshness requirements when recommending carried state for a fraud/payment check
- Doesn't consider fan-out (number of consumers) as a factor in the decision
- Conflates 'we publish domain events' with 'the event log is our system of record'