You're deciding whether a 'user updated their shipping address' change should trigger a synchronous API call to the shipping service or publish an AddressUpdated event instead. What's the actual trade-off you're making, beyond 'events are more scalable'?
answer
- three axes: availability/temporal/consistency
- confirmation gap
- outbox + idempotent consumer
- sync check before irreversible action
basics
~20 sIf you call directly, you know right away it worked, but you depend on that service being up. If you publish an event, you don't depend on it being up right now, but you don't know immediately whether it processed the change, and it might be out of sync for a bit.
solid answer
~40 sThe trade is between an immediate consistency/feedback guarantee and independence between the two services. A synchronous call gives the caller a definite yes/no right now, but couples the caller's success and latency to the callee's availability — if shipping is down or slow, the address update fails or stalls too. Publishing AddressUpdated decouples them: the address update succeeds regardless of shipping's health, and shipping catches up whenever it can — but now there's a window where shipping doesn't yet know the new address, and the calling code gets no direct confirmation that shipping processed it, so you need a separate mechanism (reconciliation, a status flag, a callback event) if that confirmation actually matters downstream.
go deeper
Should recognize there's a real trade-off (not just 'events are better') and describe it in plain terms: sync = immediate answer + dependency, async = no dependency + delayed answer.
Should name eventual consistency explicitly and give at least one concrete risk if the destination system needs a synchronous point of truth.
Should propose concrete mitigations (outbox, idempotency, reconciliation) and identify the one step in the flow, if any, that genuinely needs to stay synchronous.
Should reason about this per-step rather than per-service — deciding which specific transitions in a larger flow can tolerate staleness and which cannot — and weigh the operational cost of running mitigation machinery against the coupling cost of staying synchronous.
## Why "it scales better" is not the whole story The headline reason engineers reach for events — "it scales better" — is real but incomplete, and treating it as the whole story leads to picking async in places where the missing piece (confirmation) actually matters. The trade-off that matters day to day sits along three coupled dimensions: - **Availability coupling** - **Temporal coupling** - **Feedback/consistency guarantees** Events move you in a specific direction on all three simultaneously, not a free win on one axis. ## Availability coupling A synchronous call to the shipping service makes the address-update operation's success conditional on shipping being reachable and responsive at that exact moment. If shipping is deploying, degraded, or network-partitioned, the address update either fails outright or the user's request hangs until a timeout fires. Publishing `AddressUpdated` removes that dependency: the address change is durably recorded and the user gets a fast, successful response regardless of shipping's health. This is the **resilience half of the trade** — one service's outage no longer takes down an unrelated user-facing flow. ## Temporal coupling and the confirmation gap Temporal coupling and the confirmation gap is the part "more scalable" glosses over. | Style | What the caller knows | |---|---| | Synchronous call | The caller learns synchronously whether shipping accepted the new address — success or failure is known before the response returns to the user. | | Published event | The caller has no idea, in that request, whether shipping ever processed the change; it only knows the event was durably queued. | If nothing downstream cares whether shipping specifically got it, this gap is irrelevant. But if correctness depends on shipping being current — e.g., a package must never ship to a stale address — that gap is a real production risk, and "the event was published" is **not** the same guarantee as "shipping applied it." ## The window where the two services disagree Between the moment the event is published and the moment the shipping service's consumer processes it, the two services disagree about the customer's address. That window is usually milliseconds to seconds under healthy conditions, but it is unbounded under a consumer outage or backlog. If a warehouse worker packs and ships during that window, the package goes to the old address, and no exception was ever thrown anywhere; the system behaved "correctly" by its own internal logic while producing a wrong real-world outcome. This is the **sharp edge of eventual consistency**: failures don't look like errors, they look like correct-looking systems quietly disagreeing with each other. ## The compensating machinery teams add Because the trade-off is manageable, not avoidable, teams that pick events for a case like this typically add compensating machinery: 1. A **transactional outbox**, so the event is published exactly alongside the DB write (avoiding "update saved but event lost" from a crash between two separate writes). 2. **Idempotent consumers**, so at-least-once redelivery doesn't double-apply an update. 3. A **reconciliation job**, or a stored `lastConfirmedAddress` timestamp shipping can check before an irreversible action like dispatch. 4. And, if the calling flow genuinely needs synchronous confirmation, a request-driven "is this address current?" check right before the irreversible step, even though propagation itself stays event-driven. Mature event-driven designs often aren't "purely async everywhere" — they use events for propagation and a narrow synchronous check at the one point where staleness is actually dangerous. ## When to stay synchronous instead Not every case should reach for events. - **The user needs to know right now.** If the operation is user-facing and the user needs to know right now whether it worked (e.g., "was my new card accepted"), synchronous request/response is correct because the UX requires an immediate answer, and the coupling cost is acceptable since such services are usually built for high availability anyway. - **Two pieces of state must change together.** Similarly, if correctness requires that two pieces of state change together or not at all (e.g., debit one account and credit another), a synchronous, transactional interaction — or a carefully designed saga — beats a bare fire-and-forget event, because "eventually consistent" isn't good enough when a customer is staring at their balance. ## What a real platform does with this address A real e-commerce platform typically does both for exactly this address case: the address-change endpoint publishes `AddressUpdated` for cheap, resilient propagation to shipping, analytics, and email preview services, but the actual "create shipping label" step — which happens later, at dispatch time — makes a synchronous read of the current address from the source of truth (the user service) rather than trusting whatever stale copy shipping's own event-driven projection holds. The fast, resilient path (events) handles routine propagation, and the one irreversible, correctness-critical step (label creation) uses a synchronous read to eliminate the staleness window exactly where it would cause real harm.
- What's the transactional outbox pattern and why does it matter here?It's writing the event to an outbox table in the same database transaction as the business data change, then having a separate relay process publish it to the broker. It matters because without it, you either publish-then-save (risking a published event for a change that then fails to save) or save-then-publish (risking a saved change whose event is lost if the process crashes in between), both of which silently desynchronize the two services.
- How would you detect that the shipping service's view of an address has gone stale due to consumer lag?Monitor consumer lag on the AddressUpdated topic/queue directly (most brokers expose this natively), and separately track an 'age since last successful event processed' metric per consumer. Alerting on lag crossing a threshold catches the problem before it causes a wrong shipment, rather than discovering it from a customer complaint.
- Why isn't 'just retry the synchronous call' a full substitute for the resilience events give you?Retrying a synchronous call still blocks the caller for the duration of the retries, or requires the caller to build its own async retry queue, effectively reinventing an event system, and it doesn't help if the callee is down for an extended period — the caller either keeps failing or has to give up and tell the user something went wrong, whereas an event just waits durably until the consumer recovers.
It's the difference between texting someone 'heads up, I moved' (they'll see it eventually, you don't wait) and calling them to confirm they've updated your address in their system before you let them mail something irreplaceable.
saying these in an interview costs you the question
- Says events are chosen purely for throughput/scalability with no mention of the confirmation/consistency trade-off
- Assumes 'the event was published' means 'the downstream service applied the change'
- Proposes fire-and-forget events for a case that requires strict two-sided consistency (e.g., money transfer) without a saga or synchronous safeguard
- Doesn't mention idempotency when discussing at-least-once delivery
- Treats the choice as all-or-nothing for the whole flow instead of per-step