skip to content

When a downstream service (say, an email/notification service) goes completely offline for 30 minutes, how does the blast radius differ between a request-driven design where the upstream service calls it synchronously versus an event-driven design where the upstream service publishes an event for it?

level: seniorimportance: should knowfreq 55%

answer

  1. failure absorbed at call site vs absorbed by broker durability
  2. cascading failure via thread/connection exhaustion
  3. circuit breaker/timeout mitigate sync
  4. backlog needs retention capacity
  5. time-sensitive events aren't 'harmless delay'

basics

~20 s

If the call is direct and synchronous, the upstream service can get stuck or start failing too while the other service is down. If it's an event, the upstream service is fine — the messages just pile up and get processed once the other service comes back.

solid answer

~50 s

In the synchronous design, every upstream call to the offline notification service either times out or errors immediately, and unless the upstream has a circuit breaker or fallback, that failure can propagate — user-facing requests fail, threads/connections pile up waiting on timeouts, and the outage of one low-priority service can degrade an unrelated core flow. In the event-driven design, the upstream service publishes to a broker and moves on regardless of the notification service's health; events simply accumulate in the queue/topic during the outage and are processed once the notification service recovers, so the blast radius is contained to 'notifications are delayed,' not 'checkout is degraded.' The catch is that the event-driven version isn't free of risk either — the broker needs capacity to buffer the backlog, and if the notification service has time-sensitive side effects (e.g., expiring verification links), delayed-but-eventually-successful isn't actually equivalent to on-time.

go deeper

for a junior

Should describe the basic outcome difference in plain terms: sync can drag the caller down, async mostly doesn't.

for a middle

Should name at least one concrete synchronous mitigation (timeout or circuit breaker) and recognize that the event-driven version delays rather than eliminates the notification.

for a senior

Should explain the cascading-failure mechanism concretely (pool exhaustion) and identify at least one real limitation of the event-driven version (retention capacity, time-sensitive payloads, backlog burst on recovery).

for a principal

Should reason about this as a capacity-planning and SLA problem — sizing broker retention against realistic outage durations, distinguishing which side effects tolerate delay versus which need a separate tighter guarantee — rather than treating 'event-driven' as an unconditional fix for downstream outages.

## Where the failure gets absorbed The difference comes down to where the failure of the downstream service is absorbed: in a **synchronous design**, it's absorbed (or not) by the calling code at the moment of the call; in an **event-driven design**, it's absorbed by the broker's durability, deferred entirely out of the calling code's critical path. ## The synchronous case, concretely Walk through the synchronous case concretely. Say checkout, after successfully processing an order, makes a direct HTTP call to a notification service to send the confirmation email, and waits for that call to return before finishing its own response to the customer. When the notification service is offline, that HTTP call fails to connect or hangs until a timeout. Without protective measures, several bad things can happen at once: - Checkout's own request now takes as long as the timeout instead of returning instantly, degrading the customer's experience even though nothing about their order actually failed. - If checkout doesn't set a timeout at all, its own request thread or connection can hang indefinitely, and under load, enough hung threads/connections exhaust the pool, meaning checkout itself starts failing for unrelated customers whose orders have nothing to do with email. This is the textbook **cascading failure**: a genuinely low-priority downstream dependency ends up threatening a genuinely critical path purely because of how the two were wired together. ## The mitigations the synchronous world already has The standard synchronous-world mitigations exist precisely to contain this: - A short **timeout** bounds how long checkout waits. - A **circuit breaker** stops calling notification altogether once failures cross a threshold, failing fast instead of piling up hung calls. - And, often, teams simply decide that a synchronous call was the wrong choice for a non-critical side effect and should be moved off the critical path entirely — effectively **reinventing an event**. ## The event-driven case Now the event-driven case. Checkout, after its own transaction, publishes an `OrderConfirmationRequested` event to a broker and returns success to the customer without knowing or caring whether the notification service is up. When notification service is offline, nothing about checkout is affected at all — not latency, not thread/connection pools — because checkout's interaction with the broker (a local, fast, durable write) has already completed. The events for those 30 minutes simply accumulate in the topic/queue. When the notification service comes back online, its consumer resumes processing from wherever it left off and works through the backlog, and every customer eventually gets their confirmation email, just late. The **blast radius** of the outage is contained to "notifications are delayed," a much smaller and more tolerable failure than "checkout intermittently fails or slows down for unrelated customers." ## The risk the event-driven version still carries But the event-driven version still has real risk, not none. - **First, durability capacity.** The broker has to actually be able to buffer 30 minutes of backlog without hitting storage limits or retention expiry — if a topic's retention is shorter than the outage, events get dropped, silently converting "delayed" into "lost." - **Second, time-sensitive side effects** break the "eventually successful is fine" assumption: if the event in question is "send a password-reset link" or "send a payment-verification code" rather than "send a receipt," a 30-minute delay may render that message functionally useless, so the outage isn't harmless just because it's decoupled — it's a different kind of harm (a failed user journey) instead of a cascading one. - **Third, the backlog burst.** When the consumer recovers, it usually has to process a large backlog quickly, which can itself cause a secondary problem — a burst of many emails sent in the same minute can trip the email provider's own rate limits or look like spam behavior, so "catching up" isn't automatically graceful either. ## Why notification sending lives behind a queue This is close to the textbook justification for why most production systems put email/SMS/push-notification sending behind a queue rather than a direct synchronous call from the service that triggers it — checkout, signup, and password-reset flows all publish an event and let a dedicated, independently-scaled notification service consume it, precisely so that provider outages or rate-limit issues on the notification side never threaten the availability of core account or purchase flows. The trade they accept in return is exactly the caveat above: they must separately guard against time-sensitive tokens expiring by keeping those specific flows on a tighter, monitored SLA even though the transport is still async.

  • What's a circuit breaker and how does it change the synchronous scenario?
    A circuit breaker tracks recent failure rates for a downstream call and, once failures cross a threshold, 'opens' and fails fast locally without even attempting the call for a cooldown period. This bounds checkout's exposure to a slow/dead notification service to a brief window of slow failures rather than every single request hanging for a full timeout during the entire outage.
  • If retention on the event topic is only 15 minutes and the outage lasts 30, what actually happens to the events published in the first 15 minutes?
    Depending on the broker's retention policy, those events can be deleted before the consumer ever gets to them, silently converting a 'delayed' delivery into a permanently lost one — this is why retention/capacity planning has to account for realistic worst-case downstream outage duration, not just typical processing time.
  • Why might 'fire and forget without even publishing to a durable broker' (e.g., a raw async HTTP call with no retry) be worse than either of the two designs discussed?
    A raw fire-and-forget call with no durable intermediary gives you neither the immediate consistency of a synchronous call nor the durability guarantee of an event broker — if the notification service is down at the exact instant of the call, the message is simply lost with no backlog to replay from once it recovers, combining the downsides of both approaches with the safety of neither.

Synchronous is like a cashier who won't ring up your groceries until they've personally handed a receipt to another department that's currently on a coffee break — the whole line backs up. Event-driven is like dropping receipts in a mailbox for that department to process whenever they're back — your checkout line keeps moving.

saying these in an interview costs you the question

  • Assumes event-driven designs have zero risk from a downstream outage
  • Doesn't mention timeout/circuit-breaker as the standard mitigation for the synchronous case
  • Ignores that broker retention/capacity has limits and can silently drop backlog
  • Treats every downstream side effect as equally tolerant of delay (misses time-sensitive cases like expiring tokens)
  • Can't explain the actual mechanism of cascading failure (thread/connection pool exhaustion) beyond 'it breaks stuff'

context