skip to content

In a system of microservices, when should two services communicate through synchronous request/response calls versus asynchronous events on a broker, and what changes about failure handling in each case?

level: middleimportance: must knowfreq 74%

answer

  1. temporal coupling is the deciding question
  2. command = directed, needs answer; event = fact, past tense
  3. sync: timeouts, backoff+jitter, circuit breaker, bulkhead, fallback
  4. async: at-least-once, idempotency, DLQ, lag, per-key ordering
  5. dual write -> transactional outbox; multi-service flow -> saga

basics

~20 s

Use a synchronous call when the caller needs an answer right now to continue its work. Use an asynchronous event when you are announcing that something happened and other services can react later, without the caller waiting.

solid answer

~50 s

Decide by temporal coupling. Synchronous request/response (HTTP, gRPC) suits queries and commands where the caller cannot proceed without the result: read a price, validate a token, reserve stock before confirming. The callee must be up at that instant, so failure handling means timeouts, bounded retries with backoff and jitter, circuit breakers, bulkheads and a defined degraded response. Asynchronous events suit notifications of facts that already happened, fan-out to unknown consumers, burst absorption and long-running work: OrderPlaced consumed by shipping, billing and analytics. The broker buffers, so the producer survives a dead consumer, but you inherit eventual consistency, at-least-once delivery (consumers must be idempotent), ordering only within a partition or key, poison messages needing dead-letter queues, and harder end-to-end tracing. Two rules keep the hybrid honest: never dual-write to database and broker without a transactional outbox, and never emulate a synchronous call by publishing a message and blocking on a reply.

code

pseudocode · 17 lines
pseudocode
// Dual write - unsafe: a crash between the two lines loses the event
db.save(order)
broker.publish(OrderPlaced(order.id))

// Transactional outbox - one local transaction, relay publishes later
transaction {
    db.save(order)
    db.save(OutboxRow(topic = "orders", payload = OrderPlaced(order.id)))
}
// separate relay: read unsent outbox rows -> publish -> mark sent (at-least-once)

// Idempotent consumer
onMessage(msg) {
    if (processed.contains(msg.id)) return   // duplicate delivery is expected
    applyEffect(msg)
    processed.add(msg.id)
}

go deeper

for a junior

Say synchronous means the caller waits for an answer and asynchronous means fire-and-forget through a broker, and give one example of each.

for a middle

Frame it as temporal coupling, distinguish commands from events, and name the concrete failure-handling tools on each side: timeouts, retries, circuit breakers versus idempotency, dead-letter queues, consumer lag.

for a senior

Add consistency consequences - eventual consistency, sagas, orchestration versus choreography - and the transactional outbox for the dual-write problem; discuss availability multiplication in synchronous chains.

for a principal

Discuss it as a system-wide policy: default to asynchronous across bounded contexts and synchronous within, schema governance and compatibility rules for events, replay and dead-letter operations, and the observability investment (correlation ids, lag dashboards) that asynchronous integration obliges you to fund.

## The two interaction styles **Synchronous request/response**: service A opens a connection to service B, sends a request and blocks (logically) until a response or a timeout. Typical transports: HTTP/REST, gRPC. **Asynchronous messaging**: service A hands a message to a **broker** (Kafka, RabbitMQ, SQS, Pulsar). The broker stores it; consumers read it when they are ready. A publishes and moves on, and may not know who consumes. The deciding property is **temporal coupling**: does the caller need the other party alive *at this instant*? Synchronous calls create temporal coupling; messaging removes it by inserting durable storage between the parties. ## Commands versus events - A **command** asks a specific service to do something (`ReserveStock`). It is directed, may be rejected, and often needs an answer. Commands are usually synchronous, though they can be sent as messages to a single-consumer queue. - An **event** states a fact that already happened (`OrderPlaced`). It is past tense, has no intended recipient, and cannot be rejected. Events are naturally asynchronous. If you find yourself asking 'and what did the consumer reply?', you have a command dressed up as an event. ## Choose synchronous when - The caller literally cannot continue: authorisation checks, price lookups, availability checks before a user-visible confirmation. - The user is waiting and expects a fresh answer inside one request. - The interaction is a query with no state change. **Failure handling then means**: connection and read **timeouts** shorter than the caller's own budget; **bounded retries** only for idempotent operations, with exponential backoff plus jitter; **circuit breakers** to stop hammering a failing dependency; **bulkheads** so one slow dependency does not consume the whole thread or connection pool; and a documented **fallback** (cached value, degraded feature, explicit error). The pathology to avoid is a deep synchronous chain: with five hops at 99.9% availability each, best-case composite availability is about 99.5%, and latencies add. ## Choose asynchronous when - One fact must fan out to several consumers, including ones added later. - Traffic is bursty and the consumer should absorb it at its own rate. - The work is slow (reports, media encoding) or must survive consumer downtime. - You want to decouple deployment: consumers can be down for a rolling upgrade without failing the producer. **Failure handling then means**: **at-least-once delivery** is the default, so consumers must be **idempotent** (dedupe by message id, or make the operation naturally repeatable); **ordering** is guaranteed only per partition or key, so design for out-of-order arrival; **poison messages** need retry limits and a **dead-letter queue** plus an alert and a replay path; **consumer lag** must be monitored as a first-class metric; and **schema evolution** must be backward compatible because old events stay in the log. ## The dual-write trap A service that commits a row and then publishes an event performs two writes to two systems with no shared transaction. A crash between them loses the event (or, if reversed, publishes a fact that never became true). Distributed two-phase commit across a database and a broker is rarely available or advisable. The standard fix is the **transactional outbox**: write the event into an `outbox` table inside the same local transaction as the state change, then a separate relay process (poller or change-data-capture on the transaction log) publishes it and marks it sent. Delivery becomes at-least-once, which is exactly what idempotent consumers already handle. ## Consistency and workflow Asynchronous integration means **eventual consistency**: for a while, one service knows something another does not. A multi-service business transaction is then modelled as a **saga** - a sequence of local transactions with compensating actions - either **orchestrated** (one coordinator drives the steps, easy to see, central coupling) or **choreographed** (each service reacts to events, loosely coupled, harder to trace). Pick orchestration when the process is complex and needs visibility; choreography when the steps are few and independent. ## Anti-patterns - **Request/reply over a broker while blocking a user thread** - you kept temporal coupling and added a broker's latency and failure modes. - **Events as remote getters** ('send me your data') - that is a query; use a synchronous call or replicate the data. - **Broker as a shared database** - consumers reaching into another service's internal state via fine-grained events, which recreates coupling on the payload. - **No dead-letter path** - one bad message stalls a partition forever.

  • A service writes to its database and then publishes an event. What can go wrong and how do you fix it?
    It is a dual write across two systems with no shared transaction: a crash after the commit but before the publish silently loses the event, and reversing the order can announce a fact that was never persisted. The fix is the transactional outbox - store the event in an outbox table in the same local transaction, then relay it to the broker with a poller or change-data-capture. Delivery becomes at-least-once, so consumers must be idempotent.
  • Why is a chain of five synchronous service calls per user request considered a design smell?
    Latencies add and availabilities multiply, so five 99.9% hops give roughly 99.5% composite availability, and any one slow hop stalls the whole request while holding threads and connections. It also implies the boundaries are wrong - work that always happens together probably belongs in one service, or later steps should be triggered asynchronously after an early response.
  • Your consumers can receive the same event twice. What are the options?
    Make the operation naturally idempotent (set a value rather than increment), or deduplicate by a stable message or business key stored transactionally with the effect, or use conditional writes and unique constraints so the second attempt is a no-op. Exactly-once end to end is generally not available across systems; effectively-once is achieved by at-least-once delivery plus idempotent processing.

A synchronous call is a phone call: both parties must be available at the same moment, and you get an answer immediately. An event is posting a notice on a board: you continue with your day, several people may read it later, and nobody replies to you.

saying these in an interview costs you the question

  • Saying 'events are always better than REST' with no reference to temporal coupling or consistency needs
  • Publishing a message and blocking on a reply, then calling the design decoupled
  • Assuming brokers provide exactly-once delivery, so consumers need no idempotency
  • Writing to the database and the broker separately with no outbox or change-data-capture
  • Assuming global ordering of messages instead of per-partition or per-key ordering
  • Having no dead-letter queue, retry limit or replay procedure for poison messages

context