skip to content

What's the difference between the request/reply and publish/subscribe messaging patterns for service-to-service communication, and what coupling trade-off does each impose?

level: middleimportance: must knowfreq 80%

answer

  1. addressed vs broadcast
  2. correlation ID for async replies
  3. fan-out without publisher changes
  4. invisible vs visible coupling

basics

~20 s

Request/reply is one service asking another a direct question and getting a direct answer, like calling a specific person. Publish/subscribe is one service announcing something happened, and any number of other services that care can listen in, like posting an announcement on a notice board.

solid answer

~50 s

Request/reply is a point-to-point pattern: a caller sends a request addressed to a specific service and expects a specific reply back, whether that exchange happens synchronously over HTTP/gRPC or asynchronously via a reply queue and correlation ID. The caller knows exactly who it's talking to, which makes reasoning and debugging straightforward but couples the caller to that specific service's identity, availability, and contract. Publish/subscribe is broadcast: a publisher emits an event to a topic without knowing or caring who, if anyone, is listening, and any number of subscribers can independently react. This decouples the publisher from having to know its consumers at all, letting new consumers be added later with zero changes to the publisher, but it also means the publisher gets no direct answer and has to give up control over how or whether consumers use the data, which makes tracing 'who reacted to this and how' much harder.

go deeper

for a junior

Should correctly distinguish 'talking to one specific service' from 'broadcasting to anyone listening' with a simple example of each.

for a middle

Should explain the fan-out benefit of pub/sub (new subscribers without publisher changes) and the debuggability cost, and know that request/reply can be implemented asynchronously via correlation IDs.

for a senior

Should articulate the schema/contract coupling that persists even in pub/sub despite the identity decoupling, and describe concrete production debugging differences between the two failure modes.

for a principal

Should discuss organizational implications: how pub/sub changes ownership and change-management (who reviews an event schema change, how consumers are discovered), and when the invisible-coupling cost of pub/sub outweighs its decoupling benefit at scale.

## Two questions a message answers Request/reply and publish/subscribe answer two different questions about a service interaction: **"who is this message for?"** and **"does the sender expect something specific back?"** In **request/reply**, the sender addresses a specific service instance or logical service by name or endpoint, and expects a specific answer shaped by a known contract: - an HTTP `GET` to an orders service returns that exact order; - an RPC to a pricing service returns that exact quote. This holds whether the transport is synchronous, like a blocking HTTP call, or asynchronous, like a message sent to a reply queue tagged with a correlation ID so the response can be matched up later even though the caller didn't block waiting for it. What defines the pattern is not the transport but the addressing: one sender, one intended recipient, one expected reply shape. ## How publish/subscribe flips the addressing Publish/subscribe flips the addressing model entirely. A publisher emits an event, such as `OrderPlaced` or `InventoryAdjusted`, onto a named topic or channel and has no idea, and no need to know, which services are subscribed to it. Zero subscribers, one subscriber, or twenty subscribers can all receive the same event without the publisher's code changing at all. This exists specifically to solve the problem of **fan-out**: in a request/reply world, if the order service needs to notify inventory, shipping, loyalty, and analytics that an order happened, it would have to know about and call all four explicitly, and adding a fifth interested service would require changing the order service's code. Pub/sub inverts that dependency: new subscribers opt in on their own, and the publisher's code and deploy cadence never have to change to accommodate them. ## The coupling trade-off The coupling trade-off is the central thing to understand here. - **Request/reply creates tight, visible coupling.** The caller explicitly depends on the callee's identity, its uptime, and its response contract, and if that contract changes incompatibly, the break is usually immediate and traceable to a specific call site. That visibility is actually valuable for debugging: a failed request/reply call shows up as a clear error with a clear culprit. - **Publish/subscribe creates loose, invisible coupling.** The publisher doesn't depend on any specific subscriber's existence or health, but the system as a whole becomes harder to reason about, because you can no longer look at the publisher's code and know everyone who consumes its output. Two teams can each independently subscribe to the same event and build very different assumptions about what fields mean or how reliably it fires, and neither the publisher nor the other subscriber necessarily knows the other exists. ## How each one breaks In production: | Pattern | What a break looks like | |---|---| | **Request/reply** | Failures are loud and localized: a 500 or a timeout on a specific call, attributable to a specific dependency, is easy to alert on and easy to trace with a distributed tracing tool. | | **Pub/sub** | Failures are diffuse: if a subscriber has a bug that silently drops half the events it receives, there's no error at the publisher, no error at any other subscriber, and the first sign of trouble might be a downstream report weeks later that "loyalty points have been wrong since March." | This is why teams that lean heavily on pub/sub invest in event schemas with versioning discipline, consumer-side dead-letter handling, and often a way to audit which services actually consume which topics, since that knowledge otherwise lives only in each subscriber's own code. ## Where each one shows up A concrete real-world pattern: many e-commerce systems use request/reply for anything the checkout flow needs a direct, synchronous or near-synchronous answer to, such as "is this payment method valid," while using publish/subscribe for everything that happens as a side effect of a completed order, such as: - updating a recommendation engine; - sending a receipt email; - adjusting a fraud-detection model's training data. The order service publishes one `OrderPlaced` event and never has to know that a sixth subscriber, say a new customer-loyalty microservice, was added eight months later; that's the coupling win pub/sub buys. The cost is that if the loyalty service has a bug, nobody automatically finds out from the order service's perspective, because the order service's job was done the moment the event was published.

  • How can a request/reply exchange still be asynchronous, if 'reply' implies getting an answer back?
    By using a correlation ID: the caller sends a request message tagged with a unique ID and a reply-to address, then continues without blocking; when the callee finishes, it sends the response tagged with the same correlation ID to that reply address, and the original caller matches the incoming reply to its outstanding request by that ID. This gets the decoupling benefits of async transport while preserving a specific one-to-one question-and-answer contract.
  • If a team adds a fifth subscriber to an existing 'OrderPlaced' topic, what's the biggest operational risk they're introducing that the publisher won't see coming?
    Load and behavior on the publishing side is usually fine since publishing cost doesn't scale with subscriber count in most broker designs, but the new subscriber could start making assumptions about event fields, ordering, or delivery frequency that the publisher never explicitly promised or tested against, so a future 'harmless' change to the event schema can silently break that subscriber with no visibility for the publishing team.
  • How would you debug a pub/sub-based system where a downstream side effect (like a confirmation email) sometimes doesn't happen?
    Check the publisher's logs to confirm the event was actually emitted, check the broker's delivery and dead-letter metrics for that topic and consumer group to see if delivery or processing failed, and check the specific subscriber's logs and idempotency/retry handling, since a common cause is the consumer throwing on a malformed or unexpected event shape and the message ending up in a dead-letter queue nobody's monitoring.

Request/reply is calling a specific coworker and waiting for their answer; publish/subscribe is posting an announcement on a shared bulletin board where anyone who's interested can read it, and you never find out who did or didn't.

saying these in an interview costs you the question

  • Thinks publish/subscribe means the publisher waits for all subscribers to finish
  • Can't explain how an async request/reply exchange matches a response back to its original request
  • Assumes pub/sub has no coupling at all, ignoring the implicit schema/contract coupling
  • Believes request/reply must always be synchronous
  • Can't name a concrete reason to prefer fan-out over calling each interested service directly

context