When should you choose Kafka request-reply (RPC-style) over plain pub/sub event choreography, and what are the architectural trade-offs?
answer
- events = default (loose, resilient, replay)
- request-reply = caller must wait for the answer
- RPC over Kafka re-adds temporal coupling + blocking
- at-least-once, not exactly-once RPC
- often HTTP/gRPC is the better synchronous path
basics
~20 sUse request-reply only when the caller truly needs the answer before it can proceed and you specifically want it to flow over Kafka. It re-adds temporal coupling, blocking, reply topics, correlation, and timeouts. Prefer pub/sub events when the caller can react later — it's looser, more resilient, and replayable. Often a synchronous HTTP/gRPC call is a better RPC than Kafka.
solid answer
~50 sRequest-reply suits cases where a caller must have a result to continue (e.g. a synchronous validation/lookup the UI is waiting on) and you want Kafka's properties — durability, partitioning, a shared transport, or an existing event backbone — rather than opening a new HTTP path. The cost: you re-introduce temporal coupling (both services must be up at once), caller blocking, plus the machinery of reply topics, correlation IDs, pending maps, and timeouts; and you get only at-least-once, not exactly-once, semantics. Pub/sub choreography is the default: the publisher fires an event and moves on, consumers react independently, you gain loose coupling, independent scaling, buffering, and replay, at the cost of eventual consistency and no direct return value. Critically, if you just need synchronous RPC, native HTTP/gRPC is usually simpler and lower-latency than Kafka request-reply — choose Kafka request-reply mainly when you specifically want the message log's durability/audit or to reuse the existing topic infrastructure.
go deeper
Know that events are fire-and-forget; request-reply makes the caller wait, so use it only when you need the answer right away.
List the trade-offs: temporal coupling, blocking, extra topics/correlation vs loose coupling and replay.
Give clear selection criteria and name concrete Kafka-specific costs (reply-topic routing, at-least-once, latency) and when HTTP/gRPC wins.
Make it a design-governance stance: events by default, request-reply as a justified exception; weigh audit/log/transport reuse against latency and coupling for the whole platform.
## Two communication styles **Pub/sub event choreography:** a service *publishes* a fact ('OrderPlaced') to a topic and forgets it. Any number of consumers subscribe and react on their own schedule. There is no reply and no waiting. This is Kafka's native, intended model. **Request-reply (RPC-style):** a caller *asks* and *waits* for an answer, simulating a function call across services (correlation ID + reply topic + pending map + timeout, as in `ReplyingKafkaTemplate`). ## When request-reply is the right call - The caller **cannot proceed** without the result *now* (a synchronous lookup, a validation a user is blocked on, a price quote). - You want the answer to ride the **existing Kafka backbone** for its durability, partitioned scaling, audit log, or because every service already speaks Kafka and you don't want a second network path / service-discovery story. - You need the **buffering / load-leveling** of a topic in front of a responder that can't take direct synchronous load spikes. - Scatter-gather where one request fans out to many responders and you aggregate replies (Spring's `AggregatingReplyingKafkaTemplate`). ## When pub/sub is better (the default) - The caller can **continue and react later** — true asynchronous workflows. - You want **loose coupling**: producer and consumer deploy, scale, and fail independently; the producer doesn't even know who consumes. - You value **resilience** (consumer down ≠ producer blocked), **buffering**, **replay** (re-read the log), and **fan-out** to new consumers without touching the producer. - Eventual consistency is acceptable. ## The trade-offs request-reply re-introduces 1. **Temporal coupling:** both services must be available simultaneously; the responder being down turns into caller timeouts — exactly the coupling pub/sub removes. 2. **Blocking / latency:** the caller waits, holding a future (and often a thread or request slot). End-to-end latency now includes broker round-trips on *two* topics — typically worse than a direct HTTP/gRPC call. 3. **Operational machinery:** reply topics to provision and route, correlation IDs, an in-memory pending map (leak/scaling concerns), and timeouts to tune. 4. **Weaker delivery semantics:** at-least-once messaging means duplicate/late replies and 'unknown outcome' timeouts; there is no exactly-once RPC. 5. **Scaling friction:** shared reply topics mis-route replies across caller instances unless you use unique reply topics/groups or `REPLY_PARTITION`. ## The often-overlooked alternative If you simply need synchronous request/response, **native HTTP/gRPC** is usually simpler, lower-latency, and has mature client tooling (timeouts, retries, load balancing). Reach for Kafka request-reply specifically when you want the **durable log / audit / shared-transport** benefits or to keep one communication substrate — not merely because Kafka is present. ## Rule of thumb Default to events; use request-reply as a deliberate exception where a blocking answer is genuinely required and Kafka's properties add value over a direct RPC.
- If you just need a synchronous answer, why not always use HTTP/gRPC instead of Kafka request-reply?Often you should — HTTP/gRPC is simpler and lower-latency for pure RPC. Choose Kafka request-reply when you specifically want the durable/auditable log, topic-based buffering/load-leveling in front of the responder, scatter-gather aggregation, or to keep a single shared transport so you don't add a new network path and service-discovery concern.
- Name a concrete downside of request-reply that pub/sub does not have.Temporal coupling: both services must be up at the same time, so a down responder becomes a caller timeout. Pub/sub lets the producer publish even when consumers are offline; they catch up from the log later.
saying these in an interview costs you the question
- Using request-reply as the default because it 'feels like a normal function call' — it inherits all of pub/sub's complexity plus blocking.
- Claiming Kafka request-reply gives stronger consistency than HTTP RPC (it's at-least-once and higher latency).
- Ignoring that a slow/down responder turns every caller into a timeout (temporal coupling).
- Reaching for Kafka request-reply when a direct HTTP/gRPC call would be simpler and faster.