You're designing an API for an operation that typically completes in under 300ms but occasionally, for large inputs, takes 10+ seconds. A colleague proposes using the Asynchronous Request-Reply pattern, HTTP 202 plus polling, for every call to keep the API uniform. What are the arguments against reflexively applying this pattern here, and what alternatives would you weigh instead?
answer
- adds a mandatory extra round trip for the common fast case
- hybrid: sync response if fast, 202 fallback if slow
- alternatives: long polling, SSE, WebSockets for push
- extra infra cost: job store, workers, cleanup
- not a substitute for fire-and-forget eventing
basics
~20 sFor work that's usually fast, forcing every client through an accept-then-poll dance adds unnecessary round trips and complexity for the common case. Better options: keep it synchronous with a generous timeout for the fast path, or only switch to async for inputs that are actually large or slow.
solid answer
~50 sApplying 202-plus-polling uniformly punishes the vast majority of requests that finish in 300ms by forcing at least two round trips, submit then poll, and client-side complexity like a state machine and backoff, for no benefit. A hybrid design keeps the endpoint synchronous up to a bounded wait, returning the result directly if it's ready in time, and only falls back to 202 for the rare slow case. Other tools for the same underlying goal of decoupling client wait time from processing time include long polling, Server-Sent Events, or WebSockets for push-based updates, and true fire-and-forget messaging when the caller genuinely doesn't need a reply tied to its specific request. Also weigh the added operational cost: a job store, cleanup jobs, a worker fleet, and idempotency handling are all overhead a purely synchronous endpoint doesn't need.
go deeper
Can say that forcing a fast operation through polling feels like unnecessary extra steps.
Can propose returning the result directly for fast calls and only going async for slow ones as a rough idea.
Articulates the hybrid bounded-wait design concretely and names concrete alternatives, such as long polling, SSE, or WebSockets, with their fit.
Weighs the pattern against the full alternative portfolio and the added operational and infrastructure cost at the system level, and can justify a hybrid or purely tiered design, fast sync plus slow async, as an organizational default.
## The cost it solves and the cost it adds The Asynchronous Request-Reply pattern is a solution to a specific cost: an operation is slow enough, and unpredictably enough, that holding an HTTP connection open for it risks timeouts and wastes resources. Applying it reflexively to every endpoint 'for consistency' ignores that this solution has its own cost, and that cost is paid on every single call, including the overwhelming majority that never needed it in the first place. ## What forcing the pattern costs the fast path For an operation that is fast almost all the time, forcing 202-plus-polling means every caller now pays for at least two round trips, one to submit and at least one to poll, instead of one. It also pushes real complexity onto every client: - a state machine for `pending/running/complete/failed` - a backoff strategy - handling for jobs that get lost or stuck None of this buys anything for the 99% of calls that would have finished well within a normal request timeout anyway; it only pays off for the rare slow outlier. ## The hybrid design A much better fit for this specific shape, common input usually small and fast with an occasional large and slow outlier, is a **hybrid design**: 1. The server starts processing synchronously and holds the response open up to a bounded threshold, say 2 seconds. 2. If the work finishes within that window it returns 200 with the result directly, a single round trip exactly like a normal synchronous endpoint. 3. If the threshold is exceeded, the server switches mid-flight to returning 202 with a status URL, and the client falls back to polling only for the outlier case. This keeps the common path cheap while still degrading gracefully instead of timing out on the rare slow input. ## Other tools worth weighing Beyond the hybrid approach, there are other tools worth weighing depending on what the caller actually needs. - **Long polling and Server-Sent Events** let the server hold a connection open and push an update the moment it's ready, giving lower latency than fixed-interval polling without the client needing to run an inbound-reachable server the way a webhook does. - **WebSockets** go further, offering a persistent bidirectional channel, useful when a client needs many updates over time rather than a single eventual result. None of these substitute for genuine fire-and-forget event publishing, though, which serves a fundamentally different need: a publisher broadcasting a fact to any number of unknown subscribers with no expectation of, or channel back to, a reply correlated to its own request. Async request-reply exists specifically because the original caller does need to learn the outcome of their own specific request eventually, which requires a correlation id and a defined mechanism, poll or callback, to retrieve that particular outcome; that requirement is exactly what distinguishes this pattern from plain event publishing, and it's why swapping in fire-and-forget messaging is not a like-for-like substitute even when it looks superficially similar. ## The infrastructure cost Finally, adopting async request-reply is not free at the infrastructure level even before considering client complexity. It requires: - a durable job or status store that survives worker restarts - monitoring for stuck or orphaned jobs - a cleanup or TTL process so the store doesn't grow unbounded - typically a separate worker fleet or queue consumer decoupled from the request-accepting tier A purely synchronous endpoint needs none of this, since its entire 'state' is the in-flight HTTP call itself, which disappears the moment the response is sent. A concrete illustration of getting this trade-off right is how most REST APIs handle CRUD operations synchronously by default and reserve the async pattern for specific, genuinely long-running operations, such as GitHub returning results synchronously for the vast majority of its API surface but switching to 202-plus-polling specifically for operations like repository migrations or large data exports, which are known in advance to be slow rather than occasionally slow.
- How would a hybrid design let most calls stay synchronous while still handling the occasional slow one?The server starts processing synchronously and holds the response open up to a bounded threshold, say 2 seconds; if the work finishes within that window, it returns 200 with the result directly. If the threshold is exceeded, it switches to returning 202 with a status URL for the client to poll the rest of the way. This keeps the fast path a single round trip while still degrading gracefully for slow outliers.
- Why isn't fire-and-forget event publishing a substitute for async request-reply here?Fire-and-forget means the publisher has no expectation of, or channel back to, a specific reply correlated to its own request — it suits broadcasting facts to any number of unknown subscribers. Async request-reply exists specifically because the original caller does need to learn the outcome of their own request eventually, which requires a correlation id and a defined way, poll or callback, to retrieve that specific outcome.
- What ongoing operational cost does adopting async request-reply add compared to a purely synchronous endpoint, even before considering client complexity?You now need a durable job or status store, monitoring for stuck or orphaned jobs, a cleanup or TTL process, and typically a separate worker fleet or queue consumer, none of which a synchronous request/response endpoint requires, since its state is just the in-flight HTTP call itself.
It's like making every customer at a fast-food counter take a buzzer and wait at a table even when their order is ready in 20 seconds — reasonable for the occasional big catering order, wasteful for a regular burger.
saying these in an interview costs you the question
- proposes making every endpoint async 'for consistency' without weighing the fast-path cost
- can't describe a hybrid sync-with-async-fallback design
- conflates this pattern with fire-and-forget eventing
- ignores the operational cost of a job store, cleanup, or worker fleet
- assumes polling is free for the client