skip to content

questions

5

A client calls a REST endpoint that kicks off a report that takes 2 minutes to generate. Instead of holding the connection open, the server responds immediately with HTTP 202 Accepted and a URL the client can check later. What is this approach called, and why use it instead of just blocking the connection until the report is ready?

level: juniorimportance: must knowfreq 70%

answer

  1. 202 Accepted + Location header
  2. poll a status resource
  3. decouple accept from complete
  4. job/task/correlation id
  5. worker processes in background

basics

~20 s

The server says 'got it, working on it' right away instead of making the client wait. It hands back a link the client can check later to see if the work is done. This keeps the connection short and frees the client to do other things while it waits.

solid answer

~30 s

This is the Asynchronous Request-Reply pattern. The server accepts the request, starts the long-running work in the background (e.g. on a queue or worker pool), and immediately returns 202 Accepted with a Location header pointing to a status resource. The client polls that URL until it sees a terminal state (completed or failed) carrying the result or an error. This avoids tying up an HTTP connection/thread for minutes, sidesteps client, proxy, and gateway timeouts that are usually tens of seconds, and lets the request-accepting tier scale independently from the work-processing tier.

go deeper

for a junior

Can explain that 202 means 'accepted, not done yet' and that the client checks back later via a URL; doesn't need header-level detail.

for a middle

Should know to include a Location header and describe polling with backoff; can sketch job state transitions such as pending, running, complete, failed.

for a senior

Discusses trade-offs versus holding the connection open, idempotency of the initial submission, and how to avoid poll storms.

for a principal

Reasons about this pattern as one option in a portfolio — polling versus SSE/WebSocket versus webhook — and how job-store durability and cleanup affect overall system cost and reliability.

## What the pattern is The Asynchronous Request-Reply pattern decouples the moment a client submits a request from the moment its result is ready, by having the server return an immediate **acknowledgment** instead of blocking the HTTP connection until processing finishes. ## How it proceeds mechanically Mechanically it proceeds in clear steps. 1. The client sends a request, typically a `POST`, to start the operation. 2. The server validates the input, generates a **job id** (also called a task id or correlation id), enqueues the actual work for a background worker or queue, and responds right away with `HTTP 202 Accepted` plus a `Location` header pointing to a status resource such as `/jobs/{id}`, sometimes with a `Retry-After` header suggesting when to check back. 3. The client then periodically issues `GET` requests against that status URL. 4. While work is in progress the status resource reports something like `{"status": "running"}`; once finished it reports either the completed result (inline or via a link to a separate result resource) or an error describing why the job failed. Some implementations skip polling entirely and instead let the client register a **webhook** URL that the server calls back once the job finishes. ## Why the pattern exists The pattern exists because HTTP was designed around short synchronous request/response cycles, but many real operations take anywhere from seconds to hours: - video encoding - bulk data import - third-party settlement - ML batch inference - large report generation Holding a connection and a server thread open for that long is **wasteful and fragile**: browsers, mobile OS network stacks, load balancers, and API gateways all enforce timeouts that are usually well under a minute, so a synchronous call for a two-minute job would simply fail before completion. Returning control to the client immediately frees connection and thread resources on both ends, lets the server process the work on infrastructure sized and scaled independently of the request-accepting tier, and lets the client update its UI, do other work, or resubmit polling with backoff rather than sitting frozen. ## The trade-off The trade-off is that responsibility shifts from 'keep a connection open' to 'the server must durably remember job state.' - The server needs a **job/status store** (a database row, a cache entry, etc.) that survives the worker crashing or restarting. - The client takes on the burden of implementing a polling loop with reasonable backoff, or of exposing a reachable endpoint if using callbacks instead. Compared to simply waiting on a long synchronous call, this trades a simple mental model for more moving parts, in exchange for actually working at scale and under realistic network timeout constraints. It also differs from plain **fire-and-forget** event publishing: the original caller here still gets a specific, correlated answer to their own request, which a broadcast-style publish does not provide by itself. ## Failure patterns in production In production, watch for failure patterns such as: - clients polling too aggressively and overloading the status endpoint - jobs whose worker crashed and never update status again, leaving the client polling forever - duplicate job creation if a client that never received the 202 retries its original submission A concrete example: the Azure Architecture Center documents this exact pattern for services like document or video processing, where a client uploads a file, receives 202 with a status URL, polls it, and eventually gets a link to the finished output in blob storage; identity-verification and background-check APIs use the same shape because the underlying check can take minutes to hours and depends on external systems outside the API's control.

  • Why return 202 Accepted specifically instead of 200 OK when the server accepts the async request?
    202 explicitly signals 'accepted for processing, not yet complete,' distinguishing it from 200 which implies the response body is the final result. This lets clients, proxies, and API gateways treat the response semantically differently — they know to look for a status link rather than treat the body as a finished payload, and caching layers won't mistake it for a stable final resource.
  • What should the status endpoint return once the job is done — the result directly or a redirect?
    Both are used in practice. Returning the result inline in the 200 response body keeps the client model simple, one shape of status plus optional payload. Returning a 303 See Other with a Location pointing to a separate result resource cleanly separates status polling from result retrieval, which is useful when the result is large or has its own caching or access-control needs.
  • How do you prevent a slow poll interval from making the user wait too long after the job actually finishes?
    Use a Retry-After header the server tunes dynamically, short at first and growing, or based on an estimated remaining time, so clients neither over-poll nor under-poll. For clients that need near-immediate notification, offer a push-based alternative such as Server-Sent Events, WebSockets, or a webhook callback instead of pure polling.

Like ordering food at a fast-food counter — you get a receipt with an order number and are told to watch the display board, instead of standing at the counter until your food is cooked.

saying these in an interview costs you the question

  • says the client just 'waits' for 202 to turn into the final response on the same connection
  • doesn't mention a status/polling resource or Location header
  • confuses this with plain fire-and-forget messaging
  • thinks 202 means the operation already succeeded
  • no mention of how the client knows when to stop polling

context

open as a page

When you design the status-polling endpoint for an async request-reply API — the resource a client hits after receiving HTTP 202 — what should it return across the job's lifecycle, from just-submitted to complete, in terms of status codes, body shape, and headers?

level: middleimportance: must knowfreq 60%

basics

~20 s

The status endpoint should say 'still working' while the job runs, then either hand back the finished result or an error once it's done, and it can include a hint about how long to wait before checking again.

open as a page

In an async request-reply system, a client's initial POST times out on the network before the client receives the 202 response, so the client doesn't know whether the server actually accepted the job. What should the API design do to prevent this from silently creating duplicate jobs, and what other failure modes does a production implementation of this pattern need to guard against?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Let the client attach a unique 'idempotency key' to its request; if it retries after a timeout, the server recognizes the same key and returns the existing job instead of starting a new one. Beyond that, watch out for stuck jobs, crashed workers, and clients polling forever.

open as a page

An async request-reply API can notify completion either by having the client poll a status endpoint, or by having the client register a webhook URL that the server calls back with the result. What are the operational trade-offs between these two delivery mechanisms, and when would you choose one over the other?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Polling means the client keeps checking back for updates; a callback (webhook) means the server calls the client back when the work is done. Polling is simpler to set up on the client side; callbacks are faster and less wasteful but need the client to run something reachable.

open as a page

You're designing an API for an operation that typically completes in under 300ms but occasionally, for large inputs, takes 10+ seconds. A colleague proposes using the Asynchronous Request-Reply pattern, HTTP 202 plus polling, for every call to keep the API uniform. What are the arguments against reflexively applying this pattern here, and what alternatives would you weigh instead?

level: principalimportance: should knowfreq 40%

basics

~20 s

For work that's usually fast, forcing every client through an accept-then-poll dance adds unnecessary round trips and complexity for the common case. Better options: keep it synchronous with a generous timeout for the fast path, or only switch to async for inputs that are actually large or slow.

open as a page