skip to content

A client calls a REST endpoint that kicks off a report that takes 2 minutes to generate. Instead of holding the connection open, the server responds immediately with HTTP 202 Accepted and a URL the client can check later. What is this approach called, and why use it instead of just blocking the connection until the report is ready?

level: juniorimportance: must knowfreq 70%

answer

  1. 202 Accepted + Location header
  2. poll a status resource
  3. decouple accept from complete
  4. job/task/correlation id
  5. worker processes in background

basics

~20 s

The server says 'got it, working on it' right away instead of making the client wait. It hands back a link the client can check later to see if the work is done. This keeps the connection short and frees the client to do other things while it waits.

solid answer

~30 s

This is the Asynchronous Request-Reply pattern. The server accepts the request, starts the long-running work in the background (e.g. on a queue or worker pool), and immediately returns 202 Accepted with a Location header pointing to a status resource. The client polls that URL until it sees a terminal state (completed or failed) carrying the result or an error. This avoids tying up an HTTP connection/thread for minutes, sidesteps client, proxy, and gateway timeouts that are usually tens of seconds, and lets the request-accepting tier scale independently from the work-processing tier.

go deeper

for a junior

Can explain that 202 means 'accepted, not done yet' and that the client checks back later via a URL; doesn't need header-level detail.

for a middle

Should know to include a Location header and describe polling with backoff; can sketch job state transitions such as pending, running, complete, failed.

for a senior

Discusses trade-offs versus holding the connection open, idempotency of the initial submission, and how to avoid poll storms.

for a principal

Reasons about this pattern as one option in a portfolio — polling versus SSE/WebSocket versus webhook — and how job-store durability and cleanup affect overall system cost and reliability.

## What the pattern is The Asynchronous Request-Reply pattern decouples the moment a client submits a request from the moment its result is ready, by having the server return an immediate **acknowledgment** instead of blocking the HTTP connection until processing finishes. ## How it proceeds mechanically Mechanically it proceeds in clear steps. 1. The client sends a request, typically a `POST`, to start the operation. 2. The server validates the input, generates a **job id** (also called a task id or correlation id), enqueues the actual work for a background worker or queue, and responds right away with `HTTP 202 Accepted` plus a `Location` header pointing to a status resource such as `/jobs/{id}`, sometimes with a `Retry-After` header suggesting when to check back. 3. The client then periodically issues `GET` requests against that status URL. 4. While work is in progress the status resource reports something like `{"status": "running"}`; once finished it reports either the completed result (inline or via a link to a separate result resource) or an error describing why the job failed. Some implementations skip polling entirely and instead let the client register a **webhook** URL that the server calls back once the job finishes. ## Why the pattern exists The pattern exists because HTTP was designed around short synchronous request/response cycles, but many real operations take anywhere from seconds to hours: - video encoding - bulk data import - third-party settlement - ML batch inference - large report generation Holding a connection and a server thread open for that long is **wasteful and fragile**: browsers, mobile OS network stacks, load balancers, and API gateways all enforce timeouts that are usually well under a minute, so a synchronous call for a two-minute job would simply fail before completion. Returning control to the client immediately frees connection and thread resources on both ends, lets the server process the work on infrastructure sized and scaled independently of the request-accepting tier, and lets the client update its UI, do other work, or resubmit polling with backoff rather than sitting frozen. ## The trade-off The trade-off is that responsibility shifts from 'keep a connection open' to 'the server must durably remember job state.' - The server needs a **job/status store** (a database row, a cache entry, etc.) that survives the worker crashing or restarting. - The client takes on the burden of implementing a polling loop with reasonable backoff, or of exposing a reachable endpoint if using callbacks instead. Compared to simply waiting on a long synchronous call, this trades a simple mental model for more moving parts, in exchange for actually working at scale and under realistic network timeout constraints. It also differs from plain **fire-and-forget** event publishing: the original caller here still gets a specific, correlated answer to their own request, which a broadcast-style publish does not provide by itself. ## Failure patterns in production In production, watch for failure patterns such as: - clients polling too aggressively and overloading the status endpoint - jobs whose worker crashed and never update status again, leaving the client polling forever - duplicate job creation if a client that never received the 202 retries its original submission A concrete example: the Azure Architecture Center documents this exact pattern for services like document or video processing, where a client uploads a file, receives 202 with a status URL, polls it, and eventually gets a link to the finished output in blob storage; identity-verification and background-check APIs use the same shape because the underlying check can take minutes to hours and depends on external systems outside the API's control.

  • Why return 202 Accepted specifically instead of 200 OK when the server accepts the async request?
    202 explicitly signals 'accepted for processing, not yet complete,' distinguishing it from 200 which implies the response body is the final result. This lets clients, proxies, and API gateways treat the response semantically differently — they know to look for a status link rather than treat the body as a finished payload, and caching layers won't mistake it for a stable final resource.
  • What should the status endpoint return once the job is done — the result directly or a redirect?
    Both are used in practice. Returning the result inline in the 200 response body keeps the client model simple, one shape of status plus optional payload. Returning a 303 See Other with a Location pointing to a separate result resource cleanly separates status polling from result retrieval, which is useful when the result is large or has its own caching or access-control needs.
  • How do you prevent a slow poll interval from making the user wait too long after the job actually finishes?
    Use a Retry-After header the server tunes dynamically, short at first and growing, or based on an estimated remaining time, so clients neither over-poll nor under-poll. For clients that need near-immediate notification, offer a push-based alternative such as Server-Sent Events, WebSockets, or a webhook callback instead of pure polling.

Like ordering food at a fast-food counter — you get a receipt with an order number and are told to watch the display board, instead of standing at the counter until your food is cooked.

saying these in an interview costs you the question

  • says the client just 'waits' for 202 to turn into the final response on the same connection
  • doesn't mention a status/polling resource or Location header
  • confuses this with plain fire-and-forget messaging
  • thinks 202 means the operation already succeeded
  • no mention of how the client knows when to stop polling

context