skip to content

Design an API for an operation that takes several minutes to finish, using HTTP 202 Accepted plus a separate status resource. How do client retries fit into that design?

level: seniorimportance: should knowfreq 38%

answer

  1. 202 = accepted, outcome undetermined
  2. Location → status resource; poll is a safe GET
  3. Validate synchronously, persist before returning 202
  4. Sticky terminal states + retention window
  5. Workers must be idempotent: queues are at-least-once

basics

~20 s

The submit call validates, persists the job, and returns 202 with a Location header pointing at a status resource. The client polls that URL for state and follows a link to the result on completion. Retrying the submit must resolve to the same job; polling a status URL is always a safe GET.

solid answer

~50 s

Split acceptance from completion. - **Submit**: `POST /exports` (or `PUT /exports/{client-uuid}` if you want the submit itself replay-safe) validates synchronously, writes the job durably, and returns **202 Accepted** with `Location: /exports/{id}` plus a small body carrying the job id and state. - **Status resource**: `GET /exports/{id}` → `{ state: queued|running|succeeded|failed, progress, error, resultUrl }`. It is a plain resource, so polling is a safe `GET`. Use the HTTP `Retry-After` response header to pace clients. - **Result**: on success the status resource embeds it or links to `/exports/{id}/result`; a common variant redirects with `303 See Other`. Retry rules: retrying the *submit* must land on the same job — that is what a client-supplied id (or the sibling idempotency-key mechanism) buys you. Retrying a *poll* is always safe. And the worker itself must tolerate redelivery, because at-least-once queues re-run tasks.

code

http · 16 lines
http
POST /exports HTTP/1.1
Content-Type: application/json

{"range":"2026-07"}

HTTP/1.1 202 Accepted
Location: /exports/8f21
Retry-After: 5

{"id":"8f21","state":"queued"}

GET /exports/8f21 HTTP/1.1

HTTP/1.1 200 OK

{"id":"8f21","state":"succeeded","resultUrl":"/exports/8f21/result"}

go deeper

for a junior

Explain 202 means accepted-not-done, that the response points at a status URL via Location, and that the client polls that URL.

for a middle

Add the state machine, Retry-After pacing, a separate result URL, and why polling is safe while re-submitting is not.

for a senior

Cover durability before returning 202, sticky terminal states and retention, making the submit replay-safe, and worker-level idempotence under at-least-once queues.

for a principal

Weigh polling against webhooks or streaming at fleet scale, set retention and cancellation policy, and standardise the async job contract across services.

## Why 202 exists `200` promises the work is done. If the work takes minutes, holding the connection open ties up a socket and a thread, dies at the first proxy idle timeout, and leaves the client with no way to learn the outcome after a disconnect. **202 Accepted** says: I have durably taken responsibility for this request, the outcome is not yet determined, and here is where to look. It converts one long call into a short accept plus an observable resource — which is also what makes the whole flow retryable. ## The three pieces **1. Submit.** Do all *synchronous* validation here — schema, authorization, quota, business preconditions. Anything rejected after you return 202 becomes an asynchronous failure the client must discover by polling, which is a much worse experience. Then persist the job record before returning; a 202 for work that exists only in an in-memory queue is a lie a pod restart exposes. **2. Status resource.** `GET /exports/{id}` returns a state machine: `queued`, `running`, `succeeded`, `failed`, `cancelled`. Include `progress` where meaningful, a structured `error` for failures, and a link to the result. Design principles: - Terminal states must be *sticky* — once `succeeded`, always `succeeded`, so a late poll gets a definitive answer. - Keep completed jobs queryable for a documented retention window; a `404` on a status URL is indistinguishable from "never existed". - Return `Retry-After` (or a `pollAfter` field) so clients pace themselves instead of hammering at 100 ms. - Consider `303 See Other` on completion, pointing at the result resource, so a generic client can follow it. **3. Result.** A separate URL keeps the status payload small and lets the result carry its own caching and content type — for instance a signed download URL for a large export. ## How retries interact Three retry paths, three rules: - **Retrying the submit.** The risky one: a timeout on submit leaves the client unsure whether a job exists. Solve it as with any create — a client-chosen job id with `PUT /exports/{id}`, a natural key, or the sibling idempotency-key header — so the retry returns the *existing* job (202 again, same `Location`) instead of starting a second export. Without that, a flaky network produces duplicate multi-minute jobs. - **Retrying a poll.** Always safe: a `GET` on a resource. This is a major reason to model status as a resource rather than as a callback-only flow. - **Redelivery to the worker.** Queues are at-least-once, so the *task* can run twice even if the API accepted it once. The job handler needs its own idempotence — a state transition guarded by a conditional update (`WHERE state = 'queued'`), or effects keyed by the job id so a re-run overwrites rather than appends. ## Failure and cancellation Distinguish "the job failed" from "the request was bad": the former is `state: failed` inside a `200` status response with a structured reason, not an HTTP error on the poll — the poll itself succeeded. Offer cancellation as `DELETE /exports/{id}` or a state transition, and make cancel idempotent too. Decide what happens when a cancel races completion; usually completion wins and the state stays terminal. ## Alternatives and when to prefer them Webhooks or server-sent events push the result instead of being polled, which scales better with many waiting clients — but they need endpoint registration, delivery retries with their own at-least-once semantics, and signature verification. The pragmatic default is 202 plus a status resource *plus* an optional callback, because polling remains a reliable fallback when a webhook delivery is lost. Long polling on the status URL is a cheap middle ground. ## The common anti-patterns Returning 202 with no `Location`; returning 200 with `{status:"pending"}`, which lies about completion and defeats generic tooling; losing accepted work on restart; and status URLs that vanish immediately after completion so a reconnecting client can never learn the outcome.

  • A client retries the submit call after a timeout. How do you prevent a second job from starting?
    Give the submit a stable identity the retry can reuse — a client-generated job id used with PUT, a natural key such as (tenant, report, period), or the sibling idempotency-key mechanism. The server then returns 202 with the same Location for the existing job instead of enqueuing another. Without a stable key, the retry is indistinguishable from a genuine second request.
  • Why model status as a resource instead of returning 200 with a 'pending' field from the original endpoint?
    Because a resource is independently addressable, so any client — including one that reconnected, restarted, or is a different process — can learn the outcome by GETting the URL, and polling it is a safe idempotent read. A 200 with a pending field misreports completion to generic tooling and caches, and ties the outcome to a single call that may never be repeated.
  • The worker receives the same task twice from the queue. What protects you?
    Idempotence inside the job handler, because queues deliver at least once regardless of how carefully the API accepted the request. Guard the state transition with a conditional update such as WHERE state = 'queued', and key any produced artifacts by the job id so a re-run overwrites the same output instead of appending a second one.

saying these in an interview costs you the question

  • Returning 202 without a Location or any way to discover the outcome
  • Returning 200 with a 'pending' status field instead of 202 plus a status resource
  • Accepting the job into an in-memory queue and losing it on restart
  • Deleting the status resource the moment the job completes, so a reconnecting client gets 404
  • Assuming the worker cannot run twice because the API deduplicated the submit

context