skip to content

When you design the status-polling endpoint for an async request-reply API — the resource a client hits after receiving HTTP 202 — what should it return across the job's lifecycle, from just-submitted to complete, in terms of status codes, body shape, and headers?

level: middleimportance: must knowfreq 60%

answer

  1. state machine: pending -> running -> complete/failed
  2. Retry-After header on in-progress responses
  3. Cache-Control: no-store on the status resource
  4. terminal state = inline result or 303 redirect
  5. GET is safe/idempotent, so polling is free of side effects

basics

~20 s

The status endpoint should say 'still working' while the job runs, then either hand back the finished result or an error once it's done, and it can include a hint about how long to wait before checking again.

solid answer

~30 s

While pending, GET on the status resource returns 200 with a body like {status: "running"} and ideally a Retry-After header suggesting the next poll time. Once terminal, it returns 200 with {status: "completed", result: ...}, or a 303 redirect to a separate result resource, or {status: "failed", error: ...}. The response must be marked non-cacheable (Cache-Control: no-store) so intermediaries don't serve a stale 'running' response after completion. Some designs also support DELETE on the resource to cancel a job, and apply a TTL so old job records eventually get cleaned up.

go deeper

for a junior

Knows the status endpoint returns something different once the job is done versus while it's running.

for a middle

Can design the state enum (pending/running/completed/failed), pick status codes and headers like Retry-After and no-store, and describe inline-versus-redirect result delivery.

for a senior

Adds a cleanup/TTL strategy, cancellation, and correct cache-control behavior under real proxy or CDN layers; reasons about payload size driving inline versus redirect.

for a principal

Weighs this against alternative status-delivery mechanisms at the API-portfolio level and sets org-wide conventions, such as a standard job-status envelope reused across many async endpoints.

## Treat job progress as a state machine Designing the status resource well means treating job progress as an explicit **state machine** and mapping each state onto a concrete HTTP response shape, rather than leaving the client to guess. A typical state machine has at minimum `pending`, `running`, `completed`, and `failed`, and the status body should carry this as an explicit field (e.g. `{"status": "running"}`) rather than relying purely on HTTP status codes, since a `GET` against an existing, valid resource is naturally still `200 OK` in most states — the distinguishing information belongs in the body, not the transport-level code. ## While the job is in flight - **`Retry-After`.** While the job is in flight, the response can include a `Retry-After` header (either a fixed number of seconds, or a smarter estimate based on the actual work remaining) so well-behaved clients know when to check again instead of guessing an interval. - **Cache headers.** The response must also be marked non-cacheable — `Cache-Control: no-store`, or `no-cache` with `must-revalidate` — because a stable URL with a volatile body is exactly the case reverse proxies and CDNs get wrong by default, silently serving a client an old 'running' response long after the job actually finished. ## Two shapes for a terminal state Once the job reaches a terminal state, the design has two common shapes. 1. The first embeds the result directly in the status body, which keeps the client model simple: one `GET`, one response shape, whether pending or done. This works well when the result is small, like a JSON summary or a handful of fields. 2. The second returns `303 See Other` with a `Location` header pointing at a separate result resource, decoupling the lifecycle and access rules of the result (which might be a large file, have its own cache headers, or need independent sharing) from the lifecycle of the job-status record itself. Failure should also be represented explicitly, typically `{"status": "failed", "error": {...}}` with enough detail for the client to decide whether to retry the original submission or give up, while the HTTP status of the status `GET` itself usually stays 200, since the poll succeeded even though the underlying job did not. ## What happens to the record afterwards A subtler but important design point is what happens after the client is done with the record. Job-status entries must not live forever: without a TTL or a scheduled cleanup job, a busy API accumulates unbounded storage for jobs nobody will ever query again. Production designs commonly: - expire records some window after completion, often 24 to 72 hours - sometimes exposing an explicit `DELETE` on the job resource so clients or operators can clean up proactively This matters operationally because the job store is now a piece of durable state the team owns and must monitor, unlike a purely synchronous endpoint whose only state is the in-flight HTTP call. ## Why polling is a sound design `GET` on the status resource is inherently **safe and idempotent** — issuing it any number of times has no side effects and always reflects current truth — which is exactly why polling is a sound design: clients can poll freely without coordination concerns, unlike retrying a `POST`. A concrete real-world example is GitHub's asynchronous migration/export APIs, which accept a request to generate a repository export, return 202 with a status URL, and require the client to poll that URL until it reports the export is ready, at which point it supplies a download link separate from the status resource itself — exactly the redirect-to-result-resource shape described above.

  • Should the status endpoint return the result payload inline, or a link to a separate result resource? What determines the choice?
    Inline works well when the result is small, such as a short JSON object, and keeps client logic simple with a single response shape. A separate result resource via a 303 redirect is better when the result is large, like a generated file or video, or needs its own caching and access-control rules independent of the job-status record's lifetime.
  • How do you keep a CDN or reverse proxy from serving a stale 'still running' response after the job has actually completed?
    Mark the status response with Cache-Control: no-store, or no-cache with must-revalidate, so intermediaries always forward the request to origin instead of serving a cached copy. A stable URL with a changing body is exactly the scenario caching layers mishandle unless told explicitly not to cache.
  • What happens to the job/status record after the client has retrieved the final result — does it live forever?
    No, production designs apply a TTL or a scheduled cleanup job that removes status and result records some time after completion, both to bound storage growth and to avoid serving stale results long after they matter. Some APIs also let clients or operators explicitly DELETE the job resource to clean up early.

Like a package tracking page — while in transit it shows 'in transit, check back tomorrow,' and once delivered it shows 'delivered' with a link to the signature; refreshing the page never re-ships the package.

saying these in an interview costs you the question

  • treats the status endpoint as returning a fixed 202 forever with no terminal state
  • doesn't mention how the client distinguishes running vs done vs failed
  • forgets caching concerns on the status resource
  • assumes the job record can be kept indefinitely with no cleanup
  • conflates the status resource with the result resource with no rationale

context