In an async request-reply system, a client's initial POST times out on the network before the client receives the 202 response, so the client doesn't know whether the server actually accepted the job. What should the API design do to prevent this from silently creating duplicate jobs, and what other failure modes does a production implementation of this pattern need to guard against?
answer
- idempotency key on submission to prevent duplicate jobs
- heartbeat/visibility timeout catches stuck workers
- poll storm mitigation via Retry-After and backoff
- TTL/cleanup for orphaned job records
- at-least-once callback delivery needs dedupe on the client
basics
~20 sLet the client attach a unique 'idempotency key' to its request; if it retries after a timeout, the server recognizes the same key and returns the existing job instead of starting a new one. Beyond that, watch out for stuck jobs, crashed workers, and clients polling forever.
solid answer
~40 sRequire an idempotency key, a client-generated UUID, on the submission endpoint; the server stores it alongside the job, and a retried POST with the same key returns the existing job's status or location instead of creating a duplicate. Beyond duplicate submission, guard against workers crashing mid-job and leaving status stuck at 'running' forever, fixed with heartbeats or timeouts that flip it to 'failed'; poll storms from many clients without backoff, fixed with Retry-After plus rate limiting; orphaned job records that never get cleaned up, fixed with a TTL; and, if a webhook callback is used, ambiguous result delivery since at-least-once delivery can duplicate notifications, fixed with dedupe on delivery id.
go deeper
Recognizes that retrying after a network timeout could accidentally submit the same job twice.
Can describe using an idempotency key to prevent duplicate submissions.
Covers the full set: idempotency keys, stuck-worker detection via heartbeat or visibility timeout, poll storms, and TTL cleanup, and can explain the mechanism behind each fix.
Designs the failure-handling conventions as a reusable platform pattern — a standard idempotency-key header, a standard job-status schema, standard visibility-timeout defaults — applied consistently across many async endpoints, and reasons about the cost of each safeguard against the risk it mitigates.
## The shape of the problem Production async request-reply systems fail in a small, well-understood set of ways, and each one maps to a specific, well-established mitigation: 1. duplicate submission 2. the worker crashing mid-processing 3. the poll storm 4. orphaned job records 5. duplicate webhook delivery ## First — the duplicate submission The first failure mode is the one in the question: a client submits the initial `POST`, the network drops the response before the client sees the 202, and the client — reasonably, since it cannot distinguish 'server never got it' from 'server got it but the ack was lost' — retries the exact same submission. Without any safeguard this creates a second, duplicate job doing the same work twice, which is wasteful at best and dangerous at worst if the operation has real-world side effects like charging a card or sending an email. The standard fix is a **client-generated idempotency key**, often a UUID, attached to the submission request. The server stores this key alongside the job it creates, and if it later receives another `POST` carrying the same key, it recognizes the retry and returns the existing job's status or location rather than starting new work. This differs from deduplicating by request-body content, because two requests with identical-looking payloads might legitimately be separate operations submitted on purpose; the idempotency key represents a specific attempt, letting the server say 'same key equals same attempt' without having to infer intent from the payload. ## Second — the worker crashing mid-processing A second failure mode is the worker crashing mid-processing after the job's status has already been flipped to 'running,' leaving the status stuck there forever with no update. - **Visibility timeout.** Queue-backed architectures address this with a visibility timeout: when a worker pulls a job off the queue, the message becomes invisible to other workers for a configured period so it isn't processed twice concurrently, and if the worker crashes before finishing and deleting the message, the message reappears once that timeout expires so another worker can retry it. - **Heartbeat or overall timeout.** This must be paired with the status store itself reflecting stalled jobs, typically via a heartbeat the worker writes periodically, or an overall timeout after which a job stuck in 'running' too long is flipped to 'failed' so a polling client isn't left waiting indefinitely for something that will never finish. ## Third — the poll storm A third failure mode is the poll storm: many clients polling a status endpoint aggressively with no backoff, especially when a batch UI action starts many jobs at once and all their clients begin polling in near lock-step. This is mitigated with a `Retry-After` header the server tunes to suggest sensible intervals, combined with: - enforced **exponential backoff and jitter** on the client side - potentially **rate limiting** on the status endpoint itself as a backstop ## Fourth — orphaned job records A fourth is orphaned job records: without a TTL or scheduled cleanup, job and result records accumulate indefinitely even though no client will ever query most of them again, which is purely an operational and cost problem rather than a correctness one, but it compounds over time in a busy system. ## Fifth — duplicate webhook delivery A fifth failure mode is specific to systems using webhook callbacks rather than polling: at-least-once delivery means the same completion notification can be delivered more than once, so the client must dedupe by job id or delivery id and treat duplicate delivery as normal rather than as an error. ## What to return for a job the server has lost Finally, consider what a status endpoint should return if the server genuinely has no record of a job at all, for instance after a database failover lost recent writes — it should return an explicit 'unknown' or `404` response rather than a generic error or an indefinite hang, so the client can clearly decide whether to resubmit rather than being stuck unable to tell 'never existed' from 'transient failure, keep polling.' A well-known real-world instance of the idempotency-key safeguard is Stripe's API, which accepts an `Idempotency-Key` header on `POST` requests specifically so that network retries of payment-creation calls never double-charge a customer.
- How does an idempotency key differ from simply deduplicating by request body content?Two logically identical-looking requests might legitimately be separate operations, for example charging $10 to the same card twice on purpose, so content alone can't distinguish a retry of the same attempt from a new, coincidentally identical request. A client-generated idempotency key represents a specific attempt id, letting the server safely say 'same key equals same attempt, return the existing result' without guessing intent from the payload.
- What is a 'visibility timeout' in the context of a queue-backed async worker, and how does it relate to stuck jobs?When a worker pulls a job message off a queue, the message becomes invisible to other workers for a configured visibility timeout so it isn't processed twice concurrently. If the worker crashes before finishing and before deleting the message, the message reappears once the timeout expires so another worker can retry it, which prevents a crashed worker from permanently stranding a job, provided the status store stays consistent with that retry.
- If a client polls a status endpoint for a job the server has no record of, what should the API return, and why does it matter?It should return a clear 404 or an explicit 'unknown/lost' status rather than hanging or returning a generic 500, so the client can distinguish 'this job never existed or was lost' from 'transient error, keep polling.' An ambiguous response leaves the client unable to decide whether to resubmit, risking either abandoning real work or creating unwanted duplicates.
Like calling a hotline, hanging up before hearing confirmation, and calling again — if the hotline doesn't recognize you as the same caller, it opens a second ticket for the exact same problem.
saying these in an interview costs you the question
- only mentions duplicate submissions and nothing else when asked about failure modes broadly
- doesn't know what an idempotency key is or confuses it with a plain job id
- assumes workers never crash mid-processing
- no mention of poll storms or backoff
- treats webhook delivery as guaranteed exactly-once