Your HTTP client times out waiting for a response to a POST request. What do you actually know about whether the server applied it, and how should the client behave?
answer
- Timeout = unknown, not failed
- Never-sent vs ambiguous vs refused
- 502/504 and 500 are ambiguous, 4xx is not
- Reconcile by client-supplied reference
- Client timeout below server p99 manufactures duplicates
basics
~20 sAlmost nothing: a timeout means you lost the answer, not that the work did not happen. The request may have been fully applied. Retrying a non-idempotent POST can duplicate the effect, so either make the operation retry-safe first, or reconcile by querying before retrying.
solid answer
~50 sA timeout is an **ambiguous failure** — the outcome is unknown, not negative. Classify failures instead of lumping them: - **Definitely not applied**: DNS failure, connection refused, TLS handshake failure. Retry freely, even a POST. - **Ambiguous**: read timeout, connection reset mid-flight, 502/504 from a proxy, a 500 from a handler that may have partially committed. The server may have committed and lost the response. - **Deliberately refused**: 4xx other than 429 — the server explicitly rejected it, so an identical retry fails identically. For ambiguous failures on an idempotent method, retry with backoff. On a `POST` you have three options: retry anyway if the operation is naturally deduplicated; **reconcile** by querying for the effect (search for the order carrying your client reference) before deciding; or design the endpoint so replay is safe from the start. The systemic answer is to remove the ambiguity at design time rather than guess at call time.
go deeper
Say a timeout means the outcome is unknown — the server may have applied it — so retrying a POST risks a duplicate.
Classify failures into never-sent, ambiguous, and deliberately-refused, and give the correct retry decision for each.
Add server-side design that removes the ambiguity: atomic commits, a transactional outbox for effects, and a reconciliation read keyed by a client-supplied reference.
Frame it as an at-least-once system with receiver-side dedupe, and set timeout, deadline-propagation and retry conventions across services.
## The core insight The network gives you no way to distinguish "the request never arrived" from "the request was processed and the response was lost". Both look like a timeout. This is why at-most-once delivery is impossible over an unreliable channel: you can have at-least-once (retry, risk duplicates) or at-most-once (never retry, risk silent loss), and turning at-least-once into effectively-once requires the *receiver* to deduplicate. ## Classifying failures A competent retry policy branches on what the failure tells you: | Signal | Applied? | Retry a POST? | |---|---|---| | DNS failure, connect refused, TLS failure | No — never reached the app | Yes | | Client-side timeout waiting for a response | Unknown | Only if safe | | Connection reset after the request was sent | Unknown | Only if safe | | 502/504 from a gateway | Unknown — origin may have committed | Only if safe | | 503 with Retry-After from an admission layer | Usually not, but not guaranteed | Usually yes | | 429 | No — rejected before work | Yes, after the delay | | 400/404/409/422 | No — deliberate refusal | No; the retry fails identically | | 500 from your own handler | Unknown — may have partially committed | Only if safe | Two rows deserve emphasis. A **502/504** is a proxy saying *it* could not get an answer — the origin may have completed the work. And a **500** is not automatically "nothing happened": a handler that wrote a row and then threw while sending a webhook has committed part of the effect. ## Server-side design that removes the guess Because the client cannot resolve ambiguity, the server should: - **Commit atomically.** One transaction covering the state change, so partial application is impossible and the only outcomes are all or nothing. - **Publish effects transactionally.** Use an outbox rather than sending a webhook or message inside the handler, so a retry does not produce a second notification for a change that never landed. - **Give the client a way to check.** A reconciliation read — "does an order with client reference X exist?" — turns an ambiguous failure into a definite answer. This requires the client to attach its own identifier *before* the first attempt: if it only learns the ID from the response, it has no way to look up the record it may have created. ## Client-side behaviour On ambiguity with a non-idempotent operation: 1. **Reconcile if you can.** Query by your own reference, then retry only if absent. There is a small window where the record is committed but not yet visible on a replica, so use a read guaranteed to be current or tolerate a rare duplicate. 2. **Escalate to a human** for high-value irreversible operations (a payment capture, a wire transfer) rather than silently doubling it. 3. **Never blind-retry a create** with no dedupe story and no reconciliation, and never retry a 4xx. Timeouts also need to be *chosen*, not defaulted. A client timeout shorter than the server's real p99 manufactures ambiguity out of healthy requests: the server completes at 3s, the client gave up at 2s, and the duplicate order was caused entirely by configuration. Set the timeout above the server's worst realistic latency, propagate a deadline, and let the server abandon work that has exceeded it. ## What good sounds like "A timeout is unknown, not failed. I classify failures into never-sent, ambiguous, and deliberately-refused. For ambiguous failures on non-idempotent operations I reconcile using a client-supplied reference before retrying, and long term I make the operation replay-safe so the ambiguity stops mattering."
- Is a 4xx response ever worth retrying?Almost never, because the server explicitly refused and the identical request will be refused again, so retrying just wastes capacity. The exceptions are 429, which is an instruction to retry later and usually carries Retry-After, and 401 when your client can refresh a token and try again with new credentials. Everything else in the 4xx range needs a changed request, not a repeat.
- Why can a client timeout set too aggressively actually create duplicate records?If the timeout is below the server's real tail latency, healthy slow requests are abandoned by the client while the server goes on to commit them. The client then retries and creates a second record. The fix is to set timeouts from measured p99 or higher, propagate an explicit deadline so the server can abandon work nobody awaits, and never treat a timeout as evidence of failure.
Posting a letter and hearing nothing back: the silence tells you the reply went missing, not that the recipient never opened the envelope.
saying these in an interview costs you the question
- Treating a timeout as proof the request was not processed
- Retrying every 5xx identically, including a 500 that may have partially committed
- Retrying 4xx responses such as 400 or 422
- Assuming a 500 means nothing was written when the handler committed before failing
- Planning to reconcile by an identifier the client only receives in the response it never got