An HTTP call to a payment service times out after the client has finished sending the request body. Is it safe to retry, and what determines the answer?
answer
- Timeout = unknown, not failed
- Safe class: DNS/refused/connect timeout/TLS/reset with 0 bytes read
- Unknown class: read timeout after write, mid-response reset, 502/503/504
- Method idempotency is a claim; check what the handler does
- Bounded retries + backoff with jitter + retry budget; never retry 4xx
basics
~20 sA timeout means unknown, not failed - the server may have processed the request and lost the response. Retry only if the failure proves the request never ran (connect failure, refused connection, reset with no bytes received) or the operation is genuinely idempotent server-side.
solid answer
~60 sClassify the failure first: - **Never reached the application** - DNS failure, connection refused, connect timeout, TLS handshake failure, or a reset on a pooled connection with zero response bytes received. Retrying is safe. - **Reached the application, outcome unknown** - a read/response timeout after the request was written, or a reset mid-response. The server may have completed the charge and failed only to deliver the answer. Retrying blindly risks a duplicate. Then consider idempotency. GET, HEAD, PUT and DELETE are idempotent *by specification*, POST is not - but specification idempotency is a claim about intent, and what matters is whether **your** implementation actually produces the same result on repetition. A PUT that appends to a list is not idempotent in practice. For a non-idempotent operation such as a payment, the correct answer is not "retry" or "do not retry" but **make retrying safe**: the server must deduplicate repeated submissions of the same logical operation, or you must be able to query whether it succeeded before resubmitting. Add bounded retries with exponential backoff and jitter, a retry budget to avoid amplifying an overload, and no retries on 4xx.
go deeper
Know that a timeout does not tell you whether the server did the work, and that retrying a payment can charge twice.
Classify failures into 'definitely not processed' versus 'unknown', and check whether the operation really is idempotent rather than trusting the method.
Design the retry policy: bounded attempts, backoff with jitter, a retry budget, no retries on 4xx, deadline-aware, plus server-side deduplication for non-idempotent operations.
Own the failure model across services - who resolves ambiguity and how, reconciliation paths for unknown outcomes, and retry amplification as a systemic stability risk with explicit budgets and load shedding.
## Timeouts are ambiguous by construction When a request times out, the client learns exactly one thing: no response arrived in time. It does not learn whether the server received the request, whether it executed it, or whether it succeeded and the response was lost on the way back. In distributed-systems terms the outcome is *unknown*, and treating unknown as failed is how duplicate charges, double-shipped orders, and doubled ledger entries happen. ## Step one: what does the failure prove? Different failures carry different amounts of information. **Proves the request was not processed** (safe to retry regardless of method): - DNS resolution failure - no connection at all. - `ECONNREFUSED` - nothing listening; the SYN was rejected. - Connect timeout - the handshake never completed. - TLS handshake failure - no application data was exchanged. - A reset or EOF on a connection reused from the pool with **zero response bytes received** - the classic stale-connection race, where the server had already closed before reading. **Leaves the outcome unknown** (retry only if the operation tolerates repetition): - Read or response timeout after the request was fully written. - Reset mid-response. - 502, 503 or 504 from a gateway - the upstream may have executed the work before the gateway gave up. 503 with `Retry-After` is the closest thing to an explicit "safe to retry" signal, and even then only for the operation's semantics. A useful client-side rule: automatically retry only failures in the first group; escalate the second group to explicit application policy. ## Step two: is the operation idempotent? HTTP defines GET, HEAD, PUT and DELETE as idempotent - repeating them has the same effect on server state as doing them once (the *response* may differ; a second DELETE may return 404). POST is not idempotent by definition. But the method is a declaration, not a guarantee. Ask what the handler actually does: - `PUT /orders/123` that overwrites a record - idempotent in practice. - `PUT /counters/x` that increments - not idempotent, despite the method. - `POST /payments` that creates a new charge per call - the dangerous case. So the safe formulation is: retry when repetition is provably harmless in the implementation, not when the verb is in the idempotent list. ## Step three: make retries safe rather than avoiding them For a payment, never retrying is also wrong: a transient timeout would strand a customer. The durable answer is server-side deduplication - the client sends a stable identifier for the logical operation, and the server recognises a repeat and returns the original outcome instead of executing again. Alternatively, provide a read path that answers "did operation X succeed?", so the client can reconcile before deciding. Either way the ambiguity is resolved at the application layer, where the business meaning lives; the transport cannot resolve it. ## Step four: retry mechanics - **Bound attempts.** One or two retries; beyond that you are queueing work against a service that is already struggling. - **Exponential backoff with jitter.** Fixed delays synchronise clients into retry waves; jitter spreads them. - **Retry budget.** Cap retries as a fraction of total requests (a few percent). Without a budget, a partial outage triggers retries that multiply load and turn a degradation into a metastable collapse that persists after the original cause is gone. - **Do not retry 4xx.** A 400 or 422 will fail identically; 429 is the exception and should honour `Retry-After`. - **Respect the deadline.** Total time including retries must fit inside the caller's budget, or you are burning capacity on work nobody is waiting for. - **Retry on a fresh connection**, and prefer a different backend instance if your client can steer. ## Step five: observe Separate metrics for attempts and logical requests, and count retries by failure class. If the retry rate rises without a corresponding error rate at the callee, the client is generating load that will eventually cause the outage it is trying to survive.
- Which specific failures prove the request was never processed, so retrying is safe even for a non-idempotent operation?Failures before the request could reach application code: DNS resolution failure, connection refused, connect timeout, and TLS handshake failure. Also a reset or EOF on a pooled connection where zero response bytes were received, since that signature means the server closed before reading. Anything after the request was fully written - notably a read timeout - leaves the outcome unknown and does not qualify.
- Why do retries need a budget rather than just a maximum attempt count per request?A per-request cap still lets every client retry simultaneously when a dependency degrades, multiplying offered load by the retry factor exactly when the system has least headroom. A budget caps retries as a fraction of total traffic across the client, so retries help isolated failures but cannot amplify a widespread one. Without it a partial outage can become a self-sustaining metastable failure that persists after the trigger is gone.
- A PUT is idempotent by specification. When is retrying one still unsafe?When the handler's behaviour is not actually idempotent - for example a PUT that increments a counter, appends to a collection, or triggers a side effect such as sending an email on each call. Idempotency is a property of the implementation; the method only declares intent. Conditional requests using ETags add another wrinkle, since a retried conditional PUT may legitimately fail with 412 after the first one succeeded.
Posting a letter and getting no reply: you cannot tell whether it never arrived or the reply was lost. Sending a second letter is fine for "please confirm my address" and dangerous for "please charge my card" - unless the recipient can spot that it is the same letter.
saying these in an interview costs you the question
- Treating a timeout as proof the request failed
- Assuming any GET, PUT or DELETE is safe to retry because the specification calls the method idempotent
- Retrying with unbounded attempts or fixed-interval backoff, producing synchronised retry storms
- Retrying 400-class errors that will deterministically fail again
- Retrying the request on the same connection that just reset instead of a fresh one