When a remote procedure call returns an error, how do you tell a transient failure worth retrying from a permanent one?
answer
- would the same request ever succeed
- cause in the request or the processing
- Sender versus Receiver faults
- transient is not the same as safe
basics
~20 sAsk whether the identical request could ever succeed. Transient errors come from processing conditions — overload, an unreachable upstream — and may clear; permanent ones come from the request itself — bad arguments, unknown method, refused credentials — and recur until it changes.
solid answer
~50 sClassify by where the cause lives. A **permanent** error comes from the request, or from state that will not change by itself — malformed parameters, an unknown procedure, missing authorisation, a business rule such as an insufficient balance — so an identical resend fails the same way; SOAP 1.2 calls this an `env:Sender` fault, 'not to be resent without change'. A **transient** error comes from the processing side — an upstream that did not answer, overload, a restart — and SOAP 1.2's `env:Receiver` fault says the message 'could succeed if resent at a later point in time'; HTTP's `503` with `Retry-After` says the same. Two cautions: transient does not mean safe, because the failed attempt may have done part of the work; and some conflicts need the caller to redo a whole read-modify-write sequence, not resend the one call.
go deeper
Recall the test: could the identical request succeed later? Bad input or an unknown method will not; an overloaded or restarting server might.
Explain how protocols signal the classes — SOAP Sender versus Receiver faults, JSON-RPC's request-side codes, HTTP 503 — and why a conflict needs a fresh read-modify-write.
Keep transient and retry-safe apart in practice: a transient error on a non-idempotent call still needs the outcome resolved, and ambiguous internal errors need classifying.
Standardise an error contract across services that states each error's class and whether work was applied, so client teams stop guessing and stop retrying permanent failures.
## Two questions, not one When a **remote procedure call** fails with an error, a caller deciding whether to resend has to answer two separate questions: 1. **Could the same request succeed later?** That is the **transient versus permanent** question. It depends on where the cause lives. 2. **Is it safe to send again?** That depends on whether the failed attempt may already have applied its effect, and whether the procedure is idempotent. | | Safe to resend | Not safe to resend | |---|---|---| | **Transient** | resend after a delay | resolve the outcome first, then resend if it did not run | | **Permanent** | resending is pointless | resending is pointless | Mixing the two questions is the usual mistake: "it was a transient error, so I resent the payment" can still double-charge. ## Where the cause lives - **Permanent** — the request is wrong, or the state it needs will not change on its own: invalid arguments, an unknown procedure, refused credentials, a violated business rule. An identical resend fails identically. The fix is to *change* something — the input, the credentials, the state — or to report the error. - **Transient** — the request is fine, but processing it failed for reasons outside it: the server is overloaded or restarting, an upstream did not answer, a connection broke. The same request may succeed later. - **Ambiguous** — timeouts and generic internal errors could be either, and may have happened after the procedure ran. They need the unknown-outcome treatment before any resend. ## How the specifications name the classes | Protocol | Permanent signal | Transient signal | |---|---|---| | SOAP 1.2 | `env:Sender` fault: "not to be resent without change" | `env:Receiver` fault: "could succeed if resent at a later point in time" | | SOAP 1.1 | `Client` faultcode: "should not be resent without change" | `Server` faultcode: "may succeed at a later point in time" | | JSON-RPC 2.0 | `-32700` Parse error, `-32600` Invalid Request, `-32601` Method not found, `-32602` Invalid params: all describe the request itself | none defined; `-32603` Internal error and the implementation-defined `-32000` to `-32099` server errors carry no retry meaning | | HTTP (RFC 9110) | most `4xx` client errors (`408` Request Timeout is an exception: the client MAY repeat the request) | `503` Service Unavailable, a temporary overload or maintenance, optionally with `Retry-After` | Frameworks often draw finer lines. gRPC's status-code guidance, for example, separates a failure the client "can retry just the failing call" (`UNAVAILABLE`), one where it "should retry at a higher level" by restarting a read-modify-write sequence (`ABORTED`), and one where it "should not retry until the system state has been explicitly fixed" (`FAILED_PRECONDITION`) — and warns that even `UNAVAILABLE` is "not always safe to retry" for non-idempotent operations. That third class is worth naming on its own: an **optimistic-concurrency conflict** is transient in a sense, but resending the same write with the same stale version fails again. The caller has to re-read, recompute and send a *new* request. ## A decision procedure 1. **Did the server reject the request before running it** (parse error, unknown method, invalid parameters, refused credentials)? Permanent — fix and report; do not resend. 2. **Is it a conflict on state the caller read earlier?** Re-run the whole read-modify-write, not the single call. 3. **Is it a declared transient condition** (receiver fault, overload, unavailable)? Candidate for resending — go to step 5. 4. **Is it a timeout or a generic internal error?** Treat the outcome as unknown — go to step 5. 5. **Could the attempt have applied its effect?** If yes and the procedure is not idempotent, resolve the outcome first (status lookup, reconciliation); otherwise resend after a delay. How long to wait between attempts, how many to make and how to spread them out are retry-policy subjects of their own. ## What servers owe their callers - **Pick the class honestly.** Returning a generic internal error for invalid input turns a permanent failure into an apparently transient one, and clients retry it pointlessly. - **Document retryability per error** in the interface contract; JSON-RPC 2.0, for one, leaves the meaning of its server-error range to each implementation. - **Say whether the work ran.** An error that guarantees nothing was applied lets the client resend freely.
- Is a JSON-RPC 2.0 -32603 Internal error transient or permanent?The specification does not say. It defines -32603 only as an internal JSON-RPC error, and -32000 to -32099 as implementation-defined server errors, with no retry meaning attached. The caller must treat them as unknown in class — and possibly executed — unless the server's own documentation classifies them.
- Why can resending on a permanent error make an incident worse?Every resend fails the same way, so retries multiply load on a server that is answering correctly, fill logs with noise and delay the caller's real handling — fixing input, refreshing credentials or reporting to the user. In a chain of callers that each retry, one bad request becomes many identical failures.
saying these in an interview costs you the question
- Every error from the server deserves one retry, just in case.
- A transient error means the call did not run, so resending is safe.
- A SOAP 1.2 Sender fault tells the client to resend the message later.
- JSON-RPC 2.0 defines which of its error codes are retryable.
- A version conflict is fixed by resending the identical write.