skip to content

When a remote procedure call returns an error, how do you tell a transient failure worth retrying from a permanent one?

level: middleimportance: should knowfreq 24%

answer

  1. would the same request ever succeed
  2. cause in the request or the processing
  3. Sender versus Receiver faults
  4. transient is not the same as safe

basics

~20 s

Ask whether the identical request could ever succeed. Transient errors come from processing conditions — overload, an unreachable upstream — and may clear; permanent ones come from the request itself — bad arguments, unknown method, refused credentials — and recur until it changes.

solid answer

~50 s

Classify by where the cause lives. A **permanent** error comes from the request, or from state that will not change by itself — malformed parameters, an unknown procedure, missing authorisation, a business rule such as an insufficient balance — so an identical resend fails the same way; SOAP 1.2 calls this an `env:Sender` fault, 'not to be resent without change'. A **transient** error comes from the processing side — an upstream that did not answer, overload, a restart — and SOAP 1.2's `env:Receiver` fault says the message 'could succeed if resent at a later point in time'; HTTP's `503` with `Retry-After` says the same. Two cautions: transient does not mean safe, because the failed attempt may have done part of the work; and some conflicts need the caller to redo a whole read-modify-write sequence, not resend the one call.

go deeper

for a junior

Recall the test: could the identical request succeed later? Bad input or an unknown method will not; an overloaded or restarting server might.

for a middle

Explain how protocols signal the classes — SOAP Sender versus Receiver faults, JSON-RPC's request-side codes, HTTP 503 — and why a conflict needs a fresh read-modify-write.

for a senior

Keep transient and retry-safe apart in practice: a transient error on a non-idempotent call still needs the outcome resolved, and ambiguous internal errors need classifying.

for a principal

Standardise an error contract across services that states each error's class and whether work was applied, so client teams stop guessing and stop retrying permanent failures.

## Two questions, not one When a **remote procedure call** fails with an error, a caller deciding whether to resend has to answer two separate questions: 1. **Could the same request succeed later?** That is the **transient versus permanent** question. It depends on where the cause lives. 2. **Is it safe to send again?** That depends on whether the failed attempt may already have applied its effect, and whether the procedure is idempotent. | | Safe to resend | Not safe to resend | |---|---|---| | **Transient** | resend after a delay | resolve the outcome first, then resend if it did not run | | **Permanent** | resending is pointless | resending is pointless | Mixing the two questions is the usual mistake: "it was a transient error, so I resent the payment" can still double-charge. ## Where the cause lives - **Permanent** — the request is wrong, or the state it needs will not change on its own: invalid arguments, an unknown procedure, refused credentials, a violated business rule. An identical resend fails identically. The fix is to *change* something — the input, the credentials, the state — or to report the error. - **Transient** — the request is fine, but processing it failed for reasons outside it: the server is overloaded or restarting, an upstream did not answer, a connection broke. The same request may succeed later. - **Ambiguous** — timeouts and generic internal errors could be either, and may have happened after the procedure ran. They need the unknown-outcome treatment before any resend. ## How the specifications name the classes | Protocol | Permanent signal | Transient signal | |---|---|---| | SOAP 1.2 | `env:Sender` fault: "not to be resent without change" | `env:Receiver` fault: "could succeed if resent at a later point in time" | | SOAP 1.1 | `Client` faultcode: "should not be resent without change" | `Server` faultcode: "may succeed at a later point in time" | | JSON-RPC 2.0 | `-32700` Parse error, `-32600` Invalid Request, `-32601` Method not found, `-32602` Invalid params: all describe the request itself | none defined; `-32603` Internal error and the implementation-defined `-32000` to `-32099` server errors carry no retry meaning | | HTTP (RFC 9110) | most `4xx` client errors (`408` Request Timeout is an exception: the client MAY repeat the request) | `503` Service Unavailable, a temporary overload or maintenance, optionally with `Retry-After` | Frameworks often draw finer lines. gRPC's status-code guidance, for example, separates a failure the client "can retry just the failing call" (`UNAVAILABLE`), one where it "should retry at a higher level" by restarting a read-modify-write sequence (`ABORTED`), and one where it "should not retry until the system state has been explicitly fixed" (`FAILED_PRECONDITION`) — and warns that even `UNAVAILABLE` is "not always safe to retry" for non-idempotent operations. That third class is worth naming on its own: an **optimistic-concurrency conflict** is transient in a sense, but resending the same write with the same stale version fails again. The caller has to re-read, recompute and send a *new* request. ## A decision procedure 1. **Did the server reject the request before running it** (parse error, unknown method, invalid parameters, refused credentials)? Permanent — fix and report; do not resend. 2. **Is it a conflict on state the caller read earlier?** Re-run the whole read-modify-write, not the single call. 3. **Is it a declared transient condition** (receiver fault, overload, unavailable)? Candidate for resending — go to step 5. 4. **Is it a timeout or a generic internal error?** Treat the outcome as unknown — go to step 5. 5. **Could the attempt have applied its effect?** If yes and the procedure is not idempotent, resolve the outcome first (status lookup, reconciliation); otherwise resend after a delay. How long to wait between attempts, how many to make and how to spread them out are retry-policy subjects of their own. ## What servers owe their callers - **Pick the class honestly.** Returning a generic internal error for invalid input turns a permanent failure into an apparently transient one, and clients retry it pointlessly. - **Document retryability per error** in the interface contract; JSON-RPC 2.0, for one, leaves the meaning of its server-error range to each implementation. - **Say whether the work ran.** An error that guarantees nothing was applied lets the client resend freely.

  • Is a JSON-RPC 2.0 -32603 Internal error transient or permanent?
    The specification does not say. It defines -32603 only as an internal JSON-RPC error, and -32000 to -32099 as implementation-defined server errors, with no retry meaning attached. The caller must treat them as unknown in class — and possibly executed — unless the server's own documentation classifies them.
  • Why can resending on a permanent error make an incident worse?
    Every resend fails the same way, so retries multiply load on a server that is answering correctly, fill logs with noise and delay the caller's real handling — fixing input, refreshing credentials or reporting to the user. In a chain of callers that each retry, one bad request becomes many identical failures.

saying these in an interview costs you the question

  • Every error from the server deserves one retry, just in case.
  • A transient error means the call did not run, so resending is safe.
  • A SOAP 1.2 Sender fault tells the client to resend the message later.
  • JSON-RPC 2.0 defines which of its error codes are retryable.
  • A version conflict is fixed by resending the identical write.