skip to content

Invocation Semantics Under Failure

A timed-out remote call may have run zero times, once or twice, so an RPC system must promise at-most-once or at-least-once invocation. Interviewers probe which calls are then safe to retry.

part ofAPI stylesoverview, primer and where to startread it →
on this pageshow

questions

6

A remote procedure call to transfer funds times out with no reply. What can the caller conclude about whether the transfer happened?

level: juniorimportance: must knowfreq 38%

answer

  1. silence is not a no
  2. three places a call can die
  3. request lost versus reply lost
  4. still running after the caller quits

basics

~20 s

Nothing certain: a timeout only says no reply arrived in time. The request may have been lost before running, the transfer may have run and its reply been lost, or it may still be running, so the outcome is unknown.

solid answer

~50 s

A timeout is a statement about the caller's clock, not about the server. Three different failures look identical from the client: the request never reached the server (the transfer ran zero times), the server ran it but the reply was lost or late (it ran once), or the server is still working and may finish after the caller gave up. If the caller already resent once, it may even have run twice. So the honest result is *unknown*, and the caller has to resolve it — ask the server for the outcome, retry only if the operation is safe to repeat, or hand it to reconciliation — instead of reporting *failed* and inviting the user to click again. Even over TCP, RFC 5531 warns, a caller that receives no reply "cannot assume that the remote procedure was not executed".

go deeper

for a junior

Recall that a timeout is the caller giving up, and name the three cases: request lost, reply lost, call still running. Say the outcome is unknown, not failed.

for a middle

Explain why each case looks identical to the client, why a resend after one of them can move the money twice, and why a reliable transport does not change that.

for a senior

Show how you turn unknown into known in a running system: a caller-chosen operation identifier, a status lookup, reconciliation for non-idempotent calls and a pending state shown to users.

for a principal

Argue for making outcome-unknown a first-class result in interface contracts across teams, so client libraries never collapse it into a plain failure that invites a blind resend.

## What a timeout actually measures A **remote procedure call** (RPC) makes a call on another machine look like a local function call: the caller passes arguments, waits, and gets a result. A **timeout** is the caller deciding it has waited long enough. It is a fact about the caller's patience, measured on the caller's clock. It says nothing direct about what the server did. That matters most for a call with a side effect, such as `transferFunds(from, to, amount)`. When the call times out, the caller wants to know one thing — did the money move? — and the timeout cannot answer it. ## Three failures that look the same From the caller's side there is only silence. Behind that silence, at least three different things can have happened: 1. **The request was lost or never delivered.** The network dropped it, or the server was down. The transfer ran **zero** times. 2. **The reply was lost or delayed.** The server received the request, moved the money, and the response was dropped or arrived after the caller stopped waiting. The transfer ran **once**. 3. **The server is still working.** It is slow — queued behind other work, waiting on a database lock, calling another service. It may finish successfully a moment after the caller gave up. If the client has already resent the request once after an earlier timeout, a fourth outcome joins them: both copies ran, and the transfer happened **twice**. | What happened | Transfer ran | What the caller saw | |---|---|---| | Request lost | 0 times | no reply | | Reply lost or late | 1 time | no reply | | Still running | 0 so far, maybe 1 later | no reply | | Earlier resend also ran | 2 times | no reply | The caller cannot tell these rows apart. That is the core of the problem: a remote call has a third outcome — **unknown** — that a local call never has. ## Why "failed" is the dangerous label The common mistake is to map a timeout to *failed*. Once the client tells its user "transfer failed", the user (or an automatic retry) sends the transfer again. In rows 2 and 3 that second request moves the money a second time. Some points that often get missed: - **A reliable transport does not remove the ambiguity.** TCP stops lost or reordered packets within one connection, but it cannot tell the client whether the server process ran the procedure before the connection broke or the machine crashed. RFC 5531, the ONC RPC specification, says that over TCP a received reply lets the caller infer the procedure ran exactly once, "but if it receives no reply message, it cannot assume that the remote procedure was not executed". - **A longer timeout only moves the line.** It makes row 3 rarer; it does nothing about rows 1 and 2. - **Frameworks say the same thing in their own words.** gRPC's status-code documentation, for example, warns that a deadline-exceeded status may be returned for a state-changing operation even when the operation completed successfully. - **Not every error is ambiguous.** An explicit error that the server returns *before* dispatching the call — a malformed request, an unknown method, refused credentials — tells the caller the work did not run. The ambiguity belongs to silence and to errors raised partway through processing. ## Turning unknown into known A well-built caller treats *unknown* as its own result and resolves it deliberately: - **Look the outcome up.** If the call carried a caller-chosen operation identifier (for example a transfer reference generated before the first send), the caller can ask the server "what happened to transfer T-81?" and act on the answer. - **Retry only what is safe to repeat.** A read, or a write that sets an absolute state, can be resent. A non-idempotent transfer can be resent only if the server recognises the repeat and does not execute it twice. - **Reconcile.** For money and other irreversible effects, a background job compares both sides' records and settles the difference, rather than guessing at call time. - **Tell the user the truth.** A "pending — we are confirming" state is better than a false "failed" that invites a double charge. Which invocation guarantee the RPC system offers — whether it resends automatically, and whether the server filters duplicates — decides which of these tools the caller can lean on. The ambiguity itself never goes away; the design decides who resolves it and how.

  • Would running the call over a reliable transport such as TCP remove the ambiguity?
    No. TCP stops lost and reordered packets inside one connection, but it cannot tell the client whether the server ran the procedure before the connection broke or the server crashed. RFC 5531 notes that over TCP a reply implies exactly one execution, yet with no reply the caller still cannot assume the procedure did not run.
  • How should the caller resolve the unknown outcome of the funds transfer?
    Ask before acting again. If the call carried a caller-chosen transfer identifier, query the server for that transfer's status and report what it says. If the server cannot be asked, and it does not filter repeated requests, the transfer goes to reconciliation rather than an automatic resend, and the user sees it as pending.
  • Is a fast error reply just as ambiguous as a timeout?
    Not always. An error the server returns before dispatching the call — malformed request, unknown method, refused credentials — tells the caller the work did not run. A generic internal error raised while the handler was running can still leave effects behind, so the caller needs to know at which stage the error was produced.

Posting a letter that asks your bank to move money and hearing nothing back: the letter may be lost, the bank may have moved the money and its confirmation got lost, or it sits in a pile. Silence means check, not that nothing happened.

saying these in an interview costs you the question

  • A timeout means the server did not process the request.
  • Over TCP, a timed-out call is guaranteed not to have run.
  • Just retry the transfer; the worst case is another error.
  • Raising the timeout value removes the ambiguity.
  • If the server had run it, the caller would have received a reply.
open as a page

In an RPC system, what distinguishes maybe, at-most-once and at-least-once invocation semantics, and what does each need from client and server?

level: middleimportance: must knowfreq 30%

basics

~20 s

Maybe sends once and never resends, so the call ran zero or one times. At-least-once resends until a reply arrives, so it may run repeatedly. At-most-once also resends, but the server spots duplicates by request identifier and replays its saved reply.

open as a page

After a remote procedure call fails without a reply, which procedures may the client safely re-invoke, and how do you decide?

level: middleimportance: must knowfreq 32%

basics

~20 s

A call is safe to re-invoke when two runs leave the same intended state as one: reads and absolute writes like 'set address to X'. 'Add 10 points', 'create order' or 'send email' are unsafe unless the server filters repeats.

open as a page

When a remote procedure call returns an error, how do you tell a transient failure worth retrying from a permanent one?

level: middleimportance: should knowfreq 24%

basics

~20 s

Ask whether the identical request could ever succeed. Transient errors come from processing conditions — overload, an unreachable upstream — and may clear; permanent ones come from the request itself — bad arguments, unknown method, refused credentials — and recur until it changes.

open as a page

An RPC framework advertises exactly-once invocation. What is it really combining to deliver that, and where does the guarantee quietly break?

level: seniorimportance: should knowfreq 18%

basics

~20 s

No protocol makes a lossy network execute a call exactly once. 'Exactly-once' is at-least-once resending plus execution that absorbs repeats — an idempotent procedure or server-side duplicate filtering — and it breaks wherever that filter forgets or misses a repeat.

open as a page

A client crashes and restarts while its remote procedure calls are still running on the server. What are these orphaned calls, and how can an RPC system deal with them?

level: seniorimportance: nice to knowfreq 9%

basics

~20 s

Orphans are server executions whose caller crashed or stopped waiting. They waste resources, hold locks and can apply effects the restarted client does not expect. Systems cancel them on request, abort them by client epoch, or let them expire.

open as a page