An RPC framework advertises exactly-once invocation. What is it really combining to deliver that, and where does the guarantee quietly break?
answer
- two halves, not one mechanism
- resend plus absorb the repeat
- the filter's memory can fail
- effects beyond the server
basics
~20 sNo protocol makes a lossy network execute a call exactly once. 'Exactly-once' is at-least-once resending plus execution that absorbs repeats — an idempotent procedure or server-side duplicate filtering — and it breaks wherever that filter forgets or misses a repeat.
solid answer
~50 sOver a network that can lose messages, the client cannot tell a lost request from a lost reply, so it must resend, and resending alone means duplicates. What a framework calls exactly-once is two parts: **at-least-once delivery** (keep resending until a reply) and **at-most-once effect** (an idempotent procedure, or a server that recognises a repeated request identifier and replays its stored reply). The promise holds only while both parts hold. It breaks when the server's record of executed requests is lost in a crash or restart, or evicted before the client's last resend; when a resend lands on a replica that does not share that record; when a restarted client resends under a new identifier; and when the procedure has effects beyond the server — another service, an email — that the filter cannot reach. RFC 5531 promises only 'some degree' of at-most-once from remembered transaction identifiers.
code
pseudocode · 9 lineson request(req):
entry = executed.lookup(req.request_id)
if entry is not none:
return entry.reply // repeat: replay, do not re-run
reply = procedure(req.args) // the side effect happens here
// a crash here leaves the effect done but unrecorded,
// so the client's resend finds no entry and runs it again
executed.store(req.request_id, reply)
return replygo deeper
Recall that no network can promise exactly-once by itself; what is sold under that name is resending plus a way to make repeats harmless.
Explain the two halves: the client's resend loop and the server's duplicate filter or idempotent procedure, and why the client must reuse one identifier.
List where the filter breaks in production — restarts, eviction, replicas, client restarts, non-atomic recording, concurrent duplicates, downstream effects — and how you close each.
Decide where an organisation should pay for effectively-once — durable shared duplicate records versus idempotent operation design — and how to state the real guarantee in contracts.
## Why the network cannot do it alone An **RPC framework** carries a call from a client to a server and a reply back. When the reply does not arrive, the client cannot tell whether the request was lost (the call never ran) or the reply was lost (the call ran). It has two choices: give up, which risks **zero** executions, or resend, which risks **two**. No choice made by the client alone produces exactly one, because the information it would need — did the server run it? — is exactly what got lost. So "exactly-once invocation" is never a transport property. It is a promise built from two halves, and you can always ask which two. ## The two halves | Half | What it does | Who provides it | |---|---|---| | **At-least-once delivery** | Resends until some reply arrives, so the call is not silently dropped | the client (its retry loop) | | **At-most-once effect** | Makes the second and later executions harmless or skips them | the server: an idempotent procedure, or a duplicate filter keyed by request identifier | With both, the *observable effect* happens once, even though the request may have crossed the network several times. A more honest name is **effectively once**. The filter half is the classic **duplicate-request cache**: the client reuses one request identifier for all resends of a call, and the server remembers executed identifiers together with their replies, answering a repeat from memory. RFC 5531 (ONC RPC) describes this and calls the result "some degree of execute-at-most-once semantics" — the hedge is deliberate. ## Where the filter fails The guarantee is only as good as the filter's memory and reach. It quietly breaks when: 1. **The server crashes or restarts** and its record of executed identifiers was in memory. A resend after the restart looks new and runs again. 2. **The record is evicted too early.** If entries expire after two minutes and a client resends after three, the repeat runs. 3. **The resend reaches another replica** that keeps its own record. Behind a load balancer, attempt one and attempt two can hit different servers. 4. **The client restarts and loses its identifier.** A resend under a fresh identifier is, to the server, a different call. 5. **Execution and recording are not atomic.** If the server runs the procedure, then crashes before storing the identifier and reply, the resend runs it again (see the snippet). 6. **A duplicate arrives while the original is still running.** Without an *in progress* marker, both copies miss the lookup and both execute. 7. **The effect leaves the server.** If the procedure calls another service or sends an email, filtering at this server does nothing about repeats there; each downstream hop needs its own guarantee. ## Making it hold for one call - Choose the **request identifier on the client** before the first send, and keep it durable if a client restart must not lose it. - Store the identifier and the reply **in the same durable transaction** as the procedure's own effect, so neither can exist without the other. - Record **in progress** before executing, so a concurrent duplicate waits or is told to retry later. - Keep records **at least as long as the longest retry horizon** a client may use, and share them across every replica that can serve the call. - Prefer procedures that are **naturally idempotent** — set an absolute value, apply only if a version still matches — so a missed duplicate does no harm. The storage design for such records — key scope, retention and concurrent-duplicate handling — is a deduplication subject of its own; here the point is that the RPC system's guarantee depends on it. ## Reading a vendor or teammate's claim When someone says a framework gives exactly-once invocation, ask: - Where is the duplicate record kept, and does it survive a server restart? - How long are entries kept, compared with the client's longest resend window? - Is the record shared by every replica, or per process? - What happens to effects outside the server — downstream calls, messages, emails? If the answers are "in memory", "a while", "per process" and "not covered", the claim is at-least-once with a best-effort filter.
- Why is recording the request identifier before running the procedure not a fix either?It moves the crash window instead of closing it. If the server records 'done' and crashes before the effect commits, a resend is answered as a repeat and the effect never happens — at-most-once with a lost call, not exactly-once. Only recording the identifier and reply atomically with the effect, in one durable transaction, closes the window.
- A duplicate arrives while the original call is still executing. What should the server do?Not start a second execution. The filter needs an in-progress state for the identifier: the duplicate either waits for the original to finish and receives its reply, or gets a 'still in progress, retry later' answer. Without that state, two concurrent copies both miss the lookup and both run.
saying these in an interview costs you the question
- TCP delivers each request exactly once, so RPC over it is exactly-once.
- An in-memory reply cache keeps exactly-once across server restarts.
- Exactly-once means the client never has to resend.
- Filtering repeats at one server also stops repeated downstream calls and emails.
- A fresh request identifier per resend is fine if the body is identical.