A client crashes and restarts while its remote procedure calls are still running on the server. What are these orphaned calls, and how can an RPC system deal with them?
answer
- work with no one waiting
- locks held, effects applied
- kill, epoch or lease
- killing does not undo
basics
~20 sOrphans are server executions whose caller crashed or stopped waiting. They waste resources, hold locks and can apply effects the restarted client does not expect. Systems cancel them on request, abort them by client epoch, or let them expire.
solid answer
~50 sAn **orphan** is a remote execution with no live caller: the client crashed, rebooted or abandoned the call, but the server keeps working. Orphans burn CPU and connections, hold locks or reservations that block other callers, may send replies to a client that no longer recognises them, and — worst — can commit an effect after the restarted client has already resent the call or decided differently, producing a duplicate or a contradiction. The main strategies: the client records outstanding calls durably and, after restarting, asks the servers to cancel them; the client announces a new **epoch** (a boot counter) and servers abort work tagged with an older epoch and discard its late replies; or each call runs under a **lease** that expires unless the caller renews it. None of them can undo an effect that already committed, so the restarted client still treats those calls as outcome unknown.
go deeper
Recall what an orphan is: server work still running for a caller that crashed, rebooted or stopped waiting.
Explain the harm beyond wasted CPU — held locks, late replies, effects committed after the client moved on — and why nested calls spread it.
Compare cancellation logs, epochs and leases, say what each costs, and walk through what a restarted client must do before resending anything.
Set a fleet-wide rule that every call carries a bound the server enforces and every caller incarnation is identifiable, and weigh that cost against tolerating orphans.
## What makes a call an orphan In a **remote procedure call**, a client asks a server to run a procedure and waits for the reply. An **orphan** is the server-side execution of a call whose caller is no longer waiting for it. The server does not know this: from its side, a request arrived and is being processed. Orphans appear when: - the **client process crashes** mid-call; - the **client machine reboots** and its new incarnation knows nothing of the old one's calls; - the client **abandons** the call — it timed out and moved on without telling the server; - a **network partition** cuts the client off, so the reply has nowhere to go. Nested calls multiply the problem: if the orphan itself called other servers, those calls become orphans too, two hops from the dead caller. ## Why orphans hurt - **Wasted capacity.** CPU, memory, threads and database connections go to work nobody will read. - **Held resources.** An orphan holding a row lock or a seat reservation blocks live callers until it finishes or is killed. - **Confusing replies.** A late reply can reach a restarted client that reuses the same connection or request identifiers and be matched to the wrong call. - **Contradicting effects.** The restarted client sees an unknown outcome and may resend, cancel, or take a different path. The orphan then commits its own effect on top — a duplicate transfer, or an order placed after the user chose to abandon it. ## Strategies | Strategy | Mechanism | Cost | Limit | |---|---|---|---| | **Explicit cancellation** | The client logs each outstanding call durably; after restarting, it asks each server to cancel them | a durable log write per call | the log must survive the crash; cancellation can arrive after the effect committed | | **Epoch (boot counter)** | Every call carries the client's epoch; after a reboot the client announces a higher epoch, and servers abort work and drop late replies from older epochs | one counter per client, a broadcast on reboot | servers must be told; aborting mid-procedure needs clean rollback | | **Lease / expiry** | Each call is granted a time limit; the client must renew it for long calls, and the server stops work whose lease lapsed | renewal traffic | the server needs a safe stopping point; clock behaviour matters | | **Ignore and tolerate** | Let orphans finish; make procedures idempotent and discard late replies by checking the request identifier and epoch | none at run time | resources are still wasted and locks still held | A deadline that travels with the call bounds an orphan's life in much the same way as a lease. How deadlines are set and propagated across hops is a resilience-pattern subject of its own; for orphans, the point is that some bound must reach the server. ## The restarted client's view A client that crashed with calls in flight should, on restart: 1. Treat **every** call that was outstanding as outcome unknown — some may have run, some may still be running. 2. **Announce itself** as a new incarnation (a new epoch, or cancellation requests from its durable log), so servers stop or disown the old work. 3. **Resolve each outcome** — query the server for the result by the call's identifier, or reconcile — before resending anything that is not idempotent. 4. **Reject late replies** whose epoch or identifier belongs to the old incarnation. ## What none of them fixes - **Committed effects stay committed.** Killing an orphan stops further work; it does not roll back a write that already happened. - **Kills must cascade.** Stopping the first server's orphan leaves its own downstream calls running unless the abort or the bound travels with them. - **Orphans and duplicates meet.** An orphan that finishes after the client resent the same non-idempotent call is the classic double effect; at-most-once filtering by request identifier is what catches it.
- Why is killing an orphan not enough to make the system consistent?Because the orphan may already have committed part or all of its effect before the kill arrived. Killing stops further work and frees resources, but the restarted client still faces an unknown outcome for that call and must query or reconcile before resending anything that is not idempotent.
- What happens to calls the orphan itself made to other servers?They become orphans too, two hops from the dead caller. A kill or epoch change has to cascade down the call tree, or each nested server must bound its own work, for example with a lease or a deadline that travels with the call; otherwise the deeper servers keep working for nobody.
saying these in an interview costs you the question
- When a client dies, the server notices at once and stops the call.
- Killing an orphan undoes whatever it already wrote.
- Orphans only waste CPU; they cannot change any data.
- After a reboot, the client can safely resend all its unfinished calls.
- Orphans arise only from client crashes, never from abandoned calls.