Using Waldo et al.'s 1994 critique of RPC location transparency, what goes wrong when a team designs local objects and decides later which become remote?
answer
- A Note on Distributed Computing
- four differences, not one
- the interface must change
- failure and concurrency stay
basics
~20 sWaldo et al. argue local and remote objects differ in latency, memory access, partial failure and concurrency, so an interface designed for local use cannot simply be made remote later; failure and concurrency must be designed in from the start.
solid answer
~50 sWaldo et al.'s *A Note on Distributed Computing* (1994) argues against exactly this plan. Making an object remote changes four things its local interface never had to express: **latency** - a chatty, fine-grained interface that was free in-process now costs a round trip per call; **memory access** - arguments are copied, not shared, so reference and side-effect semantics stop holding; **partial failure** - the server or the network can fail while the caller lives, and the caller may not know what happened; and **concurrency** - the remote object is called by independent clients at once and must handle that itself. Faster networks and a pointer-free programming model can tame the first two; the last two must be visible in the interface and the design - coarse-grained operations, explicit failure outcomes, concurrency control. The paper's verdict is that systems which paper over the difference fail basic requirements of robustness and reliability.
go deeper
Recall the four differences between local and remote calls named in A Note on Distributed Computing: latency, memory access, partial failure and concurrency.
Explain each difference with a concrete symptom: chatty interfaces, copied arguments, a callee that fails while the caller lives, concurrent callers.
Explain why partial failure and concurrency cannot be hidden by a better stub, and show how they change an interface: coarse operations, explicit outcomes, time limits.
Discuss where an organisation should draw remote boundaries and how much transparency its frameworks offer, separating transparent addressing from transparent semantics.
## The plan being critiqued **Location transparency** is the promise that a caller need not know whether an object is local or remote: the stub makes both look the same. Taken to its end, it suggests a development plan - write the system as ordinary local objects, get it working, then decide which objects run on which machines, letting generated stubs absorb the difference. Waldo et al.'s *A Note on Distributed Computing* (1994) argues that this plan fails. Its abstract states the thesis: objects that interact in a distributed system must be treated in ways intrinsically different from objects in a single address space, because programmers must be aware of latency, work with a different model of memory access, and take concurrency and partial failure into account; systems that paper over the distinction fail basic requirements of robustness and reliability. ## The four differences | Difference | Local call | Remote call | What it forces in the design | |---|---|---|---| | **Latency** | Nanoseconds to microseconds | Orders of magnitude slower, plus network variance | Coarse-grained operations; fewer round trips | | **Memory access** | Shared address space; references and side effects work | No shared memory; arguments are copied | Explicit results instead of mutated arguments | | **Partial failure** | The call returns or the whole process dies | Callee or network can fail while the caller lives | Failure outcomes in the interface; timeouts | | **Concurrency** | The program controls its own threads | Independent clients call at once, with no common control | The remote object handles concurrent callers itself | RFC 5531, writing about ONC RPC, lists overlapping differences in plain terms: failures of the remote server or network must be handled, the server has no access to the client's address space, and remote procedures usually run "one or more orders of magnitude slower" than local ones. ## Why a better stub does not fix it - **Latency** can be reduced by faster networks and placement, and its cost hidden by caching - but an interface that makes a round trip per getter stays slow by design. - **Memory access** can be tamed by a programming model that never passes raw references, making copies the norm everywhere. - **Partial failure** cannot be engineered away. In one process a call either returns or the process is gone; across a network the caller can be alive while the callee, or the link, is not - and often cannot tell which. Mapping that onto an ordinary exception does not tell the caller what to do. - **Concurrency** is equally structural: a remote object is shared by callers that do not coordinate with each other. Because the last two differences live in the system, not in the stub, an interface written as if they did not exist has no place to express them. Adding them later changes the interface - and every caller. ## What designing for distribution looks like 1. **Coarse-grained operations** that do a meaningful unit of work per call, returning what the caller needs in one round trip. 2. **Explicit outcomes** in the signature or contract: what errors a call can return, and which calls may safely be repeated. What a timed-out call means for the remote work is its own subject. 3. **Time limits** on every call, so a caller is never parked forever on a silent peer. 4. **Concurrency control** inside the remote object: no assumption that one caller owns it. 5. **Visible remoteness** where it helps - asynchronous or streaming call shapes instead of pretending every call is an instant local one. ## What the critique does not say - It does not say RPC is wrong. Stubs and remote calls remain useful; the claim is that remote interfaces must be designed as remote. - It does not say transparency of **addressing** is wrong. Not caring which host serves a call is different from pretending the call cannot fail. - It does not say local and remote code must share nothing - only that the boundary between them is a design decision made up front, not a deployment detail made later. ## Common confusions - "Latency is the only difference" - it is the most visible, and the most fixable. - "A fast enough network makes remote calls local" - it shrinks latency; partial failure and concurrency remain. - "Turning network errors into exceptions hides partial failure" - it only renames it; the caller still does not know what happened on the other side.
- Why does a chatty, fine-grained interface hurt far more once it is remote?Each call becomes a round trip, and RFC 5531 notes remote procedures are usually one or more orders of magnitude slower than local ones. A loop of getters that cost nothing in-process multiplies that delay. The fix is a coarse-grained operation that returns everything the caller needs in one round trip.
- What does partial failure mean that a single-process program never faces?In one process, a call returns or the whole program is gone. Across a network, the server or the link can fail while the caller keeps running, and the caller often cannot tell whether the server crashed, is slow, or the reply was lost. What that means for retrying the call is a separate question.
- Is location transparency ever a reasonable goal?Transparency of addressing is: callers need not know which host or replica serves a call. What Waldo et al.'s critique targets is transparency of semantics - pretending a remote call has local latency, memory sharing, failure behaviour and concurrency. An interface designed for distribution can still hide where it runs.
saying these in an interview costs you the question
- Latency is the only real difference between local and remote calls.
- A fast enough network makes remote calls equivalent to local ones.
- Wrapping network errors in an ordinary exception fully hides partial failure.
- An interface can be designed locally and made remote later without changing its signature.
- Waldo et al. argued that remote procedure calls should never be used.