skip to content

What does 'boundary-crossing cost' mean, and how should the cost of a particular boundary change the interface you design across it?

level: middleimportance: must knowfreq 58%

answer

  1. ns → µs → ms: same call, different price
  2. cost = latency + serialization + partial failure + versioning
  3. expensive boundary ⇒ coarse, batched, intent-shaped calls
  4. chatty / N+1 / distributed monolith
  5. 'don't distribute your objects'

basics

~20 s

Crossing a boundary costs something: a function call is nearly free, a call to another process costs serialization, network latency and possible failure. Cheap boundaries can have fine-grained, chatty interfaces; expensive ones need coarse-grained calls that return everything needed at once.

solid answer

~60 s

Boundary-crossing cost is what you pay each time control or data moves across a separation line. It rises by orders of magnitude along a spectrum: same-object call (~nanoseconds) → in-process module call, possibly with data mapping → cross-thread/queue hop → cross-process on one host → network call to another service (sub-millisecond to tens of milliseconds, plus partial failure, retries, timeouts, versioning, and observability gaps). The cost is not only latency: it includes serialization, schema/version compatibility, error semantics (a remote call can time out with an unknown outcome, so operations should be idempotent), and cognitive/operational overhead. Interface granularity must match: expensive boundaries need coarse, intent-revealing, batched operations that return a complete result — never property-at-a-time access or N+1 loops. Cheap boundaries can afford fine-grained interfaces. Crucially, the fact that a call *can* be made local later does not mean a remote-shaped interface should be designed as if it were local — that is the core of the 'distributed objects' failure and Fowler's First Law of Distributed Objects: don't distribute your objects.

code

pseudocode · 10 lines
pseudocode
// Chatty across an expensive (remote) boundary: 1 + 3N round trips
ids = catalog.listIds(query)
for (id in ids) {
  p = catalog.getName(id)
  q = catalog.getPrice(id)
  r = catalog.getStock(id)
}

// Coarse + batched: 1 round trip, complete result, retry-safe
page = catalog.searchProducts(query, pageSize = 50)  // returns name+price+stock per item

go deeper

for a junior

Say that a local call is cheap and a network call is expensive, and that expensive boundaries need fewer, bigger calls.

for a middle

Quantify the spectrum, name chattiness/N+1, and give concrete mitigations: batching, coarse documents, caching, async.

for a senior

Add the non-latency costs — partial failure, lost transactions, versioning, tracing — and design idempotent, tolerant contracts; discuss moving the line instead of optimizing a bad one.

for a principal

Reason about placing boundaries where the interaction graph is thinnest, keeping logical boundaries strong while deferring physical ones, and the organizational cost of every distributed contract.

## The cost spectrum Every boundary has a *crossing cost* — what you pay each time a call or a piece of data goes over it. Rough orders of magnitude (hardware-dependent, but the ratios matter more than the numbers): | Crossing | Typical latency | Extra cost | |---|---|---| | Direct call inside a module | ~1 ns | none | | Call through an interface, plus DTO mapping | tens–hundreds of ns | allocation, mapping code | | In-process queue / thread hop | µs | ordering, backpressure, concurrency bugs | | Same-host inter-process (socket, IPC) | tens–hundreds of µs | serialization, process lifecycle | | Network call, same datacenter | 0.5–2 ms typical | serialization, partial failure, retries, TLS | | Cross-region network call | 30–150 ms | all of the above plus latency you cannot optimize away | So the same logical decomposition can be nearly free or catastrophically expensive depending on where you place it physically. ## Cost is more than latency Designers under-count boundary cost because they only count milliseconds. The full bill: - **Serialization / marshalling** — CPU and allocation on both sides; data must have a wire representation. - **Schema and version compatibility** — the two sides now deploy separately, so every change must be backward/forward compatible for at least one release window (expand-then-contract / tolerant reader). - **Partial failure** — the defining property of a remote boundary. A local call either returns or throws; a remote call can *time out with an unknown outcome*. That forces timeouts, retries, idempotency keys, circuit breakers, and de-duplication that a local call never needs. - **Loss of transactions** — a boundary that is also a process/database boundary usually kills the ability to commit both sides atomically. You move to sagas, outbox patterns, and eventual consistency, plus compensating actions. - **Observability gaps** — no shared stack trace; you need correlation IDs and distributed tracing to reconstruct one logical operation. - **Cognitive and process cost** — two repos, two deploy pipelines, cross-team coordination for a contract change. The **fallacies of distributed computing** (Deutsch/Gosling) name the assumptions that break here: the network is reliable, latency is zero, bandwidth is infinite, the network is secure, topology doesn't change, there is one administrator, transport cost is zero, the network is homogeneous. ## Granularity must match cost The design rule: **the more expensive the boundary, the coarser and more intent-revealing the interface.** - Cheap (in-process): fine-grained accessors are tolerable; readability wins. - Expensive (remote): each call should do a *whole unit of business work* and return everything the caller needs. Anti-patterns that appear when granularity is mismatched: - **Chatty interface** — the caller makes many small calls to complete one task. `getName()`, `getAddress()`, `getStatus()` over the network. - **N+1 across a boundary** — fetch a list of 200 ids, then call once per id. Fix with batch operations (`getAll(ids)`) or by pushing the join to the owning side. - **Distributed monolith** — services split physically but still requiring synchronous fan-out to every peer for any request; you pay all the distribution cost and get none of the independence. - **Anemic remote CRUD** — exposing `setField` over the wire, which puts the business rule in the caller and makes the contract impossible to evolve. Martin Fowler's **First Law of Distributed Object Design**: *don't distribute your objects.* The 1990s distributed-object platforms (CORBA, DCOM, early EJB remote interfaces) tried to make a remote call look exactly like a local one; the result was chatty designs whose failure and latency characteristics were invisible until production. The corollary is that remote interfaces should be *deliberately different in shape* from local ones — coarse, message/document-oriented, explicit about failure. ## Mitigation strategies (when the boundary must be expensive) 1. **Batch and bulk** — accept collections, return collections. 2. **Coarse documents** — return a complete view (`OrderSummary` with lines, totals, and shipping) rather than forcing follow-up calls. Note the trade-off: bigger payloads and more coupling to the caller's needs; BFF (backend-for-frontend) or GraphQL-style shaping exist for this. 3. **Caching / read models** — replicate the data you read constantly onto your own side, accepting staleness. This trades consistency for latency and availability. 4. **Asynchronous messaging / events** — remove the call from the request path entirely; the caller no longer waits or fails when the peer is down. Cost: eventual consistency and duplicate delivery. 5. **Idempotency** — make retries safe with client-supplied idempotency keys, so 'unknown outcome' can be resolved by retrying. 6. **Timeouts, bulkheads, circuit breakers** — bound the damage of a slow peer so its latency doesn't become yours. 7. **Move the line** — the best fix for a hot, chatty boundary is often to merge the two sides back together. Boundary placement should follow the *interaction* graph: put the line where traffic is thinnest. ## Logical boundary, cheap crossing An important nuance: keeping a *strong logical* boundary while keeping the crossing *cheap* is a legitimate and often optimal design — a modular monolith. You get information hiding, one-way dependencies, and independent reasoning at in-process cost, and you retain the option to promote the boundary to a process boundary later if a real driver (independent scaling, isolation, team autonomy) appears. Designing the interface as if it might one day be remote (no shared mutable state, no cross-module transactions, coarse operations, its own data) makes that promotion cheap without paying network cost today.

  • If in-process boundaries are so much cheaper, why ever pay for a process boundary?
    For properties you cannot get in-process: independent deployment and release cadence, independent scaling of a hot component, fault and resource isolation (one component's memory leak or crash doesn't take the rest down), separate security/compliance domains, and team autonomy. Pay when you need one of those, not by default.
  • A remote call times out. Why is that fundamentally harder than a local exception?
    Because the outcome is unknown — the peer may have completed the work, partially completed it, or never received it. You cannot safely assume failure, so you need idempotent operations (deduplication keys), a reconciliation path, or explicit compensation.

Talking to a colleague at the next desk versus mailing them letters. At the next desk you can ask twenty tiny questions. By post, each question costs days and letters get lost, so you write one letter containing everything you need and a plan for what to do if the reply never arrives. Same conversation, completely different shape.

context