skip to content

An engineer is designing an API endpoint that, per request, does one in-memory cache lookup, and is deciding whether to also add a synchronous call to a database in the same region versus a synchronous call to a service in a different geographic region. What rough latency numbers should they know off the top of their head to reason about this trade-off, and how do those numbers actually change the design decision?

level: seniorimportance: should knowfreq 70%

answer

  1. memory ~100ns, SSD ~100us, same-DC RTT ~0.5ms, cross-continent RTT ~100-150ms+
  2. roughly 6 orders of magnitude memory to cross-region
  3. speed of light is a hard floor for network latency
  4. sum critical-path terms, find the dominant one
  5. cache/async/replicate to avoid paying cross-region cost per request

basics

~20 s

Memory is nanosecond-fast, an SSD read is a fraction of a millisecond, a same-datacenter round trip is about half a millisecond, and a round trip to another continent is over 100ms - a cross-region call will dominate response time.

solid answer

~40 s

Rough numbers worth memorizing: L1/L2 cache reference, nanoseconds; main memory reference, ~100ns; SSD random read, ~100 microseconds (roughly 1,000x slower than memory); same-datacenter round trip, ~0.5ms; cross-country round trip, ~40-60ms; cross-continent/intercontinental round trip, 100-150ms+. Applying these: an in-memory cache lookup costs essentially nothing relative to network calls, so it's not the bottleneck. A same-region DB call adds roughly half a millisecond - often negligible for most APIs. A cross-region synchronous call adds 100ms+ on top of everything else, likely becoming the dominant, and possibly unacceptable, component of total latency, especially if it happens on the critical path for every request rather than being cached, batched, or made asynchronous.

go deeper

for a junior

Should know, at least roughly, that memory is much faster than disk and that network calls are the slowest of the three, even without precise numbers.

for a middle

Should recall approximate orders of magnitude for memory, SSD, and same-datacenter network round trip, and identify that a network call usually dominates end-to-end latency.

for a senior

Should recall cross-region/intercontinental round-trip figures as well, explain the speed-of-light basis for the gap, and translate the numbers directly into a design recommendation (cache/async/replicate) for a given scenario.

for a principal

Should reason about how these numbers inform organization-wide architectural decisions - multi-region data placement strategy, SLA budgeting across a request's full dependency chain, and when the operational cost of avoiding cross-region calls (replication complexity, consistency trade-offs) is or isn't worth it for a given product.

## The numbers worth memorizing The 'latency numbers every engineer should know' - popularized by Jeff Dean's widely circulated internal Google slide and later interactive visualizations like 'Latency Numbers Every Programmer Should Know' - are a small set of orders-of-magnitude figures for how long common operations take: | Operation | Rough figure | |---|---| | an L1 cache reference | (~1ns) | | a main memory reference | (~100ns) | | an SSD random read | (tens to low hundreds of microseconds) | | a round trip within the same data center | (roughly 0.5ms) | | a round trip across a continent or ocean | (100-150ms or more) | The value of memorizing these isn't precision - actual numbers shift with hardware generations and network paths - it's having an intuitive sense of the relative gaps between them, because the gaps span roughly six orders of magnitude from memory access to intercontinental network round trip, and design decisions almost always hinge on which of these regimes an operation falls into, not its exact millisecond value. ## Using them in a design decision The mechanism for using these numbers in a design decision is straightforward but easy to skip under pressure: sum the latency of every operation on a request's critical path, weighted by how many times each happens (serially or in parallel), and identify which term dominates the total. In the scenario given: - **An in-memory cache lookup** costs on the order of 100ns - completely negligible next to anything involving a network hop, so optimizing it further is very unlikely to move the needle on end-to-end latency. - **Adding a synchronous call to a database in the same region** adds roughly 0.5ms of network round-trip time plus the database's own processing time (which itself might be another 1-10ms depending on query complexity and whether it's served from the DB's own cache or requires disk I/O) - meaningful, but usually well within budget for a typical API with a 100-300ms latency target. - **Adding a synchronous call to a service in a different geographic region** is qualitatively different: 100ms+ round trip alone can consume the entire latency budget of an interactive API, before any processing happens on the other end. ## Why the gaps exist Why these gaps exist is rooted in physics and hardware, not just software design: - **Memory access** is bound by the speed of electrical signaling across a few centimeters of circuit board and is measured in nanoseconds. - **Disk/SSD access** adds mechanical or flash-controller overhead measured in microseconds. - **Network round trips** are fundamentally bound by the speed of light over physical distance plus the overhead of every hop (routers, TCP handshakes if a fresh connection, TLS negotiation) in between. A round trip from, say, the US East Coast to Europe covers roughly 6,000-11,000 km depending on route, and light in fiber travels at about two-thirds the speed of light in a vacuum, giving a theoretical minimum round-trip time well above 40ms before any processing - real-world routing, congestion, and protocol overhead push it higher still. This is a hard physical floor: no amount of engineering cleverness removes speed-of-light latency the way caching removes disk I/O. ## The trade-off and the alternatives The trade-off this creates is central to distributed system design: keeping a call synchronous and cross-region guarantees the caller has fresh, consistent data, but it means every request pays the full network latency cost on the critical path, and that cost is irreducible. The alternatives all trade something away to avoid paying it repeatedly: - **Caching the cross-region result locally** trades data freshness for speed (the cache can be stale). - **Making the cross-region call asynchronous** (fire-and-forget, or eventual consistency via an event) trades immediate consistency for lower perceived latency. - **Replicating the data into the local region entirely** trades storage cost and replication complexity for eliminating the cross-region hop altogether. None of these are free - the right choice depends on how stale the caller can tolerate the data being and how critical low latency is to the product experience. ## The failure mode The failure mode this knowledge guards against shows up constantly in real incidents: a service adds what looks like an innocuous synchronous dependency (an auth check, a feature-flag lookup, a fraud-scoring call) to another region without accounting for the network cost, and suddenly p99 latency for every request in that path balloons by 100ms+, or - worse - if that remote region has any instability, every local request now blocks on a resource with a much larger blast radius than a local dependency would have. A well-known industry example is why content delivery networks (CDNs) and multi-region active-active deployments exist at all: companies like Netflix and major cloud providers deliberately place data and compute physically close to end users specifically to avoid paying cross-region or cross-continent round-trip latency on the critical path of every request, accepting the operational complexity of running data in multiple regions as the price of keeping that ~100ms+ term out of the user-facing latency budget.

  • If cross-region network latency has a hard physical floor from the speed of light, what are the main engineering strategies to avoid paying it on every request?
    The main strategies are: caching the remote data locally so most requests are served from the near copy; replicating data into every region so reads never leave the region (accepting write-propagation complexity and eventual consistency); and moving the dependency off the synchronous critical path entirely via async messaging or event-driven updates, so the caller doesn't block waiting for the round trip.
  • Why might a database call within the same data center still add several milliseconds of latency beyond the roughly 0.5ms network round trip?
    The network round trip is only the transport cost; the database itself still needs to parse and plan the query, acquire any necessary locks, and potentially perform disk I/O if the requested data isn't already cached in the database's own memory (buffer pool), each of which can add low-single-digit to double-digit milliseconds depending on query complexity and cache hit rate.
  • How does connection reuse (keep-alive) versus establishing a fresh TCP/TLS connection per request change the latency numbers being discussed?
    A fresh connection adds extra round trips for the TCP handshake and, if encrypted, TLS negotiation - potentially doubling or tripling the effective round-trip cost for that single request - whereas a reused, already-established connection pays only the base network round-trip time, which is why connection pooling and keep-alive are standard practice for latency-sensitive synchronous calls.

It's like the difference between grabbing a tool from your own workbench (memory), walking to a shelf in your garage (SSD), driving across town to a friend's house (same-datacenter network call), and flying to another continent to borrow something (cross-region call) - no amount of walking faster changes the fact that the flight dominates your total trip time, so you either keep a copy nearby (caching) or accept you don't need it right this second (async).

saying these in an interview costs you the question

  • Has no rough sense of the order-of-magnitude gap between memory, disk, and network latency
  • Optimizes an in-memory operation while ignoring a much larger network call on the same critical path
  • Treats same-region and cross-region network calls as roughly equivalent costs
  • Doesn't recognize that cross-region latency has a physical (speed-of-light) floor that caching/replication work around rather than eliminate
  • Proposes 'just make the network faster' as a fix for cross-region latency without discussing caching, async, or replication

context