skip to content

A page render needs 60 values from a tier in another zone within a 50 ms budget; how do you check it fits?

level: middleimportance: must knowfreq 64%

answer

  1. budget first, then trips
  2. measure one crossing at that placement
  3. high percentile, not the mean
  4. count nodes, not operations

basics

~20 s

Measure one round trip from the real caller to the real tier at a high percentile, multiply by the crossings the path makes, and compare with the share of the 50 ms the tier gets. Sixty sequential crossings will not fit.

solid answer

~50 s

Start from the budget, not the store. A 50 ms promise is usually a high-percentile promise, and the tier only gets a share of it — the rest belongs to the database, the template and serialisation. Measure one crossing at that exact placement from the caller's side, take a high percentile rather than the mean, and multiply by the trip count the path actually makes. At around 1 ms per crossing across zones, 60 sequential calls are about 60 ms: over the whole budget before anything else has run. Then spend the budget down by removing crossings — fewer keys, one operation taking many keys, or a run sent without waiting for each reply. Note that on a tier whose keyspace is split across nodes, one operation over many keys may become one crossing per node holding some of them, so count nodes rather than assuming one.

go deeper

for a junior

The arithmetic is simple: how many times does this request cross the network, and how long does one crossing take? Learn to ask both questions before guessing whether something fits.

for a middle

Do the multiplication with a measured number and state the placement it came from. Know that a run sent without waiting collapses the waiting but gives you no atomicity, and that the values still have to fit in the reply buffer at each end.

for a senior

Make the partition map part of the arithmetic — the crossing count is the number of nodes the caller talks to, not the number of operations in the code — and build the budget on a high percentile because that is what the promise is made at.

for a principal

Decide what share of a user-facing deadline this dependency is entitled to before any team writes the path, and make the placement a stated requirement rather than something discovered after deployment.

## What the budget is actually made of A 50 ms request budget is a promise about **the caller's wall time**, and three things have to be pinned down before any arithmetic is honest: - **Which statistic.** A budget stated as a mean and a budget stated at a high percentile are different promises, and paths that meet the first routinely miss the second. Assume the high percentile unless someone says otherwise. - **Whose share.** The tier is one dependency among several. If the page also queries a durable engine, renders a template and serialises a response, the tier's allowance might be 10 ms of the 50, not 50. - **Which placement.** "Another zone" is the input that turns a trip count into milliseconds. Same host, same zone, another zone and another region differ by orders of magnitude. ## Getting the per-trip number honestly 1. **Measure at the caller.** The store's own accounting reports service time, which is microseconds and is not what the budget is spent on. What you need is the caller's wall time for one trivial single-key call. 2. **Measure from the real host to the real tier.** A number taken from a developer machine, from a different zone, or against a locally running store answers a different question. 3. **Take a high percentile, not the mean.** A budget is missed at the tail, and the tail of a network crossing is where retransmits, queueing and scheduling delay live. 4. **Redo it per placement.** One measurement is not portable across a zone boundary. ## The multiplication With one crossing measured at roughly 1 ms, the same 60 values cost wildly different amounts depending only on how they are asked for: | Access shape | Crossings | Approximate wall time | Share of a 50 ms budget | |---|---|---|---| | 60 sequential single-key reads | 60 | ~60 ms | over the whole budget | | One operation taking all 60 keys, single-node tier | 1 | ~1 ms plus reply transfer | a few per cent | | One run of 60 sent without waiting for replies | 1 | ~1 ms plus reply transfer | a few per cent | | Keys spread over 4 nodes, asked per node | 4 | ~4 ms | under a tenth | The store's service time — 60 operations at a handful of microseconds — is under half a millisecond in every row. It never appears in the answer. ## What the partition map does to the trip count On a tier whose keyspace is split across nodes, "one operation over many keys" is not automatically one crossing: - Some stores refuse an operation whose keys are not all served by one node, and the caller must group the keys itself. - Some client libraries accept it and quietly split it into one request per node, so the trip count equals the number of distinct nodes involved. - Some deployments put a proxy in front, so the caller makes one crossing and the fan-out happens on the far side — cheaper for the caller, but the proxy hop is itself a crossing. So the honest input to the arithmetic is **the number of crossings the caller makes**, which you get by asking how the keys map to nodes and whether the caller knows that map — not by counting the operations in your source code. ## Spending the budget down In rough order of how much they buy: - **Drop keys the path does not use.** Sixty values fetched and forty rendered is forty crossings of pure waste. - **Ask for many keys in one operation**, where the store offers one and the keys can be served together. - **Send a run of operations without waiting for each reply**, which collapses the waiting even where no multi-key operation exists. It buys latency only — no atomicity, no isolation, and nothing is rolled back if one operation in the run fails. - **Move the decision closer to the data**, so one answer crosses the network instead of sixty inputs. - **Move the tier into the caller's zone**, which shrinks every crossing but removes none. ## What the arithmetic assumes, and where stores differ - **An operation taking many keys may not exist.** Some stores in this class answer only single-key reads and writes, and then the only lever is sending a run without waiting. - **Reply size becomes a real term once trips collapse.** Sixty values in one reply must be buffered at both ends; a few hundred are fine, a few hundred thousand are a different problem. - **The execution model does not change this answer.** Whether the store executes one operation at a time or serves requests from a thread pool matters for who is delayed behind an expensive call; the term being multiplied here is negligible under either model. - **The measured crossing is a property of your deployment**, not of the store, and it changes when anything in the path between the two changes.

  • The tier is moved into the caller's own zone and a crossing now measures 0.3 ms. Is the 60-crossing design fine?
    It now fits — about 18 ms of a 50 ms budget — but it is fragile rather than correct. Sixty crossings is sixty chances of a slow tail, sixty dependencies on one network path, and a design that breaks again the moment the placement changes or the keyspace is split. Collapsing the crossings costs little and removes the whole class of risk.
  • Why is the mean round-trip time the wrong number to build the budget from?
    Because a request budget is normally a promise at a high percentile, and multiplying a mean by sixty produces a number that describes a request nobody complained about. The tail of one crossing is where queueing and retransmits live, and a path making many crossings is likelier to include at least one of them, so the path's tail is worse than its per-call tail suggests.

saying these in an interview costs you the question

  • Builds the budget on the mean round-trip time and ignores the tail
  • Assumes one operation over many keys always costs exactly one crossing
  • Quotes the store's advertised operations per second as a latency answer
  • Gives the tier the whole 50 ms and forgets the rest of the path
  • Measures a crossing from a developer machine and uses it for production