How would you set and enforce a per-request trip budget for a shared in-memory tier across many teams?
answer
- count trips per path, not calls per second
- the budget belongs to the path
- placement is a policy, not an accident
- removed crossings beat tuned ones
basics
~20 sState, per request path, how many crossings it may make and at what placement, derived from the user-facing deadline and a measured per-crossing cost. Then measure trips per path, not calls per second, and treat a removed crossing as better than a tuned one.
solid answer
~40 sA latency target alone is unenforceable, because it names neither input that decides the answer. A usable contract states, per request path, the maximum number of crossings and the required placement, both derived from the user-facing deadline: take the deadline, subtract what the other dependencies are allowed, and divide the remainder by a measured per-crossing cost at that placement. Enforcement is a measurement change — report the pair (wall time for the path, number of calls made) rather than the tier's operations per second, and reject any path whose trip count scales with the size of a collection. The review bias should be toward removing a crossing rather than shortening it: a crossing not made costs nothing and cannot fail. The contract prescribes no mechanism and says nothing about what belongs in the tier.
go deeper
The useful habit is to count how many times your code calls this tier while serving one request, and to notice when that count depends on the size of a list.
Be able to derive an allowance rather than quote one: deadline, minus the other dependencies' shares, divided by a measured per-crossing cost at the stated placement.
Show how you would make crossings visible in production — the count reported beside the path's wall time — and which review findings you would reject outright, such as a trip count that grows with the data.
The judgment call is where the contract stops. State crossings and placement, keep it separate from capacity allocation and from what the tier is allowed to hold, and be ready to defend removing a crossing over tuning one when both are on the table.
## What the contract states A per-request trip budget is a short, checkable statement per request path, not a latency target with a nicer name. It names: - **A maximum number of crossings** the path may make to this tier while serving one request. - **A required placement** — same host, same zone, same region — because a trip count means nothing without one. - **The statistic and window** the budget is judged at, normally a high percentile over a stated period, because that is what a user-facing promise is made at. - **Whether the trip count is fixed or scales**, and with what. A path whose crossings grow with the size of a collection has no budget at all; it has a budget until the collection grows. | What it fixes | A latency target | A per-request trip budget | |---|---|---| | The number of crossings | unstated | stated, per path | | The placement assumed | unstated | stated as a requirement | | What breaches look like | a slow request, after the fact | a trip count that grew, at review | | What it is judged on | one number | wall time beside the call count | ## Deriving the number rather than inventing it 1. **Start at the user-facing deadline** for the path — the promise someone outside actually made. 2. **Subtract the other dependencies' allowances.** The durable engine, any downstream service, template rendering and serialisation all take a share. What is left is the tier's allowance. 3. **Divide by a measured per-crossing cost at the required placement**, taken at the caller and at a high percentile. The quotient is the trip allowance, and it is often a small single-digit number. Writing the arithmetic down makes the trade-off explicit: where a crossing costs a millisecond, a 10 ms allowance buys a handful of crossings, and a path that wants forty must change its placement, its share, or its shape. ## Making crossings visible An unmeasured contract is advice. The measurement change is small and specific: - **Count calls per request path**, and report that count beside the path's wall time. A tier-level operations-per-second figure cannot show it, because the count is spread over requests. - **Alert on the count, not only on the time.** A path that grows from four crossings to forty is a defect the moment it ships, even if it still fits today. - **Look at the shape.** A long row of narrow, equal-width, strictly sequential calls is the signature, and it is recognisable without reading any labels. - **Catch the scaling pattern in review.** One crossing per item in a list is the single most common way a path acquires an unbounded trip count. ## Removing a crossing beats tuning one A crossing not made costs nothing, cannot fail, cannot have a bad tail, and cannot be made worse by a placement change later. A crossing that has been tuned still costs a round trip and still carries every one of those risks. So the review question is not "can this call be faster" but "why is this call happening at all": - Is every key fetched actually used by the response? - Could one question be asked instead of fetching the inputs to answer it locally? - Could the keys be asked for together, in one operation or in a run sent without waiting for each reply? Each is a direction, and the mechanism behind each has its own trade-offs — larger replies to buffer, no atomicity across a run, keys that may not all live on one node. ## Placement as policy rather than accident Placement is usually decided by whoever created the deployment and discovered by whoever is paged. Making it part of the contract means: - The required placement is stated before the path is written, so the trip allowance is known while the code is being designed. - A change of placement is a change to the contract, and the paths affected are the ones whose budgets were computed from it. - Paths that cannot meet their budget at the available placement are identified as design problems rather than as tuning problems. ## Where the contract stops It does not say what belongs in the tier, what a value should look like, or how long anything should live there. It does not pick the mechanism for collapsing crossings — that depends on the store and on how the keys map to nodes. And it is not a capacity allocation: a path can be inside its trip budget and still be a heavy consumer, which is a different conversation with different numbers. ## What varies, and what the contract must not assume - **Not every store offers an operation taking many keys.** A contract whose only route to a low trip count assumes one is unenforceable on a store that answers single-key reads and writes alone. - **The per-crossing cost is a property of your deployment**, not of the store, and it must be re-measured when anything between caller and tier changes. - **Where the keyspace is split across nodes**, the crossing count is the number of nodes the caller talks to, so the budget must be written in crossings rather than in operations. - **The execution model is irrelevant to this budget** — it decides who is delayed behind an expensive call, not how many times a caller crosses the network.
- Which single review finding most often predicts a path will breach its trip budget later?A trip count that scales with something the path does not control — one crossing per item in a collection, per row returned, or per element of a user-supplied list. It passes review while the collection is small and breaches silently as the data grows. Budgets should therefore state whether the count is fixed, and unbounded counts should be rejected rather than measured.
- A team meets its trip budget by fetching one very large value in a single crossing. Is the contract satisfied?The trip budget is, and that is exactly why it is not the only contract. One crossing carrying a very large reply moves the cost into transfer time and into the reply buffer at each end, and on some stores into the server's own service time. A trip budget governs crossings; reply size and the cost of the work asked for need their own limits alongside it.
saying these in an interview costs you the question
- Sets one latency target for the tier and calls that a budget
- Optimises the per-call path instead of removing calls from it
- Lets each team pick placement independently and discover the cost in production
- Counts operations per second per team instead of crossings per request path
- Writes the budget from a mean rather than the percentile the promise uses
- Treats the budget as guidance with nothing measuring it