A team proposes a shared in-memory store to fix a 900 ms p99, though its system of record answers reads in 4 ms warm; what would you measure before agreeing?
answer
- where did the 900 ms go
- count source reads per request
- baseline: warm engine, not cold disk
- repeats within a lifetime, or nothing
- the tail is a different population
basics
~20 sAttribute the 900 ms: where the time goes, how many source reads a request makes, what baseline the speed claim uses, and what share of reads repeat. A tier removes only time spent in the reads it absorbs.
solid answer
~50 sThe proposal assumes the time is in the reads, and a 4 ms warm read makes that unlikely. So: break the slow request down into application work, waiting on the system of record, waiting on other services, queueing when the service is near saturation, and runtime pauses. Count reads per request - one 4 ms read and 180 of them are different problems, and a tier moves the same round trips onto a new component rather than removing them. Fix the baseline: memory is far faster than a *cold* disk read, not necessarily than an engine already serving from its own memory one hop away. Then measure the repeat rate, because a derived copy pays only on reads that repeat. And look at the tail specifically; a p99 is a different population of requests, not a slow average.
go deeper
The takeaway is that a component is not added because it is fast but because a measurement shows where the time goes. Ask what the request actually spends its time on before agreeing to anything.
Be able to break a request's latency into its parts and to explain why a per-read time of 4 ms and a per-request read count of 180 point at completely different remedies.
Show that you can deliver the unpopular answer: the measurement often says the time is elsewhere, and a tier added on a guess is a permanent dependency regardless of whether it helped.
Frame the decision as one the organisation has to live with. State in advance what the tier is expected to remove and how that will be verified, so the team can tell later whether it earned its place.
## The arithmetic hidden in the proposal A 900 ms tail and a 4 ms warm read do not obviously belong to the same problem. If reads of the **system of record** account for a small slice of the request, then removing every one of them removes only that slice - and the proposal adds a hop of its own in exchange. The first duty is not to argue about the tier; it is to attribute the 900 ms. Adding a **volatile tier** on a guess is close to permanent: once traffic depends on it, removing it is a project, and the failure domain stays whether or not it helped. ## What to measure, in order 1. **Attribution.** Split the slow request into where its time actually goes: the application's own work, waiting on the system of record, waiting on other services, queueing for a thread or a connection when the service is near saturation, and runtime pauses. Until the 900 ms is attributed, every remedy is a guess dressed as a design. 2. **Reads per request.** A request issuing one 4 ms read and a request issuing 180 of them share a per-read number and share no problem at all. A high per-request read count is a shape defect: a tier in front of it absorbs the same number of round trips onto a different component. 3. **The baseline for the speed claim.** Memory is orders of magnitude faster than a **cold** read from disk. It is not orders of magnitude faster than an engine already answering from its own memory, one hop away - which is exactly what the scenario describes. Decide which comparison is being made before anyone quotes a ratio. 4. **The repeat rate.** A **derived copy** pays only for reads that repeat. Measure what share of reads ask for an entry read recently enough to still be present. Reads that never repeat give nothing back: each becomes a miss, plus a write, plus the hop. 5. **The tail itself.** A p99 is not a slow average; it is a different population of requests. Look at what those particular requests did - the largest payloads, the coldest paths, the ones behind a saturated pool - because the median request's profile may say nothing whatsoever about them. ## Reading the result | What the measurement shows | What it implies | | --- | --- | | Source reads are ~12 ms of 900 ms | The tier can return at most those 12 ms, minus the hop it adds. Look elsewhere. | | 180 source reads per request | Reduce the read count first; a tier absorbs the round trips, it does not remove them. | | Few reads, each expensive, repeated across users | The strongest honest case for a derived copy. | | Reads that essentially never repeat | The tier returns nothing: every read is a miss plus a write plus the hop. | | The tail is a different code path entirely | The remedy belongs on that path, whatever it turns out to be. | ## The honest no The result that most often survives measurement is that the time was somewhere else: in the application, in a chatty dependency, in queueing under load, in work that could simply not be done per request. Saying so is the senior move, and it is unpopular, because "put memory in front of it" is a faster sentence than a profile. The discipline is the same one used for any other dependency: name what you expect it to remove, in milliseconds, before you take on a component you will operate for years. ## What varies, and what does not - **Placement varies.** Some deployments put the store on the same host as the application, which shrinks the hop; others put it across a zone boundary, which widens it enough to matter. - **Server-side work varies.** Stores in this class differ in how much they do per request, so the portion of the time that is not network is not a constant across the class either. - **The source's own warmth varies by the hour.** An engine that is warm under steady traffic is cold after a restart or a deploy, and a measurement taken in one state does not describe the other. - **What does not vary** is that the speed claim needs a named baseline: which engine, warm or cold, in-process or across a hop. An argument that leaves the baseline unstated is not measuring anything. ## What the measurement does not decide No latency number tells you whether the entries are **derivable**. If the proposal quietly includes state that no system of record holds, that is a different decision with a different cost - a loss of facts rather than a loss of speed when the tier is empty - and it has to be taken on its own terms. Nor does a measurement settle who will run the tier: a single self-run instance, a replicated pair someone must fail over, and a **managed in-memory service** carry very different operational bills, and the comparison should say which one is being proposed.
- The profile shows one request issues 180 reads to the system of record. Does that argue for the tier?It argues first for issuing fewer reads. The request's shape is the defect: a tier in front absorbs the same 180 round trips onto a new component, so the per-read time falls while the count - and the added hop on each - stays. Fix the count, then re-measure and decide.
- What read pattern makes a derived copy pointless no matter how fast the store is?Reads that do not repeat within an entry's lifetime: a pass over distinct rows, or a per-user read taken once. Every read is then a miss, plus a write, plus the hop, so the tier adds latency and a dependency while returning nothing.
saying these in an interview costs you the question
- The p99 is slow, so put memory in front of the database.
- Memory is a hundred times faster, so the tier will fix it.
- No need to profile; everyone puts a store in front.
- A hit ratio can be assumed rather than measured from the read pattern.
- A p99 proves the source is the bottleneck under load.