A team says memory beats disk by orders of magnitude, so a shared tier will cut latency; what must that claim name first?
answer
- a ratio needs a baseline
- warm source or cold source
- both paths pay the same hop
- nanoseconds hide inside hundreds of microseconds
basics
~20 sName the baseline: which medium the comparison is against, whether the source is already answering from its own memory, and whether a network hop is being added. Against a warm source one hop away, the measured gap can be nil.
solid answer
~50 sThe ratio is a property of the media in isolation and means nothing until three things are fixed. First, which medium is on the other side - a spinning disk, a solid-state device and a page already resident in the source's memory are three baselines, orders of magnitude apart from each other. Second, whether the source is warm: if its working set is resident, the comparison is memory against memory and the medium advantage has already been collected by the source, for free. Third, whether a hop is added: a memory access is around a hundred nanoseconds and a datacentre round trip a few hundred microseconds, so once both paths cross the same network the medium difference is a rounding error inside the hop. Against cold reads the argument is real; against a warm source one hop away it can measure as nothing.
go deeper
Remember that a speed comparison needs two sides. Ask what the memory read is being compared with - a cold device read or a page the source already holds in memory - before agreeing with any ratio.
Explain the arithmetic: a memory access is nanoseconds, a network round trip is hundreds of microseconds, so once both paths cross the network the medium difference is swallowed by the hop.
Demonstrate that you would ask for the current latency breakdown - device wait against server time against network - and that you know a warm source one hop away can leave no gap to close at all.
Treat the ratio as a claim requiring evidence, and separate the latency argument from the load and ownership arguments so the design review debates the one that is actually doing the work.
## The claim, and the three things it hides 'Memory is orders of magnitude faster than disk' is true of the media in isolation and says almost nothing about a system until three things are named. 1. **Which medium is on the other side.** A spinning disk, a solid-state device and a page already resident in the source's own memory are three different baselines, separated by orders of magnitude from each other and not only from RAM. 2. **Whether the source is warm.** A disk engine keeps recently used pages in an in-memory area. If the pages this workload touches are already there, the comparison is memory against memory, and the medium advantage has already been collected - by the engine, without a new component. 3. **Whether a network hop is added.** A shared in-memory tier is a separate component with its own address. Reaching it costs a round trip. If the existing source is also one hop away, both end-to-end latencies are dominated by the same hop. ## The comparison once the baseline is named | baseline for one small read | rough end-to-end cost | what dominates | |---|---|---| | cold page on a spinning disk | several milliseconds | seek and rotation | | cold page on a solid-state device | hundreds of microseconds | device access | | resident page in the source, one hop away | a few hundred microseconds | the hop | | entry in a shared in-memory tier, one hop away | a few hundred microseconds | the hop | | entry held inside the calling process | under a microsecond | nothing external | The first two rows are where the slogan is true and the medium argument is strong. The middle two are the common case in a service whose source is well provisioned, and there the measured difference can be **nil**. The last row is a different arrangement altogether, and it is the only one that removes the hop rather than paying it again. ## How the gap goes to zero, or below The arithmetic is unforgiving. A memory access costs about a hundred nanoseconds; a datacentre round trip costs a few hundred microseconds - roughly three orders of magnitude more. Once both paths pay a hop, the medium difference disappears inside it. The comparison can even come out **negative**: - the new component may sit further away in the network than the source does; - a request that does not find what it wants there has still spent the hop before it continues; - a resident single-entry read at the source often costs tens of microseconds of server time - small enough that removing it changes nothing a caller can perceive. None of that says the component is pointless. It says the **latency argument** for it is unsound in this regime. There are other arguments - relieving a source that is expensive per request, or holding state that has no other home - but they are different arguments resting on different evidence, and reaching a verdict on them is a separate question from stating what the medium gap actually is. ## What to ask for instead of the ratio When the claim is made, the useful follow-ups are measurements rather than opinions: - what does the source's read latency look like now, at the median and at the high percentiles? - how much of that time is device wait, how much is server work, and how much is network? - how much of what this workload touches is already resident in the source? - where would the new component sit relative to the caller - same host, same rack, same region? If most of the time is the hop, a component one hop away cannot remove it. If most of the time is device wait on cold reads, the medium argument is real and the numbers will show it plainly. ## What varies, and must be said out loud The shape of this argument is not identical for every store or every source. Stores in this class differ in whether they are reached over a network at all; the same component held inside the calling process pays no hop, which inverts the arithmetic above. Sources differ in how much memory they were given and therefore in how often they touch a device. And some stores in this class can keep colder values on a solid-state device, which puts a device access back into a component that is otherwise all memory. Scoping the claim is not hedging - it is the content. A candidate who says 'against cold reads the gap is real, against a warm source one hop away it is close to nothing, and here is the breakdown I would ask for to tell which we have' has answered the question. One who repeats the ratio has not.
- The source's median is fine but its high percentiles are terrible. Does the medium argument apply there?That is exactly where it can bite: the slow tail is usually the device accesses the source's memory did not cover, and those are the reads a resident component would answer differently. It is still a claim about the tail, not the median, and acting on it is a separate decision from stating it.
- What single measurement would show that the hop, not the medium, dominates?A breakdown of one read into device wait, server time and network time. If network time is the majority and device wait is near zero, the source is warm and the medium comparison has already been settled in its favour; another component one hop away cannot improve on a cost it also pays.
- Why does a payload's size change which baseline matters?Because the ratios being quoted are ratios of access time, which is roughly constant per operation. Once a payload is large enough that moving the bytes dominates, both paths spend most of their time on transfer over the same link, and naming the access-time baseline stops changing the answer much.
saying these in an interview costs you the question
- Quotes a memory-to-disk ratio without saying which device or whether it is warm
- Assumes the source reaches its device on every request it serves
- Forgets that the new component is reached over the same network as the source
- Treats a benchmark of the component alone as an end-to-end latency
- Believes adding a memory component can only lower latency, never raise it