How does a store that answers every read from RAM differ from a disk engine holding most pages in RAM?
answer
- guaranteed versus probable residency
- median similar, tail different
- tails come from device accesses
- the guarantee has a capacity bill
basics
~10 sAn in-memory store guarantees every read it serves is a memory read; a disk engine usually gets one. The gap shows in the tail - a miss costs the engine a device access.
solid answer
~50 sBoth components answer from memory most of the time, so the difference is not the slogan. An in-memory store keeps its dataset resident, so its read path is a lookup and a copy, and its latency distribution is narrow. A disk engine keeps its data in files and a large in-memory area of recently used pages; a read whose page is resident costs about the same as the store's read, and a read whose page is not resident pays a device access first. So the engine has two latencies and the store effectively has one. The honest statement is about **predictability**: guaranteed residency versus probable residency. The guarantee is paid for in RAM, which is why an in-memory tier normally holds a chosen subset. Some stores in this class can keep colder values on a solid-state device, which trades part of that guarantee for capacity.
go deeper
Recall the promise each component makes: one answers from memory by construction, the other keeps recently used pages in memory and goes to the device when it has to. Say that before saying anything about speed.
Explain the consequence rather than the ratio: similar medians once the engine is warm, very different high percentiles, and a capacity bill on the side that guarantees residency. Name which side of the comparison you are describing.
Show that you would measure the distribution, not the average, and that you know the engine's residency is a property of that deployment. Scope claims to stores that actually behave that way instead of asserting one design as the model.
Frame it as an economics and failure-domain question: a residency guarantee means buying RAM in proportion to what is held, and owning a component whose promise only holds while that budget does.
## Two components, not one dial An **in-memory store** keeps the dataset it is holding in the memory of the process that serves it. Answering a read means addressing the entry and handing back the bytes; for the entries it holds resident there is no step in that path that reaches a storage device. The store therefore makes a **residency guarantee**: what it currently holds, it holds in RAM. A **disk engine** makes a weaker claim. Its data lives in files, and it keeps a large in-memory area of recently used pages. A read whose page is already in that area is also a memory read, and costs the server roughly what the in-memory store's read costs. A read whose page is not there pays a device access first. The engine has two latencies, and which one a request gets depends on what has been touched recently and on how much memory the engine was given. That is the real distinction, and it is not the slogan: **guaranteed residency against probable residency**. In a healthy, well-provisioned system both components answer from memory most of the time. ## Where the difference actually appears | what is read | rough cost of the access | what governs it | |---|---|---| | an entry already in RAM | around a hundred nanoseconds | memory latency | | a page on a solid-state device | tens to a few hundred microseconds | device access latency | | a page on a spinning disk | several milliseconds | seek and rotation | | anything across a datacentre network | a few hundred microseconds | the hop, whatever the medium | The figures are rough and move with every hardware generation; the **ordering** is the stable part. Read the table as the shape of the comparison, not as numbers to quote back. Now put a workload through it. Suppose an engine answers ninety-nine reads in a hundred from resident pages. Its **median** read is a memory read, hard to tell apart at the server from the in-memory store's. Its **high percentiles** are a different animal: the hundredth read pays a device access two to four orders of magnitude slower than the median. The in-memory store's distribution is narrower, because for resident entries it has no miss to have. So what the medium buys is best stated as the shape of the latency distribution: - the median is often similar once the engine is warm; - the tail is where the residency guarantee shows; - the tail is what a request timeout, a queue depth or a high-percentile objective is actually made of; - neither component escapes the network: if it is reached over a hop, the hop is usually larger than the medium difference being argued about. ## What the guarantee costs RAM is the most expensive storage in the machine, and a machine takes only so many modules. A component that promises residency is therefore promising that what it holds fits in memory somebody paid for. That bill is the other half of the medium story, and it is why an in-memory tier normally holds a chosen subset of what the system knows rather than all of it. How large that subset should be, what one entry costs beyond its payload, and what the store does once it fills are separate subjects with their own answers; the medium only forces the question to exist. ## Where stores in this class differ None of the above should be stated as though every store behaved identically: - **Not every store keeps everything in RAM.** Some in this class can keep colder values on a solid-state device and fetch them back on access. That trades part of the residency guarantee for capacity, and the tail comes back with it. - **Placement changes the arithmetic.** The component may be reached over a network or may sit in the application process. Only the second avoids a hop, and the hop dominates most medium differences. - **Engines differ enormously in residency.** One deployment never misses its engine's memory; another misses constantly. Any claim about the engine side of this comparison is a claim about a particular deployment, not about disk. ## Saying it in an interview Lead with what each component promises rather than with a ratio: the store guarantees a memory read, so its distribution is narrow and its capacity is bounded by what RAM costs; the engine usually answers from memory too, and its interesting number is the tail rather than the median. Then name the baseline before quoting any figure - which device, warm or cold, and whether a network hop sits in the path. Whether a given design should own such a component at all is a separate question, decided on failure domains and measured evidence, not on the medium alone.
- Does the residency guarantee mean reads from such a store never slow down?No. The guarantee is about where the bytes are, not about queueing, server work or the network. A request still pays whatever hop separates it from the component, still waits behind other work on that server, and on stores that keep colder values on a solid-state device a device access can reappear for exactly those values.
- Why do the rough figures for these media keep moving while the argument does not?Because hardware generations change the constants, not the ordering. Memory access stays far below device access, device access stays far below a spinning-disk seek, and a datacentre round trip stays orders of magnitude above a memory access. Quote the ordering and the reason; treat any specific nanosecond or microsecond number as an approximation of its era.
A shop that stocks every item on the shelf against a shop that stocks the popular ones and sends to the warehouse for the rest. For anything on the shelf, both serve you at the same speed. The difference is what happens to the rare request - and the rent on a shelf big enough for everything.
saying these in an interview costs you the question
- Says a disk engine reaches the device on every read it serves
- Treats in-memory as a tuning flag rather than where the data lives
- Compares medians and ignores that only one component has a miss tail
- Assumes every store in this class keeps its whole dataset in RAM
- Thinks the residency guarantee comes without a capacity bill