Only 5 million of a store's 200 million entries, all copies of database rows, are read hourly — which number sizes it?
answer
- two numbers, only one is a size
- touched in a window, not stored
- a source of truth changes the arithmetic
- entries times measured cost per entry
basics
~20 sSize for the working set — the entries actually touched in a window — not total data. A database holds the rows, so an absent entry costs one extra read, and memory bought for untouched entries is never read.
solid answer
~50 sTwo different numbers both get called "the size". **Total data** is everything that exists; the **working set** is the distinct entries actually touched inside a window you choose. Here the rows live in a database, so an entry the store does not hold costs one extra read — which means paying for the 195 million entries nobody touches buys memory nobody reads. Size for the working set, and treat the number as a measurement rather than a fraction: count distinct entries touched in the window you care about, multiply by the memory one entry actually costs, and check the result against the ceiling you intend to set. The arithmetic inverts when the store holds the **only** copy of something — sessions, locks, counters — because then an entry that is not held is gone, not re-read, and total data becomes the number you must size for.
go deeper
Recall that an in-memory store usually holds a fraction of what exists, and that the fraction worth holding is what gets touched. Say "working set" and "total data" as two different numbers, and say which one you would size for.
Explain the conversion: distinct entries touched in a named window, multiplied by the memory one entry actually costs, checked against the ceiling. Show that you know a window has to be chosen and that a nightly sweep is not the window.
Demonstrate the condition behind the rule — the working set is the right number only because a source of truth exists. Name the content on your own tier that has no source, and show that it is sized differently.
Frame it as a purchase with two failure directions of unequal cost: over-buying shows up on an invoice and in restart time, under-buying shows up in the miss rate and in pressure on the system of record. Say which of the two you would rather be wrong in and why.
## Two numbers that both get called "the size" Every capacity conversation about an in-memory store opens with two numbers, and they are usually an order of magnitude apart: - **Total data** — everything that exists and could conceivably be held: every row, document or object in the system of record. - **The working set** — the distinct **entries** actually touched inside a window you name: the last hour, the last day, one checkout flow. An *entry* is the key, the value and whatever bookkeeping the store keeps for it, taken together; the working set is counted in entries and then converted to bytes, not guessed as a share of the total. | The number | What it counts | When it is the right one to size for | |---|---|---| | Total data | Everything that exists | The store holds the only copy, or the whole set genuinely must be resident for a latency target | | The working set | Distinct entries touched in a chosen window | A source of truth exists elsewhere, so an entry that is absent is re-read rather than lost | ## Why the working set is the number here In the situation described, the rows exist in a database. That single fact decides the arithmetic. The worst thing that happens to an entry the store never held is that a caller waits for one read against the database and the entry is written on the way back. So memory spent holding entries nobody touches buys nothing measurable: it is not faster, not safer, and not cheaper. The two ways of getting it wrong are not symmetric: 1. **Sizing for total data** over-buys. The cost shows up as a bigger machine, a longer restart, and a larger blast radius when the node is lost — never as a wrong answer. 2. **Sizing below the working set** under-buys, and the workload tells you so: at flat traffic, the share of reads served from the store falls, because the tier can no longer hold one window's worth of touched entries at once. That asymmetry is why the working set is the *starting* number and the miss rate is the correction. ## When the arithmetic inverts A store of this class is not always holding copies. The same component is routinely asked to hold state that exists nowhere else — sessions between application instances, leases, quotas, deduplication markers. For that content: - there is no source to re-read, so an entry the store does not hold is **data loss**, not a miss; - "the working set" is not a meaningful reduction, because every live entry is one you must still be able to produce; - the number to size for is total live data, and the honest lever is how long entries stay alive, not how many of them you hold. A tier that mixes both kinds of content inherits the stricter rule for the part that has no source of truth. Deciding which content deserves to be held at all is a separate design question; the sizing rule only asks whether a source exists. ## Choosing the window, and how the window goes wrong The working set is not a property of the data — it is a property of the data *and a window*. Two traps follow: - **A window too short** flatters the number. Ten minutes of traffic on a site with a daily rhythm counts a fraction of what a day touches, and the tier under-performs outside the window you measured. - **A sweep inflates the number.** A nightly job that reads every entry once makes the distinct count over twenty-four hours equal to total data, while telling you nothing about the interactive traffic you are actually sizing for. Measure the window that serves the latency you care about, and treat a full sweep as its own workload. The working set also moves in bytes without moving in entries: if each stored value grows, the entry count is unchanged and the memory required is not. ## What varies between stores The concept is universal; the surrounding machinery is not. - Some stores expose a **memory ceiling** you configure; on others no such setting exists and the machine or container limit is the only bound, which makes the sizing decision a node-choice decision instead. - What a store counts as the cost of one entry differs — its own bookkeeping, how the value is laid out, and how the allocator rounds a request all differ between implementations, so the per-entry figure must come from a measurement on your own data rather than from a number you remember. - Some stores in this class offer no way to observe which entries were touched, so the working set is inferred from the miss rate rather than read off directly. The defensible answer in an interview is therefore not a number but a shape: *the working set, measured over a named window, converted to bytes by a measured cost per entry, and checked against the miss rate afterwards* — unless the tier is the only copy, in which case it is total live data.
- A reporting job reads every entry once a night. Does that make the working set equal to total data?Only for a twenty-four-hour window, which is the wrong window to size interactive traffic with. Measure distinct entries touched in the window whose latency you care about, and treat the sweep as a separate workload that will disturb the tier while it runs rather than as evidence about the size the tier needs.
- Traffic is flat and the entry count is unchanged, but each value grew from 300 bytes to 900. Which number moved?The working set's size in bytes, not its size in entries. Sizing is entries touched multiplied by the memory one entry costs, and only the second factor moved — so the same access pattern now needs roughly three times the memory, and a ceiling that was comfortable is not.
- Is it ever right to size for total data when a database holds every row?Yes, when the point of the tier is that no read may ever reach the database — a latency floor, or protecting a system of record that cannot take the traffic. That is a deliberate decision to buy memory for entries nobody reads, and it should be stated as such rather than arrived at by accident.
saying these in an interview costs you the question
- Sizes for the full row count because any row might be requested
- Assumes an absent entry is lost even when a database holds the row
- Treats the working set as a fixed fraction, like ten percent
- Sizes for the working set even when the store holds the only copy
- Names a window without measuring it — just says 'the hot data'