How do you measure the real per-entry cost of your own entries in an in-memory store rather than estimating it?
answer
- a delta, not a lookup
- write a realistic sample, divide
- one constant per entry shape
- check more than one sample size
- entry accounting, never resident size
basics
~20 sWrite a large sample of realistically shaped entries into a store in a known state, read its own accounting of memory attributed to entries before and after, and divide the difference by the number written. Repeat per entry shape.
solid answer
~50 sThe procedure is a delta, not a lookup. Start from a store in a known state, note its own accounting of memory attributed to entries, write tens of thousands of entries whose keys are real keys at real lengths and whose values follow the real size distribution — with a lifetime attached if production attaches one — then read the figure again and divide the difference by the count. Do this once per distinct entry shape and weight the results by the production mix, because one blended average hides whichever shape dominates. Take readings at more than one sample size: the lookup structure grows in steps, so a single measurement straddling a growth step reads high. Divide the store's own accounting of entry memory, never what the operating system reports the process holding. Re-measure after a version upgrade, a build change, or a move to a different store.
go deeper
The takeaway is the direction of the method: you get this number by writing real entries and looking at what happened, not by finding it written down. Knowing that the figure is specific to a store and a shape is most of the value at this stage.
Be able to run the delta end to end and say which figure you are reading — the store's accounting of entry memory rather than what the operating system reports the process holding — and why a sample of a hundred entries produces a nonsense constant.
Show the judgment around the measurement: one constant per entry shape weighted by the mix, several sample sizes because the lookup structure grows in steps, the lifetime present if production has it, and a re-measure triggered by an upgrade.
Make the constant an auditable input rather than a fact. Decide what provenance a capacity number must carry, and what signal tells you it has aged — a drift between the modelled cost per entry and the one production is currently demonstrating.
## Why a published figure cannot be the answer Every component of per-entry overhead is a property of the store you are about to run, not of the data you are about to store: - the size and contents of the entry header; - how much spare capacity the lookup structure carries per entry; - the step the allocator rounds each request up to; - whether a deadline costs a header field that exists anyway or a record in a separate structure; - whether small values sit inside the descriptor or are reached through a reference. A number someone published is a measurement of their store, their build, their key lengths and their lifetime posture. It is not wrong so much as **about something else**. The senior habit this question looks for is the reflex of saying "write a thousand real entries and divide" before offering any figure at all. ## The procedure 1. **Build a representative sample.** Real keys at their real lengths and with their real variety — not ten thousand copies of one key pattern with a counter appended, unless that is genuinely what production writes. Values drawn from the real size distribution, not from its mean. A lifetime attached exactly as production attaches one. 2. **Get to a known state.** Note what the store's own accounting says it attributes to entries before you begin, so the measurement is a delta rather than an absolute. 3. **Write a large N.** Tens of thousands at least, so that one-off costs — the structure's initial allocation, the connection you measured over — wash out against the per-entry term. 4. **Read and divide.** The difference in the store's own entry accounting, divided by N, is the measured cost per entry. 5. **Repeat per shape and weight by the mix.** A workload with sessions, counters and small collections has three constants, not one, and the blended average will mislead you about whichever shape grows fastest. 6. **Take more than one point.** Measure at several values of N and check the increments agree. The lookup structure grows in steps, so a single measurement that happens to straddle a growth step reads high, and one taken entirely between steps reads low. ## The two readings you must not mix up | Figure | What it reports | Use it for | |---|---|---| | The store's accounting of entry memory | What the store believes its entries cost | Deriving the per-entry constant | | Resident size | What the operating system reports the process holding | A different diagnosis entirely | Dividing resident size by entry count is the classic corrupted measurement. Resident size includes memory the process holds for reasons that have nothing to do with your entries, and it carries the history of everything the process did earlier. A constant derived from it moves whenever that history changes, which makes it useless for projecting forward. ## What to record alongside the number A measured constant with no provenance decays into a published figure within a quarter. Record with it: - the store and version, and the build if that is something that varies in your environment; - the entry shape measured — key length distribution, value size distribution, whether a lifetime was attached; - N, and the several values of N you checked; - the date and the command-free description of how to re-run it. That turns the number into something a reviewer can challenge and a successor can reproduce, which is the difference between a measurement and a rumour. ## When to re-measure - After a version upgrade of the store — entry representations change between versions. - After a change in how entries are shaped: longer keys, a different value size distribution, lifetimes introduced or removed. - Before a migration to a different store in this class, where essentially none of the constant transfers. - Periodically against production reality: compare the modelled cost per entry against reported data size divided by reported entry count, and treat a drift as a signal that the model has aged. ## The traps - **A sample that is too uniform.** Identical key lengths and identical value sizes give a clean number that production will not reproduce. - **A sample that is too small.** At small N the structure's fixed allocation dominates and the constant reads absurdly high. - **Measuring without the lifetime.** If production attaches a deadline and the sample does not, the constant may be understated on stores where a deadline costs something. - **Measuring on an idle store and planning for a busy one.** The constant itself is fine; what changes under load is memory held for reasons other than entries, which belongs in the headroom part of a capacity plan rather than in this constant. - **Treating the result as universal.** It is specific by construction. That specificity is the entire value. ## What good looks like in an interview The answer that lands is short and procedural: build a realistic sample, delta the store's own entry accounting across a large write, divide, repeat per shape, record the provenance, re-measure on upgrade. The answer that does not land begins with a number.
- Why not simply divide what the operating system reports the process holding by the entry count?Because that figure includes memory held for reasons unrelated to your entries and carries the process's whole history, including memory freed earlier but not returned. A constant derived from it changes when that history changes rather than when your entries do, so it cannot be projected forward. Use the store's own accounting of what it attributes to entries, and treat the gap between the two as a separate diagnosis.
- The workload has several entry shapes. Is one blended average acceptable?Only if the mix is stable and no shape is growing faster than the others, which is rarely true. Measure each shape separately and weight by the production mix. The reason is practical: a blended constant hides the shape whose count is about to multiply, and that is exactly the shape a capacity projection needs to be right about.
- What does it mean if the measured cost per entry differs noticeably between two sample sizes?Most often that one of the measurements straddled a growth step in the lookup structure, which is allocated in jumps rather than smoothly. Measure at several sizes and use the increments between them rather than a single absolute. If the difference persists across many sizes, suspect that the sample itself is not uniform in shape — a wider key or value distribution in one run will move the constant legitimately.
saying these in an interview costs you the question
- Trusts a per-entry figure read somewhere instead of measuring the real entry shape.
- Divides what the operating system reports the process holding by the entry count.
- Measures with a hundred entries and treats the result as the constant.
- Uses one key pattern and one value size as a representative sample.
- Omits the lifetime from the sample when production entries carry one.
- Records the number with no store version, entry shape or date beside it.