A store of database copies at its memory ceiling saw its hit ratio slide from 95% to 72% at flat traffic — evidence of what?
answer
- the working set is not directly observable
- flat traffic rules out the easy answer
- entries grew, or bytes per entry grew
- a high ratio proves less than it seems
basics
~20 sEvidence that the working set no longer fits under the ceiling: either more distinct entries are touched in the same window, or each entry now costs more memory. It is a size signal, not a policy verdict.
solid answer
~50 sYou cannot observe a working set directly — the store will not hand you the list of entries the workload touched. The miss rate stands in for it, and a fall at **flat traffic** says the tier now holds less than one window's worth of touched entries. Three explanations fit: the distinct count grew, the bytes per entry grew, or something sweep-shaped widened the window's distinct count without widening the interactive traffic. All three are size statements, and all three are answered by re-measuring entries touched and memory per entry, then comparing against the ceiling. What the number does *not* tell you is which entries deserve to be held or how the store should choose what to drop — and it is only a signal at all because these entries are copies: where a store holds the only copy, a miss is an incident rather than a measurement.
go deeper
Know that the share of reads answered by the store is the everyday signal that it is big enough, and that a drop with no traffic change means it is now holding less of what people ask for.
Explain the inference in both directions: add memory until the ratio stops rising, and read a fall at flat traffic as the working set outgrowing the ceiling. Name the two factors — distinct entries and bytes per entry — and say how you would tell them apart.
Show the discipline of testing downward: a very high ratio is not evidence of right-sizing, and stepping the ceiling down until the ratio moves is how the real edge of the working set gets found. Mention the competing sweep workload as a live cause.
Draw the line the signal must not cross. Capacity evidence and design decisions about what belongs on the tier are different conversations, and a team that answers a falling ratio by changing how entries are dropped has substituted the second for the first.
## Why the miss rate is a measuring instrument The working set — the distinct **entries** touched in a window — is the number that decides how much memory an in-memory store needs. It is also a number you usually cannot read anywhere. Most stores of this class will tell you how many entries they hold and how much memory they attribute to them, but not which entries the workload asked for, over which window, and how many distinct ones that was. So the working set is inferred, and the instrument is the share of reads served from the store. The inference runs in both directions: - **Add memory and the hit ratio keeps rising** → the tier was holding less than one window's worth of touched entries. You were below the working set. - **Add memory and the hit ratio stops moving** → you have covered the window. Memory beyond that point holds entries nobody asks for. - **Hold memory constant and the hit ratio falls at flat traffic** → the working set grew past what the ceiling holds. That is the situation described. The flat part of that curve is the only honest evidence you get that a ceiling is big enough, and its knee is the number the sizing argument is really about. ## What the fall is consistent with "Flat traffic" rules out the boring explanation, so what is left is a size change, on one of two axes, plus one impostor: 1. **More distinct entries touched.** New tenants, a new feature addressing rows that were never read before, an identifier scheme that split one entry into several. The entry count of the working set grew. 2. **More memory per entry.** The same access pattern, but the values grew, so the same ceiling holds fewer of them. The working set in entries is unchanged and in bytes is not. 3. **A sweep-shaped workload.** A job that reads broadly and once inflates the distinct count in the window and displaces entries that interactive traffic will ask for again. Here the interactive working set did not really grow — a second workload arrived and is competing for the same memory. | Observation | What it is evidence of | What it is not evidence of | |---|---|---| | Miss rate rising, traffic flat | The working set no longer fits under the ceiling | A defect in the store | | Miss rate flat while memory is added | The ceiling already covers the window | That the ceiling is minimal | | Miss rate very high and stable | Nothing about whether the ceiling is right-sized | That the ceiling could not be much smaller | That last row is the one candidates get wrong. A hit ratio of 99% is equally consistent with a ceiling that exactly covers the working set and with one that covers it four times over. High is not the same as *right*, and a very high ratio is an invitation to test downward, not a proof of correct sizing. ## The line this evidence does not cross The miss rate is being read here as a statement about **size**. It is not a target to be maximised, and it does not answer: - which entries deserve to be held in the first place; - how the store should choose what to drop when it is full; - whether the tier should exist at all for this workload. Those are design questions with their own answers. Treating a falling hit ratio as a prompt to change how the store removes entries, rather than as a prompt to re-measure the working set against the ceiling, is the classic misreading — the store is telling you about capacity, and it is being asked about policy. ## When there is no signal to read The entire method depends on a miss being *recoverable*: these entries are copies, and a miss means one extra read against the system of record. On the same class of component, holding state that exists nowhere else — sessions, leases, quotas, deduplication markers — there is no hit ratio worth reading: - an entry that is absent is not a miss, it is loss; - the feedback arrives as an incident, not as a metric trend; - so the size has to be established in advance from live state, and watched by how close the store's own accounting sits to its ceiling rather than by a ratio. A mixed tier gets the strict treatment for the part that has no source of truth. ## What varies between stores - Some stores report a hit and miss count natively; others leave the application to count them, and a tier fronted by a client-side layer will report a ratio measured somewhere other than where you think. - Whether the store even has a configurable ceiling differs, and where it does, what that number counts differs — which is why the comparison is always ceiling against the store's own accounting, not against what the process holds. - What the store does once it is full differs too, and that changes the *shape* of the evidence: on a store configured to remove entries, growth past the ceiling shows up as a falling hit ratio, whereas on one configured to refuse writes it shows up as failing writes with the ratio on already-held entries barely moving. Read the signal your posture actually produces.
- The hit ratio is 99.5% and steady. What experiment would tell you whether the tier is oversized?Reduce the ceiling in steps and watch the ratio. As long as it stays flat, the memory you removed was holding entries nobody asked for; the point where it starts to fall is the edge of the working set. Do it on the tier's own traffic, in a window that includes the busiest period, and stop well before the ratio moves.
- How would you separate 'the working set grew' from 'the entries grew' with the same falling ratio?Compare the two factors independently. Count distinct entries touched in the window and compare with last month's count, then divide the memory the store attributes to its entries by the number of entries held and compare that. One of the two moved; the fix differs, since the first is a workload change and the second is a payload change.
saying these in an interview costs you the question
- Reads a falling hit ratio as a prompt to change how entries are dropped
- Treats a high hit ratio as proof the ceiling is correctly sized
- Assumes the store is broken rather than too small for the window
- Ignores that a sweep-shaped job inflates the distinct count
- Reads a hit ratio on a tier holding state with no source of truth