An origin database behind a shared in-memory tier suddenly takes the full read load while the tier's memory graph stays flat - which signal moved first?
answer
- flat is the expected shape at the ceiling
- ask which number moved first
- removals lead, hit ratio follows
- a silent signal is not a healthy one
- removals flat plus growth means the keys changed
basics
~20 sOrder the signals by which one leads: if the working set outgrew the ceiling, the pressure-removal rate moved first, the hit ratio fell next, and origin load rose last. Flatness at the ceiling is expected, not reassuring.
solid answer
~50 sA flat memory line is the *expected* shape for a tier pinned at its ceiling, so it rules nothing out. Three candidate causes each have a different lead signal. **The working set outgrew the ceiling**: the pressure-removal rate moves first, then the hit ratio falls, then origin load rises, while stored-data size stays flat at the bound throughout. **The tier restarted or failed over to an empty node**: stored-data size collapses toward zero and climbs again, connection counts churn as callers reconnect, and the pressure-removal rate stays at zero. **A deploy changed which keys callers ask for**: the hit ratio falls immediately with no change in removals at all, and stored-data size keeps growing because the old entries are still there and the new ones are being added beside them. Reading the three counters together discriminates between them before anyone touches the keyspace.
go deeper
Know that a flat memory line at the ceiling proves nothing, and that a store can be fast and useless at the same time. Start from which number moved first rather than which number looks high.
Explain the chain in order - removals, then hit ratio, then origin load - and be able to name the combination of counters that distinguishes an undersized tier from an emptied one and from a changed key population.
Show that your alerting sits on the leading signal rather than the symptom, that you handle the store which refuses writes instead of removing, and that you state plainly where the counters stop licensing conclusions.
Decide which of these signals pages a human and which only lands on a dashboard, and make the ordering itself the shared model, so that on-call reasoning across teams starts from the same chain rather than from whichever graph is open.
## Why 'the memory graph is flat' is not evidence A store holding at its memory ceiling produces a **flat line**. It is making room continuously, accepting a write for every entry it deletes, and the figure barely moves. That is the shape of a tier under sustained pressure and also the shape of a tier with nothing happening, which is why the memory graph is the wrong place to start. The question to ask of any symptom here is not "what is high" but **which number moved first**, because the lead signal is what separates causes that look identical once the symptom is visible. ## The ordering when the working set outgrows the ceiling 1. **Pressure-removal rate rises.** The server begins deleting entries to accept writes. Nothing user-visible has happened yet. 2. **Hit ratio falls** - first as counted at the server, since the entries being looked up are the ones just deleted. 3. **Origin load rises**, because every miss becomes a request to whatever sits behind the tier. 4. **Origin latency rises**, and only then do callers complain. 5. **Stored-data size does nothing at all**, because it was already at the bound before step one. The gap between step one and step four is the entire alerting opportunity, and it exists only if the pressure-removal counter is on a pager rather than on a dashboard nobody opens. ## Three causes, three lead signals | Cause | Moves first | Moves next | Stays flat, and that is the tell | |---|---|---|---| | **Working set outgrew the ceiling** | pressure-removal rate | hit ratio, then origin load | stored-data size, pinned at the ceiling | | **Restart or failover to an empty node** | stored-data size collapsing to near zero; connection count churning as callers reconnect | hit ratio, then origin load | pressure-removal rate, at zero - there is nothing to make room for | | **A deploy changed the keys callers ask for** | hit ratio, immediately and steeply | origin load | removal rates, unchanged; stored-data size keeps climbing as new keys are added beside the old ones | All three present identically at the origin. They separate on the tier's own counters, and they separate *before* anyone inspects a single key. ## The signals that are silent here, and what that is worth - **Latency at the tier need not move at all.** Removing an entry is cheap and serving a miss is cheap; the tier can be perfectly fast while being useless. A candidate who expects the tier's own latency percentiles to lead this failure has the model backwards. - **Connection count leads a different family of failure** - callers retrying, a pool leaking, or a reconnect storm after a restart - and it is nearly the only signal that leads the restart case from the caller's side. - **Replication lag leads the failover case** in the minutes *before* it, not after: a copy that had been drifting is the one that gets promoted holding less. - **A signal that goes to zero is not the same as a signal that is healthy.** A counter that stops being reported because a node disappeared reads as a flat line on most dashboards, which is how a dead node can look calm. ## Where stores in this class differ - On a store that **refuses writes at its ceiling** rather than removing entries, the first row of the table is wrong: no removal counter moves, writes start failing at the caller, and the hit ratio may stay high because the entries that are there stay there. The lead signal becomes the caller's write-error rate. - On a store that **reclaims deadlines only when an entry is touched**, a population can be entirely past its deadline and still occupying memory, so stored-data size is a weaker leading indicator than it appears. - On a **partitioned keyspace** all of these counters are per node, so a single node's collapse is diluted in any fleet-level graph. The lead signal on that shape is very often *the spread between nodes*, not the aggregate. ## Where this reading stops The signal set will tell you **which cause family you are in** and license one safe action - freeze the deploy, raise the ceiling if the instance has headroom and one owner, add a node if the shape allows. It will not tell you **which population grew**; naming that means walking the keyspace, which is a different exercise with its own cost on a live tier. Nor does it tell you what hit ratio this design should have been getting in the first place. The discipline is to state exactly what the numbers license and to stop there, rather than narrating a cause the counters never supported.
- The hit ratio fell steeply, removal rates did not move, and stored-data size is still climbing. Which cause does that combination point at?Callers asking for keys that nobody is writing - typically a deploy that changed a key naming scheme or a prefix. Nothing is being deleted, so removal counters are flat; the old entries remain, and the new key population is being written alongside them, so memory grows rather than holding. It is the one cause in this family where the tier is doing nothing wrong at all.
- Why should the alert sit on the pressure-removal rate rather than on origin load?Because origin load is the last signal in the chain, by which point users are already affected. The removal rate moves before any miss reaches the origin, which is the only window in which a capacity action is still cheap. Origin load stays useful as a symptom-level page, but as the sole signal it converts a warning into an incident report.
- On a keyspace partitioned across nodes, why can a fleet-level graph miss this failure entirely?Because every counter here is per node, and averaging dilutes one node's collapse across the rest. A single node at its ceiling, removing continuously while nine others idle, barely shifts the aggregate - yet every key assigned to it is affected. On that shape the leading indicator is the spread between nodes rather than the fleet number.
saying these in an interview costs you the question
- Reads a flat memory graph as evidence the tier is healthy
- Expects the tier's own latency percentiles to lead this failure
- Alerts on origin load only, after users are already affected
- Assumes every store removes entries rather than refusing writes at its ceiling
- Treats a counter that stopped reporting as a counter reading zero
- Jumps to inspecting individual keys before reading the counters together