Your in-memory store emits separate counters for entries removed under memory pressure and entries removed at their deadline - what does each rising alone mean?
answer
- promised removal versus unpromised removal
- one returns memory, one competes for it
- the deadline path says nothing about pressure
- zero to non-zero is the event
- some stores refuse writes instead of removing
basics
~20 sA rising deadline-removal rate means more entries reached the end of their lifetime, which says nothing about memory pressure. A rising pressure-removal rate means the server is deleting entries nobody asked it to delete, to make room.
solid answer
~50 sThe two counters answer different questions and must never be summed. **Time-based removal** is the store honouring a lifetime someone attached to an entry: the rate rising tracks the write pattern and the lifetimes chosen, and it is consistent with the tier having plenty of room - memory is being returned, not squeezed. **Pressure removal** is the server choosing what to delete so that a new write can be accepted, and every one of those deletions is data a caller expected to still be there - including, depending on how the store is configured, entries that carry no lifetime at all and were never intended to disappear. So a rising deadline rate is usually a change in traffic; a rising pressure rate is a capacity statement plus an ownership question - *whose* entries is the server choosing, and did that owner agree they were disposable?
go deeper
Recall that an entry leaving the store at its deadline and an entry deleted to make room are two different events with two different counters, and that only the second means the tier is short of room.
Explain what each counter actually tracks and why they must not be summed, including why a rising deadline rate is a statement about the write pattern rather than about capacity.
Show that you alert on the step from zero to non-zero pressure removals, that you cover the store that refuses writes instead, and that you ask whose entries are being chosen before calling it benign.
Treat pressure removal as a budget being spent on someone's behalf: decide in advance which populations are allowed to be victims, which must never be, and what that implies about running them on the same instance.
## Two removals, two counters, one common confusion An entry can leave the store for two structurally different reasons, and a tier worth operating reports them as two numbers: - **Time-based removal** - the entry carried a lifetime, its deadline arrived, and the store removed it. The removal was *promised* when the entry was written. - **Pressure removal** - the store needed room for an incoming write and deleted something to get it. Nobody promised this removal to anyone. Collapsing them into one 'entries removed' figure destroys the only distinction that matters, because the first is the system working as designed and the second is the system spending someone's data to stay up. | Counter rising | What it licenses you to conclude | What it does not mean | |---|---|---| | **Deadline removals** | more entries reached the end of their lifetime in this window - more writes with lifetimes, shorter lifetimes, or a batch written together reaching its deadline together | nothing about memory pressure; this path *returns* memory rather than competing for it | | **Pressure removals** | the server is now choosing what to delete so that writes can be accepted | not necessarily that the keyspace grew - one caller writing much larger values can force this with no change in entry count | | **Both rising together** | ordinary growth in a tier that is already at its bound | not that one caused the other; they are independent paths | | **Pressure removals at a steady non-zero rate** | the tier is deliberately or accidentally smaller than its working set, continuously | not, by itself, an incident - it is the normal operating state of an intentionally undersized tier | ## Why the pressure counter is the one that needs an owner When the server removes an entry to make room, three facts follow at once: 1. **Somebody's read will now miss** that would previously have hit. If that entry was a replaceable copy of something durable, the cost is an origin lookup. If it was state that exists nowhere else - a session, a lease, a deduplication record - the cost is a user-visible failure. 2. **The choice of victim is a policy, not a fairness guarantee.** Removal families order candidates by recency of use, by frequency of use, arbitrarily, or restrict the candidate pool to entries that carry a deadline. None of them know which team owns which entry, so a quiet-but-critical population can be deleted to make room for a noisy one. 3. **On a shared tier the counter has no owner**, because the instance-wide figure does not say whose entries went. Attribution is a separate exercise; the signal alone tells you removals happened, not to whom. ## Where stores in this class differ - check before you read the graph This is the part that makes a confident answer wrong on the wrong store: - **Not every store removes under pressure.** Some refuse the write instead, in which case the pressure-removal counter stays at zero permanently and the signal that actually moves is a write-error rate at the caller. If you alert only on removals there, you will never be paged. - **Reclaiming a deadline is not universal in shape.** Some stores reclaim on touch only, some run a background sweep, some reclaim during allocation. Where reclaim happens on touch, the deadline-removal rate tracks *traffic to expired entries* rather than the passage of time, so an idle population that has all expired may register almost nothing. What is universal is only that memory is not returned at the instant the deadline passes. - **Some stores do not separate the two counters at all,** or publish only one of them, which means the discrimination has to be made from the memory figures and the write-error rate instead. ## Turning it into an alert - **Alert on the appearance of pressure removals where there were none**, not on their absolute rate - the step from zero to non-zero is the meaningful event on a tier that was sized to fit. - **On a tier deliberately smaller than its working set**, invert it: alert on a sustained multiple of the usual rate rather than on any removal at all. - **Alert on the write-error rate too**, so that a store which refuses writes rather than removing entries is covered by the same page. - **Leave deadline removals on a dashboard, not a pager.** A change there is information about the write pattern, not about the tier's health, and paging on it teaches people to ignore the pager. - **Do not sum the two.** A combined 'keys removed' graph is the one artifact guaranteed to hide the transition from working-as-designed to spending-someone's-data.
- The pressure-removal rate jumped but the number of entries held barely changed. What can explain that?Value size rather than entry count. One caller writing much larger values consumes the same ceiling with far fewer entries, so the server must delete existing entries to accept each new one while the total count stays roughly flat. The signal reading stops there: knowing which population grew is a separate investigation over the keyspace.
- Why can a store's deadline-removal counter stay near zero while a large population has certainly expired?Because on stores that reclaim a deadline when an entry is next touched, the counter tracks traffic to expired entries rather than the passing of time. A population nobody is reading can be entirely past its deadline and still register almost no removals - and, importantly, still be occupying memory. Where a background sweep exists the shape differs, which is why the reclaim behaviour has to be confirmed per store.
- Which signal replaces the pressure-removal counter on a store that refuses writes at its ceiling?The caller's write-error rate, plus stored-data size sitting pinned at the ceiling. On that posture the server never deletes anything, so no removal counter will ever move; the failure surfaces at the writing caller instead. An alerting scheme built only on removals is blind to the whole failure on such a store.
saying these in an interview costs you the question
- Sums the two counters into one 'entries removed' graph
- Reads a rising deadline-removal rate as evidence of memory pressure
- Assumes every store in this class removes entries rather than refusing writes
- Believes memory returns at the instant an entry's deadline passes
- Treats any non-zero pressure-removal rate as an incident by itself
- Thinks only entries carrying a lifetime can ever be removed under pressure