Callers of an in-memory store report steadily worse latency, and then the process disappears with no error in its own log: what happened?
answer
- the slowdown and the ending share a cause
- nothing in the store's own log
- pages went somewhere slower
- the store's ceiling never got to decide
basics
~20 sThe operating system's out-of-memory kill ended the process, and the slow phase before it was paging to disk. Nothing is in the store's own log because the decision was external, and its own ceiling never engaged.
solid answer
~50 sTwo symptoms, one cause. The latency was **paging to disk**: as the process approached the machine limit, memory it was using moved to disk, so operations that normally take microseconds began waiting on a disk. The disappearance was **the operating system's out-of-memory kill** — an external decision, which is exactly why there is nothing in the store's own log; it did not choose to stop. The structural finding is that neither of the store's own ceiling behaviours ever engaged: it was never going to refuse a write or remove an entry, because it never reached a ceiling of its own. On stores that have no ceiling concept at all, this is the only ending available. One caveat: where paging is disabled or absent, common in containers, the kill arrives with no slow phase to warn you.
go deeper
The key idea: the process was killed from outside, and the slowness beforehand was memory moving to disk. Do not expect the store itself to have logged anything about it.
Explain the mechanism in both halves — paging to disk explains the latency, the operating system's out-of-memory kill explains the disappearance — and say why the store's own ceiling behaviours never got a chance to run.
Demonstrate the diagnosis: compare what the store attributes to its entries against what the process actually holds, look outside the store for evidence of the kill, and know that where paging is disabled there is no warning phase at all.
Make the point that this outcome exists only when the store's own ceiling could not engage, and that the gap between attributed and resident memory is what decides whether a configured ceiling is really protecting the process.
This is the third of the three endings memory exhaustion can have, and it is the one that is not the store's decision. Reading it correctly matters because both of its symptoms get filed under something else: the slow phase gets investigated as a performance problem, and the disappearance gets filed as a crash. ## The slow phase is paging, not the store working harder As a process approaches the memory the machine or container will give it, the operating system starts moving its pages to disk to keep serving allocations. The store's code is unchanged and its own counters look ordinary; what changed is that touching memory can now mean waiting on a disk. The effect on an in-memory store is brutal in a way it is not for most software. **The entire value proposition of this tier is that a lookup is a memory access.** Once those accesses go to disk, a call that took microseconds takes milliseconds, and because callers usually hold a connection while they wait, the slowdown propagates: connections pile up, pools drain, and the symptom reported to you is a general application slowdown rather than anything pointing at memory. This is why **"it just got slow" is usually the third outcome arriving**, and why an in-memory store that is paging is already in an incident even though it is still answering every call correctly. ## The kill is external, which is why the log is empty When the process cannot be satisfied any longer, the operating system ends it. Two consequences follow directly: - **The store writes nothing about it.** Refusals and removals are the store's own decisions and can be counted and logged by the store; the kill is not a decision the process makes or observes. You find it in the operating system's or the orchestrator's records, not in the store's. - **The loss is total.** It is not a subset of entries chosen by some rule — the whole process image goes at once, and everything the store was holding in memory goes with it. Whether anything is recoverable afterwards is a separate subject, the store's durability posture. ## Why the store's own ceiling never engaged A store with a memory ceiling **below** the machine limit reaches its own limit first and then does whatever it was configured to do: refuse writes, or remove entries. Getting killed instead means one of a small set of things was true: 1. The store has no ceiling concept at all, and this is simply the only ending it has. Some stores in this class are like this. 2. A ceiling exists but is unset, or is set at or above the machine limit, so the operating system always wins the race. 3. A ceiling is set below the machine limit, but it governs only the memory the store attributes to its **entries**, while the operating system counts everything the **process** holds — buffers held on behalf of connections and followers, and memory the allocator has freed internally but not returned. If that gap is wide enough, resident size crosses the machine limit while data size is still comfortably under the ceiling. The third case is the one that surprises people, because the store's own reporting genuinely says it is fine. Choosing the ceiling and the headroom under it is its own subject; the point here is only that outcome three happens precisely when outcomes one and two were never reachable. ## Where paging is not available The slow phase is not guaranteed. In containers with paging disabled, and on hosts with none configured, there is nowhere for pages to go, so the sequence is not "slow, then dead" but simply **dead**, at full speed, with the last health check green. | | Paging available | Paging disabled | |---|---|---| | Warning before the kill | Minutes to hours of degradation | None | | Symptom reported | Latency, timeouts, pool exhaustion | Abrupt disappearance | | Chance to intervene | Real, if anyone is watching latency | Only before the fact, via alerting on memory | ## What to take from it - Treat rising latency on an in-memory tier as a **memory** symptom until proven otherwise, not only as a load symptom. - Compare what the store attributes to its entries against what the operating system says the process holds; a widening gap means the ceiling is not protecting you as much as it appears to. - Expect no evidence inside the store. Look at the operating system's or the platform's record of the kill, and note that a fast restart can hide the whole event from anything that only polls the store's own health. - Remember the ordering: a store that can be killed was never going to refuse or remove, so the two gentler outcomes are not available as a fallback.
- The store's own memory reporting said it was well under its ceiling. How can it still be killed?The ceiling governs what the store attributes to its entries; the operating system counts the whole process, including buffers held for connections and followers and memory the allocator freed internally but never returned. When that gap is wide, resident size crosses the machine limit while data size still looks safe.
- Does the kill always give you a slow phase first?No. The slow phase is paging to disk, so it only exists where paging is available. In containers with paging disabled, there is nowhere for pages to go and the process is killed at full speed, with the last health check green and no degradation to alert on.
- Why does an in-memory store suffer so much more from paging than a typical service?Because its whole premise is that a lookup is a memory access. Software that already reads from disk absorbs paging as a modest extra cost; a store whose operations are measured in microseconds can slow by three orders of magnitude, and callers holding connections while they wait spread the delay into the application.
saying these in an interview costs you the question
- Calls it a crash and looks for a bug in the store
- Treats the slow phase as an unrelated earlier incident
- Assumes the store logged why it stopped
- Believes the kill removes only some entries
- Thinks a store under its own ceiling cannot be killed
- Expects a warning phase even where paging is disabled