skip to content

A running in-memory store exhausts the memory available to it: what three outcomes can follow, and how does each appear to the caller?

level: juniorimportance: must knowfreq 76%

answer

  1. one condition, several possible endings
  2. ask what happens to reads
  3. somebody chose this, probably by default
  4. the third one arrives as latency first

basics

~20 s

Memory exhaustion has three configured endings: the store refuses writes while reads still work, it removes entries to make room, or the operating system kills the process. Which one you get was a choice, usually a default nobody made deliberately.

solid answer

~50 s

One condition — the store needs bytes it cannot get — has three quite different endings, and which one happens was configured in advance. **Refusal**: writes fail while reads of existing entries generally keep succeeding, so the caller sees a partial outage on one path and nothing already stored is lost. **Removal (eviction)**: the store drops entries to stay under its memory ceiling and tells nobody, so the caller meets unexplained absence rather than an error, and a later read cannot distinguish a removed entry from one never written. **The kill**: if the store has no ceiling of its own, or one at or above the machine limit, the operating system ends the process and the whole keyspace goes with it. The asymmetry is the point: refusal is loud and recoverable, removal is silent and lossy, the kill is total.

go deeper

for a junior

Remember that running out of memory has more than one ending: writes can start failing, entries can quietly disappear, or the process can be killed outright. Knowing that a disappearing entry is not always an error message is most of the value here.

for a middle

Explain each outcome from the caller's seat and say which of the two limits caused it — the store's own memory ceiling, or the machine limit underneath it. Be precise that removal under pressure is eviction, not expiry.

for a senior

Show that you can read the outcome back from a production symptom: write errors with healthy reads, silent absence with no errors, or climbing latency followed by a vanished process. Note that where paging is unavailable the third gives no warning at all.

for a principal

Frame the posture as a blast-radius decision rather than a setting. Say what varies across implementations — default posture, whether a choice exists, whether the store has a ceiling at all — instead of presenting one behaviour as the model.

Memory is the one resource in an in-memory store whose exhaustion is **configured rather than fated**. A disk fills and writes fail; a network partitions and calls time out. Memory runs out and the store does one of three quite different things — and which one it does was decided in advance, usually by whoever accepted a default without reading it. ## One condition, three endings The condition itself is dull: the store needs bytes it cannot get. What makes it interesting is that two limits are in play, not one. - **The memory ceiling** is the store's own configured limit, enforced by the store's own code against the memory it attributes to its entries. - **The machine limit** is what the operating system or the container will actually let the process hold. A store that has a ceiling below the machine limit reaches its own ceiling first, and its own code chooses the outcome. A store with no ceiling concept at all, or with one set at or above the machine limit, never gets to choose — the operating system chooses for it. That single structural fact is why the third outcome exists. ## Outcome one: writes are refused The store keeps everything it holds and starts failing the calls that would need more memory. From the caller's seat this is a **partial outage**: the write path errors, while reads of entries already stored generally continue to succeed, because serving an existing entry needs no new memory for it. (The honest qualifier: a read that has to assemble a large reply still needs memory, so refusal is not a guarantee that every read survives.) This outcome is **loud and recoverable**. Nothing already stored is lost, the failure arrives at the caller with a return value attached, and somebody's monitoring can see it immediately. What the application should do about a refused write — retry it, drop it, route it elsewhere — is its own subject. ## Outcome two: entries are removed The store makes room by deleting entries it still holds. This is **eviction**: removal triggered by the ceiling. It is not the same event as **expiry**, which is removal because an entry's lifetime elapsed — same disappearance, different trigger, and only the first one is caused by memory pressure. The caller is told nothing. The write that triggered the removal succeeds normally. The loss surfaces later, somewhere else, as an entry that is simply not there — indistinguishable from one that was never written. This outcome is **silent and lossy**, and whether "lossy" matters turns on one question: *does the removed entry have a source of truth?* If it is a copy of something stored durably elsewhere, removal costs a slower read later. If it is the only copy — a session, a lease, a counter, a record that a job already ran — removal is data loss with nothing to re-read. ## Outcome three: the process is killed When the process reaches the machine or container limit, the operating system ends it. Where paging to disk is available, this is preceded by paging: the process's memory moves to disk, and operations that took microseconds start taking milliseconds. **This is why "it just got slow" is usually the third outcome arriving.** Where paging is disabled or unavailable — common in containers — there is no slow phase at all and the kill is abrupt. The kill is **total**: the whole keyspace goes at once, and the store logs nothing about it, because the decision was made outside the process. ## The three side by side | Outcome | What the caller sees | What is lost | How loud | |---|---|---|---| | Writes refused | Errors on writes, reads still served | Nothing already stored | Loud, immediate | | Entries removed | Nothing at the time; absence later | The removed entries | Silent | | Process killed | Rising latency, then no store at all | The entire keyspace | Total, external | ## Reading the outcome back from the symptom 1. Write calls failing while reads keep working for minutes on end → the store is refusing. 2. Entries missing that nobody deleted, with no errors anywhere → the store is removing. 3. Every operation slowing, then the process gone with nothing in its own log → the operating system killed it (abruptly, where paging is unavailable). ## What varies between stores None of this is uniform across the class, and asserting one store's behaviour as the model is the classic mistake: - whether a refusal posture exists at all — some stores in this class only ever remove; - whether the default is to refuse or to remove, which differs between implementations; - whether the store has a ceiling concept of its own, or leaves the limit entirely to the operating system; - whether removal may take any entry or only entries that carry a lifetime; - whether the ceiling is enforced against what the store attributes to its entries rather than against the process's resident size — where those two diverge, the process can exceed the machine limit while the store still believes itself under its ceiling, which turns an expected outcome one or two into an unexpected outcome three.

  • If reads keep succeeding while writes are refused, is the store "up"?
    It is half up, and that is worse than it sounds. Anything read-only carries on looking healthy, so a dashboard built on read success reports green while every write path is failing. Treat refusal as an outage of the write path specifically, and alert on write errors rather than on reachability.
  • Why is a removed entry harder to notice than a refused write?
    A refusal is returned to the caller that caused it, at the moment it happened. A removal is not returned to anyone: the write that forced it succeeds, and the loss appears later as a read that finds nothing, at a different caller, in a different code path, usually with no way to tell removal from an entry that was never written.
  • Can a store configured to remove entries still be killed by the operating system?
    Yes. The ceiling is enforced against the memory the store attributes to its own entries, while the operating system counts everything the process holds — buffers for connections and followers, and memory the allocator has not returned. If that gap is large enough, the process can cross the machine limit while the store still considers itself under its ceiling.

A full car park has three configured endings. The barrier stays down and new drivers are turned away, but every parked car is untouched. Or the attendant tows cars to free spaces, and owners find out only when they come back. Or the whole structure is condemned and every car inside is gone. Which ending is acceptable depends entirely on whether the cars are replaceable.

saying these in an interview costs you the question

  • Assumes reaching the ceiling always means eviction
  • Says the store simply crashes when it runs out of memory
  • Thinks a removed entry is reported to the caller as an error
  • Believes a store refusing writes is entirely unavailable
  • Treats every removed entry as a reloadable copy
  • Calls the slow phase before a kill a separate incident