skip to content

Every entry on your tier is written with a lifetime, yet its reported memory keeps climbing — what explains that, and what would you check?

level: seniorimportance: should knowfreq 50%

answer

  1. a lifetime is not a memory bound
  2. resident equals live plus unreclaimed
  3. nobody reads it, so nothing reclaims it
  4. does memory fall under pressure?

basics

~20 s

Most likely the tier holds entries whose deadlines passed but which nothing has reclaimed: nobody reads them back, and a sampled sweep is not keeping up. Check whether anything reads, and whether memory falls under pressure.

solid answer

~50 s

A lifetime on every entry does not put a ceiling on memory, because the deadline only marks an entry dead — the bytes return on a separate schedule. If the workload writes entries that are never read back, reclaim on access never fires for them, and a background sweep that samples can fall behind a high write rate. The result is a resident population of live entries plus an unbounded backlog of dead ones. Checks worth making, in order: is anything reading these entries at all; does the entry count exceed what you believe is live; and does memory drop sharply when the tier is pushed toward its ceiling, which confirms dead entries were simply waiting for pressure. Rule out two different causes before settling: entries that carry no lifetime at all, and memory the process has freed but not handed back.

go deeper

for a junior

Take away the headline: putting a lifetime on everything does not cap memory, because the bytes come back later than the deadline does.

for a middle

Explain why the workload matters. Entries nobody reads back never trigger reclaim on access, so only a sampled sweep and allocation pressure can return their memory.

for a senior

Show the diagnosis as a sequence: confirm nothing reads these entries, compare resident count against expected live count, then watch what memory does as the tier is pressed.

for a principal

Frame it as a sizing contract. If headroom assumes prompt reclaim, either buy the headroom for the dead backlog or introduce an explicit mechanism that returns memory on your schedule.

## Why "everything has a lifetime" is not a memory bound The intuition behind the alarm is reasonable: if every entry dies after an hour, surely memory tracks one hour of writes. It does not, because a lifetime governs **serving**, not **occupancy**. At the deadline the entry becomes unservable and stays exactly as resident as it was. It becomes free only when one of three things happens: a caller touches it, the store's background sweep reaches it, or the store needs the room. So the resident population is always **live entries plus dead entries not yet reclaimed**, and the size of that second group is a property of the workload, not of the lifetime. ## The workload shape that produces this The symptom shows up almost exclusively on one shape: **entries written and never read back**. Records of work already done, fan-out state for a consumer that may not arrive, per-request artefacts written defensively — anything where the write is the point and the read is rare. - **Reclaim on access contributes nothing**, because no access ever happens. - **The background sweep is sampled**, so it works through the dead population at a pace set by its own budget, not by how much there is to do. A write rate above that pace grows the backlog. - **Reclaim under allocation pressure is the backstop**, and it fires only when the tier is actually pressed. Until then, memory climbs. On a store with no background pass at all, this is not a fault condition; it is the expected steady state, and the tier is behaving exactly as designed. ## What to check, in order 1. **Is anything reading these entries?** If the answer is no, you have the explanation and everything else is confirmation. 2. **How many entries are resident against how many you believe are live?** A large excess is dead entries waiting for a reclaim route. This is a one-off comparison for diagnosis, not a number to build alerting on. 3. **What does memory do when the tier approaches its ceiling?** If it falls sharply and then serving continues normally, the store was holding dead entries and released them when it needed the room — which confirms the diagnosis and tells you the backstop works. 4. **Rule out the two look-alikes.** Entries that were written without a lifetime, or that lost one, are never reclaimed by any of the three routes — a different subject with a different fix. And memory the process has freed internally but not handed back to the operating system is a third subject again: the store's own accounting has already dropped it, so the two are distinguishable by which number moved. ## The trap in the obvious fix The reflex is to shorten the lifetime. On this workload it does nothing for memory. A shorter lifetime makes entries dead sooner; it does not make anything reclaim them sooner, because none of the three routes is triggered by how long ago the deadline passed. You end up with the same resident bytes and a shorter window in which the data would have been useful had anyone wanted it. ## What actually helps | approach | what it does | when it is right | |---|---|---| | size for live plus dead | accepts the backlog and pays for it in headroom | the backlog is bounded and affordable | | delete explicitly when done | returns memory on your schedule, not the store's | the writer knows when the entry stops mattering | | a deliberate pass that touches the entries | drives reclaim on access on purpose | you cannot delete, but can afford to read | | spread deadlines across the population | smooths reclaim work rather than reducing it | many entries would otherwise die together | The judgment underneath all four is the same one: **a deadline expresses when data stops being meaningful, and it is not an instrument for managing capacity.** If a capacity plan depends on memory coming back at a particular time, that plan needs an explicit mechanism, because no store in this class offers a bound between the deadline passing and the bytes returning.

  • Would halving every lifetime bring the memory down?
    Not on a keyspace nothing reads back. Entries become dead sooner, but none of the reclaim routes is triggered by how long ago a deadline passed, so the same bytes stay resident until a sweep samples them or the store needs the room. You lose useful data and keep the footprint.
  • How would you tell this apart from memory the process freed but has not returned to the operating system?
    By which number moved. In the case here, the store's own accounting still counts the entries, because they are genuinely still held. In the other, the store's accounting has already dropped them and only the process footprint stays high. Comparing the two figures separates the diagnoses immediately.
  • Is it ever acceptable to just let the backlog sit there?
    Yes, when it is bounded and the headroom is affordable. Reclaim under allocation pressure is a real backstop: the store gives up dead entries before anything live, so the tier degrades gracefully. The decision is whether you can price the extra headroom, not whether the state is tidy.

saying these in an interview costs you the question

  • Shortens the lifetime and expects the memory to come back sooner.
  • Concludes the store is leaking memory.
  • Assumes a background pass promptly clears every deadline that has passed.
  • Treats a lifetime on every entry as a bound on total memory.
  • Blames fragmentation before checking whether the entries are still held.