A store up 40 days shows resident size at twice its entry total, and the ratio has climbed every day this week — how do you tell unreturned memory from a leak?
answer
- one reading decides nothing
- which number is moving?
- find the event behind the step
- flat gap versus climbing gap
- restart, then watch the curve return
basics
~20 sThe snapshot cannot decide it; the trend and the workload history can. Establish which of the two figures is moving, what the workload did before the gap opened, whether the gap is flat or climbing, and whether a restart returns it and the climb then resumes on the same curve.
solid answer
~50 sA ratio read once has at least four defensible explanations, so the first move is to stop reading it once. Ask which number is growing: if the entry total is climbing too, you are storing more than you think and the gap is a side story. If the entry total is flat and only resident size climbs, the memory is unreturned, trapped, or genuinely leaked, and history separates those. A gap that jumped at an event — a mass deletion, a large expiry wave, a change in value sizes — and has been flat since is aftermath. A gap that climbs steadily with no event behind it and no change in entry count is the one worth chasing. The restart test is the blunt discriminator: if resident size comes back to near the entry total and then climbs the same curve again, the cause is in the workload, not in a defect.
go deeper
The key idea is that one memory reading cannot diagnose anything. Two numbers over several days, plus knowing what the workload did in between, is the minimum evidence anyone can reason from.
Be able to name the candidate explanations — memory freed but not returned, space trapped by a change in value sizes, a real leak, and simply storing more than you thought — and say which observation separates each from the others.
Run the procedure: split the entry total from resident size over weeks, hunt for the event behind any step, judge flat against climbing, then use a restart as the discriminator and watch whether the curve comes back the same way.
Land on a decision, not a label. Unreturned memory is not a defect and still costs headroom and an eventual restart, so the output is a scheduling and capacity call with a stated blast radius rather than a bug report.
## Why the snapshot cannot answer "Resident size is twice the entry total" is compatible with at least four different worlds, and they demand opposite responses: 1. **Unreturned memory.** A mass deletion or a large expiry wave freed space inside the process, and the operating system never got it back because the freed blocks are interleaved with survivors. Nothing is wrong. 2. **Trapped memory.** The workload's value sizes changed, and free space now sits in places the current sizes cannot use. Nothing is broken, but nothing will reclaim it either until the sizes shift back. 3. **A genuine leak.** Something allocates and never frees — in the store, in an extension it loads, or in the buffers it holds for connections. This is a defect and it ends in the operating system's out-of-memory kill. 4. **Data you forgot you were storing.** Entries written without a lifetime, a collection nothing ever trims, a diagnostic structure that only grows. The process is holding exactly what you asked it to hold. The ratio is identical in all four. The evidence that separates them is time and workload, which is why any useful version of this question hands you both. ## The discrimination procedure 1. **Split the movement.** Chart the store's entry total and the process's resident size on the same axis over weeks. Three shapes matter: both climbing, only resident climbing, and a gap that stepped once and went flat. 2. **Both climbing** points at case 4 before anything else. You are storing more; the gap is proportional and probably innocent. Look at what is being written without a lifetime and what collection has no upper bound before you suspect the store. 3. **Only resident climbing, entry total flat** narrows it to unreturned, trapped or leaked memory — which the remaining steps separate. 4. **Look for the event.** Ask what the workload did. A bulk purge, an expiry wave, a migration that rewrote every value at a new size, a traffic shape the process had never seen before. A gap that opened at an identifiable event and then stopped moving is aftermath, not a defect. 5. **Trend, do not snapshot.** *Flat and wide* is the signature of unreturned or trapped memory. *Climbing with no event and a flat entry count* is the signature of a leak. A week of daily readings settles more than any single measurement. 6. **Apply the restart test.** Restart the process and watch what happens. If resident size returns to near the entry total and then climbs the same curve over the following days, the cause is workload-shaped and reproducible; a defect that consumes a fixed amount per unit of a particular operation will show the same slope, but the curve's relationship to traffic is what tells them apart. A restart that does not return the memory says the entries were real all along. ## Reading the shapes | Entry total | Resident size | Gap over days | Most likely | |---|---|---|---| | Flat | Flat, wide gap | Flat since an event | Unreturned after a deletion or expiry wave | | Flat | Flat, wide gap | Flat, no event found | Trapped by a change in value sizes | | Flat | Climbing | Climbing steadily | Leak, or buffers held for connections that never drain | | Climbing | Climbing | Roughly proportional | You are storing more than you think | ## What varies between stores The procedure survives the differences, but the evidence available does not. How much a store reports about its own memory, and how it labels it, differs widely across this class — some publish a detailed breakdown, others little beyond a total. Where a store allocates from **fixed-size classes**, trapped space is counted inside the store's own entry total rather than in the gap, so case 2 above can present with a perfectly healthy ratio and the discrimination has to run on the entry total instead. And whether the process ever returns a large free span at all depends on the allocator underneath, not on the store. So describe the procedure in terms of what is moving, not in terms of one figure you expect to exist. ## The evidence the procedure needs Before any of the above is possible, three things have to be on hand, and an investigation that starts without them will produce a confident wrong answer: - **Uptime.** The same gap means different things on a process started yesterday and one started six weeks ago. - **A history of both figures**, not their ratio — the ratio hides which of the two is moving. - **A workload log.** What was bulk-deleted, what expired en masse, what changed size, what shipped. ## Why the answer is an availability judgment Unreturned and trapped memory are not defects, but they are not free either: they occupy real memory that the machine cannot give to anything else, and they consume headroom the process may need later. The honest conclusion often looks like "this is not a leak, and it still forces a restart eventually" — which is a scheduling decision rather than a bug report. The one thing not to do is the inverse of each error: opening a defect report on the aftermath of a purge, or explaining away a genuinely climbing gap as "just fragmentation" until the operating system's out-of-memory kill settles it for you. Note also that this procedure works at the level of process totals; *which* entries or prefixes are holding the bytes is a separate investigation with separate tooling.
- What does a restart tell you that the ratio cannot?It removes all history at once. A fresh process holds only what the current workload puts in it, so if resident size drops to near the entry total, whatever the gap contained was accumulated rather than required. Watching the days after the restart is the real test: a gap that rebuilds on the same curve is workload-shaped and will need the same remedy again.
- Both the entry total and resident size are climbing — what do you check first?What is being written that nobody removes. Entries stored with no lifetime at all, a collection under one key that only ever gains members, a diagnostic structure that grows with traffic. The store is holding exactly what it was asked to hold, so this is a workload question rather than a memory-management one, and no restart will fix it.
- The gap is wide but has been flat for a month — is there anything to do?It is not a defect, and it is still memory the machine cannot use for anything else, so it consumes the headroom the process has to grow into. The choices are to leave it and track it, to change the workload that produced it, or to schedule a restart as an availability event. Doing nothing is a defensible answer only once it is a decision.
- Can buffers held for connections look exactly like a leak?Yes, and they are the classic false positive. A client that reads slowly, or a follower that cannot keep up, makes the process accumulate memory on their behalf while the entry total stays perfectly flat — which is the leak signature. The tell is that it tracks connection or follower behaviour rather than time, and it drains when they do.
saying these in an interview costs you the question
- Declares a leak from a single ratio reading.
- Treats the aftermath of a bulk purge as a defect.
- Skips the workload history and goes straight to tooling.
- Explains a steadily climbing gap away as fragmentation.
- Never checks whether the entry total is climbing too.
- Believes any ratio above two is abnormal by itself.