In a runtime that reclaims unreachable memory automatically, how can a program still leak memory?
answer
- collection frees only the unreachable
- reachable does not mean wanted
- something long-lived still points at it
- a container that only ever grows
- added on a hot path, removed never
basics
~20 sAutomatic reclamation frees only what nothing can reach. A leak there is memory that stays reachable from something live - an unbounded cache, a growing registry, a collection nobody trims - but will never be used again.
solid answer
~50 sA tracing collector answers one question: is this object still reachable from a live starting point? Anything reachable is kept, whether or not the program will ever look at it again. So the everyday leak in a collected runtime is not memory the collector missed - it is memory the program is still holding on purpose and has stopped wanting. The classic shape is a map used as a cache with no eviction and no expiry: every key ever requested stays reachable from a long-lived field, so the live set grows with the number of distinct keys ever seen and never comes back down. Nearly every leak of this kind is reachable memory; the exception is a resource held outside the managed heap through a handle nobody closed. The fix is always the same in form - find what still points at it, and stop pointing.
go deeper
Recall the one sentence that matters: collection frees what nothing can reach, so memory that is still referenced is kept even if the program will never use it again. Name one shape, such as a cache that is never trimmed.
Explain the mechanism: reachability from live starting points is the whole rule, so a leak is an asymmetry between a path that adds an entry and a path that should remove it and does not.
Show how you would confirm it in production - trend retained memory after collection under steady traffic, ask whether growth tracks concurrent load or cumulative distinct keys, then locate the container that holds the path.
Frame it as a design obligation: every long-lived container in a service needs a stated bound and an owner, because unbounded retention is a capacity decision made by accident rather than a bug in the runtime.
## What automatic reclamation actually promises A tracing collector answers exactly one question about every object: **can it still be reached** by following references from a set of live starting points - the running call stacks, values held by long-lived process-wide storage, and whatever the runtime itself pins. Everything reachable is kept. Everything else is free memory, whether the program still wanted it or not. That promise is narrower than people hear it as. It removes one bug class - releasing memory somebody still points at, and the corrupted reads that follow - and it removes the chore of pairing every allocation with a matching release. It does **not** decide what the program *ought* to still want. Reachability is a mechanical property of the object graph. Usefulness is a property of your intentions, and no collector has access to those. ## So what is a leak here? In a manually released heap, a leak is memory that was never handed back and can no longer be reached in order to hand it back: genuinely lost. In a collected heap that shape is rare. The everyday leak is its mirror image - memory that is **perfectly reachable and will never be used again**. | | manual release | automatic reclamation | |---|---|---| | the leaked bytes are | unreachable and never released | reachable from a live starting point | | the reclaimer's verdict | no reclaimer involved | correctly retained; this is not garbage | | what the fix changes | adds the missing release | drops the reference that should have gone | | how it looks in a dump | orphaned blocks | a live container that keeps growing | Both are called leaks because both end the same way: a footprint that only climbs, and eventually an allocation that cannot be satisfied. But they are found differently. There is no missing release call to hunt for; there is a *path* to hunt for, from something long-lived down to the bytes you no longer want. ## The shapes that hold memory forever Almost every retention bug is one of a short list: - **A cache with no eviction and no expiry** - a map keyed by user, tenant, request or query that is written on every miss and never trimmed. Its size tracks distinct keys ever seen, not concurrent load. - **A registration nobody removed** - a handler added to a long-lived hub whose owner was discarded without unsubscribing. - **A value parked in per-thread storage on a pooled thread** - the task ends, the thread does not, and the value stays attached to it. - **A handle that was opened and never closed**, holding a buffer or an operating-system resource the managed heap does not even measure. - **A registry that only grows** - a process-wide list of every instance ever created, kept "for diagnostics", with no matching removal. The common factor is an **asymmetry**: something is added on a path that runs often, and removed on a path that runs never, or only on the happy path, or only when a shutdown that never happens arrives. ## Telling a leak from a service that is simply large 1. Look at memory **after** a collection, not at the peak. A footprint that saw-tooths between two stable levels is allocation churn, not retention. 2. Ask what the growth **tracks**. Concurrent requests is capacity. Distinct keys ever seen, or components ever created, is retention. 3. Give it time and a fixed workload. Steady traffic with steadily climbing retained memory is the signature; a warm-up that flattens after an hour is not. 4. Find the path. Whatever tool your ecosystem gives you, the answer is the same shape: one long-lived container, one reference, one place that adds without removing. ## What does not fix it - **Asking for a collection explicitly.** The data is reachable; every collection is entitled to keep it and will. - **Raising the memory limit.** A leak is a slope, not a level. A bigger ceiling changes when you hit it, never whether. - **Clearing the local variable.** If the registry, cache or hub still has an entry, dropping your own reference changes nothing. - **Restarting on a schedule.** It buys time at the cost of never learning the shape, and the leak usually gets faster as traffic grows. The honest summary for an interview: automatic reclamation guarantees you will not free memory too early. It guarantees nothing about freeing it at all, because the thing it measures - reachability - is a decision your code made, not one the runtime can second-guess.
- How is a leak different from a service that simply needs a lot of memory?A large service reaches a plateau: under steady traffic its retained memory settles at a level and stays there, rising and falling with load. A leak has a slope - retained memory climbs with cumulative work done rather than with concurrent work in flight, and never returns to its earlier floor. Measure after collection, over hours, at fixed traffic.
- Is a cache with a maximum entry count safe from this?Bounded in entries is not bounded in bytes. A cap of ten thousand entries still leaks if entries can hold arbitrarily large values, or if each value transitively retains a whole object graph. Bound what actually matters - total retained size, or a time-based expiry - and be clear that a cap turns an unbounded leak into a fixed, budgeted cost.
- Why does explicitly requesting a collection not help a leaking service?Because nothing about the leaked data is garbage. It is reachable from a live starting point, so any collection - requested or automatic - is obliged to keep it. A requested collection can be useful for measurement, since it gives a clean post-collection floor to trend, but it cannot reclaim retained memory and never will.
A left-luggage office only destroys a bag once nobody holds a ticket for it. The leak is not an abandoned bag - it is the drawer of tickets you kept for bags you will never collect.
saying these in an interview costs you the question
- Claims a runtime with automatic reclamation cannot leak memory
- Thinks leaked objects are unreachable and the collector simply missed them
- Says an explicit collection request will reclaim the growing data
- Blames the allocator or fragmentation before finding what holds the reference
- Treats a cache as bounded because its entries feel temporary
- Says raising the memory limit fixes it rather than delaying it