skip to content

questions

24

In a runtime that reclaims unreachable memory automatically, how can a program still leak memory?

level: juniorimportance: must knowfreq 74%

answer

  1. collection frees only the unreachable
  2. reachable does not mean wanted
  3. something long-lived still points at it
  4. a container that only ever grows
  5. added on a hot path, removed never

basics

~20 s

Automatic reclamation frees only what nothing can reach. A leak there is memory that stays reachable from something live - an unbounded cache, a growing registry, a collection nobody trims - but will never be used again.

solid answer

~50 s

A tracing collector answers one question: is this object still reachable from a live starting point? Anything reachable is kept, whether or not the program will ever look at it again. So the everyday leak in a collected runtime is not memory the collector missed - it is memory the program is still holding on purpose and has stopped wanting. The classic shape is a map used as a cache with no eviction and no expiry: every key ever requested stays reachable from a long-lived field, so the live set grows with the number of distinct keys ever seen and never comes back down. Nearly every leak of this kind is reachable memory; the exception is a resource held outside the managed heap through a handle nobody closed. The fix is always the same in form - find what still points at it, and stop pointing.

go deeper

for a junior

Recall the one sentence that matters: collection frees what nothing can reach, so memory that is still referenced is kept even if the program will never use it again. Name one shape, such as a cache that is never trimmed.

for a middle

Explain the mechanism: reachability from live starting points is the whole rule, so a leak is an asymmetry between a path that adds an entry and a path that should remove it and does not.

for a senior

Show how you would confirm it in production - trend retained memory after collection under steady traffic, ask whether growth tracks concurrent load or cumulative distinct keys, then locate the container that holds the path.

for a principal

Frame it as a design obligation: every long-lived container in a service needs a stated bound and an owner, because unbounded retention is a capacity decision made by accident rather than a bug in the runtime.

## What automatic reclamation actually promises A tracing collector answers exactly one question about every object: **can it still be reached** by following references from a set of live starting points - the running call stacks, values held by long-lived process-wide storage, and whatever the runtime itself pins. Everything reachable is kept. Everything else is free memory, whether the program still wanted it or not. That promise is narrower than people hear it as. It removes one bug class - releasing memory somebody still points at, and the corrupted reads that follow - and it removes the chore of pairing every allocation with a matching release. It does **not** decide what the program *ought* to still want. Reachability is a mechanical property of the object graph. Usefulness is a property of your intentions, and no collector has access to those. ## So what is a leak here? In a manually released heap, a leak is memory that was never handed back and can no longer be reached in order to hand it back: genuinely lost. In a collected heap that shape is rare. The everyday leak is its mirror image - memory that is **perfectly reachable and will never be used again**. | | manual release | automatic reclamation | |---|---|---| | the leaked bytes are | unreachable and never released | reachable from a live starting point | | the reclaimer's verdict | no reclaimer involved | correctly retained; this is not garbage | | what the fix changes | adds the missing release | drops the reference that should have gone | | how it looks in a dump | orphaned blocks | a live container that keeps growing | Both are called leaks because both end the same way: a footprint that only climbs, and eventually an allocation that cannot be satisfied. But they are found differently. There is no missing release call to hunt for; there is a *path* to hunt for, from something long-lived down to the bytes you no longer want. ## The shapes that hold memory forever Almost every retention bug is one of a short list: - **A cache with no eviction and no expiry** - a map keyed by user, tenant, request or query that is written on every miss and never trimmed. Its size tracks distinct keys ever seen, not concurrent load. - **A registration nobody removed** - a handler added to a long-lived hub whose owner was discarded without unsubscribing. - **A value parked in per-thread storage on a pooled thread** - the task ends, the thread does not, and the value stays attached to it. - **A handle that was opened and never closed**, holding a buffer or an operating-system resource the managed heap does not even measure. - **A registry that only grows** - a process-wide list of every instance ever created, kept "for diagnostics", with no matching removal. The common factor is an **asymmetry**: something is added on a path that runs often, and removed on a path that runs never, or only on the happy path, or only when a shutdown that never happens arrives. ## Telling a leak from a service that is simply large 1. Look at memory **after** a collection, not at the peak. A footprint that saw-tooths between two stable levels is allocation churn, not retention. 2. Ask what the growth **tracks**. Concurrent requests is capacity. Distinct keys ever seen, or components ever created, is retention. 3. Give it time and a fixed workload. Steady traffic with steadily climbing retained memory is the signature; a warm-up that flattens after an hour is not. 4. Find the path. Whatever tool your ecosystem gives you, the answer is the same shape: one long-lived container, one reference, one place that adds without removing. ## What does not fix it - **Asking for a collection explicitly.** The data is reachable; every collection is entitled to keep it and will. - **Raising the memory limit.** A leak is a slope, not a level. A bigger ceiling changes when you hit it, never whether. - **Clearing the local variable.** If the registry, cache or hub still has an entry, dropping your own reference changes nothing. - **Restarting on a schedule.** It buys time at the cost of never learning the shape, and the leak usually gets faster as traffic grows. The honest summary for an interview: automatic reclamation guarantees you will not free memory too early. It guarantees nothing about freeing it at all, because the thing it measures - reachability - is a decision your code made, not one the runtime can second-guess.

  • How is a leak different from a service that simply needs a lot of memory?
    A large service reaches a plateau: under steady traffic its retained memory settles at a level and stays there, rising and falling with load. A leak has a slope - retained memory climbs with cumulative work done rather than with concurrent work in flight, and never returns to its earlier floor. Measure after collection, over hours, at fixed traffic.
  • Is a cache with a maximum entry count safe from this?
    Bounded in entries is not bounded in bytes. A cap of ten thousand entries still leaks if entries can hold arbitrarily large values, or if each value transitively retains a whole object graph. Bound what actually matters - total retained size, or a time-based expiry - and be clear that a cap turns an unbounded leak into a fixed, budgeted cost.
  • Why does explicitly requesting a collection not help a leaking service?
    Because nothing about the leaked data is garbage. It is reachable from a live starting point, so any collection - requested or automatic - is obliged to keep it. A requested collection can be useful for measurement, since it gives a clean post-collection floor to trend, but it cannot reclaim retained memory and never will.

A left-luggage office only destroys a bag once nobody holds a ticket for it. The leak is not an abandoned bag - it is the drawer of tickets you kept for bags you will never collect.

saying these in an interview costs you the question

  • Claims a runtime with automatic reclamation cannot leak memory
  • Thinks leaked objects are unreachable and the collector simply missed them
  • Says an explicit collection request will reclaim the growing data
  • Blames the allocator or fragmentation before finding what holds the reference
  • Treats a cache as bounded because its entries feel temporary
  • Says raising the memory limit fixes it rather than delaying it
open as a page

A container's memory ceiling kills a worker whose managed heap sits well under its configured maximum — why?

level: middleimportance: must knowfreq 62%

basics

~20 s

A container ceiling counts every byte the whole process holds; the configured maximum heap bounds only one region inside it. Thread stacks, runtime metadata, allocator caches and buffers living outside the managed heap are charged against the ceiling too.

open as a page

On a service's memory footprint chart, why is the post-collection floor the line you trend to decide whether it is leaking?

level: middleimportance: must knowfreq 72%

basics

~20 s

The post-collection floor is the live set: the bytes still reachable once the collector has finished. Peaks only record how much garbage piled up since the previous collection, so only a rising floor shows that memory is being retained.

open as a page

A gateway's footprint saw-tooths between 1.2 GB and 3.0 GB every forty seconds while every post-collection floor sits at 1.2 GB — what is going on?

level: middleimportance: must knowfreq 60%

basics

~20 s

That is churn, not a leak. The flat 1.2 GB floor says the live set is constant; the 1.8 GB reclaimed every forty seconds says the service allocates roughly 45 MB a second of short-lived objects. The fix is allocation rate, not a retainer hunt.

open as a page

In a heap snapshot, what does an object's retained size measure that its shallow size does not?

level: middleimportance: must knowfreq 66%

basics

~20 s

Shallow size is the bytes of the object itself: header, fields and its reference slots, but not their targets. Retained size adds everything that would become unreachable if that object did, so it measures what the object keeps alive.

open as a page

Why does a discarded component stay in memory after it registered a handler with a long-lived event hub?

level: middleimportance: must knowfreq 58%

basics

~20 s

The hub outlives the component and still holds its handler. The handler captures the component, so a live path runs hub to handler to component, and reclamation keeps the whole object graph hanging off it.

open as a page

Why can a long-running renderer fail one 64 MB contiguous allocation while its memory report still shows 40% of the heap free?

level: middleimportance: must knowfreq 62%

basics

~20 s

An allocation needs one run of consecutive addresses, so it fails when no single free block is that large. The total-free figure is a sum over scattered holes and never promises a 64 MB run.

open as a page

Why does a heap snapshot's dominator tree tell you which single edge, if cut, would free a whole subgraph?

level: seniorimportance: must knowfreq 54%

basics

~20 s

An object x dominates y when every path from a root to y passes through x. Cutting the reference that makes x reachable therefore makes the whole subtree x dominates unreachable at once, which is what its retained size counts.

open as a page

Why does crossing a container's hard memory ceiling stop a process outright, while crossing the runtime's configured maximum heap does not?

level: middleimportance: should knowfreq 46%

basics

~20 s

The heap maximum is enforced inside the process by the runtime, which can collect, refuse an allocation and report the failure. A hard ceiling is enforced outside the process by the party that granted the memory, and nothing inside gets a turn to react.

open as a page

How do you work backwards from a hard five-hundred-megabyte container ceiling to the maximum heap you configure for a worker?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Treat the heap maximum as a remainder, not a choice. Measure every non-heap line at peak concurrency — stacks, runtime metadata, buffers outside the managed heap, allocator overhead — add an explicit margin, subtract the total from the ceiling, then check the remainder against the live set.

open as a page

How long must you watch a freshly restarted long-lived service before a rising post-collection floor entitles you to call it retention?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Long enough that every legitimate slow fill has plateaued — bounded caches, pools, lazily built structures — and then across several complete traffic cycles, comparing floors at the same phase of each cycle. Retention does not plateau; warm-up does.

open as a page

You have two heap snapshots of a search-indexing service taken an hour apart; how do you diff them to find what is growing?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Compare aggregates, not individual objects: per-type instance counts and shallow bytes, then retained size along the same root paths. The growing type plus a holder count that stayed flat points at one container accumulating entries.

open as a page

A leak is confirmed and memory grows with distinct items ever seen rather than with concurrent load - which retention shape does that point at?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Growth that tracks cumulative distinct items, not concurrency, points at a keyed container with no eviction - a cache or registry keyed by user, tenant or query. Confirm it by correlating the footprint with that container's entry count.

open as a page

Why can a value parked in per-thread storage leak when the thread comes from a pool rather than ending with the request?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Per-thread storage is scoped to the thread, not to the task. A pooled thread never ends, so a value set during one request stays attached to it - and reachable - until something explicitly removes it or a later task overwrites it.

open as a page

Why does a service's footprint stay at its spike-time peak after traffic returns to normal, even though its post-collection live set is back to the pre-spike value?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Footprint is a high-water mark. Freed objects return their bytes to the process's own free pool, not to the operating system, and a unit of memory can only go back if nothing live remains inside it.

open as a page

Your measured memory budget for a worker exceeds its container ceiling by twenty percent — which lever do you pull, and on what evidence?

level: principalimportance: should knowfreq 36%

basics

~20 s

Pick the lever from the shape of the budget, not from habit. Only shrinking the live set or the per-request working memory reduces real demand; raising the ceiling and re-splitting the work relocate it, and re-splitting multiplies fixed per-process costs.

open as a page

For a fleet of long-lived services, what would you page a human on: peak footprint, percentage of the limit, or the trend of the post-collection floor?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use two signals with different urgencies: the post-collection floor's slope opens a ticket days ahead because it detects retention early, and proximity to the limit pages, because it means failure is imminent. Peak footprint alone pages on healthy churn and should not.

open as a page

For a fleet of search-indexing services, what standing memory-diagnostics policy do you set when a full heap snapshot stalls the process for seconds and writes a file the size of the live set?

level: principalimportance: should knowfreq 33%

basics

~20 s

Keep something cheap always on and make the expensive capture deliberate: low-rate allocation sampling fleet-wide, snapshots taken from a drained instance, one at a time, rate-limited on the fatal path, and the files treated as a data export.

open as a page

A fleet fails one large contiguous allocation weekly per host with ample total free, so how do you choose between reshaping the request, recycling hosts and adopting a relocating manager?

level: principalimportance: should knowfreq 30%

basics

~20 s

Price the failure first, then remove the cause if you can: a request that need not be one run stops failing. Everything else — early reservation, recycling, continuous consolidation — treats a symptom at a different price.

open as a page

Why can raising a worker's configured maximum heap, inside an unchanged container ceiling, make the worker die sooner?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A maximum heap is permission to grow, not a reservation of what the program needs. Raise it and the runtime reclaims later, keeps unreclaimed objects resident longer, and lets the process footprint drift up toward a ceiling the heap setting does not know about.

open as a page

Why can a post-collection floor drift upward for a fortnight even though the service's live set never grew?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Because not every floor is the same measurement. A floor taken after a collection that visited only part of memory still contains garbage the collection never looked at, so a series of such floors can drift while the reachable set is flat. Compare floors of equal scope.

open as a page

Where does a sampled allocation profile mislead you compared with recording every allocation in a service?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Sampling one allocation per fixed interval of bytes estimates bytes per site well, but hides any site whose total allocation is far below one interval and makes object counts noisy. Both forms profile birth, not survival.

open as a page

Why can a service that never closes the file and socket handles it opens grow its footprint while its managed heap stays flat?

level: seniorimportance: nice to knowfreq 29%

basics

~20 s

A handle is a small managed object attached to a large resource the collector does not measure - kernel structures and off-heap buffers. Heap pressure therefore never rises, no collection is provoked, and the real memory is released only by an explicit close.

open as a page

Why can the largest free block keep shrinking in a heap whose manager is able to relocate objects to consolidate free space?

level: seniorimportance: nice to knowfreq 28%

basics

~10 s

Relocation needs permission. An object whose address has escaped the manager, or that is too costly to copy, is pinned in place, and one immovable object stops its whole unit being emptied and merged.

open as a page