skip to content

Why can a service that never closes the file and socket handles it opens grow its footprint while its managed heap stays flat?

level: seniorimportance: nice to knowfreq 29%

answer

  1. two accountings, one invisible
  2. small name, large resource
  3. no heap pressure, no collection
  4. unreachable is not released
  5. descriptor count fails before memory does

basics

~20 s

A handle is a small managed object attached to a large resource the collector does not measure - kernel structures and off-heap buffers. Heap pressure therefore never rises, no collection is provoked, and the real memory is released only by an explicit close.

solid answer

~50 s

Two accountings are in play. The collector sees a handle object of a few dozen bytes; the process holds a socket buffer, a mapped region or a kernel file description that may be orders of magnitude larger and that the heap accounting never counts. Because the cheap side is all the collector measures, opening handles in a loop creates almost no heap pressure, so collection is rarely provoked and there is nothing to prompt a release even where the runtime offers a last-resort cleanup hook - and ecosystems differ on whether they offer one at all, so depending on it is never sound. Meanwhile the process footprint climbs and a per-process limit on open descriptors eventually refuses new work, which usually shows up as connection failures before it shows up as memory. The release must be explicit and unconditional: close on a scope-exit path that runs on the error route too, and treat the open descriptor count as the metric that detects it.

go deeper

for a junior

Recall that automatic reclamation covers memory the runtime allocated, not files, sockets or buffers owned by the operating system; those need an explicit close.

for a middle

Explain the size mismatch: a tiny managed handle names a large external resource, so heap pressure never rises and collection is never provoked to help.

for a senior

Diagnose it from outside the heap - a rising open-descriptor count and growing non-managed memory against a flat managed heap - and know that the dropped-handle variant has no retainer to find.

for a principal

Set the standard: every external resource is acquired in a scope-bound construct with one owner, every pool has a maximum, and descriptor counts are monitored as a first-class capacity signal.

## Two accountings, one of which is invisible A handle is a small managed object that names something much larger living elsewhere: a kernel file description and its buffers, a socket with its send and receive queues, a mapped region, a decompressor's working memory, a driver-side statement. The managed heap accounts for the **name**. The operating system accounts for the **thing**. | | the handle object | the resource behind it | |---|---|---| | typical size | tens of bytes | kilobytes to megabytes | | who accounts for it | the managed heap | the kernel or a non-managed arena | | what releases it | becoming unreachable | an explicit close | | what a heap dump shows | a small object | nothing at all | This mismatch is the whole answer. Every mechanism that would normally rescue you is driven by the number the collector can see, and that number barely moves. ## Why nothing triggers a cleanup 1. **No heap pressure means no collection.** Collection is provoked by allocation pressure in the region the collector manages. Ten thousand handles might be a few hundred kilobytes of managed objects, which provokes nothing, while the resources behind them are hundreds of megabytes. 2. **Reachability is the wrong question for a resource.** Even if the handle object does become unreachable, unreachable means "may be reclaimed", not "is released now". The bytes the collector reclaims are the name's bytes. 3. **A last-resort cleanup hook is not a plan.** Some ecosystems run a cleanup action when an abandoned object is finally reclaimed; others decline to offer one at all, and where it exists it runs at an unspecified time, on an unspecified thread, possibly never before exit. Correctness that depends on it is correctness that depends on a collection that may not happen. 4. **The limit that bites first is not memory.** A per-process cap on open descriptors is usually reached long before the footprint becomes alarming, so the failure arrives as "cannot open" errors on new connections rather than as an out-of-memory condition - and that misdirection is why this shape is often diagnosed as a networking problem first. ## The retention shape underneath There are really two variants, and they call for different fixes: - **The handle is still reachable** - it was stashed in a list, a map or a per-connection context that nothing trims. This is an ordinary retention bug with an unusual multiplier: one small entry pins a large non-managed resource. Fix the container, and the close. - **The handle became unreachable without being closed** - the code simply dropped it, perhaps on an error path. Nothing in the object graph retains anything, yet the resource is still held, because releasing it was never the collector's job. The second variant is the one that makes this shape worth understanding: it is a leak with no retainer to find. Searching a heap snapshot for what holds the memory yields nothing, because the memory was never on the heap. ## Making the release deterministic 1. **Bind the close to a scope.** Acquire in a construct whose exit always releases, so the release runs on the error path as well as the success path. This is the resource-acquisition-is-initialization discipline, and it is the only approach that does not depend on anyone remembering. 2. **Give every handle exactly one owner.** Ambiguous ownership produces both double closes and no closes; a handle passed to a callee either transfers ownership explicitly or is borrowed for the duration of the call. 3. **Close in the reverse order of opening,** so that a wrapper never outlives what it wraps and closing the outer layer cannot leave an inner one dangling. 4. **Monitor the right number.** Track open descriptors per process and non-managed memory alongside heap usage. A monotonically rising descriptor count is the earliest and cleanest signal this shape produces. 5. **Cap the pool.** Where handles are pooled deliberately, the pool's maximum is the bound on the resource; without a maximum, a pool is just a leak with a friendly name. The general lesson generalises past files and sockets: **automatic reclamation is about memory the runtime allocated, and nothing else.** Any resource whose release the operating system or a non-managed allocator owns needs an explicit, unconditional release path, no matter how automatic the rest of the memory story is.

  • Why can a heap snapshot fail to explain this leak at all?
    Because the memory was never on the managed heap. In the variant where the handle was dropped without being closed, there is no retainer to find - no container, no path, nothing holding anything. The snapshot honestly reports a small, healthy heap while the process footprint and the descriptor count both climb.
  • Which signal detects this shape earliest?
    The count of open descriptors for the process, trended over time. It rises immediately and monotonically, long before the footprint is alarming, and it distinguishes this shape from managed-heap retention in one graph. Non-managed memory usage is the natural companion metric.
  • Why is a cleanup hook that runs when an abandoned handle is reclaimed not a sufficient fix?
    It runs at an unspecified time, on an unspecified thread, and only if a collection actually reclaims the object - which handle-heavy code barely provokes, since the managed objects are tiny. Ecosystems also differ on whether such a hook exists at all. Treat it as a last-resort safety net that reports a bug, never as the release mechanism.

A cloakroom ticket weighs nothing, but each one holds a full locker. Throwing the tickets away does not empty the lockers; only handing them back does.

saying these in an interview costs you the question

  • Assumes automatic reclamation releases operating-system resources
  • Thinks dropping the last reference to a handle closes it
  • Relies on a last-resort cleanup hook as the normal release path
  • Expects the managed heap graph to show the growth
  • Closes only on the success path, not on the error path
  • Calls an unbounded handle pool a cache rather than a leak