skip to content

A leak is confirmed and memory grows with distinct items ever seen rather than with concurrent load - which retention shape does that point at?

level: seniorimportance: should knowfreq 41%

answer

  1. ask in step with what, not how much
  2. rises and falls means capacity
  3. cumulative distinct keys means a cache
  4. per lifecycle cycle means a registration
  5. export every container's entry count

basics

~20 s

Growth that tracks cumulative distinct items, not concurrency, points at a keyed container with no eviction - a cache or registry keyed by user, tenant or query. Confirm it by correlating the footprint with that container's entry count.

solid answer

~50 s

The shape of the growth curve names the suspect before you open any dump. Growth proportional to **distinct keys ever seen** means something keyed is being written on a miss and never trimmed: a cache without eviction or expiry, or a registry keyed by identity. Growth proportional to **objects created and destroyed** means a registration nobody removed, since each lifecycle cycle leaves one entry behind. Growth that steps up as workers are first used and then plateaus means a value parked in per-thread storage on a pooled thread, where the ceiling is pool size times value size. Growth that tracks **resources opened** rather than anything in the managed heap means a handle nobody closed. The confirmation is to export the entry count of each candidate container as a metric and check which one moves in lockstep with the footprint; that correlation is cheaper and less ambiguous than reasoning about sizes.

go deeper

for a junior

Remember the first question to ask about growing memory: what does the growth track? Something that rises and falls with load is normal working set, not a leak.

for a middle

Map each growth law to its shape - distinct keys to an untrimmed cache, lifecycle cycles to a forgotten registration, worker warm-up to per-thread storage - and say what each ceiling is.

for a senior

Demonstrate the method under incident pressure: instrument container sizes, correlate rather than reason, then reproduce with a synthetic workload that varies key novelty independently of concurrency.

for a principal

Make the signal a standing requirement: every long-lived container in the fleet exports its entry count and declares a bound, so retention is caught by a monotonic-growth alert rather than by an outage.

## The curve names the suspect Once a leak is established, the useful next question is not "how much" but **"in step with what"**. Retention shapes have different independent variables, and the variable is far more diagnostic than the size. Plot retained memory against the candidates - concurrent requests in flight, requests served, distinct keys ever requested, components created, handles opened - and one of them will be a straight line. | growth tracks | shape it points at | ceiling | |---|---|---| | distinct keys ever seen | a keyed cache or registry with no eviction or expiry | none; grows with the key space explored | | objects created and destroyed | a registration on a long-lived hub that is never removed | none; one entry per lifecycle cycle | | workers used for the first time | a value parked in per-thread storage on a pooled thread | climbs to a ceiling of pool size x retained value, then flattens | | resources opened | a handle that is never closed | bounded only by the limit that eventually refuses to open more | | concurrent requests in flight | nothing - this is capacity, not retention | falls again when load falls | That last row is the control. A number that rises **and falls** with load is the working set. Retention is defined by the not-falling. ## Why distinct keys is such a strong signal A cache written on every miss has exactly the growth law of the key space the workload explores. That produces a curve nothing else produces: - It is **sub-linear in traffic**. Doubling request volume against the same popular keys barely moves it, because those keys are already resident. Teams conclude "it is not traffic-related" and look elsewhere, when in fact it is a cache. - It is **linear in novelty**. A new tenant onboarding, a crawler walking an identifier space, a report that queries by date, or a retry storm with unique correlation identifiers each add keys that will never be requested again. - It **accelerates at the worst moment**, because incident traffic is unusual traffic, which is to say novel keys. The classic version is a cache keyed by something with unbounded cardinality - a full request, a generated identifier, an error message string - rather than by a bounded domain. A key space that is effectively infinite means the cache has no maximum size even in principle. ## Confirming it without guessing 1. **Instrument the candidates.** Export `size()` of every long-lived map, list and registry as a metric. This is a few lines and it is the single highest-value thing you can add to a service with a retention problem. 2. **Correlate, do not reason.** Overlay each entry count on the retained-memory trend. The container whose count moves in lockstep is the leak; sizes and averages can wait until you know which container it is. 3. **Check the key, not the value.** Once identified, ask what the key's cardinality actually is over a year. "User" sounds bounded and is not, if users churn. "Error message" sounds small and is unbounded once messages interpolate values. 4. **Test the hypothesis cheaply.** Drive a synthetic workload that touches many distinct keys at low concurrency. A cache-shaped leak grows immediately; a registration-shaped one does not move at all. ## Where two shapes look alike Two cases genuinely confuse: - **A registry keyed by instance identity** behaves like both shapes at once, because each new object is both a lifecycle event and a new key. Distinguish by whether the entry survives after the object is discarded - if the container keys on identity and never removes, it retains the key itself. - **A per-thread value whose contents accumulate** looks unbounded rather than plateaued, because the ceiling is pool size times a value that is itself growing. The tell is that the number of retained objects stays equal to the worker count while their size climbs. The discipline underneath all of this is that a retention bug is an **asymmetry between an add path and a remove path**, and the growth curve tells you which add path is running. Find the container, count its entries, and the missing removal is usually one line away from the addition.

  • Why can a cache-shaped leak look unrelated to traffic volume?
    Because its growth law is distinct keys, not requests. Doubling traffic against already-resident popular keys barely moves it, so load tests that repeat a small key set show nothing. It grows with novelty instead - new tenants, generated identifiers, date-ranged queries - which is why it often surfaces after an onboarding or a crawl rather than after a traffic increase.
  • A registry is keyed by object identity and entries are never removed. Which growth signature does it show?
    Both at once, which is what makes it confusing: every new object is a lifecycle event and a new key. The distinguishing test is whether entries survive the objects - an identity-keyed container that never removes retains the key itself, so the count keeps rising even though no object of that generation is still in use.
  • What is the cheapest instrumentation that makes these shapes visible?
    Export the entry count of every long-lived container as a metric and alert on monotonic growth over a window. It costs a few lines, needs no dump, and identifies the container by correlation rather than by inference - and it moves well before retained memory moves enough to page anyone.

saying these in an interview costs you the question

  • Reads a peak that falls again as evidence of retention
  • Assumes any growth must be proportional to request volume
  • Calls a per-thread value on a pooled worker unbounded growth
  • Treats an identifier-keyed cache as bounded because identifiers feel small
  • Guesses the container from object counts instead of correlating with entry counts
  • Believes a cache is safe once it has a maximum entry count, whatever the values hold