skip to content

questions

5

On a service's memory footprint chart, why is the post-collection floor the line you trend to decide whether it is leaking?

level: middleimportance: must knowfreq 72%

answer

  1. look at the low points
  2. peaks mix two different quantities
  3. teeth are about pacing, not keeping
  4. measure right after a collection finishes
  5. a rising bottom edge is the signal

basics

~20 s

The post-collection floor is the live set: the bytes still reachable once the collector has finished. Peaks only record how much garbage piled up since the previous collection, so only a rising floor shows that memory is being retained.

solid answer

~40 s

In a collected runtime, memory in use only rises between collections, because nothing is released the moment it becomes garbage. So the **peak** is the live set *plus* every byte allocated since the last collection — it moves with allocation rate as readily as with retention, and a change in it tells you nothing about which one moved. The **floor**, the value right after a collection completes, is approximately what was still reachable: the live set. Trending floors across days separates the two quantities. A floor that is flat while peaks are tall means the service churns; a floor that climbs means the reachable set is growing. To be worth trending, the floors must come from comparable collection events and must exclude the start-up period.

go deeper

for a junior

Remember that memory rising between collections is normal, and that the chart's low points, not its high points, say how much the service is really keeping.

for a middle

Be able to say why the peak is ambiguous: it is the live set plus everything allocated since the last collection, so it moves with allocation rate as easily as with retention.

for a senior

Show the method, not the impression: floors only, comparable collection events, warm-up discarded, same phase of the traffic cycle compared, and a growth rate in bytes per day rather than 'it looks like it climbs'.

for a principal

Decide what evidence your organisation accepts before anyone may declare a leak, and make sure the floor series is retained long enough that a fortnight of history already exists when the argument starts.

## Two readings on the same chart A service that runs on a collected runtime does not release memory at the instant an object becomes garbage. The runtime hands out memory as the program asks for it, and periodically a **tracing collector** decides which objects are still reachable from the program's roots and reclaims everything else. Between two collections the used-memory series is monotone: every allocation adds to it and nothing subtracts. That is why the chart has teeth rather than a smooth line. The two readings on that chart measure different things: - The **peak**, immediately before a collection, is `live set + all garbage created since the previous collection`. It is a sum of two independent quantities. - The **floor**, immediately after a collection completes, is approximately the **live set**: the bytes that were still reachable and therefore could not be reclaimed. Because the peak is a sum, a change in it is ambiguous. A service whose peak rises by a third may be keeping a third more objects alive, or may simply be allocating a third more temporary objects per request while keeping exactly the same amount reachable. The floor is the only one of the two readings that isolates retention. | reading | what it contains | what moves it | what it proves | |---|---|---|---| | peak before collection | live set plus accumulated garbage | allocation rate, collection spacing, headroom | how much room the runtime is allowed to fill | | floor after collection | the live set, approximately | retention only | whether reachable memory is growing | | average of the series | a blend of both | everything at once | nothing on its own | ## Why retention is the thing the floor exposes A leak in a collected runtime is not memory the collector forgot. The collector reclaims exactly what is unreachable; if something survives, it survived because a reference to it was still findable from a root. A leak is therefore **retention**: an object the program no longer needs but has not stopped referring to. Retention is visible as a live set that grows without bound — which is exactly the floor. This is also why the shape of the teeth is not evidence. Tall teeth over a level floor are a completely healthy service with a high allocation rate; the collector is doing its job perfectly and reclaiming everything each cycle. A shallow chart with a slowly rising floor is the dangerous one, even though it looks calmer. ## Reading the chart in practice 1. Take the used-memory value **immediately after** each collection, not a sampled gauge that may land anywhere in a cycle. 2. Keep only collections of **comparable scope** — a floor after a collection that visited the whole heap and one after a collection that visited part of it are not the same measurement. 3. Discard the start-up segment, where caches, pools and lazily built structures are still filling toward a legitimate steady state. 4. Fit a slope over days and compare points at the **same phase of the traffic cycle**, so a daily load pattern is not mistaken for a trend. 5. Classify: a flat floor with tall teeth is churn; a floor that climbs in steps that never come back down is retention. Stated as a rate, the result is actionable in a way that "memory looks high" never is. A floor that moves from 1.2 GB to 2.1 GB over about thirteen days is climbing roughly 90 MB a day, and that number both proves the diagnosis and says how long there is before the limit arrives. ## Traps around the measurement itself - **The process total is not the floor.** Runtimes differ in whether they hand reclaimed pages back to the platform; many keep what they have already taken. So the total a platform reports for the process can stay flat while the live set falls. Trend the runtime's own post-collection number. - **One collection is a point, not a trend.** Two floors an hour apart cannot distinguish retention from ordinary variation in what a service holds at different moments. - **A restart destroys the evidence.** It also "fixes" the symptom for exactly as long as it takes the retention to rebuild, which is why restarts are a mitigation and never a diagnosis. - **Pooled and cached objects are legitimately reachable.** A pool that has grown to its intended size raises the floor and is not a defect; what distinguishes retention is that it does not stop. The discipline is compact enough to state in one line, and it is the line an interviewer is listening for: judge memory by the bottom edge of the chart, never the top.

  • The runtime's post-collection number falls, but the memory total the platform reports for the process does not. Is that a leak?
    No. Runtimes differ in how eagerly they return reclaimed pages to the platform, and many hold on to memory they have already taken so they do not have to ask for it again. The process total therefore tracks the high-water mark of what the runtime asked for, not the live set. Trend the runtime's own post-collection value for the retention question.
  • You only have two collection events in your observation window. What can you conclude?
    Essentially nothing about retention. Two floors give you a line through two points, and normal variation — a burst of in-flight requests, a cache near its eviction horizon, a different point in the daily load pattern — easily moves a floor by more than a slow leak does in an hour. You need floors across several full traffic cycles before a slope means anything.

saying these in an interview costs you the question

  • Reads the peak as the amount of memory the service actually needs
  • Calls any climbing memory graph a leak without waiting for a collection
  • Believes a saw-tooth shape is itself evidence of a leak
  • Treats the memory the platform reports for the process as the live set
  • Restarts the service and reports the problem as fixed
open as a page

A gateway's footprint saw-tooths between 1.2 GB and 3.0 GB every forty seconds while every post-collection floor sits at 1.2 GB — what is going on?

level: middleimportance: must knowfreq 60%

basics

~20 s

That is churn, not a leak. The flat 1.2 GB floor says the live set is constant; the 1.8 GB reclaimed every forty seconds says the service allocates roughly 45 MB a second of short-lived objects. The fix is allocation rate, not a retainer hunt.

open as a page

How long must you watch a freshly restarted long-lived service before a rising post-collection floor entitles you to call it retention?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Long enough that every legitimate slow fill has plateaued — bounded caches, pools, lazily built structures — and then across several complete traffic cycles, comparing floors at the same phase of each cycle. Retention does not plateau; warm-up does.

open as a page

For a fleet of long-lived services, what would you page a human on: peak footprint, percentage of the limit, or the trend of the post-collection floor?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use two signals with different urgencies: the post-collection floor's slope opens a ticket days ahead because it detects retention early, and proximity to the limit pages, because it means failure is imminent. Peak footprint alone pages on healthy churn and should not.

open as a page

Why can a post-collection floor drift upward for a fortnight even though the service's live set never grew?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Because not every floor is the same measurement. A floor taken after a collection that visited only part of memory still contains garbage the collection never looked at, so a series of such floors can drift while the reachable set is flat. Compare floors of equal scope.

open as a page