Why does a service's footprint stay at its spike-time peak after traffic returns to normal, even though its post-collection live set is back to the pre-spike value?
answer
- footprint records the past
- high-water mark, not current need
- freed is not returned
- release needs a wholly empty unit
- one survivor keeps a unit mapped
basics
~20 sFootprint is a high-water mark. Freed objects return their bytes to the process's own free pool, not to the operating system, and a unit of memory can only go back if nothing live remains inside it.
solid answer
~40 sTwo different things are being plotted. The live set is what the program needs now; the footprint is the largest demand it has ever had to satisfy. Reclaiming the spike's objects put their bytes back into the manager's own pool, where the next allocation can use them and nobody outside the process can. Handing space back to the operating system happens in whole units, and a unit is only releasable when no live object sits in it. After a spike the survivors are sprinkled across almost every unit, so a heap that is mostly free can hand back nothing. The reading is therefore: a permanent step in footprint with no step in the live set is retained free space, not something being kept alive.
go deeper
Learn that memory a program has taken is not automatically given back when objects die. The bytes go into the program's own pool for reuse, so the number the operating system reports can stay high after the work is finished.
Explain the release rule: space goes back in whole units and a unit needs nothing live inside it. Then explain why a spike leaves survivors scattered across nearly every unit, so a mostly-empty heap can release nothing.
Demonstrate the reading. Put footprint and post-collection live set on one axis, annotate the traffic event, and state what the step in one without a step in the other means before anyone starts hunting through the code for a retained reference.
Decide whether the retained high-water mark costs the fleet anything. Weigh host density and headroom against the cost of releasing and reacquiring space at every burst, and set the policy deliberately rather than inheriting a default.
## Footprint records the peak, not the present The memory a process holds is space it has taken from the operating system. Taking it is a request the program makes; giving it back is a separate request the program must also make, and managers make it rarely. Freeing an object returns its bytes to **the manager's own pool of free space**, where they are available to the next allocation inside the same process and to nothing else. So the footprint curve records the largest demand the process has ever had to satisfy, while the post-collection live set records what it needs now. During the spike the program genuinely needed the peak. When the spike ends the objects die, reclamation runs, the live set falls back to its old value, and the pool that grew to hold those objects stays exactly as large as it grew. ## Why the pool does not shrink back Space goes back to the operating system in whole units: pages, or the larger runs a manager carves its heap into. A unit can be released only if **nothing live is inside it**, and after a spike that condition is rarely met. - The spike's short-lived objects were interleaved with ordinary long-lived ones, because allocation order is arrival order. - Each unit therefore ends the spike holding a handful of survivors and a lot of free space. - A unit that is 98 percent free is still not releasable; one survivor is enough to keep it. - Multiplied over hundreds of units, a heap that is mostly empty can hand back nothing at all. - Releasing also costs. A manager that returns space eagerly must take it again at the next burst, so many deliberately keep it. | Curve | What it measures | What a permanent step up means | |---|---|---| | Process footprint | Space taken from the operating system, ever | Peak demand has moved; nothing is necessarily wrong | | Post-collection live set | Bytes still reachable after reclamation | Something is being kept that was not kept before | | The gap between them | Free space retained inside the process | Space is held and is probably scattered | ## Reading the two curves 1. Plot footprint and post-collection live set on one axis. A step in the first with no step in the second is the signature of this symptom. 2. Check that the step coincides with a known event. A step with no spike behind it is a different story with a different owner. 3. Ask whether the retained gap is usable. It serves ordinary allocations perfectly well, so throughput and pause behaviour are usually unaffected. 4. Ask whether it is contiguous. If the gap is large but the largest free block is small, the retained space has also shattered, and the next large request is at risk. ## When the gap is harmless and when it is not Most of the time this is benign and even desirable: the process paid once for peak capacity and will not pay again at the next spike. Teams nonetheless treat any flat-high footprint as a defect, so it is worth being explicit about the three situations where it genuinely bites. - **Density.** Hosts are packed by footprint, so memory the process will never hand back is memory another workload cannot use. - **Ceilings.** A limit chosen from steady-state observation now sits permanently close to the process's own high-water mark. - **Shape.** The retained space can be large and still unable to serve one large run, which is the failure that makes this symptom visible in the first place. Ecosystems differ in the default: some runtimes return space eagerly and pay to reacquire it, others hold it for the life of the process by design, and some expose the choice as a policy. The reading above is identical in all of them; only the default answer changes. ## The wrong conclusion to draw The tempting conclusion is that the spike leaked. It did not, and the flat live set is what says so: the objects it allocated were reclaimed. What survived the spike is not objects but **shape** — a pool that grew, and a scattering of survivors that stops it shrinking again. Naming that correctly matters, because the two diagnoses send you to completely different work: one into the code, hunting for something still referenced, the other to the allocation pattern, the release policy and the free-space distribution. The practical follow-through is small. Record footprint and post-collection live set together, annotate the graph with traffic events, and treat a step in one without a step in the other as expected behaviour to be understood rather than a bug to be hunted — unless density, a ceiling or a large contiguous request makes the retained space actually cost something.
- Does the gap between footprint and live set ever close on its own?Usually only partially. New allocations are served from the retained pool, so the footprint simply stays flat while work continues. It falls only when the manager can concentrate survivors and hand whole units back, which needs consolidation plus a release policy willing to pay for reacquiring the space later.
- What would make the same graph worth escalating?Three things: the host is packed by footprint so the retained space blocks other workloads, the limit for the process now sits close to the new high-water mark, or the largest contiguous free block has fallen even though the gap is large. The last one is the warning that a big allocation will fail next.
- Why does a redeploy appear to fix it?Because a fresh process starts with an unbroken heap and no history, so its footprint reflects current demand instead of the largest demand ever seen. That is why a restart is a diagnostic signal rather than a fix: it resets the high-water mark and the free-space shape together, and both will drift back with uptime.
saying these in an interview costs you the question
- Calls a flat-high footprint a leak without checking the live set
- Assumes freeing an object hands its bytes straight to the operating system
- Expects the footprint curve to follow the live set down as a rule
- Thinks a partly-free unit can be returned to the operating system
- Reads the gap as wasted memory that must be a bug