skip to content

A gateway's footprint saw-tooths between 1.2 GB and 3.0 GB every forty seconds while every post-collection floor sits at 1.2 GB — what is going on?

level: middleimportance: must knowfreq 60%

answer

  1. two features: the teeth and the floor
  2. flat bottom, tall teeth
  3. amplitude divided by cycle time
  4. no growth means nothing retained
  5. the lever is bytes per request

basics

~20 s

That is churn, not a leak. The flat 1.2 GB floor says the live set is constant; the 1.8 GB reclaimed every forty seconds says the service allocates roughly 45 MB a second of short-lived objects. The fix is allocation rate, not a retainer hunt.

solid answer

~40 s

The two axes of the shape say different things. The **floor** is flat, so the reachable set is not growing — nothing is being retained. The **amplitude**, 1.8 GB reclaimed per cycle in forty seconds, is an allocation rate of about 45 MB per second, essentially all of it dying before the next collection. That is a healthy collector doing a lot of work. The cost is not growth but the processor time spent collecting and the headroom needed to hold a cycle's worth of garbage. The fix is to allocate less per unit of work: reuse buffers instead of allocating one per request, avoid materialising whole payloads to read a field, stop copying between representations. Changing how much room the runtime may fill re-paces the teeth; it does not move the floor.

code

pseudocode · 18 lines
pseudocode
floors = empty list
reclaimed = 0

for each collection event e in window:
    if e.scope is not WHOLE_HEAP:
        continue                          # a partial floor still holds uncollected garbage
    floors.append(pair(e.time, e.bytes_after))
    reclaimed = reclaimed + (e.bytes_before - e.bytes_after)

trend = slope_per_day(floors)             # how fast the live set is growing
churn = reclaimed / duration(window)      # bytes reclaimed per second

if trend > noise_band:
    report "retention: live set climbing", trend
else if churn > allocation_budget:
    report "churn: live set flat, allocation rate high", churn
else:
    report "flat floor, modest churn: nothing to chase"

go deeper

for a junior

Learn to look at the bottom of the saw-tooth first. A flat bottom with tall teeth means the service creates a lot of short-lived objects and the collector is keeping up.

for a middle

Compute the rate out of the picture: amplitude divided by cycle time is the allocation rate, and be able to say why that number, not the peak, is what a code change moves.

for a senior

Demonstrate that you choose the investigation from the shape — churn sends you to the hot request path and its per-request allocations, retention sends you looking for what still holds a reference.

for a principal

The judgment is whether the churn is worth engineering time at all: weigh the processor cost of the collection work against the cost of the code changes that would cut allocation per request.

## Two shapes, two diagnoses, two different fixes A footprint chart of a long-lived service has exactly two features worth naming, and they are independent: - **The teeth** — the rise and fall between collections. Their height is the amount of garbage created per cycle; their spacing is how long that takes to accumulate. - **The floor** — the sequence of values immediately after each collection. That is the live set. | shape | what it says | what drives it | the fix | |---|---|---|---| | saw-tooth over a level floor | churn: high allocation rate, constant live set | bytes allocated per unit of work | allocate less per request | | staircase floor that never returns | retention: the reachable set is growing | references that are never dropped | find and break the retaining reference | | level floor, shallow teeth | neither | nothing to chase | leave it alone | Confusing the two wastes days. Hunting a retainer in a service that is merely churning finds nothing, because nothing is being retained. Cutting allocation rate in a service that is genuinely retaining delays the arrival of the limit slightly and fixes nothing, because the floor keeps climbing at the same slope. ## Reading the numbers out of this chart In the chart described, each cycle reclaims `3.0 GB - 1.2 GB = 1.8 GB`, and a cycle takes forty seconds. So the service allocates and discards about `1800 MB / 40 s = 45 MB` per second. Every one of those bytes is dead by the time the next collection looks at it, otherwise the floor would have moved. Across a minute that is roughly 2.7 GB of allocation supporting a live set of 1.2 GB — a turnover ratio, and the ratio, not the absolute footprint, is the thing to argue about. Three consequences follow from a high rate at a flat floor: 1. **Processor time.** Collection work scales with what has to be traced and copied, and a fast allocation rate means more cycles per hour, each doing that work again. 2. **Room.** The runtime must have somewhere to put 1.8 GB of garbage between collections. A tighter limit does not reduce the allocation; it makes the same allocation force more frequent collections. 3. **Nothing about failure.** Churn at a flat floor does not trend toward the limit. It can be expensive, and it can be entirely acceptable. ## What actually reduces churn The lever is bytes allocated per unit of work, and it lives in the code path, not in configuration: - **Reuse a buffer per worker** rather than allocating a fresh one per request, so the buffer becomes part of the live set once instead of garbage every time. - **Parse lazily.** Materialising an entire message to read one field allocates the whole structure for the lifetime of one comparison. - **Stop re-encoding.** Converting between representations on the way in and again on the way out doubles the garbage for no change in meaning. - **Avoid defensive copies** on paths where the data is never mutated afterwards. A subtlety worth stating: some of these fixes deliberately move bytes *into* the live set. A reused buffer per worker raises the floor a little and removes a large amount of churn. That is a good trade and it is also a reminder that a floor moving once, to a new level, and then staying there is not retention. Retention is a floor that keeps moving. ## Why this is the first question to settle Runtimes differ in how much of this they hide. Some make short-lived allocation extremely cheap and collect young objects so efficiently that 45 MB a second is unremarkable; others make the same rate painful. What does not differ is the reading: the floor answers "is anything being kept?" and the teeth answer "how fast is work creating garbage?". Settle which of the two you are looking at before opening any tool, because the two questions have no diagnostic steps in common. The honest conclusion for the chart above is that the service is not leaking, and whether the churn is worth attacking depends on what the collection work is costing — an argument you make with a rate, not with a screenshot of the chart's tallest point.

  • Would giving the runtime more room to fill make this chart better?
    It would change the pacing, not the diagnosis. More room means more garbage accumulates before each collection, so the teeth grow taller and arrive further apart; the same 45 MB per second is still being allocated and the floor still sits at 1.2 GB. It buys fewer, larger collection cycles rather than less work, and it consumes footprint to do so.
  • The floor steps once from 1.2 GB to 1.35 GB after a change and then stays level for a week. Is that a leak?
    No. Retention is a floor that keeps moving; this one moved once and settled, which is what a larger steady-state structure looks like — a bigger pool, a cache with a higher ceiling, an extra per-worker buffer. Record the new level as the baseline and keep trending from there. Only a slope that persists across cycles is retention.
  • What does the ratio of allocated bytes to live bytes tell you that neither number tells you alone?
    It tells you how much garbage the service creates to maintain a given working set, which is the quantity a code change can actually move. A 1.2 GB live set is a design fact; 2.7 GB of allocation a minute to serve it is a code-path fact, and comparing the two across releases catches a regression that neither absolute number reveals.

A bathtub with the taps running hard and the plug out holds a steady depth: alarming to listen to, but the water level is not rising. A tub filling a centimetre a day with the taps barely on is the one that floods the room.

saying these in an interview costs you the question

  • Calls a saw-tooth footprint a leak because the peaks are high
  • Starts hunting a retaining reference while the floor is flat
  • Thinks a higher memory limit reduces how much the code allocates
  • Assumes high allocation rate must eventually exhaust the limit
  • Reports the peak as the service's memory requirement