skip to content

questions

4

Why can a long-running renderer fail one 64 MB contiguous allocation while its memory report still shows 40% of the heap free?

level: middleimportance: must knowfreq 62%

answer

  1. two free-space numbers, not one
  2. the sum is not the shape
  3. contiguity, not quantity
  4. check the largest free block
  5. request exceeds every single hole

basics

~20 s

An allocation needs one run of consecutive addresses, so it fails when no single free block is that large. The total-free figure is a sum over scattered holes and never promises a 64 MB run.

solid answer

~40 s

The manager is not checking the request against total free space; it is looking for one contiguous block of at least 64 MB. On a process that has run for days, free space arrives as many separate holes, because every object that outlives its neighbours permanently divides the space around it. The sum of those holes can be huge while the biggest single one is 44 MB, and a 64 MB request then fails. The give-away is that the failures are size-selective: small allocations keep succeeding and only the largest request dies. The number to look at is the `largest contiguous free block`, trended against the largest request the program makes — total free can stay perfectly flat while that one collapses.

code

pseudocode · 12 lines
pseudocode
free_blocks = [12MB, 8MB, 30MB, 6MB, 44MB, 10MB]   // total free = 110MB
request     = 64MB

best = 0
for each b in free_blocks:
    if size(b) > best:
        best = size(b)            // best = 44MB

if request > best:
    fail(out_of_memory)           // fires: 64MB > 44MB, with 110MB free
else:
    return carve(best_block, request)

go deeper

for a junior

Remember that memory is handed out in one continuous piece. A program asking for a big buffer needs that many bytes side by side, so a number saying how much is free in total does not tell you whether the request can be met.

for a middle

Explain the mechanics: free space is a set of holes, holes merge only when adjacent, surviving objects act as permanent dividers, and the request is matched against the largest hole. Name the measure you would look at instead of total free.

for a senior

Show the diagnosis. Capture the failing request size, compare it with the largest contiguous free block at that moment, trend both, and demonstrate that failures are size-selective before proposing anything. Say what you would alert on.

for a principal

Frame it as a capacity question with an unusual variable: uptime. Decide whether the service pays for consolidation, reshapes its requests, or recycles hosts, and insist on the leading indicator that shows whichever choice was made is working.

## The request is for a run, not for a quantity A memory manager hands out **contiguous** space: a request for 64 MB is a request for 64 MB at consecutive addresses, because the caller will index into it with address arithmetic. So the question the manager asks is not *do I have 64 MB free anywhere* but *do I have one free block of at least 64 MB*. Those are different questions, and on a long-running process they drift apart. Two measures describe the same heap, and a third explains it: | Measure | What it answers | How it misleads | |---|---|---| | **Total free bytes** | How much space is unused in total | Says nothing about the shape of that space; it is a sum over holes of every size | | **Largest contiguous free block** | The biggest single request that can still be served | Can collapse while total free stays perfectly flat | | **Free-block size histogram** | How the free space is distributed | Needs trending to be useful; one sample rarely alarms | The dashboard almost always shows the first one. The allocation is decided by the second. ## How ample free space becomes unusable free space - Objects with different lifetimes are allocated next to each other, because allocation order is arrival order, not lifetime order. - When a short-lived object is freed it leaves a hole exactly its own size. - Two holes that touch merge into one larger hole; two holes separated by a surviving object do not. - Survivors are therefore dividers, and every survivor that outlives its neighbours permanently splits the space around it. - Over hours and days the distribution shifts from a few large holes towards many small ones while the total barely moves. - Small requests keep succeeding throughout, because a small request has many holes that can serve it. That is why the first allocation to fail is almost always the largest one the program makes, and why the failure looks absurd next to a graph showing 40 percent free. ## A worked example A heap reports 110 MB free, held as blocks of 44, 30, 12, 10, 8 and 6 MB. A 64 MB request arrives: - The sum, 110 MB, is comfortably above the request, and that is the number on the graph. - The largest single block is 44 MB, which is 20 MB short. - Nothing in the manager's bookkeeping can combine the 44 and the 30 into a 74 MB run, because live objects sit between them. Blocks merge only when they are adjacent. - The request fails while the process still has 40 percent of its heap free. Halve the request to 32 MB and it succeeds against the very same heap; halve it again and six different blocks could serve it. The failure is a property of the pair (request size, free-space shape), not of either one alone. ## Reading the symptom in order 1. Record the size of the request that failed. A failure with no size attached cannot be diagnosed. 2. Compare that size with the largest contiguous free block at the moment of failure. If the block is smaller than the request, the diagnosis is already made. 3. Trend both numbers over days. **Total free flat, largest free block declining** is the signature of free space shattering. 4. Check whether the failures are size-selective. If small allocations keep succeeding while only the big buffer dies, the heap is not short of space, it is short of runs. ## What this symptom is not - It is not the host running out of memory. That is a different accounting with a different owner; here the program's own free space is ample by its own report. - It is not on its own evidence that anything is being kept forever. A heap can shatter with a completely flat set of live objects. - It is not automatically fixed by raising the ceiling. Extra space helps only if it restores a run large enough for the request, which is worth measuring rather than assuming: a bigger heap that shatters at the same rate buys time and nothing else. ## Why it belongs to long-lived processes A process restarted every few hours never shows this, because a fresh heap is one unbroken run and holes have had no time to accumulate. The failure belongs to processes that live for days: a renderer, a long-lived worker, a cache host. The unintuitive consequence is that **uptime is a variable in whether an allocation succeeds**, which contradicts the mental model of an allocation as a pure function of the size requested and the memory installed. The operational answer follows from the reading: put the largest contiguous free block on the same graph as the largest request the program makes, and alert on the gap between them rather than on total free. That one pairing turns an inexplicable weekly outage into a slope you can watch approach.

  • Which single number would you add to a dashboard to see this failure coming?
    The largest contiguous free block, sampled per process and trended over days, plotted against the largest request the program makes. Total free is the number teams already have and it is the one that stays flat while the failure develops. A free-block size histogram is a useful second view, but the gap between the largest block and the largest request is what actually predicts the outage.
  • Does the same failure mode exist for small allocations?
    In principle yes, in practice far more rarely. The smaller the request, the more holes can serve it, so the probability of finding a fit stays high until the free space is shattered into truly tiny pieces. That is why failures are size-selective and why the first casualty is almost always the biggest buffer the program asks for.
  • Does a heap that cannot serve a large request always recover after a collection?
    Not necessarily. Reclaiming dead objects returns bytes but does not by itself move the survivors, so the holes stay where they were and the largest block can be unchanged after the cycle. Only consolidation that relocates live objects rebuilds large runs, and whether the manager does that at all is the thing to establish before assuming the next cycle will fix it.

A car park with two hundred scattered empty bays is still full for a bus: the total says yes, and the bus needs eleven bays in a row.

saying these in an interview costs you the question

  • Assumes an allocation is checked against total free bytes
  • Concludes the host is out of memory and just raises the limit
  • Thinks several free blocks can be chained to serve one request
  • Calls every allocation failure a leak without checking sizes
  • Believes a fresh collection cycle always restores a large run
  • Ignores that only the largest requests are failing
open as a page

Why does a service's footprint stay at its spike-time peak after traffic returns to normal, even though its post-collection live set is back to the pre-spike value?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Footprint is a high-water mark. Freed objects return their bytes to the process's own free pool, not to the operating system, and a unit of memory can only go back if nothing live remains inside it.

open as a page

A fleet fails one large contiguous allocation weekly per host with ample total free, so how do you choose between reshaping the request, recycling hosts and adopting a relocating manager?

level: principalimportance: should knowfreq 30%

basics

~20 s

Price the failure first, then remove the cause if you can: a request that need not be one run stops failing. Everything else — early reservation, recycling, continuous consolidation — treats a symptom at a different price.

open as a page

Why can the largest free block keep shrinking in a heap whose manager is able to relocate objects to consolidate free space?

level: seniorimportance: nice to knowfreq 28%

basics

~10 s

Relocation needs permission. An object whose address has escaped the manager, or that is too costly to copy, is pinned in place, and one immovable object stops its whole unit being emptied and merged.

open as a page