skip to content

Where does a sampled allocation profile mislead you compared with recording every allocation in a service?

level: seniorimportance: nice to knowfreq 28%

answer

  1. cheap and continuous versus exact and costly
  2. sampled by bytes, not by calls
  3. small total, never sampled
  4. counts noisier than byte totals
  5. birth, not survival

basics

~20 s

Sampling one allocation per fixed interval of bytes estimates bytes per site well, but hides any site whose total allocation is far below one interval and makes object counts noisy. Both forms profile birth, not survival.

solid answer

~50 s

An allocation profiler records the call site of allocations. **Full recording** sees every one, at a cost: the instrumentation slows the program, produces enormous traces, and can suppress optimisations that would have removed an allocation entirely, so the trace lists objects the uninstrumented program never creates. **Sampling** takes one allocation per interval of bytes allocated. That makes total-bytes estimates per site reasonably faithful, because the chance of being sampled scales with the bytes a site allocates — but a site whose whole lifetime allocation is far below one interval is usually never sampled, and instance *counts* are far noisier than byte totals. The deeper limit applies to both: an allocation profile says where objects were **born**, not what keeps survivors reachable. A leak is a retention fact, so it is found in a snapshot, not in an allocation profile.

code

pseudocode · 14 lines
pseudocode
interval  = 512 KiB               // one sample per this many bytes
countdown = random_gap(mean = interval)

on allocate(site, size):
    countdown = countdown - size
    if countdown > 0:
        return                    // not sampled; this site stays invisible
    weight = max(size, interval)  // scale the sample back up
    record(site, estimated_bytes = weight)
    countdown = random_gap(mean = interval)

// A site whose TOTAL allocation is far below one interval
// almost never drives the countdown to zero, so it is absent
// from the profile without ever being inactive.

go deeper

for a junior

Know the distinction in one line: an allocation profile says where objects are created, a heap snapshot says what is still holding them. They answer different questions and you need the second one for a leak.

for a middle

Explain the sampling rule — one sample per interval of bytes — and what follows from it: byte totals per site are reasonably faithful, small-volume sites go missing, and instance counts are the noisiest thing you can read off it.

for a senior

Show when you reach for which. Continuous sampling to watch churn and spot a rising site; a snapshot when the question is survival. Say plainly that full recording perturbs the program, including by preventing optimisations that would have removed allocations.

for a principal

Decide the standing fidelity for a fleet: what interval, retained for how long, and what evidence promotes a suspicious profile into an expensive snapshot. State that estimates are estimates and that a snapshot outranks them when the two disagree.

## Two ways to watch allocation Allocation profiling attributes bytes to the code that requested them. There are two implementations and they fail differently. **Full recording** instruments every allocation. It is exact, and that exactness costs: - per-allocation overhead on the hottest path in the program, which distorts timing and therefore behaviour; - trace volume that makes long captures impractical; - an observer effect at the optimiser level — escape analysis and scalar replacement can remove an allocation altogether in an uninstrumented run, and instrumentation that forces every allocation to be observable can prevent exactly that, so the trace reports allocations the real program never performs. **Sampling** takes one allocation per interval of bytes — the countdown is decremented by each allocation's size and a sample is recorded when it runs out. Overhead is a fraction of a percent and capture can stay on continuously. ## What sampling gets right Because the sampling decision is driven by *bytes*, a site's chance of being sampled scales with the bytes it allocates. Estimated bytes per site therefore converge on the truth as the run lengthens. In practice, byte-interval sampling is trustworthy for: - which call sites dominate total allocation volume; - the relative ranking of the top sites; - large, infrequent allocations, which are sampled almost whenever they happen. ## Where it skews | Situation | What sampling reports | Why | |---|---|---| | A site allocating far below one interval in total | nothing at all | the countdown rarely reaches zero on its allocations | | Many tiny objects, few bytes | understated instance count | samples are proportional to bytes, not to calls | | One huge array | reported, with a large weight | size alone can exhaust the countdown | | A short capture window | noisy rankings below the top few | too few samples for the estimate to settle | | A shared factory or helper frame | bytes attributed to the helper | stack depth too shallow to reach the caller | Two of these deserve emphasis. First, **counts are not bytes**: converting a byte estimate into an object count divides by an assumed average size, and a site with a mixture of sizes gives a count you should not quote. Second, **absence is not evidence**: a site missing from a sampled profile has allocated few bytes, which is not the same as being inactive. ## The limit both forms share Even a perfect allocation profile answers the wrong question for a leak: 1. It records the **birth** of an object, attributed to the stack that requested it. 2. A leak is a fact about **survival** — something continues to reference the object after its purpose ended. 3. The code that allocates and the code that retains are routinely far apart: a generic parsing routine allocates the entries, and a completely different component keeps a map of them. So allocation profiles and heap snapshots are complements, not substitutes: - **Allocation profile** — cheap, continuous, answers *who is producing garbage* and therefore *where the churn is*. - **Heap snapshot with retained sizes** — expensive, point-in-time, answers *who is holding memory* and therefore *where the leak is*. The sequence a practised engineer uses is: sampled profiling always on, so that a rise in allocation volume at a suspicious site is visible immediately; and a snapshot taken when the question is what is still alive, because no allocation profile of any fidelity will ever answer that. ## A note on interpreting weights Sampled profiles report *estimated* bytes, scaled up from the samples taken. Comparing two such estimates is fine; treating one as an exact measurement is not. When an estimate and a snapshot disagree, the snapshot is the measurement — it counted the objects that were actually there.

  • Why does lowering the sampling interval not simply make the profile correct?
    It trades overhead for resolution. A smaller interval means more samples, more instrumentation cost on the hot path, and a larger profile, moving back toward full recording's distortion. It also cannot fix the attribution limit: more samples still describe where objects were born, never what retains them.
  • A site is absent from a sampled profile. What have you actually learned?
    Only that it allocated few bytes during the window. It may run constantly on tiny objects, or run rarely, or have run before capture started. Absence bounds a site's byte volume; it says nothing about call frequency, about survival, or about whether that site is involved in a leak.

saying these in an interview costs you the question

  • Hunts a leak in an allocation profile instead of a snapshot
  • Reads a sampled byte estimate as an exact measurement
  • Treats absence from a profile as proof a site is inactive
  • Converts sampled bytes into an object count and quotes it
  • Assumes full recording is simply the accurate version of sampling