Your training job reports 22 GB reserved but 9 GB allocated, then fails to allocate — why?
answer
- two numbers, one device
- the gap is cached, not free
- one tensor needs one contiguous span
- shape churn leaves unusable holes
basics
~20 sReserved is what the caching allocator holds from the driver; allocated is what live tensors use. The gap is cached blocks of the wrong sizes, and a new tensor needs one contiguous block none of them can supply.
solid answer
~50 sThe two numbers measure different things. `allocated` is the sum of live tensors — parameters, gradients, optimizer state, saved activations. `reserved` is the total the process has taken from the driver and not returned; the gap is the allocator's own cache of free blocks it keeps so it does not pay the driver's cost on every step. An allocation fails when the requested size needs one contiguous block, no cached block is that large, and the driver has no unclaimed memory left to carve a new segment from. So the error means "no contiguous block this big", not "no free bytes". Variable-length batches are the usual aggravator: a new shape every step splits blocks into remainders that never recombine. Fix it by making the requested shapes repeat, or by allocating the largest shape first so the big block exists and is recycled.
go deeper
Be ready to say that the two numbers are not the same: one counts live tensors, the other counts what the process holds from the driver. Knowing the gap is a cache, not a leak, is enough at this level.
Explain the mechanics: why a caching allocator exists, why a tensor needs a contiguous block, and how shape churn splits blocks into remainders that never recombine. This is the tier where the interviewer expects the word fragmentation with a reason attached.
Show the diagnosis. Compare the failing request against the gap, decide fragmentation versus a genuine ceiling, and pick the matching fix — repeating shapes and a peak-shaped warm-up for one, the memory levers for the other.
Own the policy. Decide whether jobs may share a device at all, what shape budget the data pipeline guarantees, and whether the team's fitting margin is measured against the peak rather than the average — a run that fits only when the cache is tidy is not reproducibly sized.
## Two numbers that measure different things A deep-learning training process does not hand every tensor request straight to the device driver. It runs a **caching allocator** in between, and that allocator publishes two totals which interviewers love to contrast. - **Allocated** is the sum of the tensors that are alive right now: parameters, gradients, optimizer state, the activations saved for the backward pass, the current batch, and anything your own code is still holding a reference to. - **Reserved** is the total the process has taken from the driver and has *not* given back. Reserved is always greater than or equal to allocated. The difference — 13 GB in the numbers above — is memory the process owns but is not currently lending to any tensor. It is *cached*, not leaked and not in use by anyone else. ## Why the cache exists Asking the driver for memory is expensive and tends to synchronize the device, which stalls the pipeline. A training loop allocates and frees the same handful of shapes thousands of times per epoch, so the allocator grabs large segments once, cuts them into blocks, and recycles those blocks across steps. In steady state you see allocated oscillate within a step — rising through the forward pass as activations accumulate, falling through the backward pass as they are consumed — while reserved climbs to a high-water plateau and stays there. ## Why an allocation still fails with free memory on the device A tensor needs a **contiguous** span of device memory. When a request arrives, the allocator looks for a cached free block big enough; if none fits, it asks the driver for a fresh segment; if the driver has nothing unclaimed left, the request fails. That is the whole story of the confusing case: several gigabytes are free *in aggregate*, spread across blocks of 40 MB, 180 MB and 600 MB, and the request wants 1.5 GB in one piece. This is **fragmentation**, and the message it produces is indistinguishable from a genuine capacity problem unless you read the two numbers. ## What makes fragmentation bad Repeating shapes are the allocator's friend: the same block is freed and re-taken every step forever. Fragmentation comes from *shape churn*. A corpus of variable-length text where every batch has a different sequence length produces a slightly different request every step; blocks get split, the leftover remainders are too small for the next request, and they never coalesce back into a large block because a long-lived tensor sits in the middle of the segment. Mixed input resolutions do the same thing. Two processes sharing one device compound it, because the driver has less to grant when a new segment is needed. ## How to tell fragmentation from a real capacity problem Log three numbers per step: allocated, reserved, and **peak** allocated. - Large reserved-minus-allocated gap and a *modest* failing request → fragmentation. Attack the allocation pattern. - Allocated itself sitting just under the device's capacity at the moment of failure → a genuine capacity problem. No allocator trick saves you; you need the memory levers (smaller micro-batch, accumulation, recomputation, lower precision, sharding). - Reserved creeping up epoch after epoch while allocated stays flat → the distribution of requested shapes is widening over time. ## What actually helps **Make the shapes repeat.** Fix a per-batch token or pixel budget so the allocator sees a small set of sizes instead of a continuum. This is the highest-leverage fix and it costs nothing at runtime. **Allocate the peak shape first.** Running one warm-up step at the largest shape you will ever see forces the big block into existence early, while the segment is still uncarved; every later step reuses it. **Release the cache back to the driver — sparingly.** Returning cached blocks helps when *another process* needs the memory, or when a new, larger segment must be formed. Inside a single process it is usually pointless: the allocator would have reused those blocks anyway, and you pay the driver's cost to re-acquire them. Doing it every step is a classic way to make a run slow without making it fit. ## What not to conclude The gap is not a leak, it is not other processes, and it is not the size of your model. And a run that fits only because the cache happened to be tidy is not really fitting — the peak, not the average, is the number that has to fit.
- When does releasing the allocator's cached blocks back to the driver actually help?When something outside your process needs the memory — a second training job, an evaluation process, a data-loading worker on the same device — or when a new, larger segment has to be formed and the driver has nothing left to grant. Inside one process it is usually a no-op that costs re-acquisition time, because the allocator would have reused those blocks itself.
- Why do variable-length text batches fragment the allocator worse than fixed-size image batches?Fixed shapes mean the same request every step, so the same block is freed and re-taken indefinitely. Variable lengths mean a new request size every step: blocks are split, the remainders are too small for the next request, and they cannot coalesce while a long-lived tensor pins the middle of the segment. Free bytes accumulate in pieces nobody can use.
- How would you prove a failure is fragmentation rather than a genuine capacity ceiling?Compare the size of the failing request with the reserved-minus-allocated gap at the moment of failure. If the gap comfortably exceeds the request, the bytes exist but not contiguously — fragmentation. If allocated is already brushing the device's capacity, it is a real ceiling and only the memory levers help.
A car park can have 300 free spaces and still turn away a coach: the coach needs six adjacent bays, and the free spaces are scattered one at a time across four floors.
saying these in an interview costs you the question
- Says the reserved-minus-allocated gap is leaked memory
- Assumes free bytes on the device mean an allocation will succeed
- Blames model size whenever an out-of-memory error appears
- Clears the allocator cache every step and calls it a fix
- Thinks reserved memory is held by other processes on the device