In PyTorch, why can keeping one small slice of a big tensor retain all its memory?
answer
- descriptor and buffer are separate objects
- refcounting happens at the storage level
- slices are views, fancy indexing copies
- compare nbytes with storage nbytes
- clone at the boundary where lifetimes diverge
basics
~20 sBasic indexing and slicing return views that share the original tensor's storage buffer. That buffer is freed only when the last tensor referencing it is gone, so one saved row pins the entire allocation. Call .clone() to detach the slice onto its own storage.
solid answer
~50 sA tensor is a shape/stride descriptor over a storage buffer, and refcounting is per **storage**, not per tensor. `row = big[0]` is a view: even though `row` reports a few kilobytes from `numel() * element_size()`, `row.untyped_storage().nbytes()` still reports the whole allocation, and deleting `big` frees nothing while `row` lives. The production shape of this is a loop that appends `output[0]`, `batch[:, :10]` or a cropped patch to a results list — each entry pins a full batch, and memory grows linearly with iterations even though the retained data is tiny. On GPU this shows up as a slow climb toward CUDA out-of-memory rather than a leak in Python's heap. The fix is to force a copy: `.clone()`, or `.detach().clone()` when the slice also carries a graph reference, or `.contiguous()` when the slice happens to be non-contiguous. Advanced indexing (`x[mask]`, `x[idx_tensor]`) already copies, so it does not have this problem.
code
python · 9 linesimport torch
big = torch.zeros(10_000, 512)
row = big[0]
print(row.numel() * row.element_size()) # 2048 logical bytes
print(row.untyped_storage().nbytes()) # 20480000 bytes still held
row_copy = big[0].clone()
print(row_copy.untyped_storage().nbytes()) # 2048 - independent buffergo deeper
Know that slicing a tensor gives you a view onto the same data, not a copy, and that .clone() is how you get an independent tensor. That alone prevents the most common version of this bug.
Explain that storage is refcounted separately from the tensor descriptor, classify which ops are views and which copy, and show the nbytes-versus-storage-nbytes comparison that proves it.
Walk the diagnosis end to end on a real service: sample allocated bytes per iteration, find the long-lived container, prove the retention by inspecting storage size, and fix it by cloning at the lifetime boundary rather than everywhere.
Make it a design rule rather than a bug hunt — define where results cross out of a batch's lifetime and require a copy at that boundary, and keep memory-per-iteration in the standard metrics so this class of regression is caught by monitoring, not by an outage.
## Two objects, not one Every PyTorch tensor consists of a lightweight descriptor — shape, strides, storage offset, dtype, device — and a reference to a storage buffer holding the raw bytes. Many descriptors can point at the same buffer. Python's reference counting applies to both levels: the descriptor dies when nothing references the tensor, but the buffer only dies when no surviving descriptor points at it. That decoupling is what makes views free, and it is also what makes this retention bug possible. ## Which operations produce a view Basic indexing and slicing — `x[0]`, `x[2:5]`, `x[:, :10]`, `x[..., 0]` — all return views, as do `narrow`, `select`, `transpose`, `permute`, `squeeze`, `unsqueeze`, `expand`, `view`, `t()` and `detach`. `reshape` and `flatten` return a view when the layout permits and a copy otherwise. Advanced indexing is the important exception: `x[torch.tensor([0, 3, 7])]` and `x[mask]` gather elements into a fresh buffer, so they copy. So does `clone`, `repeat`, `cat`, `stack` and any arithmetic that produces a new result. A candidate who can classify these three groups on demand is telling you they have debugged real memory problems. ## Seeing the retention The measurement that makes it obvious is comparing the logical size of a tensor with the size of the buffer behind it: - logical: `t.numel() * t.element_size()`, or the `t.nbytes` attribute; - actual: `t.untyped_storage().nbytes()`. For a row of a `(10000, 512)` float32 tensor those differ by four orders of magnitude. `t.data_ptr()` and `t.untyped_storage().data_ptr()` let you confirm two tensors share a buffer. On CUDA, `torch.cuda.memory_allocated()` reports live allocated bytes and `torch.cuda.memory_reserved()` reports what the caching allocator holds from the driver; a retention bug moves the first number, so watching `memory_allocated()` across iterations is the fast triage. ## The failure mode in production The canonical version is an evaluation or inference loop: ``` for batch in loader: out = model(batch) results.append(out[0]) # a view of the whole batch output ``` Each appended element is a few kilobytes of data holding an entire batch's buffer hostage. After a few thousand iterations the process is out of memory, and the heap profile looks bizarre — a list of small tensors consuming gigabytes. The same shape appears when caching a cropped image patch from a large decoded frame, when storing `logits[:, -1]` from every step of a generation loop, or when a metric object keeps per-batch predictions. It also appears across the NumPy boundary: `Tensor.numpy()` shares the buffer with the tensor, so a NumPy array made from a slice keeps the tensor's storage alive too, and vice versa via `torch.from_numpy`. ## Fixing it Copy at the point of retention, not at the point of use. `.clone()` allocates a buffer sized to the slice and copies the elements; the original storage is then free to die with its last real reference. If the slice came out of a graph-connected computation, `.detach().clone()` also drops the reference to the autograd graph, which can retain far more than the tensor itself. On CPU workloads where the slice is non-contiguous, `.contiguous()` gives you a compact copy as a side effect, but say what you mean and use `clone()`. Copying costs memory bandwidth, so do not clone reflexively — clone at boundaries where a small object outlives a large one. Inside a computation, views are exactly what you want. ## Diagnosing it when you did not write the loop Watch live allocated bytes per iteration; if it climbs monotonically and a garbage collection does not reclaim it, look for long-lived containers holding tensors. For each suspect, compare its logical bytes with its storage bytes. When the storage is much larger than the tensor, you have found it. On GPU, remember that `memory_reserved` staying high after the leak is fixed is normal caching allocator behaviour, not a second leak; `torch.cuda.empty_cache()` returns it to the driver but is rarely needed and does not fix retention.
- How would you confirm at runtime that a tensor is holding a much larger buffer than its shape suggests?Compare `t.numel() * t.element_size()` (or `t.nbytes`) against `t.untyped_storage().nbytes()`. A large gap means the tensor is a view into a bigger allocation. `t.data_ptr()` and `t.untyped_storage().data_ptr()` confirm which tensors share the same buffer, and on GPU `torch.cuda.memory_allocated()` sampled per iteration shows whether the retention is actually growing.
- Does x[mask] have the same problem as x[0]?No. Boolean masks and index tensors are advanced indexing, which cannot generally be expressed with strides, so PyTorch gathers the selected elements into a fresh buffer. The result is already a copy and holds nothing else alive. That also means advanced indexing is not free — it allocates — which is the tradeoff in the other direction.
- Why is .detach().clone() often the right form rather than plain .clone() when saving results from a loop?`clone()` copies the data but the result is still connected to the computation that produced it, so a graph-connected tensor kept in a list can retain the whole graph and its saved activations — far more memory than the buffer itself. `detach()` cuts that link. In an evaluation loop, running under an inference context and detaching before saving avoids both retention paths.
- Your GPU memory_reserved stays high after you fixed the retention. Is that a second leak?No. PyTorch's caching allocator keeps freed blocks reserved from the driver so later allocations are fast, so `memory_reserved` is a high-water mark rather than live usage. Watch `memory_allocated` instead. `torch.cuda.empty_cache()` returns cached blocks to the driver but costs performance and only matters when another process needs that GPU memory.
saying these in an interview costs you the question
- Thinks a small tensor can only hold small memory
- Assumes deleting the parent tensor frees its buffer
- Believes all indexing returns a copy
- Clones every tensor defensively to be safe
- Blames the caching allocator for a retention bug