In PyTorch, what does an in-place op like add_() risk that add() does not?
answer
- trailing underscore means mutate in place
- returns the same tensor, not a new one
- every view of that storage changes too
- x += 1 and x = x + 1 differ
- version counter guards the backward pass
basics
~20 sA trailing underscore means the op overwrites the tensor's own storage instead of allocating a result. That saves an allocation but mutates every view and alias of that tensor, and it destroys values that a later backward pass may still need.
solid answer
~50 sBy convention every PyTorch method ending in an underscore — `add_`, `mul_`, `clamp_`, `zero_`, `copy_`, `relu_` — writes into the receiver's existing storage and returns that same tensor, whereas `add`, `mul`, `torch.clamp` and friends allocate a fresh result and leave the input untouched. Two risks follow. First, **aliasing**: views share storage, so `x[0].add_(5)` modifies `x`, and a tensor you handed to a caller, cached, or exposed through `Tensor.numpy()` changes underneath them. Second, **autograd**: if the overwritten value was recorded for the backward pass, PyTorch raises a RuntimeError saying one of the variables needed for gradient computation has been modified by an in-place operation. In-place is worth it where it demonstrably cuts peak memory — long activation chains, `nn.ReLU(inplace=True)`, optimizer updates — but it is not free savings, because the framework may keep the pre-op value anyway.
code
python · 10 linesimport torch
x = torch.ones(2, 3)
row = x[0]
row.add_(5) # writes through the view
print(x[0]) # tensor([6., 6., 6.])
y = torch.ones(2, 3)
y[0].add(5) # out-of-place: result discarded
print(y[0]) # tensor([1., 1., 1.])go deeper
Know the naming convention: a trailing underscore means the tensor is changed in place and no new tensor is returned. Be able to point at add_ versus add and say which one leaves the input alone.
Explain both hazards concretely — writing through a view mutates the parent, and overwriting a value autograd saved triggers the version-counter error — and note that x += 1 is in-place while x = x + 1 is not.
Show when you would actually choose it: optimizer-style updates, inference activation chains where peak memory was measured, buffers you allocated yourself, and never on a tensor handed in by a caller.
Set the team convention. Default to out-of-place in shared model code, allow in-place only behind a measured memory justification, and treat the autograd in-place error as a correctness bug to be diagnosed rather than silenced with a clone.
## The naming convention PyTorch marks every mutating method with a trailing underscore. `t.add_(1)` adds one to `t` and returns `t` itself; `t.add(1)` returns a new tensor and leaves `t` alone. The same pairing exists for `mul_`, `div_`, `sub_`, `clamp_`, `zero_`, `fill_`, `normal_`, `uniform_`, `relu_`, `sigmoid_`, `copy_`, `masked_fill_` and `scatter_`. Assignment forms are also in-place: `x += 1` calls `add_`, and `x[0] = 3` writes into `x`'s storage. Notably `x = x + 1` is **not** in-place — it rebinds the name to a new tensor while the old one lives on if anything else references it. Explaining that distinction is often the whole question. Some functional APIs take an `inplace` flag instead of the underscore: `torch.nn.ReLU(inplace=True)`, `torch.nn.Dropout(inplace=True)` and `torch.nn.functional.relu(x, inplace=True)`. Many ops also accept an `out=` argument, which writes into a preallocated destination — a third mutation channel with the same aliasing implications. ## Risk one: aliasing A view shares storage with its parent, so an in-place write through a view is a write to the parent. `row = x[0]; row.zero_()` zeroes the first row of `x`. That is often the intent, and it is how you fill a preallocated buffer efficiently. It becomes a bug when the two references belong to different owners: a tensor returned from a function and cached by the caller, a batch element stored in a metrics accumulator, or a NumPy array produced by `Tensor.numpy()`, which shares the same buffer and will change when the tensor is mutated. The discipline is that a function should mutate only tensors it created or ones the caller explicitly passed as a destination. Anything else is a hidden side effect on data with an unknown number of aliases. ## Risk two: autograd correctness Autograd records the tensors it needs to compute gradients. Some backward formulas need the operation's *input*, some need its *output* (the derivative of a sigmoid, for example, is expressed in terms of the sigmoid's output). If you overwrite a tensor that was saved for backward, PyTorch does not silently compute a wrong gradient: it version-counts every storage and raises at `backward()` time, reporting that a variable needed for gradient computation has been modified by an in-place operation and naming the expected version. That check is a feature — treat the error as a real correctness violation, not an obstacle to route around. The common triggers are mutating a tensor after using it in a computation, reusing one buffer across iterations, and `inplace=True` on an activation whose input is also consumed by a skip connection. ## What in-place actually saves It saves an allocation and the memory traffic of writing a second buffer. In a deep chain of elementwise activations during **inference**, that can meaningfully reduce peak memory, which is why `inplace=True` is a standard trick in vision backbones. During **training** the picture is muddier: if autograd needs the pre-op value, it keeps a copy anyway, and you have paid the complexity without the saving. Optimizers are the clean case — `param.add_(grad, alpha=-lr)` inside a no-grad update is exactly what in-place is for, and it is what the built-in optimizers do internally. Speed gains are usually small on modern hardware, since the dominant cost is memory bandwidth either way and the allocator is a caching one. Reach for in-place when a memory measurement calls for it, not as a default style. ## A practical rule set - Prefer out-of-place by default; the code is easier to reason about and safe under autograd. - Use in-place for buffers you own and just allocated, for optimizer-style updates under no-grad, and for inference-time activation chains where you have measured the peak-memory win. - Never mutate a tensor you received as an argument unless the API says it is an output parameter. - Remember that `x[i] = v`, `x += v`, `out=` and `inplace=True` are all the same mechanism wearing different clothes. - When an in-place op triggers the autograd version-counter error, fix it by making the op out-of-place, not by cloning blindly until the error goes away — the clone may be hiding a genuine ordering mistake.
- Is x += 1 the same as x = x + 1 in PyTorch?No. `x += 1` calls `add_` and writes into `x`'s existing storage, so every view and alias sees the change and the version counter increments. `x = x + 1` allocates a new tensor and rebinds the name; the old tensor is untouched and survives if anything else still references it. Under autograd, the first can raise a version-counter error where the second never will.
- When is inplace=True on an activation layer actually worth it?Mainly at inference or in a long chain of elementwise activations where you have measured peak memory and the saved intermediate is large. In training it often saves nothing, because autograd keeps the value it needs regardless, and it breaks outright when the same input feeds a skip connection or another consumer. Measure before adopting it as a default.
- What does an in-place op do to a NumPy array obtained from the same tensor?It changes it. `Tensor.numpy()` shares the underlying buffer rather than copying, so mutating the tensor mutates the array and vice versa. If you hand a NumPy view to code that assumes a stable snapshot — a logging thread, a plotting call, a queue — copy explicitly with `.clone().numpy()` or `np.array(...)` at that boundary.
saying these in an interview costs you the question
- Thinks in-place is always faster and lower memory
- Says x += 1 and x = x + 1 are identical
- Forgets that views share storage with the parent
- Clones randomly until the autograd error disappears
- Mutates a tensor received as a function argument