In PyTorch, how do tensor.detach() and torch.no_grad() differ?
answer
- one tensor versus one code region
- shared storage, cleared metadata
- no intermediates saved is the memory win
- stop-gradient for teachers and targets
- neither one touches dropout
basics
~20 sdetach() acts on one tensor: it returns a view sharing the same data but cut off from the graph. torch.no_grad() is a context manager that switches off graph recording for every operation inside it. Use detach to stop gradients at a point; use no_grad to run a whole block untracked.
solid answer
~50 s`x.detach()` is **per-tensor and surgical**. It returns a new tensor that shares `x`'s storage — same memory, no copy — but with `requires_grad=False` and no `grad_fn`, so backward stops there. Everything computed *from* the detached tensor is untracked, while operations elsewhere in the same expression still build a graph. That is how you implement a stop-gradient: target networks, teacher outputs in distillation, or logging a value without keeping the graph alive. `torch.no_grad()` is **scoped and wholesale**. Inside the block, autograd records nothing at all: outputs come back with `requires_grad=False` and no intermediates are saved, which is the real memory saving. It is the right tool for evaluation, inference, and manual parameter updates. `torch.inference_mode()` is the stricter, faster variant for pure inference — it also skips version tracking, but tensors created inside it cannot later participate in autograd.
code
python · 13 linesimport torch
x = torch.ones(3, requires_grad=True)
d = x.detach()
print(d.requires_grad, d.data_ptr() == x.data_ptr()) # False True
with torch.no_grad():
y = x * 2
print(y.requires_grad, y.grad_fn) # False None
z = x * 2
print(z.requires_grad, z.grad_fn is not None) # True Truego deeper
Know that no_grad wraps a block of code you do not want gradients for, such as evaluation, while detach() cuts a single tensor out of the graph.
Explain that detach shares storage and only clears autograd metadata, and that no_grad's real benefit is not saving intermediate activations. Say clearly that neither affects dropout or batch-norm.
Give real stop-gradient use cases — target networks, distillation teachers, truncated BPTT — and pick inference_mode for serving while knowing its restriction on tensors that re-enter autograd.
Own it as a memory and cost policy: which regions of a system are recorded decides evaluation batch size and serving throughput. Standardise inference paths on inference_mode and make stop-gradient boundaries explicit in model code rather than incidental.
## Two different questions The distinction is scope. `detach()` answers "should gradient flow *through this tensor*?" `no_grad()` answers "should anything in *this region of code* be recorded at all?" They frequently produce the same visible result — a tensor with `requires_grad=False` — which is why they get confused, but they solve different problems. ## detach(): a stop-gradient marker ``` y = model(x) z = y.detach() * 2 + other(x) ``` `y.detach()` returns a tensor pointing at the *same underlying storage* as `y`. No data is copied. What differs is the metadata: the result is a leaf with `requires_grad=False` and `grad_fn=None`. Autograd, walking backwards from a loss built on `z`, will happily flow into `other(x)` but hits a wall at the detached branch. Because storage is shared, two consequences matter in practice: - **Mutating one mutates the other.** `y.detach().zero_()` zeroes `y` too. If you need an independent copy, use `.detach().clone()`. - **In-place edits on a detached tensor can corrupt a pending backward pass.** Autograd keeps a version counter per storage; if a tensor that a backward node saved is modified in place, backward raises `RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation`. Detaching does not exempt you from that check, because the storage is the same. Canonical uses: freezing the target branch of a DQN or an EMA teacher; passing a generator's output into a discriminator without updating the generator; truncated backpropagation through time, where you detach the hidden state between chunks; and `loss.detach()` (or `loss.item()`) when accumulating metrics. ## no_grad(): switching the recorder off ``` with torch.no_grad(): preds = model(val_batch) ``` Inside the block, autograd is disabled thread-locally. Ops on grad-requiring tensors still compute correct numbers, but no `grad_fn` is attached and, crucially, **no intermediate activations are saved**. That is where the memory win comes from: a forward pass under `no_grad` holds only what the next op needs, rather than everything the backward pass would have wanted. On a large model this is often the difference between a batch size of 8 and 64 at evaluation time. It is also the standard wrapper for hand-written parameter updates, because writing `p -= lr * p.grad` on a leaf that requires grad would otherwise raise `RuntimeError: a leaf Variable that requires grad is being used in an in-place operation`. `torch.no_grad()` can be used as a decorator (`@torch.no_grad()`) as well as a context manager, and its counterpart `torch.enable_grad()` re-enables recording inside a `no_grad` region — useful when a library disables grad globally but you need one differentiable sub-call. `torch.set_grad_enabled(flag)` takes a boolean, which is handy for a shared train/eval function. ## inference_mode(): the stricter cousin `torch.inference_mode()` goes further than `no_grad`: it also skips the per-tensor version counter and view-tracking bookkeeping, so it is measurably faster. The price is that tensors *created* inside an inference-mode block are marked as inference tensors and cannot later be used in autograd or mutated outside; you get a runtime error if you try. Use it for serving paths where the result is consumed and discarded, and stay on `no_grad` when a produced tensor may re-enter a training graph — for example when generating pseudo-labels. ## What neither of them does They do not change layer behaviour. Dropout still drops and batch-norm still updates its running statistics under `no_grad`; that is governed by the module's train/eval mode, an entirely separate switch. A candidate who says "I used no_grad so dropout is off" is describing a bug, and it is the single most common confusion around this pair. They also do not free memory that is already held. If a stale graph is being kept alive by a Python reference elsewhere, wrapping later code in `no_grad` does nothing for it. ## Picking one A rough rule: if the answer to "which tensors" is *these specific ones, mid-graph*, use `detach()`. If it is *everything I am about to do*, use `no_grad()` — and `inference_mode()` when that region is a pure inference path whose outputs never re-enter training. Nesting them is fine and common: an evaluation loop under `no_grad`, and a training step that detaches a teacher output while the rest stays fully tracked.
- Does detach() copy the tensor's data?No. The returned tensor shares the same storage; only the autograd metadata differs. That makes it cheap, but it means an in-place edit on either one is visible through the other, and modifying a tensor that a backward node saved will trigger the version-counter error. When you need an independent buffer — say, to stash a snapshot of weights — use `.detach().clone()`.
- Why use torch.inference_mode() over torch.no_grad() in a serving path?Inference mode disables version-counter and view bookkeeping in addition to graph recording, so it has less per-op overhead. The restriction is that tensors created inside are tagged as inference tensors and cannot later be used in an autograd graph or mutated outside the block. For a request handler that returns numbers and forgets them, that restriction costs nothing; for pseudo-label generation that feeds training, stay on no_grad.
- You wrapped evaluation in torch.no_grad() but results still vary between runs. What did you miss?no_grad only turns off gradient recording — it does not change layer behaviour. Dropout is still sampling and batch-norm is still using and updating batch statistics unless the module is in eval mode. The two switches are independent and both are required: eval mode for deterministic layer behaviour, no_grad for the memory and speed saving.
- Inside a no_grad block, how do you re-enable gradients for one sub-computation?Wrap it in `torch.enable_grad()`, which is the inverse context manager and restores recording within the enclosing no_grad region. `torch.set_grad_enabled(bool)` does the same with a runtime flag, which suits a shared function that runs in both training and evaluation. Neither can resurrect tracking for tensors already produced under no_grad — those are plain leaves now.
saying these in an interview costs you the question
- Saying no_grad disables dropout or batch-norm updates
- Believing detach() returns an independent copy of the data
- Wrapping a training step in no_grad and wondering why nothing learns
- Using detach() where the whole eval loop should be under no_grad, so memory still balloons
- Thinking inference_mode is just an alias for no_grad