Why is .grad None on a PyTorch tensor after calling backward()?
answer
- only leaves keep a gradient
- grad_fn present means intermediate
- retain_grad if you want it kept
- detach and numpy sever the chain
- None is not the same as zero
basics
~20 sAutograd only fills .grad on leaf tensors that require grad. A tensor produced by an operation is a non-leaf, so its gradient is used during the backward pass and discarded — call retain_grad() on it to keep a copy. .grad is also None if the tensor never required grad or was cut off from the graph.
solid answer
~40 sThere are three usual causes. First, the tensor is a **non-leaf**: it was produced by an operation, so it has a `grad_fn`, and autograd computes its gradient only as an intermediate and then throws it away. `t.retain_grad()` before `backward()` makes it keep one. Second, the tensor never had `requires_grad=True` — a plain `torch.randn` or a parameter you froze — so no backward node was ever created for it. Third, the chain from the loss to the tensor is **broken**: something called `.detach()`, wrapped the forward in `torch.no_grad()`, converted through NumPy or a Python float, or you reassigned the parameter instead of updating it in place. Diagnose by printing `t.is_leaf`, `t.requires_grad` and `t.grad_fn`: a leaf that requires grad and still shows `None` means the loss simply does not depend on it.
code
python · 14 linesimport torch
a = torch.tensor([2.0], requires_grad=True)
b = a * 3 # non-leaf: has grad_fn
b.retain_grad() # ask autograd to keep b's gradient
c = (b ** 2).sum()
c.backward()
print(a.is_leaf, a.grad) # True tensor([36.])
print(b.is_leaf, b.grad) # False tensor([12.])
d = torch.tensor([1.0], requires_grad=True)
(d.detach() * 5).sum() # nothing recorded: d is cut off
print(d.grad) # Nonego deeper
Remember that .grad is filled only on leaf tensors you created with requires_grad=True, and that a tensor produced by an operation will show None.
Explain leaf vs non-leaf via grad_fn and is_leaf, name retain_grad() and torch.autograd.grad() as the ways to see an intermediate gradient, and list what severs a graph.
Show a diagnostic routine: sweep named_parameters() for None grads to locate the disconnected subtree, and separate a None gradient (broken wiring) from a zero gradient (dead units or masking).
Insist that training code fails loudly on this class of bug rather than degrading quietly: an assertion that every expected-trainable parameter received a gradient after the first step turns a silent week-long regression into an immediate failure, and it belongs in the shared training harness rather than in each team's loop.
## Leaves and non-leaves Autograd divides tensors into two classes. A **leaf** is a tensor you created directly — `torch.randn(3, requires_grad=True)`, an `nn.Parameter`, an input batch. Its `is_leaf` is `True` and its `grad_fn` is `None`, because no recorded operation produced it. A **non-leaf** is anything that came out of an op on a grad-requiring tensor; it carries a `grad_fn` and `is_leaf` is `False`. When you call `.backward()`, autograd computes a gradient for *every* node it walks through — it has to, that is how the chain rule composes. But it only **stores** the result on leaves that require grad. The intermediate gradients are passed to the next node and released, because keeping them would multiply memory for no benefit in the normal training case, where only parameter gradients are consumed by the optimizer. So this is expected behaviour, not a bug: ``` a = torch.tensor([2.0], requires_grad=True) b = a * 3 # non-leaf c = (b ** 2).sum() c.backward() a.grad # tensor([36.]) b.grad # None, plus a UserWarning ``` Modern PyTorch emits a warning when you access `.grad` on a non-leaf precisely because this trips people up. The fix, when you genuinely want the intermediate gradient (probing activations, visualizing saliency, debugging a layer), is `b.retain_grad()` *before* the backward call. An alternative that avoids storing anything is `torch.autograd.grad(c, b)`, which returns the gradient as a value instead of writing it into `.grad`. ## Cause two: the tensor never required grad `.grad` stays `None` forever on a tensor whose `requires_grad` is `False`, because no backward node referencing it was ever created. This bites in two situations. One is a freshly created tensor where you forgot the flag. The other is a parameter you froze earlier — `p.requires_grad_(False)` — and then wondered why the optimizer is not moving it. A related trap: `.grad` is `None`, not zeros, before the first backward pass. Code like `if p.grad.norm() > 1:` crashes with `AttributeError: 'NoneType' object has no attribute 'norm'` on the very first iteration. Modern `zero_grad()` also sets gradients back to `None` rather than to zero tensors by default, so `None` is a normal steady state between steps, not evidence of a broken graph. ## Cause three: the graph is severed The most interesting failure. The tensor is a leaf, it requires grad, and `.grad` is still `None` — which means the loss you called `backward()` on genuinely does not depend on it, through a *recorded* path. Ways that happens: - **An explicit `.detach()`** somewhere in the forward path. Detach returns a tensor sharing the same storage with `requires_grad=False`, so numerics are unchanged and nothing raises — the gradient just stops there. - **A `torch.no_grad()` block** wrapping part of the model. Everything computed inside is untracked. - **A round trip through NumPy or Python scalars** — `.numpy()`, `.item()`, `.tolist()`, or `int(...)`. Autograd cannot follow numbers out of the framework and back. - **Rebinding instead of mutating.** Writing `layer.weight = layer.weight - lr * g` replaces the `nn.Parameter` object with a plain non-leaf tensor; the optimizer still points at the old object. In-place updates under `no_grad`, or just using an optimizer, avoid this. - **A branch that was never taken.** With a dynamic graph, a parameter used only inside an `if` that was false this iteration gets no gradient at all. That is correct, and it is why some sparse or mixture models legitimately show `None` for some parameters on some steps. ## The diagnostic sequence When a gradient is missing, do this in order: 1. Print `t.is_leaf`, `t.requires_grad`, `t.grad_fn`. That immediately separates "non-leaf", "not tracked" and "severed". 2. Print `loss.grad_fn` and walk `loss.grad_fn.next_functions` a level or two, or simply check `loss.requires_grad` — if the loss itself does not require grad, the whole forward pass was untracked. 3. Sweep the model: `[n for n, p in model.named_parameters() if p.grad is None]` after a backward pass names exactly which subtree is disconnected, which usually points straight at the offending line. 4. Distinguish `None` from **zero**. A gradient of `0.0` means the path exists but the derivative vanished — a saturated activation, a dead ReLU, a multiplication by a zeroed mask. That is a modelling problem, not a wiring problem, and the fix is completely different. That last distinction is the one interviewers listen for. `None` means autograd never got there; zero means it got there and the math said zero.
- How would you get the gradient of an intermediate activation without changing your training loop?Two options. `activation.retain_grad()` before `backward()` makes autograd store the gradient on that tensor. Or use `torch.autograd.grad(loss, activation)`, which returns the gradient directly and does not populate any `.grad` field — cleaner for one-off inspection because it leaves no state behind. A tensor hook registered with `t.register_hook(fn)` is the third route when you want to observe every backward pass without keeping the values.
- A parameter's grad is all zeros rather than None. What does that tell you?That the path from loss to parameter exists and autograd traversed it — the derivative genuinely evaluated to zero. Typical causes are saturated activations, dead ReLUs whose inputs are all negative, a mask or padding that zeroes the contribution, or a loss term with a zero coefficient. It is a modelling or data issue, whereas None is a wiring issue, and confusing the two sends you debugging the wrong half of the system.
- Why does converting a tensor to a NumPy array mid-model break gradients silently?Autograd only records operations executed by PyTorch ops on PyTorch tensors. Once you call `.numpy()` you are in NumPy's world, and the tensor you build from the result is a fresh leaf with `requires_grad=False`. Nothing raises because the forward numbers are correct; the loss simply no longer depends on anything upstream, so those parameters get `None`.
saying these in an interview costs you the question
- Assuming every tensor stores a gradient after backward()
- Treating .grad of None as identical to a zero gradient
- Not knowing retain_grad() exists for intermediates
- Blaming a vanishing-gradient problem when the graph is actually detached
- Expecting .grad to hold zeros before the first backward pass