Why does PyTorch's backward() accumulate into .grad instead of replacing it?
answer
- the chain rule sums over paths
- plus-equals, not assignment
- shared weights need contributions merged
- stale values persist silently
- set to None beats zeroing
basics
~20 sAutograd adds each backward pass's result to whatever is already in .grad. That is what lets several losses, several graph branches, or several forward passes contribute to one gradient. The cost is that stale values persist until something explicitly clears them.
solid answer
~50 sAccumulation is the correct default because the chain rule itself sums. When a tensor is used in more than one place in a graph — a shared embedding, a weight-tied decoder, a residual connection — its total derivative is the *sum* of the contributions from every consumer, so autograd's leaf nodes are additive by construction. Extending that behaviour across separate `backward()` calls makes multi-loss and multi-pass patterns work without any special API: call backward on each loss and the gradients combine correctly. The price is that `.grad` is persistent state. Nothing clears it for you, so a forgotten clear silently mixes this step's gradient with the last one's and training degrades without any error. You reset with `p.grad = None` or `p.grad.zero_()`; `None` is preferred and is what modern `zero_grad()` does by default, since it skips a kernel launch and lets the memory be released.
code
python · 13 linesimport torch
w = torch.tensor([1.0], requires_grad=True)
(w * 2).sum().backward()
print(w.grad) # tensor([2.])
(w * 3).sum().backward()
print(w.grad) # tensor([5.]) - accumulated, not replaced
w.grad = None # the cheap reset; w.grad.zero_() also works
(w * 3).sum().backward()
print(w.grad) # tensor([3.])go deeper
Remember that gradients add up across backward() calls and must be cleared explicitly, otherwise old values leak into the next update.
Explain that accumulation mirrors the chain rule summing over every path a tensor feeds, and that the same mechanism makes weight tying and multi-loss backward passes work without special handling.
Diagnose the silent version of this bug — training that merely underperforms with no error — and know why set_to_none is the default, including how optimizers treat a None gradient differently from a zero one.
Treat mutable gradient state as an API hazard worth fencing in shared training code: a reviewed step function or framework wrapper, so that no team rediscovers the missing-clear bug through a week of degraded runs.
## The rule `backward()` does not assign to `.grad`; it does `+=`. If `.grad` is `None`, autograd allocates a tensor and stores the result. If a tensor is already there, the new gradient is added element-wise on top. ``` w = torch.tensor([1.0], requires_grad=True) (w * 2).sum().backward() w.grad # tensor([2.]) (w * 3).sum().backward() w.grad # tensor([5.]) -- 2 + 3, not 3 ``` ## Why summing is the right default This is not an arbitrary API choice, it falls out of the mathematics. If a variable feeds several downstream expressions, the multivariable chain rule says its derivative is the sum over all paths. Consider a weight-tied model where the same embedding matrix is used by the input layer and the output projection: gradient arrives at that matrix along two distinct routes and the correct update uses both. Inside a single graph, autograd's accumulation node for that leaf simply adds each incoming contribution as it arrives. Making the cross-call behaviour identical means there is no seam between "one backward with two paths" and "two backwards with one path each". That uniformity buys several patterns for free: - **Multiple loss terms computed separately.** `loss_a.backward(retain_graph=True)` then `loss_b.backward()` gives the same parameter gradients as `(loss_a + loss_b).backward()`, which is useful when the two terms are produced by different parts of a system. - **Splitting a batch that will not fit in memory.** Run several smaller forward/backward passes before updating; the summed gradient corresponds to the larger batch. (Whether to scale the loss by 1/N depends on whether your loss already averages over the batch.) - **Multi-task heads** whose losses are computed and back-propagated at different points in the step. ## The cost: persistent mutable state Because `.grad` survives across iterations, it behaves like an accumulator that you own. Forget to clear it and step N's update is computed from the sum of steps 1..N. Nothing raises. The loss curve is merely worse than it should be — often just "unstable, needs a lower learning rate", which is why this bug can live in a codebase for months. Two ways to clear: - `p.grad = None` — drops the reference. The next backward allocates fresh. This is cheaper (no elementwise kernel), lets the memory be reclaimed, and means untouched parameters read as `None` rather than as stale zeros. - `p.grad.zero_()` — writes zeros in place, keeping the buffer allocated. Necessary only if something else holds a reference to that exact tensor. In current PyTorch, `zero_grad()` defaults to the `None` form (`set_to_none=True`). One consequence catches people out: any code that assumes `.grad` is a tensor must handle `None`, and optimizers with momentum treat a `None` gradient as "no update for this parameter this step" rather than as an update with a zero gradient — which for some optimizers is a meaningfully different thing, since a zero gradient still decays momentum and can still apply weight decay. ## Related autograd behaviours worth separating Accumulation is often confused with two other things: - **`retain_graph=True`** controls whether the *graph* survives a backward pass, not whether gradients accumulate. Gradients accumulate whether or not you retain the graph; you need `retain_graph` only when you will call backward through the *same* graph again. - **`create_graph=True`** makes the backward pass itself differentiable, so you can take gradients of gradients. It also means the resulting `.grad` tensors carry their own graph, which holds memory — a common leak source in meta-learning code. And if you want a gradient *without* touching `.grad` at all, `torch.autograd.grad(outputs, inputs)` returns the values directly and accumulates nothing. That is the cleaner primitive for computing a gradient penalty or an influence score inside a step that also does ordinary training, precisely because it does not contaminate the accumulator the optimizer will read.
- Why is p.grad = None preferred over p.grad.zero_()?Setting to None skips an elementwise kernel per parameter and lets the gradient buffer be freed, which lowers peak memory. It also keeps the semantics honest: a parameter that received no gradient this step reads as None rather than as a zero tensor. Modern zero_grad() uses the None form by default. The tradeoff is that any code reading .grad must tolerate None, and some optimizers treat a None gradient as skip-this-parameter rather than as a zero-gradient update.
- Does calling backward() on two separate losses give the same gradients as backward() on their sum?Yes for the parameter gradients, because accumulation is addition and differentiation is linear. The practical differences are elsewhere: two backward passes over the same graph require retain_graph=True on the first, and doing them separately costs an extra traversal. Summing first is usually cheaper; back-propagating separately is useful when the terms are produced by different components or when you want to inspect each contribution.
- What is the difference between retain_graph=True and gradient accumulation?They are unrelated switches. Accumulation is about the destination — .grad is added to rather than overwritten — and happens on every backward call. retain_graph is about the source: after a backward pass, autograd frees the buffers the graph saved, and retain_graph=True keeps them so you can traverse the same graph again. You can accumulate across many graphs without ever retaining one.
saying these in an interview costs you the question
- Saying each backward() overwrites the previous gradient
- Confusing retain_graph with gradient accumulation
- Thinking PyTorch clears .grad automatically after an optimizer step
- Believing accumulation is a workaround rather than a consequence of the chain rule
- Assuming a None gradient and a zero gradient are interchangeable for an optimizer