skip to content

In PyTorch, what does requires_grad=True do to a tensor?

level: juniorimportance: must knowfreq 78%

answer

  1. opt a tensor into the tape
  2. metadata flag, not a mode
  3. the flag spreads to outputs
  4. grad_fn appears on results
  5. gradient lands in .grad after backward

basics

~10 s

requires_grad=True tells PyTorch to record every operation that uses the tensor, so a backward pass can compute a derivative with respect to it. After backward(), the gradient shows up in that tensor's .grad attribute.

solid answer

~40 s

Setting `requires_grad=True` opts a tensor into autograd. From that point, every operation that consumes it is recorded as a node in a dynamic graph built during the forward pass, and the result tensor carries a `.grad_fn` pointing back at the operation that produced it. The flag is *infectious*: any output of an op with at least one `requires_grad=True` input also has `requires_grad=True`, so you only need to set it on the things you actually want derivatives for. When you call `.backward()` on a scalar output, autograd walks that graph in reverse and accumulates the derivative into `.grad` on each leaf tensor that requires grad. You can flip it after creation with `t.requires_grad_(True)`, and `nn.Parameter` sets it for you, which is why model weights are trainable without you touching the flag.

code

python · 10 lines
python
import torch

x = torch.tensor([2.0], requires_grad=True)
print(x.is_leaf, x.grad_fn, x.grad)   # True None None

y = (x ** 3).sum()
print(y.requires_grad, y.grad_fn)     # True <SumBackward0 ...>

y.backward()
print(x.grad)                          # tensor([12.])

go deeper

for a junior

Be able to say plainly that requires_grad=True makes PyTorch track operations on the tensor so backward() can fill in .grad, and that model parameters get the flag automatically.

for a middle

Explain that the flag propagates to any output of an op with a grad-requiring input, that the graph is built during the forward pass, and what grad_fn and is_leaf tell you.

for a senior

Connect the flag to memory: grad-requiring ops save activations for backward, so freezing layers with requires_grad_(False) cuts both backward compute and graph memory. Know why backward on a non-scalar raises.

for a principal

Frame it as a cost lever across a training platform: which tensors carry gradients decides activation memory, and that decides batch size, checkpointing policy and hardware choice. Push teams to audit what is actually differentiable rather than tuning batch size blindly.

## What the flag actually switches on A PyTorch tensor is a block of numbers plus some metadata. `requires_grad` is one piece of that metadata, a boolean. When it is `False` (the default for `torch.tensor`, `torch.zeros`, `torch.randn` and friends), operations on the tensor are pure numerics: the result is computed and nothing else is remembered. When it is `True`, PyTorch additionally records *what was done*, so it can later replay the chain in reverse and apply the chain rule. ``` x = torch.tensor([2.0], requires_grad=True) y = x ** 3 ``` Here `y` is not just the number 8. It also carries `y.grad_fn`, a reference to a `PowBackward0` node that knows how to turn a gradient with respect to `y` into a gradient with respect to `x`. That node holds references to whatever it needs for the derivative, which in this case includes `x` itself. ## The graph is built during the forward pass This is what "define-by-run" means. There is no separate compile step and no static graph declared ahead of time. The graph is a side effect of running your Python code, so `if`, `for`, recursion and data-dependent shapes all work naturally: whatever path the forward pass took this iteration is exactly the path the backward pass will walk. The flip side is that the graph is rebuilt from scratch every iteration, and it is freed as soon as `.backward()` has consumed it. ## The flag propagates forward You almost never set `requires_grad` on intermediates. If any input to an operation requires grad, the output does too: ``` a = torch.randn(3, requires_grad=True) b = torch.randn(3) # requires_grad False c = a + b c.requires_grad # True ``` That is why setting the flag on model weights (which `nn.Parameter` does automatically) is enough to make the entire loss differentiable with respect to them. Input data usually does *not* require grad, because you do not want a derivative with respect to the pixels — though for adversarial examples or input optimization you would set it deliberately. The converse is also useful: if *no* input to an op requires grad, the output does not, and no graph node is created at all. Freezing a backbone by setting `p.requires_grad_(False)` on its parameters therefore saves both the backward computation and the memory the graph would have held. ## Where the gradient lands `.backward()` on a scalar seeds the reverse pass with an implicit gradient of 1.0 and walks the graph backwards. At each *leaf* tensor that requires grad — a tensor you created directly, rather than one produced by an operation — autograd adds the computed derivative into `.grad`. So: ``` x = torch.tensor([2.0], requires_grad=True) (x ** 3).sum().backward() x.grad # tensor([12.]) because d(x^3)/dx = 3x^2 = 12 ``` Before any backward pass, `.grad` is `None`, not a tensor of zeros. That distinction matters when you write code that inspects gradients. ## Common gotchas - **Non-scalar backward.** `.backward()` with no argument only works on a scalar; on a non-scalar tensor it raises `RuntimeError: grad can be implicitly created only for scalar outputs`. You must pass a gradient of matching shape, which is why training code reduces to a scalar loss first. - **Integer and boolean tensors cannot require grad.** Autograd needs a floating-point (or complex) dtype; asking for `requires_grad=True` on an int tensor raises. - **You cannot set the flag on a non-leaf.** `requires_grad_()` on a tensor that already has a `grad_fn` raises, because its grad-ness is determined by its inputs. - **The flag is not "training mode".** Turning it off on parameters freezes them; it is unrelated to dropout or batch-norm behaviour, which is governed separately. ## Reading the metadata Three attributes tell you everything about a tensor's autograd status: `requires_grad` (will a derivative be tracked), `is_leaf` (is this an endpoint autograd will write `.grad` to), and `grad_fn` (which operation produced it, `None` for leaves). When something is not differentiating the way you expect, printing those three on the suspect tensor is the first diagnostic.

  • If only the weights require grad, why does the input tensor still end up inside the graph?
    Because backward nodes save whatever they need for their derivative. For a matmul, the gradient with respect to the weight is a function of the input activation, so that activation is kept alive by the graph until backward runs. That is why activation memory, not parameter memory, usually dominates a training step, and why freezing early layers frees memory only if nothing downstream still needs those saved tensors.
  • What happens if you call backward() on a tensor that isn't a scalar?
    It raises `RuntimeError: grad can be implicitly created only for scalar outputs`. Autograd needs a starting gradient to seed the reverse pass; for a scalar it can assume 1.0, but for a vector output there is no canonical choice. You either reduce to a scalar first, or pass an explicit gradient of the same shape, e.g. `y.backward(torch.ones_like(y))`, which computes a vector-Jacobian product.
  • How do you freeze part of a model so its weights get no gradients?
    Set `p.requires_grad_(False)` on the parameters you want frozen. Autograd then skips building backward nodes for those operations entirely, so you save backward compute and graph memory as well as keeping the weights fixed. Remember to pass only the still-trainable parameters to the optimizer, otherwise it holds state for tensors whose `.grad` stays `None`.

saying these in an interview costs you the question

  • Thinking requires_grad must be set on every intermediate tensor
  • Saying requires_grad switches the model into training mode
  • Believing .grad holds zeros before the first backward pass
  • Claiming the graph is built when backward() is called, not during forward
  • Assuming an integer tensor can require gradients

context