skip to content

PyTorch

PyTorch is the research default and increasingly the production one: tensors, autograd, nn.Module models, a training loop you write yourself, and export paths for serving. This subtree follows that order, from the tensor up to deployment.

on this pageshow

explore

questions

page 1 of 2

In PyTorch, what does requires_grad=True do to a tensor?

level: juniorimportance: must knowfreq 78%

answer

  1. opt a tensor into the tape
  2. metadata flag, not a mode
  3. the flag spreads to outputs
  4. grad_fn appears on results
  5. gradient lands in .grad after backward

basics

~10 s

requires_grad=True tells PyTorch to record every operation that uses the tensor, so a backward pass can compute a derivative with respect to it. After backward(), the gradient shows up in that tensor's .grad attribute.

solid answer

~40 s

Setting `requires_grad=True` opts a tensor into autograd. From that point, every operation that consumes it is recorded as a node in a dynamic graph built during the forward pass, and the result tensor carries a `.grad_fn` pointing back at the operation that produced it. The flag is *infectious*: any output of an op with at least one `requires_grad=True` input also has `requires_grad=True`, so you only need to set it on the things you actually want derivatives for. When you call `.backward()` on a scalar output, autograd walks that graph in reverse and accumulates the derivative into `.grad` on each leaf tensor that requires grad. You can flip it after creation with `t.requires_grad_(True)`, and `nn.Parameter` sets it for you, which is why model weights are trainable without you touching the flag.

code

python · 10 lines
python
import torch

x = torch.tensor([2.0], requires_grad=True)
print(x.is_leaf, x.grad_fn, x.grad)   # True None None

y = (x ** 3).sum()
print(y.requires_grad, y.grad_fn)     # True <SumBackward0 ...>

y.backward()
print(x.grad)                          # tensor([12.])

go deeper

for a junior

Be able to say plainly that requires_grad=True makes PyTorch track operations on the tensor so backward() can fill in .grad, and that model parameters get the flag automatically.

for a middle

Explain that the flag propagates to any output of an op with a grad-requiring input, that the graph is built during the forward pass, and what grad_fn and is_leaf tell you.

for a senior

Connect the flag to memory: grad-requiring ops save activations for backward, so freezing layers with requires_grad_(False) cuts both backward compute and graph memory. Know why backward on a non-scalar raises.

for a principal

Frame it as a cost lever across a training platform: which tensors carry gradients decides activation memory, and that decides batch size, checkpointing policy and hardware choice. Push teams to audit what is actually differentiable rather than tuning batch size blindly.

## What the flag actually switches on A PyTorch tensor is a block of numbers plus some metadata. `requires_grad` is one piece of that metadata, a boolean. When it is `False` (the default for `torch.tensor`, `torch.zeros`, `torch.randn` and friends), operations on the tensor are pure numerics: the result is computed and nothing else is remembered. When it is `True`, PyTorch additionally records *what was done*, so it can later replay the chain in reverse and apply the chain rule. ``` x = torch.tensor([2.0], requires_grad=True) y = x ** 3 ``` Here `y` is not just the number 8. It also carries `y.grad_fn`, a reference to a `PowBackward0` node that knows how to turn a gradient with respect to `y` into a gradient with respect to `x`. That node holds references to whatever it needs for the derivative, which in this case includes `x` itself. ## The graph is built during the forward pass This is what "define-by-run" means. There is no separate compile step and no static graph declared ahead of time. The graph is a side effect of running your Python code, so `if`, `for`, recursion and data-dependent shapes all work naturally: whatever path the forward pass took this iteration is exactly the path the backward pass will walk. The flip side is that the graph is rebuilt from scratch every iteration, and it is freed as soon as `.backward()` has consumed it. ## The flag propagates forward You almost never set `requires_grad` on intermediates. If any input to an operation requires grad, the output does too: ``` a = torch.randn(3, requires_grad=True) b = torch.randn(3) # requires_grad False c = a + b c.requires_grad # True ``` That is why setting the flag on model weights (which `nn.Parameter` does automatically) is enough to make the entire loss differentiable with respect to them. Input data usually does *not* require grad, because you do not want a derivative with respect to the pixels — though for adversarial examples or input optimization you would set it deliberately. The converse is also useful: if *no* input to an op requires grad, the output does not, and no graph node is created at all. Freezing a backbone by setting `p.requires_grad_(False)` on its parameters therefore saves both the backward computation and the memory the graph would have held. ## Where the gradient lands `.backward()` on a scalar seeds the reverse pass with an implicit gradient of 1.0 and walks the graph backwards. At each *leaf* tensor that requires grad — a tensor you created directly, rather than one produced by an operation — autograd adds the computed derivative into `.grad`. So: ``` x = torch.tensor([2.0], requires_grad=True) (x ** 3).sum().backward() x.grad # tensor([12.]) because d(x^3)/dx = 3x^2 = 12 ``` Before any backward pass, `.grad` is `None`, not a tensor of zeros. That distinction matters when you write code that inspects gradients. ## Common gotchas - **Non-scalar backward.** `.backward()` with no argument only works on a scalar; on a non-scalar tensor it raises `RuntimeError: grad can be implicitly created only for scalar outputs`. You must pass a gradient of matching shape, which is why training code reduces to a scalar loss first. - **Integer and boolean tensors cannot require grad.** Autograd needs a floating-point (or complex) dtype; asking for `requires_grad=True` on an int tensor raises. - **You cannot set the flag on a non-leaf.** `requires_grad_()` on a tensor that already has a `grad_fn` raises, because its grad-ness is determined by its inputs. - **The flag is not "training mode".** Turning it off on parameters freezes them; it is unrelated to dropout or batch-norm behaviour, which is governed separately. ## Reading the metadata Three attributes tell you everything about a tensor's autograd status: `requires_grad` (will a derivative be tracked), `is_leaf` (is this an endpoint autograd will write `.grad` to), and `grad_fn` (which operation produced it, `None` for leaves). When something is not differentiating the way you expect, printing those three on the suspect tensor is the first diagnostic.

  • If only the weights require grad, why does the input tensor still end up inside the graph?
    Because backward nodes save whatever they need for their derivative. For a matmul, the gradient with respect to the weight is a function of the input activation, so that activation is kept alive by the graph until backward runs. That is why activation memory, not parameter memory, usually dominates a training step, and why freezing early layers frees memory only if nothing downstream still needs those saved tensors.
  • What happens if you call backward() on a tensor that isn't a scalar?
    It raises `RuntimeError: grad can be implicitly created only for scalar outputs`. Autograd needs a starting gradient to seed the reverse pass; for a scalar it can assume 1.0, but for a vector output there is no canonical choice. You either reduce to a scalar first, or pass an explicit gradient of the same shape, e.g. `y.backward(torch.ones_like(y))`, which computes a vector-Jacobian product.
  • How do you freeze part of a model so its weights get no gradients?
    Set `p.requires_grad_(False)` on the parameters you want frozen. Autograd then skips building backward nodes for those operations entirely, so you save backward compute and graph memory as well as keeping the weights fixed. Remember to pass only the still-trainable parameters to the optimizer, otherwise it holds state for tensors whose `.grad` stays `None`.

saying these in an interview costs you the question

  • Thinking requires_grad must be set on every intermediate tensor
  • Saying requires_grad switches the model into training mode
  • Believing .grad holds zeros before the first backward pass
  • Claiming the graph is built when backward() is called, not during forward
  • Assuming an integer tensor can require gradients

context

open as a page

In PyTorch, what must a custom map-style Dataset implement, and what does DataLoader add?

level: juniorimportance: must knowfreq 80%

basics

~10 s

A map-style Dataset implements len, returning the number of samples, and getitem(index), returning one sample (often a tensor/label tuple). DataLoader wraps it and produces shuffled, collated mini-batches, optionally loaded by worker processes.

open as a page

How do you load a GPU-trained PyTorch checkpoint on a CPU-only inference box?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Pass map_location to torch.load, for example torch.load(path, map_location="cpu"), so saved tensors are restored onto the CPU instead of the CUDA device they were saved from. Then load_state_dict into a model and move the model once with .to(device).

open as a page

In PyTorch, why call model(x) instead of model.forward(x)?

level: juniorimportance: must knowfreq 66%

basics

~10 s

Calling the module runs nn.Module.call, which fires registered forward pre-hooks and forward hooks around forward(). Calling model.forward(x) directly skips all of that machinery, so hook-based features silently stop working.

open as a page

Does torch.from_numpy() copy the NumPy array or share memory with it?

level: juniorimportance: must knowfreq 62%

basics

~10 s

torch.from_numpy() shares the array's buffer, so writing through either side is visible to the other. torch.tensor(arr) always copies. Tensor.numpy() shares as well, and only works for CPU tensors.

open as a page

In PyTorch, in what order do zero_grad, forward, backward and step go?

level: juniorimportance: must knowfreq 88%

basics

~10 s

Per batch: optimizer.zero_grad(), then the forward pass, then compute the loss, then loss.backward() to fill each parameter's .grad, then optimizer.step() to apply the update. Zeroing must precede backward; the update must follow it.

open as a page

In PyTorch, how do tensor.detach() and torch.no_grad() differ?

level: middleimportance: must knowfreq 80%

basics

~20 s

detach() acts on one tensor: it returns a view sharing the same data but cut off from the graph. torch.no_grad() is a context manager that switches off graph recording for every operation inside it. Use detach to stop gradients at a point; use no_grad to run a whole block untracked.

open as a page

Why is .grad None on a PyTorch tensor after calling backward()?

level: middleimportance: must knowfreq 70%

basics

~20 s

Autograd only fills .grad on leaf tensors that require grad. A tensor produced by an operation is a non-leaf, so its gradient is used during the backward pass and discarded — call retain_grad() on it to keep a copy. .grad is also None if the tensor never required grad or was cut off from the graph.

open as a page

What does DataLoader's collate_fn do in PyTorch, and when must you write your own?

level: middleimportance: must knowfreq 60%

basics

~20 s

collate_fn turns a list of individual samples into one batch. The default stacks equally-shaped tensors along a new dimension 0 and recurses through tuples and dicts. Write your own for variable-length data, custom padding, or ragged structures.

open as a page

What does setting DataLoader's num_workers above 0 actually change in PyTorch?

level: middleimportance: must knowfreq 70%

basics

~20 s

It moves sample fetching and collation into that many separate processes, each prefetching batches ahead of the training loop. The costs are process startup per epoch, duplicated memory, a picklable dataset requirement, and worse tracebacks when something fails.

open as a page

How does torch.compile differ from torch.jit.script for shipping a PyTorch model?

level: middleimportance: must knowfreq 60%

basics

~20 s

torch.compile is a just-in-time accelerator that stays inside Python: TorchDynamo captures graphs from bytecode, a backend generates kernels, and unsupported code simply breaks the graph and runs eagerly. It emits no portable artifact, so it speeds a process up rather than getting the model out of Python.

open as a page

In PyTorch, when does torch.jit.trace silently produce a wrong TorchScript model?

level: middleimportance: must knowfreq 70%

basics

~20 s

torch.jit.trace records only the operations one example input actually executed, so a branch or loop that depends on tensor values is frozen as whatever ran that day. torch.jit.script compiles the Python source instead and keeps the control flow.

open as a page

In PyTorch, what breaks if you store layers in a plain Python list inside nn.Module?

level: middleimportance: must knowfreq 72%

basics

~10 s

Submodules in a plain list are never registered, so model.parameters() misses them, the optimizer never updates them, and .to(device) leaves them on the CPU. Use nn.ModuleList or nn.ModuleDict instead.

open as a page

In PyTorch, what is the difference between an nn.Parameter and a registered buffer?

level: middleimportance: must knowfreq 78%

basics

~10 s

An nn.Parameter is trainable state: it shows up in model.parameters(), so the optimizer updates it. A buffer registered with register_buffer is untrained state that still travels with .to(device) and is saved in state_dict().

open as a page

In PyTorch, how does broadcasting decide the output shape of a + b?

level: middleimportance: must knowfreq 72%

basics

~20 s

PyTorch aligns the two shapes from the trailing dimension backwards. Each pair must be equal, or one of them must be 1, or missing entirely; size-1 and missing dimensions are stretched to the other operand's size. Anything else raises.

open as a page

In PyTorch, when does .view() fail on a tensor where .reshape() works?

level: middleimportance: must knowfreq 66%

basics

~20 s

view() only succeeds when the requested shape can be expressed with the tensor's existing strides, so it fails on non-contiguous tensors such as the output of transpose() or permute(). reshape() returns a view when it can and silently copies when it cannot.

open as a page

Why does PyTorch's nn.CrossEntropyLoss expect raw logits, not softmax output?

level: middleimportance: must knowfreq 62%

basics

~10 s

nn.CrossEntropyLoss applies log_softmax internally, so it must receive unnormalized scores. Feeding it softmax probabilities applies softmax twice, which flattens the distribution, shrinks gradients and slows or stalls training. No error is raised.

open as a page

What do model.train() and model.eval() change in a PyTorch model?

level: middleimportance: must knowfreq 76%

basics

~20 s

They flip a boolean flag that certain layers read. Dropout zeroes activations only in train mode; BatchNorm normalizes with batch statistics and updates its running averages in train, and uses the stored averages in eval. Neither call affects gradient tracking.

open as a page

Your PyTorch job resumed from a checkpoint and the loss jumped — what did it omit?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Almost always the optimizer state. Saving only model.state_dict() discards Adam's moment estimates and momentum buffers, so the first steps after resume behave like a freshly initialized optimizer. A resumable checkpoint also needs the scheduler, the AMP scaler and the step counter.

open as a page

Why does PyTorch's backward() accumulate into .grad instead of replacing it?

level: middleimportance: should knowfreq 68%

basics

~20 s

Autograd adds each backward pass's result to whatever is already in .grad. That is what lets several losses, several graph branches, or several forward passes contribute to one gradient. The cost is that stale values persist until something explicitly clears them.

open as a page

In PyTorch, when is IterableDataset right, and how do you shard it across workers?

level: middleimportance: should knowfreq 45%

basics

~20 s

Use IterableDataset when data arrives as a stream with no cheap random access — a remote log, a compressed shard, a queue. Each worker runs the whole iter, so without sharding by get_worker_info().id every sample is emitted once per worker.

open as a page

What does torch.export.export() return, and how is it stricter than tracing?

level: middleimportance: should knowfreq 50%

basics

~20 s

torch.export.export() returns an ExportedProgram: one whole-graph ATen representation plus a signature describing which inputs are parameters, buffers and user arguments, and symbolic shape constraints. Unlike tracing, it refuses to silently specialize — unsupported dynamism raises instead.

open as a page

How does PyTorch initialize nn.Linear weights by default, and when do you override it?

level: middleimportance: should knowfreq 50%

basics

~10 s

nn.Linear initializes itself in reset_parameters: the weight with nn.init.kaiming_uniform_ (a=sqrt(5)) and the bias uniformly within ±1/sqrt(fan_in). Override with torch.nn.init functions applied through model.apply when the default's implied nonlinearity does not match yours.

open as a page

In PyTorch, how does type promotion pick the result dtype of a mixed-dtype op?

level: middleimportance: should knowfreq 50%

basics

~20 s

PyTorch picks the dtype that can represent both operands: the higher category first (bool below integer below float below complex), then the wider type within that category. Python scalars and 0-dim tensors count only weakly and never widen a dimensioned operand.

open as a page

In PyTorch, what does an in-place op like add_() risk that add() does not?

level: middleimportance: should knowfreq 55%

basics

~20 s

A trailing underscore means the op overwrites the tensor's own storage instead of allocating a result. That saves an allocation but mutates every view and alias of that tensor, and it destroys values that a later backward pass may still need.

open as a page

How do torch.amp.autocast and GradScaler fit into a PyTorch training loop?

level: middleimportance: should knowfreq 54%

basics

~20 s

Run only the forward pass and the loss under torch.amp.autocast, which executes eligible operations in a 16-bit type while the weights stay 32-bit. With float16, route the backward pass through torch.amp.GradScaler so small gradients do not underflow to zero.

open as a page

When do you call a PyTorch lr_scheduler's step() — every batch or every epoch?

level: middleimportance: should knowfreq 48%

basics

~20 s

It depends on the schedule. Epoch-granularity schedules such as StepLR are stepped once per epoch; warmup and OneCycleLR are stepped once per optimizer step. Either way the call goes after optimizer.step(), and ReduceLROnPlateau also needs a metric argument.

open as a page

In PyTorch, why does keeping the loss tensor grow memory each iteration?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A loss tensor holds a reference to the whole computation graph that produced it, including every saved activation. Storing it in a list or adding it to a running total keeps that graph alive, so each iteration's activations pile up. Use loss.item() or loss.detach() instead.

open as a page

Your GPU idles between batches in PyTorch training — how do you find and fix the input bottleneck?

level: seniorimportance: should knowfreq 55%

basics

~20 s

First prove it: time a pass that iterates the DataLoader with no model. If that alone matches epoch time, the pipeline is the limit. Then raise num_workers, add persistent_workers, pin_memory with non_blocking copies, and cut per-sample decode cost.

open as a page

Your ONNX export of a PyTorch model returns different numbers — how do you debug it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Separate a real bug from floating-point noise first by comparing with a tolerance, not equality. Then check the usual causes in order: the module was not in evaluation mode at export, shapes were baked because no dynamic axes were declared, and unsupported operators changed semantics. Bisect by exporting submodules.

open as a page

showing 1–30 of 37