skip to content

In PyTorch, what is the difference between an nn.Parameter and a registered buffer?

level: middleimportance: must knowfreq 78%

answer

  1. one is optimized, one is not
  2. both still move with .to(device)
  3. BatchNorm: weight versus running_mean
  4. register_buffer(..., persistent=False)
  5. bare attribute is in neither dict

basics

~10 s

An nn.Parameter is trainable state: it shows up in model.parameters(), so the optimizer updates it. A buffer registered with register_buffer is untrained state that still travels with .to(device) and is saved in state_dict().

solid answer

~40 s

Both are tensors an `nn.Module` owns, and both move when you call `.to(device)` and both land in `state_dict()` — the difference is who writes to them. An `nn.Parameter` is created with `requires_grad=True` by default and is yielded by `model.parameters()`, which is exactly the iterable you hand to an optimizer, so gradient descent updates it. A buffer is registered with `self.register_buffer("name", tensor)`, is yielded by `model.buffers()` instead, and is never seen by the optimizer; you update it yourself inside `forward()` or it stays constant. The canonical example is `nn.BatchNorm2d`, whose `weight`/`bias` are parameters while `running_mean`, `running_var` and `num_batches_tracked` are buffers. Use `register_buffer(..., persistent=False)` when the tensor is derived state (a cached causal mask, precomputed RoPE tables) that you would rather not carry inside every checkpoint.

code

python · 17 lines
python
import torch
import torch.nn as nn

class Scaler(nn.Module):
    def __init__(self, dim):
        super().__init__()
        self.scale = nn.Parameter(torch.ones(dim))
        self.register_buffer("running_mean", torch.zeros(dim))
        self.register_buffer("mask", torch.ones(dim), persistent=False)

    def forward(self, x):
        return (x - self.running_mean) * self.scale * self.mask

m = Scaler(4)
print([n for n, _ in m.named_parameters()])  # ['scale']
print([n for n, _ in m.named_buffers()])     # ['running_mean', 'mask']
print(list(m.state_dict().keys()))           # ['scale', 'running_mean']

go deeper

for a junior

Know that nn.Parameter is the trainable kind and register_buffer is the non-trainable kind, and that model.parameters() is what you hand to the optimizer.

for a middle

Explain the mechanics: assignment registers a parameter, register_buffer registers a buffer, and both are walked by .to() and saved in state_dict. Name BatchNorm's running_mean as the classic buffer.

for a senior

Show the debugging instinct: a bare tensor attribute silently skips device movement and optimization. Be able to diagnose a device-mismatch error back to an unregistered mask in three lines.

for a principal

Own the checkpoint contract. Decide which derived state is persistent and which is persistent=False, so that changing sequence length or model config later does not invalidate every saved checkpoint in the team's registry.

## The two kinds of state a module owns An `nn.Module` in PyTorch is a container for three things: submodules, parameters, and buffers. Submodules are other `nn.Module` instances. Parameters and buffers are both tensors, and the distinction between them is one of the most reliably asked PyTorch questions, because getting it wrong produces bugs that never raise an exception. A **parameter** is a tensor wrapped in `nn.Parameter`. Wrapping does two things: it sets `requires_grad=True` by default, and it makes `nn.Module.__setattr__` register the tensor in the module's internal `_parameters` dict when you assign it to an attribute. Registered parameters are what `model.parameters()` and `model.named_parameters()` yield — and that iterable is precisely what you pass to an optimizer, e.g. `torch.optim.Adam(model.parameters(), lr=1e-3)`. So "is it a parameter?" is operationally the same question as "will the optimizer update it?" A **buffer** is a plain tensor you register explicitly with `self.register_buffer("running_mean", torch.zeros(dim))`. It goes into the module's `_buffers` dict. It is yielded by `model.buffers()` / `model.named_buffers()`, never by `model.parameters()`, and therefore no optimizer ever touches it. If it changes, it changes because your own code assigned to it. ## What the two have in common This is the half candidates usually miss. Both parameters and buffers are: - **moved by `.to()`, `.cuda()`, `.float()`, `.half()`** — `nn.Module.to()` walks `_parameters` and `_buffers` recursively through all submodules and swaps the tensors in place. A tensor stored as a bare attribute is in neither dict, so it is silently left behind on the CPU. - **included in `state_dict()`** — checkpoints contain both, which is why loading a BatchNorm checkpoint restores its running statistics and not just its affine weights. - **visible to `model.apply(fn)` traversal indirectly**, in the sense that the module tree they hang off is walked. ## The failure mode ``` self.mask = torch.tril(torch.ones(n, n)) # bare attribute ``` This looks harmless. Then `model.cuda()` moves every parameter and buffer to the GPU but not `self.mask`, and the first forward pass raises `RuntimeError: Expected all tensors to be on the same device`. Worse, if you had assigned a bare tensor that you *intended* to train, there is no error at all — it simply never appears in `model.parameters()`, the optimizer never updates it, and the model trains around a frozen tensor. That silent version is the one interviewers like, because it costs a real training run to discover. The symmetric mistake is wrapping something in `nn.Parameter` that should not be learned — a positional-encoding table or a normalisation constant. Now the optimizer drifts it, and results change subtly between runs. ## persistent=False `register_buffer(name, tensor, persistent=False)` registers the buffer for device movement but excludes it from `state_dict()`. Use it for anything that is a pure function of the config and can be rebuilt at construction time: causal attention masks, rotary embedding tables, precomputed windows. It keeps checkpoints smaller and, more importantly, keeps them loadable when you later change the sequence length the mask was sized for. ## Freezing versus buffering A related question: if you want a tensor that is trainable-shaped but currently frozen, do you make it a buffer? No. Keep it an `nn.Parameter` and set `param.requires_grad = False`, or simply exclude it from the parameter group you pass to the optimizer. It stays a parameter — same `state_dict` key, same semantics — and you can unfreeze it later. Turning it into a buffer changes the checkpoint key and the meaning of the attribute. ## What to check when debugging Three one-liners settle almost every question in this area: - `[n for n, _ in model.named_parameters()]` — is the tensor going to be optimized? - `[n for n, _ in model.named_buffers()]` — is it going to move with the module? - `list(model.state_dict().keys())` — is it going to survive a save/load round trip? If a tensor appears in none of the three, it is a bare attribute, and both device movement and checkpointing will skip it.

  • What does the persistent=False argument to register_buffer change?
    The buffer is still registered for device and dtype movement with `.to()`, but it is excluded from `state_dict()`, so it is neither saved nor expected on load. It is the right choice for state that is fully derived from the config — causal masks, rotary tables, precomputed windows — because it keeps checkpoints smaller and avoids shape mismatches when you later change sequence length.
  • If you want a weight frozen for a few epochs, should it become a buffer?
    No. Keep it an `nn.Parameter` and set `requires_grad = False`, or leave it out of the optimizer's parameter group. It keeps the same `state_dict` key and the same meaning, and you can unfreeze it later by flipping the flag. Converting to a buffer renames the checkpoint entry and permanently changes what the attribute is.
  • How do you tell, at runtime, whether a tensor you assigned was actually registered?
    Print `[n for n, _ in model.named_parameters()]` and `[n for n, _ in model.named_buffers()]`, or just `list(model.state_dict().keys())`. If the name appears in none of them, it is a bare Python attribute: the optimizer will not update it and `.to(device)` will not move it.

saying these in an interview costs you the question

  • Says buffers are just parameters with requires_grad=False
  • Thinks buffers are excluded from state_dict by default
  • Assumes any tensor attribute moves with model.cuda()
  • Wraps constants in nn.Parameter so the optimizer drifts them
  • Cannot name a real buffer such as BatchNorm running_mean

context