skip to content

In PyTorch, when does .view() fail on a tensor where .reshape() works?

level: middleimportance: must knowfreq 66%

answer

  1. metadata change versus possible copy
  2. strides must permit the new walk
  3. transpose leaves the tensor non-contiguous
  4. reshape falls back to contiguous().view()
  5. aliasing guarantee is the real difference

basics

~20 s

view() only succeeds when the requested shape can be expressed with the tensor's existing strides, so it fails on non-contiguous tensors such as the output of transpose() or permute(). reshape() returns a view when it can and silently copies when it cannot.

solid answer

~50 s

A tensor is a shape plus strides over a flat storage buffer. `view()` never touches data — it only rewrites shape and strides — so it requires the new shape to be compatible with the current stride layout. After `transpose()`, `permute()` or a strided slice the tensor is no longer contiguous, and `view()` raises `RuntimeError: view size is not compatible with input tensor's size and stride ... use .reshape(...) instead`. `reshape()` tries `view()` first and falls back to `contiguous().view()`, i.e. a real copy, when the strides do not allow it. That convenience costs you a guarantee: the result of `reshape()` may or may not alias the input, so writing into it may or may not be visible in the original. Use `view()` when you *want* an alias and want an error otherwise; use `contiguous().view()` or `reshape().clone()` when you want the copy to be explicit; check with `is_contiguous()`.

code

python · 12 lines
python
import torch

x = torch.arange(6).reshape(2, 3)
xt = x.t()
print(xt.is_contiguous())        # False
print(xt.stride())               # (1, 3)
try:
    xt.view(6)
except RuntimeError as e:
    print("view failed:", e)
print(xt.reshape(6))             # works - copies
print(xt.contiguous().view(6))   # works - copy made explicit

go deeper

for a junior

Recall that view() can raise a RuntimeError telling you to use reshape() instead, and that transposing is what usually triggers it. Knowing the error message and the one-line fix is enough at this level.

for a middle

Explain the mechanism: a tensor is shape plus strides over one buffer, view is a pure metadata change so it needs compatible strides, and reshape falls back to a copy. Name contiguous() and is_contiguous().

for a senior

Show the judgment call — pick view when aliasing is load-bearing so a bad layout errors loudly, and treat a contiguous() call in a hot loop as a copy you should be able to justify or design away.

for a principal

Frame it as an API-contract question for shared model code: decide whether internal helpers guarantee views or permit copies, document that, and keep layout decisions (channels-last, transposed attention buffers) explicit rather than emergent.

## Shape, strides and storage A PyTorch tensor is a thin descriptor over a one-dimensional storage buffer. The descriptor carries a shape, a stride per dimension (how many elements to skip to advance one step along that axis) and a storage offset. Reshaping is therefore not inherently a data operation: if the elements you want, in the order you want them, can be walked with some set of strides over the existing buffer, the reshape is free. A tensor is **contiguous** when its strides match the row-major layout implied by its shape — the last dimension has stride 1, and each earlier stride is the product of the sizes to its right. `torch.arange(6).reshape(2, 3)` has strides `(3, 1)` and is contiguous. Transposing it swaps both the sizes and the strides, giving shape `(3, 2)` with strides `(1, 3)`: the same memory, walked in a different order, and no longer contiguous. `Tensor.is_contiguous()` reports this, and `Tensor.stride()` shows why. ## What view() promises `view()` promises to return a tensor **sharing storage** with the input. That is a strong guarantee, and it is why the method is restricted. It can only merge or split dimensions whose elements are already adjacent in memory in the required order. Merging the two dimensions of a transposed tensor into one flat run is impossible with any stride, because in memory the elements alternate — the error message says exactly that: at least one dimension spans across two contiguous subspaces. Because of the guarantee, `view()` is also the right tool when aliasing is the point: reshaping a parameter buffer you intend to write through, or flattening activations you will later fill in place. If the layout does not permit it, an exception is far better than a silent copy that quietly breaks the aliasing your code depends on. ## What reshape() does instead `reshape()` is defined as "give me this shape, I do not care how". Internally it attempts the view path and, when the strides do not allow it, materialises a contiguous copy and views that. The result is that `reshape()` never raises for a valid element count, but its aliasing behaviour is **data-dependent**: the same line of code may return a view on Monday and a copy on Tuesday, depending on whether an upstream op happened to transpose. Any code that writes into the reshaped tensor and expects the original to change is a latent bug. `Tensor.flatten()` and `Tensor.ravel()` behave like `reshape` in this respect. `Tensor.contiguous()` is the explicit form: it returns `self` when the tensor is already contiguous and a fresh copy otherwise, so `x.contiguous().view(...)` documents that a copy may occur while still ending with a hard aliasing guarantee. ## Where this bites in real code The most common instance is attention-style reshaping: `x.transpose(1, 2).view(batch, seq, -1)` raises, and the fix is either `.contiguous().view(...)` or `.reshape(...)`. Another is flattening a channels-last image tensor produced by `permute`. A third is a strided slice like `x[:, ::2]`, which is non-contiguous even though nothing was transposed. There is also a performance dimension. `contiguous()` on a large activation tensor is a full copy — extra memory traffic and extra peak memory. In an inner loop it can be worth restructuring the computation (for example, choosing a layout that avoids the transpose, or using `einsum` or `torch.permute` combined with ops that accept arbitrary strides) rather than paying for a copy every step. Conversely, some kernels are much faster on contiguous input, so a deliberate `contiguous()` can be a win. Neither is a rule; measure. ## Related shape ops and how they classify Always views, never copies: `view`, `transpose`, `permute`, `squeeze`, `unsqueeze`, `expand`, basic slicing, `narrow`, `select`, `Tensor.t()`, `detach`. Maybe a view: `reshape`, `flatten`, `contiguous`, `Tensor.to(dtype)` when the dtype is unchanged. Always copies: `clone`, `repeat`, advanced (fancy) indexing with an index tensor or boolean mask, `torch.cat`, `torch.stack`. ## The interview-ready summary Say that `view` is a pure metadata change with an aliasing guarantee and therefore a stride precondition; that `reshape` relaxes the precondition by silently copying when needed and therefore gives up the guarantee; that `transpose`/`permute`/strided slices are the usual reason the precondition fails; and that you pick between them by asking whether you need the alias, not by which one avoids the error message.

  • Why is losing the aliasing guarantee from reshape() actually dangerous, rather than just untidy?
    Because whether you get a view depends on the upstream layout, which can change when someone adds a transpose or a strided slice earlier in the pipeline. Code that writes into the reshaped tensor and relies on the original updating will work in tests and fail in production, with no exception. If you need the alias, use `view` so an incompatible layout is an error.
  • What does calling contiguous() cost on a large activation tensor in a training loop?
    A full copy: one allocation the size of the tensor plus a read and a write of every element, added to peak memory at exactly the point where activations are already largest. In an inner loop that can be a measurable fraction of step time, so it is worth checking whether a different layout, einsum, or an op that accepts strided input avoids the transpose that forced it.
  • How do you tell at runtime whether a given tensor can be viewed to a new shape?
    Check `Tensor.is_contiguous()` for the common case, and `Tensor.stride()` when you want to reason about the layout precisely. A contiguous tensor can always be viewed to any shape with the same element count; a non-contiguous one can sometimes be viewed and sometimes not, so the practical test is to call `view` and let it raise, or just call `contiguous()` first when you do not mind the copy.

saying these in an interview costs you the question

  • Says view and reshape are interchangeable aliases
  • Thinks reshape always returns a view
  • Believes transpose copies the underlying data
  • Calls contiguous() everywhere without noticing the copy cost
  • Confuses element count compatibility with stride compatibility

context