Does torch.from_numpy() copy the NumPy array or share memory with it?
answer
- from_numpy wraps, tensor() copies
- mutation crosses the boundary both ways
- numpy defaults float64, torch float32
- numpy() needs CPU and no grad
- reversed arrays have negative strides
basics
~10 storch.from_numpy() shares the array's buffer, so writing through either side is visible to the other. torch.tensor(arr) always copies. Tensor.numpy() shares as well, and only works for CPU tensors.
solid answer
~40 s`torch.from_numpy(arr)` wraps the array's existing memory in a tensor — no copy, and mutating the array changes the tensor. `torch.tensor(arr)` is documented as always copying, and `torch.as_tensor(arr)` copies only when it has to (wrong dtype or device). In the other direction `Tensor.numpy()` also shares the buffer, so the same aliasing applies. Two details bite. **dtype**: NumPy defaults to float64 while PyTorch defaults to float32, so `from_numpy` on a default array gives you a float64 tensor that will raise a dtype-mismatch error against float32 model weights — convert explicitly with `.float()` or `astype(np.float32)`. **Device**: `.numpy()` on a CUDA tensor raises and tells you to call `.cpu()` first, and on a tensor that requires grad it tells you to `detach()` first; `.numpy(force=True)` does whatever is needed and may return a copy.
code
python · 12 linesimport numpy as np
import torch
arr = np.zeros(3) # float64 by default
shared = torch.from_numpy(arr) # zero-copy, dtype float64
copied = torch.tensor(arr) # copy, still float64
arr[0] = 5.0
print(shared[0], copied[0]) # tensor(5., ...) tensor(0., ...)
back = shared.numpy() # shares again, CPU only
print(back.dtype) # float64go deeper
Recall the pairing: from_numpy shares memory, torch.tensor copies, and .numpy() shares back. Also know that .numpy() fails on a GPU tensor and that the fix is .cpu() first.
Explain the dtype seam — NumPy float64 versus torch float32 — and why detach().cpu().numpy() is the standard idiom for extracting results, plus what as_tensor adds over the other two.
Show where you place the copy in a real pipeline: which side owns the buffer, where mutation could cross the boundary into another thread or a cache, and how a shared NumPy view can pin a large tensor's storage.
Own the boundary convention for the codebase — where dtype is fixed, whether ingest is zero-copy for throughput or copied for safety, and how that choice interacts with memory-mapped data and multi-process loading.
## The three constructors and what each promises Going from NumPy into PyTorch there are three routes with different copy semantics: - `torch.from_numpy(arr)` — zero-copy. The tensor and the array reference the same memory, and the tensor's dtype is inferred from the array's. - `torch.tensor(arr)` — always copies. This is the safe default and the one to reach for when you want an independent tensor; it also accepts `dtype=` and `device=` arguments. - `torch.as_tensor(arr)` — copies only if it must. If the dtype and device already match, it behaves like `from_numpy`; otherwise it converts, which allocates. Going the other way, `Tensor.numpy()` returns an array sharing the tensor's buffer. So a round trip `torch.from_numpy(arr).numpy()` gives you back something aliasing the original. ## Why sharing is both a feature and a hazard Sharing is what makes preprocessing pipelines cheap: a decoded image, a memory-mapped array, or a large feature matrix can enter the framework without doubling memory. The hazard is that mutation crosses the boundary invisibly. Writing `arr[0] = 5` changes the tensor; an in-place tensor op like `t.add_(1)` changes the array. If one side is being read by another thread, cached as a snapshot, or handed to code that assumes immutability, this is a real bug. Copy explicitly at the boundary — `torch.tensor(arr)` inbound, `t.clone().numpy()` or `np.array(t.numpy())` outbound — whenever the lifetimes or owners differ. Sharing also means the retention rule applies across the boundary: a small NumPy array that is a view of a big tensor keeps that tensor's whole storage alive. ## The dtype seam This is the detail that actually costs people an afternoon. NumPy's default floating type is float64; PyTorch's default is float32, and `torch.get_default_dtype()` will confirm that. So `torch.from_numpy(np.zeros(10))` produces a float64 tensor. Feed it to a model whose parameters are float32 and you get a dtype-mismatch RuntimeError from the first matmul — or, worse, an upcast that silently doubles memory and halves throughput in a part of the pipeline that does tolerate it. Integer arrays behave similarly: NumPy's default integer is platform-dependent (int64 on Linux and macOS, int32 on Windows), so a tensor built from an index array can differ across machines. Be explicit: `arr.astype(np.float32)` before conversion, or `.float()` / `.to(torch.float32)` after, and set the dtype at the earliest point in the pipeline so one conversion covers everything downstream. ## Device and autograd restrictions on .numpy() `Tensor.numpy()` requires CPU memory, because NumPy cannot address GPU memory. On a CUDA tensor it raises a TypeError telling you to copy to host first with `.cpu()`. If the tensor requires grad, it raises telling you to call `.detach()` first — the point being that a NumPy array cannot participate in the autograd graph, so the framework refuses to let you silently drop out of it. The combined idiom for pulling results out is therefore `t.detach().cpu().numpy()`. Since PyTorch 1.13 there is also `t.numpy(force=True)`, which performs whatever detaching, host copy or conversion is needed and may therefore return a copy rather than a shared view — convenient in logging code, wrong when you specifically wanted the alias. ## Layout constraints `torch.from_numpy` cannot wrap an array with **negative strides**, which is exactly what `arr[::-1]` or a flipped image produces; it raises a ValueError about a negative stride and the fix is `arr.copy()` (or `np.ascontiguousarray(arr)`). It also supports only the dtypes PyTorch has — object arrays, most structured dtypes and NumPy's extended types will not convert. ## The summary you should be able to give "`from_numpy` shares, `torch.tensor` copies, `as_tensor` copies only when needed, and `.numpy()` shares back. Watch two things: NumPy is float64 by default while torch is float32, and `.numpy()` needs a CPU tensor that does not require grad, hence `detach().cpu().numpy()`."
- Why does torch.from_numpy on a default NumPy array often break a model forward pass?NumPy's default float dtype is float64 while PyTorch's is float32, and from_numpy preserves the array's dtype. The resulting float64 tensor hits float32 parameters and raises a dtype-mismatch error at the first matmul. Fix it once, early: `arr.astype(np.float32)` before conversion or `.float()` right after, rather than sprinkling casts through the model.
- What does .numpy(force=True) change compared with plain .numpy()?Plain `.numpy()` refuses anything it cannot share: a CUDA tensor, a tensor requiring grad, or a non-convertible layout. `force=True`, available since PyTorch 1.13, performs whatever detach, host copy or conversion is needed, so it always succeeds but may return a copy rather than an alias. It is convenient in logging and plotting code and wrong wherever you specifically wanted shared memory.
- Why does torch.from_numpy reject arr[::-1]?Reversing a NumPy array produces a view with a negative stride, and PyTorch tensors do not support negative strides, so from_numpy raises a ValueError naming that. Materialise a normal layout first with `arr[::-1].copy()` or `np.ascontiguousarray(...)`. The same applies to some flipped-image pipelines in OpenCV-style preprocessing.
- How do you safely hand a model's output to a plotting or logging thread?Copy at the boundary: `t.detach().cpu().numpy().copy()`, or `t.detach().cpu().clone().numpy()`. Otherwise the array aliases tensor memory that a later in-place op or a reused buffer can overwrite while the other thread reads it, and it also pins the tensor's full storage even if you sliced out one row.
saying these in an interview costs you the question
- Assumes every conversion between NumPy and torch copies
- Forgets NumPy defaults to float64 and torch to float32
- Calls .numpy() on a CUDA tensor and blames the driver
- Uses torch.tensor() on a huge array in a hot loop unnecessarily
- Thinks .numpy() detaches the tensor from autograd for you