skip to content

PyTorch

PyTorch is the research default and increasingly the production one: tensors, autograd, nn.Module models, a training loop you write yourself, and export paths for serving. This subtree follows that order, from the tensor up to deployment.

on this pageshow

explore

questions

page 2 of 2

How do you capture a PyTorch model's intermediate activations without editing its forward()?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Register a forward hook on the submodule you want: handle = layer.register_forward_hook(fn), where fn receives (module, args, output). Detach what you store, and call handle.remove() afterwards or the hook keeps firing and holding tensors alive.

open as a page

In PyTorch, why can keeping one small slice of a big tensor retain all its memory?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Basic indexing and slicing return views that share the original tensor's storage buffer. That buffer is freed only when the last tensor referencing it is gone, so one saved row pins the entire allocation. Call .clone() to detach the slice onto its own storage.

open as a page

Where does torch.nn.utils.clip_grad_norm_ go in a PyTorch loop, and what does it clip?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Call it after loss.backward() and before optimizer.step(). It computes one global L2 norm across all the gradients you pass it and, if that norm exceeds max_norm, scales every gradient by the same factor so the combined norm equals max_norm.

open as a page

How do you choose between ONNX Runtime, Triton, and ExecuTorch for a PyTorch model?

level: principalimportance: should knowfreq 32%

basics

~20 s

Let the target decide. On-device means ExecuTorch; a heterogeneous CPU or accelerator fleet with simple pre-processing favours ONNX Runtime; a GPU service with Python-shaped pre- and post-processing favours a Triton PyTorch or Python backend. TorchServe is retired and should not anchor a new design.

open as a page

Why does renaming a submodule attribute in an nn.Module break old checkpoints?

level: principalimportance: should knowfreq 32%

basics

~20 s

state_dict keys are built from the attribute path down the module tree, so self.fc produces fc.weight. Rename the attribute and every key under it changes, making a saved checkpoint's keys unrecognisable to the new class.

open as a page

When do you write a custom torch.autograd.Function, and how?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Subclass torch.autograd.Function with static forward and backward methods, save what backward needs via ctx.save_for_backward, and invoke it with .apply(). You need one only when autograd cannot derive the gradient itself: a non-differentiable step, a custom CUDA or C++ kernel, or a hand-written backward that is faster or more numerically stable.

open as a page

Why do PyTorch DataLoader workers sometimes produce identical random augmentations?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Because a random generator created in the parent process is copied unchanged into every worker. PyTorch seeds each worker's torch, Python random and NumPy global generators from base_seed plus worker id, but an RNG object your Dataset holds, or a third-party library's state, is duplicated.

open as a page

showing 31–37 of 37