skip to content

Why does tf.GradientTape.gradient() return None in TensorFlow, and how do you fix it?

level: middleimportance: should knowfreq 52%

answer

  1. no recorded path, not a zero
  2. tapes watch variables, not constants
  3. the with block is the boundary
  4. numpy in the middle severs the chain
  5. integers and argmax have no gradient

basics

~20 s

None means the tape found no recorded path from the loss back to that source. Usual causes: the source is not a watched variable, the computation happened outside the tape's context, or the path runs through a non-differentiable op. Fix the connection, not the symptom.

solid answer

~50 s

`tape.gradient` returns `None` for a source when there is no recorded differentiable chain from the target to it. Four causes cover almost every real case. First, the source is not watched: the tape auto-watches trainable `tf.Variable` objects touched inside the context, but a `tf.constant`, a NumPy array or an input tensor needs an explicit `tape.watch(x)`. Second, the op ran outside the `with` block, so it was never recorded — including the classic mistake of computing the loss after the block closes. Third, the path leaves TensorFlow: calling `.numpy()` on an intermediate, or using a plain Python or NumPy operation, breaks the recording. Fourth, the path crosses something non-differentiable — `tf.stop_gradient`, an `argmax`, a cast to an integer dtype, or a variable created with `trainable=False`. Debug it by shortening the chain: differentiate an intermediate tensor against the source and see where the None first appears.

code

python · 12 lines
python
import tensorflow as tf

x = tf.constant(3.0)

with tf.GradientTape() as tape:
    y = x ** 2
print(tape.gradient(y, x))          # None - x is not a Variable and was never watched

with tf.GradientTape() as tape:
    tape.watch(x)
    y = x ** 2
print(tape.gradient(y, x).numpy())  # 6.0

go deeper

for a junior

Recall that the tape only auto-watches trainable variables, so a gradient with respect to a constant or input needs tape.watch(x) inside the with block.

for a middle

Enumerate the causes — unwatched source, computation outside the context, a NumPy detour, a non-differentiable op — and explain why None is not the same as zero. Expect a code snippet where you must point at the break.

for a senior

Demonstrate the bisect routine: keep intermediates, take several gradients from a persistent tape, and check dtypes along the path. Recognize the partial case where only one submodule's variables come back None.

for a principal

Insist on a guard rather than a convention: a training step that asserts no unexpected None gradients on the first step catches a severed subgraph before a week of compute is spent training a model whose head never moved.

## What None actually means `tape.gradient(target, sources)` returns a structure matching `sources`. A `None` entry is not "the gradient is zero"; it is "I have no recorded path from the target to this source, so I cannot even say it is zero". The distinction matters, because a genuine zero gradient is a modelling problem while a `None` is almost always a wiring bug in your code. Downstream, `None` is contagious: passing a list containing `None` to `optimizer.apply_gradients` at best warns and skips that variable, at worst fails outright — either way that weight never trains, which is a silent accuracy bug if only part of the model is affected. ## Cause 1: the source is not watched The tape watches trainable `tf.Variable` objects that are accessed inside its context. It does *not* watch constants, `tf.Tensor` inputs, NumPy arrays or Python floats, because tracking every tensor would be wasteful. If you need a gradient with respect to an input — for adversarial examples, saliency maps, input optimization — you must call `tape.watch(x)` inside the block before using `x`. Two related traps: a variable created with `trainable=False` is not auto-watched, and constructing the tape with `watch_accessed_variables=False` disables auto-watching entirely, so every source must be watched by hand. That flag is useful for restricting a backward pass to a subset of the model, and confusing when someone else set it. ## Cause 2: the op ran outside the context Only ops executed while the context is open are recorded. The most frequent version of this bug is closing the block after the forward pass and computing the loss on the next line. Another is computing part of the model — an embedding lookup, a preprocessing step — before entering the block, then differentiating with respect to something on that side of the boundary. ## Cause 3: the chain leaves TensorFlow Autograd only knows about TensorFlow ops. The moment you call `.numpy()` on an intermediate tensor, or hand a tensor to a NumPy or plain-Python function that returns a new array, the graph of recorded ops is severed. Execution continues happily and the numbers may even look right on the forward pass, which is why this one is hard to spot. The fix is to stay inside TensorFlow ops for anything on the differentiable path; use `.numpy()` only for logging, and only outside the tape. ## Cause 4: a non-differentiable link Some ops legitimately have no gradient. `tf.stop_gradient` exists specifically to cut the path — sometimes deliberately, sometimes left in from an experiment. Index-producing ops such as `tf.argmax` and `tf.math.argmin` return integers and are not differentiable. Casting to `tf.int32` or `tf.int64` anywhere on the path kills it, as does any op whose output dtype is integer or boolean. Control flow that picks a branch by a hard comparison can also produce a path with no gradient with respect to some inputs. When your loss involves a discrete decision, the answer is usually a differentiable relaxation (a softmax or a soft mask) rather than trying to force a gradient through the integer op. ## A debugging routine 1. Print which entries are `None` alongside `variable.name`. A pattern — for example every variable in one submodule — usually identifies the severed layer immediately. 2. Shorten the chain: inside the tape, keep intermediate tensors, then differentiate each intermediate against the source. The first intermediate that yields `None` sits just downstream of the break. Use `persistent=True` so you can take several gradients from one forward pass. 3. Check dtypes along that stretch. An unexpected integer or boolean tensor is the signature of cause 4. 4. Confirm the block boundaries. Re-read where the `with` closes relative to where the loss is computed. ## The one legitimate None If you are training a multi-head model and deliberately compute a loss that touches only one head, the other head's variables genuinely have no path to that loss, and `None` is correct. Handle it by passing only the relevant variable subset to `tape.gradient`, or by filtering the pairs before calling `apply_gradients` — not by silently converting the `None` to zeros, which hides the case where the disconnection was accidental.

  • How is a None gradient different from a zero gradient?
    None means no recorded differentiable path exists — a structural problem in how the computation was written. A zero gradient means the path exists and the derivative genuinely evaluates to zero, which is a modelling signal: a saturated activation, a dead ReLU, or a loss term that is locally flat. Converting None to zeros hides real wiring bugs, so filter deliberately rather than by default.
  • How would you compute the gradient of a loss with respect to the input image?
    Inside the tape context call `tape.watch(image)` before the forward pass, since an input tensor is not a trainable variable and is not auto-watched. Then `tape.gradient(loss, image)` gives per-pixel sensitivity — the basis of saliency maps and FGSM-style adversarial examples. Without the explicit watch you get None.
  • What does watch_accessed_variables=False give you?
    It turns off automatic watching, so the tape records ops only for variables you pass to `tape.watch` explicitly. It is useful when you want a backward pass restricted to part of a large model — a fine-tuned head over a frozen trunk — and it saves the memory the tape would otherwise spend recording the frozen part. The cost is that forgetting a watch call yields None.

saying these in an interview costs you the question

  • Assumes the tape watches every tensor automatically
  • Reads None as a gradient of zero
  • Calls .numpy() mid-forward and still expects gradients
  • Replaces None with zeros to make the error go away
  • Puts tf.argmax on the differentiable path

context