skip to content

How does a tf.GradientTape training step compute and apply gradients in TensorFlow?

level: middleimportance: must knowfreq 76%

answer

  1. a recorder, not a graph you build
  2. the with block is the scope
  3. three calls: record, gradient, apply
  4. no gradient-zeroing call needed
  5. one backward pass per tape

basics

~10 s

tf.GradientTape records the forward pass that runs inside its with block. tape.gradient(loss, model.trainable_variables) replays that recording backwards to produce gradients, and optimizer.apply_gradients(zip(grads, variables)) writes the update into the variables.

solid answer

~50 s

A TensorFlow training step has three moving parts. Inside `with tf.GradientTape() as tape:` you run the forward pass and compute the loss — the tape records every op on watched tensors, and anything computed outside that block is invisible to it. After the block you call `grads = tape.gradient(loss, model.trainable_variables)`, which walks the recorded tape backwards and returns one gradient per variable, in the same order. Then `optimizer.apply_gradients(zip(grads, model.trainable_variables))` applies the optimizer's update rule and mutates the variables in place. There is no separate gradient-zeroing call: each step builds a fresh tape, so nothing accumulates between steps unless you accumulate it yourself. A non-persistent tape releases its recording after the first `gradient()` call, so you get one backward pass per tape. Wrapping the whole step in `tf.function` is the normal way to make it fast.

code

python · 18 lines
python
import tensorflow as tf

model = tf.keras.Sequential([tf.keras.layers.Dense(10)])
loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)
optimizer = tf.keras.optimizers.Adam(learning_rate=1e-3)

@tf.function
def train_step(x, y):
    with tf.GradientTape() as tape:
        logits = model(x, training=True)
        loss = loss_fn(y, logits)
    grads = tape.gradient(loss, model.trainable_variables)
    optimizer.apply_gradients(zip(grads, model.trainable_variables))
    return loss

x = tf.random.normal((32, 4))
y = tf.random.uniform((32,), maxval=10, dtype=tf.int32)
print(train_step(x, y).numpy())

go deeper

for a junior

Be able to name the three calls in order — record inside the with block, tape.gradient, optimizer.apply_gradients — and say that the loss must be computed inside the block.

for a middle

Explain what the tape actually records, why trainable_variables is the right source list, and why no gradient-zeroing step exists in TensorFlow. Expect to be asked what a non-persistent tape does after its first gradient call.

for a senior

Show you can extend the pattern: gradient accumulation for oversized batches, persistent tapes for multi-optimizer steps, and wrapping the step in tf.function so the loop is not Python-bound. Be ready to spot a broken loop in review.

for a principal

Own the convention. Decide whether the team writes bespoke steps at all, and if so, whether one shared, tested step utility is enforced so that training=True, correct variable lists and clipping do not have to be rediscovered in every project.

## What a tape is `tf.GradientTape` is a recorder. While its context is open, every TensorFlow op executed on a *watched* tensor is appended to a tape, along with enough information to compute that op's derivative. When you later ask for a gradient, TensorFlow walks that recording in reverse and applies the chain rule. The tape auto-watches trainable `tf.Variable` objects that are touched inside the context — which is exactly the set of things a model's weights are — so you rarely have to watch anything explicitly during ordinary training. This is the mechanism behind eager-mode differentiation: there is no static graph you compile in advance and no global backward call. The tape exists for exactly the region of code you wrap. ## The three lines ``` with tf.GradientTape() as tape: logits = model(x, training=True) loss = loss_fn(y, logits) grads = tape.gradient(loss, model.trainable_variables) optimizer.apply_gradients(zip(grads, model.trainable_variables)) ``` **The `with` block** must contain the forward pass *and* the loss. This is the single most common beginner error: computing the loss after the block closes leaves the loss disconnected from the recording, and you get gradients of `None`. **`tape.gradient(target, sources)`** returns a structure matching `sources`. Pass `model.trainable_variables` — not `model.variables`, which also includes non-trainable state such as batch-normalization moving statistics that no optimizer should update. The returned list is positionally aligned with the variable list, which is why `zip` is the idiomatic pairing. **`optimizer.apply_gradients`** takes an iterable of `(gradient, variable)` pairs, applies the optimizer's rule (momentum, Adam moments, weight decay, any gradient clipping configured on the optimizer), increments the optimizer's internal step counter, and assigns the new values. It returns nothing useful; the effect is the mutation of the variables. ## No zero_grad Frameworks that accumulate gradients into the parameter objects require you to clear them before each backward pass. TensorFlow does not: `tape.gradient` returns a fresh list of tensors, and `apply_gradients` consumes it immediately. A candidate who insists TensorFlow needs a gradient-zeroing call is describing a different framework. If you *want* accumulation — for example to simulate a larger batch than fits in memory — you build it yourself: keep a list of `tf.Variable` accumulators, add each micro-batch's gradients into them, apply once every N micro-batches, and reset the accumulators. ## One tape, one backward pass A default tape is non-persistent: it frees its recording as soon as `gradient()` returns, and a second call raises a `RuntimeError`. If you genuinely need several independent gradients from the same forward pass — two losses against different variable subsets, a GAN-style step, or a second-order derivative — construct it as `tf.GradientTape(persistent=True)` and delete the tape when finished so the buffers are released. Nesting two tapes gives higher-order derivatives: the outer tape records the inner tape's gradient computation. ## Speed and correctness details Wrapping the step function in `tf.function` moves it from op-by-op eager execution into a compiled graph, which typically buys a large speedup for small models where Python overhead dominates. Call the model with `training=True` inside the tape so training-time layer behaviour is used. Keep everything in TensorFlow ops: converting an intermediate tensor to NumPy inside the block silently breaks the recorded chain, because NumPy operations are not recorded. Finally, remember the ordering contract. `tape.gradient` does not update anything and `apply_gradients` does not compute anything. Mixing those two responsibilities up in an interview answer — "then I call `tape.gradient` and the weights are updated" — is an immediate tell that the candidate has only ever used `fit()`.

  • What happens if you call tape.gradient twice on the same tape?
    A default (non-persistent) tape releases its recording after the first call, so the second raises a RuntimeError. Construct it with `tf.GradientTape(persistent=True)` when you need several gradients from one forward pass — for two losses, disjoint variable sets, or higher-order derivatives — and `del tape` afterwards so the retained buffers are freed.
  • How would you accumulate gradients over several micro-batches before applying them?
    Keep a list of `tf.Variable` accumulators shaped like the trainable variables. For each micro-batch, run a tape, and `assign_add` the resulting gradients into the accumulators. After N micro-batches, divide by N if your loss is a mean, call `apply_gradients` once with the accumulators, then zero them. TensorFlow does not accumulate for you, so this is explicit.
  • Why pass model.trainable_variables rather than model.variables?
    `model.variables` includes non-trainable state such as batch-normalization moving mean and variance, and any weight a layer has frozen. Those are updated by their own layer logic or not at all; handing them to an optimizer either errors on a None gradient or corrupts them. `trainable_variables` is exactly the set the optimizer is supposed to own.

saying these in an interview costs you the question

  • Says TensorFlow needs a gradient-zeroing call each step
  • Computes the loss outside the tape's with block
  • Thinks tape.gradient itself updates the weights
  • Passes model.variables instead of model.trainable_variables
  • Reuses one non-persistent tape for two gradient calls

context