skip to content

What must be created inside a tf.distribute strategy.scope(), and why?

level: middleimportance: must knowfreq 80%

answer

  1. It is a variable-creation context
  2. Ask what owns state, not what runs
  3. Optimizer slots count too
  4. Datasets and fit stay outside
  5. Loading a checkpoint creates variables as well

basics

~20 s

Anything that creates variables: the model, the optimizer, and any metric or custom tf.Variable. Inside the scope those become mirrored variables with one synchronized copy per replica. Datasets, fit calls, and the training step itself go outside.

solid answer

~40 s

`strategy.scope()` is a variable-creation context, not a training context. Every `tf.Variable` created inside it becomes a distributed variable — a `MirroredVariable` under `MirroredStrategy`, with one copy per replica kept identical by the strategy. So the model, its optimizer, and any stateful metrics or hand-made variables must be built inside the scope, which in Keras usually means both the model construction and `model.compile` go there. What does **not** belong inside: building the `tf.data.Dataset`, calling `fit`, `evaluate` or `predict`, and defining the `@tf.function` train step — those work fine outside because the variables they touch were already created correctly. If a variable is created outside the scope and then updated from replica context, it is an ordinary single-copy variable, and TensorFlow raises an error rather than silently training four unsynchronized models.

code

python · 20 lines
python
import tensorflow as tf

strategy = tf.distribute.MirroredStrategy()

# Inside: everything that creates variables.
with strategy.scope():
    model = tf.keras.Sequential([
        tf.keras.layers.Dense(16, activation="relu"),
        tf.keras.layers.Dense(1),
    ])
    model.compile(
        optimizer=tf.keras.optimizers.Adam(),
        loss="mse",
        metrics=[tf.keras.metrics.MeanAbsoluteError()],
    )

# Outside: data and the fit call.
x = tf.random.normal((512, 8))
y = tf.random.normal((512, 1))
model.fit(x, y, batch_size=64, epochs=1)

go deeper

for a junior

Be able to say that the model and optimizer are created inside strategy.scope() and that the dataset and fit call are not, and to recognize the pattern in code you are shown.

for a middle

Explain that the scope intercepts variable creation so each tf.Variable becomes a mirrored copy per replica, and name the state that qualifies: weights, optimizer slots, metric accumulators, custom variables.

for a senior

Show you have debugged this: the error when a non-distributed variable is updated from replica context, models built at import time before any scope exists, and reloading a checkpoint outside the scope before fine-tuning.

for a principal

Set the code convention that keeps this from recurring — a build_model() factory called inside the scope, no module-level model construction, and checkpoint restore paths that are scope-aware by design.

## What a scope actually is `with strategy.scope():` sets a thread-local default strategy. Its one real job is to intercept **variable creation**. While the scope is active, `tf.Variable(...)` does not produce a plain variable; it routes through the strategy's variable creator and produces a *distributed variable*. Under `MirroredStrategy` that is a `MirroredVariable`: one physical copy on each replica's device, presented as a single logical variable, with updates applied to every copy so they never diverge. Under `TPUStrategy` the mechanism differs in detail but the contract is the same. Nothing else about the scope is magic — it does not run anything in parallel, it does not distribute data, and leaving it does not stop distributed training. ## What goes inside The practical rule is: *anything that owns state that training will update.* - **The model.** Layer weights are created either at construction time or lazily on first build; building the model inside the scope covers both. - **The optimizer.** Optimizers own slot variables — momentum, Adam's moment estimates, the iteration counter. These are variables too, and they must be mirrored. - **Metrics.** A Keras metric is stateful; it holds accumulator variables. - **Any custom `tf.Variable`** you create for a schedule, a counter, a moving statistic. - **`model.compile`.** Compilation is where the optimizer and metrics get attached and, in practice, where some of their variables come into existence, so the conventional pattern is to put construction and compilation together inside one `with` block. - **Checkpoint objects** you intend to restore into, so restoration writes to the mirrored copies. ## What goes outside - **Dataset construction.** `tf.data` pipelines contain no trainable variables. Build the dataset normally and, in a custom loop, wrap it with `strategy.experimental_distribute_dataset(...)` or `strategy.distribute_datasets_from_function(...)`. - **`fit` / `evaluate` / `predict`.** Keras remembers which strategy the model was built under; calling `fit` inside the scope is harmless but unnecessary, and the common instruction "put everything inside the scope" is why people wrongly believe the scope is what makes training distributed. - **Your `@tf.function` train step.** It is called through `strategy.run(...)`, which enters *replica context* itself. ## Cross-replica context versus replica context These two contexts are worth naming, because error messages mention them. Inside `strategy.scope()` but outside `strategy.run(...)` you are in **cross-replica context**: you see one logical variable and can read its aggregated value. Inside a function passed to `strategy.run(...)` you are in **replica context**: your code runs once per replica, sees that replica's slice of the input, and writes to that replica's copy of the variable. `tf.distribute.get_replica_context()` returns non-`None` only in the second case. Gradient application and all-reduce are the bridge between them. ## The failure mode Create a model outside the scope and then compile it inside, and the weights are ordinary variables while the optimizer's slots are mirrored. When training tries to update a non-distributed variable from replica context, TensorFlow raises an error rather than proceeding — this is one of the few distribution mistakes the framework catches loudly, precisely because silently training divergent copies would be catastrophic. The lesson is not "the error is annoying" but "variable identity is the thing the scope establishes, so the scope must wrap creation, not usage." A subtler variant: a module-level model built at import time, then a scope opened later in `main()`. The weights were created long before, so no amount of later scoping fixes them. The fix is always to move creation inside, often by wrapping construction in a `build_model()` function called from within the `with` block. ## Loading a saved model Restoring a model or a checkpoint must also happen inside the scope, for the same reason — `keras.saving.load_model` creates the variables as it reads them, and outside the scope it creates undistributed ones. This trips people who train inside a scope, save, and then reload for fine-tuning without one. ## Nesting and the default strategy When no scope is active, TensorFlow uses a default no-op strategy, which is why unmodified single-device code works: `tf.distribute.get_strategy()` always returns *something*. That default is also what makes the "write once, scale later" story work — the same model code is correct with or without a strategy, provided creation and usage are cleanly separated.

  • Does calling model.fit outside the scope mean training is not distributed?
    No. The model remembers the strategy it was built under, so `fit` distributes either way. The scope governs variable creation, not execution — which is exactly why people mis-model it as an "enable distribution" switch. Calling `fit` inside the scope is allowed and harmless, it just is not what makes distribution happen.
  • Where must you load a previously saved model if you intend to keep training it under a strategy?
    Inside `strategy.scope()`. Loading creates the variables as it deserializes them, so loading outside produces ordinary, undistributed variables and the first training step fails. The same applies to restoring a `tf.train.Checkpoint` — create the checkpointed objects inside the scope, then restore.
  • What is the difference between cross-replica context and replica context?
    Inside the scope but outside `strategy.run`, you are in cross-replica context: one logical variable, aggregated reads. Inside the function passed to `strategy.run`, you are in replica context: the code executes once per replica against that replica's slice and its own variable copy. `tf.distribute.get_replica_context()` returns non-None only in the latter.

The scope is like the registry office where a variable gets its identity papers: it matters where the variable was born, not where it later goes to work.

saying these in an interview costs you the question

  • Believing the scope is what turns on distributed training
  • Building the model outside and only compiling inside
  • Creating the tf.data pipeline inside the scope for safety
  • Loading a saved model outside the scope before fine-tuning
  • Forgetting that optimizer slot variables need the scope too

context