In Keras, what does model.compile() configure before you call fit()?
answer
- configuration, not architecture
- three arguments you almost always pass
- strings are shorthand for objects
- metrics reported, loss optimised
- predict() needs none of it
basics
~20 smodel.compile() attaches the training configuration to a model: the optimizer that updates the weights, the loss that fit() minimises, and the metrics reported each epoch. Until you call it, fit() and evaluate() have nothing to optimise and raise an error.
solid answer
~40 s`compile()` is where you declare *how* the model trains, separately from *what* the model is. The three arguments you almost always pass are `optimizer` (how weights move — `"adam"` or `keras.optimizers.Adam(learning_rate=1e-3)`), `loss` (the scalar fit() minimises), and `metrics` (extra numbers reported but never optimised). Strings are shorthand for the default-configured object; pass the object itself whenever you need to set a hyperparameter such as a learning rate or `from_logits=True`. `compile()` also stores training-execution options like `jit_compile` and `run_eagerly`, and it instantiates the metric objects, so calling it twice resets them. `fit()` and `evaluate()` both require it; `predict()` and calling the model directly do not, because neither needs a loss.
code
python · 12 linesimport keras
model = keras.Sequential([
keras.layers.Dense(64, activation="relu"),
keras.layers.Dense(10),
])
model.compile(
optimizer=keras.optimizers.Adam(learning_rate=1e-3),
loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=["accuracy"],
)go deeper
Be able to name the three arguments — optimizer, loss, metrics — and say in one sentence what each does. Know that fit() fails without compile() and that predict() does not need it.
Explain that strings are shorthand for default-configured objects, and that you pass the object whenever a hyperparameter matters. Be ready to say why a loss must be differentiable and a metric need not be.
Show that you know compile() rebuilds stateful metric objects, that it creates no weights, and that the compiled metric names are what callbacks and History keys are written against. Mention run_eagerly as the debugging switch.
Frame compile() as the seam that keeps architecture reusable across training regimes — the same graph compiled with different optimizers, losses and jit settings for pretraining, fine-tuning and evaluation, so configuration lives in one auditable place rather than scattered through a loop.
## The split Keras makes Keras separates a model's *structure* from its *training configuration*. Layers and their connections are the structure; the optimizer, the loss function and the reported metrics are the configuration. `compile()` is the call that attaches the second to the first. This is why the same `Sequential` or Functional model can be compiled two different ways — once with a small learning rate for fine-tuning, once with a large one for a from-scratch run — without rebuilding a single layer. ## The arguments `Model.compile()` in Keras 3 accepts, among others: - **`optimizer`** — the algorithm that turns gradients into weight updates. `"adam"`, `"sgd"`, `"rmsprop"` are string shorthands for the corresponding class with all defaults. Pass an instance (`keras.optimizers.Adam(learning_rate=3e-4)`) when you need to change anything. - **`loss`** — the scalar quantity training minimises. Again a string (`"binary_crossentropy"`, `"mse"`, `"sparse_categorical_crossentropy"`) or an object (`keras.losses.SparseCategoricalCrossentropy(from_logits=True)`). A model with several outputs can take a list or a dict of losses keyed by output name, combined with `loss_weights`. - **`metrics`** — a list of quantities *reported* but never optimised. `"accuracy"` is the common one; anything in `keras.metrics` can be passed as an object. - **`weighted_metrics`** — the same, but these receive `sample_weight` when you pass one to `fit()`; plain `metrics` do not. - **`run_eagerly`** — force step-by-step Python execution so you can put a `print` or a breakpoint inside a custom step. Slower; a debugging switch only. - **`jit_compile`** — whether to compile the step with XLA where the backend supports it. Defaults to `"auto"`. - **`steps_per_execution`** — how many batches to run per outer execution call, which reduces per-step overhead on small models. ## What happens at the call Two practical consequences follow from `compile()` being a real state change rather than a declaration. First, the strings you pass are **resolved into objects**, and the metric objects are stateful variables that live on the model. Calling `compile()` a second time throws those away and builds new ones, so metric history restarts. This is normally what you want when you change the configuration, and a surprise if you call `compile()` inside a loop and wonder why numbers reset. Second, `compile()` performs **no shape inference and no weight creation**. A subclassed model still has no weights after compiling; those appear when the model first sees data (in `fit()` or a manual call). So `compile()` succeeding tells you nothing about whether your loss matches your label shape — that error surfaces on the first batch. ## Loss versus metric Candidates often blur these. The loss must be differentiable, because gradients flow through it; a metric need not be, which is exactly why accuracy — a step function of `argmax` — can be a metric but never a loss. The loss drives learning; metrics exist for humans and for callbacks that monitor them. If a metric is what you actually care about (F1, AUC), you still train on a differentiable surrogate such as cross-entropy and *report* the metric. ## Which calls need it - `fit()` — requires `compile()`; it needs an optimizer and a loss. - `evaluate()` — requires it; it needs a loss and metrics to report. - `predict()` and `model(x)` — do **not**. A loaded inference-only model can go straight to `predict()`. If you load a model with `keras.saving.load_model()` from a `.keras` file, the compiled configuration is restored too, so you can resume training without re-compiling. ## The reporting contract `fit()` returns a `History` object whose `.history` attribute is a dict of lists — one entry per compiled metric plus `"loss"`, and the same keys prefixed with `val_` when validation data is supplied. Those key names come directly from what you passed to `compile()`, which is why callbacks that monitor a metric are written against the compiled name (`"val_loss"`, `"val_accuracy"`) and silently do nothing when the name does not exist. ## A minimal, correct call ``` model.compile( optimizer=keras.optimizers.Adam(learning_rate=1e-3), loss=keras.losses.SparseCategoricalCrossentropy(from_logits=True), metrics=["accuracy"], ) ``` Read it as three declarations: move the weights this way, minimise this number, and print that one alongside it.
- Does predict() work on a model that was never compiled?Yes. `predict()` only runs the forward pass, and calling the model directly as `model(x, training=False)` does the same. Neither needs an optimizer or a loss. Only `fit()` and `evaluate()` require a compiled model, because they need something to minimise and something to report.
- When would you pass an optimizer object instead of the string "adam"?Whenever you need to set anything on it: a non-default `learning_rate`, a learning-rate schedule object, `clipnorm`/`clipvalue` for gradient clipping, or `weight_decay`. The string `"adam"` is exactly `keras.optimizers.Adam()` with every default, which is rarely the learning rate you want for fine-tuning.
- Why can accuracy be a metric but not a loss?Accuracy is computed from `argmax` over the outputs, so it is a step function — its gradient is zero almost everywhere and undefined at the jumps. Training needs a differentiable signal, so you minimise cross-entropy and report accuracy. That gap is also why loss and accuracy can move in opposite directions within an epoch.
saying these in an interview costs you the question
- Thinking compile() builds the weights or infers shapes
- Believing predict() requires a compiled model
- Treating metrics as things training optimises
- Assuming re-compiling preserves accumulated metric state
- Passing "adam" then claiming a custom learning rate was set