In Keras, when do you use ReduceLROnPlateau instead of LearningRateScheduler?
answer
- fixed curve versus reactive cut
- one reads epochs, one reads metrics
- factor multiplies, cooldown suppresses
- must fire before early stopping does
- optimizer schedules do not compose
basics
~20 sLearningRateScheduler applies a rate you compute from the epoch index — a curve fixed before the run starts. ReduceLROnPlateau is reactive: it multiplies the current rate by factor only after a monitored metric fails to improve for patience epochs. Use it when you cannot pick the curve in advance.
solid answer
~40 sBoth are callbacks that write to `model.optimizer.learning_rate`, but they decide the value differently. `keras.callbacks.LearningRateScheduler(schedule)` calls your function at the start of each epoch with the epoch index and the current rate, and installs whatever float you return — step decay, exponential decay, warm-up, anything expressible as a function of the epoch. `keras.callbacks.ReduceLROnPlateau(monitor="val_loss", factor=0.5, patience=3, min_lr=1e-6)` instead watches a metric at epoch end and cuts the rate by `factor` only when it stalls, with `cooldown` epochs of silence afterwards so it does not fire repeatedly on the same plateau. The key gotcha is composition: if the optimizer was built with a `keras.optimizers.schedules` object, that schedule computes the rate itself per step, and a callback assigning a float over `optimizer.learning_rate` replaces it. Pick one mechanism per run.
code
python · 19 linesimport keras
def step_decay(epoch, lr):
return lr * 0.1 if epoch in (30, 60) else lr
fixed = keras.callbacks.LearningRateScheduler(step_decay, verbose=1)
reactive = keras.callbacks.ReduceLROnPlateau(
monitor="val_loss",
mode="min",
factor=0.5,
patience=3,
cooldown=1,
min_lr=1e-6,
verbose=1,
)
stopper = keras.callbacks.EarlyStopping(monitor="val_loss", patience=8)go deeper
Know that one callback applies a rate you compute from the epoch number, while the other cuts the rate only when a watched metric stops improving, and that both are passed to fit through callbacks.
Explain the arguments that matter — schedule signature, factor, patience, cooldown, min_lr — and that both callbacks work by writing to model.optimizer.learning_rate at epoch boundaries.
Show the coordination judgment: patience ordering against EarlyStopping so the reducer can actually fire, why an optimizer-level schedule and a rate callback cannot both own the value, and how you log the effective rate to prove what ran.
Own the choice of mechanism across a training platform — reproducible fixed recipes versus adaptive reduction for unfamiliar data — and how that choice interacts with batch size, run length and the cost of an unfinished sweep.
## Two different questions Both callbacks change the learning rate during `fit()`, but they answer different questions. `LearningRateScheduler` answers "what should the rate be at epoch *n*?" — the answer is known before the run starts. `ReduceLROnPlateau` answers "has progress stalled, and should I cut the rate now?" — the answer depends on what the run does. ## LearningRateScheduler `keras.callbacks.LearningRateScheduler(schedule, verbose=0)` takes a callable. Keras invokes it in `on_epoch_begin` with the zero-based epoch index and the current learning rate, and assigns the returned float to the optimizer. Because it is your function, any epoch-indexed shape is available: ``` def step_decay(epoch, lr): return lr * 0.1 if epoch in (30, 60) else lr ``` The granularity is per epoch, not per step. If you want a rate that changes every batch — a warm-up over the first 500 steps, say — a per-epoch callback cannot express it, and you want an optimizer-level schedule instead. ## ReduceLROnPlateau `keras.callbacks.ReduceLROnPlateau(monitor="val_loss", factor=0.1, patience=10, mode="auto", min_delta=1e-4, cooldown=0, min_lr=0.0)` runs in `on_epoch_end`. It keeps a best value and a wait counter exactly like `EarlyStopping`: improvement by more than `min_delta` resets the counter, and when the counter reaches `patience` the callback multiplies the current rate by `factor` (0.1 by default — an aggressive order-of-magnitude cut; 0.5 is the more common hand-set value), clamps it at `min_lr`, and enters a `cooldown` period during which the counter is not incremented. Cooldown exists because a rate cut takes a few epochs to show up in the metric, and without it the callback would fire again immediately and drive the rate to `min_lr` in a handful of epochs. Because the mechanics mirror EarlyStopping, the same configuration errors apply: `mode` must match the metric direction, the `monitor` key must actually exist in the epoch logs, and `min_delta` must be large enough that noise does not count as improvement. ## Using both, and the ordering with EarlyStopping ReduceLROnPlateau and EarlyStopping are frequently used together on the same metric, and the only rule that matters is that the reducer's patience must be shorter than the stopper's. If EarlyStopping has `patience=3` and ReduceLROnPlateau has `patience=5`, the run always stops before the rate is ever cut, and the reducer is dead code. A typical pairing is `patience=3` on the reducer and `patience=8` on the stopper, giving the model two or three rate cuts to escape a plateau before the run is abandoned. ## The conflict with optimizer-level schedules Keras offers a second, independent mechanism: `keras.optimizers.schedules` objects such as `ExponentialDecay` or `CosineDecay`, passed as the optimizer's `learning_rate`. These are evaluated by the optimizer itself on every step, using the iteration counter, which is why they can express per-step warm-up and decay that a per-epoch callback cannot. The two mechanisms do not compose. A schedule object *is* the value of `optimizer.learning_rate`; a callback that assigns a float there replaces the object, so the schedule stops applying from that moment on. Likewise `LearningRateScheduler` reads the current rate to hand to your function, and a schedule object is not a plain float. The resulting behaviour is confusing rather than obviously broken — hence the rule: either the optimizer owns the schedule, or the callbacks do. Never both. ## How to verify what is actually happening The rate is readable in a callback as `self.model.optimizer.learning_rate`, so a three-line custom callback that appends it to the logs dict in `on_epoch_end` puts the true rate into `history.history` and into any CSV log. Both built-in callbacks accept `verbose=1`, which prints on each change. Given how easy it is to configure a scheduler that silently never fires, printing or logging the rate is worth the two lines on any run where the schedule matters. ## Which to reach for Use a fixed schedule when you are reproducing a known recipe, when the run is long and well-characterised, or when you need sub-epoch shaping like warm-up. Use ReduceLROnPlateau when you are exploring, when the dataset is new, or when run length varies — it needs no tuning of the decay points, only a sensible `factor` and a `patience` shorter than your stopping rule.
- You use ReduceLROnPlateau and EarlyStopping together. What relationship must their patience values have?The reducer's patience must be clearly shorter than the stopper's, or training ends before the rate is ever cut and the reducer is dead code. A common pairing is patience 3 on the reducer against 8 on the stopper, which allows two or three cuts to rescue a plateau before the run is abandoned. Check the printed or logged rate to confirm the reducer actually fired.
- Why does cooldown exist on ReduceLROnPlateau?A rate cut needs several epochs before its effect shows up in the monitored metric. Without a cooldown the callback would see the still-flat metric on the very next epoch, fire again, and repeat — collapsing the rate to min_lr within a few epochs of the first plateau. Cooldown suspends the patience counter for that many epochs so each cut gets a fair trial.
- Your team wants learning-rate warm-up over the first 500 steps. Can LearningRateScheduler do it?Not cleanly. That callback is invoked once per epoch, at epoch begin, so its finest granularity is one value per epoch — with 500 steps sitting inside a single epoch there is nowhere to express the ramp. Sub-epoch shaping belongs to an optimizer-level schedule from keras.optimizers.schedules, which the optimizer evaluates per step from its iteration counter.
- How do you record the actual learning rate that was used each epoch?Read `self.model.optimizer.learning_rate` in a small custom callback's `on_epoch_end` and write it into the logs dict. Because History copies the epoch logs, the value lands in `history.history` and in any CSVLogger output. Both built-in rate callbacks also accept verbose=1 and print on each change, which is the quickest way to confirm a scheduler is firing at all.
saying these in an interview costs you the question
- Thinks ReduceLROnPlateau follows a preset decay curve
- Sets reducer patience longer than early-stopping patience
- Expects a callback and an optimizer schedule to compose
- Believes factor is subtracted from the learning rate
- Assumes LearningRateScheduler can change the rate per batch