In Keras fit(), how does validation_split differ from validation_data?
answer
- one you carve, one you supply
- positional, not random
- shuffling happens after the cut
- arrays only, not datasets
- sorted data poisons the held-out slice
basics
~20 svalidation_split=0.2 tells fit() to carve a validation set out of the arrays you passed, taking the last 20% of the rows before any shuffling. validation_data hands fit() a set you split yourself. The split argument works only for in-memory arrays.
solid answer
~40 sBoth produce the `val_`-prefixed numbers reported at the end of each epoch, but they differ in who chooses the rows. `validation_split=0.2` makes `fit()` slice the **last** fraction of `x` and `y` — and it does that *before* the `shuffle=True` pass, so shuffling only ever reorders the training remainder. On data sorted by class, date or user, that hands you a validation set drawn from one region of the distribution, and the val metrics are meaningless. It also only works for in-memory arrays: a `tf.data.Dataset`, a generator or a `keras.utils.PyDataset` cannot be sliced that way. `validation_data=(x_val, y_val)` is the explicit form — you did the split, so you control stratification, grouping and time ordering, and it accepts a dataset. Prefer it for anything but a quick smoke test.
code
python · 11 linesimport numpy as np
import keras
x = np.random.rand(1000, 8)
y = np.sort(np.random.randint(0, 2, 1000)) # sorted by label
model = keras.Sequential([keras.layers.Dense(1, activation="sigmoid")])
model.compile(optimizer="adam", loss="binary_crossentropy", metrics=["accuracy"])
# validation set = rows 800-999, all label 1, chosen before any shuffling
model.fit(x, y, validation_split=0.2, epochs=2, verbose=0)go deeper
Know that both arguments produce val_loss and val_accuracy at the end of each epoch, and that validation_split lets fit() do the split for you while validation_data means you did it yourself.
Explain that validation_split takes the last fraction positionally and before shuffling, and that it does not accept datasets or generators. Name what goes wrong on sorted data.
Demonstrate that you choose the split policy from the problem — stratified, grouped by entity, or chronological — and that you can explain the training-versus-validation measurement gap from dropout, batch-norm and the running epoch average.
Own the evaluation protocol as a contract: which split policy the domain demands, how it is generated reproducibly outside the training call, and why a convenience argument that silently produces a non-representative held-out set is not acceptable in a pipeline other people trust.
## Two ways to get val_ numbers When `fit()` runs an epoch it can, at the end of that epoch, run a forward-only pass over a held-out set and report the compiled loss and metrics again under `val_`-prefixed names (`val_loss`, `val_accuracy`). Those names land in `History.history` and are what monitoring callbacks watch. There are two ways to tell `fit()` what that held-out set is. ## validation_split — convenience with a sharp edge `validation_split` is a float in `[0, 1)`. `fit()` interprets it as *the fraction of `x`/`y` to set aside*. Three properties of that slice matter, and interviewers ask about all three: 1. **It is the last fraction, positionally.** With 10,000 rows and `validation_split=0.2`, rows 8,000–9,999 become validation. There is no sampling and no stratification. 2. **It is taken before shuffling.** `fit(shuffle=True)` is the default, but the shuffle applies to the training portion each epoch — it does not re-draw the validation slice. So the same rows are validation for the whole run, and the whole run's validation signal comes from whatever happens to live at the end of your array. 3. **It only works on sliceable in-memory data.** NumPy arrays and tensors are fine. A `tf.data.Dataset`, a Python generator or a `keras.utils.PyDataset` is not — Keras cannot slice a stream it has not consumed, so passing `validation_split` with one of those is an error. Point 2 is where real runs go wrong silently. Data exported from a database is very often ordered: by label (all the negatives, then all the positives), by ingestion date, by customer. Take the last 20% of a label-sorted file and your validation set is a single class; accuracy is either 0 or 1 and never moves, or `val_loss` is wildly detached from `loss`. Nothing raises. The fix is to shuffle the arrays *yourself* before calling `fit()`, or to stop using `validation_split` at all. ## validation_data — explicit, and the default choice `validation_data` accepts a `(x_val, y_val)` tuple, a `(x_val, y_val, sample_weight_val)` triple, or a dataset/`PyDataset` that yields those. Because you build it, you control the split policy: stratify by class, group by user so no user appears on both sides, or cut by time so validation is strictly in the future. Those policies are exactly what `validation_split` cannot express. When both arguments are passed, `validation_data` wins and the split is ignored. ## The related knobs - **`validation_freq`** — run validation every *n* epochs instead of every one. Useful when validation is expensive relative to an epoch. - **`validation_steps`** — with an infinite or streaming validation source, how many batches count as one validation pass. Without it, `fit()` would run forever at the first epoch boundary. - **`validation_batch_size`** — a separate batch size for the validation pass; handy when validation activations are larger (no gradients are stored, so it can usually be bigger, not smaller). ## Why the val numbers legitimately differ from the training numbers Even with a perfect split, `loss` and `val_loss` are not measured the same way, and a candidate who explains this stands out: - The **training** number printed for an epoch is a running average over the batches of that epoch, accumulated while the weights were still changing. The **validation** number is computed once, at the end, with the final weights. Early in training, when the model improves fast within an epoch, `val_loss` below `loss` is normal and is not a bug. - Layers behave differently. `Dropout` is active during training and inert during validation; `BatchNormalization` uses batch statistics while training and its moving averages otherwise. Both push training loss up relative to validation loss. ## Practical rule Use `validation_split` for a five-line smoke test on data you know is already shuffled. Use `validation_data` for anything whose numbers you will quote, and build the split with the policy the problem demands — stratified, grouped or chronological. When you are stuck with `validation_split`, shuffle the arrays with a fixed seed first so at least the slice is representative, and remember that a fixed slice means every epoch is validated against the same rows.
- Your data is sorted by label and you used validation_split=0.2 — what do you see?A validation set made of one class. `val_accuracy` pins at a constant, `val_loss` sits far from `loss` and barely moves, while training metrics look healthy. Nothing raises, because Keras did exactly what you asked. Shuffle the arrays before `fit()` or, better, build an explicit stratified split and pass it as `validation_data`.
- Why can val_loss be lower than the training loss reported for the same epoch?Two reasons that both favour validation. The training number is a running average accumulated across the epoch while weights were still improving, whereas validation is measured once with the final weights. And `Dropout` is on during training but off during validation, and `BatchNormalization` switches from batch statistics to its moving averages. Early in a run this gap is expected.
- You pass a tf.data.Dataset to fit() and set validation_split — what happens?It fails. `validation_split` needs to slice the input positionally, and a dataset, generator or `keras.utils.PyDataset` is a stream Keras cannot index before consuming. Split the pipeline yourself and pass the held-out dataset as `validation_data`, adding `validation_steps` if that source does not terminate.
saying these in an interview costs you the question
- Believing validation_split samples rows at random
- Thinking shuffle=True reshuffles the validation slice
- Assuming validation_split works with a tf.data pipeline
- Calling a lower val_loss than loss a bug
- Expecting validation_split to stratify by class