skip to content

Online and Incremental Learning

Batch learning refits on a whole dataset; online learning updates per example or mini-batch, which you need when data arrives faster than a retrain. Interviewers probe drift and cost here.

on this pageshow

questions

4

How does refitting a model in batch differ from updating it incrementally as new data arrives?

level: middleimportance: must knowfreq 58%

answer

  1. keep the data, or keep the model?
  2. one full run versus many tiny edits
  3. sufficient statistics can simply be added
  4. a split needs every row at the node
  5. path-dependent artifact, no clean rollback

basics

~20 s

Batch refitting discards the current model and fits a new one over the whole dataset. Incremental (online) learning keeps the existing parameters and nudges them with each new example or chunk, without revisiting the older rows.

solid answer

~50 s

Batch means you periodically fit a fresh model on the whole dataset; incremental means the current model absorbs each new example or chunk as a small parameter update and never touches the old rows again. Batch is reproducible — the artifact is a function of the data plus the seed — and any algorithm can be used, but it costs a full training run and new data only reaches predictions at the next cycle. Incremental reflects new data in seconds and needs one example in memory at a time, but only some families support it: anything whose fit is running counts or a gradient-updated parameter vector can, while a random forest or a single CART tree cannot, because each split was chosen from statistics over all rows at that node. Boosted ensembles sit in between: you can append trees, but the ones already there stay frozen.

go deeper

for a junior

Be ready to state the plain difference: a batch refit builds a new model over all the data, an incremental update edits the model you already have using only the new rows. Know that not every algorithm supports the second.

for a middle

Explain the mechanics of why: which families keep sufficient statistics or a gradient-updated parameter vector and can absorb a row, and why a tree's split, chosen from all rows at a node, cannot be revised by one more example.

for a senior

Show the operational side. Talk about what you lose when the artifact is path-dependent, how you snapshot and roll back parameters, and why teams usually run incremental updates on top of a periodic cold refit rather than instead of one.

for a principal

Own the framing: staleness tolerance, data volume and reproducibility requirements are what pick the regime, not novelty. Be able to argue that an online learner is an operational commitment — snapshots, replay, incident recovery — and say when that cost is not worth paying.

## Two ways to keep a model current A model is a summary of the data it was fitted on. When more data arrives, you have two structurally different options. **Batch refitting (the default).** You keep the historical data, and every so often you throw the current model away and run the whole fitting procedure again over everything — old rows plus new. The output is a new artifact that has no memory of the previous one. This is what almost every tabular pipeline does. **Incremental (online) updating.** You keep the *model*, not the data. Each new example, or each chunk of examples, is used to adjust the parameters you already have, and is then discarded. After the update the row is gone; its influence survives only inside the parameters. Note what this axis is *not* about. Whether one gradient step is computed from one row, 256 rows, or the entire training set is a question about the optimizer inside a single fit. Batch-versus-online here is about *when the model is refit as the world produces new data* — a deployment question, not an optimizer question. Candidates routinely conflate the two, and interviewers listen for it. ## Which algorithms can actually be updated Incremental updating is not a property you can bolt on; it depends on how the method stores what it learned. - **Count-based models.** If the fit is a table of counts — class frequencies and per-feature frequencies — those counts are *sufficient statistics*: merging two datasets is just adding two tables. New data means incrementing counters, and the result is bit-for-bit what a full refit would have produced. - **Gradient-updated parameter vectors.** A linear or logistic model fitted by stochastic updates has a natural per-example update rule: `w <- w - lr * grad`. Feeding it a new row is exactly what training already did. - **Centroid models.** A streaming k-means assigns the incoming point to its nearest centroid and moves that centroid a small step toward it. - **Stream-designed trees.** Hoeffding trees (very fast decision trees) were built for this: they accumulate statistics at each leaf and use a statistical bound to decide when enough examples have been seen to commit to a split. And which cannot: - **CART trees and random forests.** A split threshold is chosen by scanning candidate cut points against the statistics of *all* rows reaching that node. One extra row cannot revise that choice; honouring it means regrowing the tree. "Add data" and "refit" are the same operation here. - **Gradient-boosted ensembles** (including XGBoost, LightGBM and CatBoost as algorithms) are a partial case. You can continue training by appending trees fitted to the current residuals, but every tree already in the ensemble is frozen and was fitted against the old data's residuals. Append long enough on new-only data and you get a hybrid whose early stages describe a world that no longer exists. ## What each costs you | | Batch refit | Incremental update | |---|---|---| | Data you must keep | all of it | only the current row/chunk | | Compute per update | a full training run | one cheap step | | Time for new data to reach predictions | one cycle | seconds | | Reproducibility | rerun on the data, get the model back | depends on the exact order of every update ever applied | | Algorithm choice | anything | only updatable families | | Recovery from bad data | refit without it | restore a parameter snapshot; the bad rows are unrecoverable | That reproducibility row is the one people underrate. A batch artifact can be regenerated from the data; an online model is a path-dependent object, and if a corrupt hour of events got folded in, there is no "remove those rows" — only rolling back to a saved snapshot of the parameters and replaying. ## A worked case A ride-hailing surge model predicts a multiplier from recent demand, supply and weather. Refitting it from scratch every hour on the trailing weeks is simple, reproducible and fully auditable, and an hour of staleness is usually tolerable — but it is a full training run every hour, and during a stadium emptying out, an hour is an eternity. Updating the same model on each chunk of completed trips as they close costs almost nothing per update and tracks the surge as it happens, at the price of an artifact whose exact state depends on the order chunks arrived in, and which no one can regenerate from the warehouse. The usual production answer is not one or the other: incremental updates between cycles for freshness, plus a periodic cold refit from scratch as the anchor that resets accumulated path dependence and restores reproducibility. ## How to decide Ask three questions. How stale can a prediction be before it is wrong in a way that costs money? Does the data even fit in memory for a batch fit? And can you afford an artifact you cannot regenerate? If staleness is cheap and the data fits, batch refitting is the boring correct answer, and choosing online learning without one of those pressures is complexity you will pay for at the first incident.

  • Why can a count-based model be updated exactly, while a tree cannot?
    Because counts are sufficient statistics: the fit depends on the data only through totals, and totals are additive, so incrementing them gives exactly the model a full refit would have produced. A tree's split is a choice made by comparing candidate thresholds across all rows at a node — the choice itself, not just a total, would have to change, and that means regrowing the node.
  • If you do run incremental updates, why still schedule an occasional refit from scratch?
    The cold refit is the anchor. It resets path dependence, so the artifact becomes reproducible from the stored data again; it lets you change the feature set, the encoding or the hyperparameters, which an in-place update cannot; and it flushes out any bad data that got folded into the parameters and can no longer be subtracted.
  • Where does 'the model keeps learning' actually live in a serving system?
    Usually not in the serving process. The updates are applied by a separate learner that emits a versioned parameter snapshot, and serving loads snapshots. That keeps a bad update from being irreversible, keeps every replica scoring with identical parameters, and means a rollback is loading the previous snapshot rather than untraining a live process.

Batch refitting is redrawing the whole map from every survey ever taken. Incremental updating is pencilling in each new road as the surveyor phones it in, and never looking at the old surveys again.

saying these in an interview costs you the question

  • Confuses batch refitting with full-batch gradient descent inside one fit
  • Claims any model can be updated one example at a time
  • Says an incremental update always equals a full refit's parameters
  • Thinks online learning is mainly a speed optimization
  • Cannot name a single family that supports incremental fitting

context

open as a page

Why does a model that is warm-started and then updated only on recent data forget older patterns?

level: middleimportance: should knowfreq 40%

basics

~20 s

A warm start carries old data in only as a starting point. Once every later update uses recent rows alone, each step moves the parameters further from that start, so an old example's influence decays away.

open as a page

How do you train a model on a 400 GB file that never fits in memory?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Fit out-of-core: stream the file in fixed-size chunks and update a learner that has a per-chunk update rule, so only one chunk sits in memory. First check whether a sample or fewer columns removes the problem.

open as a page

When is a deliberately frozen model the right choice over one that keeps updating itself?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Freeze the model when validating a new version costs more than staleness does, when a past decision must be explainable against a specific artifact, or when its outputs shape future labels. Freezing trades silent decay for reviewable behaviour.

open as a page