skip to content

Why does a model that is warm-started and then updated only on recent data forget older patterns?

level: middleimportance: should knowfreq 40%

answer

  1. the start point is not the data
  2. convex plus converged equals same answer
  3. old rows leave the objective entirely
  4. influence decays with every later step
  5. step size sets the memory length

basics

~20 s

A warm start carries old data in only as a starting point. Once every later update uses recent rows alone, each step moves the parameters further from that start, so an old example's influence decays away.

solid answer

~50 s

A warm start means initializing the next fit from the current parameter values instead of from scratch. On the same data with a convex objective run to convergence, that changes only the time taken, not the answer. The forgetting comes from what you fit on afterwards: if each update's objective contains only the new rows, the old data survives solely through the initialization, and every subsequent step pushes the parameters away from it. With a fixed step size an example's influence decays geometrically, so the model has an effective memory window rather than a full history. Sometimes that is exactly what you want — a fashion catalogue that turns over every six weeks should forget. It is a bug when the recent chunk is small, seasonal or unrepresentative, and one quiet week can wash out a year of signal.

go deeper

for a junior

Know what a warm start is: the next fit begins from the current parameter values instead of from scratch. Remember that training only on the newest rows means the model gradually stops reflecting older ones.

for a middle

Explain the mechanism precisely: on a converged convex fit a warm start changes only runtime, and the forgetting comes from removing old rows from the objective, after which influence decays with each later step at a rate set by the step size.

for a senior

Show you have operated this. Discuss mixing sampled history into each update, evaluating on a frozen holdout that predates the updates, capping step sizes so one bad chunk cannot move the model, and periodic cold refits as the anchor.

for a principal

Own the tradeoff between adaptivity and stability as a product decision: how much recency the domain actually rewards, whether recency should be an explicit weighting you chose rather than an emergent decay, and what the failure looks like when the recent slice is unrepresentative.

## What a warm start actually is Fitting an iterative model requires a starting point for the parameters. A cold start uses zeros or a random draw. A **warm start** uses the parameter values from the previous fit. That is the whole idea — it changes the initialization, nothing else. Two consequences follow, and confusing them is the usual interview mistake. **On the same data, with a convex objective, run to convergence, a warm start changes nothing but the runtime.** A convex problem — say a linear model under log loss with an L2 penalty — has a single minimum, and every starting point reaches it. Warm-starting the hourly refit on the full trailing window is therefore a pure speed optimization: fewer iterations, identical answer, and it is the safest form of warm starting there is. **Two caveats to that.** If the objective is not convex, the starting point selects which optimum you land in, so warm and cold starts genuinely differ. And if you stop early rather than converging — a fixed iteration budget, or early stopping on a validation score — the initialization leaks into the result even on a convex problem, because you never got to the point where the start stops mattering. That is why comparing a warm-started model against a cold-started one in an experiment is not a clean comparison: pin the initialization or you are measuring two things at once. ## Where the forgetting comes from None of the above forgets anything. Forgetting appears when you change **what the update is computed from**. A warm start plus a fit on the full history is just a fast refit. A warm start plus updates on *only the new rows* is a different animal: at that point the old data appears in the objective not at all, and enters the result only through where the parameters happened to be when the update began. Watch what happens to one example. It contributes an update that moves the weight vector by some amount. The next update, computed from newer rows, moves it again — partly along the same direction, partly against it. With a constant step size, the contribution of an example is multiplied down by roughly a constant factor per subsequent update, so its influence decays geometrically and the model effectively remembers a trailing window whose length is set by the step size. Big steps mean a short memory and a twitchy model; small steps mean a long memory and a sluggish one. There is no setting that gives you both. The step-size schedule is the lever, and it has a trap in each direction. A decaying schedule (steps shrinking as more examples are seen) is what makes stochastic optimization converge — but a model that has converged has also stopped listening, so a learner meant to track a moving world needs a floor under its step size, or it will freeze into last quarter's picture. A constant step size never freezes, but never settles either: the parameters keep jittering around the current optimum forever. ## When forgetting is the feature A fashion marketplace whose catalogue turns over every six weeks *wants* an aggressively short memory: last season's relationships between price, imagery and conversion are not merely stale, they are about items nobody can buy. A surge-pricing model wants to weight this evening far above last month. In those settings the decay is doing the job you would otherwise have to write by hand with sample weights. It turns into a bug when the recent slice is not representative: - **Small chunks.** A quiet overnight window carries little signal but still gets a full round of updates, so noise gets the same authority as a busy afternoon. - **Seasonality.** Train on December for a month and January's model has a strong opinion about Christmas. - **Rare classes and rare events.** Whatever is rare is absent from most chunks and is forgotten between appearances — precisely the cases you care about in fraud or failure prediction. - **A bad hour of data.** Corrupt events fold into the parameters and cannot be subtracted afterwards; only a snapshot restore removes them. ## Controlling it - **Mix history into every update.** Keep a reservoir of sampled historical rows and blend them into each chunk, so the objective is never purely recent. This is the standard defence and costs a little storage. - **Make recency explicit.** Weight examples by age deliberately, rather than letting the step size decide the memory length as a side effect. A weight you chose can be reasoned about and tuned; an implicit decay cannot. - **Anchor with a periodic cold refit** on the full retained history, so accumulated path dependence is reset on a known cadence and the artifact is regenerable again. - **Evaluate against a frozen holdout that predates the updates.** A model that has forgotten looks fine on the recent data it just trained on, and only a stable reference set exposes it. - **Cap the step size and validate before promoting a snapshot,** so no single anomalous chunk can move the model far. ## The one-line version Warm starting is about *initialization* and is usually harmless. Forgetting is about *which rows are in the objective*, and it is a design decision — either an intended, tuned recency preference or an accident with a decay rate nobody chose.

  • A marketplace's catalogue turns over every six weeks, so the model keeps scoring items it never trained on. Do incremental updates solve that?
    Only partly. Updates adjust parameters, but the feature space is usually fixed at fit time: a category encoding built from last quarter's catalogue has no slot for a new item, and unseen values collapse into a default bucket. Fixing that needs a scheme that tolerates new values — hashing into a fixed width, or features describing the item's attributes rather than its identity — which is a refit-level change, not an update.
  • How would you tell that an online model has forgotten something rather than genuinely improved?
    Score it on a frozen holdout that predates the updates as well as on recent data. A model that has simply adapted holds up on both; one that has forgotten looks strong on the recent slice and has quietly lost ground on the stable reference, usually concentrated in segments or classes that are rare in the recent chunks.
  • Is warm-starting an experiment comparison ever a problem?
    Yes, whenever fitting stops before convergence. Under a fixed iteration budget or early stopping, the initialization still influences the final parameters, so a warm-started candidate and a cold-started baseline differ by two things at once. Pin both to the same initialization, or run both to convergence, before attributing the difference to the change you were testing.

It is a conversation where you only remember the last ten minutes: nothing was deleted on purpose, but each new sentence pushes an older one out of reach.

saying these in an interview costs you the question

  • Says a warm start changes the solution of a converged convex fit
  • Treats warm start and incremental learning as the same thing
  • Assumes updating on new data keeps old data equally weighted
  • Never mentions the step size as the memory control
  • Evaluates an updated model only on the most recent data

context