skip to content

How would you serve a lazy instance-based model under a 10 ms budget as its store keeps growing?

level: seniorimportance: nice to knowfreq 36%

answer

  1. Lazy means paying at every request
  2. Store size drives memory and latency together
  3. Watch the tail, not the mean
  4. Cap, reduce, or distil into fixed size
  5. Distilling costs you instant freshness

basics

~20 s

A lazy model does its work at request time against everything it has stored, so memory and per-query cost both scale with the store. Bound what is stored, or distil the same data into a fixed-size eager model.

solid answer

~50 s

First name where the bill lands: instance-based methods defer everything to prediction, so the store is the model and query cost tracks its size. That is why 'training is free' is misleading — you pay per request, forever. Under a hard 10 ms budget I would measure p99 rather than mean latency against store size, then pick from a short list: cap the store with a recency window or a stratified sample, replace raw rows with a reduced set of prototypes, cache scores for repeated queries, or train a fixed-size eager model on the same data so scoring becomes constant-cost arithmetic. The last option is usually the honest answer under a tight budget, and its cost is what you give up: new rows no longer take effect the moment they are written, so you need a retrain and deploy cycle instead.

go deeper

for a junior

Know the basic shape: a lazy model stores examples and does its work when a prediction is asked for, so it needs memory for the data and gets slower as more of it accumulates.

for a middle

Be able to contrast where the compute sits in lazy versus eager methods and to state that a fitted fixed-size model scores at constant cost while a stored-example model's cost tracks how much data it has kept.

for a senior

Show operating judgment: measure p99 against store size, cap or reduce the store with a measured accuracy curve, and name what distillation into a fixed-size model costs you in freshness and per-segment accuracy.

for a principal

Own the framing that the growing store is a permanent per-request tax bought in exchange for freshness. Decide whether that property is worth it for this domain, set the SLO-linked cap policy, and define the review trigger when volume steps up.

## Where the cost lives A **lazy** or instance-based method does essentially no work when data arrives and all of its work when a prediction is requested: the stored examples *are* the model. An **eager** method pays once during fitting and produces a self-contained object that scores requests with a fixed amount of arithmetic. The consequence for a service is blunt: - **Memory** for a lazy model is the retained dataset, and it grows every day you keep writing to it. - **Latency** grows with the store too, because each request is answered by consulting stored data. - **Adding data is instant** — write a row and it is live. There is no fit step to run. That last property is the genuine attraction, and it is why teams choose these models. When 5 million new rows arrive there is nothing to retrain; the system simply knows more. The bill is that every one of those rows makes the store bigger and every subsequent query a little slower. An eager fixed-size model inverts all three: constant memory, constant per-request work, and new data that changes nothing until you refit and redeploy. ## Two shapes of the same constraint **A 10 ms bidding-style budget.** The service must return a score within a hard deadline or the response is worthless. Here the danger is not the average — it is the tail. Mean latency can sit comfortably at 3 ms while the p99 blows through 10 ms, because tail requests are the ones that touch cold memory, miss the cache, or land while the store no longer fits in RAM and starts paging. A service that misses its deadline on 1% of requests is failing 1% of the business, so **measure and alert on p99/p999 against store size, not on the mean**. **An 8 MB offline mobile budget.** The app must ship the model inside a fixed download size and score with no network. A fitted linear model with 40 coefficients is a few hundred bytes and trivially fits. Shipping 200,000 training rows so the device can look up neighbours does not fit, cannot be trimmed by a compression trick without changing what the model is, and gets worse with every data refresh. There is also a governance angle worth raising: shipping raw training rows to a device means shipping the underlying records themselves, which may be exactly what your privacy commitments forbid. A fixed-size fitted model is a lossy summary and does not carry the rows. ## The option list, roughly in order of how much you give up 1. **Bound the store.** Keep a recency window, or a stratified sample that preserves rare classes and rare regions. Measure accuracy as a function of store size: the curve usually plateaus far below "everything ever collected", and the plateau point is your cap. 2. **Reduce rather than truncate.** Replace many near-duplicate rows with a smaller set of representative prototypes or cluster summaries. You keep coverage of the input space at a fraction of the size — the wins here are usually larger than naive sampling because redundancy in dense regions is what dominates the store. 3. **Cache.** If the query distribution is skewed, caching scores for repeated or bucketed inputs removes most of the traffic from the expensive path. This helps hit rates, not worst-case latency, so it never rescues a p99 on its own. 4. **Index or approximate the lookup.** Specialised structures and approximate lookups reduce the constant and often the growth rate, but the fundamental coupling between store size and query cost remains — treat it as buying headroom, not as removing the problem. 5. **Distil into an eager model.** Train a fixed-size model on the same data (optionally on the lazy model's own outputs as targets, so it learns the behaviour you already validated). Scoring becomes constant-cost, memory becomes a known number, and the deployment story becomes ordinary. ## What you must say out loud when you distil Moving from lazy to eager is a real trade, not a free win: - **Freshness changes shape.** Instant incorporation of new rows is replaced by a retrain-and-deploy cadence. Decide that cadence deliberately and monitor for the drift it exposes. - **Local idiosyncrasy gets smoothed.** A stored-example model can serve a strange but real pocket of the input space perfectly by remembering it. A fixed-form model may average that pocket away. Check per-segment error, not just the global metric, before and after the switch. - **The failure mode moves.** A lazy system degrades gradually and visibly (latency creeps). An eager system fails quietly (a stale model keeps answering confidently after the world moves). Your monitoring has to change with the model. ## The decision framing that lands well Say explicitly what the store is buying you. If the value of instant data incorporation is high — a fast-moving domain where yesterday's model is genuinely wrong — then pay for the store, cap it, and engineer the tail. If the data distribution is stable, the freshness argument is weak and you are paying a permanent per-request tax for a property you do not use; distil and move on. Either way, tie the store cap to the latency SLO with a measured curve rather than a guess, and re-measure it when data volume steps up.

  • What do you give up by replacing the instance-based model with a fitted one?
    Instant incorporation of new data. With a store, writing a row makes it live immediately; with a fitted model you need a retrain and deploy cycle, so freshness becomes a cadence you must choose and monitor. You also risk smoothing away small but real pockets of the input space that the store served correctly by remembering them, so check per-segment error, not just the global metric.
  • How would you decide the cap on how many rows to keep?
    Measure two curves against store size: validation accuracy and p99 latency. Accuracy usually plateaus well before the store gets large, and latency rises steadily; the cap is the smallest store on the accuracy plateau that leaves headroom under the SLO. Re-measure after any step change in traffic or data volume, and prefer a stratified cap so rare classes are not thinned out.
  • Why can mean latency look healthy while the service still misses its budget?
    Because deadlines are violated in the tail. Typical requests hit warm memory and cached paths, while the slow ones touch cold data, miss the cache, or arrive once the store exceeds RAM and the system starts paging. A 3 ms mean with a 40 ms p99 fails 1% of requests, which for a bidding-style service means losing 1% of the business. Alert on p99 and p999.
  • When is keeping the growing store still the right call?
    When freshness is the product. In fast-moving domains where a model trained last week is measurably wrong today, instant incorporation of new rows is worth a real per-request tax, and the engineering job becomes bounding the store and defending the tail rather than eliminating the design. If the distribution is stable, you are paying that tax for a property you never use.

A lazy model is answering every question by searching the filing cabinet; an eager model wrote a one-page cheat sheet first. The cabinet is always up to date, but every question takes longer as it fills.

saying these in an interview costs you the question

  • Says instance-based models are cheap because training is free
  • Quotes mean latency instead of the p99 tail
  • Assumes storing every row forever is harmless
  • Thinks adding replicas fixes per-query cost that scales with data
  • Switches to a fitted model without mentioning the freshness it costs
  • Ignores that shipping raw rows to a device ships the records themselves

context