skip to content

questions

20

In a design round for checkout delivery-date estimates, what counts as the baseline a proposed model must beat?

level: juniorimportance: must knowfreq 62%

answer

  1. ask what decides the number today
  2. incumbent rule, not notebook comparator
  3. lane and service-level median table
  4. same parcels, same window
  5. lift is meaningless without its comparator

basics

~20 s

The baseline is whatever already decides the number in production - here a lane-and-service-level median transit table refreshed weekly from recent actuals - read on live traffic. A comparator invented for the write-up, such as a network-wide mean, is not the bar.

solid answer

~50 s

Something already prints an arrival window at checkout, so the baseline is that thing, not a placeholder. In this carrier it is a median transit table keyed by origin region, destination region and service level, computed over the last eight weeks of delivered parcels, refreshed weekly, with thin cells carrying forward last week's value. The bar is that table's accuracy on the same parcels, over the same window, against the same outcome the model would be scored on. Quoting a lift over a network-wide mean, or over a constant `3 days`, measures ground the rule already covers, so the number flatters the model. The first move in the round is to ask what produces the date today and how often it is right - a design that cannot state that number has not framed the problem yet.

go deeper

for a junior

Recall that a baseline is the estimator already producing the number in production, and that a comparison needs the same parcels and the same period on both sides.

for a middle

Explain how the incumbent is built - a median per lane and service level over a recent window, refreshed weekly, backing off or carrying forward where a cell is thin - and why that construction is already decent.

for a senior

Show the judgment of measuring the incumbent first, including its fallback cases, and of refusing a lift number whose comparator is unnamed. Say what margin over the rule would justify the build.

for a principal

Frame the bar as the thing that makes the whole design decidable: every later cost and constraint is argued against the margin over the incumbent, and a team that cannot state that number is not ready to commit engineers.

## What a baseline is in a design round A **baseline** is the estimator whose job the model is applying for. In delivery-date estimation that is not an abstraction: something already prints an arrival window on the checkout page, and whatever it is sets the bar. Before any boxes are drawn, ask *what produces that number right now, and how often is it right?* A design round that skips this question tends to end with a model that is worse than the rule it replaced and nobody noticed, because nobody wrote the rule's number down. Three families of heuristic baseline turn up, and which one you inherit decides how hard the bar is: - **A heuristic table** - a transit estimate keyed by a few attributes, computed from recent observed outcomes. - **A popularity or majority rule** - quote the network's most common transit time for that service level regardless of lane. Cheap, and stronger than teams expect wherever most volume moves the same way. - **A carry-forward rule** - this period's estimate is the last period's observed actual for the same key, held over when the key is too thin to recompute. Where week-to-week persistence is strong, this is usually the hardest of the three to beat. The carrier in this setting runs a mix of all three: a **median transit table** per (origin region, destination region, service level) over the last eight weeks of delivered parcels, refreshed weekly, backing off to a coarser key or carrying last week's value where a cell is thin. The promised window is that median plus a padding. That table, at its live accuracy, is the baseline. ## The comparator trap The common failure is quiet: the team compares against something that is not what ships. | comparator | where it comes from | what beating it proves | |---|---|---| | the incumbent table on live traffic | the system printing dates today | the model is worth deploying | | a network-wide mean transit time | one line in an analysis notebook | little - the table already beats it | | a figure published by another network | a write-up about a different population | nothing measurable here | | a constant such as three days | convenience | nothing; it is a straw man | A lift quoted against any row but the first is arithmetic about an estimator nobody would ship. It is also the reason a model can pass an internal review and then fail its launch read: the gap it closed was the gap between a straw man and the rule, not between the rule and the ceiling. ## Making the two numbers comparable Once you have named the incumbent, four things make its number and the model's number the same kind of number: 1. **Same parcels.** Score both on the identical population. If the model only covers lanes with enough history, either restrict the rule's read to those lanes or give the model a documented fallback to the rule elsewhere and score the combined system end to end. 2. **Same window.** Read both over the same calendar period. Transit times move with season, weather and network load, so a rule measured in a quiet month against a model measured in peak is not a comparison. 3. **Same outcome definition.** Both estimators are scored against the same delivery outcome, counted the same way, including how a parcel that was never delivered is treated. 4. **Same treatment of fallbacks.** The rule's thin cells fall back to a coarser key; those parcels still get a promise and those promises still count. Quietly dropping them lifts the rule's number, and dropping the model's low-confidence cases lifts the model's. ## Why the rule is usually stronger than it looks A heuristic that has been in production for years has absorbed corrections nobody wrote down: padding tuned by operations after a bad peak season, a service level that quietly quotes an extra day, cells overridden for lanes that cross a customs border. Those are real accuracy, and the model inherits none of them for free. It is also cheap in a sense the model is not - no labelled training set, no retraining, no owner on call for a strange estimate. Stating the bar honestly is therefore not modesty, it is the thing that makes the rest of the design decidable. Every later argument - what to log, what to serve, what the arrival window costs to compute - is measured against the margin over that number. ## What a good answer sounds like "Today the date comes from a median transit table per lane and service level, refreshed weekly from the last eight weeks of delivered parcels, carrying forward thin cells. It puts about four parcels in five inside the promised window. That is the bar; I want the model read on the same parcels over the same weeks, and I want to know what margin over it would justify the work." That is thirty seconds, and it reframes the round from *what model* to *what is worth building*.

  • The rule quotes a date for every parcel, but the model only covers lanes with enough history. How do you compare them?
    Score them on one population. Either restrict the rule's number to the lanes the model covers, or make the fallback to the rule part of the design and score the combined system on all parcels. What you must not do is read the rule on everything and the model on its easy subset.
  • What if there is no rule and a human planner quotes the date?
    The planner is the baseline. Sample their quotes on the same parcels and measure them the same way. A human incumbent usually scores better than teams expect, and it also carries a running cost the model would remove, which belongs in the comparison alongside the accuracy.
  • The incumbent table has no recorded accuracy. What do you do first?
    Measure it before building anything. Replay the promises it made against the delivery outcomes already stored, per lane and service level. That read is cheap, it is the bar every later decision leans on, and its per-lane spread tells you where any margin would have to come from.

The model is applying for a job that someone already holds. The interview is against the person currently doing the work, not against an empty chair.

saying these in an interview costs you the question

  • Comparing the model against a global average nobody would ship
  • Treating no model exists as no baseline exists
  • Reading the rule and the model on different weeks or different parcels
  • Assuming a hand-written rule is too crude to be a serious bar
  • Reporting lift without stating what the lift is over
open as a page

In a churn model for a subscription service, what must one row of the training table fix before any feature is picked?

level: juniorimportance: must knowfreq 66%

basics

~20 s

One row must fix three things: the prediction unit (one account in one billing month), the horizon the label covers (cancels within the next 30 days), and the as-of timestamp that separates feature territory from label territory.

open as a page

In a warehouse pick-path service, how would you split a 250 ms p99 budget across its stages before choosing a model?

level: middleimportance: must knowfreq 64%

basics

~20 s

Reserve first what you cannot spend — the caller's network round trip, fixed request handling and explicit headroom — then allocate the remainder across feature lookup, inference and post-processing. The inference slice becomes a ceiling any candidate model must fit.

open as a page

In an inpatient deterioration-risk service, which decisions do the offline sweep metric and the online launch metric each gate?

level: middleimportance: must knowfreq 72%

basics

~20 s

The offline sweep metric ranks candidate models during training - a ranking score over a frozen retrospective cohort. The online launch metric decides whether the service ships: what the deployed worklist changed in care, read on live shifts.

open as a page

Why ship the rule-based arrival window first and instrument it, rather than waiting until the model is ready?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Shipping the rule first delivers the product immediately and starts the clock on the evidence a model needs: every promise logged with the inputs as they stood at decision time, joined later to the actual delivery. Without that log, a model has no honest training set and no measured bar.

open as a page

A pick-path service is quiet mid-shift but every handheld requests at wave release — what must its latency clause state?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The arrival rate and the burst window the percentile must hold at, not just the percentile. A latency number without a load attached is met at the shift average and missed at the minute the business actually cares about.

open as a page

An inpatient deterioration model's offline AUC rose while escalations that changed care fell after launch - what divergences explain that?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Three divergences do it: the offline score integrates over a ranking the ward never reads past row twenty, the retrospective cohort is not the population the live path scores, and changed care counts action taken, not ranking accuracy.

open as a page

In a subscription churn table, why does a feature snapshotted at cancellation time rather than at the as-of stamp break the model live?

level: seniorimportance: must knowfreq 74%

basics

~20 s

Because the feature records a consequence of the cancellation, not a cause. Offline it separates the classes almost perfectly; at prediction time the account has not cancelled yet, so the field holds a neutral value and the model scores confidently on nothing.

open as a page

In a warehouse pick-path service, what does a 250 ms p99 budget cover beyond the model's inference time?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A p99 budget is measured end to end at the caller, so it also covers the handheld's network round trip, request handling, feature lookup, post-processing and serialisation. Inference is one line item inside that total, not the total itself.

open as a page

A new parcel lane carries about forty shipments a week - why is its own median transit time a weak baseline?

level: middleimportance: should knowfreq 45%

basics

~20 s

Forty parcels a week is too little evidence to key an estimate on: transit days are whole numbers, so a small shift in the week's mix flips which day sits in the middle and the quoted median hops. A support floor with a backoff to a coarser key fixes it.

open as a page

Why does a pick-path envelope state a freshness tolerance per input signal rather than one number for all?

level: middleimportance: should knowfreq 46%

basics

~20 s

Because the signals age at wildly different rates: slot inventory is wrong within seconds, aisle congestion within tens of seconds, item velocity within a day, bin geometry within a month. One shared number either breaks the fastest signal or over-builds for the slowest.

open as a page

In a subscription churn model, which event is the churn label when many accounts simply never renew?

level: middleimportance: should knowfreq 57%

basics

~20 s

There is no single event: an explicit cancellation is logged with a timestamp, but a silent non-renewal is an absence. The label must be a predicate over billing state, declared positive once the renewal date plus a grace period has passed with no new paid period.

open as a page

With a 30-day churn horizon plus a 14-day grace period, which subscription rows cannot be labelled yet?

level: middleimportance: should knowfreq 49%

basics

~20 s

Every row whose as-of timestamp is newer than 44 days ago. The horizon plus the grace period is the label maturity lag, so the training cut sits 44 days in the past and the most recent six weeks of data carries features but no trustworthy outcome.

open as a page

Two deterioration rankers: the candidate wins on pooled weekly precision but loses in both wards - which reading do you trust?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Neither reading is wrong; they answer different questions. The reversal comes from the candidate moving worklist rows toward the higher-prevalence ward, so the honest read is whichever matches how worklist capacity is actually allocated - and both get reported.

open as a page

The transit rule already puts 82% of parcels inside the promised window - what would make you decline to replace it with a model?

level: principalimportance: should knowfreq 42%

basics

~20 s

Decline when the reachable margin is small against the upkeep a learned estimator adds permanently: a maintained training set, a refit rota, monitoring, and someone on call to explain a bad date. State the margin that would change the answer before the offline sweep, not after.

open as a page

Your deterioration service's online read needs a quarter to accumulate while the offline sweep reports hourly - how much launch weight should the offline number carry?

level: principalimportance: should knowfreq 42%

basics

~20 s

As much weight as its track record earns. Pair each past launch's offline delta with the online delta that followed; that history says whether the proxy can license a launch or only veto a candidate, and what a launch carried by it must commit to.

open as a page

A carrier consolidates its hubs this month - why can a recent-window transit rule beat a model trained on a year of parcels?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A hub consolidation is a structural break: most of the model's training rows describe routings that no longer happen, while an eight-week rolling median only remembers eight weeks and re-converges as post-change parcels fill its window. Short memory wins across a break.

open as a page

Before a pick-path model is chosen, where does the ceiling on cost per prediction come from?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

From the value side: what one prediction changes, net of the rule already in place. For pick-path routing that is picker seconds saved, times the fraction of picks whose order the model actually alters, valued at the loaded labour rate.

open as a page

Why is team-draft interleaving unavailable as a cheap online read for a deterioration worklist shared by a ward team?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Team-draft interleaving needs many independent impressions and an action attributable to one ranker's item. A ward worklist is one shared list per team per shift, worked top to bottom, and each action changes the patient - so neither condition holds.

open as a page

What does a subscription service give up by training on a usage-decline proxy instead of the cancellation outcome?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

It gives up the guarantee that the target is the business outcome. The proxy matures in weeks and fires while an account is still winnable, but it flags seasonal dips, misses steady users who leave over price, and moves whenever the product changes.

open as a page