skip to content

questions

5

In a design round for checkout delivery-date estimates, what counts as the baseline a proposed model must beat?

level: juniorimportance: must knowfreq 62%

answer

  1. ask what decides the number today
  2. incumbent rule, not notebook comparator
  3. lane and service-level median table
  4. same parcels, same window
  5. lift is meaningless without its comparator

basics

~20 s

The baseline is whatever already decides the number in production - here a lane-and-service-level median transit table refreshed weekly from recent actuals - read on live traffic. A comparator invented for the write-up, such as a network-wide mean, is not the bar.

solid answer

~50 s

Something already prints an arrival window at checkout, so the baseline is that thing, not a placeholder. In this carrier it is a median transit table keyed by origin region, destination region and service level, computed over the last eight weeks of delivered parcels, refreshed weekly, with thin cells carrying forward last week's value. The bar is that table's accuracy on the same parcels, over the same window, against the same outcome the model would be scored on. Quoting a lift over a network-wide mean, or over a constant `3 days`, measures ground the rule already covers, so the number flatters the model. The first move in the round is to ask what produces the date today and how often it is right - a design that cannot state that number has not framed the problem yet.

go deeper

for a junior

Recall that a baseline is the estimator already producing the number in production, and that a comparison needs the same parcels and the same period on both sides.

for a middle

Explain how the incumbent is built - a median per lane and service level over a recent window, refreshed weekly, backing off or carrying forward where a cell is thin - and why that construction is already decent.

for a senior

Show the judgment of measuring the incumbent first, including its fallback cases, and of refusing a lift number whose comparator is unnamed. Say what margin over the rule would justify the build.

for a principal

Frame the bar as the thing that makes the whole design decidable: every later cost and constraint is argued against the margin over the incumbent, and a team that cannot state that number is not ready to commit engineers.

## What a baseline is in a design round A **baseline** is the estimator whose job the model is applying for. In delivery-date estimation that is not an abstraction: something already prints an arrival window on the checkout page, and whatever it is sets the bar. Before any boxes are drawn, ask *what produces that number right now, and how often is it right?* A design round that skips this question tends to end with a model that is worse than the rule it replaced and nobody noticed, because nobody wrote the rule's number down. Three families of heuristic baseline turn up, and which one you inherit decides how hard the bar is: - **A heuristic table** - a transit estimate keyed by a few attributes, computed from recent observed outcomes. - **A popularity or majority rule** - quote the network's most common transit time for that service level regardless of lane. Cheap, and stronger than teams expect wherever most volume moves the same way. - **A carry-forward rule** - this period's estimate is the last period's observed actual for the same key, held over when the key is too thin to recompute. Where week-to-week persistence is strong, this is usually the hardest of the three to beat. The carrier in this setting runs a mix of all three: a **median transit table** per (origin region, destination region, service level) over the last eight weeks of delivered parcels, refreshed weekly, backing off to a coarser key or carrying last week's value where a cell is thin. The promised window is that median plus a padding. That table, at its live accuracy, is the baseline. ## The comparator trap The common failure is quiet: the team compares against something that is not what ships. | comparator | where it comes from | what beating it proves | |---|---|---| | the incumbent table on live traffic | the system printing dates today | the model is worth deploying | | a network-wide mean transit time | one line in an analysis notebook | little - the table already beats it | | a figure published by another network | a write-up about a different population | nothing measurable here | | a constant such as three days | convenience | nothing; it is a straw man | A lift quoted against any row but the first is arithmetic about an estimator nobody would ship. It is also the reason a model can pass an internal review and then fail its launch read: the gap it closed was the gap between a straw man and the rule, not between the rule and the ceiling. ## Making the two numbers comparable Once you have named the incumbent, four things make its number and the model's number the same kind of number: 1. **Same parcels.** Score both on the identical population. If the model only covers lanes with enough history, either restrict the rule's read to those lanes or give the model a documented fallback to the rule elsewhere and score the combined system end to end. 2. **Same window.** Read both over the same calendar period. Transit times move with season, weather and network load, so a rule measured in a quiet month against a model measured in peak is not a comparison. 3. **Same outcome definition.** Both estimators are scored against the same delivery outcome, counted the same way, including how a parcel that was never delivered is treated. 4. **Same treatment of fallbacks.** The rule's thin cells fall back to a coarser key; those parcels still get a promise and those promises still count. Quietly dropping them lifts the rule's number, and dropping the model's low-confidence cases lifts the model's. ## Why the rule is usually stronger than it looks A heuristic that has been in production for years has absorbed corrections nobody wrote down: padding tuned by operations after a bad peak season, a service level that quietly quotes an extra day, cells overridden for lanes that cross a customs border. Those are real accuracy, and the model inherits none of them for free. It is also cheap in a sense the model is not - no labelled training set, no retraining, no owner on call for a strange estimate. Stating the bar honestly is therefore not modesty, it is the thing that makes the rest of the design decidable. Every later argument - what to log, what to serve, what the arrival window costs to compute - is measured against the margin over that number. ## What a good answer sounds like "Today the date comes from a median transit table per lane and service level, refreshed weekly from the last eight weeks of delivered parcels, carrying forward thin cells. It puts about four parcels in five inside the promised window. That is the bar; I want the model read on the same parcels over the same weeks, and I want to know what margin over it would justify the work." That is thirty seconds, and it reframes the round from *what model* to *what is worth building*.

  • The rule quotes a date for every parcel, but the model only covers lanes with enough history. How do you compare them?
    Score them on one population. Either restrict the rule's number to the lanes the model covers, or make the fallback to the rule part of the design and score the combined system on all parcels. What you must not do is read the rule on everything and the model on its easy subset.
  • What if there is no rule and a human planner quotes the date?
    The planner is the baseline. Sample their quotes on the same parcels and measure them the same way. A human incumbent usually scores better than teams expect, and it also carries a running cost the model would remove, which belongs in the comparison alongside the accuracy.
  • The incumbent table has no recorded accuracy. What do you do first?
    Measure it before building anything. Replay the promises it made against the delivery outcomes already stored, per lane and service level. That read is cheap, it is the bar every later decision leans on, and its per-lane spread tells you where any margin would have to come from.

The model is applying for a job that someone already holds. The interview is against the person currently doing the work, not against an empty chair.

saying these in an interview costs you the question

  • Comparing the model against a global average nobody would ship
  • Treating no model exists as no baseline exists
  • Reading the rule and the model on different weeks or different parcels
  • Assuming a hand-written rule is too crude to be a serious bar
  • Reporting lift without stating what the lift is over
open as a page

Why ship the rule-based arrival window first and instrument it, rather than waiting until the model is ready?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Shipping the rule first delivers the product immediately and starts the clock on the evidence a model needs: every promise logged with the inputs as they stood at decision time, joined later to the actual delivery. Without that log, a model has no honest training set and no measured bar.

open as a page

A new parcel lane carries about forty shipments a week - why is its own median transit time a weak baseline?

level: middleimportance: should knowfreq 45%

basics

~20 s

Forty parcels a week is too little evidence to key an estimate on: transit days are whole numbers, so a small shift in the week's mix flips which day sits in the middle and the quoted median hops. A support floor with a backoff to a coarser key fixes it.

open as a page

The transit rule already puts 82% of parcels inside the promised window - what would make you decline to replace it with a model?

level: principalimportance: should knowfreq 42%

basics

~20 s

Decline when the reachable margin is small against the upkeep a learned estimator adds permanently: a maintained training set, a refit rota, monitoring, and someone on call to explain a bad date. State the margin that would change the answer before the offline sweep, not after.

open as a page

A carrier consolidates its hubs this month - why can a recent-window transit rule beat a model trained on a year of parcels?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A hub consolidation is a structural break: most of the model's training rows describe routings that no longer happen, while an eight-week rolling median only remembers eight weeks and re-converges as post-change parcels fill its window. Short memory wins across a break.

open as a page