skip to content

The transit rule already puts 82% of parcels inside the promised window - what would make you decline to replace it with a model?

level: principalimportance: should knowfreq 42%

answer

  1. compare margin against permanent upkeep
  2. bound the reachable band first
  3. some misses are unreachable at dispatch
  4. refit rota, monitoring, a named owner
  5. state the threshold before the sweep

basics

~20 s

Decline when the reachable margin is small against the upkeep a learned estimator adds permanently: a maintained training set, a refit rota, monitoring, and someone on call to explain a bad date. State the margin that would change the answer before the offline sweep, not after.

solid answer

~50 s

Two numbers decide it. The first is **headroom**: a share of misses comes from weather, customs holds and failed first attempts, which nothing logged at dispatch predicts, so the honest question is how far above 82% any estimator on these signals could get, and how much of that gap the model actually closes. The second is **upkeep**: the rule needs a weekly refresh and nothing else, while the model adds a labelled training set that must keep being assembled point-in-time, a refit rota, a monitoring surface, a rollback path and a named owner who can answer for a strange date at midnight. At two million parcels a week, moving 82% to 87% avoids 100,000 misses a week and 5.2 million a year - worth owning if misses are expensive, easy to decline if the extra points change nothing downstream. Where the gap is thin, sharpening the rule is usually the better trade.

code

pseudocode · 10 lines
pseudocode
weeklyParcels  = 2000000
ruleMissRate   = 0.18          // rule holds 82% inside the window
modelMissRate  = 0.13          // expected reach, read off logged residuals

weeklyMissesAvoided = weeklyParcels * (ruleMissRate - modelMissRate)   // 100000
yearlyMissesAvoided = weeklyMissesAvoided * 52                         // 5200000
yearlyContacts      = yearlyMissesAvoided * contactsPerMiss            // 260000 at 1 in 20

build = value(yearlyContacts, goodwill) >
        buildCost + yearlyUpkeep(trainingSet, refitRota, monitoring, onCall)

go deeper

for a junior

Recall that a model is not free after it is built: it needs data, periodic refreshes and someone responsible for it, and that cost belongs in the decision to build it.

for a middle

Explain the two sides of the comparison - the reachable gain over the current rule, and the recurring obligations a learned estimator adds that a weekly table recompute does not.

for a senior

Show that you bound the reachable band from the rule's own residuals rather than assuming the ceiling, and that you keep the heuristic as the fallback even after a model ships.

for a principal

Own the decision: state the margin that would justify the build before the evidence arrives, name who carries the rota and the pager, and be willing to sharpen the rule and decline the model.

## The shape of the decision This is the judgment the category exists for, and it is not "can we build it". It is: *does the margin learning adds exceed what it costs to own, permanently?* Both sides of that comparison have to be numbers before the sweep starts, because after the sweep the decision gets made by whichever result looks nicest. ## Headroom: what is actually reachable The rule puts 82% of parcels inside the window, so 18% miss. That 18% is not one thing: - **Structured error** - lanes systematically under-padded, dispatch after a cut-off, days when the sorting hub is backed up. These are visible in the signals already logged and are what a learned estimator claims. - **Error from signals nobody captures** - a customs hold, a road closure, a recipient who is not home for a first attempt. No estimator built on what is known at dispatch removes this share, though a better system might reduce it by capturing new signals rather than by learning harder on the old ones. So the first task is to bound the reachable band, not to assume the ceiling is 100%. The rule's own logged residuals do this: the share of error that is structured by lane, hour and backlog is roughly what is on offer. If a careful read says a good estimator lands near 87%, the decision is about five points, not eighteen. ## Upkeep: what the model adds forever | obligation | heuristic table | learned estimator | |---|---|---| | periodic refresh | a weekly recompute over a window | a refit run, with the data assembled point-in-time | | training data | none | a maintained labelled set, joined on matured outcomes | | correctness surface | support and value age per cell | input and prediction monitoring, plus outcome quality once labels mature | | failure handling | fall back down the ladder | a rollback path to a prior version, and a fallback estimator anyway | | people | the team that owns the table | a named owner who can explain a bad date and act on an alarm | | explaining one estimate | read the cell and the rung | reconstruct inputs, version and score | The last two rows are the ones that get waved away in design rounds and dominate the real cost. The rule can be explained by a support agent reading a table. A learned estimate cannot, which means the organisation acquires a standing dependency on the few people who can interpret it. Notice also that the fallback row does not disappear: a model still needs the rule underneath it for lanes it cannot cover and for the minutes when scoring is unavailable. Building the model does not retire the heuristic; it adds to it. ## Sizing the margin At two million parcels a week, 18% is 360,000 misses a week and 13% is 260,000, so five points is 100,000 fewer misses a week, or 5.2 million a year. If roughly one miss in twenty produces a support contact, that is about 260,000 fewer contacts a year. Now the decision is arguable rather than aesthetic: that volume of avoided misses either pays for a permanent team obligation or it does not, and the answer differs between a network at this scale and one a hundredth of the size doing the identical work. ## Four conditions that argue for declining 1. **Thin headroom.** The structured share of the residuals is small, so the reachable gain is a point or two. 2. **Shifting ground.** The network reorganises often enough that a fitted estimator spends much of its life describing a layout that has changed. 3. **Nothing downstream changes.** If the promise is padded to a whole day and displayed as a range, an estimator that is half a day sharper may not move a single customer-visible number. 4. **No owner.** If no team will hold the refit rota, the monitoring and the pager, the model will be built, will decay quietly and will end up worse than the rule it replaced. ## Sharpening the rule instead The alternative that gets skipped is improving the heuristic, which usually buys part of the same margin at a fraction of the upkeep: finer cells where support allows, a padding chosen per lane from the spread of observed transit times rather than one global constant, a backlog term on the hub, a separate treatment of the days around a peak season. These stay explainable, need no labelled training set and no refit rota, and they raise the bar the model would have to clear - which is exactly the point. If the rule can be pushed most of the way, the case for learning gets weaker, and that is a good outcome, not a failed project. ## How to say it in the round Commit to the threshold before the evidence arrives: "I would build it if the logged residuals say a model reaches at least four points above the rule on the lanes carrying most volume, and if a team signs up for the refit rota and the pager. Below that I would sharpen the padding per lane and keep the table." That is the answer a lead is expected to give - a stated bar, the upkeep named honestly, and a willingness to conclude that the model is not worth building.

  • The offline read comes back at three points above the rule, below your stated bar. What now?
    Hold the bar, and say why in writing. Then ask whether the gap is thin because the signals are weak rather than the estimator: if a new input would plausibly move it, plumb that input into the rule's log and revisit. Moving the threshold after seeing the number is how upkeep gets acquired by accident.
  • Does building the model let you retire the heuristic?
    No. It still answers for lanes with too little history, for the minutes when scoring is unavailable, and as the thing a support agent can read. Budget for maintaining both, which is part of why the margin has to be worth it.
  • How do you bound the reachable share without building anything?
    Read the rule's logged residuals. Group them by lane, dispatch hour and backlog and see how much of the error those groups explain; what remains unexplained by every signal you actually capture is what no estimator on those signals will reach. It is a rough bound, and rough is enough to decide.

saying these in an interview costs you the question

  • Comparing build cost only, ignoring permanent upkeep
  • Assuming any remaining error is reachable by a better model
  • Expecting the heuristic to be retired once a model ships
  • Choosing the decision threshold after seeing the offline number
  • Building without a team that owns the refits and the pager