skip to content

Should you predict delivery ETA in minutes and threshold it, or classify 'late by 30+ minutes' directly?

level: seniorimportance: should knowfreq 52%

answer

  1. ask what the decision actually consumes
  2. will the cutoff ever move?
  3. average error is not decision accuracy
  4. a point estimate carries no probability
  5. predict the distribution, read the tail off

basics

~20 s

Model the number when you need the ETA itself or the cutoff may move; classify the flag directly when the cutoff is fixed and only the decision matters. A regressor optimises accuracy everywhere, not around the 30-minute boundary.

solid answer

~50 s

The two framings optimise different things. A regressor on minutes learns from the full target - a 31-minute delay and a three-hour delay are different examples - serves any cutoff you later choose, and gives you the ETA that customers actually want to see. But its objective spends capacity minimising error across the whole range, including extreme delays that no decision depends on, and a point prediction of 28 minutes is not a probability of being late. A direct classifier optimises exactly the decision you make, outputs a probability of lateness, and is unbothered by extreme values - at the cost of discarding magnitude, hard-coding 30 minutes into the label, and needing a relabel and retrain if the promise changes. If I need both, I model the conditional distribution of delivery time - for example by predicting several quantiles - and read the chance of exceeding 30 minutes off it.

go deeper

for a junior

Be ready to describe both routes plainly: predict the number then compare it to 30, or label history as late/not and classify. Knowing that the number can be re-thresholded later while the label cannot is the key takeaway.

for a middle

Explain where each training objective spends its effort - squared error over the whole range versus separating two classes at one boundary - and why a point prediction cannot supply a probability of exceeding the cutoff.

for a senior

Demonstrate the diagnosis: check the density of true durations near the cutoff before believing an aggregate error figure, and check that the label's reference point is not another model's output. Offer the quantile or distributional framing as the compromise.

for a principal

Own the systems tradeoff: one distributional model serving many cutoffs versus several purpose-built classifiers, and who pays for the extra complexity. Be ready to argue about label definitions that create feedback loops with the promises the business makes.

## The two framings You have historical deliveries with their actual duration in minutes, and a decision to make: warn the customer when a delivery will be more than 30 minutes late. **Framing A - regress, then threshold.** Train a model to predict the duration, then apply `predicted > 30` to get the flag. **Framing B - classify directly.** Derive a binary label from history (`was it over 30 minutes?`) and train a classifier on that label. Both are defensible. The interview is about knowing what each buys and what each costs. ## What regression buys you **Full-information target.** Every example carries a real number, not one bit. A delivery at 31 minutes and one at 180 minutes are the same class under Framing B, but different training signal under Framing A. That extra information typically makes the model more data-efficient, which matters most when late deliveries are rare. **One model, many cutoffs.** Operations wants 30 minutes, the customer-refunds team wants 45, and next quarter marketing promises 20. A regressor serves all three without retraining; a classifier trained on a 30-minute label serves exactly one. **The number is itself the product.** If the app displays an ETA, you need the regression model regardless, and a separate classifier is a second model to train, deploy and monitor. **Ranking.** Predicted duration orders deliveries by expected lateness, which is enough for `intervene on the worst 200 orders this hour` even without any cutoff. ## What regression costs you **The objective is misaligned with the decision.** Squared error is minimised by getting the whole range right, so the fit is dragged by the long tail of 3-hour disasters, while the decision depends entirely on the narrow band of deliveries near 30 minutes. The model spends capacity where the decision does not care, and can be beaten near the boundary by a classifier that only ever cared about that region. **Average error does not translate into decision accuracy.** An average error of six minutes sounds excellent, but the flag's accuracy depends on how densely true durations pile up just either side of 30 minutes. If most deliveries land at 25-35 minutes, six minutes of error is catastrophic for the flag; if durations are bimodal at 15 and 60, the same model flags almost perfectly. **A point prediction is not a probability.** Predicting 28 minutes tells you nothing about the chance of exceeding 30, so you cannot weigh the cost of an unnecessary warning against the cost of a missed one. Thresholding the point prediction implicitly asserts a probability of one on the wrong side of the cutoff. **Symmetric error is the wrong assumption near a hard boundary.** Under-predicting by five minutes and over-predicting by five minutes are equally penalised by squared error, but around a promise they have very different business consequences. ## What direct classification buys you The objective *is* the decision, so the model concentrates capacity on separating late from on-time. The output is a probability of lateness, which is exactly the input a cost-weighted decision needs. Extreme durations cannot dominate the fit, because a 3-hour delivery is just another positive. And you can use features that only relate to lateness - a road closure flag, a store's backlog - without worrying that they distort the ETA estimate. ## What direct classification costs you **Magnitude is discarded.** A 31-minute and a 300-minute delay are identical labels, so the model cannot learn that the second is worse or prioritise it. **The cutoff is baked into the label.** Changing the promise to 20 minutes means redefining the target and retraining - and every historical evaluation you have done is about a threshold that no longer exists. **You may manufacture an imbalance problem.** If only 3% of deliveries are late, you turn a well-behaved continuous target into a rare-positive classification task with all the difficulties that brings. ## The framing that gets both Model the **conditional distribution** of duration rather than its mean. Predict several quantiles - say the 50th, 80th, 90th and 95th - or estimate the spread of the residuals and integrate. Then the probability of exceeding 30 minutes is read off the predicted distribution, and moving the promise to 20 minutes is a lookup rather than a retrain. You keep the displayed ETA, you keep the ability to change the cutoff, and you gain the probability that a cost-based decision requires. The price is a heavier model and a harder evaluation story. ## Two traps in defining the label at all **Leakage through the label definition.** `Late by more than 30 minutes` is usually relative to a promised time - and the promise is often produced by an earlier model. Your label then depends on a system output, so improving the promise silently changes the label distribution, and features derived from the promise leak. **Availability at prediction time.** Every feature must be knowable when the prediction is made, at order placement. Courier assignment or pickup delay may not exist yet, and using them makes an offline model look brilliant and a live one useless. ## The rule of thumb Regress when the number itself is consumed, when the cutoff may move, or when several consumers want different cutoffs. Classify when the cutoff is stable, the only consumer is one decision, and magnitude beyond the cutoff is irrelevant. When you need a defensible probability at a moving cutoff, model the distribution.

  • The promise changes from 30 minutes to 20 next quarter - which framing survives?
    The regression framing, and any distributional version of it: the target is unchanged, so you re-apply the new cutoff and carry on. A classifier trained on a 30-minute label must be relabelled and retrained, and every historical evaluation refers to a threshold that no longer exists. If cutoffs move regularly, that alone usually decides the design.
  • How would you get a probability of lateness out of a model that predicts delivery minutes?
    Model the spread, not just the centre. Fit several quantiles of the duration and interpolate the chance of exceeding the cutoff, or estimate the residual distribution conditional on the features and integrate its tail. What you must not do is treat the point prediction as certainty, or convert distance from the cutoff into a probability by an arbitrary transformation.
  • Your ETA model has an average error of about six minutes - is that enough to know the late flag will be accurate?
    No. Flag accuracy depends on how many true durations sit within a few minutes of the cutoff. Six minutes of error is ruinous if most deliveries cluster at 25 to 35 minutes, and nearly harmless if durations are bimodal far from 30. Always look at the distribution of true values around the boundary before trusting an aggregate error number.
  • What is the risk if 'late' is defined relative to a promised time that another model produced?
    Your label becomes a function of another system's output, so it shifts whenever that system is retrained, and any feature derived from the promise partially encodes the label. It also creates a feedback loop: a more conservative promise mechanically reduces lateness without any real improvement. Prefer a label anchored to something the business controls independently, and record the promise version alongside every row.

A regressor is a thermometer and a classifier is a fire alarm. The thermometer serves any threshold you later care about; the alarm is tuned to exactly one and tells you nothing about how hot it actually is.

saying these in an interview costs you the question

  • Assumes a low average error automatically yields an accurate late flag
  • Treats a point ETA prediction as a probability of being late
  • Bakes a business cutoff into the label without asking if it moves
  • Ignores that squared error is dominated by extreme delays
  • Builds a classifier on features unavailable when the order is placed

context