skip to content

A manager treats your model's point prediction of 42 minutes as a promise — how do you respond?

level: seniorimportance: should knowfreq 40%

answer

  1. a point prediction is a centre
  2. half of cases land above it
  3. one case or many cases
  4. promise a bound, not the average
  5. name the assumptions with the number

basics

~20 s

Replace the single number with an interval and say what it covers: the best estimate is 42 minutes, and about 95 of 100 individual orders like this land between 27 and 57. Then agree which number the promise uses.

solid answer

~50 s

A point prediction is the centre of a distribution, not a floor or a ceiling, and roughly half of individual cases land above it by construction. So the answer is to give the manager the shape, not just the middle: best estimate 42 minutes; for a single order, a 95% prediction interval of roughly 27 to 57 minutes; for the *average* over many orders like it, a much tighter range of about 40 to 44. Which one they need depends on the decision — staffing a day of similar orders is an averaging question, promising one customer is a single-case question. If they want a commitment, quote an upper bound rather than the centre, so the promise is one the model expects to keep most of the time. State the conditions too: the interval assumes the model form is right, residual spread is constant, and this order is inside the fitted range.

go deeper

for a junior

Remember that a predicted value is an average, not a maximum, and that quoting it alone hides how much individual cases vary. Always attach a range when someone plans around the number.

for a middle

Be able to produce all three quantities from one fitted model — the point estimate, the interval for the average case, and the interval for a single new case — and say which question each answers.

for a senior

Demonstrate that you pick the object from the decision, convert a promise into a bound with a stated keep rate, and name the assumptions behind the interval before anyone asks.

for a principal

Own how uncertainty is reported across the organisation: what model outputs carry by default, who chooses the risk level behind a commitment, and how to resist pressure to publish tighter-looking numbers that answer an easier question.

## Why the point number is the wrong object to hand over A regression's fitted value at a set of inputs is an estimate of the *mean* outcome for cases like that. It is the centre of a distribution of plausible outcomes. If the model is well calibrated, about half of individual cases will come in above it. A promise built on the centre is therefore a promise you expect to break roughly half the time — which is not a modelling failure, it is what a mean is. The manager is not making a statistical error out of carelessness; a single number simply carries no visible uncertainty. The fix is to change what you hand over. ## The three numbers, and what each is for From the same fitted model at the same inputs you can produce three different statements: 1. **Point estimate — 42 minutes.** The model's best single guess for the average case. Useful for arithmetic that will be aggregated anyway. 2. **Interval for the mean response — say 40 to 44 minutes.** Where the true average sits for all orders with these characteristics. Narrow, because averaging cancels individual noise. 3. **Prediction interval for one order — say 27 to 57 minutes.** Where the *next single* order will land. Wide, because it must carry that order's own variability as well. The most common failure in this conversation is quoting number 2 for a number 3 question. The narrow band looks better and is what most default model summaries emphasise, but promising 40 to 44 minutes to one customer would be wrong most of the time. ## Matching the number to the decision Ask what the figure will be used for before choosing which one to say. - **Capacity and staffing for a day of similar orders.** This is about the average across many cases, so the mean-response interval is the right object, and the individual variability largely cancels. - **A commitment to one customer, an alert threshold on one transaction, an SLA on one request.** These are single-case decisions. The prediction interval is the right object, and the centre is close to useless as a promise. - **A financial plan built from many predictions.** Aggregate, so again about means — but note that if errors are correlated across cases, that cancellation is weaker than it looks and the aggregate uncertainty is larger than a naive calculation suggests. ## How to turn a prediction into an honest commitment If the organisation needs one number to promise, do not give them the centre. Give them a bound with a stated success rate: "the model expects about 90% of orders like this to arrive within 55 minutes." That converts a statistical object into an operational one that can actually be held to, and it makes the tradeoff explicit — a tighter promise means a lower keep rate, and that is a business choice rather than a modelling choice. A useful move is to lay out the tradeoff as a small table of promise thresholds against expected keep rates and let the decision-maker pick. You have then given them the shape of the distribution in a form they can act on, and the choice of risk level sits with the person who owns the consequence. ## Stating the conditions Every interval you quote comes with assumptions, and a senior answer names them without being asked: - **The model form is right for this region.** If the order's characteristics sit outside the range the model was fitted on, the interval understates the risk, because it prices sampling error only. - **The residual spread is roughly constant.** If long-distance orders scatter far more than short ones, one global interval is too wide at one end and too narrow at the other, and the promise fails asymmetrically. - **The residuals are not badly skewed.** Delivery times usually have a long right tail; a symmetric interval will then understate how bad the worst cases get, which is exactly the direction that hurts a promise. - **Cases are independent.** If a single incident delays many orders at once, per-order intervals do not describe the aggregate risk on a bad day. ## When the response is "that interval is too wide to be useful" This is the real test of the conversation. The width is a finding, not a presentation problem. Three legitimate responses, and one illegitimate one: - Reduce it honestly — better predictors, more data at the relevant operating point, or a segment-specific model where the residual spread is smaller. - Change the decision so it tolerates the uncertainty — a buffer, a range-based promise, a fallback when the actual outcome runs long. - Accept it and price the risk explicitly, choosing the keep rate deliberately. - What you do not do is switch to the narrower mean-response interval because it looks better. That is not a tighter estimate; it is an answer to a different question, and the gap shows up later as broken commitments rather than as visible statistical error. ## The compact version to say out loud "42 is the centre, so about half of orders will run longer. For one order the model puts 95% of outcomes between roughly 27 and 57 minutes. If we need something to promise, I would quote a bound with a stated keep rate rather than the average — and all of this assumes this order looks like the ones the model was fitted on."

  • Staffing for tomorrow's volume of similar orders — which interval do you quote?
    The interval for the mean response, because staffing is driven by the average across many orders and individual variation largely cancels. It is much narrower than the single-order interval and is the correct object here. The one caveat is correlated shocks, such as a weather event, which break the cancellation and are not captured by that interval.
  • The manager says the interval is too wide to plan around. What do you do?
    Treat the width as information rather than a presentation problem. Offer real options: better predictors or more data at that operating point, a segment-specific model, a buffer in the decision, or an explicitly chosen keep rate. What I would not do is quote the mean-response interval instead, since that answers a different question and hides the risk rather than reducing it.
  • What would you check before quoting any interval for this order at all?
    That the order's characteristics fall inside the range the model was fitted on, that residual spread does not grow with the fitted value, and that the residuals are not badly skewed. Delivery times often have a long right tail, and a symmetric interval then understates the worst cases — the direction that matters most for a promise.

saying these in an interview costs you the question

  • Quotes the mean-response interval as a per-order guarantee
  • Presents an interval without saying whether it covers one case or an average
  • Narrows the interval to make the number more palatable
  • Repeats the point prediction with no uncertainty attached
  • Treats interval width as a presentation problem rather than a finding

context