skip to content

Before a pick-path model is chosen, where does the ceiling on cost per prediction come from?

level: seniorimportance: nice to knowfreq 34%

answer

  1. derive it, do not price it
  2. value of what one prediction changes
  3. net of the incumbent rule
  4. multiply by how often it changes anything
  5. say what one prediction is

basics

~20 s

From the value side: what one prediction changes, net of the rule already in place. For pick-path routing that is picker seconds saved, times the fraction of picks whose order the model actually alters, valued at the loaded labour rate.

solid answer

~40 s

The ceiling is derived, not negotiated. Suppose better routing saves about 4 seconds of walking on a pick whose order it changes, and that the incumbent rule already produces the same order for four picks in five. The expected saving is then 0.8 seconds per prediction. At a loaded labour cost of roughly `$21.60` per hour — `$0.006` per second — one prediction is worth about `$0.0048`. If the design wants the capability to return several times its running cost, the ceiling lands near a tenth of a cent. That number is written into the envelope before any model exists, because it rules out whole classes of design on arithmetic rather than on taste, and no accuracy improvement rescues a prediction that costs more than it saves.

go deeper

for a junior

Recall that a prediction is only worth what it changes, and that a system costing more per prediction than it saves cannot be fixed by making the predictions better.

for a middle

Walk the derivation end to end: saving when it applies, how often it applies, the rate that converts it to money, and the unit the resulting ceiling is stated per.

for a senior

Insist on counting value net of the incumbent rule and on naming the prediction unit, since both silently change the ceiling by an order of magnitude or more.

for a principal

Use the total — value per prediction times volume — to ask early whether the capability justifies the system that produces it, while the answer can still change the plan.

## The ceiling comes from the value side, not the bill A cost clause written at framing time answers one question: **what is one prediction allowed to cost?** The tempting way to answer it is to price the serving path and call that the number. That is backwards — it measures what a design happens to cost rather than what the problem can afford, and it has no way of saying no. The defensible derivation runs the other way, from the behaviour a prediction changes: 1. **What does one prediction change?** Here, the order in which the remaining picks of a wave are walked. 2. **What is that change worth when it happens?** Say 4 seconds of walking saved on a pick whose order it alters. 3. **How often does it actually change anything?** If the incumbent rule already yields the same order four times in five, only one pick in five is affected. 4. **What is a second worth?** At a loaded labour cost of about `$21.60` per hour, one second is `$0.006`. ## A worked number | step | value | |---|---| | saving when the order changes | 4 seconds | | fraction of picks actually re-ordered | 1 in 5 | | expected saving per prediction | 0.8 seconds | | loaded labour cost per second | `$0.006` | | value of one prediction | about `$0.0048` | | value at 10,000 picks per day | about `$48` per day | The last row is the sentence that changes a design conversation. A capability worth tens of dollars a day cannot carry a per-prediction cost in the cents, however good its scores are — and every architecture that would imply such a cost is ruled out before anybody trains anything. The ceiling itself is then set below the value with margin: if the design wants the capability to return, say, five times what it costs to run, the ceiling is about `$0.001` per prediction. ## Three things that move the ceiling, and one that does not - **The prediction unit.** The ceiling is per prediction, so the envelope must say what one is. Scoring 20 candidate bins as 20 separate predictions rather than producing one ordering per pick divides a `$0.001` ceiling into `$0.00005` per call. Two designs can have identical monthly bills and completely different per-prediction economics purely from how the unit is defined. - **The baseline.** Value is counted **net of the rule already in production**, never gross. Counting the whole 4 seconds instead of the 0.8 inflates the ceiling fivefold, and everything downstream inherits the error. - **The rate the seconds are valued at.** A loaded labour cost is an assumption; state it in the envelope so that the ceiling can be recomputed when it changes, instead of being silently wrong. - **Accuracy does not move it.** A more accurate model can raise the fraction of picks it usefully re-orders and so raise the value side, but it cannot make an over-budget prediction affordable. Cost and quality are separate clauses, and the cost clause binds first. ## Why this is written before a model is chosen Because afterwards, the number stops being a constraint and becomes a justification. Written first, the ceiling behaves exactly like the latency ceiling: it is a published figure any candidate design is checked against, and it rules things out cheaply — on arithmetic in a design round rather than after a quarter of building. It also forces the awkward, useful question early: if the whole capability is worth roughly `$48` a day in saved labour, what may the system that produces it reasonably cost to build and keep running? That comparison belongs at the start, when the answer can still change the plan. ## Where this clause stops The envelope sets the ceiling. It does not trace where money actually goes once the architecture exists — splitting fixed spend that is incurred whether or not a request arrives from the spend that scales per request, attributing shared infrastructure, or measuring the realised cost of a prediction against the bill. Those come later, against this number, and they are a separate exercise with its own methods. The framing round's contribution is the figure they are all checked against.

  • Why does the fraction of picks the model actually re-orders belong in the ceiling?
    Because value accrues only where the prediction changes behaviour. If the rule already in production produces the same order for four picks in five, the model is decorative on those four, and the expected saving per prediction is a fifth of the per-pick saving. Counting the gross saving instead of the increment over the baseline inflates the ceiling by exactly that factor and licenses a design the problem cannot afford.
  • How does the choice of prediction unit move the ceiling?
    Directly, because the ceiling is stated per prediction. One ordering produced per pick and twenty candidate bins scored individually describe the same capability but differ a hundredfold in calls, so a `$0.001` ceiling becomes `$0.00005` per call under the second definition. The envelope has to name the unit, or the number it states means nothing to whoever designs the serving path.

saying these in an interview costs you the question

  • Deriving the cost ceiling from the infrastructure bill rather than the value.
  • Assuming an accuracy gain always justifies a higher cost per prediction.
  • Stating a cost ceiling without saying what one prediction is.
  • Counting the gross saving instead of the gain over the incumbent rule.
  • Ignoring that the existing rule already gets most picks right.
  • Leaving the labour rate implicit so the ceiling cannot be recomputed.