How would you argue that $0.0006 per refreshed catalogue embedding is cheap, rather than just quoting the number?
answer
- a cost number alone is unarguable
- same denominator on both sides
- value is skewed, cost is uniform
- margin per embedding against cost per embedding
- cadence proportional to value, with a floor
basics
~20 sPut it against what a refreshed embedding earns, on the same denominator. Taking the refresh's attributed margin as $30,000 a month over 20 million embeddings, each costs $0.0006 and returns $0.0015 - and that return is very unevenly spread.
solid answer
~40 sA cost figure alone cannot be argued with or against; a **ratio of value to cost** can. Divide the attributed incremental margin for the same window by the same denominator: `$30,000 / 20,000,000 = $0.0015` of margin per refreshed embedding against `$0.0006` of fully-loaded cost, so the refresh returns about 2.5x. Then refuse to stop at the average. If roughly 90% of that margin comes from the ~5% of listings that actually get surfaced, those earn about `$0.027` each while the remaining 19 million earn about `$0.00016` - below what they cost. The design answer is not 'the refresh is worth it' but **cadence proportional to value**: refresh surfaced listings often, the long tail rarely, with a staleness floor so a cold listing is still current the first time it is shown.
go deeper
Remember that a cost per prediction means nothing alone. It becomes an argument only when it sits beside what one prediction earns, measured over the same window and the same count.
Build the ratio explicitly: attributed margin and fully-loaded cost on the identical denominator, then state the multiple. Say why the fully-loaded figure, not the marginal one, is the side that faces revenue.
Go past the average to the distribution, and turn it into a design change - cadence proportional to value with a staleness floor - while knowing that dropping the tail removes only its marginal spend.
Own the policy: which value figure the organisation puts against unit cost, on what window, and who re-measures it after the cadence changes. Without that, every team argues from a different average.
## Why the number needs a second number Six hundredths of a cent, six ten-thousandths of a dollar - a per-prediction cost has no natural scale, so on its own it is either dismissed as free or feared as a line item, depending on the room. The only way to make it arguable is to put the value of one prediction beside the cost of one prediction, **computed on the same denominator and the same window**. Mismatching those is the most common error: monthly margin against a weekly count, or revenue attributed to the whole search experience against the embeddings refreshed by one job. Take as given that the refresh's attributed incremental margin is **$30,000 a month** - how that figure is established is a measurement question, not a costing one, and the costing answer must simply use the same window it was measured over. ## The ratio, then the distribution | | monthly | per refreshed embedding | |---|---|---| | fully-loaded cost | $12,000 | $0.0006 | | attributed margin | $30,000 | $0.0015 | | net | $18,000 | $0.0009 | An average return of 2.5x looks like the end of the conversation. It is not, because **cost is spread uniformly and value is not**. A rolling window refreshes every listing on the same cadence, so every listing costs the same; but the margin arrives only where a listing is actually retrieved and shown. If the surfacing distribution is the usual heavy-headed one: - the top ~5% of listings, about 1 million embeddings, carry ~$27,000 of the margin - roughly **$0.027 each**, some 45 times their cost; - the remaining 19 million carry ~$3,000 - roughly **$0.00016 each**, well under the $0.0006 they cost. The programme is strongly profitable *and* most of the individual work inside it loses money. Both statements are true and only the second one suggests a change. ## What follows from that 1. **Make cadence follow value.** Refresh frequently where artwork changes are seen, rarely where they are not. The same budget buys more margin without any efficiency work at all. 2. **Keep a staleness floor.** A listing with no traffic today may be surfaced tomorrow, and cold listings are exactly the ones whose first impression matters. A value-weighted cadence with no floor quietly guarantees that every new listing is shown with a stale vector. 3. **Do not expect the tail's cost to leave with the tail.** Dropping 19 million refreshes a month removes only their **marginal** share - `19,000,000 x $0.00006 = $1,140` - while the reservation, the amortised training run and the platform lines stay. The remaining 1 million embeddings then carry all of it, so their fully-loaded unit cost jumps. Capacity has to be resized for the saving to be real. 4. **Re-check the attribution after you change the cadence.** The measured margin was produced under uniform refreshing; a value-weighted schedule changes the thing being measured, so the old number describes a system that no longer exists. ## The traps - **Different denominators on the two sides.** Cost per embedding against margin per user session is a comparison of nothing. Force both onto the same count. - **Averaging away a skew.** A fleet-wide mean over a heavy-tailed distribution says the tail is fine when the tail is the problem. - **Treating the marginal figure as the cost side.** Against revenue, the fully-loaded figure is the honest one; the marginal figure belongs to incremental decisions inside the current shape of the system. - **Assuming value is proportional to freshness.** Some listings' artwork never changes, so refreshing them produces an identical vector and earns nothing at any cadence; that is a change-detection case, not a value-weighting one. ## In the room Give the ratio first, because it answers the question asked. Then volunteer the distribution, because that is what turns a cost figure into a design decision - and a cadence proportional to value, floored for cold listings, is a better answer than any argument about whether six ten-thousandths of a dollar is a lot of money.
- Dropping the tail refresh removes 19 million embeddings a month. How much spend actually leaves?About $1,140 - their marginal compute at $0.00006 each. The reservation, the amortised training run, the platform lines and most of the stored-vector cost all stay, so the bill barely moves and the fully-loaded cost of the embeddings that remain rises sharply. The saving is only realised if the committed capacity is resized to the smaller job.
- What does a uniform rolling window give you that a value-weighted cadence does not?A bounded worst-case staleness for every listing, including ones with no traffic history. A value-weighted cadence optimises for listings that are already being surfaced, so without an explicit staleness floor it systematically starves new and cold listings - the ones whose first impression is most likely to be the one that matters.
saying these in an interview costs you the question
- Quotes a per-prediction cost with no value figure beside it
- Assumes every refreshed listing earns the same margin
- Uses different denominators for the cost and value sides
- Concludes the tail is profitable from the fleet-wide average
- Expects dropping the tail to remove its fully-loaded share of cost
- Compares attributed margin against the marginal rather than fully-loaded cost