Your per-embedding figure amortises a $90,000 training run over 18 months, but the embedding model is replaced after five - what breaks?
answer
- the horizon is an assumption
- denominator is a guess about lifetime
- short life, fat amortised share
- successor invalidates every stored vector
- cutover recompute is one-off, not steady state
basics
~10 sThe amortised share rises 3.6x, from $0.00025 to $0.0009 per embedding, carrying the fully-loaded figure from $0.0006 to $0.00125. Worse, the successor cannot reuse the old vectors, so the whole catalogue must be re-embedded.
solid answer
~50 sThe horizon was an assumption and it was wrong, so re-derive rather than defend. Planned: `$90,000 / (18 months x 20M) = $0.00025` per embedding. Realised: only `5 x 20M = 100M` embeddings ever carried that run, so the share is `$0.0009` - 3.6 times higher - and the fully-loaded figure moves from $0.0006 to $0.00125. The second half is the one people miss: distances are only meaningful between vectors from the **same** embedding model, so retiring it invalidates all 200 million stored vectors. Before the successor can serve, the entire catalogue must be re-embedded - `200M / 50,000 = 4,000 accelerator-hours`, about $12,000 of compute in one window against $1,200 in a normal month. The fix is to amortise over the observed replacement interval, or to stop amortising and treat each training run as a period cost.
code
pseudocode · 12 linesplanned_share = training_run_cost / (planned_life_months * monthly_embeddings)
// 90000 / (18 * 20e6) = 0.00025
realised_share = training_run_cost / (actual_life_months * monthly_embeddings)
// 90000 / ( 5 * 20e6) = 0.00090
// a successor model cannot reuse vectors from the retired one
if model_version(stored_vector) != model_version(query_encoder):
recompute_required = true
cutover_embeddings = catalogue_images // 200e6, not 20e6
cutover_hours = cutover_embeddings / embeddings_per_accelerator_hour // 200e6 / 50000 = 4000
cutover_compute = cutover_hours * hourly_rate // 4000 * 3 = 12000go deeper
Recall that amortising divides a one-off cost by how many predictions will carry it, and that the lifetime in that division is a guess. A shorter life means a bigger share per prediction.
Recompute both figures on the realised life and state the direction and the multiple. Explain why a monthly amortisation line of $18,000 rather than $5,000 moves the fully-loaded unit figure so much.
Go past the arithmetic to the operational consequence: a new embedding model voids every stored vector, so cutover means recomputing the whole catalogue in one window with change-skipping unavailable. Budget that separately from steady state.
Set the convention. Decide whether training runs are amortised over an evidence-based interval or expensed in period, require the horizon to be published with any unit figure, and make model replacement a trigger to re-issue it.
## The assumption hiding in the denominator Amortising a training run looks like accounting and behaves like a forecast. `training_run_cost / (life_months x monthly_volume)` contains two guesses, and the one that is almost never examined is **life_months**. A model's life is set by things outside the cost model: a better pretrained backbone appears, the catalogue's photography standard changes, a quality regression forces a retrain, a team reorganises. Eighteen months is a plausible number and a wholly invented one. This is the usual place a per-prediction figure goes wrong, and it fails silently: nothing errors, no alarm fires, the published number simply stops being true and stays on the slide. ## Re-deriving on the realised life | | planned | realised | |---|---|---| | training run | $90,000 | $90,000 | | life | 18 months | 5 months | | embeddings carrying the run | 360,000,000 | 100,000,000 | | amortised share per embedding | $0.00025 | $0.0009 | | monthly amortisation line | $5,000 | $18,000 | | fully-loaded cost per embedding | $0.0006 | $0.00125 | The unit figure roughly doubles even though the compute per embedding never changed, and the amortised share alone - $0.0009 - is now **fifteen times the marginal compute** of $0.00006. Any argument that was won on the strength of $0.0006 has to be re-run. ## The consequence the cost model never had a line for A vector is only comparable with vectors produced by the same model. Two embedding spaces placed in one approximate nearest-neighbour index give neighbours that are arbitrary, so a model replacement is not a rolling upgrade; it is a **full catalogue recompute before cutover**: 1. Re-embed all 200 million images with the successor - `200,000,000 / 50,000 = 4,000 accelerator-hours`, about **$12,000** of compute, against $1,200 in an ordinary month. 2. Build or populate a second index from those vectors while the old one keeps serving. 3. Cut reads over, then retire the old vectors and reclaim the storage. During that window the system pays for two copies of the vectors and two indexes, and the change-skipping optimisation that normally suppresses most of the work is void - nothing was produced by the new model, so nothing can be skipped. ## What to do about it - **Amortise over an observed interval, not a convention.** If the last three embedding models lived 7, 5 and 6 months, the defensible horizon is around six, not eighteen. - **Publish the horizon with the figure.** '$0.0006 per embedding, assuming an 18-month model life' is auditable; '$0.0006' is not. - **Give a sensitivity band.** Quoting the number at both a short and a long life shows the reader how much of it is assumption. - **Or stop amortising.** Treating each training run as a period cost in the month it ran removes the forecast entirely, at the price of a lumpy monthly figure. Either convention is defensible; an unstated one is not. - **Budget the cutover separately.** The full recompute is a one-off project cost, not part of the steady-state unit figure, and folding it into the monthly number makes both harder to read. ## The trap in reverse The same arithmetic bites the other way when a model outlives its horizon: the published figure then **overstates** cost and can kill a refresh that is actually cheap. The lesson is not 'assume a short life' but 'the horizon is an input, so treat it like one' - review it whenever a model is replaced, and re-issue the unit figure rather than letting the old one circulate. ## In the room Say the number, then say what it assumes. When the interviewer moves the life from eighteen months to five, recompute out loud, name the direction and magnitude, and then volunteer the recompute of the catalogue - that second step is what separates someone who has replaced an embedding model in production from someone who has only costed one.
- How do you pick a defensible amortisation horizon in the first place?From the observed interval between this team's last few model replacements, not from a depreciation convention borrowed from hardware. If the history is too thin to average, publish the figure with the horizon stated and a sensitivity band at a short and a long life, or expense each training run in the month it ran and accept a lumpy series.
- Why can't the successor serve while the old vectors are still in the index?Distances are only meaningful within one embedding space. Mixing vectors from two models in a single index makes the returned neighbours arbitrary - the query encoder's output has no defined relationship to the old vectors. So the catalogue is re-embedded into a second index, reads cut over, and only then are the old vectors retired.
- The model outlives its horizon instead. What is the failure then?The published figure overstates cost, because the run is spread over fewer embeddings than actually carried it. That is the quieter failure: it can argue a cheap refresh out of existence, or make an in-house capability look worse than an alternative. Re-issue the figure on model replacement in either direction.
saying these in an interview costs you the question
- Treats the amortisation horizon as a measured quantity
- Assumes the successor model can reuse the retired model's vectors
- Keeps quoting the original figure after an early retirement
- Folds the one-off cutover recompute into the steady-state unit cost
- Amortises over the catalogue's life rather than the model's