Which costs does a per-embedding figure computed from the catalogue re-embedding workers' compute bill alone leave out?
answer
- two numbers, two different questions
- marginal answers one more tonight
- fully loaded answers keep or kill
- a reservation bills its idle hours
- the training run amortises into every unit
basics
~20 sThe amortised training run, the reserved-but-idle share of the accelerator fleet, vector storage and index writes, and orchestration and monitoring. On the worked catalogue refresh they turn $0.00006 of marginal compute into about $0.0006 fully loaded.
solid answer
~50 sFour lines are missing, all of them real money: the **training run** that produced the embedding model, amortised across the embeddings it will serve; the **hours of the reservation the job did not consume** but that were billed anyway; **vector storage and the index writes** every produced embedding triggers; and the **orchestration and monitoring** around the job. On the worked refresh the compute actually used is $1,200 a month while the fully-loaded bill is $12,000, so the unit figure moves from `$0.00006` to `$0.0006` - ten times. Both numbers are correct and they answer different questions: the marginal one answers *should we embed one more batch tonight*, since the fixed lines are already committed; the fully-loaded one answers *is this capability worth keeping*, and is the only one you may put beside a revenue figure.
code
pseudocode · 12 linesmonthly_fully_loaded(window):
reserved = reserved_accelerator_hours(window) * hourly_rate // 1440 * 3 = 4320
amortised = training_run_cost / planned_life_months // 90000 / 18 = 5000
storage = vector_storage_cost(window) + index_write_cost(window) // = 1600
platform = orchestration_cost(window) + monitoring_cost(window) // = 1080
return reserved + amortised + storage + platform // = 12000
produced = embeddings_produced(window) // 20,000,000
used = consumed_accelerator_hours(window) * hourly_rate // 400 * 3 = 1200
fully_loaded_per_embedding = monthly_fully_loaded(window) / produced // 0.0006
marginal_per_embedding = used / produced // 0.00006go deeper
Remember that the compute bill is not the whole cost. Training, idle reserved hours, storage and the platform around the job are all real spend and all land in the per-prediction figure.
Build the roll-up line by line and show the arithmetic from each line to its per-embedding share. Explain why the reservation's unused hours count and why the fully-loaded figure is roughly ten times the marginal one here.
Pick the right figure for the decision on the table and say why: marginal for incremental work inside the current shape, fully loaded for keep-or-kill and for any comparison against revenue. Name the two largest lines unprompted.
Decide which lines the organisation counts and publish that definition, including people time and idle reservations. Teams that define the denominator differently will reach opposite conclusions from the same system.
## Two figures, two questions A per-prediction cost is not a single quantity. **Marginal cost** is what producing one more prediction adds to the bill right now. **Fully-loaded cost** is the total cost of having the capability at all, spread over what it produced. On a catalogue image-embedding refresh these differ by roughly an order of magnitude, and an engineer who can only produce one of them will misprice at least one decision. The marginal figure from the worked setting: the bulk re-embedding workers consumed 400 accelerator-hours at a given $3 an hour, producing 20 million embeddings, so `$1,200 / 20,000,000 = $0.00006` each. ## The roll-up to fully loaded | line | monthly | per embedding | why it belongs | |---|---|---|---| | reserved accelerator fleet (1,440 hours) | $4,320 | $0.000216 | billed whether the job runs or not; only 400 hours were consumed | | amortised training run ($90,000 over 18 months) | $5,000 | $0.00025 | no embeddings exist without it | | vector storage and index writes | $1,600 | $0.00008 | every embedding produced is stored and indexed | | orchestration and monitoring | $1,080 | $0.000054 | the job does not run itself | | **total** | **$12,000** | **$0.0006** | ten times the marginal figure | The largest single line is not the compute. It is the **amortised training run**, followed by the **idle share of the reservation** - $4,320 billed against $1,200 consumed. That ordering is typical of a bulk embedding workload and it is why quoting only the compute understates so badly. ## Why the reservation counts even when idle A reservation buys hours, not work. If two accelerators are held for the month, 1,440 accelerator-hours are paid for; the job used 400. The remaining 1,040 hours are a cost of *having the capability available on this cadence*, and they belong in the fully-loaded figure exactly the way an empty seat belongs in an airline's cost per passenger. They are also the lever nobody pulls: the usual response to a cost review is to make the job cheaper, when the honest first move is to notice that two-thirds of the reserved hours produced nothing. ## Using the right one - **Marginal** - incremental decisions inside the current shape of the system: re-embed one more batch tonight, add a category to the rolling window, back-fill a supplier's listings. The fixed lines do not move, so including them argues against work that is genuinely almost free. - **Fully loaded** - decisions about the capability: keep it, cut its cadence, compare it against the margin it earns, or set an internal charge for another team that wants to use it. A figure quoted without its label is the defect. 'Six hundredths of a cent per embedding' sounds like a rounding error and will win an argument it should not; '$0.0006 fully loaded, of which the training run is $0.00025' invites the right next question. ## The denominator moves too Because four of the five lines are fixed within a month, the fully-loaded figure is **inversely sensitive to output**. Halve the embeddings produced without changing the reservation and the fully-loaded unit cost nearly doubles while the bill barely falls. That is not a paradox; it is the definition working correctly, and it is why any claim that an efficiency change reduced cost per prediction has to say which of the two figures moved, and whether the fixed lines were resized to match. ## What to say in the room Give the marginal number, label it, then give the fully-loaded number with its two or three largest lines named. Say which one you would use for the decision on the table. If you are asked for one number, give the fully-loaded one and its window: it is the one that cannot be gamed by pushing spend into a line you chose not to count.
- Which of the two do you quote when deciding whether to re-embed one extra batch tonight?The marginal one. The reservation is already paid for the month and the training run is already spent, so the only money that moves is the accelerator time the extra batch consumes. Quoting $0.0006 there argues against work that actually costs $0.00006 and leaves reserved hours idle.
- Why can the fully-loaded figure rise in a month when the bill fell?Because most of it is fixed within the month. If output falls faster than spend - fewer images changed, a shorter window, a paused category - the denominator shrinks while the reservation, the amortised run and the platform lines do not. Cost per embedding rises even though less money left the account.
- Where does engineering time belong in this roll-up?In the fully-loaded figure if you are arguing keep-or-kill, as a stated line rather than folded silently into another one; out of it if you are arguing an incremental refresh. Whichever you choose, say so - an unstated people line is the usual reason two teams quote different unit costs for the same system.
A restaurant's cost per meal is not the price of the ingredients on the plate. Serving one more portion costs the ingredients; running the kitchen costs the rent, the empty tables and the chef's training, whether or not that portion is ordered.
saying these in an interview costs you the question
- Treats the compute bill as the total cost of the capability
- Amortises the training run but ignores idle reserved hours
- Forgets that produced vectors cost money to store and index
- Uses the fully-loaded figure to decide whether to refresh one more batch
- Expects the marginal and fully-loaded figures to agree