skip to content

A bulk catalogue re-embedding job runs 400 accelerator-hours a month at $3 an hour for 20 million embeddings - what is the cost per embedding?

level: juniorimportance: must knowfreq 62%

answer

  1. spend over output, nothing cleverer
  2. name the denominator before dividing
  3. hours times hourly rate first
  4. quote per thousand for readability
  5. marginal compute only, not fully loaded

basics

~10 s

Six hundredths of a cent. 400 hours at $3 is $1,200 of compute, divided by 20 million embeddings gives $0.00006 each, or $0.06 per thousand. That figure covers marginal compute only.

solid answer

~40 s

Multiply the hours by the rate to get the numerator, then divide by what the window actually produced: `400 x $3 = $1,200`, and `$1,200 / 20,000,000 = $0.00006` per embedding. I would quote it as **$0.06 per thousand embeddings**, because a number with four leading zeros is hard to argue about. Two things have to be said out loud with it. First, which denominator: embeddings produced in the window, not images in the catalogue and not search requests served later. Second, which spend: this is the bulk workers' compute for that window, so it excludes the amortised training run, the idle share of any reservation, and vector storage. It also implies 50,000 embeddings per accelerator-hour, which is the sanity check a reviewer will run.

go deeper

for a junior

Recall the shape: spend for a window divided by what the window produced, with both sides on the same window. Say the unit out loud - per embedding, per thousand - and say that it is compute only.

for a middle

Explain why the denominator is a choice: embeddings produced, catalogue images, accelerator-hours and queries all give different units. Derive throughput from the same two inputs and use it as a sanity check.

for a senior

Label the figure marginal, and name in one breath what a fully-loaded version would add: amortised training, idle reservation, vector storage, platform. Quote the number at a scale that cannot be rounded to zero.

for a principal

The unit you publish shapes every later decision. Fix one definition with its window and its exclusions written down, so refresh cadence, fleet changes and value comparisons are all argued against the same denominator.

## The figure a design round is asking for When an interviewer asks what a prediction costs, they want **one number with a stated unit**, not a shrug and not a shopping list. The arithmetic is deliberately trivial: a spend over a count. The judgment is in naming *which* spend and *which* count, because a catalogue image-embedding refresh has several of each and quoting the wrong pair is how this question is actually failed. The worked setting throughout: a catalogue of **200 million listing images** is re-embedded on a rolling window so that search and recommendations see current artwork. The bulk re-embedding workers produced **20 million embeddings** last month and consumed **400 accelerator-hours** at a given rate of **$3 per accelerator-hour**. ## The two inputs 1. **The numerator** - `400 accelerator-hours x $3/hour = $1,200`. Take it from what was billed for the window, not from a per-image estimate, because the billed hours already include retried batches, warm-up and the tail of the last run. 2. **The denominator** - `20,000,000` embeddings produced in that same window. Same window on both sides, always: a month of spend over a week of output is the classic off-by-four. `$1,200 / 20,000,000 = $0.00006` per embedding, which is **$0.06 per thousand** or **$60 per million**. ## Choosing the denominator deliberately Four counts are available in this system and each yields a different, legitimate unit: | denominator | the unit it produces | what it answers | |---|---|---| | embeddings produced in the window | cost per embedding | what one more embedding costs to make | | images in the catalogue | cost per catalogue image per month | what keeping the whole catalogue current costs | | accelerator-hours consumed | cost per accelerator-hour | the input rate you already had | | search requests the vectors served | cost per query | a serving-side figure this job does not own | Only the first is *cost per prediction*. The second is the one an executive usually means, and the two diverge the moment the job stops re-embedding everything every month. ## The throughput sanity check Dividing output by hours gives **50,000 embeddings per accelerator-hour**, about 14 per second. Quote it. A reviewer holds a rough sense of how long an image forward pass takes, and if their estimate and yours disagree by an order of magnitude then one of the two inputs is wrong - usually the hours, because a reservation was billed and only part of it was used. Having the throughput number also makes the next question answerable without new arithmetic: a one-off recompute of the full 200-million-image catalogue is `200,000,000 / 50,000 = 4,000 accelerator-hours`, which is $12,000 of compute, ten times a normal month. ## What this number is not This is a **marginal** figure: the compute consumed producing those embeddings and nothing else. It excludes, at minimum: - the **training or fine-tuning run** that produced the embedding model, amortised across the embeddings it will ever serve; - the **idle share of a reservation** - hours that are paid for whether the job runs or not; - **vector storage and the index writes** each produced embedding triggers; - the **orchestration and monitoring** that runs the job and watches it. Each of those is real money, and adding them produces a **fully-loaded** figure that on this system is roughly ten times larger. Both numbers are legitimate; they answer different questions, and quoting one while the audience is asking the other is the most common way this discussion goes wrong. The correct move in an interview is to state the marginal figure, label it as marginal, and say in one sentence what would have to be added to make it fully loaded. ## Units and presentation Per-prediction costs are small by construction, so pick a scale that survives being read aloud: **per thousand** for cheap inference, **per million** for very cheap inference, and always with the time window attached. A figure quoted as $0.00006 invites rounding to zero, and 'inference is basically free' is a conclusion that stops the cost conversation exactly where it should start.

  • What throughput does that figure imply, and why quote it alongside?
    `20,000,000 / 400 = 50,000` embeddings per accelerator-hour, about 14 a second. It is the sanity check: a reviewer's own sense of per-image compute either agrees or tells you one input is wrong, usually the hours, because a reservation was billed and only part of it consumed. It also makes the next estimate free - a full 200-million-image recompute is 4,000 accelerator-hours.
  • The job re-ran 10% of its batches after failures. Does that go on top of the $1,200?
    Not if the $1,200 came from the bill. Billed accelerator-hours already contain the retries, so adding a failure allowance on top double-counts. You add it only when the numerator was built bottom-up from a per-image compute estimate, which is the case where retries, warm-up and idle gaps between batches are all invisible.

saying these in an interview costs you the question

  • Quotes an hourly rate when asked for a per-prediction figure
  • Divides by catalogue size instead of embeddings produced in the window
  • Uses a month of spend against a week of output
  • Rounds the figure to zero and concludes inference is free
  • Cannot say which spend the number covers when asked