skip to content

Skipping catalogue images whose bytes are unchanged cuts monthly re-embedding compute by 80% - what happens to cost per embedding?

level: seniorimportance: should knowfreq 44%

answer

  1. check both halves of the ratio
  2. same fraction off top and bottom
  3. pick the unit matching the decision
  4. fixed spend does not follow the work
  5. throughput levers, not count levers

basics

~10 s

Marginal cost per embedding produced does not move: numerator and denominator fall together, so it stays $0.00006. Cost per listing kept current falls fivefold, and the fully-loaded figure per embedding actually rises.

solid answer

~40 s

Compute falls from $1,200 to $240 and output falls from 20 million to 4 million embeddings, so `$240 / 4,000,000` is still `$0.00006` - the marginal unit figure is unchanged, because the lever removed work rather than making work cheaper. What did improve is a different unit: **cost per listing kept current**, which goes from `$1,200 / 20M = $0.00006` to `$240 / 20M = $0.000012`, five times better, since all 20 million listings are still covered. The fully-loaded figure moves the other way: the reservation, the amortised training run and the platform lines do not shrink, so about $10,920 now divides by 4 million embeddings - roughly `$0.0027` each. The saving is real only if the reservation is resized to match the smaller workload.

go deeper

for a junior

Recall that cost per prediction is a ratio. If a change removes work, both the spend and the count fall together and the ratio can stay exactly where it was.

for a middle

Compute the three figures and keep them apart: per embedding produced, per listing kept current, and fully loaded. Explain why the third rises when output falls faster than committed spend.

for a senior

Report which figure moved and resize the committed capacity so the compute cut reaches the bill. Know that the skip state is scoped to the embedding model and dies with it.

for a principal

Require efficiency claims to name their denominator. A programme that reports 'cost per prediction' without saying which one will accumulate savings that never appear in the invoice.

## Check both halves of the ratio A unit cost is a ratio, so any lever has to be classified by **which half it touches**. Change-skipping - hashing each listing image and re-embedding only those whose bytes differ from last pass - removes work. It cuts the numerator and the denominator by the same 80%, which is why the marginal figure is untouched: | | before skipping | after skipping | |---|---|---| | embeddings produced | 20,000,000 | 4,000,000 | | accelerator-hours | 400 | 80 | | marginal compute | $1,200 | $240 | | marginal cost per embedding produced | $0.00006 | $0.00006 | | cost per listing kept current | $0.00006 | $0.000012 | | fully-loaded per embedding produced | $0.0006 | about $0.0027 | All three rows at the bottom are honest; they are three different questions. Quoting the first to claim a saving is wrong, quoting the third to claim a regression is wrong, and the second is the one that matches what the lever actually did. ## Which unit matches the decision - **Cost per embedding produced** - the price of making one. Use it when comparing models, batch shapes or hardware, because those change the cost of the work itself. - **Cost per listing kept current** - the price of the *outcome* the refresh exists for. Use it for anything that changes how much work is done: change-skipping, a longer window, a narrower catalogue slice. - **Fully-loaded per embedding** - the price of the capability, spread over what it produced. Use it against revenue, and expect it to rise whenever output falls faster than committed spend. ## Levers that really do move cost per embedding produced Each of these raises **embeddings per accelerator-hour**, which is the only quantity the marginal figure divides by: 1. **A smaller embedding model** - fewer operations per image, so more images per hour, at whatever retrieval quality the smaller model gives. 2. **Larger inference batches** - more images in flight per pass, which raises throughput until it plateaus. 3. **Higher accelerator utilisation across the window** - keeping the devices fed rather than idle between batches, by overlapping decode and transfer with compute. Notice that none of these reduces the number of images embedded, and change-skipping does nothing to throughput. They are complementary, and they are also the reason a cost review should ask *which figure moved* before congratulating anyone. ## The second-order effects of skipping - **Fixed spend does not follow the work.** With the reservation untouched, the monthly bill falls only from $12,000 to about $10,920 - roughly 9% - because the compute was never the large line. Resizing the reservation to the smaller workload is what converts the 80% compute cut into a real saving. - **The skip check has its own cost.** A hash read and comparison per listing image runs for all 20 million, including the 16 million that are skipped. It is far cheaper than a forward pass here, but on a very cheap model or a very high change rate it stops paying. - **The skip state is model-scoped.** The stored hashes say 'this image was already embedded by *this* model'. Replace the embedding model and every entry is void; the next pass is a full recompute at full cost. - **Changed bytes are not the same as changed meaning.** Re-encoding or a watermark flips the hash without changing what the image shows, so some of the remaining 20% is wasted work; conversely a listing whose *photo* is untouched genuinely needs no new vector, which is why the lever works at all. ## Reporting it honestly State the lever, then the three figures with their units: compute down 80%, cost per listing kept current down fivefold, cost per embedding produced unchanged, fully-loaded per embedding up unless the fleet is resized. That framing survives the follow-up question. A single sentence claiming 'we cut cost per prediction by 80%' does not, and it is the version most likely to be repeated in a review where nobody can reconstruct which denominator was used.

  • Which levers do move marginal cost per embedding produced?
    The ones that raise embeddings per accelerator-hour: a smaller embedding model, larger inference batches, and keeping the accelerators fed across the window instead of idling between batches. Change-skipping is not one of them - it removes work rather than making work cheaper, which is why it leaves the ratio exactly where it was.
  • When does the skip check cost more than it saves?
    When the per-image hash read and comparison approaches the per-image embedding cost - a very small model, tiny inputs, or a storage layer where reading the bytes to hash them is itself expensive. It also degrades as the change rate rises, since the check runs for every image but only avoids work on the unchanged fraction.
  • The monthly bill fell only 9% after an 80% compute cut. What would you change?
    Resize the committed capacity to the new shape, because the compute was never the large line: the reservation, the amortised training run and the platform costs are. Either shrink the reservation to the smaller workload or fill the freed hours with work that was previously queued, so the paid hours produce something.

saying these in an interview costs you the question

  • Claims change-skipping lowers marginal cost per embedding produced
  • Reports an 80% compute cut as an 80% cut in the bill
  • Leaves the reservation untouched and still claims the full saving
  • Forgets the skip state is void after an embedding model change
  • Treats cost per embedding and cost per listing kept current as one unit