skip to content

In a video-moderation service, which half of the spend grows with upload volume: the quarterly training run or per-clip scoring?

level: juniorimportance: must knowfreq 62%

answer

  1. one is paid once, one repeatedly
  2. traffic multiplies only one half
  3. fixed with respect to traffic, not cadence
  4. milliseconds times billions is not free
  5. predictions per retrain flips it

basics

~20 s

Per-clip scoring grows with upload volume; the quarterly training run does not. Training is paid once per model version, while scoring is paid again for every clip that arrives, so only the serving half is multiplied by traffic.

solid answer

~40 s

Training is a **fixed** cost with respect to traffic: one run produces one model version, and it consumes the same accelerator-hours whether that model later scores a thousand clips or a billion. Per-clip scoring is a **variable** cost: every upload buys another forward pass, so this half is multiplied by upload volume. That asymmetry is why the split flips with scale — a model in a pilot looks training-dominated, and the same model under real upload volume is overwhelmingly serving-dominated. Note what "fixed" does and does not mean: the training half still moves if you retrain more often or run longer. It simply never moves because more users uploaded video.

go deeper

for a junior

Recall the shape of each half: the training run is paid once per model version, scoring is paid once per clip. When traffic grows, only the second one grows with it.

for a middle

Explain the mechanics: put both halves into accelerator-hours, note that the training half is fixed with respect to traffic but not to cadence or run length, and name the crossover as the clip count where they are equal.

for a senior

Show that you have watched the split flip in a live system, and that you know the serving half is really paid as provisioned machine-hours rather than per forward pass, so a quiet night does not shrink it much.

for a principal

Frame it as which half deserves engineering effort and which unit each half is forecast in — a fixed cost is budgeted per release, a variable cost turns the traffic forecast into the cost forecast.

## The two shapes of spend A production ML tier buys compute in two shapes, and almost every cost argument in a design round turns on telling them apart. **Fixed spend** is paid once per model version. A video-moderation service retrains its policy classifier on accelerator hardware a few times a quarter. That run consumes what it consumes whether the model it produces goes on to score one clip or a billion. **Variable spend** is paid again for every item scored. Every upload buys another forward pass through the model, so this half is multiplied by traffic. "Fixed" here means *fixed with respect to traffic*, not fixed in every sense. A longer run, a larger model or a faster retraining cadence all move the training half. What never moves it is more users uploading more video. ## Why the intuition points the wrong way Training feels like the expensive half for reasons that have nothing to do with arithmetic: - it is **lumpy and visible** — one scheduled job, one owner, one number in a review; - it runs on the **most expensive machines** in the estate, so the hourly rate is the one people remember; - per-clip inference is quoted in **milliseconds**, and milliseconds sound free; - during development traffic really is tiny, so the early bill really is training-dominated, and that early impression survives launch. The serving half hides because it is spread across millions of small, boring events. It is a floor, not a spike, and nobody reviews a floor. ## The ratio that flips the split Put both halves in one unit — accelerator-hours — and a single number settles the argument: **predictions served per retrain interval**. 1. Take the retrain's accelerator-hours as a given. 2. Multiply the clips scored in that same interval by the per-clip accelerator time, and convert to hours. 3. Compare the two. The **crossover** is the clip count at which they are equal. Below the crossover the system is training-dominated. Above it, it is serving-dominated, and every further doubling of traffic pushes it further out. A moderation tier at real upload volume sits far above the crossover, which is why "training is the expensive part" is a true statement about a pilot and a false one about production. | | The quarterly training run | Per-clip scoring | |---|---|---| | Paid | once per model version | once per clip scored | | Grows with | run length, model size, cadence | upload volume | | Shape over a month | one spike | a continuous floor | | Effect of 10x traffic | none | 10x | | Where a saving lands | once per retrain | on every clip, forever | ## The same system at three stages 1. **Prototype.** A few thousand evaluation clips. Training is effectively the whole bill, and effort spent on inference efficiency is wasted effort. 2. **Launch.** Traffic approaches the crossover. The halves are comparable and either is worth attacking. 3. **Scale.** Traffic is tens of times the crossover. Serving *is* the bill, and the training run is a rounding error that still receives most of the attention. Nothing about the model changed between those three stages. Only the denominator did. ## What the distinction buys you in a design round - It says **which half to attack**. A percentage shaved off the dominant half is worth far more than the same percentage off the other, and the ratio names the dominant half before any optimisation work begins. - It says **which half a proposal touches**. "Retrain weekly instead of quarterly" moves the fixed half; "score every clip with a second model as well" moves the variable half against the whole of traffic. - It sets the **unit you forecast in**. Fixed spend is forecast per release; variable spend is forecast per unit of demand, which quietly turns a traffic forecast into a cost forecast. - It explains why an internal or invite-only deployment behaves so differently from the public one: identical code, opposite cost structure. ## Two honest caveats Two refinements keep the simple picture from becoming wrong: - The serving half is only *approximately* per-clip. Machines are rented by the hour, not by the forward pass, so what you actually pay tracks **provisioned capacity**; a fleet held at peak size still bills through a quiet night. - The training half is not one number either. Failed runs, sweeps and the evaluation passes around a release all belong to it, and a team that counts only the winning run understates it. Neither caveat changes the direction of the answer at scale. Upload volume multiplies one half and leaves the other untouched, and that asymmetry is the whole point.

  • If the training half is fixed, does that mean it never changes?
    It is fixed only with respect to traffic. It still moves with retraining cadence, run length, model size, failed runs and the sweeps around a release. Doubling how often you retrain doubles it; doubling how many clips are uploaded does not touch it.
  • Why does the same model look training-dominated during a pilot and serving-dominated at launch?
    Only the clip count changed. The pilot serves far fewer predictions per retrain interval than the crossover, so the one-off run outweighs the accumulated forward passes; at launch the interval carries orders of magnitude more clips, and the accumulated per-clip compute overtakes the run.
  • Does adding a second model to score every clip change the fixed half or the variable half?
    Mostly the variable half. The second model adds another one-off training run, but it also adds a forward pass to every single upload, so its cost is multiplied by traffic while the first model's run is not.

Cutting a film costs the same whether twelve people watch it or twelve million; the projector hours do not. Beyond a certain audience, the projectors cost more than the edit did.

saying these in an interview costs you the question

  • Training is always the expensive half of an ML system.
  • Serving is cheap because one forward pass takes milliseconds.
  • Doubling upload volume doubles the training bill too.
  • An idle model and a saturated one cost about the same overall.
  • The split is a property of the model, not of the volume served.