skip to content

A policy classifier costs 600 accelerator-hours per quarterly retrain and scores 1.9 billion clips at 40 ms each — which half dominates?

level: middleimportance: must knowfreq 57%

answer

  1. convert both to one unit first
  2. clips times service seconds
  3. then divide by 3600
  4. crossover equals hours times 3600 over service time
  5. 54 million clips per retrain here

basics

~20 s

Serving dominates, by roughly 36 to 1. At 40 ms per clip, 1.9 billion clips need about 21,600 accelerator-hours against the retrain's 600, and the two halves would only be equal at about 54 million clips per retrain interval.

solid answer

~30 s

Convert both halves into the same unit, accelerator-hours, then compare. Serving: `1.94e9 clips x 0.040 s = 77.8 million seconds`, which is about **21,600 accelerator-hours** for the quarter. Training: **600 accelerator-hours**, once. Serving is therefore about **36x** the training run. The number worth carrying away is not 36 but the **crossover**: `600 h x 3600 / 0.040 s = 54 million clips`. Below 54 million clips per retrain interval the run dominates; above it, scoring does. This system sits 36x past the crossover, which is why any effort aimed at the training run is aimed at under 3% of the bill.

code

pseudocode · 16 lines
pseudocode
training_accel_hours     = 600        // one quarterly retrain, given
service_seconds_per_clip = 0.040      // accelerator time per clip, given
clips_this_quarter       = 1_940_000_000

serving_accel_hours = clips_this_quarter * service_seconds_per_clip / 3600
// 1_940_000_000 * 0.040 / 3600 = 21_556 hours, call it 21_600

crossover_clips = training_accel_hours * 3600 / service_seconds_per_clip
// 600 * 3600 / 0.040 = 54_000_000 clips per retrain interval

if clips_this_quarter > crossover_clips then
    dominant = "serving"        // 1.94e9 is 36x the crossover, so this branch fires
    ratio    = serving_accel_hours / training_accel_hours   // about 36
else
    dominant = "training"
    ratio    = training_accel_hours / serving_accel_hours

go deeper

for a junior

The step to remember is the unit conversion: clips times seconds per clip, divided by 3600, gives hours you can compare with the training run's hours.

for a middle

Derive the crossover rather than just the ratio: training hours times 3600, divided by per-clip seconds. Then say which inputs move it and which merely move where the system sits relative to it.

for a senior

Qualify the inputs before trusting the result — failed runs inflate the fixed half, provisioned rather than consumed hours inflate the variable half, and accelerator time is not machine time.

for a principal

Use the ratio to set a ceiling on effort: with the fixed half under 3% of compute, the argument for where the next engineering quarter goes is settled before anyone profiles anything.

## Put both halves in one unit The halves are quoted in incompatible units — a training run in hours, an inference in milliseconds — and that mismatch is exactly why people guess the answer wrong. The fix is mechanical: express both as **accelerator-hours over the same interval**, which here is one retrain interval, a quarter. 1. **Serving.** Clips in the interval times per-clip accelerator time. About 1.94 billion clips (250 per second averaged over a 90-day quarter) at `0.040 s` each is `77,760,000 s`, or **21,600 accelerator-hours**. 2. **Training.** One run, given as **600 accelerator-hours**. 3. **Compare.** `21,600 / 600 = 36`. The serving half is thirty-six times the training half. ## The crossover, not the ratio The ratio 36 is specific to this quarter's traffic. The reusable number is the **crossover clip count** — the predictions per retrain interval at which the halves are equal: `crossover = training_hours x 3600 / service_seconds_per_clip = 600 x 3600 / 0.040 = 54,000,000 clips` That single figure answers a family of questions without redoing the arithmetic. Fifty-four million clips per quarter is the break-even; anything above it is serving-dominated, and the further above, the more decisively. | Clips per retrain interval | Serving hours | Dominant half | |---|---|---| | 5 million | 56 | training, heavily | | 54 million | 600 | equal — the crossover | | 540 million | 6,000 | serving, 10x | | 1.94 billion | 21,600 | serving, 36x | ## What actually moves the crossover Only two inputs appear in the formula, so only two things move it: - **The retrain's accelerator-hours** move it *proportionally*. A run twice as long pushes the crossover to 108 million clips — the fixed half now needs twice the traffic before it is outweighed. - **Per-clip accelerator time** moves it *inversely*. Cutting 40 ms to 20 ms doubles the crossover to 108 million; raising it to 80 ms halves the crossover to 27 million. And two things that feel relevant but are not: - **Upload volume** does not move the crossover. It moves *where you sit* relative to it, which is a different statement. - **Retraining cadence** does not move it either. Halving the interval halves both the clips in that interval and the number of intervals, so the break-even *per interval* is unchanged; what changes is that you cross it later in each cycle. ## Reading the result honestly Three refinements stop the number from being quoted too confidently: - **Failed runs and sweeps count.** The 600 hours is the winning run. If a release typically burns three abandoned runs and a hyperparameter sweep, the honest fixed half might be two or three times the headline figure — which moves the ratio from 36x to roughly 12x, and still does not change the verdict. - **Accelerator time is not machine time.** A clip that occupies 40 ms of accelerator work may occupy far more wall-clock time on a machine that is also decoding video, fetching features and waiting on the network. Compare like with like, or state which one you are counting. - **You pay for provisioned hours, not consumed ones.** A fleet held at peak size bills through the trough, so the *paid* serving hours exceed the 21,600 consumed. That correction pushes the ratio further in serving's favour, never back toward training. ## Why an interviewer asks this The question is not arithmetic for its own sake. It checks three things at once: - that you **convert to a common unit** before comparing, rather than arguing from the feel of the numbers; - that you know **which lever belongs to which half**, so you do not propose a change to the fixed half as a fix for a variable-half problem; - that you can state the **ceiling on a saving** before doing the work. With training at 600 of roughly 22,200 accelerator-hours, even deleting the training run entirely saves under 3%. That is the whole argument for pointing the next optimisation at the per-clip path instead. The answer an interviewer wants is therefore two sentences and a number: serving dominates by about 36x, the crossover is 54 million clips per retrain interval, and here is the division that produced both.

  • Does retraining monthly instead of quarterly move the crossover?
    No. The crossover is `training_hours x 3600 / service_seconds` and cadence appears in neither term, so it stays at 54 million clips per interval. What changes is that a month carries only a third of the clips, so each interval sits closer to the crossover while the annual training total triples.
  • The team measures 40 ms of accelerator time but 180 ms of machine time per clip. Which number belongs in this comparison?
    Whichever one the bill is denominated in, used consistently on both sides. If you rent whole machines, count machine time on the serving side and machine time on the training side; mixing accelerator time for one half and machine time for the other silently inflates whichever half got the larger unit.
  • How much would halving per-clip accelerator time save here?
    About 10,800 accelerator-hours a quarter, roughly eighteen retrains' worth. It also doubles the crossover to 108 million clips, so the system moves from 36x past break-even to 18x — still decisively serving-dominated.

saying these in an interview costs you the question

  • Comparing hours against milliseconds without converting units.
  • Assuming 40 ms per clip must be negligible at any volume.
  • Thinking more traffic shifts the crossover point itself.
  • Counting only the winning training run, not the failed ones.
  • Mixing accelerator time on one side with machine time on the other.