skip to content

A video-moderation scoring fleet is provisioned for peak; overnight uploads fall to a fifth of peak, so why does the monthly bill barely move?

level: seniorimportance: should knowfreq 52%

answer

  1. the invoice tracks capacity, not requests
  2. peak-sized fleet, rented all night
  3. half the day's slots score nothing
  4. the trough is waste, not saving
  5. fill it with deferrable work

basics

~20 s

Because the serving half is paid as provisioned machine-hours, not as forward passes. Capacity held for the evening peak is billed through the quiet hours, so scoring fewer clips overnight consumes less compute without buying back any of the capacity already rented.

solid answer

~40 s

The "per-request" half is only per-request in the *consumption* sense. What you are billed for is **capacity held**, and a fleet sized for the evening peak holds that capacity for all 86,400 seconds of the day. With peak at 500 clips per second, the trough at 100 and the daily mean at 250, the fleet buys the equivalent of `500 x 86,400` clip-slots and scores about half of them; the rest is paid-for capacity that produced nothing. So the overnight drop shows up as idle accelerators, not as a smaller invoice. The two levers that do move the bill are shrinking the capacity you hold and **filling the trough with deferrable work** — the back-catalogue re-scan is exactly that kind of work.

code

pseudocode · 14 lines
pseudocode
provisioned_capacity_clips_per_sec = 500      // held all day for the evening peak
mean_arrival_clips_per_sec         = 250      // averaged over the whole day
trough_arrival_clips_per_sec       = 100      // overnight, a fifth of peak
seconds_per_day                    = 86_400

paid_capacity_clips   = provisioned_capacity_clips_per_sec * seconds_per_day  // 43_200_000
clips_actually_scored = mean_arrival_clips_per_sec * seconds_per_day          // 21_600_000

idle_fraction = 1 - (clips_actually_scored / paid_capacity_clips)             // 0.50

// the bill tracks paid_capacity_clips; the overnight drop moves only
// clips_actually_scored, so it widens idle_fraction instead of shrinking the bill
deferrable_slots_in_trough = (provisioned_capacity_clips_per_sec
                              - trough_arrival_clips_per_sec) * seconds_per_day / 3

go deeper

for a junior

Remember that machines are rented by the hour. Scoring fewer clips overnight uses less of the capacity you already bought; it does not give any of it back.

for a middle

Compute the gap: capacity held over the whole day against clips actually scored, and state the idle fraction. That number is the difference between consumed compute and the invoice.

for a senior

Name the levers that actually move it — hold less capacity, shorten per-clip service time, or schedule latency-insensitive work such as a catalogue re-scan into the paid-for trough — and say which part of the idle gap is deliberate.

for a principal

Decide how much idle capacity the latency budget and the loss of a zone are worth, and hold that as a stated allowance rather than letting it be discovered in an invoice review.

## You rent hours, not forward passes The fixed-versus-variable model says serving scales with traffic, and it does — but it scales with traffic *as consumption*, while the invoice tracks *capacity*. Between those two sits the gap that surprises teams in their second month of production. A scoring tier for a video-moderation service holds enough accelerators to absorb the evening peak, because that is when the clips arrive and that is when the latency budget must still be met. Those machines are rented by the hour. At 03:00 they are still rented. ## The arithmetic of one day Take a diurnal curve with a peak of 500 clips per second, a trough of 100, and a daily mean of 250: - capacity held all day: the equivalent of `500 x 86,400 = 43.2 million` clip-slots; - clips actually scored: `250 x 86,400 = 21.6 million`; - **idle fraction: about 50%** of the day's paid capacity produced no prediction at all. The overnight hours contribute the deepest part of that gap — at a fifth of peak, four-fifths of the capacity in those hours is doing nothing — but the shortfall exists in every hour that is not the peak hour. ## Why the split gets worse, not better This correction pushes the training-versus-serving comparison further in serving's favour: - Consumed serving compute for a quarter was about **21,600 accelerator-hours** against a 600-hour retrain, a ratio of 36. - Paid serving capacity, at roughly twice consumption, is nearer **43,000 hours**, a ratio closer to 72. The training half cannot be corrected the same way, because a training run is scheduled work that saturates its machines for its duration. The serving half is demand-shaped work on machines that must be there for the worst minute of the day. That asymmetry is the real reason serving dominates a mature ML system's bill. ## What actually moves the number Given that the bill follows held capacity, only a few things move it: 1. **Hold less capacity.** Anything that lowers the peak the fleet must absorb — smoothing arrivals, accepting a slower scoring path for a share of uploads, moving a share of clips to a cheaper first-pass check — lowers the floor the whole day pays for. 2. **Fill the trough.** Latency-insensitive work scheduled into the idle hours consumes capacity that is already paid for. The back-catalogue re-scan is the obvious candidate: it has a completion date, not a latency budget. 3. **Cut per-clip service time.** This one is doubly good: it lowers consumption *and* lowers the capacity the peak requires, so it shrinks both terms. 4. **Do less work per clip.** Skipping re-scoring for inputs whose content and deciding policy have not changed removes forward passes outright rather than making them cheaper. Notice what is *not* on the list: waiting for traffic to fall. The trough is not a saving; it is the shape of the waste. ## Where the honest limits are - **Elastic capacity narrows the gap but never closes it.** A fleet that releases machines when demand falls pays for less idle time, but it still holds a floor for correctness and recovery, and it still has to be sized against a peak it has not seen yet. How much headroom that floor needs, and what utilisation target to size to, is a fleet-sizing question rather than a spend-shape one. - **A capacity commitment lowers the rate, not the shape.** Committing to a long-term block of capacity changes what an hour costs; it does not change how many idle hours you bought. - **Idle is not always waste.** Some of the gap is the price of the latency budget, and some is the price of surviving the loss of a zone. The question to answer is how much of the gap is deliberate, not how to drive it to zero. ## The sentence to say in the design round "The serving half is billed as provisioned machine-hours, so the monthly number tracks the peak we hold, not the clips we score. Our lever is not fewer clips at night — it is holding less capacity, or putting the catalogue re-scan into the hours we have already paid for."

  • Does the same reasoning apply to the training half?
    No, and that asymmetry is the point. A training run is scheduled work that saturates its machines for its duration, so its paid hours and consumed hours are close. Serving is demand-shaped work on machines that must exist for the worst minute of the day, so its paid hours run well above what it consumes.
  • If the fleet scales in overnight, is the problem solved?
    It narrows the gap rather than closing it. Released capacity stops billing, but a floor remains for correctness and for absorbing a burst, and the capacity still has to be held against a peak that has not happened yet. The bill falls somewhat; it does not follow the traffic curve down.
  • Which correction should you apply to a cost-per-clip figure because of this?
    Divide by paid machine-hours rather than consumed accelerator time. A figure derived from consumption understates what a clip actually costs by roughly the idle fraction, which here is about a factor of two.

A restaurant staffed for the dinner rush still pays those cooks at three in the afternoon. The quiet hours are not cheaper; they are the same wage buying no covers.

saying these in an interview costs you the question

  • Quiet hours are cheap because fewer clips are scored.
  • The serving bill follows the request curve up and down.
  • Elastic capacity makes idle time disappear entirely.
  • Idle accelerators overnight are always pure waste.
  • A capacity commitment removes the cost of unused hours.