skip to content

A policy change forces re-scanning a 4-billion-clip back catalogue — why is that spend neither the training half nor the per-upload half?

level: seniorimportance: should knowfreq 41%

answer

  1. a third bucket, not either half
  2. catalogue size drives it, not traffic
  3. one policy edit, months of compute
  4. no user waiting on it
  5. scope, dedupe by content hash, defer

basics

~20 s

A bulk re-scan is a third bucket: per-item like serving, but triggered by a policy decision rather than by traffic, and unbounded by any user-facing latency. At 40 ms per clip it costs about 44,400 accelerator-hours, roughly two quarters of live scoring in one pulse.

solid answer

~40 s

It is **per-item** like the serving half — four billion clips at `0.040 s` each is about **44,400 accelerator-hours** — but nothing about upload traffic caused it. A policy edit did, so its frequency is a governance decision, not a demand curve, and it cannot be forecast from a traffic model. It is also **one-off and lumpy** like a training run, yet it is two orders of magnitude larger than one: about 74 retrains, or twice a whole quarter of live scoring. The saving grace is that it carries **no user-facing latency budget**, which makes it the only compute in the system that you can throttle, defer, run on cheaper interruptible capacity, or scope down. Budget it per policy change, not per month.

go deeper

for a junior

Notice that this compute is per clip, like serving, but no user triggered it and no user is waiting for it. That combination is what makes it a separate bucket.

for a middle

Do the arithmetic against both neighbours: catalogue size times per-clip time, compared with one retrain and with a quarter of live scoring, so the magnitude of a single policy change is explicit.

for a senior

Show the controls you would actually build — scoping by the policy's reach, deduplication by content and policy version, deferral into the overnight trough, and a stated priority against the live upload path.

for a principal

Treat it as a governance cost, not an engineering one: the decision bar for editing a policy should reflect that a wording change can cost more compute than two quarters of live scoring.

## A third bucket, not a variation on the other two The fixed-versus-variable picture has room for exactly two kinds of compute: the run that produces a model version, and the forward pass that serves a request. A back-catalogue re-scan is neither, and a cost model that forces it into one of the two slots will be wrong in both directions. - Slot it as **serving** and your per-clip figures become unforecastable, because they now contain spikes no traffic model predicts. - Slot it as **training** and you have a "fixed" cost two orders of magnitude larger than the run it sits beside, firing on a schedule nobody owns. ## What it costs here The arithmetic is the same shape as the serving half, with the catalogue in place of the traffic: - 4 billion clips at `0.040 s` of accelerator time each is `160,000,000 s`, about **44,400 accelerator-hours**. - Against one retrain at 600 hours, that is roughly **74 retrains** in a single pass. - Against a quarter of live scoring at 21,600 hours, it is **about two quarters' worth**, delivered in whatever window the policy team asks for. A single policy edit can therefore cost more than everything else the system does in six months. This is the number teams most often discover after the fact. ## How it behaves unlike each half | | Training run | Live per-upload scoring | Back-catalogue re-scan | |---|---|---|---| | What triggers it | staleness or a model change | a user uploading a clip | a policy or model change | | Scales with | run length and model size | upload volume | catalogue size | | Latency budget | a deadline, in hours or days | user-facing, milliseconds | none beyond a completion date | | Forecastable from | the release calendar | a traffic model | neither | | Can be paused midway | not usefully | never | yes, and resumed | The last two rows are the ones worth saying out loud in a design round. **Nothing forecasts it** — that is the budgeting problem. **It can be paused** — that is the lever. ## The lever: it has no user waiting on it Every other forward pass in this system has someone waiting. A re-scan has a completion date instead of a latency budget, and that difference is what makes the spend tractable: 1. **Scope it.** Re-scan only the clips the policy could plausibly touch — a category, a locale, a date range — rather than the whole catalogue by reflex. This is usually the largest single saving available and it is free. 2. **Skip what cannot have changed.** Key past verdicts by the content hash of the clip and the version of the model and policy that produced them. A clip whose hash is unchanged and whose deciding policy is unchanged does not need re-scoring; only the intersection with the edited policy does. 3. **Fill the trough with it.** The live scoring tier is provisioned for the evening peak and idle overnight. Scheduling re-scan work into that trough consumes capacity that is already being paid for, which makes part of the re-scan effectively free. 4. **Accept interruption.** Because there is no user waiting, the work can run on cheaper reclaimable capacity and simply resume; it needs restartable batches and a durable cursor, not a latency guarantee. ## Where this goes wrong in practice - **It is run on the live scoring tier without extra capacity**, and upload scoring degrades exactly when a policy team is watching. The two workloads compete for the same accelerators; give them separate pools or an explicit priority, and state which one loses. - **It is triggered by reflex.** Every policy wording change fires a full catalogue pass because no one has the scoping mechanism, so the largest compute event in the system has the lowest decision bar. - **It is absent from the forecast** because the forecast is built from a traffic model, and the re-scan is not traffic. - **It is repeated** because the previous pass's verdicts were never keyed by model and policy version, so nothing knows what has already been decided under the current rules. ## What to say when the interviewer asks what it costs Give the three numbers and the shape: catalogue size times per-item compute, expressed against both a retrain and a quarter of live scoring so the magnitude lands; then say that it is budgeted per policy change rather than per month, and that its only real cost control is scoping and deduplication, because unlike the live path it has no latency to trade away.

  • Why not just run the re-scan on the live scoring tier and let it soak up spare capacity?
    You can, but only with an explicit priority. The two share the same accelerators, so an unthrottled re-scan will queue behind or in front of user uploads and degrade the path that has a latency budget. Separate pools, or a strict priority with the re-scan yielding, makes the trade-off deliberate rather than accidental.
  • What has to be recorded during a re-scan so the next one can be cheaper?
    The verdict keyed by the content hash of the clip together with the model version and the policy version that produced it. That triple lets the next pass ask what has already been decided under the current rules and re-score only the genuine intersection with the edit.
  • How should this bucket appear in a capacity forecast?
    As an event-driven line, not a monthly average. Forecast it as expected policy changes per year times the scoped clip count per change, and hold the resulting compute as a stated allowance; averaging it into the monthly serving figure hides the peak that actually has to be provisioned or deferred.

saying these in an interview costs you the question

  • A re-scan is just more serving traffic on a busy day.
  • It is a one-off, so it need not be budgeted.
  • Every policy wording change requires a full catalogue pass.
  • Run it on the live scoring tier; there is always slack.
  • Nothing needs recording, since the next pass will redo it anyway.