A team proposes retraining a policy classifier yearly rather than quarterly to cut the accelerator bill — what caps that saving?
answer
- compute the share before optimising
- Amdahl's Law, applied to spend
- under 3% of compute is the ceiling
- staleness is charged to another budget
- attack the ninety-seven percent
basics
~20 sThe training half's share caps it. At roughly 2,400 of about 88,800 accelerator-hours a year, training is under 3% of compute, so dropping from four retrains to one saves at most about 2% — while the staleness it buys is paid on the serving side and outside the accelerator bill entirely.
solid answer
~40 sThis is Amdahl's Law applied to spend: the ceiling on any optimisation is the share of the whole it can touch. Four retrains at 600 accelerator-hours is **2,400 hours a year** against roughly **86,400 hours** of live scoring — about **2.7%** of compute. Going to one retrain a year removes three of those runs, **1,800 hours**, or **2.0% of the annual bill**, and that is the best case before anything goes wrong. Measured against *paid* rather than consumed serving hours the share is smaller still. Meanwhile the cost of a staler model lands elsewhere: more missed and more wrongly-flagged clips per upload, which is paid in harm and in human review, not in accelerator-hours. A 2% ceiling is not worth that.
go deeper
The first step is not an opinion but a division: what fraction of the bill is the training half? That fraction is the most the proposal can ever save.
Show the ceiling explicitly — annual training hours over total hours — and note that a peak-provisioned serving fleet makes the denominator larger and the ceiling smaller.
Argue both ledgers: the capped compute saving against the per-clip error a staler model adds, which is charged to review capacity and to users rather than to accelerators.
Own the framing that the cheapest half is rarely worth the engineering quarter, and name the conditions — stable inputs, scarce human effort, release risk — under which a slower cadence is right for reasons other than money.
## The ceiling is the share, not the idea Before arguing about whether a yearly retrain is enough, settle what winning would be worth. The saving from optimising any one component is capped by that component's share of the total — the cost form of **Amdahl's Law**. If the training half is 2.7% of compute, then no change to the training half, however aggressive, can return more than 2.7%, and deleting it entirely still leaves 97.3% untouched. ## The arithmetic On the worked system — 600 accelerator-hours per retrain, 40 ms per clip, about 1.94 billion clips a quarter: | | Accelerator-hours per year | Share | |---|---|---| | Four quarterly retrains | 2,400 | 2.7% | | Live per-clip scoring | 86,400 | 97.3% | | Total | 88,800 | 100% | Dropping to one retrain a year removes **1,800 hours**, which is **2.0%** of the annual total. Two corrections make it smaller: - Measured against *paid* serving capacity rather than consumed compute — a peak-provisioned fleet bills roughly twice what it consumes — the training share falls to about 1.4%, and the saving with it. - A back-catalogue re-scan triggered by a single policy change is around 44,400 hours on its own. One such event in the year makes the training half about 1.8% of compute before the cadence change is even considered. ## The cost that moves the other way Retraining cadence is a genuine money lever — it is one of the few things that moves the fixed half at all — but it is not free on the other side of the ledger, and the costs it creates are not denominated in accelerator-hours: - **A staler model is a less accurate model** on drifting inputs, and in a moderation tier that error is paid per clip: more violations missed, more benign clips flagged. Those land in harm and in human review capacity, which is a different and usually more expensive budget. - **The gap between a policy change and its enforcement grows.** With a yearly cadence, a policy the model has never been trained on can stand unenforced for months unless something else covers it. - **Each run gets bigger and riskier.** Fewer, larger jumps between model versions make a regression harder to attribute and a rollback coarser, because more changed at once. So the honest framing is not "2% saving versus nothing", it is "2% of the accelerator bill, against a quality cost charged to a budget that is not this one". ## Where the same effort pays more The arithmetic does not say "do not optimise". It says point the effort at the 97%: 1. **Per-clip service time.** Taking 40 ms to 30 ms saves about 5,400 accelerator-hours a quarter — nine retrains' worth, every quarter — and it lowers the capacity the peak requires as well. How a model is made cheaper per clip is a modelling subject; the arithmetic only says that is where the money is. 2. **Work not done at all.** Skipping re-scoring for clips whose content and deciding policy are unchanged removes forward passes rather than shortening them, and removal beats optimisation. 3. **Idle capacity.** A peak-provisioned fleet can be paying for roughly twice the compute it consumes; that gap is larger than the entire training half several times over. ## When stretching the cadence is actually right There are honest reasons to retrain less often, and none of them is the accelerator bill: - inputs are **genuinely stable**, and measurement shows quality does not decay between runs; - each run consumes scarce **human** effort — labelling, review, sign-off — which is the real constraint rather than the compute; - the release process around a model version carries **risk** that is worth taking less often. And there are reasons to move the other way: on a fast-drifting problem, retraining more often can *reduce* total cost by cutting the per-clip error rate that the expensive half is paying for. ## The answer in one move Compute the share before debating the change. Two numbers — the training half's annual hours and the serving half's — settle whether the proposal can matter at all, and here they say the ceiling is about 2%. Then check whether the saving is even net positive once the quality cost is counted, and it usually is not.
- Is there a case where retraining more often lowers total cost?Yes, on a fast-drifting problem. Extra runs add to a small fixed half while a fresher model lowers the per-clip error rate, and those errors are paid on the large side of the ledger — in human review and in harm. If the error reduction is real and measured, more frequent retraining can be net cheaper.
- The proposal saves 2% of compute. Why is that not simply free money?Because the 2% is a gross figure that ignores what a staler model costs per clip. The error it adds is charged to review capacity and to users, not to the accelerator bill, so the change can look positive on one budget while being negative overall.
- What would change the answer to this proposal?A much larger training half — a run consuming tens of thousands of accelerator-hours, or a sweep of many runs per release — could push the fixed share into double digits, at which point cadence becomes a genuine lever. Recompute the share rather than assuming the ratio transfers between systems.
saying these in an interview costs you the question
- Retraining is the expensive part, so retrain less often.
- A 2% saving is free money since nothing else changes.
- Model staleness costs nothing until quality alarms fire.
- Optimise whichever half is easiest to change.
- The same training-to-serving ratio holds for any ML system.