How would you split extra compute across pretraining, RL post-training and test time?
answer
- three curves, not one
- which failure class does each fix
- capital cost versus per-request cost
- only one of them can be applied selectively
- reward design binds before FLOPs do
basics
~20 sTreat them as three curves with different economics. Pretraining raises the base capability, reinforcement learning after it sharpens behaviour on checkable tasks, and test-time compute buys accuracy on the hard request — but you pay that one on every call, forever.
solid answer
~50 sStart by diagnosing where the errors come from, because each regime fixes a different failure. If the model lacks knowledge or basic competence, that is a pretraining problem and no amount of deliberation at inference will fix it. If it demonstrably *can* do the task but does so unreliably or in the wrong format, that is what reinforcement learning against verifiable rewards is for. If it succeeds when it works longer on a problem, that is test-time compute. The economics differ sharply: pretraining and RL are one-time costs amortised across every future request, while test-time compute is a recurring per-request charge with a latency bill attached. So the same quality gain has very different lifetime prices depending on which curve you buy it from. Be honest that the relative returns are genuinely unsettled in 2026 — and that RL is usually bottlenecked on verifiable tasks and reward design, not on FLOPs.
go deeper
Know that compute gets spent in more than one place — training the model, refining it afterwards, and letting it think longer per request — and that these are separate decisions.
Be able to say what each regime fixes: pretraining supplies base knowledge, reinforcement learning improves reliability on checkable tasks, and extra inference-time work helps when the model can solve a problem given more effort.
Demonstrate diagnosis. Attribute production failures to a regime before proposing spend, and separate one-time training cost from recurring per-request cost when arguing for a fix.
Own the allocation as a strategy call: refuse to rank the regimes in the abstract, argue from a failure-class distribution and a traffic forecast, name the non-compute bottleneck in each, and state plainly where the public evidence runs out.
## Why there are three curves now The original scaling story was one curve: loss against pretraining compute. That framing has been superseded. Compute is now discussed as three layered regimes, each with its own inputs, its own returns and its own cost structure: 1. **Pretraining compute** — next-token prediction over a huge corpus. Sets the base of everything: world knowledge, language, latent skills. 2. **Reinforcement-learning post-training compute** — training after pretraining against a reward, most productively where outcomes can be checked automatically (code that passes tests, maths with a known answer, a tool call that succeeds). 3. **Test-time compute** — letting the model do more work on a single request before answering: longer deliberation, sampling multiple attempts, verifying and revising. They are layered, not alternatives. RL sharpens capabilities that pretraining put there; test-time compute exploits behaviours that RL taught the model to use productively. A model that was never trained to use extra deliberation well does not get much from being given more of it. ## Diagnose before allocating The allocation question is unanswerable in the abstract and quite tractable once you look at failures. A rough triage: - **The model does not know the thing.** Missing domain facts, an unfamiliar language, a niche codebase. This is a data-and-pretraining problem. Neither RL nor deliberation invents knowledge that was never present. - **The model can do it but does not do it reliably.** It succeeds sometimes, fails in ways that a checker could catch, wanders off format, gives up early. This is the RL regime's home ground, provided you can *score* success automatically. - **The model succeeds when it works longer.** Accuracy climbs visibly with more deliberation or more sampled attempts. Test-time compute is buying real accuracy here. - **The model succeeds but the harness fails it.** Bad tool descriptions, missing context, poor retrieval. This is an engineering problem masquerading as a compute problem, and it is by far the cheapest one to fix. It is also the most commonly misdiagnosed. ## The cost asymmetry This is the part that separates a principal-level answer. Pretraining and RL are capital costs: paid once, amortised across every request the model will ever serve. Test-time compute is an operating cost: paid on every single call, in money and in latency, forever. The consequence is that a quality improvement worth a fixed amount has wildly different prices depending on which curve you buy it from. At high traffic, moving a capability from test time into the weights — by training the behaviour in rather than eliciting it per request — is worth a large one-time investment. At low traffic or during exploration, the reverse holds: test-time compute needs no training run, ships immediately, and can be dialled per request class, which makes it the right first move for anything you might change your mind about. Test-time compute also has a distinctive property: it can be applied *selectively*. You do not have to pay it uniformly. Hard requests get more, easy ones get less. Nothing in the training-side regimes offers that granularity, and it substantially changes the effective economics. ## Where each regime actually binds None of the three is purely FLOP-limited: - Pretraining increasingly binds on the supply of high-quality tokens rather than on accelerators. - RL binds on **verifiable tasks and reward design**. Building a large set of problems whose success can be judged automatically, without exploitable shortcuts, is the hard part; reward hacking is the standard failure. Adding rollout compute to a badly specified reward makes the model better at gaming it. - Test-time compute binds on latency tolerance and on returns that flatten: past some depth of deliberation, extra thinking stops converting into accuracy, and on some tasks it actively hurts by talking the model out of a correct first answer. ## Be honest about the uncertainty The relative returns across the three regimes are not settled. RL post-training at scale is young enough that public evidence is thin and mostly vendor-reported. Test-time scaling curves are strong on verifiable domains — competition maths, code — and much weaker on open-ended judgement work where there is nothing to check against. Anyone offering you a clean allocation formula is overselling. What you can defend is the method: diagnose the failure class, cost the fix against traffic volume, prefer the cheapest regime that addresses the actual defect, and measure. ## What an interviewer wants to hear Not a ranking. They want to see you refuse to answer in the abstract, ask what is failing, distinguish capital from operating cost, note that test-time compute is the only one you can apply selectively per request, name the non-FLOP bottleneck in each regime, and admit where the evidence runs out.
- When is test-time compute clearly the wrong place to spend?When the failure is a knowledge gap rather than an effort gap — deliberating longer about a fact the model never learned only produces a more elaborate wrong answer. Also when latency is a hard product constraint, when traffic volume is high enough that a recurring per-request charge dwarfs a one-time training investment, or when measurement shows accuracy flat or declining with more deliberation, which happens on tasks where the first instinct was already right.
- What binds RL post-training before compute does?The supply of tasks whose outcome can be checked automatically, and the quality of the reward that checks them. Building large sets of verifiable problems is slow, and any reward with an exploitable shortcut gets found — the model optimises the metric rather than the goal. Adding rollout compute to a poorly specified reward buys faster reward hacking, not better behaviour. Verifier coverage is usually the binding constraint.
- How does this framing change what you monitor in production?You stop tracking a single quality number and start attributing failures to a regime. Split errors into knowledge gaps, reliability and format failures, insufficient-deliberation cases, and harness defects, and track their shares over time. That distribution is the actual input to the allocation decision, and it moves — a fix in one regime shifts the mix, so the right next investment changes with it.
saying these in an interview costs you the question
- Ranks the three regimes without asking what is failing
- Treats test-time compute as a substitute for missing knowledge
- Ignores that inference-time cost recurs on every request
- Assumes RL post-training is limited mainly by available FLOPs
- Presents the relative returns as a settled, known ordering