skip to content

With a fixed per-step simulator cost, how do you plan a 10M-step deep RL training budget?

level: principalimportance: should knowfreq 34%

answer

  1. two currencies, not one
  2. price a single environment step first
  3. compute cost or calendar cost
  4. the search eats more than the final run
  5. more actors raise throughput, not efficiency

basics

~20 s

Treat environment steps and wall-clock time as two separate budgets. Price one step, estimate steps-to-threshold from short pilot runs, reserve most of the budget for search, and pick the algorithm family by whichever resource is scarce.

solid answer

~50 s

Start by pricing a single environment step in seconds and in money, because that number decides everything else. Then estimate steps-to-threshold from short pilot runs and extrapolate the early learning curve rather than assuming the published figure transfers to your environment. Plan the budget as a portfolio: the final run is a minority of it, since configuration search consumes far more steps than the run you eventually report. Choose the family by which currency is scarce — an expensive or slow step argues for a replay-based learner with a high update-to-data ratio, while a microsecond step in a vectorised simulator argues for an on-policy method with many actors. Parallel actors buy wall-clock time, not sample efficiency: they raise steps per second, but the agent still needs the same number of steps, and they stop helping once the learner update rather than the simulator is the bottleneck.

go deeper

for a junior

Know the two units involved: environment steps consumed and wall-clock hours spent, and that adding parallel environment copies speeds up the clock without reducing the steps needed.

for a middle

Be able to convert a per-step simulation cost into a run duration, and explain why parallel actors raise throughput only until the learner update becomes the bottleneck.

for a senior

Show that you would pilot first and extrapolate steps-to-threshold in your own environment, book evaluation and search consumption explicitly, and pick the algorithm family from the measured per-step cost.

for a principal

Own the tradeoff between buying steps and buying compute, and be ready to say what you would cut — scope, threshold, or search breadth — when the pilot shows the budget cannot reach the target.

## Two budgets, not one A training plan has two independent currencies and confusing them is the classic planning mistake. - The **sample budget** is the number of environment steps the agent may consume — 10M in this case. It is fixed by what interaction costs: licence fees, hardware wear, engineer time, or the calendar. - The **wall-clock and compute budget** is how long the run takes and what it costs in accelerators. It is governed by steps per second and by the cost of each gradient update. A method can be excellent on one and terrible on the other. Deciding which one binds is the first act of planning, and it is what a principal-level answer leads with. ## Price one step Measure, do not guess: seconds of simulation per step, cost per step, and how that scales as you run copies of the environment in parallel. Ten million steps at fifty milliseconds each is about 139 hours of pure simulation on one copy; the same ten million at fifty microseconds is under ten minutes. The plan for those two worlds shares nothing. Also ask whether the per-step cost is *compute* or *calendar*. Consider a sensor-driven building-climate agent whose every environment step is a fixed five-minute control interval. Ten million steps is roughly ninety-five years of continuous operation on one unit. No amount of hardware compresses that, because the cost is time passing in the world. Either you run a fleet of units genuinely in parallel, or you learn from far fewer steps — and the algorithm choice stops being a preference and becomes a constraint. ## Estimate steps-to-threshold before committing Do not plan against a number from a paper on a different environment. Run short pilots, plot returns against environment steps, and extrapolate: RL learning curves are noisy but their early slope carries real information about whether the threshold is plausibly inside the budget. If a pilot at one tenth of the budget is nowhere near the shape of a curve that would arrive, the honest conclusion is that the budget is wrong or the problem formulation is, and both are cheaper to fix now. ## Budget the search, not just the run The single reported run is the small part. Configuration search, ablations and evaluation rollouts dominate consumption, and evaluation itself costs steps that must be booked. State the split up front — for instance a majority of the budget to search, a reserve to the final runs, and a fixed slice to evaluation — because an unbooked search silently eats the run you actually needed. ## Match the family to the scarce currency - **Expensive or slow step.** Use a replay-based off-policy learner and push the update-to-data ratio, with the ensemble and reset machinery that a high ratio requires. You are deliberately spending compute, which you have, to buy back steps, which you do not. - **Cheap, vectorisable step.** Use an on-policy method with many parallel actors. The step count will be larger and it will not matter, because throughput is the thing you can buy and the method is far less fussy to tune. ## Where parallel actors help and where they do not This is the part candidates most often get wrong. Running more environment copies raises **steps per second**; it does **not** reduce the number of steps the agent needs. If interaction is billed per step, more actors change the wall-clock schedule and nothing about the invoice. And the wall-clock benefit itself runs out. Doubling actors stops helping once the **learner update**, not the simulator, is the bottleneck — past that point the extra actors sit waiting on the next parameter broadcast. Synchronous collection also runs at the pace of the slowest actor, so heterogeneous episode lengths erode the speedup. There is even a sample-efficiency cost hiding in the throughput win: with more actors and a fixed number of updates per iteration, each transition contributes to fewer gradient updates, so a heavily parallelised run can need *more* total steps than a modest one. ## What a good answer sounds like Name the two budgets, price one step, say whether the cost is compute or calendar, get steps-to-threshold from a pilot instead of a paper, book the search explicitly, then choose the algorithm family from the scarce currency — and finish by stating what you would cut if the pilot said 10M steps is not enough. That last part is the judgment the question is looking for.

  • When does doubling the number of parallel actors stop shortening wall-clock time?
    Once the learner update or the parameter broadcast, rather than simulation, is the bottleneck — extra actors then idle waiting for fresh parameters. Synchronous collection also caps the speedup at the pace of the slowest actor, so uneven episode lengths eat into it. And nothing about extra actors reduces the number of environment steps required, so a per-step billed environment sees no saving at all.
  • Each step of a building-climate agent is a five-minute control interval. What does that do to the plan?
    It converts the interaction budget from a compute cost into a calendar cost: ten million steps is on the order of ninety-five years on a single unit. The only real levers are running a fleet of units as genuinely parallel actors and choosing the most sample-efficient family available, so that the step count needed drops by an order of magnitude rather than the step rate rising.
  • How would you decide the split between configuration search and final runs?
    Work backwards from steps-to-threshold measured in a pilot. Multiply it by the number of final runs you must produce, add the evaluation rollouts those need, and whatever is left is the search allowance — which then dictates how many configurations you can afford to try. If that number is one or two, the honest move is to cut the search space rather than pretend the budget covers it.

saying these in an interview costs you the question

  • Plans against a step count taken from a paper
  • Says more parallel actors improve sample efficiency
  • Books only the final run and forgets the search
  • Ignores whether the per-step cost is compute or calendar
  • Assumes throughput scales linearly with actor count

context