How do you decide how many draws a Monte Carlo simulation needs?
answer
- start from the decision's tolerance
- error falls with the square root of n
- halving error is not doubling draws
- four times the draws, half the error
- extra draws never fix a wrong model
basics
~20 sPick the precision the decision actually needs, then buy draws to reach it. Simulation error falls like 1/sqrt(n), so halving the error costs four times the draws and one extra decimal digit costs a hundred times.
solid answer
~50 sWork backwards from the decision. Decide the tolerance first: if a launch call flips at a five-point difference in win probability, an error of half a point is plenty and an error of two points is not. Then price it. The Monte Carlo error of an average over `n` independent runs shrinks in proportion to `1/sqrt(n)`, so the arithmetic is fixed: four times the draws to halve the error, a hundred times the draws to gain one decimal digit. For a simulated probability near 0.6, ten thousand runs gives a margin of roughly half a percentage point. Estimate the spread from a short pilot run, solve for `n`, then run it. Report the Monte Carlo margin alongside the number, and remember that draws only shrink simulation noise: they do nothing about a wrong model or biased inputs.
go deeper
Recall that more draws mean less noise and that the payoff is slow: four times the runs to halve the error. Never quote a simulated number as if it were exact.
Do the sizing arithmetic out loud. From a pilot run's spread, solve spread over sqrt(n) below tolerance for n, and show the 4x and 100x consequences of the square-root rate.
Separate the three error sources when reporting: simulation noise you buy away with draws, input uncertainty you address by re-running across plausible parameters, and model error you address by validation.
Own the compute-versus-precision tradeoff for the organisation. Decide when a half-point answer today beats a tenth-point answer next week, and push teams toward variance reduction or an analytic shortcut before approving a hundredfold compute bill.
## Draws are a purchase, not a default The most common weak answer to "how many runs" is a round number chosen by habit: ten thousand because it sounds serious, a million because the machine can. The professional answer names a tolerance and derives the count. ## The rate that prices everything A simulation estimate is an average of `n` independent runs of the same random experiment. Its typical error is proportional to `1/sqrt(n)`, with the constant of proportionality set by how variable a single run is. Everything about simulation budgeting follows from that square root. - Doubling accuracy costs 4x the runs. - Ten times the accuracy costs 100x the runs. - Going from a 1% margin to a 0.1% margin on an overnight job turns it into a hundred-night job. This is why simulation is described as easy to start and expensive to finish. A crude answer arrives in seconds; a precise one may not be affordable at all, which is a genuine reason to look for an analytic shortcut or a variance-reduction technique instead of more compute. ## Working the arithmetic For a simulated probability, the per-run quantity is a 0/1 indicator with variance `p(1-p)`, so the simulation error is `sqrt(p(1-p)/n)`. Suppose a strategy simulation returns a win probability near 0.62. At `n = 10,000`: sqrt(0.62 * 0.38 / 10000) = sqrt(0.0000236) = 0.0049 So roughly half a percentage point, and a rough 95% interval of about plus or minus one point. If the decision needs a quarter-point margin, that is four times the runs, 160,000. If it needs a tenth of a point, that is 100 times the original, a million runs. Note that `p(1-p)` is largest at `p = 0.5` and small near 0 or 1, so rare-event probabilities have a small absolute error but a terrible relative error: estimating a probability of 0.001 well enough to know it within 10% needs on the order of a hundred thousand runs before a single meaningful event pattern emerges. For a simulated average that is not a probability, run a short pilot of a few thousand iterations, measure the spread of the per-run values, and solve `spread / sqrt(n) <= tolerance` for `n`. Pilot-then-size is the standard workflow and costs almost nothing. ## Sequential stopping An alternative to fixing `n` up front is to monitor the running Monte Carlo margin and stop when it drops below tolerance. That is efficient, and it is legitimate here in a way it is not when peeking at experimental data, because the quantity being estimated is a fixed property of a known model and more draws only ever tighten the same target. Two cautions: the margin estimate is itself noisy early on, so impose a minimum run count before allowing a stop, and the stopping decision must be based on precision, never on the estimate having reached a number someone hoped for. ## What more draws cannot fix This is the point that separates a middle answer from a senior one. Draws shrink only the simulation's own sampling noise around the answer that the model implies. They do not touch: - **Model error.** If the simulated rules leave out a real mechanism, a billion runs converge precisely on the wrong number. - **Input error.** Parameters estimated from data carry their own uncertainty. A win rate simulated from a hit probability that is itself only known to plus or minus three points cannot be trusted to half a point, no matter how many runs are made. Propagating input uncertainty means re-running the simulation across plausible parameter values, not adding iterations at one fixed value. - **Bugs.** A miscoded rule is systematic, and averaging never removes a systematic error. A good habit is to report two numbers: the simulation margin, which you control with draws, and the sensitivity of the answer to the parameters you are least sure of, which you control only by better inputs. ## Reproducibility and reporting Record the seed, the number of runs, and the version of the model alongside the result. A simulated figure that cannot be regenerated is not evidence, and a figure quoted without its run count invites a reader to treat noise as signal. ## What to say out loud Name the tolerance the decision needs, quote the `1/sqrt(n)` rate and the 4x/100x consequences, size `n` from a pilot, report the margin with the number, and finish by separating simulation noise from model and input error.
- A simulation gives a 0.62 win probability from 10,000 runs; how precise is that?The simulation error is `sqrt(0.62*0.38/10000)`, about 0.005, so roughly plus or minus one percentage point at conventional 95% coverage. That is precise enough to distinguish 0.62 from 0.58 but useless for distinguishing 0.62 from 0.615. Quote the margin whenever you quote the number.
- Does adding draws help when the simulated model itself is wrong?No. Extra draws reduce only the noise around whatever answer the model implies, so a model missing a real mechanism converges more and more precisely on a wrong number. Model error is fixed by validating the simulated rules against observed outcomes, and input uncertainty by re-running across plausible parameter values.
- Is it acceptable to stop a simulation as soon as the running margin is small enough?Yes, provided the stopping rule watches precision and not the value. The target is a fixed property of a known model, so more draws only tighten the same quantity. Impose a minimum number of runs first, since an early margin estimate is itself noisy, and never stop because the estimate reached a hoped-for number.
saying these in an interview costs you the question
- Picks a round number of runs with no tolerance in mind
- Thinks doubling draws halves the simulation error
- Believes more draws can correct a biased model
- Quotes a simulated figure with no margin
- Stops the run when the estimate reaches a desired value