skip to content

How does variance reduction cut Monte Carlo error without buying more draws?

level: seniorimportance: nice to knowfreq 26%

answer

  1. attack the spread, not the sample size
  2. mirror each uniform draw
  3. subtract something whose mean you know
  4. one minus the squared correlation
  5. share one stream across compared designs

basics

~20 s

Variance reduction reshapes what you average rather than how much. Antithetic variates pair each uniform U with 1-U so their errors cancel; control variates subtract a correlated quantity whose true mean is known; common random numbers reuse one stream across competing designs.

solid answer

~50 s

Since Monte Carlo error is the per-draw spread divided by `sqrt(n)`, you can attack the numerator instead of the denominator. **Antithetic variates** run each simulation twice, once on a uniform draw `U` and once on `1-U`. When the simulated output moves monotonically with `U`, the pair is negatively correlated, so averaging them has less variance than two independent runs. **Control variates** use a quantity `Y` that is correlated with the target and whose true mean is known exactly; the estimator becomes `X - c*(Y - E[Y])`, and with the optimal coefficient the variance is multiplied by `1 - rho^2`, so a correlation of 0.9 removes about 81% of it. **Common random numbers** compare two designs on the identical seeded stream, so the shared noise cancels out of their difference. Each buys precision that would otherwise cost a multiple of the compute, and each has a failure mode worth stating.

go deeper

for a junior

Recall that precision comes from more runs, and that clever pairing of random draws can substitute for some of them. Knowing the names antithetic and control variates is enough here.

for a middle

Explain the covariance term in the variance of an average and why mirrored uniforms help only when the output moves monotonically with the driving randomness.

for a senior

Choose and validate a technique on a real simulation: check monotonicity before pairing, source a control's mean from theory, and measure the achieved spread against the plain estimator on a pilot run.

for a principal

Own the compute policy. Argue when engineering effort on variance reduction beats a hundredfold compute bill, and insist that comparisons across configurations always share a random stream so decisions are not made on seed luck.

## Why bother Monte Carlo error is roughly `sigma/sqrt(n)`, where `sigma` measures how much a single simulated run varies. Buying accuracy through `n` is brutal: four times the compute to halve the error. Variance reduction attacks `sigma` instead, and a technique that halves `sigma` is worth exactly as much as quadrupling the run count, usually for a fraction of the cost. ## Antithetic variates Simulations are driven by uniform random numbers. Run the simulation once on a stream of uniforms `U`, then run it again on the mirrored stream `1-U`, which is equally valid uniform randomness. Average the two outputs into one paired observation. The variance of an average of two variables is Var((X1 + X2)/2) = (Var(X1) + Var(X2) + 2*Cov(X1, X2)) / 4 With independent runs the covariance is zero and you get half the variance of a single run, exactly what two draws should buy. With antithetic pairs the covariance is negative whenever the output is a monotone function of the driving uniforms, so the paired variance is strictly less than that, and the same compute buys more precision. Monotone dependence is common: a longer service time makes a queue longer, a higher uniform makes an inverse-transform draw larger. The failure mode is precise and worth naming. If the output is a symmetric function of `U`, so that `f(U) = f(1-U)`, the pair is perfectly positively correlated and the two runs contribute exactly as much information as one. You have paid for two runs and received one. For strongly non-monotone outputs, antithetic pairing can therefore increase variance per unit of compute. ## Control variates Suppose alongside the quantity of interest `X` you can also compute, in the same simulated run, a related quantity `Y` whose true expectation `E[Y]` you know exactly from theory. Then Z = X - c*(Y - E[Y]) has the same expectation as `X` for any constant `c`, because the correction term has mean zero. Its variance is Var(Z) = Var(X) - 2*c*Cov(X, Y) + c^2 * Var(Y) Minimising over `c` gives the optimal `c* = Cov(X, Y)/Var(Y)`, and substituting it yields Var(Z) = Var(X) * (1 - rho^2) where `rho` is the correlation between `X` and `Y`. So the reduction depends entirely on correlation: 0.5 removes 25% of the variance, 0.9 removes 81%, 0.99 removes 98%. Weakly correlated controls are barely worth the code. Two practical points. `E[Y]` must be known exactly, not estimated from the same run, or the correction reintroduces the very noise it is meant to cancel. And `c*` is itself estimated from the simulated output, which adds a small bias at tiny sample sizes; with thousands of runs it is negligible, and a pilot run can supply `c` for the main run to avoid the issue entirely. ## Common random numbers When the question is not "what is the win rate" but "is design B better than design A", simulate both on the identical seeded stream of random numbers. Whatever luck the stream contains hits both designs, so it largely cancels in the difference `B - A`, and the comparison becomes far sharper than two independently seeded runs would allow. Formally, the variance of a difference is `Var(A) + Var(B) - 2*Cov(A, B)`, and positive induced correlation shrinks it. This is where seeding stops being merely a reproducibility habit and becomes a statistical tool. Note the distinction plainly: a fixed seed on a single run buys reproducibility and nothing else; the same seed deliberately shared across two compared configurations buys precision on their difference. The technique requires the two designs to consume the random stream in the same order, which in practice means giving each source of randomness its own generator so a change in one design does not shift every downstream draw. ## Importance sampling, briefly When the quantity of interest is dominated by rare outcomes, sampling from the natural distribution wastes almost every run on irrelevant cases. Importance sampling draws instead from a distribution that visits the interesting region often, then reweights each draw by the ratio of the true density to the sampling density to keep the estimate unbiased. It can cut the variance of a rare-event estimate by orders of magnitude, and it can also blow the variance up spectacularly if the sampling distribution has lighter tails than the target, which makes a few enormous weights dominate the average. ## Choosing among them Try the cheapest applicable technique first. Common random numbers costs almost nothing when comparing configurations and is nearly always right. Antithetic variates costs one line for a monotone simulation. Control variates need a known-mean companion quantity, which is a modelling insight rather than a code change. Importance sampling needs care and a good proposal, and belongs to rare-event work. Always verify the reduction empirically by comparing the spread of the reduced estimator against the plain one on a pilot run, since a technique applied to the wrong integrand can quietly make things worse. ## What to say out loud Explain that error is spread over `sqrt(n)` and that these techniques attack the spread, give antithetic pairing with its monotonicity condition, give control variates with the `1 - rho^2` formula, and distinguish reproducibility seeding from common random numbers.

  • When do antithetic variates fail to reduce variance?
    When the simulated output is not monotone in the driving uniforms. In the worst case the output is symmetric, so `f(U)` equals `f(1-U)`, the pair is perfectly positively correlated, and two runs deliver the information of one. Check the monotonicity assumption, or measure the paired spread against the independent one on a pilot before adopting it.
  • How much variance does a control variate correlated 0.9 with the target remove?
    With the optimal coefficient the variance is multiplied by `1 - rho^2`, so `1 - 0.81` leaves 19% and about 81% is removed. Matching that with extra draws alone would take roughly five times the runs. Correlation 0.5 removes only a quarter, which is often not worth the extra code and the risk of a mis-specified control.
  • Is fixing the random seed itself a variance-reduction technique?
    Not on its own. A seed on a single run buys reproducibility: the same figure can be regenerated, with the same error it always had. It becomes variance reduction only when the same stream is deliberately shared across two configurations being compared, because the shared randomness then cancels out of their difference.
  • Why must the control variate's mean be known exactly rather than estimated?
    The correction term `Y - E[Y]` works because it has mean zero and contributes no bias. Substituting an estimate of `E[Y]` computed from the same runs makes the correction correlated with its own error, which reintroduces noise and can bias the result. Use a mean known from theory, or one obtained from an independent pilot run.

Weighing two nearly identical parcels on the same slightly miscalibrated scale gives a very accurate difference, even though each individual reading is off. Sharing the error source makes the comparison sharp.

saying these in an interview costs you the question

  • Thinks antithetic pairing always reduces variance
  • Uses a control variate whose mean is estimated from the same run
  • Calls a fixed seed a variance-reduction technique on its own
  • Expects large gains from a weakly correlated control
  • Claims variance reduction makes the estimator biased

context