When does ensembling independent LLM samples stop improving accuracy?
answer
- cancelling needs independence
- agreement can just be shared bias
- spread measures variance, never bias
- temperature is the weakest diversity lever
- compare best-of-N against the voted answer
basics
~20 sAggregation only cancels errors that are independent. Samples drawn from one model on one prompt fail in the same direction, so five runs agreeing means the model is consistent, not correct — and more samples then only sharpen a biased estimate.
solid answer
~50 sVoting or averaging over N runs helps under two conditions: each run must be better than chance, and the runs' errors must be at least partly uncorrelated. Take five independently sampled demand forecasts aggregated by median — the median is robust to one wild run, but if all five come from the same model with the same prompt and the same context, they inherit the same prior and cluster tightly around the same wrong number. The tight cluster then reads as confidence when it is really shared bias. Diversity has to be engineered: temperature and seed give the shallowest variation, prompt and framing variants more, genuinely different base models the most. I measure this rather than assume it — track per-sample accuracy, the aggregate's accuracy, and the gap between the best-of-N and the voted answer, which tells you how much aggregation is leaving on the table.
code
python · 12 linesfrom statistics import median
def aggregate(samples):
return median(samples)
# five independently sampled demand forecasts, one of them pathological
forecasts = [1180.0, 1205.0, 1195.0, 1620.0, 1190.0]
print(aggregate(forecasts))
# but a shared context gap biases every sample the same way
biased = [1180.0, 1205.0, 1195.0, 1200.0, 1190.0] # truth was 1500.0
print(aggregate(biased))go deeper
Know what ensembling is — run the task several times and combine the results by majority or median — and that combining only helps when the individual runs are usually right. Recognising that identical prompts produce similar mistakes is the key idea to carry.
Explain the two preconditions, better-than-chance samples and uncorrelated errors, and name where diversity comes from, ranking temperature below prompt variation and model variation. Be able to say why a median is used for numbers and a vote for labels.
Show the diagnostic instinct: measure per-sample accuracy, plot the aggregate as N grows, and compare best-of-N against the voted answer to separate a selection failure from a capability failure. Expect to explain how an ensemble can manufacture false confidence in a production system.
Own the spend decision — N-way sampling multiplies cost linearly with a curve that flattens, so argue for selective application by stake or confidence, and for spending the same budget on grounding or a stronger model when best-of-N shows the answer is rarely generated at all.
## The mechanism, stated properly Ensembling runs the same task N times and collapses the results: majority vote for a discrete label, median or trimmed mean for a number, a judge or an aggregator agent for free text. The intuition borrowed from classical statistics is that independent errors point in different directions and partly cancel, so the aggregate is more accurate than a typical single sample. That intuition rests on two assumptions, and both are load-bearing. **Each sample must be better than chance.** If the per-sample probability of being right is below one half on a binary decision, majority voting drives accuracy down, not up, and more samples make it worse. The ensemble amplifies whatever the base rate is; it does not create competence. **Errors must be at least partly independent.** This is the assumption that fails in practice with language models. Two samples from the same model on the same prompt are not two independent observers. They share weights, training data, and — crucially — the framing of the prompt and the contents of the retrieved context. When the model has a systematic bias, every sample carries it. ## The demand-forecast case Consider a system that samples five demand forecasts for next quarter and aggregates them by median. The median is a genuinely good aggregator for numeric output: it is robust to one run that produces an absurd figure, which a mean is not. So the design is sound as far as it goes. Now suppose the retrieved context omitted a promotional campaign. All five samples read the same context, all five miss the same driver, and all five come back low by roughly the same amount. The five figures cluster tightly. Downstream, the tightness is reported as agreement and therefore as confidence — and the ensemble has manufactured false certainty about a number that is systematically wrong. Adding five more samples narrows the cluster further and changes nothing about the error. The lesson generalises: **spread across samples measures the model's variance, never its bias.** An ensemble can only fix variance. ## Where diversity actually comes from Ranked from weakest to strongest source of decorrelation: - **Temperature and seed.** The cheapest lever and the shallowest. It perturbs the sampling path, so it catches errors that are artefacts of one unlucky decoding, but leaves every systematic prior intact. - **Prompt and framing variants.** Asking the same question three different ways — different decomposition, different output format, different emphasis — decorrelates more, because framing drives a large share of model behaviour. - **Different context or tool paths.** Letting each run retrieve independently, or approach the problem with different tools, decorrelates the inputs as well as the reasoning. This is where an omitted-driver error can actually be caught. - **Different base models, ideally from different providers.** The strongest available decorrelation, because the shared training data and shared architecture bias is broken. It is also the most operationally annoying: different formats, different failure modes, different costs. If you take one thing into an interview: diversity has to be designed in, and the cheap lever is the weak one. ## Choosing the aggregator The aggregation rule should match the output type. Majority vote for discrete labels. Median or trimmed mean for numbers, because both resist a single outlier run. For free-form text there is no arithmetic aggregator, so you either pick a representative sample by similarity to the others, or hand the candidates to a judge or aggregator agent — which introduces a new component with its own failure modes. And where a checker exists — the code runs, the constraint holds, the total reconciles — filter with the checker before aggregating anything, because verification beats voting outright. ## Measuring instead of assuming Three numbers make this concrete on your own data: - **Per-sample accuracy**, so you know whether you are above the threshold where voting helps at all. - **Aggregate accuracy at N**, plotted as N grows. The curve flattens; find where, and stop there rather than at a round number. - **Best-of-N (oracle) accuracy versus the aggregate.** A large gap means the right answer was in the pool and the aggregation rule failed to select it — a selection problem, fixable with a verifier or a better judge. A small gap means the right answer was rarely generated at all — a capability or grounding problem, and no amount of sampling will fix it. That last diagnostic is the strongest thing you can say on this topic, because it tells you whether to invest in aggregation or abandon it. ## Cost Ensembling multiplies cost by N with no compounding benefit once the curve flattens. Unlike sequential debate, the samples parallelise, so latency is roughly one call plus aggregation — which is why ensembling is often the cheaper accuracy lever in wall-clock terms even when it is the more expensive one in tokens. Selective application, where only low-confidence or high-stakes requests get the full N, keeps the bill proportional to the value at risk.
- How would you tell whether your ensemble's ceiling is a selection problem or a capability problem?Compare best-of-N accuracy against the aggregated answer's accuracy. If best-of-N is much higher, the correct answer is being generated and thrown away — a selection failure, worth fixing with a verifier or a better aggregator. If best-of-N is barely above a single sample, the model rarely produces the right answer at all, and more sampling is wasted money; invest in grounding, retrieval or a stronger model instead.
- Why prefer a median over a mean when aggregating numeric samples?The median is robust to a single pathological run. One sample that returns a figure an order of magnitude off drags a mean badly while barely moving a median. A trimmed mean is a reasonable middle ground when you want to use more of the distribution. None of this helps when every sample is biased the same way — robustness handles outliers, not shared error.
- Does raising temperature to force disagreement improve the ensemble?Only marginally, and it cuts both ways. Higher temperature decorrelates the decoding path but also lowers per-sample accuracy, and voting only helps while samples stay better than chance. Push it far enough and you degrade the aggregate. Decorrelating through prompt variants, independent retrieval, or different base models raises diversity without paying for it in per-sample quality.
Five thermometers with the same manufacturing defect will agree closely and all read three degrees low. Their agreement measures their similarity, not the temperature.
saying these in an interview costs you the question
- Treats tight agreement across samples as evidence the answer is correct
- Assumes samples from one model on one prompt are independent
- Believes more samples always improve the aggregate regardless of per-sample accuracy
- Relies on temperature alone as the source of ensemble diversity
- Votes on tasks where a cheap deterministic checker could filter the samples first