skip to content

What is the difference between hard voting and soft voting in a classifier ensemble?

level: juniorimportance: should knowfreq 52%

answer

  1. One counts, the other averages
  2. Labels versus predicted probabilities
  3. Confidence can outvote a bare majority
  4. Averaging assumes a shared, calibrated scale

basics

~10 s

Hard voting takes each model's predicted class label and picks the majority. Soft voting averages the models' predicted class probabilities and picks the highest average, so a confident model outweighs several hesitant ones.

solid answer

~50 s

Hard voting counts labels: each model votes for one class and the majority wins, so a model that was 99% sure and one that was 51% sure count the same. Soft voting averages the predicted probabilities and takes the arg max, so confidence carries weight. They can disagree on the same row: with predicted probabilities of class 1 at `0.45`, `0.45` and `0.90`, hard voting says class 0 (two labels to one) while soft voting says class 1 (mean `0.60`). Soft voting usually scores better because it uses the margin rather than just its sign, and it gives you a probability for thresholding downstream. It relies on the members being reasonably calibrated, though — one systematically overconfident model can dominate the average. Weighted soft voting, with weights fitted on out-of-fold predictions, is exactly a stack with a linear meta-learner.

code

python · 10 lines
python
probs = [0.45, 0.45, 0.90]  # each model's P(class 1) for one row

votes_for_1 = sum(1 for p in probs if p >= 0.5)
hard = 1 if votes_for_1 > len(probs) / 2 else 0

mean_p = sum(probs) / len(probs)
soft = 1 if mean_p >= 0.5 else 0

print("hard voting ->", hard)              # 0: two of three labels are class 0
print("soft voting ->", soft, round(mean_p, 2))  # 1: mean probability is 0.60

go deeper

for a junior

Be able to state both rules in one sentence each and work a three-model example by hand. Know that soft voting needs probabilities and hard voting only needs labels.

for a middle

Explain why soft voting normally scores higher — it uses the margin, not just its sign — and name calibration as the assumption it rests on. Be ready to describe weighted voting.

for a senior

Show judgment about when to trust an average: which members are calibrated, what the tie rule is, and why plain soft voting is a baseline that a learned combination must beat before you take on its complexity.

for a principal

Frame voting as the zero-training combiner with nothing to overfit and no leakage surface, and argue when that safety is worth more to the organisation than the extra points a fitted blend might buy.

## Two ways to combine votes Suppose three trained classifiers each score the same row for a binary target. **Hard voting** asks each model for its final class label and takes the majority. Each model gets exactly one vote, and how sure it was is discarded. **Soft voting** asks each model for its predicted probability of each class, averages those probabilities across models, and picks the class with the highest average. A model that is 95% sure now pulls harder than a model that is 51% sure. ## Where they disagree Take three models whose predicted probabilities of class 1 are 0.45, 0.45 and 0.90. - Hard voting: the first two predict class 0 (both below the 0.5 threshold), the third predicts class 1. Majority is **class 0**. - Soft voting: the mean probability is (0.45 + 0.45 + 0.90) / 3 = 0.60, which is above 0.5, so the answer is **class 1**. Two lukewarm votes lose to one confident one. Whether that is right depends entirely on whether that confidence is earned. ## Why soft voting usually wins — and when it does not Soft voting normally scores better because it uses strictly more information: the margin, not just its sign. It also gives you a probability out the other end, which you need if downstream logic sets a threshold, ranks candidates, or feeds an expected-value calculation. Hard voting can only ever hand you a label. The catch is **calibration**. Averaging probabilities only makes sense if the numbers are comparable across models. A model whose scores are pushed toward 0 and 1 — many margin-based and heavily overfitted models behave this way — will dominate the average whenever it is confident, including when it is confidently wrong. Two defences: calibrate each base model's scores before averaging so that a stated 0.8 means roughly an 80% hit rate for all of them, or move to **weighted** soft voting, where each model's contribution is scaled by a weight you set or fit. Fitting those weights on out-of-fold predictions is exactly stacking with a linear meta-learner. Plain soft voting is the special case where every weight is fixed at 1/M and nothing is learned — which is why it is a strong, hard-to-beat baseline: there are no weights to overfit. ## Practical details worth knowing - **Ties.** With an even number of voters, hard voting can deadlock; you need a rule (fall back to the strongest single model, or to the soft-voted answer). Odd counts avoid the problem for binary targets but not for multiclass, where a three-way split can leave no majority at all — the usual convention is plurality. - **When hard voting is the only option.** Some models expose only a decision, or scores on an arbitrary, unbounded scale that cannot be averaged with probabilities. Either convert them to probabilities via calibration or vote on labels. - **Regression.** The analogue of soft voting is simply averaging the predicted values; a median is the robust variant when one member occasionally produces wild outputs. - **Neither one learns anything.** Voting has no training step of its own, so unlike stacking there is no meta-learner to leak into. That makes it the cheap and safe first thing to try, and the benchmark any learned combination should have to beat.

  • When is hard voting the only combination rule available to you?
    When some members expose only a decision rather than a probability, or produce scores on an unbounded, model-specific scale that is not comparable with the others'. Averaging incomparable numbers is meaningless, so you either calibrate each member's scores onto a common probability scale first, or fall back to voting on labels.
  • Why can one badly calibrated member ruin soft voting?
    Soft voting is an average, so a member that reports 0.02 and 0.98 where the truth is nearer 0.3 and 0.7 swings the mean far more than its accuracy deserves — including on the rows where it is confidently wrong. Fix it by calibrating each member's scores before averaging, or by down-weighting that member with weights fitted on out-of-fold predictions.
  • How do you break a tie in hard voting?
    With an even number of binary voters a tie is possible, and in multiclass problems even an odd count can split with no majority. Standard options are to take the plurality, fall back to the soft-voted answer, or defer to the single strongest member. Pick one deliberately rather than inheriting whatever the tie-break happens to be.

Hard voting is a show of hands; soft voting is asking each judge for a score out of ten and averaging. The second only works if the judges mark on the same scale.

saying these in an interview costs you the question

  • Says soft voting is always better, without mentioning calibration
  • Thinks soft voting averages the predicted labels
  • Assumes voting members must all be the same algorithm
  • Ignores ties when the number of voters is even
  • Believes voting has a training step that can leak

context