In a bagged classifier, when does averaging predicted probabilities beat majority voting?
answer
- threshold first, or combine first
- voting discards how sure each model was
- three at 0.51 versus two at 0.05
- averaging gives a movable threshold
- voting is safer when scores are not comparable
basics
~20 sAveraging predicted probabilities beats majority voting when the base models differ in confidence: a few strongly negative models can outweigh a bare majority of barely-positive ones. Voting collapses each model to one ballot and discards that information.
solid answer
~50 sMajority voting thresholds each base model's output to a label first, then counts labels; averaging combines the probabilities and thresholds once at the end. The two disagree whenever confidence is unevenly distributed. Take five bagged trees outputting 0.51, 0.51, 0.51, 0.05 and 0.05 for the positive class: three of five vote positive, but the mean is `(3*0.51 + 2*0.05)/5 = 0.326`, so averaging predicts negative. Averaging is the usual default because it preserves how sure each model was and produces a continuous score you can rank or re-threshold; voting throws that away and gives you only B+1 discrete levels. Voting earns its place when the base models emit labels only, or when their probability estimates are so badly miscalibrated that one wildly overconfident model would dominate the mean — a ballot caps every model's influence at one vote.
go deeper
Know both rules by name and what each produces: a count of labels versus a mean score. Remember that averaging is the more common default in bagged classifiers.
Be ready to construct a small example where the two disagree and explain the order of operations — thresholding before combining throws away confidence that combining could have used.
Show judgment about which rule fits the deployment. If the score feeds a ranked queue or a cost-sensitive cutoff you need the continuous output; if the base models' scores are not comparable, argue for the vote and say why.
Own the downstream consequences of the choice. A continuous ensemble score invites threshold tuning, monitoring and calibration work; a vote count is simpler to govern and explain. Decide which the organisation can actually maintain.
## Two ways to combine the same predictions A bagged classifier has B base models, each able to output a score for the positive class. There are two standard ways to turn B outputs into one decision. **Majority voting** (often called hard voting) asks each base model for a label — usually by comparing its score to 0.5 — and predicts whichever label wins the count. **Probability averaging** (soft voting) takes the mean of the B scores and applies the decision threshold once, to that mean. The order of operations is the whole difference. Voting thresholds first and combines second; averaging combines first and thresholds second. Thresholding is a lossy step, so doing it first destroys information that the combination step could have used. ## The disagreement, concretely Five bagged trees produce these positive-class probabilities for one row: 0.51, 0.51, 0.51, 0.05, 0.05. - **Voting:** three trees are above 0.5 and two are below, so the vote is 3 to 2 for the positive class. - **Averaging:** the mean is `(0.51 + 0.51 + 0.51 + 0.05 + 0.05) / 5 = 1.63 / 5 = 0.326`, comfortably below 0.5, so the prediction is negative. Which is right? Read what the models are saying. Three of them are on the fence — 0.51 is a coin flip that happened to land on the positive side. Two of them are close to certain the answer is negative. Averaging hears the certainty and the hesitation; the vote hears only five equal ballots. In most bagged ensembles the averaged answer is the better bet, and empirically soft aggregation is the more common default for exactly this reason. ## Why averaging is usually the default Beyond the accuracy of the single decision, averaging gives you a **continuous ensemble score**. That matters for three practical reasons. You can rank rows by risk, which any ranking metric needs. You can move the decision threshold after the fact when the cost of a false positive and a false negative differ. And the averaged score of many base models is smoother than any one model's score — an individual unpruned tree often reports a crude leaf frequency such as 0.0 or 1.0, and averaging hundreds of those crude numbers produces something far more granular than any single tree could offer. Majority voting, by contrast, produces only B+1 possible outputs: the count of positive votes, from 0 to B. You can use that count as a coarse score, but it is a quantised version of the information averaging already gives you cleanly. ## When voting is the right call There are real cases for it: - **The base models only emit labels.** Some learners have no natural probability output. If a label is all you have, a vote is all you can do. - **The scores are not comparable across models.** Averaging assumes the B numbers live on the same scale. If one base model is systematically overconfident — pushing everything to 0.99 or 0.01 — its opinions dominate the mean, while a vote caps it at one ballot. Voting is the more robust aggregation when calibration is doubtful, and in plain bagging the base models are all the same learner on similar data, so this is more of a worry when mixing model types. - **You need a defensible, explainable rule.** Twelve of twenty models agreed is easier to communicate to a non-technical reviewer than a mean score of 0.58. ## The regression counterpart For regression there is no thresholding step, so the question becomes which summary statistic to use. The mean is the default and the one the variance argument is built on. The median is the alternative when a small number of base models can produce wild values — it sacrifices a little efficiency for resistance to those outliers. Ties into the same theme: the aggregation rule is a small decision with real consequences, and it should be chosen rather than inherited by accident. ## Getting it right in an interview The trap is asserting that the two rules agree. They agree often, which is why the difference is easy to overlook, and they disagree exactly in the cases where the ensemble is uncertain — the rows where the decision matters most. Being able to construct a five-model example where they diverge, and to say which answer you would trust and why, is the whole content of the question.
- Why does averaging give you a threshold you can move and voting does not?Averaging outputs a continuous score, so you can set any cutoff you like after training to match the cost of each error type. Voting outputs a count from 0 to B, which gives only B+1 coarse levels and ties the decision to the base models' own internal 0.5 cutoffs.
- When is hard majority voting the safer choice?When the base models emit labels only, or when their scores are not on a comparable scale. A single overconfident model can drag an average a long way, while a vote caps every model's influence at one ballot. It is also the easier rule to explain to a reviewer who wants to see agreement counts.
- How do you aggregate a bagged ensemble for regression?Average the numeric predictions — the mean is what the variance-reduction argument is built on. The median is a reasonable alternative when a few base models can emit extreme values, trading a little statistical efficiency for resistance to those outliers.
Two experts who are certain the patient is fine outweigh three who are only just leaning the other way. Averaging hears the certainty; a show of hands counts five equal opinions.
saying these in an interview costs you the question
- Assumes averaging and majority voting always give the same label
- Calls averaging thresholded 0/1 labels soft voting
- Treats a single tree's leaf frequency as a calibrated probability
- Insists the majority must win regardless of confidence
- Cannot name a case where voting is preferable