Why can a state-level correlation between income and vote share mislead about individual voters?
answer
- unit computed, unit claimed
- averaging discards within-group variation
- between-group and within-group can disagree
- aggregated coefficients are usually inflated
- Robinson named it in 1950
basics
~20 sA state-level correlation describes states, not people. Averaging discards within-state variation, so the state-level coefficient can be stronger than, or opposite in sign to, the individual-level one. Reading it as a fact about voters is the ecological fallacy.
solid answer
~40 sThe coefficient was computed over states, so it describes states. Aggregation replaces every person with a group average, which removes within-group variation and leaves only differences between groups — and those two sources of variation can point in different directions. In US voting data this is not hypothetical: richer states have leaned one way in recent presidential elections while, inside a given state, higher-income voters have leaned the other. A group-level number is also usually smoother and larger in magnitude than the individual-level one, because averaging cancels individual noise. The name for reading a group-level association as a claim about individuals is the ecological fallacy, a term introduced by W. S. Robinson in 1950. The discipline is to state the unit of analysis with every correlation and to answer individual-level questions with individual-level data.
go deeper
Be ready to say that a correlation computed on group averages describes groups, and that applying it to a single person is a named fallacy rather than a rounding error.
Explain the mechanism: aggregation removes within-group variation, leaving only between-group differences, and those two relationships need not agree in size or even sign.
Show the operational habit — name the unit of analysis on every chart and coefficient, and refuse to convert an aggregate finding into an individual-level recommendation without individual data.
Own the reporting norm across the organisation: aggregate dashboards look more persuasive than they are, and a lead sets the rule about what unit a decision may be based on.
## The two different questions There are two distinct correlations hiding behind the sentence 'income is correlated with voting': - **Between units:** take fifty states, compute each state's average income and its vote share, and correlate those fifty pairs. - **Within units:** take individual voters and correlate each person's income with how they voted. These answer different questions and can give different — even opposite-signed — answers. Only the second is a statement about people. ## Why aggregation changes the answer Every individual observation can be decomposed into a group average plus a deviation from that average. When you aggregate, the deviations disappear entirely: only the group averages survive into the analysis. Two consequences follow. First, **the noise cancels.** Individual behaviour is enormously variable, and averaging thousands of people per state washes most of that variability out. What is left is a small number of smooth points, which typically correlate much more tightly than the individuals did. Aggregated correlations are routinely far larger in magnitude than individual ones computed on the same underlying data, purely as an artefact of averaging. Second, **the surviving comparison is a different comparison.** Between-group differences reflect whatever makes those groups differ overall — for states, that includes urbanisation, industry mix, age structure, education and regional culture, all of which travel with average income. The within-group relationship reflects how one person differs from their neighbour. There is no mathematical reason these two must agree, and empirically they often do not. ## The voting case US presidential elections make this concrete. Across states, higher average income has been associated with support for one party in recent cycles; within states, higher-income individuals have tended to support the other. Both patterns are real, both are measurable, and neither is a mistake — they are answers to different questions. The mistake is only made when someone reads the state-level chart and says 'rich people vote this way'. This is exactly the setting in which the term was coined. W. S. Robinson's 1950 paper showed that correlations computed on US geographic aggregates were dramatically different from the corresponding correlations computed on individuals, and warned that the aggregate quantity cannot substitute for the individual one. The label 'ecological fallacy' comes from calling group-level data 'ecological'. ## The fallacy in the other direction The reverse error also exists: inferring a group-level relationship from individual data, sometimes called the individualistic or atomistic fallacy. If income predicts voting weakly at the individual level, it does not follow that state averages must be weakly related. Group-level relationships can be produced by group-level features — a state's tax regime, a regional economy — that no individual record contains. So the rule is symmetric and simple: **the unit of analysis you computed on is the unit you are entitled to talk about.** ## Where this bites outside politics - Comparing countries: national average sugar consumption against national diabetes rate says nothing directly about whether a given person's sugar intake predicts their diagnosis. - Product analytics: correlating per-account averages of a usage metric and a revenue metric across accounts, then briefing it as a statement about individual users inside an account. - Cohort dashboards: monthly cohort means correlate beautifully because each point aggregates thousands of users; the same relationship at user level can be almost flat. - Team and organisational metrics: correlating team-level throughput with team-level tooling adoption, then telling individual engineers what their own adoption implies. In every case the aggregated chart looks cleaner and more persuasive than the individual-level truth, which is precisely why it is dangerous in a slide deck. ## What to do Label the unit of analysis in every chart title and every reported coefficient — 'across 50 states', 'across 300 accounts' — so that the reader cannot silently substitute a different unit. When the decision is about individuals, get individual-level data, even a sample of it, rather than reasoning from aggregates. When only aggregates exist, say so and treat the finding as hypothesis-generating. And be suspicious of an unusually tidy correlation between two averages: tidiness is often the signature of aggregation rather than of a strong relationship. In an interview, the strongest answer names the mechanism (within-group variation was discarded), gives the direction of the artefact (aggregate correlations are usually inflated and can flip sign), names the fallacy, and states the remedy in one line: match the unit of analysis to the unit of the claim.
- Why are correlations between group averages usually larger than the individual-level ones?Because averaging cancels individual-level noise. Each group point summarises thousands of highly variable people, so most of the scatter that would appear in an individual plot never reaches the aggregated one. What is left is a small, smooth set of points, and smooth points correlate tightly. The tidiness is a property of the aggregation, not evidence of a strong relationship.
- Is there a mirror-image fallacy going the other way?Yes, sometimes called the individualistic or atomistic fallacy: inferring a group-level relationship from individual-level data. Group-level features such as a regional tax regime or a shared policy can create relationships between aggregates that no individual record encodes. The rule is symmetric — talk about the unit you actually computed on.
- You only have account-level aggregates but the question is about individual users. What do you do?Say so explicitly and downgrade the finding to hypothesis-generating. Then get individual-level data even on a small sample, since a few thousand user rows answer the actual question better than millions of users compressed into account means. If that is impossible, report the aggregate relationship with its unit named in the chart title and make no personal claim.
Countries with more doctors per person report more diagnoses. That is a fact about countries; it does not tell an individual that visiting a doctor makes them ill.
saying these in an interview costs you the question
- Reads a group-average correlation as a fact about individuals
- Never states the unit of analysis for a reported coefficient
- Treats a tidy aggregate scatter as a strong relationship
- Assumes group and individual relationships must share a sign
- Infers group patterns from individual data without comment