skip to content

What does direct standardisation of a mortality rate to a fixed age mix let you compare?

level: middleimportance: should knowfreq 40%

answer

  1. a crude rate carries its population's own shape
  2. one shared reference distribution
  3. same weights, each group's own rates
  4. the result is hypothetical, not a count
  5. which standard you pick changes the number

basics

~20 s

It applies both populations' age-specific rates to one shared reference age distribution, so the comparison no longer reflects the fact that one population is older. The output is a hypothetical rate built for comparison, never either population's real death count.

solid answer

~50 s

A crude rate is the age-weighted blend of age-specific rates: `crude = sum over age bands of w_i * r_i`, where the `w_i` are that population's own age structure. Two countries can differ in crude mortality purely because one is older, with no difference in how deadly any given age is. Direct standardisation replaces both weight vectors with one agreed standard age distribution and recomputes `standardised = sum w_std_i * r_i`. Now the only thing differing between the two figures is the age-specific rates themselves. Three caveats I would state without being asked. The standardised rate is a fiction — nobody actually died at that rate, so never quote it as a burden. The standard population chosen changes the number, and if the two sets of age-specific rates cross across bands it can change which country looks worse, so fix the standard in advance. And it only neutralises the variable you stratified on.

go deeper

for a junior

Know that a crude rate mixes in the population's age structure, so an older country can show higher overall mortality while being safer at every age. Recognising the need to adjust is enough here.

for a middle

Be ready to write the standardised rate as the sum of standard weights times the group's own stratum rates, compute it on a two-band example, and state that the result is hypothetical rather than an actual count.

for a senior

Show that you fix and document the standard population in advance, know it can flip the ordering when stratum rates cross, and can say when indirect standardisation is the better tool for sparse strata.

for a principal

Own the reporting convention: which standard the organisation uses, whether crude or standardised figures lead in external communication, and how you stop the choice of standard becoming a negotiable input to the conclusion.

## Crude rates carry the population's own shape A **crude rate** is total events divided by total population — deaths per 1,000 people, say. Written out over strata it is a weighted average: ``` crude = sum over strata i of w_i * r_i ``` where `r_i` is the stratum-specific rate (deaths per 1,000 among people aged 65-74, for instance) and `w_i` is that stratum's share of the population. The `w_i` are a property of the population, not of its health. So a crude rate answers "how many people died here", which is the right question for planning hospital beds, and a very poor question for "is it more dangerous here". ## A worked reversal | | share young | rate young | share old | rate old | crude | |---|---|---|---|---|---| | Country A | 0.60 | 2 / 1,000 | 0.40 | 20 / 1,000 | 9.2 / 1,000 | | Country B | 0.85 | 3 / 1,000 | 0.15 | 25 / 1,000 | 6.3 / 1,000 | Country A's crude mortality (9.2) is far higher than Country B's (6.3). Yet A's age-specific rates are **lower in both bands**: 2 against 3 among the young, 20 against 25 among the old. A simply has a much older population, and old age carries a high rate everywhere. Standardise both to a 50/50 age split: ``` A: 0.5*2 + 0.5*20 = 11.0 per 1,000 B: 0.5*3 + 0.5*25 = 14.0 per 1,000 ``` The comparison reverses, and now it means what people thought the crude rate meant. ## The mechanics of direct standardisation Pick one **standard population** — a published reference age distribution, or the pooled population of the groups you are comparing. Then for each group compute ``` standardised rate = sum over strata of w_std_i * r_i(group) ``` Every group is scored with the same weights and its own rates. Two consequences follow directly: - The difference between two standardised rates is driven **only** by differences in stratum-specific rates. - The number is counterfactual. It is what the group's death rate would be if it had the standard population's age structure. Nobody experienced it, so it must never appear in a sentence about how many people died, or in a capacity plan. ## The choice of standard matters Because the standardised rate is a weighted average of the group's rates, changing the weights changes the number. Usually it only rescales, and the ordering is stable. But if the two groups' age-specific rates **cross** — one better among the young, worse among the old — the ordering itself can flip with the standard chosen. Suppose C has rates 1 (young) and 30 (old), and D has 3 and 20. Against a young standard of 90/10: ``` C: 0.9*1 + 0.1*30 = 3.9 D: 0.9*3 + 0.1*20 = 4.7 C looks better ``` Against an old standard of 20/80: ``` C: 0.2*1 + 0.8*30 = 24.2 D: 0.2*3 + 0.8*20 = 16.6 D looks better ``` Same data, opposite conclusion. This is why the standard must be fixed **before** the comparison, documented, and kept identical across every group and every period in the report. Choosing it afterwards is a way of choosing the answer. ## Indirect standardisation When a group's stratum-specific rates are unstable — small strata, few events — direct standardisation amplifies that noise, because it multiplies a shaky rate by a large standard weight. **Indirect** standardisation reverses the roles: apply a set of standard stratum-specific rates to the group's own population structure to get expected events, then form ``` SMR = observed events / expected events ``` A value above 1 means more events than the standard rates would predict for a population of that shape. The trade-off is that SMRs computed against the same standard are not strictly comparable **with each other**, because each uses its own population's weights. ## Limits worth stating - Standardisation neutralises exactly the variable you stratified on. An age-standardised rate says nothing about differences in sex composition, income, altitude or reporting completeness. - You can standardise on several variables by cross-classifying strata (age by sex), but cells thin out quickly and the stratum rates become unstable — the practical ceiling arrives fast. - The stratum boundaries are a choice. Very wide age bands leave composition differences inside a band, which the method cannot remove. - If the two populations barely overlap in some stratum — one has almost nobody over 80 — its standardised rate rests on a rate estimated from a handful of people. ## What interviewers are checking That you can write the crude rate as a weighted average and see the weights as the culprit; that you know the standardised figure is a comparison device rather than a real rate; that you would fix the standard population in advance; and that you can name indirect standardisation and when it is preferred.

  • What is indirect standardisation and when would you prefer it?
    Indirect standardisation applies a set of standard stratum-specific rates to the study population's own age structure to get expected events, then reports the ratio of observed to expected. Prefer it when the study group's own stratum rates are unstable or unavailable — small strata, few events — because direct standardisation multiplies those shaky rates by large standard weights and amplifies the noise.
  • How do you choose the standard population?
    Either a published external reference distribution or the pooled population of the groups being compared. What matters more than the choice is fixing it before you look at results, documenting it, and reusing the identical standard for every group and every period. Otherwise the standard becomes an adjustable knob, and if the groups' age-specific rates cross, it can decide which one looks worse.
  • Can you standardise on more than one variable at a time?
    Yes, by cross-classifying strata — age by sex, for example — and weighting the resulting cells by a joint standard distribution. The ceiling arrives quickly: each added variable multiplies the cell count, cells thin out, and the stratum-specific rates that the method depends on become too noisy to weight. Past two or three variables, weighting is no longer the practical tool.

It is like re-scoring two teams against an identical fixture list instead of against whichever opponents each happened to face.

saying these in an interview costs you the question

  • Quotes a standardised rate as the actual number of deaths
  • Picks the standard population after seeing which one flatters the result
  • Believes standardising removes every difference between the populations
  • Compares crude rates across populations with very different age structures
  • Confuses direct standardisation with the observed-over-expected ratio

context