skip to content

What is the difference between sample covariance and Pearson's correlation coefficient r?

level: juniorimportance: must knowfreq 86%

answer

  1. same sign, very different scale
  2. one of them carries measurement units
  3. units of x times units of y
  4. divide by both standard deviations
  5. result is bounded to minus one and one

basics

~20 s

Sample covariance measures whether two variables move together, but it carries the product of their units, so its size means little alone. Pearson's r divides covariance by both standard deviations, giving a unitless number between -1 and 1.

solid answer

~50 s

Sample covariance is the average cross-product of deviations from the two means: `cov(x, y) = sum((xi - xbar) * (yi - ybar)) / (n - 1)`. Its sign tells you the direction of the linear relationship, but its magnitude is expressed in units of x times units of y. Measure height in centimetres and weight in kilograms and you get one covariance; measure the same people in inches and pounds and you get a completely different number for identical data. Pearson's r fixes that by standardising: `r = cov(x, y) / (sx * sy)`, where sx and sy are the sample standard deviations. Dividing by both standard deviations cancels the units, bounds the result to the interval -1 to 1, and makes different variable pairs comparable. That scale-freedom is exactly why analysts quote r and almost never quote a raw covariance.

go deeper

for a junior

Be ready to write both formulas from memory and state the one-line difference: covariance has units and no bounds, r is unitless and lives between -1 and 1. Know that the sign means direction and the absolute value means strength.

for a middle

Explain the mechanics: r is the covariance of the standardised columns, equivalently the mean product of z-scores. Show how a unit conversion multiplies covariance but cancels out of r, and why the bound -1 to 1 exists at all.

for a senior

Demonstrate judgment about which number to report. Covariance is a building block for variance and matrix work; r is what a stakeholder can read. Be able to say when a large covariance is telling you about spread rather than about association.

for a principal

Own the reporting convention. Decide when the team publishes standardised association versus effect in real units, and push back on dashboards that surface raw covariances or compare correlation values across samples with very different variable spreads.

## The quantity underneath both Suppose you have n paired observations `(x1, y1), ..., (xn, yn)` -- say height and weight for n people. Centre each column on its own sample mean and multiply the two deviations for each person: `(xi - xbar) * (yi - ybar)` A person who is above average on both variables contributes a positive product. Someone below average on both also contributes a positive product, because a negative times a negative is positive. Someone above average on one and below on the other contributes a negative product. Sum the products and divide: `cov(x, y) = sum_i (xi - xbar) * (yi - ybar) / (n - 1)` That is the **sample covariance**. The `n - 1` divisor (rather than n) is the same correction used for the sample variance: two means were estimated from the same data, so one degree of freedom is spent. Note that covariance with itself is just the variance: `cov(x, x) = var(x)`. ## Why the covariance number is hard to read The sign of a covariance is informative: positive means the two variables tend to sit on the same side of their means, negative means opposite sides, zero means the cross-products cancel. The **magnitude** is not readable, for two reasons. 1. **It carries units.** A covariance of height (cm) with weight (kg) is measured in centimetre-kilograms. There is no intuition for whether 120 cm-kg is big. 2. **It rescales with the units.** Convert the same people to inches and pounds. Each height is multiplied by about 0.394 and each weight by about 2.205, so every deviation and therefore every cross-product is multiplied by about 0.394 * 2.205 = 0.869. The covariance changes even though nothing about the people changed. In general, for `x' = a + b*x` and `y' = c + d*y`, `cov(x', y') = b * d * cov(x, y)`: shifts do nothing, scalings multiply through. Because of this you cannot compare a covariance across datasets, across variable pairs, or even across unit choices for the same pair. ## Pearson's r as standardised covariance Pearson's product-moment correlation divides the covariance by the two sample standard deviations: `r = cov(x, y) / (sx * sy)` The standard deviations carry the same units as their variables, so units cancel completely and r is a pure number. Under the same rescaling, `sx' = |b| * sx` and `sy' = |d| * sy`, so the b and d factors cancel in the ratio and `r' = sign(b * d) * r`. Rescaling by positive constants leaves r untouched; only flipping the direction of one variable flips the sign. An equivalent way to see it: r is the mean product of z-scores, `r = sum(zx_i * zy_i) / (n - 1)` where `zx_i = (xi - xbar) / sx`. Standardise both columns first and covariance and correlation become the same number. ## Reading the value r lies in -1 to 1, and the bound follows from the Cauchy-Schwarz inequality. r = 1 means every point lies exactly on a straight line with positive slope; r = -1 means the same with negative slope; r = 0 means no linear tendency in this sample. Strength is the absolute value: r = -0.8 is a stronger linear association than r = 0.5, just in the opposite direction. Common rough language calls |r| around 0.1 weak, 0.3 moderate and 0.5 or more strong, but those thresholds are field-dependent and should never be treated as rules. ## What each is good for Covariance is the building block: it is what you accumulate when you assemble a covariance matrix over many variables, and it is what enters variance formulas for sums of variables. Correlation is the reporting quantity: comparable, unitless, bounded, and immediately interpretable as tightness around a straight line. One consequence worth stating out loud: a large covariance does **not** imply a large r. If both variables have huge spread, the covariance can be enormous while the points still form a loose cloud, and r comes out small. The reverse holds too: two tightly-linked variables with tiny spread can have a near-zero covariance and an r of 0.95. ## In an interview Write both formulas, say the words "covariance has units, r does not", and give the unit-conversion example. Then close the loop: r is covariance divided by both standard deviations, which is the same as the covariance of the standardised columns.

  • If you convert weight from kilograms to pounds, what happens to the covariance and to r?
    Every weight deviation is multiplied by about 2.205, so the covariance is multiplied by about 2.205 as well. The standard deviation of weight is multiplied by exactly the same factor, so it cancels in `r = cov / (sx * sy)` and r is unchanged. Any positive linear rescaling of either variable leaves r alone; multiplying by a negative constant flips its sign but not its magnitude.
  • Can two variables have a very large covariance but a small r?
    Yes. Covariance grows with the spread of both variables, so two highly variable columns can produce a big covariance while the scatter is still a loose cloud. Dividing by sx and sy removes the spread, and r reports only how tightly the points hug a straight line. The reverse also happens: tiny spreads can give a minuscule covariance alongside an r near 1.
  • Is a relationship with r = -0.9 weaker than one with r = 0.6?
    No. The sign is direction only: negative means one variable tends to be above its mean when the other is below. Strength is the absolute value, so |-0.9| = 0.9 is a considerably tighter linear association than 0.6. Reporting the sign as if it were a penalty on strength is a classic misreading.

Covariance is like saying two runners finished 400 seconds apart -- meaningful only once you know whether the race was a sprint or a marathon. Pearson's r is the same gap expressed as a share of the race, so any two races can be compared.

saying these in an interview costs you the question

  • Uses covariance and correlation as interchangeable words
  • Reads a large covariance as a strong relationship
  • Thinks Pearson's r has measurement units
  • Says r = -0.8 is weaker than r = 0.5
  • Compares covariances computed on differently-scaled variables
  • Forgets the n - 1 divisor in the sample covariance

context