skip to content

What does the point-biserial correlation between a binary group flag and a continuous test score measure?

level: middleimportance: should knowfreq 48%

answer

  1. not a new estimator at all
  2. Pearson's r with 0/1 coding
  3. standardised mean gap times a balance factor
  4. sqrt(p times 1 minus p) peaks at a 50/50 split
  5. sign follows an arbitrary coding choice

basics

~10 s

Point-biserial r is just Pearson's correlation with the binary variable coded 0 and 1. It rescales the gap between the two group means, and shrinks as the two groups become more unequal in size.

solid answer

~50 s

Point-biserial r is not a separate estimator, it is Pearson's r applied when one variable takes only two values, so the same formula gives the same number. Algebraically it becomes `r_pb = ((M1 - M0) / s) * sqrt(p * (1 - p))`, where M1 and M0 are the two group means, s is the standard deviation of all the scores and p is the proportion in group 1. Two things drive it: the standardised gap between the group means, and group balance. With a half-standard-deviation gap and a 50/50 split, r = 0.5 x sqrt(0.25) = 0.25; with the same gap at a 10/90 split it falls to 0.5 x sqrt(0.09) = 0.15. The sign follows an arbitrary coding choice, so state which group is coded 1. It is also monotonically tied to the two-sample t statistic: `r_pb = t / sqrt(t^2 + n - 2)`.

go deeper

for a junior

Recall that a binary variable coded 0/1 can go straight into Pearson's correlation, and that the resulting number reflects how far apart the two group means sit.

for a middle

Derive or state the standardised-gap times balance-factor form, and explain why sqrt(p(1-p)) shrinks the coefficient when one group is rare.

for a senior

Show you would not compare point-biserial values across datasets with different prevalence, and that you report the standardised mean gap and the group distributions alongside the coefficient.

for a principal

Take a position on which summary a rare-event flag should be reported with across the org, so teams stop reading prevalence-driven attenuation as a weak effect in a model or metric review.

## It is Pearson's r wearing a hat When one variable is binary and the other continuous, you can still feed both into the ordinary Pearson correlation formula after coding the binary variable as 0 and 1. The result has a name, point-biserial r, but it is numerically identical to Pearson's r on that coding. Nothing new is being estimated; the name just signals the special structure, which happens to admit a much more interpretable algebraic form. ## The interpretable form `r_pb = ((M1 - M0) / s) * sqrt(p * (1 - p))` Here M1 is the mean score of the group coded 1, M0 the mean of the group coded 0, s the standard deviation of all n scores pooled together (computed with the divide-by-n convention), p = n1/n the share in group 1, and 1 - p the share in group 0. The first factor is a standardised mean difference: how far apart the groups sit, in units of overall score spread. The second is a balance factor. ## The balance factor is the part people forget sqrt(p(1-p)) is maximised at p = 0.5, where it equals 0.5, and it collapses toward 0 as either group becomes rare. Concretely, with a mean gap of one full standard deviation: a 50/50 split gives r_pb = 1.0 x 0.5 = 0.50, a 10/90 split gives 1.0 x 0.30 = 0.30, and a 1/99 split gives 1.0 x 0.0995 = about 0.10. The groups differ by exactly as much in all three cases, yet the correlation drops fivefold. This is why point-biserial r is a poor way to compare effects across datasets with different prevalence, and why a small r on a rare-group flag does not mean the groups are similar. The consequence runs the other way too: point-biserial r cannot reach 1 unless the split is balanced and the separation is complete. With a rare flag it is structurally capped well below 1, so judging it against the usual correlation rules of thumb understates the group difference. ## Sign is a coding artifact Swapping the labels, coding the other group as 1, flips the sign and leaves the magnitude untouched. Any linear recode does the same: using 1 and 2 instead of 0 and 1 changes nothing at all, since correlation is invariant to linear rescaling. So the sign carries meaning only once you say which group wears the 1, and any report of a point-biserial value must state that. ## Relationship to the two-sample comparison The same two numbers, the mean gap and the group sizes, drive both point-biserial r and the equal-variance two-sample t statistic on the same data. They are linked by `r_pb = t / sqrt(t^2 + n - 2)`, so knowing one gives the other. That identity is worth remembering because it explains a common confusion: the t statistic keeps growing as the sample grows for a fixed real gap, while r_pb stays put. One measures how strong the pattern is, the other how much data stands behind it, and inflating the sample changes only the second. ## Assumptions and traps Because it is Pearson's r, everything that troubles Pearson's r troubles this too: it is a summary of means and variances and is sensitive to outliers in the continuous variable, and it can only see a shift in level between the two groups. Two groups with identical means but very different spreads produce r_pb near 0 while being obviously different distributions, so plot the two distributions rather than trusting the coefficient alone. A related coefficient, the biserial correlation, is used when the binary variable is a dichotomised version of something underlying and continuous, such as pass/fail cut from an actual score. It answers a different question, estimating the correlation with the latent continuous variable, and generally comes out larger than the point-biserial value on the same data. Use point-biserial when the two groups are genuinely two categories, not a chopped continuum.

  • Does the value change if you code the group flag 1 and 2 instead of 0 and 1?
    No. Correlation is invariant to any linear rescaling of either variable, so 0/1, 1/2 and 10/20 all give the identical coefficient. Only reversing which group is the higher code changes anything, and that flips the sign while leaving the magnitude alone. This is why a reported point-biserial value is meaningless until you say which group carries the higher code.
  • Why can a large group difference still produce a small point-biserial r?
    Because of the balance factor sqrt(p(1-p)). At 1 percent prevalence that factor is about 0.10, so even a full one-standard-deviation gap between the groups caps r_pb near 0.10. The groups genuinely differ; the coefficient is structurally squeezed by the imbalance. Report the standardised mean gap alongside r whenever the flag is rare.
  • What does point-biserial r fail to capture about the two groups?
    Anything other than a shift in mean level. Two groups with equal means but very different variances or shapes give r_pb near 0, and a handful of extreme scores in the continuous variable can move the coefficient substantially. Always look at the two distributions side by side before letting one number stand in for the comparison.

The mean gap sets the loudness; the balance factor sets the volume knob, and a rare group turns it right down.

saying these in an interview costs you the question

  • Calls it a different estimator from Pearson's r
  • Reads the sign as meaningful without stating the coding
  • Ignores group imbalance when judging the magnitude
  • Concludes the groups are similar from a small r on a rare flag
  • Assumes a growing sample raises the correlation

context