skip to content

How does the t distribution with 4 degrees of freedom differ from the standard normal?

level: middleimportance: nice to knowfreq 33%

answer

  1. same centre, different weight in the tails
  2. variance is df over df minus two
  3. at 4 df that variance equals 2
  4. the 5% cutoff moves from 1.96 upward
  5. by a couple of hundred df they coincide

basics

~20 s

Both are symmetric bell curves centred at zero, but t with 4 degrees of freedom has heavier tails and variance 2 rather than 1. Its two-sided 95% cutoff is about 2.78 against the normal's 1.96.

solid answer

~40 s

The t family is indexed by degrees of freedom, and low degrees of freedom mean fat tails. At 4 degrees of freedom the variance is `df / (df - 2) = 2`, twice the normal's, and the extra mass sits far from centre: the two-sided 5% critical value is about 2.78 instead of 1.96. The practical consequence is that a small-sample t-test needs a much larger statistic to reject, and its confidence intervals are correspondingly wider. As degrees of freedom grow the curves converge — about 2.09 at 20 df, about 2.01 at 50, and about 1.97 at 200, which is indistinguishable from 1.96 for any real decision. That convergence is why the t-versus-normal distinction is a small-sample concern, and why using t as the default costs nothing when the sample is large.

go deeper

for a junior

Be ready to say that the t distribution is symmetric and centred at zero like the normal but has heavier tails, and that its shape depends on the degrees of freedom.

for a middle

Explain the numbers: variance df/(df-2), the two-sided cutoff of about 2.78 at 4 degrees of freedom against 1.96, and the convergence as degrees of freedom grow.

for a senior

Show where it bites in practice: catch an analysis that used a normal cutoff on a handful of observations, and quantify how much wider the correct interval should have been.

for a principal

Own how small-sample results are communicated. Decide whether tiny-sample comparisons should be reported at all, and how to present intervals that are honestly wide without inviting misreading.

## Same shape family, different tail weight The standard normal and every member of the t family share three features: they are unimodal, symmetric, and centred at zero. What separates them is how much probability sits far from the centre. The t distribution is indexed by a single parameter, its **degrees of freedom**. Small values produce a curve that is slightly shorter in the middle and noticeably thicker in the tails; large values produce a curve that is visually identical to the normal. A clean way to quantify it is the variance: ``` Var(t with v df) = v / (v - 2) for v > 2 ``` At 4 degrees of freedom that is `4/2 = 2` — twice the standard normal's variance of 1. At 20 df it is about 1.11, at 50 about 1.04, and at 200 about 1.01. For `v <= 2` the variance is not even finite, and at 1 df (the Cauchy case) the mean does not exist either. That progression is the whole story: heavy tails at small df, converging to the normal as df grows. ## What it means for critical values The number that actually changes a decision is the cutoff for a two-sided 5% test: | Reference distribution | two-sided 95% cutoff | | --- | --- | | t, 4 df | about 2.78 | | t, 10 df | about 2.23 | | t, 20 df | about 2.09 | | t, 50 df | about 2.01 | | t, 200 df | about 1.97 | | standard normal | 1.96 | Read down the column and two facts jump out. First, the small-sample penalty is severe: at 4 degrees of freedom you need a statistic roughly 40% larger than the normal cutoff to reject at the same nominal level. Second, the penalty evaporates quickly — by 200 degrees of freedom the difference between 1.972 and 1.960 is well below anything that changes a real conclusion, which is why analysts working with large datasets treat the two interchangeably. The same numbers drive interval widths. A 95% confidence interval for a mean is `xbar +/- t_crit * s / sqrt(n)`, so with five observations (4 df) the multiplier is 2.78 rather than 1.96, and the interval is about 40% wider than a naive normal calculation would suggest. ## Where the heavy tails come from A t statistic divides a deviation by an *estimated* standard deviation rather than a known one. Both parts vary from sample to sample. Occasionally a small sample happens to look unusually tight, `s` comes out too small, and dividing by that small number produces a statistic far from zero even though the underlying deviation was ordinary. Those episodes are exactly the extra tail mass. With many degrees of freedom, `s` is a precise estimate, it stops wandering, and the ratio behaves like a normal deviate again. This also explains why the correction only ever runs one way. The t distribution is *wider* than the normal, never narrower, so treating a small-sample t statistic as a z score always over-rejects — it is never an error in the safe direction. ## Reading a plot of the two curves If you overlay t with 4 df on the standard normal: - The peak of the t curve sits slightly **lower**, because the total area is fixed at 1 and more of it has moved outward. - Between roughly -1 and 1 the normal is above the t curve. - Beyond about +/- 2 the t curve is clearly above, and the gap persists further out. - Overlay t with 200 df instead and the two curves are indistinguishable at plotting resolution. ## Practical takeaways - Report degrees of freedom alongside any t statistic; without them the number cannot be turned into a p-value. - Be suspicious of any small-sample analysis that quotes 1.96 or "two standard errors" as its cutoff. - Do not bother switching to normal cutoffs once the sample is large — the t result is already the same number, and using t uniformly removes one thing to get wrong. - Heavier tails in the *reference distribution* are about uncertainty in the estimated standard deviation. They are not a statement that your data are heavy-tailed, which is a different question about the sample itself.

  • What is the variance of a t distribution with 4 degrees of freedom?
    It is df / (df - 2) = 4 / 2 = 2, twice the standard normal's variance of 1. The formula holds only for more than 2 degrees of freedom; at 2 or fewer the variance is infinite, and at 1 degree of freedom the distribution is Cauchy and has no mean either. As df grows the ratio approaches 1 from above.
  • At roughly what degrees of freedom does the t cutoff become practically indistinguishable from 1.96?
    By a couple of hundred. The two-sided 95% cutoff is about 2.09 at 20 degrees of freedom, about 2.01 at 50, and about 1.97 at 200, against the normal's 1.960. Past roughly 100 the difference is under 1% of the cutoff, which will not flip a decision that was not already on a knife edge.
  • Does the heavier tail of the t distribution mean your data are heavy-tailed?
    No. The tail weight describes the reference distribution of the test statistic, and it comes from having estimated the standard deviation rather than knowing it. The data themselves are assumed to be roughly normal. Genuine heavy tails in the sample are a separate concern about the data, not something the t distribution's shape is reporting.

It is the difference between a forecast from a well-calibrated instrument and one from a cheap gauge you also had to calibrate yourself: the second forecast has to hedge wider because the ruler itself is uncertain.

saying these in an interview costs you the question

  • Says the t distribution is skewed
  • Uses 1.96 as the cutoff for a five-observation sample
  • Thinks fat tails describe the data rather than the statistic
  • Believes t can be narrower than the normal at some df
  • Cannot say what parameter indexes the t family

context