Mathematics
The t-Distribution and Small-Sample Confidence Intervals
Quick fact
The t-distribution was published in 1908 by William Sealy Gosset under the pseudonym 'Student' because his employer, Guinness Brewery, wouldn't allow employees to publish research.
Why this is interesting
Think you can use the normal curve for any sample? When your sample is small, that curve may be quietly leading you astray. What's the hidden adjustment that keeps your results honest?
Read the full explanation
Understanding The t-Distribution and Small-Sample Confidence Intervals
Imagine you're trying to estimate the average height of students in a huge school. If you can only measure a handful of students, you're not very sure about your estimate. The normal distribution works perfectly when you know the true spread (standard deviation) of the whole population, but in real life you usually don't. Instead, you estimate it from your sample. When your sample is small, that estimate is more likely to be off, so you need to be extra cautious. The t-distribution is like a safety-fitted version of the normal curve: it's wider and has heavier 'tails'. This means it expects more extreme values than the normal would, reflecting the extra uncertainty. As your sample grows, the t-distribution gets thinner and thinner until it looks almost exactly like the normal curve. To use it, you need to know the 'degrees of freedom', which for a simple mean is just your sample size minus one.
A deeper explanation
The t-distribution arises from the ratio of a normally distributed sample mean to an estimated standard error. Because the standard error is itself a random variable (computed from the sample's standard deviation), the ratio doesn't follow a normal distribution when the sample is small. Instead, it follows a t-distribution with n−1 degrees of freedom. The extra variability of the estimated standard error makes the distribution more spread out than the normal, with heavier tails. When constructing a confidence interval for a small sample mean, you use the t critical value (from a t-table or software) instead of the z critical value from the normal. This widens the interval, reflecting the greater uncertainty. As n grows, the estimate of the standard error becomes more accurate, and t converges to z. This mechanism is why small-sample intervals based on the normal are too narrow and lead to overconfident, misleading conclusions.