Mathematics
The t-Distribution and Confidence Intervals for Small Samples
Quick fact
The t-distribution was developed by William Sealy Gosset in 1908, under the pseudonym 'Student' to avoid revealing his employer's trade secrets. Its extra-thick tails ensure that 95% confidence intervals are correct even when the sample size is small—provided the underlying data is approximately normal.
Why this is interesting
You've collected 12 measurements from a lab experiment and you want to estimate the true value. Your instinct is to use the familiar bell curve to build a confidence interval, but doing so could give you false confidence. Why does the bell curve let you down just when you need it most?
Read the full explanation
Understanding The t-Distribution and Confidence Intervals for Small Samples
Imagine you're studying the heights of a rare species of plant, and you can only measure a handful of specimens. You want to know the average height of the entire species, but your sample is tiny. A confidence interval gives you a range of plausible values for that true average. The usual method involves the normal distribution, but for small samples, the normal distribution is too optimistic—it assumes you know the population's spread precisely. In reality, with a small sample, you only have a rough estimate of that spread, which introduces extra uncertainty. The t-distribution, shaped like the normal but with fatter tails, accounts for this extra uncertainty. As your sample size increases, your estimate of the spread improves, and the t-distribution gets closer to the normal distribution.
A deeper explanation
For a normally distributed population, the sample mean follows a t-distribution with n−1 degrees of freedom. The t-distribution is defined by a parameter called degrees of freedom, which equals the sample size minus one. Its shape depends on this parameter: with fewer degrees of freedom, the tails are heavier, meaning extreme values are more likely. This heaviness compensates for the uncertainty introduced by estimating the population standard deviation from the sample. When constructing a confidence interval, the formula is: sample mean ± (t-critical value) × (sample standard deviation / √n). The t-critical value is larger than the z-critical value for the same confidence level, producing a wider interval. This wider interval correctly reflects the greater uncertainty of small samples, ensuring that if you repeated the experiment many times, 95% of the intervals would contain the true population mean.