Mathematics
The Central Limit Theorem and Why Averages Behave Predictably
Quick fact
The Central Limit Theorem guarantees that, for any population with a finite variance, the distribution of sample means approaches a normal distribution as sample size (n) grows, regardless of the population's shape—even if the population is heavily skewed or bimodal.
Why this is interesting
You've probably noticed that averages are everywhere—test scores, polls, heights. But have you ever wondered why those averages always seem to follow a familiar bell-shaped curve, even when the underlying data are wildly skewed?
Read the full explanation
Understanding The Central Limit Theorem and Why Averages Behave Predictably
Imagine you're measuring the heights of all adults in a city. The population distribution might be roughly bell-shaped, but what if you're measuring household income, which is typically right-skewed (a few very high values)? If you take many random samples, each of size 30, and compute the average for each sample, you'll get a set of sample means. Plot those averages, and you'll see a bell-shaped curve—even though the original income data were lopsided. This is the Central Limit Theorem in action. It says that the distribution of sample averages tends toward a normal distribution as the sample size increases, no matter what the original population looks like (as long as it has finite variance). The key is that we're averaging many independent observations, and the extremes tend to cancel each other out. So even if one sample has a few high values, another sample might have more low values, and over many samples, the averages become more consistent and symmetric.
A deeper explanation
Why does this happen? The CLT stems from the fact that when you sum many independent random variables, their fluctuations partly cancel. Mathematically, the standardized sum converges to a normal distribution. For sample means, the mean of the sampling distribution equals the population mean (μ), and the standard deviation (called the standard error) equals σ/√n, where σ is the population standard deviation. As n increases, the standard error shrinks, making the sample means cluster more tightly around μ. The theorem requires independence and finite variance, but it does not require the population to be normal. This is powerful because it means that even when we don't know the shape of the population, we can still use normal-theory methods for inference. In practice, the approximation becomes good for n ≥ 30 in many cases, but for heavily skewed distributions, larger samples may be needed. This theorem underlies everything from poll margins of error to quality control charts, providing a predictable behavior for averages that would otherwise seem chaotic.