Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

The central limit theorem and why averages become normal

Quick fact

The Central Limit Theorem states that the distribution of sample averages from any population with finite variance approaches a normal distribution as the sample size increases, even if the original population is highly skewed, like an exponential distribution, which is why pollsters can estimate election outcomes with just a few thousand respondents.

Why this is interesting

Ever wondered why so many natural and social phenomena—from heights to test scores—show up in the familiar bell curve? The answer lies in a powerful statistical guarantee that turns any messy distribution into a normal one, but only when you take averages.

Read the full explanation

Understanding The central limit theorem and why averages become normal

Imagine you are measuring the heights of everyone in a small town. Heights are approximately normally distributed, but suppose you instead measured something very skewed, like the number of social media friends people have—most have a few hundred, a few have millions. If you took the average number of friends from a random sample of, say, 50 people, and repeated that sampling many times, the averages themselves would start to look like a bell curve. Why? Because extreme values are diluted when you average many observations. A single person with millions of friends is just one data point; when you average it with 49 typical numbers, the extreme influence shrinks. As you increase the sample size, the averages become more and more concentrated around the true population mean, and the shape of their distribution becomes increasingly normal. This holds true regardless of the original distribution's shape, as long as the population has a finite variance (which is almost always the case in real data). So the CLT is the reason we can use the normal curve to make predictions about averages from big samples, even when the original data is far from normal.

A deeper explanation

The CLT works because of a deep mathematical principle: when you sum many independent random variables, the fluctuations in the sum (or average) are the result of the combined variations of each variable. Each individual variable may have a wild distribution, but when you add them, the extreme deviations tend to cancel out. More formally, if you have a population with mean μ and finite variance σ², and you take a sample of size n, the sample mean X̄ has its own distribution with mean μ and variance σ²/n. As n grows, the distribution of the standardized statistic Z = (X̄ – μ) / (σ/√n) converges to the standard normal distribution. This convergence is guaranteed by the Lindeberg–Lévy version of the theorem. The practical consequence is immense: we can approximately treat the sample mean as normally distributed, allowing us to compute probabilities, construct confidence intervals, and perform hypothesis tests using the normal curve, even when the population distribution is unknown or non-normal. This is why the CLT is often called the 'crowning achievement' of probability theory and the engine behind much of modern statistical practice.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.