Mathematics
The Berry-Esseen Theorem: Quantifying the Accuracy of Normal Approximation
Quick fact
The Berry-Esseen theorem guarantees that the maximum error in using the normal distribution to approximate the cumulative distribution function of a standardized sample mean is at most a constant times the third absolute moment divided by the square root of the sample size—about 0.4 to 0.5 times that ratio for independent samples.
Why this is interesting
You know that averages of random samples become bell-shaped as the sample grows—but how good is that approximation for a finite sample? The Berry-Esseen theorem provides the answer.
Read the full explanation
Understanding The Berry-Esseen Theorem: Quantifying the Accuracy of Normal Approximation
Imagine you are estimating the average height of people in a city by sampling a small group. The Central Limit Theorem tells you that the distribution of your sample average, when standardized, follows a bell curve—the normal distribution—as your sample size becomes large. But in practice, your sample size is never infinite. You might wonder: how close is my bell curve to the true distribution? The Berry-Esseen theorem answers this question by giving a precise bound on the difference between the true cumulative distribution function (CDF) and the normal CDF. It says that this difference, at any point, is at most a constant (around 0.4 or 0.5) times a quantity related to the underlying distribution's skewness, divided by the square root of the sample size. In simple terms: the larger your sample, the smaller the error—and the more skewed your original data, the larger the error. This turns the qualitative promise of the CLT into a quantitative tool.
A deeper explanation
The Berry-Esseen theorem is a refinement of the Central Limit Theorem. For independent and identically distributed random variables with mean μ and variance σ², the standardized sum Sn = (1/√n)∑(Xi - μ)/σ converges in distribution to a standard normal. The theorem bounds the Kolmogorov distance between the CDF of Sn and Φ, the standard normal CDF: supₓ |Fn(x) - Φ(x)| ≤ C·ρ/σ³·(1/√n), where ρ = E|X-μ|³ is the third absolute central moment. The constant C can be taken as 0.4748 (for i.i.d. sequences) or 0.4097 (under slightly different conditions). The key mechanism is that the error decays as 1/√n, meaning that to halve the error you must quadruple the sample size. The bound depends on the third moment, which captures the skewness and tail heaviness of the distribution: heavy-tailed or highly skewed distributions will have a larger ρ, making the approximation worse for the same n. This theorem is not just a theoretical curiosity; it provides a rigorous justification for using normal approximations in finite samples, which is vital in hypothesis testing, confidence intervals, and quality control.