Mathematics
The Normal Distribution and the Empirical Rule
Quick fact
In a normal distribution, about 68% of all values lie within one standard deviation of the mean, 95% within two, and 99.7% within three—this is the empirical rule, a quick key to understanding data spread.
Why this is interesting
Have you ever wondered why so many measurements—heights, test scores, or noise levels—form a bell-shaped curve? What if you could predict exactly how many values fall within a certain range just by knowing the average and the spread?
Read the full explanation
Understanding The Normal Distribution and the Empirical Rule
Imagine measuring the heights of a large group of adults. Most people are close to the average height, while very tall and very short individuals are rare. When you plot the frequencies, you get a symmetric, bell-shaped curve known as the normal distribution. This shape is characterized by two numbers: the mean (the peak) and the standard deviation (how spread out the curve is). The empirical rule is a shortcut that leverages this shape: if you go one standard deviation to the left and right of the mean, you capture about 68% of the data. If you go two standard deviations, you capture about 95%, and three standard deviations covers about 99.7%. This means almost all data (99.7%) lies within three standard deviations, so anything beyond that is considered unusual—a potential outlier.
A deeper explanation
The normal distribution is defined by the probability density function that sets the precise shape of the bell curve. Because of this shape, the area under the curve between any two points corresponds to the proportion of data falling in that range. The empirical rule emerges from the mathematical property that these areas are fixed for any normal curve: approximately 68% of the area lies within ±1 standard deviation, 95% within ±2, and 99.7% within ±3. This rule is incredibly useful for quick assessments of data variability and for flagging anomalies. For instance, in quality control, a process might be flagged if a measurement falls beyond three standard deviations, because such an event is rare under normal conditions. The rule also forms the basis for the z-score: the number of standard deviations a data point is from the mean. Knowing that 95% of data falls within two standard deviations directly connects to the common 95% confidence interval, a cornerstone of inferential statistics. Understanding this rule gives you a mental model for interpreting data spread and identifying outliers without extensive calculations.