Mathematics
The Law of Large Numbers and Its Two Forms
Quick fact
The law of large numbers has two formal versions: the weak law, which says the probability that the sample average deviates from the expected value goes to zero as the sample size grows, and the strong law, which guarantees that the sample average converges to the expected value with certainty (almost surely) as the sample size tends to infinity.
Why this is interesting
Why does a coin that lands heads 10 times in a row still settle at 50% heads after a million flips? The answer lies in a statistical law with two distinct versions.
Read the full explanation
Understanding The Law of Large Numbers and Its Two Forms
Imagine flipping a fair coin. In the short run, you might see many more heads than tails. But over many flips, the proportion of heads tends to get closer to 0.5. This is the law of large numbers in action. It assures us that the average from many independent trials will settle near the true expected value. There are two flavors of this law. The weak law says that for any small margin, the chance of the average being outside that margin shrinks to zero as you take more samples. The strong law goes further: with enough data, the average is certain to converge to the expected value, almost surely. The distinction is subtle: the weak law allows for the possibility that the average might occasionally deviate, but with vanishing probability, while the strong law guarantees convergence except on a set of outcomes with probability zero. To visualize, imagine a dartboard where the center is the true average. The weak law says that with many throws, the darts (sample averages) will spread closer and closer to the center, and the probability of landing far away becomes tiny. The strong law says that nearly all sequences of darts will eventually cluster at the center.
A deeper explanation
The law of large numbers arises from the interplay of independence and averaging. If you average many independent random variables with finite variance, the fluctuations tend to cancel out. More formally, let X₁, X₂, … be independent and identically distributed with mean μ. The sample average X̄ₙ = (X₁ + … + Xₙ)/n converges to μ. The weak law (also called Khinchin's law) states that for any ε 0, P(|X̄ₙ – μ| ε) → 0 as n → ∞. This is convergence in probability. It says that the probability of a large deviation becomes arbitrarily small, but it does not exclude the possibility of occasional large deviations in finite samples. The strong law (proved by Kolmogorov) states that P(limₙ→∞ X̄ₙ = μ) = 1. This is almost sure convergence. It means that with probability one, the sequence of sample averages converges to the exact expected value. The strong law is mathematically stronger and has the satisfying interpretation that the long-run frequency converges deterministically. Why does this matter? It validates the intuitive idea that probabilities can be estimated by long-run frequencies, a cornerstone of statistics. It justifies using sample means to estimate population means, and it underpins Monte Carlo methods, insurance risk pooling, and even why casinos need only a small advantage to be profitable in the long run. The two forms differ in the mode of convergence: weak law uses convergence in probability, strong law uses almost sure convergence. This distinction may seem technical, but it has implications for what can be concluded about the behavior of a single sequence of observations.