Mathematics
The Law of Large Numbers
Quick fact
The law of large numbers guarantees that if you repeatedly flip a fair coin, the proportion of heads will approach 50% as the number of flips grows. Yet, it also means that the difference between the number of heads and tails can grow without bound—so don't expect a 'balancing' effect in the short term.
Why this is interesting
Ever wondered why casinos always win in the long run? The law of large numbers is the secret behind their business model.
Read the full explanation
Understanding The Law of Large Numbers
Imagine you're flipping a fair coin. In the short run, you might get three heads in a row, but over thousands of flips, the proportion of heads settles near 50%. This is the law of large numbers at work: as you increase the number of trials, the average of your results converges to the expected value. Think of it as a tug-of-war between randomness and the underlying probability. Each flip is an independent draw from the same distribution. The sample average is the mean of these draws. The law tells us that the more flips you do, the closer that average gets to the true probability. It doesn't mean that the number of heads and tails will become equal—just that the ratio becomes more and more predictable.
A deeper explanation
The law of large numbers has two versions, differing in the strength of convergence. The weak law, proved by Bernoulli, states that for any small error, the probability that the sample average deviates from the expected value by more than that error tends to zero as the sample size increases. This is convergence in probability. The strong law, established by Borel and Kolmogorov, is stronger: it says that the sample average converges almost surely to the expected value. That means the set of outcomes where the average fails to converge has probability zero. In practice, this means that if you could observe an infinite sequence, the average would certainly approach the true mean. The key insight is that the variance of the sample average shrinks as 1/n, so the distribution becomes more concentrated. This is why averages over large samples are so predictable, even when individual observations are noisy. The law grounds statistical inference: it justifies using sample means to estimate population means, and it underpins Monte Carlo methods. It's not about compensating for past deviations—the averages don't 'remember' the past; they just accumulate more data. Understanding the difference between the two forms clarifies how statistical guarantees operate in practice.