Mathematics
Almost Sure Convergence and Convergence in Probability
Quick fact
Almost sure convergence is strictly stronger than convergence in probability: if a sequence converges almost surely, it also converges in probability, but the converse is false. For example, a sequence of random variables can converge in probability without converging at any single point of the sample space.
Why this is interesting
Flip a fair coin repeatedly. Does the proportion of heads get closer to 1/2? Yes, but there are two very different ways to understand that 'closeness'. That difference is the heart of this card.
Read the full explanation
Understanding Almost Sure Convergence and Convergence in Probability
Imagine you are throwing darts at a board. You are trying to hit the bullseye. Convergence in probability is like saying that your chance of missing by more than a tiny amount becomes smaller and smaller as you keep throwing. Almost sure convergence is a more demanding promise: it says that, with probability 1, your throws eventually all land within a tiny distance of the bullseye and stay there forever. To make this precise, think of a probability space as a collection of all possible outcomes (the sample space). A random variable is a function that assigns a number to each outcome. A sequence of random variables is a sequence of such functions. We say the sequence converges in probability to a random variable X if, for any positive tolerance ε, the probability that the distance between Xn and X exceeds ε goes to 0 as n grows. In symbols: lim{n→∞} P(|Xn - X| ≥ ε) = 0 for every ε 0. For almost sure convergence, we first need to identify the set of outcomes where the sequence of numbers Xn(ω) converges to X(ω) in the usual sense of calculus. If the probability of that set is 1, we say the sequence converges almost surely. The difference is subtle but crucial: convergence in probability allows the sequence to misbehave infinitely often, as long as those bad occasions become increasingly rare. Almost sure convergence forbids that—the sequence must eventually settle down for every outcome in a set of probability 1.
A deeper explanation
The key to distinguishing these two modes lies in the event of convergence. Let An(ε) = {|Xn - X| ≥ ε} be the event that the n-th term is at least ε away from the limit. Convergence in probability means P(An(ε)) → 0 for each fixed ε. This requires that the probability of being outside a shrinking band goes to zero, but the sequence can still oscillate in and out of that band infinitely often. Almost sure convergence requires that the set of outcomes where the sequence does not converge has probability 0. Mathematically, this is equivalent to: for every ε 0, the probability of the limsup of An(ε) is 0. The limsup is the event that An(ε) occurs for infinitely many n. In other words, almost sure convergence demands that for every ε, infinitely many large deviations occur only on a null set. This distinction reveals why almost sure convergence implies convergence in probability: if the limsup has probability 0, then the probabilities of An(ε) must shrink to 0. Conversely, a counterexample shows the reverse implication fails: consider a sequence that takes the value 1 on a small set that shrinks in probability but keeps occurring infinitely often. Specifically, define Xn to be 1 on an interval of length 1/n and 0 elsewhere. Then P(|Xn| ≥ 1) = 1/n → 0, so Xn → 0 in probability. But for every outcome, Xn jumps back and forth between 0 and 1 infinitely often, so the sequence never converges pointwise anywhere. This example is not just academic. It explains why the strong law of large numbers (which uses almost sure convergence) is a stronger statement than the weak law (which uses only convergence in probability). The distinction also underpins the Borel-Cantelli lemmas, which provide a practical tool for checking almost sure convergence by looking at the sum of probabilities of An(ε).