Mathematics
Sufficiency and the Factorization Theorem
Quick fact
A sufficient statistic can compress a dataset to a single number (or small set) without losing any information about the parameter—the Fisher information is exactly preserved.
Why this is interesting
If you had to summarize a dataset without losing any information about a key unknown, could you do it with a single number? Sometimes you can—a number that captures everything needed about the parameter.
Read the full explanation
Understanding Sufficiency and the Factorization Theorem
Imagine you have a bag of coins, and you want to know the probability of heads (the parameter). You flip a coin 100 times and record the sequence of heads and tails. The order of flips tells you nothing extra about the coin's bias; only the total number of heads matters. That total is a sufficient statistic for the probability of heads. In general, a statistic is sufficient if it captures all the information in the sample about the parameter. The factorization theorem gives a practical test: a statistic is sufficient if the joint probability of the data can be written as a product of two parts—one that depends on the data only through the statistic, and another that is free of the parameter. This means that once you compute the sufficient statistic, you can discard the raw data and lose nothing about the parameter. It's a form of data reduction that simplifies inference without sacrificing accuracy.
A deeper explanation
The factorization theorem works because of how likelihood functions separate. For a sample X1, X2, ..., Xn, their joint probability (or density) is f(x1, ..., xn | θ). A statistic T(X) is sufficient for θ if this joint can be factored as g(T(x), θ) h(x), where h(x) does not depend on θ. The factor g captures all dependence on θ, and it depends on the data only through T. Why does this work? If the likelihood factorizes this way, then the part h(x) is irrelevant for inference about θ, because it does not involve θ. The likelihood ratio between two parameter values depends only on T(x). Thus, T contains all information about θ. This concept matters because it underpins many efficient methods. For example, maximum likelihood estimators are often functions of sufficient statistics. The Rao-Blackwell theorem further shows that conditioning any estimator on a sufficient statistic can only improve its variance, providing a way to construct optimal estimators.