Mathematics
The Likelihood Ratio Test and Its Asymptotic Distribution
Quick fact
Under the null hypothesis, the likelihood ratio test statistic—twice the log of the likelihood ratio—follows a chi-squared distribution in large samples, a result known as Wilks' theorem.
Why this is interesting
You have two competing explanations for your data—one simple, one complex. How do you decide which one is better, and can you trust the verdict?
Read the full explanation
Understanding The Likelihood Ratio Test and Its Asymptotic Distribution
Imagine you are a detective with two hypotheses about a crime scene. The simpler hypothesis (null) says a single suspect acted alone. The more complex one (alternative) says there was an accomplice. You have evidence: a set of clues. The likelihood is a measure of how well each hypothesis explains the clues. The likelihood ratio test compares these two explanations by computing how much more likely the clues are under the complex hypothesis versus the simple one. If the complex hypothesis explains the evidence much better, you reject the simple one. The ratio tells you the strength of the evidence, and you compare it to a threshold to decide. In statistics, we compute a test statistic from the ratio of the maximized likelihoods under each hypothesis. Twice the natural log of this ratio is called the likelihood ratio statistic. The greater the statistic, the more the data favor the alternative hypothesis. To make a decision, we need to know how large this statistic must be to be surprising. That's where the asymptotic distribution comes in.
A deeper explanation
The likelihood ratio test (LRT) is built on the likelihood function, which gives the probability of observing the data given a set of parameters. For a null hypothesis that imposes constraints on the parameters (e.g., setting some coefficients to zero) and an alternative hypothesis that frees those parameters, we compute the maximum of the likelihood under each. The LRT statistic is \( \Lambda = -2 \ln \left( \frac{L(\hat{\theta}0)}{L(\hat{\theta}1)} \right) \), where \( \hat{\theta}0 \) and \( \hat{\theta}1 \) are the maximum likelihood estimates under the null and alternative, respectively. Wilks' theorem shows that under regularity conditions, as the sample size grows, \( \Lambda \) converges in distribution to a chi-squared random variable with degrees of freedom equal to the difference in the number of free parameters between the two hypotheses. This is a powerful result because it allows us to set critical values and compute p-values without knowing the exact distribution of the statistic. The mechanism behind this relies on the central limit theorem and a Taylor expansion of the log-likelihood around the true parameter. In large samples, the log-likelihood is well-approximated by a quadratic function, and the likelihood ratio statistic becomes equivalent to the squared distance between the null and alternative estimates, which follows a chi-squared distribution. This asymptotic result is what makes the LRT so broadly applicable, from classic ANOVA to modern machine learning model comparison.