Mathematics
Maximum Likelihood Estimation and Its Properties
Quick fact
Maximum likelihood estimation (MLE) is the method behind many familiar statistical tools, including linear regression: the least squares line is exactly the MLE for the slope and intercept when errors are normally distributed.
Why this is interesting
Imagine you have a coin that lands heads 7 times out of 10—what probability would you assign to landing heads? Maximum likelihood estimation gives you a principled way to choose that number.
Read the full explanation
Understanding Maximum Likelihood Estimation and Its Properties
Let’s start with a simple example. Suppose we have a coin that we suspect is biased, and we flip it 10 times, getting 7 heads and 3 tails. We want to estimate the probability p of heads. One natural guess is 0.7, but why? MLE formalizes this intuition: we choose p that makes the observed data (7 heads out of 10) most probable. The likelihood of observing 7 heads is given by the binomial formula: L(p) = C(10,7) p^7 (1-p)^3. The value of p that maximizes L(p) is indeed 0.7. This is MLE at its core: pick the parameter that maximizes the probability of the data we actually saw. In general, we first write the likelihood function, which is the joint probability of the observed data expressed as a function of the unknown parameter. Then we find the parameter value that makes this likelihood as large as possible. Often, it’s easier to work with the log-likelihood, because logarithms turn products into sums and do not change the location of the maximum.
A deeper explanation
MLE works because it directly optimizes the likelihood function, which quantifies how well a parameter explains the observed data. The likelihood is constructed from the probability distribution assumed for the data. For independent observations, it is simply the product of their individual probabilities (or densities). The log-likelihood is then the sum of the log of these probabilities. Taking the derivative of the log-likelihood with respect to the parameter and setting it to zero yields the maximum likelihood estimator (MLE). This estimator has several desirable properties. Under regularity conditions, as the sample size grows (n → ∞), the MLE is consistent: it converges in probability to the true parameter value. It is also asymptotically efficient: its variance is as small as possible, reaching the Cramér–Rao lower bound, which is the inverse of the Fisher information. Moreover, the MLE is asymptotically normal: the distribution of the estimator approaches a normal distribution centered at the true parameter, with variance given by the inverse Fisher information. This property is the basis for constructing confidence intervals and hypothesis tests. A particularly elegant property is the invariance principle: if θ̂ is the MLE of θ, then for any function g, the MLE of g(θ) is simply g(θ̂). This makes MLE extremely flexible across applications.