Mathematics
Linear Regression via the Normal Equations
Quick fact
The normal equations give the least-squares solution directly in one step, unlike iterative methods such as gradient descent. For a simple straight line, they reduce to basic formulas for the slope and intercept using only the means and variances of the data.
Why this is interesting
You can draw a line through almost any scatter plot, but how do you find the best one? And why does the answer come out of a single, simple formula?
Read the full explanation
Understanding Linear Regression via the Normal Equations
Imagine you have a scatter plot of points, and you want to draw a line that predicts the y-value from the x-value. Which line is best? A common choice is to minimize the total squared vertical distance between each point and the line—this is called the least-squares criterion. For a line with equation y = mx + b, the sum of squared errors is a function of m and b. The normal equations are the conditions that make this sum as small as possible: they set the derivative with respect to each parameter to zero. When you solve these two equations (for a simple line), you get formulas for the slope and intercept that involve the means, variances, and covariance of x and y. These formulas are the normal equations in their simplest form.
A deeper explanation
In multiple regression, we write the model as y = Xβ + ε, where X is the design matrix (each row is an observation, each column a predictor), β is the vector of coefficients, and ε is error. The least-squares objective is ||y - Xβ||², which is equivalent to (y - Xβ)ᵀ(y - Xβ). To minimize this, we take the derivative with respect to β and set it to zero: -2Xᵀ(y - Xβ) = 0, leading to XᵀXβ = Xᵀy. This is the normal equation. If XᵀX is invertible (which requires the predictors to be linearly independent), we get the closed-form solution β = (XᵀX)⁻¹Xᵀy. This is a beautiful result: it gives the optimal coefficients in one shot, without iteration. It works because the least-squares objective is quadratic, so there is a unique minimum. The normal equations are the exact solution to the minimization, making linear regression one of the few machine learning algorithms with an analytical solution.