Mathematics
Multiple Regression and the Problem of Multicollinearity
Quick fact
When two predictors are perfectly correlated, it is impossible to calculate separate coefficient estimates—the model literally cannot distinguish their individual effects, and the coefficients become undefined.
Why this is interesting
Ever wondered why adding more variables to your regression model sometimes makes the results less trustworthy? In multiple regression, predictors can secretly sabotage each other when they are too similar—this is the problem of multicollinearity.
Read the full explanation
Understanding Multiple Regression and the Problem of Multicollinearity
Imagine you're trying to figure out how much each of your friends' cooking skills contribute to a shared dinner, but they always cook together and you can't tell who added what. In multiple regression, we want to estimate the unique effect of each predictor on an outcome, holding others constant. But when predictors are correlated, their effects get confounded. Formally, multiple regression fits a line (or hyperplane) to predict an outcome using several predictors. The coefficients represent the expected change in the outcome for a one-unit change in that predictor, assuming all other predictors stay the same. When predictors are not independent, this 'holding constant' becomes problematic. Consider a simple example: predicting house price using square footage and number of bedrooms. These two are naturally correlated—bigger houses have more bedrooms. If we try to estimate the coefficient for bedrooms while holding square footage constant, we're asking: "What happens to price if we add a bedroom without changing the size?" That's an unrealistic scenario because it rarely occurs in practice. The data simply doesn't contain many such cases, so the estimate becomes unstable and unreliable. Multicollinearity makes the coefficient estimates sensitive to small changes in the data. Adding or removing a few observations can swing the coefficients wildly, even changing their signs. This instability is a major concern when you want to interpret coefficients as measures of importance.
A deeper explanation
The underlying problem is rooted in the mathematics of least squares estimation. If we write the regression model as Y = Xβ + ε, the ordinary least squares solution is β̂ = (X'X)⁻¹X'Y. This requires the inverse of the matrix X'X. When predictors are highly correlated, X'X becomes nearly singular—its determinant approaches zero—making the inverse numerically unstable and the coefficients highly sensitive to small changes in the data. The standard errors of the coefficients are inflated because the variance of each coefficient estimate is proportional to 1/(1 - R²) for each predictor, where R² is the coefficient of determination from regressing that predictor on all the other predictors. This factor is called the Variance Inflation Factor (VIF). If two predictors are highly correlated, their individual R² values are high, so the VIF is large, ballooning the standard errors. This means the coefficients are estimated with less precision, and statistical significance tests become unreliable. Multicollinearity does not reduce the overall fit of the model (R²) or the ability to predict the outcome, as long as the pattern of correlation remains the same in the new data. However, it severely undermines the interpretability of individual coefficients. You cannot say "this predictor has no effect" when its coefficient is not significant, because the data may simply not allow you to separate its effect from that of a correlated predictor. Detecting multicollinearity can be done by examining the correlation matrix of predictors, but a more robust method is calculating VIFs for each predictor. A common rule of thumb is that a VIF greater than 10 indicates problematic multicollinearity. Remedies include removing one of the correlated predictors, combining them into a single composite measure, or using regularized regression techniques like ridge regression that can handle correlated predictors better. Understanding multicollinearity is essential for any analyst interpreting regression output, as it guards against drawing false conclusions about the individual importance of variables.