Mathematics
Causal Inference Using Instrumental Variables
Quick fact
Instrumental variables can uncover causal effects even when we cannot measure or control for all confounders, as long as a valid instrument exists—making them a cornerstone of 'natural experiments' in economics and epidemiology.
Why this is interesting
How can we prove that a new drug really works when we can't run a randomized trial? The answer may be hiding in an unlikely variable.
Read the full explanation
Understanding Causal Inference Using Instrumental Variables
Imagine you want to know whether a new fertilizer increases crop yield. You observe that farmers who use it have higher yields, but they also tend to have better soil, which itself boosts yield. You can't randomize farmers to use the fertilizer, and you can't fully measure soil quality. This is a classic confounding problem: the effect of fertilizer is mixed with the effect of soil. An instrumental variable (IV) is a clever workaround. It's a third variable that influences whether a person uses the fertilizer but has no direct effect on crop yield except through that fertilizer use. In our example, rainfall could serve as an instrument: more rainfall makes fertilizer more effective, encouraging farmers to use it, but rainfall's only effect on yield is through the fertilizer (assuming it doesn't directly damage crops). By comparing farms that had different rainfall amounts, and hence different fertilizer use, we can isolate the true causal effect of the fertilizer, even though we never directly controlled for soil quality. The key is that the instrument only 'touches' the outcome through the treatment.
A deeper explanation
The IV method works by breaking the correlation between the treatment and the error term that contains unmeasured confounders. In a simple regression of outcome on treatment, the error term includes all omitted causes. The instrument, Z, is assumed to satisfy two conditions: relevance (Z is correlated with the treatment after controlling for other variables) and exogeneity (Z is uncorrelated with the error term). Under these conditions, we can use Z to create a 'clean' source of variation in the treatment. The most common estimation technique is two-stage least squares (2SLS). In the first stage, we regress the treatment on the instrument and other covariates, obtaining the predicted treatment values. In the second stage, we regress the outcome on these predicted values (plus the covariates). Because the predicted treatment is driven only by the instrument, it is no longer correlated with the omitted confounders, so its coefficient provides an unbiased estimate of the causal effect. The reason this works is that the instrument acts like a pseudo-random assignment: it shuffles the treatment without altering the outcome directly, mimicking what randomization would achieve. This is why IV analyses are often called natural experiments. However, the validity of the instrument is crucial; if the instrument has any direct effect on the outcome or is correlated with unmeasured confounders, the resulting estimates can be even more biased than ordinary regression. Thus, IV is a powerful tool, but it hinges on strong, often untestable assumptions.