Mathematics
Causal Inference from Observational Data Using Counterfactuals
Quick fact
Observational data, like health records or social media activity, can be used to estimate causal effects without ever running a randomized experiment—by carefully comparing what actually happened to what we predict would have happened in an alternative scenario, a 'counterfactual'.
Why this is interesting
You've probably heard that 'correlation does not imply causation.' But how can we ever know what causes what, if we can't run an experiment on everything? The answer lies in asking a clever question about a world that never happened.
Read the full explanation
Understanding Causal Inference from Observational Data Using Counterfactuals
Imagine you want to know if drinking coffee makes you live longer. You look at two groups of people: coffee drinkers and non-drinkers. You find that coffee drinkers live longer. But does coffee cause longevity? Not necessarily. Perhaps coffee drinkers also tend to exercise more, or have better jobs. This is the problem of confounding variables: factors that influence both the 'treatment' (coffee) and the 'outcome' (lifespan). The key to solving this riddle is the idea of a counterfactual: for each person, there are two possible worlds—the world where they drink coffee, and the world where they don't. We only observe one world. The causal effect of coffee is the difference between what would happen in the coffee world and what would happen in the no-coffee world. Since we never see both, we can't directly measure it. But we can use smart statistical methods. If we assume that, given a set of observable characteristics (like age, gender, fitness), the treatment assignment is 'as if random'—a big assumption called ignorability—then we can compare similar people across the two groups. By matching coffee drinkers with non-drinkers who have the same profile, we can estimate what a coffee drinker's outcome would have been had they not drunk coffee. That's the counterfactual we were missing.
A deeper explanation
The formal grounding for this reasoning is the Potential Outcomes framework (also called the Rubin Causal Model). For each unit i and a binary treatment D, we posit two potential outcomes: Yi(1) if treated, Yi(0) if not. The causal effect for that unit is Yi(1) - Yi(0). The fundamental problem of causal inference is that we only observe Yi = Di \ Yi(1) + (1-Di) \ Yi(0). So for any individual, one outcome is missing. In a randomized experiment, randomization ensures that treatment assignment is independent of the potential outcomes: D ⊥ (Y(1), Y(0)). This means the observed difference in means between treated and untreated groups is an unbiased estimate of the average causal effect. In observational data, this independence fails because treatment is chosen based on variables that also affect the outcome—confounders. The solution is to adjust for these confounders. If we collect all relevant confounders X, and if the ignorability assumption holds—D ⊥ (Y(1), Y(0)) | X—then within each stratum of X, the treatment is effectively random. We can then average the within-stratum differences, or use more sophisticated techniques like propensity score matching, weighting, or double-machine learning, to estimate the average causal effect. The entire validity of the causal conclusion hangs on this unverifiable assumption: did we measure enough right variables to make treatment assignment ignorable? This is why causal inference from observational data is both powerful and delicate.