Mathematics
The Bootstrap Method for Estimating Sampling Distributions
Quick fact
The bootstrap method lets you estimate the sampling distribution of almost any statistic by resampling from your original data thousands of times—without needing to know the population's distribution or derive complex formulas.
Why this is interesting
You have a dataset and want to know how much your sample statistic might vary if you repeated the experiment—but you can only run the experiment once. What can you do?
Read the full explanation
Understanding The Bootstrap Method for Estimating Sampling Distributions
Imagine you have a small sample of heights from a population, and you want to know the average height and how uncertain that average is. The traditional way might assume heights are normally distributed and use a formula. But what if you don't want to make that assumption? The bootstrap says: treat your sample as a mini-population. From this mini-population, you repeatedly draw new samples of the same size, with replacement, and calculate the statistic of interest (like the mean) each time. With replacement means that after you pick a value, you put it back, so the same person can be picked again. After doing this, say, 10,000 times, you have a collection of bootstrapped statistics. The spread of those bootstrapped statistics approximates the sampling distribution of the statistic. That approximation lets you compute a standard error and confidence intervals, giving you a measure of uncertainty without needing to know the true population distribution.
A deeper explanation
The bootstrap works because of the 'plug-in principle': the sample data is the best empirical approximation we have to the unknown population distribution. By resampling from the sample, we're essentially simulating what would happen if we took many samples from that empirical distribution. The mechanism relies on the law of large numbers: as the number of bootstrap resamples increases, the distribution of the bootstrapped statistic converges to the sampling distribution of the statistic under the empirical distribution. This yields a nonparametric estimate of the sampling distribution. Importantly, the bootstrap can be applied to almost any statistic—medians, correlations, regression coefficients—without needing analytic formulas. It is especially useful when the sample size is small or the data violates traditional assumptions. However, it does have limitations: it requires the original sample to be representative of the population, and it can underestimate uncertainty in some cases. Still, the bootstrap revolutionised statistics by making complex inference computationally feasible and robust.