Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

The Bootstrap Method: Resampling Data to Estimate Uncertainty

Quick fact

The bootstrap method, introduced by Bradley Efron in 1979, allows you to estimate the standard error of any statistic—even complex ones like the median—by resampling your original data with replacement. It has become a cornerstone of modern statistics because it works when traditional formulas are too complicated or don't exist.

Why this is interesting

Ever wonder how statisticians measure uncertainty from just one sample? The bootstrap achieves the seemingly impossible: creating thousands of new datasets from the data you already have.

Read the full explanation

Understanding The Bootstrap Method: Resampling Data to Estimate Uncertainty

Imagine you want to know the average height of people in a city, but you can only afford to measure a small sample of, say, 50 people. You calculate the average of your sample, but you know that a different sample might give a slightly different average. The spread (or variability) of those averages is the sampling distribution, and it's crucial for constructing confidence intervals. However, getting another sample might be too costly or impossible. The bootstrap cleverly estimates this variability by treating your original sample as a mini-population. You repeatedly draw new samples from it—each of the same size as the original—but you draw with replacement, so some individuals may appear multiple times and others may be left out. For each of these 'bootstrap samples', you compute the average. The collection of these averages forms a bootstrap distribution. The spread of this distribution is an estimate of the standard error, and you can also extract confidence intervals directly from it. The key insight is that if your sample is a good representation of the population, then resampling from the sample mimics resampling from the population.

A deeper explanation

The bootstrap's power lies in the 'plug-in principle' and the law of large numbers. The population is unknown, but the empirical distribution of your sample (which gives each observed value an equal probability) is a valid estimate of it. By resampling from this empirical distribution, you simulate drawing many samples from the estimated population. The bootstrap distribution of a statistic approximates the true sampling distribution, provided the sample is large enough and the statistic is 'smooth' (e.g., means, medians, correlations). This works because the bootstrap distribution captures the variability induced by sampling. The mechanism is simple: (1) draw a large number of bootstrap samples (e.g., 10,000), (2) compute the statistic of interest for each, (3) use the spread of these values (e.g., their standard deviation, or the 2.5th and 97.5th percentiles) to form standard errors and confidence intervals. This approach is widely used in regression, machine learning, and complex survey designs, offering a flexible alternative to theoretical formulas when assumptions are violated or the statistic is difficult to derive mathematically.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.