Mathematics
Confidence Intervals for Population Means and Proportions
Quick fact
A 95% confidence interval does NOT mean there's a 95% chance the true parameter lies in that particular interval. Instead, it means that if you repeated the sampling process many times, 95% of the intervals you'd construct would contain the true parameter.
Why this is interesting
Imagine you poll 1,000 voters and find 52% support a candidate. How confident can you be that the true support in the whole population is close to 52%? A confidence interval gives you a range, say 49% to 55%, with a measured level of certainty.
Read the full explanation
Understanding Confidence Intervals for Population Means and Proportions
When you estimate a population mean or proportion from a sample, you rarely get exactly the true value—there's always sampling error. To express this uncertainty, statisticians build a confidence interval: a range of plausible values for the true parameter. Think of it as a net cast around your sample estimate: the wider the net, the more confident you are it catches the true value. The construction relies on the sampling distribution of your statistic. Thanks to the Central Limit Theorem, for large enough samples, the sample mean or proportion follows a bell-shaped (normal) distribution centered at the true population value. The spread of this distribution, called the standard error, tells you how much your estimate might vary from sample to sample. To build a 95% interval, you go out about 2 standard errors (the exact number comes from a critical value) on either side of your sample estimate. For a confidence interval for the mean, the formula is: sample mean ± (critical value) × (standard error). For a proportion, it's similar: sample proportion ± (critical value) × (standard error of the proportion). The critical value is determined by the desired confidence level (e.g., 1.96 for 95%) and the distribution you're using (normal for large samples with known sigma, t-distribution for smaller samples or unknown sigma).
A deeper explanation
The mechanism behind confidence intervals is the repeated-sampling interpretation. Each sample produces a different interval, but the method is designed so that a certain percentage (the confidence level) of these intervals will capture the true parameter. This works because, under the Central Limit Theorem, the sampling distribution of the sample mean or proportion centers on the true value, and the critical values are chosen so that the area in the tails gives the desired coverage probability. In practice, the key steps are: 1) compute the sample statistic, 2) find the standard error (which depends on sample size and variability; for proportions it's sqrt(p(1-p)/n)), 3) choose a critical value based on the confidence level and the appropriate distribution (z for normal, t for small samples with unknown sigma), and 4) form the interval by adding and subtracting margin of error (critical value × standard error). The resulting interval's width is controlled by the level of confidence (more confidence means a wider interval) and the sample size (larger samples give narrower intervals). The confidence interval is the standard way to report estimation uncertainty across fields—medicine, public opinion, quality control—because it conveys both the estimate and its precision.