Mathematics
Bayes' Theorem and the Problem of False Positives
Quick fact
When a disease affects 1 in 1,000 people and the test is 99% accurate, a positive result means only about 9% of those who test positive actually have the disease—the other 91% are false positives. This is called the false positive paradox.
Why this is interesting
Imagine you test positive for a rare disease that affects only 1% of people. The test is 99% accurate. Would you assume you are 99% likely to have the disease? You might be surprised to learn the real probability is only about 50%.
Read the full explanation
Understanding Bayes' Theorem and the Problem of False Positives
Let's build your intuition with a simple example. Suppose you're a doctor and you have a screening test for a rare disease. The disease affects only 1% of the population (the base rate). Your test is quite good: it correctly identifies 99% of people who have the disease (sensitivity) and correctly says 'no disease' for 99% of healthy people (specificity). Now you have a new patient who tests positive. How worried should you be? Intuition says the test is 99% accurate, so it's almost certain they have the disease. But this is wrong because we'he ignoring the base rate. Think about it in terms of a large population: imagine 10,000 people. Only 100 (1%) actually have the disease. Of those 100, 99 will test positive (true positives). But among the 9,900 healthy people, 1% will also test positive—that's 99 false positives. In total, you get 198 positive results, but only 99 are real. So the probability that a positive result is a true positive is 99 / 198 = 50%. The test being accurate doesn't mean a positive result is accurate—because false positives can easily outnumber true positives when the condition is rare.
A deeper explanation
The mechanism behind this surprising result is Bayes' theorem, which formally updates the probability of a hypothesis (the disease) given new evidence (a positive test). The formula is: P(Disease|Positive) = [P(Positive|Disease) P(Disease)] / P(Positive). The numerator multiplies the sensitivity (P(Positive|Disease)) by the prior probability of disease (P(Disease)). The denominator, P(Positive), is the total probability of a positive test result, which includes both true positives and false positives: P(Positive) = (Sensitivity Base Rate) + (False Positive Rate (1 - Base Rate)). When the base rate is low, the false positive rate, even if small, makes up a large fraction of all positives. This is why the posterior probability (the probability after seeing the test) can be much lower than the accuracy. This principle extends beyond medicine—it applies to spam filters, security screenings, and any situation where we must update beliefs with imperfect evidence. Ignoring the base rate is a common cognitive error called the base rate fallacy, and Bayes' theorem gives us a rigorous way to correct it.