Mathematics
The Chi-Square Test for Independence in Contingency Tables
Quick fact
The chi-square test for independence can detect whether two categorical variables, like 'pizza preference' and 'movie taste,' are related, and it was developed by Karl Pearson in 1900, who also invented the correlation coefficient.
Why this is interesting
You're at a party and notice that people who like pineapple pizza also seem to love terrible movies. Is that a real pattern or just coincidence? How can we tell with numbers?
Read the full explanation
Understanding The Chi-Square Test for Independence in Contingency Tables
Imagine we have a table (called a contingency table) that counts how many people fall into each combination of two categories, such as 'Likes Pineapple Pizza' and 'Likes Action Movies'. The question is: are these two categories independent, meaning the chance of liking pineapple pizza is the same whether someone likes action movies or not? To test this, we calculate what the counts would be if they were independent—these are called expected frequencies. We then compare each observed count to its expected count. If the observed counts deviate a lot from the expected, that suggests the categories are related. This comparison is summed into a single number, the chi-square statistic. We then see if that number is so large that it is unlikely to happen by random chance alone, using a probability distribution called the chi-square distribution.
A deeper explanation
Mechanism: For each cell in the contingency table, the expected frequency is computed as (row total column total) / grand total. This formula arises because if the two variables are independent, then the probability of being in a particular row and column is the product of the row probability and column probability. The chi-square statistic is then Σ (observed - expected)^2 / expected. Each squared residual is weighted by the inverse of the expected count, so larger deviations from cells with small expected counts are penalized more. The resulting statistic follows a chi-square distribution with degrees of freedom equal to (number of rows - 1) (number of columns - 1). This distribution is right-skewed, and its mean equals the degrees of freedom. If the statistic is large, the p-value—the probability of getting a statistic this large or larger if the null hypothesis of independence were true—will be small. A small p-value (usually below 0.05) leads us to reject the null hypothesis and conclude the variables are associated. The test is non-parametric, meaning it does not assume a normal distribution of the data. It is widely used in fields such as biology, political science, and market research to analyze survey results, genetic crosses, and more.