Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

The Chi-Squared Test for Categorical Data and Independence

Quick fact

The chi-squared test was developed by Karl Pearson in 1900 and is one of the most widely used statistical tests in research.

Why this is interesting

Have you ever wondered if your coffee preference is truly linked to your personality type, or if it's just a coincidence? How can we be sure that two categories—like coffee and personality—are actually related in the real world?

Read the full explanation

Understanding The Chi-Squared Test for Categorical Data and Independence

Imagine you have a bag of 100 colored candies: 50 red, 30 blue, and 20 green. That's your expected distribution. Now you open a bag and count the actual colors: 40 red, 40 blue, 20 green. Are these differences just random chance, or is the candy company actually changing the proportions? The chi-squared test helps answer this by comparing what you observed to what you expected. For a test of independence, you have two categorical variables, like 'gender' (male/female) and 'preference for cats vs. dogs'. You collect data in a table (called a contingency table). The test asks: Does the pattern of counts in one category depend on the other? For example, are males more likely to prefer dogs than females? The chi-squared test calculates a number (the chi-squared statistic) that measures how far the observed counts deviate from what you'd expect if there were no relationship (i.e., independence). A large deviation means the variables are likely not independent.

A deeper explanation

The underlying principle is the multinomial distribution and the central limit theorem. Under the null hypothesis of independence, we can calculate the expected count for each cell in the contingency table using the row and column totals: expected = (row total × column total) / grand total. The chi-squared statistic is the sum over all cells of ((observed - expected)^2 / expected). This statistic approximately follows a chi-squared distribution with (rows-1)×(columns-1) degrees of freedom, when the sample is large and expected counts are not too small (usually at least 5). The p-value tells us the probability of getting a chi-squared statistic as extreme as ours (or more) if the null hypothesis were true. A small p-value (typically < 0.05) leads us to reject the null hypothesis and conclude that there is a significant association between the variables. This test is widely used in biology (e.g., genetics to test inheritance patterns), marketing (e.g., association between age group and product preference), and social sciences (e.g., relationship between education level and voting behavior). It's a cornerstone of categorical data analysis, and its logic extends to other tests like the likelihood ratio test and logistic regression.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.