Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

Chi-Square Tests: Goodness of Fit and Independence

Quick fact

The chi-square statistic is computed by summing (observed - expected)² / expected across all categories. Its distribution is defined by degrees of freedom, and it can be used for both goodness-of-fit and tests of independence.

Why this is interesting

Have you ever wondered if your favorite cereal brand really has more marshmallows than the others? Chi-square tests can tell you if such differences are just random or actually meaningful.

Read the full explanation

Understanding Chi-Square Tests: Goodness of Fit and Independence

Imagine you're a botanist counting the colors of flowers in a garden. You expect a 3:1 ratio of red to white based on genetics. You count 70 red and 30 white. Are the numbers close enough to what you expected, or do they suggest something else is going on? This is exactly what a chi-square goodness-of-fit test assesses. It compares what you observe to what you'd expect if your hypothesis (the 3:1 ratio) were true. The test calculates a single number, the chi-square statistic, that measures how far your observations are from the expectations. The larger the statistic, the more the data differ from your expectation. But 'large' is relative—it depends on how many categories you have, which is captured by the degrees of freedom. To decide if the difference is statistically significant, you compare your statistic to a critical value from the chi-square distribution.

A deeper explanation

The chi-square statistic works because it quantifies the discrepancy between observed frequencies (O) and expected frequencies (E) under the null hypothesis. The formula is χ² = Σ (O−E)² / E. Squaring makes all deviations positive, and dividing by E standardizes them, giving more weight to relative differences in categories with small expected counts. This statistic follows a chi-square distribution, which is right-skewed and depends on degrees of freedom (df). For a goodness-of-fit test, df = number of categories − 1; for a test of independence, df = (rows − 1) × (columns − 1). The expected values are calculated assuming the null hypothesis is true—for independence, you use the marginal sums to compute expected cell counts. The deeper principle is that the sum of squared standardized deviations approximates a known distribution, allowing us to calculate the probability of observing such a deviation by chance alone. This is why chi-square tests are so widely used: they provide a simple, assumption-light way to test hypotheses about categorical data.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.