Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

Poisson Regression for Modelling Count Data

Quick fact

Poisson regression assumes that the variance of the count equals its mean, which is a strong restriction known as equidispersion—yet real data often show overdispersion, requiring extensions like the negative binomial model.

Why this is interesting

Have you ever tried to predict the number of rainy days in a month or the number of accidents at an intersection? Ordinary linear regression can give predictions that are negative or nonsensical.

Read the full explanation

Understanding Poisson Regression for Modelling Count Data

Imagine you're counting events that happen randomly over time or space: customer arrivals, emergency calls, or disease cases. These counts are always non-negative integers (0, 1, 2, ...). Often, small counts are common and large ones are rare, creating a right-skewed distribution. If you try to fit a straight line with ordinary linear regression, your predictions could become negative, which is impossible for a count. The Poisson distribution is a natural choice for such data—it's a probability distribution that describes how many events occur in a fixed interval when events happen independently at a constant average rate. Poisson regression builds on this: it models the natural logarithm of the expected count as a linear combination of predictors. Why log? Because it ensures the expected count is always positive and turns multiplicative effects into additive ones. So instead of predicting the count directly, we predict how the rate changes. Each coefficient tells us how the log of the expected count changes per one-unit increase in the predictor, and exponentiating gives the rate ratio.

A deeper explanation

The core mechanism is the generalized linear model (GLM) framework. The response variable Y is assumed to follow a Poisson distribution with parameter λ (the expected count). The model links λ to predictors X via: log(λ) = β0 + β1X1 + ... + βpXp. This is called the log link. Because the Poisson distribution has the property that its variance equals its mean (λ), the model automatically handles the heteroscedasticity common in counts—variance increases as the mean increases—something ordinary least squares would fail to capture. Estimation is done by maximum likelihood, which finds the coefficients that make the observed data most probable. The coefficient βj can be exponentiated to get exp(βj), the multiplicative change in the expected count for a one-unit increase in Xj, holding others constant. This is akin to a rate ratio. The model also permits the inclusion of an offset, such as the log of exposure time, to model rates per unit of exposure. A critical limitation is the equidispersion assumption; in practice, counts often show overdispersion (variance mean) due to unobserved heterogeneity, requiring negative binomial regression or quasi-Poisson. Understanding this mechanism is essential for analyzing count data correctly across science and policy.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.