Mathematics
Logistic Regression for Binary Outcomes
Quick fact
Logistic regression doesn't fit a line to the data; it fits an S-shaped curve that always stays between 0 and 1. This curve is the logistic function, and the model estimates the probability of an event, such as tumor malignancy, based on features like tumor size or age.
Why this is interesting
You have a dataset of patients with a tumor: benign or malignant. Can a straight line predict the chance of cancer? And what if the line predicts a probability of 1.2—impossible! How do we fix that?
Read the full explanation
Understanding Logistic Regression for Binary Outcomes
Imagine you want to predict whether a student passes an exam (yes/no) based on hours studied. Linear regression would try to draw a straight line through points that are either 0 or 1, which often produces values like 0.8 or -0.2—not valid probabilities. Logistic regression solves this by compressing the line into a smooth S-curve (the sigmoid). The curve approaches 0 as the predictor decreases and approaches 1 as it increases, so every prediction is a legitimate probability. The model doesn't directly predict the outcome but the probability of the outcome. You can then choose a threshold, say 0.5, to classify: if the probability is above 0.5, predict 'pass'; otherwise, 'fail'. The key is that the relationship is not linear between the predictor and the probability; it's linear between the predictor and the logarithm of the odds (called log-odds). That’s why the coefficient means: for a one-unit increase in the predictor, the log-odds of the outcome increase by that amount. It’s an elegant way to handle a binary outcome while retaining a linear-model structure.
A deeper explanation
Logistic regression models the probability p that an event occurs as p = 1 / (1 + e^-(β₀ + β₁x₁ + ... + βₖxₖ)). This is the logistic function, which takes any real-valued linear combination and maps it to the interval (0,1). The quantity β₀ + β₁x₁ + ... + βₖxₖ is the log-odds: ln(p/(1-p)). So, the model is linear in log-odds, not in probability. The parameters β are estimated using maximum likelihood estimation (MLE), which chooses values that make the observed data most probable. Unlike linear regression, there is no closed-form solution, so we use numerical optimization (like Newton's method). The model's predictions are probabilities, and with a threshold (often 0.5) we classify into one of the two outcomes. This approach is powerful because it doesn't assume the data are normally distributed, as linear regression does. It also naturally handles nonlinear relationships between predictors and the outcome, making it a flexible tool for classification and risk prediction in medicine, finance, social sciences, and many other fields.