Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

The Method of Least Squares and Fitting a Line to Data

Quick fact

The method was developed independently by Carl Friedrich Gauss and Adrien-Marie Legendre in the early 1800s and was used to predict the orbit of the dwarf planet Ceres.

Why this is interesting

You have a scatter of points on a graph—how do you draw the 'best' straight line through them? There's a surprising mathematical answer that has powered science for over 200 years.

Read the full explanation

Understanding The Method of Least Squares and Fitting a Line to Data

Imagine you've collected data on hours studied and test scores for a group of students. Plotting the points, you see a rough upward trend. You want to draw a line that captures this trend. The 'best' line is one that is as close as possible to all the points. But how do you define 'close'? The least squares method does this by measuring the vertical distance from each point to the line—called the residual. Some residuals are positive (point above the line), some negative (point below). If you simply added them up, positives and negatives could cancel out, giving a misleadingly good fit. To avoid this, you square each residual, making all values positive, and then add them up. The line that makes this total sum of squares as small as possible is the 'best-fitting line'. This line has the equation y = mx + b, where m is the slope and b is the intercept. The least squares method provides specific formulas for m and b that guarantee the sum of squared residuals is minimized. So, you end up with a line that balances being close to all points, giving you a mathematical summary of the relationship between the two variables.

A deeper explanation

Why does minimizing the sum of squared residuals work? The underlying principle is optimization: we are choosing the line parameters (slope and intercept) to minimize a cost function—the sum of squared errors. This is a convex problem, meaning there is a unique global minimum. The solution can be found analytically using calculus: taking partial derivatives of the cost function with respect to m and b, setting them to zero, and solving the resulting equations. This yields the well-known formulas: m = (nΣxy - ΣxΣy) / (nΣx² - (Σx)²) b = (Σy - mΣx) / n This method assumes that the relationship between x and y is linear, and that the deviations from the line are random. By squaring the residuals, we give more weight to larger errors, which helps avoid allowing small errors to accumulate. The line is also the 'best' in the sense of minimizing the Euclidean distance in the vertical direction, which is appropriate when the x-values are controlled and the y-values have measurement error. Beyond simple lines, the same principle extends to fitting curves (polynomial least squares) and multiple variables (multiple linear regression). It is the backbone of most statistical modeling, allowing us to make predictions, test hypotheses, and uncover relationships in data. Understanding least squares also highlights why extreme outliers can skew results—they have large squared residuals, so the method works hard to reduce them, pulling the line toward them.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.