Mathematics
Total Variance
Quick fact
Total variance is often called the 'total sum of squares' and is the starting point for measuring how much of the variation can be explained by a model.
Why this is interesting
Imagine you have a set of exam scores, and you want to know how much students' performance varies overall. How do you capture that total spread in a single number?
Read the full explanation
Understanding Total Variance
Total variance is a measure of how scattered a dataset is. To compute it, you first find the mean of all data points. Then, for each point, you calculate its difference from the mean, square that difference (to avoid cancellation of positive and negative differences), and add all these squared differences together. This sum is the total variance. For example, if you have test scores 70, 80, 90, the mean is 80. The deviations are -10, 0, 10; squared they become 100, 0, 100; total variance = 200. This tells you the overall variability without yet explaining why it occurs.
A deeper explanation
The concept of total variance is fundamental because it sets the total amount of variation that can be explained. In statistical modeling, we partition total variance into two parts: variance explained by the model (e.g., due to different treatments) and residual (unexplained) variance. This decomposition, seen in ANOVA and regression, allows us to calculate the proportion of variance explained (R²). Total variance also appears in machine learning, such as in PCA, where the total variance is the sum of eigenvalues, and we seek to capture most of it with fewer components. Understanding total variance is crucial for interpreting how well a model fits data and for making decisions about model complexity.