Mathematics
Principal Component Analysis for Dimensionality Reduction
Quick fact
PCA transforms correlated variables into a new set of uncorrelated variables, called principal components, which are ordered by how much of the original variance they capture. The first principal component alone accounts for the largest possible share of total variation, often more than any single original variable.
Why this is interesting
How can you compress a dataset with hundreds of columns into just two or three numbers per row, losing as little information as possible? That is exactly what Principal Component Analysis (PCA) does.
Read the full explanation
Understanding Principal Component Analysis for Dimensionality Reduction
Imagine you have a cloud of data points in a high-dimensional space. PCA finds new axes, called principal components, that align with the directions of greatest spread (variance) in the data. Think of it like rotating a cloud of points so that the longest axis of the cloud becomes the first component, the next longest perpendicular axis becomes the second, and so on. By projecting the data onto just the first few components, you reduce the number of variables while preserving the most important structure. For example, in a dataset of body measurements, the first component often captures overall size, and the second captures shape—two numbers that explain most of the variation among many measures.
A deeper explanation
Mathematically, PCA works by computing the covariance matrix of the (often standardized) data. The eigenvectors of this matrix point in the directions of maximum variance, and the corresponding eigenvalues give the amount of variance along those directions. The principal components are these eigenvectors, sorted by decreasing eigenvalue. To reduce dimensionality, you select the top k eigenvectors and project the data onto them. The key insight is that projecting onto those eigenvectors minimizes the sum of squared distances between the original points and their projected versions (reconstruction error). This is equivalent to maximizing the projected variance, because the sum of variances along all original dimensions equals the sum of eigenvalues. PCA is widely used for data visualization (e.g., 2D plots of high-dimensional data), noise reduction, and as a preprocessing step to improve model performance and mitigate the curse of dimensionality.