Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

Singular Value Decomposition and Its Pseudoinverse for Data Science

Quick fact

The SVD of any matrix exists for any rectangular matrix, and its pseudoinverse (Moore-Penrose inverse) gives the least-squares solution to linear equations, which powers most regression and signal processing algorithms.

Why this is interesting

Ever wondered how a search engine ranks millions of pages or how Netflix predicts what you'll watch next? It often comes down to a simple matrix operation called the singular value decomposition—a tool that reveals the hidden structure in any dataset.

Read the full explanation

Understanding Singular Value Decomposition and Its Pseudoinverse for Data Science

Think of a matrix as a machine that takes a vector and transforms it into another vector. But any such transformation can be broken down into three simpler steps: first, rotate the input (using an orthogonal matrix U^T), then stretch or shrink the coordinates (using a diagonal matrix Σ with positive entries called singular values), and finally rotate again (using an orthogonal matrix V). That is exactly what the singular value decomposition (SVD) does: M = U Σ V^T. Intuitively, the singular values tell you how much the matrix stretches along each of the new axes, and the vectors in U and V tell you the directions of stretching. This decomposition works for any matrix—square or not—which is why it is so powerful. For a data matrix with rows as samples and columns as features, the right singular vectors (columns of V) define the principal directions in the feature space, and the singular values indicate the importance of each direction.

A deeper explanation

The SVD is not just a factorization; it reveals the intrinsic geometry of a linear map. The number of non-zero singular values equals the rank of the matrix. The largest singular values and their corresponding vectors capture the dominant patterns in the data. This leads to the best low-rank approximation: if you keep only the top k singular values and set the rest to zero, you get the matrix of rank k that is closest to the original in the Euclidean (Frobenius) norm—a result known as the Eckart–Young theorem. This is the basis for PCA, where you center the data and take the SVD; the right singular vectors are the principal components, and the singular values are related to the variance explained. When the matrix is not square or is singular, the straightforward inverse does not exist. The pseudoinverse, commonly denoted A⁺, is defined using the SVD: if A = U Σ V^T, then A⁺ = V Σ⁺ U^T, where Σ⁺ is obtained by taking the reciprocal of each non-zero singular value and leaving zeros unchanged. This pseudoinverse gives the minimum-norm least-squares solution to Ax = b. That means it solves the problem: find x that minimizes ||Ax - b||², and among all such solutions, pick the one with the smallest norm. This property makes it ideal for noisy, overdetermined, or underdetermined systems—exactly the situations found in data science, where you often have more equations than unknowns or vice versa, and measurements contain errors. By using the pseudoinverse, you can compute a solution that is robust and well defined, even when the matrix is not invertible. Moreover, for ill-conditioned problems, you can truncate the smallest singular values to get a regularized solution, which is similar to ridge regression. Thus, SVD and the pseudoinverse are not just mathematical curiosities; they are the workhorses behind many machine learning and signal processing algorithms.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.