Mathematics
The Characteristic Polynomial and the Cayley–Hamilton Theorem
Quick fact
The Cayley–Hamilton theorem states that every square matrix satisfies its own characteristic equation; for example, if A is a 2x2 matrix with characteristic polynomial p(λ)=λ² - (trace A)λ + det A, then p(A)=A² - (trace A)A + (det A)I = 0. This result holds for all matrices over any commutative ring, not just real or complex matrices.
Why this is interesting
Imagine a matrix as a machine that transforms vectors. Could there be a polynomial that, when evaluated on the matrix itself, gives zero—like a magic button that always returns the same result?
Read the full explanation
Understanding The Characteristic Polynomial and the Cayley–Hamilton Theorem
Think of a square matrix A as a transformation that can be applied to vectors. Just as a number can be plugged into a polynomial, you can plug a matrix into a polynomial by treating the constant term as a multiple of the identity matrix. The characteristic polynomial of A is defined as p(λ) = det(A - λI), where λ is a variable and I is the identity matrix. Its roots are exactly the eigenvalues of A. For example, if A is a 2x2 matrix, then p(λ) = λ² - (trace A)λ + det A. The Cayley–Hamilton theorem says that if you replace λ with A in this polynomial, the result is the zero matrix—a surprising identity that is not obvious at first glance. This theorem provides a powerful shortcut: instead of computing high powers of A directly, you can express them using lower powers, making calculations much easier.
A deeper explanation
The reason why the Cayley–Hamilton theorem holds lies in the adjugate matrix. Recall that for any square matrix A, the adjugate (or classical adjoint) satisfies A·adj(A) = det(A)·I. If we write B = A - λI, then B·adj(B) = det(B)·I = p(λ)·I. The entries of adj(B) are polynomials in λ, so we can write adj(B) = B₀ + B₁λ + ... + Bₙ₋₁λ^(n-1). Expanding and comparing coefficients in the identity (A - λI)(B₀ + B₁λ + ... + Bₙ₋₁λ^(n-1)) = p(λ)I gives a set of equations that, when combined, yield exactly p(A)=0. Furthermore, the characteristic polynomial is an invariant under similarity transformations: similar matrices share the same characteristic polynomial, meaning the theorem is not an artifact of a particular basis. The theorem is fundamental because it imposes a polynomial relation on the matrix, which limits the possible behavior of its powers and paves the way for definitions of matrix functions like exponentials and square roots via power series that terminate due to the theorem.