Mathematics
Optimization Problems with Constraints Using Lagrange Multipliers
Quick fact
The method of Lagrange multipliers transforms a constrained optimization problem into a system of equations by introducing a special variable λ (the Lagrange multiplier). The solution is found at points where the gradient of the objective function is parallel to the gradient of the constraint function.
Why this is interesting
You want to find the highest point on a hill, but you are only allowed to walk along a specific mountain trail. How can you know exactly where that highest point is?
Read the full explanation
Understanding Optimization Problems with Constraints Using Lagrange Multipliers
Imagine a smooth hill, represented by a function f(x, y) whose height is the value you want to maximize. Now imagine a path on the hill defined by a constraint g(x, y) = 0 — the trail you must follow. Walking along this path, the height changes, and you want to find the highest point along it. Key observation: At the highest point along the trail, the trail is tangent to a contour line of the hill. Contour lines are level sets, where f is constant. If the trail crossed a contour line, you could keep climbing, so the trail must be parallel to a contour at the top. This means the trail's direction is the same as a level curve's direction. Now, the gradient of f (∇f) points perpendicular to level curves, and the gradient of g (∇g) points perpendicular to the constraint curve. Because the trail and the level curve are parallel, their perpendicular directions are also parallel. Therefore, ∇f and ∇g point in the same (or opposite) direction, meaning one is a scalar multiple of the other: ∇f = λ ∇g, where λ is some number—the Lagrange multiplier. This single vector equation (which gives two equations, one for each coordinate) together with the constraint g(x, y) = 0 gives enough equations to solve for x, y, and λ.
A deeper explanation
The condition ∇f = λ ∇g arises from a fundamental principle: at a local extremum of f subject to g = 0, the point lies on the constraint surface and any infinitesimal move along the surface must not change f to first order. Consider a small displacement d along the constraint. Since d is tangent to the level set of g, we have d · ∇g = 0. For the point to be a stationary point of f along the constraint, we also need d · ∇f = 0. This must hold for every d tangent to the constraint. The space of such d is the tangent space of the constraint. For the linear functional d → d · ∇f to vanish on that tangent space, the vector ∇f must be orthogonal to that tangent space. But the orthogonal complement of the tangent space is spanned by ∇g. Hence, ∇f must be parallel to ∇g, i.e., there exists a scalar λ with ∇f = λ ∇g. This is the essence: Lagrange multipliers turn a constrained optimization problem into an unconstrained stationary condition on an augmented function, the Lagrangian L(x, λ) = f(x) - λ(g(x)). The derivative with respect to λ recovers the constraint. Solving ∇L = 0 yields candidates for extrema. The method is widely used in economics to maximize utility given a budget, in physics to find equilibrium configurations, and in machine learning for regularized optimization.