Mathematics
Partial Derivatives and the Gradient Vector in Multivariable Calculus
Quick fact
The gradient vector always points in the direction of steepest ascent of a multivariable function, and its magnitude equals the slope in that direction. This single vector encodes all the partial derivatives of the function at a point.
Why this is interesting
Imagine you're standing on a hillside in a thick fog. How can you know which way is the steepest climb without seeing the whole mountain? The answer lies hidden in a vector that calculus can compute.
Read the full explanation
Understanding Partial Derivatives and the Gradient Vector in Multivariable Calculus
For a function of two variables, f(x, y), the partial derivative with respect to x, written ∂f/∂x, measures how fast f changes as you move in the positive x-direction while holding y constant. Similarly, ∂f/∂y measures the rate of change as you move in the y-direction while holding x constant. Imagine you're on a hill: ∂f/∂x tells you the steepness if you walk due east, and ∂f/∂y tells you the steepness if you walk due north. The gradient vector, denoted ∇f = (∂f/∂x, ∂f/∂y), combines these two pieces of information into a single arrow that points in the direction of steepest uphill climb. If you could choose any direction to walk, the gradient shows you the one that gains altitude fastest. The length of the gradient arrow tells you how steep that climb is.
A deeper explanation
The reason the gradient points toward steepest ascent lies in the directional derivative. For any unit vector u, the rate of change of f in that direction is the dot product ∇f · u. This dot product is maximized when u points in the same direction as ∇f, because the cosine of the angle between them is then 1. Thus, the gradient gives both the optimal direction and the maximum slope (its magnitude). This is a direct consequence of the linear approximation of f near a point: f(x+dx, y+dy) ≈ f(x,y) + ∇f · (dx, dy). The gradient captures the first-order (linear) behavior of the function, and all higher-order changes are smaller for small steps. In machine learning, gradient descent moves opposite to the gradient to find a minimum of a loss function. In physics, forces are often the negative gradient of a potential energy field, telling particles which way to accelerate to decrease energy. The gradient is the backbone of multivariable optimization and vector calculus.