Mathematics
The Chain Rule Beyond One Dimension: Understanding Gradients
Quick fact
The gradient vector, built from partial derivatives, points exactly in the direction of the steepest ascent on a surface—a key idea used in machine learning to minimize errors.
Why this is interesting
You know how to find the slope of a hill, but what if you're on a mountain with slopes in every direction? How do you find the steepest path? The answer lies in a deceptively simple rule that extends beyond one dimension.
Read the full explanation
Understanding The Chain Rule Beyond One Dimension: Understanding Gradients
Think of a simple single-variable function like f(x) = x². Its derivative tells you how fast f changes as x changes. Now imagine a function of two variables, like the height of a landscape: h(x, y). At any point, the landscape has a slope in the x-direction and a different slope in the y-direction. These are the partial derivatives, ∂h/∂x and ∂h/∂y. But what if x and y themselves depend on another variable, like time t? For instance, you're walking on that landscape, so your coordinates change over time. How fast is your height changing? The chain rule in one dimension says df/dt = (df/dx)(dx/dt). In higher dimensions, you must consider both coordinates: dh/dt = (∂h/∂x)(dx/dt) + (∂h/∂y)(dy/dt). Each partial derivative tells you how h responds to changes in that direction, and you multiply by how fast that coordinate changes, then add them together. Now, if you want to know the overall slope in a particular direction, you combine the partial derivatives into a vector called the gradient, denoted ∇h = (∂h/∂x, ∂h/∂y). The gradient points uphill, and its magnitude tells you the steepness. The chain rule becomes a dot product: the rate of change of h along a path is the gradient dotted with the velocity of the path.
A deeper explanation
The chain rule in higher dimensions emerges from the idea of linear approximation. Near a point, a multivariable function can be approximated by a plane. The slope of that plane in any direction is determined by the gradient. When you move along a path given by a vector function, you're effectively asking how much the function changes per unit time. The chain rule formula dh/dt = ∇h · (dx/dt, dy/dt) captures this: the gradient provides the local slopes, and the velocity vector gives how fast you're moving in each coordinate direction. The dot product sums the contributions from each dimension, reflecting the principle that independent changes add. This concept is essential because it extends the idea of 'rate of change' to complex systems. In optimization, the gradient tells you which direction to adjust parameters to increase or decrease a cost function most efficiently—this is the basis of gradient descent. In physics, the chain rule connects changing coordinates to changes in fields like temperature or potential. Understanding this bridges the gap from simple derivatives to vector calculus, enabling powerful tools like the Jacobian matrix, which generalizes the gradient to transformations, and the directional derivative, which measures change along any given direction.