Mathematics
Partial Derivatives and Gradients for Functions of Several Variables
Quick fact
The gradient vector points in the direction of the steepest increase of a function, and its magnitude is exactly the rate of that steepest ascent—this fact powers everything from hill-climbing algorithms to the backpropagation in neural networks.
Why this is interesting
You know that the slope of a hill tells you how steep it is in one direction, but what if you're standing on a mountain and need to know the steepest way up? How does calculus tell you which direction to walk?
Read the full explanation
Understanding Partial Derivatives and Gradients for Functions of Several Variables
For a function of one variable, the derivative gives the slope of the tangent line. For a function of two variables, like a surface f(x, y), there are two basic slopes: one in the x-direction and one in the y-direction. These are the partial derivatives, written as ∂f/∂x and ∂f/∂y. To find ∂f/∂x, you treat y as a constant and differentiate with respect to x, just like in single-variable calculus. Similarly, ∂f/∂y treats x as a constant. The partial derivative measures how much the function's value changes when you nudge one input while holding the others frozen. Now imagine standing on a hillside. The slope under your feet depends on which way you face. The gradient, written as ∇f, packages all the partial derivatives into a single vector: ∇f = (∂f/∂x, ∂f/∂y). This vector points directly uphill, and its length gives the steepness of that uphill direction. If you were to walk in the direction of the gradient, you would climb the fastest. This is a remarkable leap from single-variable calculus: instead of a single slope, we have a whole compass of directions, and the gradient picks out the optimal one.
A deeper explanation
Why does the gradient point to the steepest ascent? The key is linear approximation. Near a point, a function of several variables looks like a flat plane tangent to the surface. The partial derivatives give the tilt of this plane along the coordinate axes. Any direction of travel can be expressed as a unit vector u. The instantaneous rate of change in that direction—called the directional derivative—is the dot product of the gradient with u: ∇f · u. By the properties of dot products, this quantity is maximized when u points in the same direction as ∇f, and its maximum value is the magnitude |∇f|. This is why the gradient is the direction of steepest ascent. The gradient also has a profound geometric meaning: it is perpendicular to level sets (contour lines). When you're on a hill, the contour lines represent constant altitude; to stay at the same altitude, you would walk along the contour, and the gradient points directly orthogonal to that path. This perpendicularity is why gradients appear in the method of Lagrange multipliers for constrained optimization, and why they are fundamental in physics (e.g., electric potential gradients) and in machine learning (gradient descent updates weights along the negative gradient). The gradient is the natural generalization of the derivative to multiple inputs, and it encapsulates all the first-order behavior of the function. It's the starting point for understanding higher derivatives (Hessian), curvature, and the full machinery of multivariable calculus.