Mathematics
The Gradient Descent Algorithm from a Calculus Perspective
Quick fact
Gradient descent is the backbone of training neural networks: every time a model learns from data, it's essentially performing thousands of tiny downhill steps guided by the gradient of a loss function—the calculus at the heart of artificial intelligence.
Why this is interesting
If you're on a foggy mountain and need to find the valley, your first instinct is to take a step downhill—but how do you mathematically know which way is down? That's exactly the problem gradient descent solves, using nothing but calculus.
Read the full explanation
Understanding The Gradient Descent Algorithm from a Calculus Perspective
Imagine you're hiking in a fog. You can't see the whole landscape, but you can feel the slope beneath your feet. To reach the lowest point, you take a step in the direction where the ground descends most steeply. Gradient descent formalizes this intuition: given a function f(x, y) (like a cost function in machine learning), its gradient is a vector that points in the direction of steepest increase. To minimize, we move opposite to the gradient—the direction of steepest decrease. We do this step-by-step: start at some point, compute the gradient at that point, take a small step in the opposite direction, then repeat. The size of each step is controlled by a parameter called the learning rate (often denoted α). This process is the essence of training many AI models.
A deeper explanation
Mathematically, for a function f(x) (where x is a vector), the gradient ∇f(x) is the vector of partial derivatives. The directional derivative along a unit vector u is ∇f(x)·u. To minimize, we choose u = -∇f(x)/‖∇f(x)‖, giving the steepest descent. The update rule is: x{new} = x - α∇f(x). This is essentially a first-order Taylor approximation: f(x + h) ≈ f(x) + ∇f(x)·h. For small h, the linear term dominates, so setting h = -α∇f(x) (for small α) guarantees that f decreases. Under suitable conditions—like f being differentiable and the learning rate not too large—this iterative process converges to a local minimum. The calculus perspective reveals that the gradient provides local linear information; the algorithm uses only this first-order info, which is why it is simple but can be slow. If the function is non-convex, gradient descent might get stuck in a suboptimal local minimum, because the gradient is zero there (like at the bottom of a valley, but not the global lowest point). The learning rate is crucial: too large, and you overshoot, increasing f; too small, and you converge very slowly. This is why adaptive learning rates (like Adam) are popular in practice.