Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

The Gradient Descent Algorithm from a Calculus Perspective

Quick fact

Gradient descent is the backbone of training neural networks: every time a model learns from data, it's essentially performing thousands of tiny downhill steps guided by the gradient of a loss function—the calculus at the heart of artificial intelligence.

Why this is interesting

If you're on a foggy mountain and need to find the valley, your first instinct is to take a step downhill—but how do you mathematically know which way is down? That's exactly the problem gradient descent solves, using nothing but calculus.

Read the full explanation

Understanding The Gradient Descent Algorithm from a Calculus Perspective

Imagine you're hiking in a fog. You can't see the whole landscape, but you can feel the slope beneath your feet. To reach the lowest point, you take a step in the direction where the ground descends most steeply. Gradient descent formalizes this intuition: given a function f(x, y) (like a cost function in machine learning), its gradient is a vector that points in the direction of steepest increase. To minimize, we move opposite to the gradient—the direction of steepest decrease. We do this step-by-step: start at some point, compute the gradient at that point, take a small step in the opposite direction, then repeat. The size of each step is controlled by a parameter called the learning rate (often denoted α). This process is the essence of training many AI models.

A deeper explanation

Mathematically, for a function f(x) (where x is a vector), the gradient ∇f(x) is the vector of partial derivatives. The directional derivative along a unit vector u is ∇f(x)·u. To minimize, we choose u = -∇f(x)/‖∇f(x)‖, giving the steepest descent. The update rule is: x{new} = x - α∇f(x). This is essentially a first-order Taylor approximation: f(x + h) ≈ f(x) + ∇f(x)·h. For small h, the linear term dominates, so setting h = -α∇f(x) (for small α) guarantees that f decreases. Under suitable conditions—like f being differentiable and the learning rate not too large—this iterative process converges to a local minimum. The calculus perspective reveals that the gradient provides local linear information; the algorithm uses only this first-order info, which is why it is simple but can be slow. If the function is non-convex, gradient descent might get stuck in a suboptimal local minimum, because the gradient is zero there (like at the bottom of a valley, but not the global lowest point). The learning rate is crucial: too large, and you overshoot, increasing f; too small, and you converge very slowly. This is why adaptive learning rates (like Adam) are popular in practice.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.