Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Mathematics

Gradient Descent for Multivariable Functions

Quick fact

Gradient descent is so fundamental that virtually every machine learning model, from linear regression to deep neural networks, uses some variant of it to learn from data.

Why this is interesting

Imagine you're blindfolded on a hillside and need to find the lowest point in the valley. You can only feel the slope under your feet—how would you step?

Read the full explanation

Understanding Gradient Descent for Multivariable Functions

Gradient descent is an iterative algorithm for finding the minimum of a function that has multiple input variables. Think of the function as a hilly landscape, where the height at any point is the function's value. Your goal is to reach the lowest valley. However, you can't see the whole landscape; you can only observe the slope at your current position. The gradient of the function at that point is a vector that points in the direction of the steepest ascent (uphill). To go downhill, you move in the opposite direction—the negative gradient. You take a step of a certain size (determined by the learning rate) and then recalculate the gradient at your new position. Repeat this process many times, and you will gradually approach a local minimum. For example, if you're minimizing f(x, y) = x^2 + y^2, the gradient is (2x, 2y). Starting at (3, 4), the negative gradient is (-6, -8), so you move toward (0, 0), the global minimum.

A deeper explanation

The algorithm hinges on the principle that the gradient of a differentiable multivariable function points in the direction of greatest increase. By updating the current point w as w ← w − η∇f(w), we take a step proportional to the negative gradient, where η (the learning rate) controls the step size. A small η ensures stable convergence but may require many iterations; a large η can overshoot the minimum or cause divergence. The process relies on the function being differentiable, and for convex functions, any local minimum is also global, making convergence reliable. In non-convex landscapes, gradient descent can get stuck in local minima or saddle points, prompting more advanced versions. This method is the workhorse of machine learning: training a model involves minimizing a cost function that measures prediction error, and gradient descent iteratively adjusts model parameters to reduce that error. Its efficiency in high-dimensional spaces makes it indispensable, and its variants (stochastic, mini-batch, momentum) address practical challenges.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.