Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Technology

The Mechanics of Gradient Descent in Machine Learning

Quick fact

Gradient descent is used to train nearly every neural network, from image recognition to language models.

Why this is interesting

You've probably seen a ball rolling down a hill to settle at the lowest point. But did you know that's essentially how machines learn?

Read the full explanation

Understanding The Mechanics of Gradient Descent in Machine Learning

Imagine you're blindfolded on a mountain and want to reach the valley. You take a step downhill by feeling the slope under your feet. That's exactly what gradient descent does. In machine learning, the 'mountain' is a mathematical surface called the cost function — it measures how wrong the model's predictions are. The goal is to find the lowest point on this surface, where the error is minimal. The model adjusts its internal parameters (like weights) step by step, moving in the direction that most reduces the error. Each step's size is controlled by the learning rate: too large and you might overshoot the valley; too small and you'll take forever to get there.

A deeper explanation

The gradient is a vector that points in the direction of steepest increase of the cost function. To minimize, we move in the opposite direction: the negative gradient. Mathematically, we update each parameter θ by subtracting the gradient times the learning rate: θ ← θ − η·∇J(θ). Here, J is the cost function, η is the learning rate, and ∇J is the gradient. Repeating this process iteratively brings the parameters to a minimum. However, the cost surface can have multiple valleys (local minima). Depending on the starting point and step size, the algorithm might settle into a local minimum, missing the global one. Variants like stochastic gradient descent use random subsets of data to speed up computation and escape some local minima. Understanding these mechanics is crucial because they underpin how models learn from data, and tweaking the learning rate or initialization can drastically change the outcome.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.