Mathematics
Using the Chain Rule in Multivariable Calculus
Quick fact
The multivariable chain rule is the engine behind backpropagation in neural networks, allowing millions of parameters to be updated in a single training step.
Why this is interesting
You know how to differentiate a function of one variable — but what happens when a change in one variable triggers a cascade of changes through many others? How do you find the total rate of change when everything is connected?
Read the full explanation
Understanding Using the Chain Rule in Multivariable Calculus
Imagine a factory assembly line where each station transforms a product. In single-variable calculus, the chain rule tells you how a change at the start affects the final output when there's one line. In multivariable calculus, imagine a network of pipes and valves, where water flows through branches that merge. If you turn a valve at one point, how does the flow at the end change? The chain rule gives you a systematic way to add up all the contributions along every path from the starting variable to the final function. For a function f(x,y) where x and y themselves depend on another variable t, the derivative of f with respect to t is the sum of the partial derivatives of f with respect to x and y, each multiplied by the derivative of that variable with respect to t: df/dt = (∂f/∂x)(dx/dt) + (∂f/∂y)(dy/dt). It's like saying: 'The total change comes from how f changes due to x, scaled by how x changes due to t, plus the same for y.'
A deeper explanation
The multivariable chain rule emerges naturally from the concept of the total differential. For a function f(x,y), the total differential is df = (∂f/∂x)dx + (∂f/∂y)dy. If x and y are functions of t, then dividing by dt gives the chain rule. Underneath, this is a direct application of the chain rule for limits: the change in f when x and y change slightly is the sum of the changes due to each variable, and each of those changes is proportional to the change in that variable. The rule extends to any number of variables, and when f is a function of variables that are themselves functions of many variables, the formula becomes a sum over all paths in a dependency graph. This structure is what makes the chain rule so powerful: it allows us to compute derivatives of compositions without explicitly substituting expressions, which is essential in optimization (e.g., gradient descent) and in machine learning's backpropagation, where the chain rule efficiently computes gradients layer by layer. Thus, the chain rule is not just a formula; it's the fundamental law of how change propagates through interconnected systems.