Backpropagation Intuition
Goal
By the end of this lesson, you can explain backpropagation as local sensitivity passed backward through a dependency chain and use a small perturbation to infer a useful parameter direction.
Start with the question backpropagation must answer
A network produces a loss at the end of a forward pass.
Training needs to know:
If I change this earlier weight a tiny amount, how will the final loss change?
A large network has many parameters, so trying every combination would be hopelessly expensive.
Backpropagation answers the question efficiently by reusing the computation dependencies from the forward pass.
A small perturbation makes the idea visible
Suppose one weight is currently w.
Compute the loss at:
w - hw + h
for a small h.
If the loss is lower at w + h, then increasing the weight locally appears useful.
If the loss is lower at w - h, decreasing the weight locally appears useful.
This is a finite-difference way to observe gradient direction. It is slow for a real network, but excellent for building intuition and checking gradient code.
Local effects multiply through a chain
Imagine this dependency:
w -> prediction -> loss
The loss changes when prediction changes, and prediction changes when w changes.
The chain rule combines those two local sensitivities:
dLoss/dw = dLoss/dPrediction × dPrediction/dw
For a longer network, the chain contains more operations, but the reasoning is the same.
If a value branches into several downstream paths, the gradient contributions from those paths add back together.
Backpropagation: forward to the loss, backward to the weight
The forward pass creates a chain from the weight to the loss. Backpropagation follows that chain backward to find how the weight affected the loss.
Forward pass: prediction = 2 × 1 = 2
The visual uses the derivative for this tiny squared-error example. The next lesson derives the derivative rules behind that calculation; here, focus on the direction of information flow and the sign of the gradient.
Probe one weight in the Lab
The Lab uses the simple function:
prediction = weight * 2.0
loss = (prediction - 5.0) ** 2
The best weight for this tiny problem is 2.5 because 2.5 × 2 = 5.
- Click Run with
w = 1.0andepsilon = 1e-4. - Read
finite-difference gradient:andgradient descent should move weight:. The gradient should be negative, so the useful local direction is to increase the weight toward2.5. - Change only
w = 1.0tow = 3.0. - Before running, reason from the target:
3.0 × 2 = 6, which is now above the target5. Predict that the gradient sign should become positive and gradient descent should move the weight downward. - Click Run and confirm the direction changes to
decrease. - Restore
w = 1.0.
Loading lab…
A useful special case about epsilon
You can also change epsilon = 1e-4 to epsilon = 0.5 and rerun with w = 1.0. In this particular Lab, the centered finite-difference estimate stays exactly the same because the loss is a quadratic function of w. That is a special property of this toy—not a rule that large perturbations are always safe.
For more complicated curved functions, a large epsilon can compare points that are too far apart to represent the slope right around the current value. The word local still matters.
A finite-difference check is evidence about the gradient. It is not the efficient training algorithm itself.
A common misconception
“Backpropagation sends the final prediction error unchanged to every parameter.”
Different parameters influence the output through different paths and local operations. A parameter multiplied by a zero activation can receive a very different gradient from a parameter on an active path.
That is why local sensitivity matters.
Debug gradient direction one link at a time
If a gradient sign surprises you, draw the dependency chain.
For each link ask:
- if the earlier value increases, does the later value increase or decrease?
- is that local effect positive, negative, or zero?
- how do the signs combine along the path?
This reasoning is often enough to find a reversed sign before reading a long formula.
Quick Check
Key Takeaways
- Backpropagation asks how each earlier value affects the final loss.
- The chain rule combines local sensitivities through dependency paths.
- Finite differences make gradient direction visible and can check implementation.
- A centered difference happens to be exact for this quadratic toy even with a larger epsilon; that does not remove the need for local checks on more complex functions.
- Gradients are path-specific; the final error is not copied unchanged to every parameter.
Next Lesson
Next, you will learn exactly enough derivative math to describe those local sensitivities directly.
References
- PyTorch, Autograd.
Completion is stored locally on this device.