Backpropagation in Code
Goal
By the end of this lesson, you can compute a simple backpropagation gradient in code, update a parameter with gradient descent, and verify the update using both direction reasoning and the new loss.
Use the smallest trainable model possible
Take this model:
prediction = w*x
and squared loss:
loss = (prediction - target)²
There is only one trainable parameter, w.
The dependency chain is:
w -> prediction -> loss
That makes every derivative visible.
Compute the local pieces
First:
dLoss/dPrediction = 2*(prediction - target)
Then:
dPrediction/dw = x
Chain them:
dLoss/dw = dLoss/dPrediction * dPrediction/dw
This product is the gradient for w.
Now gradient descent updates:
w_new = w - learning_rate * dLoss/dw
The whole training step is just forward computation, local sensitivities, chain rule, and an update.
Reason about the sign before using the formula
Suppose the current prediction is too low and x is positive.
To increase the prediction, w should increase.
A correct gradient/update combination should produce that direction for a sufficiently small learning rate.
This qualitative check catches many sign errors before you trust the exact number.
Complete one missing chain-rule step
This is the first Browser Lab in which the starter is supposed to fail some checks until you write a small piece of code.
- Click Run unchanged. The starter prints the two local derivatives, but
grad_wis still0.0. The weight does not move and the learner checks report that the exercise is incomplete. - Find:
# TODO: use the chain rule to combine the two local derivatives.
grad_w = 0.0
- Replace only the
0.0expression with the product ofd_loss_d_predictionandd_prediction_d_w. - Before running, calculate the expected gradient by hand. It should be negative for this example, so gradient descent should move
wupward. - Click Run. Confirm the gradient is
-2.4, the weight increases, the loss decreases, and all learner checks pass. - After that works, cut the learning rate in half and predict how the update distance changes.
Loading lab…
A deliberate sign failure
If you replace the update with:
w = w + learning_rate * gradient
then the parameter moves with the uphill gradient instead of against it.
For the starter example, loss should increase rather than decrease.
This is a useful debugging experiment because the code still runs. The failure is mathematical, not syntactic.
Print derivative pieces, not only the final gradient
If the final gradient is wrong, inspect:
- prediction;
- loss;
dLoss/dPrediction;dPrediction/dw;- their product;
- parameter before and after update.
Stop at the first value that disagrees with your hand reasoning.
Backpropagation reuses intermediate results
In a large network, calculating each parameter's effect from scratch would repeat enormous amounts of work.
Backpropagation is efficient because it moves backward through the computation graph and reuses already-computed local derivatives.
For the tiny chain:
w → prediction → loss
you calculate the sensitivity of loss to prediction, then combine it with the sensitivity of prediction to w.
In a deeper network, the same principle repeats across many nodes.
This is why the forward pass and backward pass are tightly connected: the backward pass follows dependencies created by the forward computation.
A syntax-correct program can still have a mathematically wrong gradient. Direction checks, finite-difference checks, and loss-after-update checks provide independent evidence instead of trusting one implementation path.
Quick Check
Key Takeaways
- Backpropagation combines local derivatives along the forward dependency path.
- A parameter gradient can be checked by sign reasoning and finite differences.
- Gradient descent subtracts a scaled gradient.
- One update should be validated by inspecting both parameter movement and loss movement.
Next Lesson
You have reached the Level 2 mini checkpoint. After it, you will move from one-example-style reasoning to batches and practical training behavior.
References
- PyTorch, Autograd.
Completion is stored locally on this device.