Skip to main content
L2.9

Backpropagation in Code

Goal

By the end of this lesson, you can compute a simple backpropagation gradient in code, update a parameter with gradient descent, and verify the update using both direction reasoning and the new loss.

Use the smallest trainable model possible​

Take this model:

prediction = w*x

and squared loss:

loss = (prediction - target)²

There is only one trainable parameter, w.

The dependency chain is:

w -> prediction -> loss

That makes every derivative visible.

Compute the local pieces​

First:

dLoss/dPrediction = 2*(prediction - target)

Then:

dPrediction/dw = x

Chain them:

dLoss/dw = dLoss/dPrediction * dPrediction/dw

This product is the gradient for w.

Now gradient descent updates:

w_new = w - learning_rate * dLoss/dw

The whole training step is just forward computation, local sensitivities, chain rule, and an update.

Reason about the sign before using the formula​

Suppose the current prediction is too low and x is positive.

To increase the prediction, w should increase.

A correct gradient/update combination should produce that direction for a sufficiently small learning rate.

This qualitative check catches many sign errors before you trust the exact number.

Complete one missing chain-rule step​

This is the first Browser Lab in which the starter is supposed to fail some checks until you write a small piece of code.

  1. Click Run unchanged. The starter prints the two local derivatives, but grad_w is still 0.0. The weight does not move and the learner checks report that the exercise is incomplete.
  2. Find:
# TODO: use the chain rule to combine the two local derivatives.
grad_w = 0.0
  1. Replace only the 0.0 expression with the product of d_loss_d_prediction and d_prediction_d_w.
  2. Before running, calculate the expected gradient by hand. It should be negative for this example, so gradient descent should move w upward.
  3. Click Run. Confirm the gradient is -2.4, the weight increases, the loss decreases, and all learner checks pass.
  4. After that works, cut the learning rate in half and predict how the update distance changes.

Loading lab…

A deliberate sign failure​

If you replace the update with:

w = w + learning_rate * gradient

then the parameter moves with the uphill gradient instead of against it.

For the starter example, loss should increase rather than decrease.

This is a useful debugging experiment because the code still runs. The failure is mathematical, not syntactic.

If the final gradient is wrong, inspect:

  • prediction;
  • loss;
  • dLoss/dPrediction;
  • dPrediction/dw;
  • their product;
  • parameter before and after update.

Stop at the first value that disagrees with your hand reasoning.

Backpropagation reuses intermediate results​

In a large network, calculating each parameter's effect from scratch would repeat enormous amounts of work.

Backpropagation is efficient because it moves backward through the computation graph and reuses already-computed local derivatives.

For the tiny chain:

w → prediction → loss

you calculate the sensitivity of loss to prediction, then combine it with the sensitivity of prediction to w.

In a deeper network, the same principle repeats across many nodes.

This is why the forward pass and backward pass are tightly connected: the backward pass follows dependencies created by the forward computation.

A syntax-correct program can still have a mathematically wrong gradient. Direction checks, finite-difference checks, and loss-after-update checks provide independent evidence instead of trusting one implementation path.

Quick Check

1. For gradient descent, how is a parameter updated?
2. Why print local derivative pieces?
3. If one supposedly correct small update raises loss unexpectedly, what should you inspect?

0 of 3 questions answered.

Key Takeaways

  • Backpropagation combines local derivatives along the forward dependency path.
  • A parameter gradient can be checked by sign reasoning and finite differences.
  • Gradient descent subtracts a scaled gradient.
  • One update should be validated by inspecting both parameter movement and loss movement.

Next Lesson

You have reached the Level 2 mini checkpoint. After it, you will move from one-example-style reasoning to batches and practical training behavior.

References

Lesson actions

Completion is stored locally on this device.

View progress