Gradient Descent
Goal
By the end of this lesson, you can describe gradient descent as repeated parameter updates toward lower loss and trace several updates on a one-parameter model.
How do we know which way to change a parameter?
Suppose our model has one adjustable value, w, and the loss is:
loss = (w - 3)²
The loss is smallest when w = 3.
If we start at w = 0, we want the parameter to move upward. If we start at w = 5, we want it to move downward.
In a simple problem like this, we can see the answer. A real model may have thousands or millions of adjustable parameters, so trying every possible setting is not practical.
Gradient descent gives us a local rule for choosing a direction.
The gradient tells us the local uphill direction
For the toy loss (w - 3)², the gradient is:
2(w - 3)
You do not need to derive that expression yet. For this Level, treat it as a small direction calculator: plug in the current w, read the sign, and use that sign to decide which way loss rises locally. In Level 2, you will build the derivative idea from small changes and see where formulas like this come from.
At w = 0, the gradient is -6.
Gradient descent uses this update:
new_w = old_w - learning_rate × gradient
With learning rate 0.1:
new_w = 0 - 0.1 × (-6) = 0.6
Subtracting a negative value moves w upward, toward 3.
At a value above 3, the gradient becomes positive, so subtracting it moves w downward.
Follow the same parameter update
The target is w = 3, matching the worked example above. Change the learning rate, then step through the path and watch both w and loss.
A useful picture is standing on a foggy hill. You cannot see the whole landscape, but you can feel the local slope. The gradient points uphill, so gradient descent takes a step in the opposite direction.
Follow several updates in the Lab
The Lab starts with start_w = -4.0 and repeatedly minimizes the same bowl-shaped loss.
- Click Run without editing anything.
- Inspect
first step:,final w:, andfinal loss:. The first gradient should be negative, so the first update moveswupward toward3. - Find the line
start_w = -4.0. - Change only
-4.0to5.0. - Before running, evaluate the sign of
2(w - 3)atw = 5: it is positive. Predict which direction the first update should move. - Click Run and inspect
first step:. Confirm that the positive gradient makes gradient descent movewdownward toward3. - Restore
start_w = -4.0.
Loading lab…
After the guided comparison, you may try another starting value. Keep target_w, learning_rate, and steps fixed so the starting point is the only changed factor.
The goal is to understand each update, not just to see a final small number.
Direction is only half the problem
The gradient tells us a local direction. We still need to decide how large a step to take. That is the job of the learning rate, which is the next lesson.
If steps are too large, the parameter can jump across a low-loss region and become unstable. If they are too small, progress can be painfully slow.
Also remember that decreasing training loss only shows that optimization is working on the training objective. It does not prove that the model generalizes to held-out data.
Debug optimization by looking at the path
When training behaves strangely, inspect the sequence rather than only the final result:
- Is the loss decreasing, flat, oscillating, or exploding?
- Are parameter updates moving in the direction you expected?
- Is the sign in the update rule correct?
- Is the gradient formula or automatic-gradient computation correct?
- Is the learning rate appropriate for the scale of the problem?
These questions turn “training failed” into specific evidence you can investigate.
Quick Check
Key Takeaways
- Gradient descent repeatedly updates parameters using local loss information.
- The gradient points toward increasing loss; gradient descent moves in the opposite direction.
- You can use the gradient formula here without deriving it yet; Level 2 builds that derivative idea from small changes.
- The learning rate controls how far each update moves.
- Training-loss progress is optimization evidence, not proof of held-out performance.
Next Lesson
Next, you will focus on the number that controls the size of every gradient step: the learning rate.
References
- scikit-learn, Stochastic Gradient Descent.
Completion is stored locally on this device.