Learning Rate
Goal
By the end of this lesson, you can explain the learning rate as optimization step size, recognize too-small and too-large rates from training behavior, and design a fair learning-rate comparison.
The same direction can succeed or fail because of step size
In the previous lesson, the gradient told us which local direction should reduce loss.
But knowing the direction is not enough.
Imagine walking downhill:
- with tiny steps, you may move safely but very slowly;
- with useful-sized steps, you can make steady progress;
- with giant jumps, you may cross the valley, bounce back, or move farther away.
The learning rate controls this step size.
The basic gradient-descent update is:
parameter = parameter - learning_rate × gradient
The gradient decides the direction. The learning rate scales the distance.
Compare two updates before thinking about a whole training run
Suppose the gradient is -4.
With learning rate 0.01:
change = -0.01 × (-4) = +0.04
With learning rate 0.5:
change = -0.5 × (-4) = +2
Both updates move in the same direction, but the second travels fifty times farther.
That simple calculation explains why changing the learning rate can completely change a training curve even when the gradient direction is unchanged.
Three common training patterns
A too-small learning rate often looks like:
- loss decreases;
- progress is steady;
- but many steps produce only tiny improvement.
A useful rate often looks like:
- loss decreases clearly;
- updates remain stable;
- progress reaches a low-loss region in a reasonable number of steps.
A too-large rate can look like:
- loss oscillates;
- loss suddenly grows;
- parameter values jump across the useful region;
- training becomes numerically unstable.
These are clues, not universal proofs. Other bugs can create similar symptoms, so isolate the learning rate when you test it.
Run the learning-rate comparison
The Lab uses the same one-parameter bowl loss with learning rates 0.02, 0.2, and 1.1.
- Click Run.
- Read
final losses: {0.02: 21.658119, 0.2: 0.001792, 1.1: 1878.542396}. Every run starts atstart_w = -4.0and takes the samesteps = 10. - Label each rate:
0.02is slow (the loss is still large),0.2is useful (the loss is almost zero), and1.1is unstable (the loss became much larger than where it started). - Find
learning_rates = [0.02, 0.2, 1.1]. Change only1.1to0.9. Keepstart_wandstepsunchanged so the learning rate is the only difference. - Before running, predict:
0.9is still large, but smaller than1.1. Will it explode, or settle somewhere? - Click Run. The new rate should finish at about
0.565. It no longer explodes, because each step overshoots the minimum by a little less than it did before, so the zig-zag slowly shrinks. It is still worse than0.2. - Press Reset afterward.
Loading lab…
Keeping the other conditions fixed matters because otherwise you cannot tell whether the changed behavior came from the learning rate or from something else.
Why one good learning rate is not universal
A value that works for one problem may fail for another.
Feature scale, model architecture, optimizer, batch size, and loss shape can all change how large an update is sensible.
Modern systems may also use a learning-rate schedule, changing the rate during training. Even then, the core idea remains the same: step size is a training setting that affects optimization and must be recorded if you want someone else to reproduce the run.
A common misconception
“If training is slow, just use the largest learning rate that does not crash immediately.”
A rate can avoid an obvious crash and still produce noisy or unreliable optimization. Look at the path of the loss, not just whether the program finishes.
Quick Check
Key Takeaways
- Learning rate is optimization step size.
- Too-small rates can be slow; too-large rates can overshoot or become unstable.
- Training curves reveal learning-rate problems better than one final number.
- Learning rate is part of the experiment record and should be changed in a controlled comparison.
Next Lesson
Next, you will see why the numerical scale of input features can change optimization behavior even when the underlying information is the same.
References
- scikit-learn, Stochastic Gradient Descent.
Completion is stored locally on this device.