본문으로 건너뛰기
L2.11

Optimizers: SGD to Adam

Goal

By the end of this lesson, you can explain what an optimizer adds beyond a gradient, compare plain SGD and Adam on the same objective, and recognize when changing optimizers would only hide a deeper bug.

The gradient gives evidence; the optimizer turns it into an update​

Backpropagation gives each parameter a gradient: a local direction and sensitivity for the current loss.

An optimizer decides how to turn that gradient information into a parameter update over time.

The simplest case is plain stochastic gradient descent, or SGD:

parameter = parameter - learning_rate * gradient

SGD can work very well. It does not need to remember Adam-style moving averages; it can use the current gradient directly.

Why keep optimizer state?​

Imagine one parameter receives gradients that change scale from step to step.

Adam keeps running summaries of:

  • the gradient itself;
  • the squared gradient.

Those summaries let it adapt effective update sizes across parameters and across time.

That can make early optimization easier on some problems, but Adam does not change the meaning of a wrong gradient. If the forward pass, loss, or gradient is broken, a more sophisticated optimizer still receives broken evidence.

Compare trajectories, not just final loss​

Suppose SGD and Adam both start from the same parameter value on the same bowl-shaped objective.

A useful comparison keeps fixed:

  • starting parameter;
  • objective;
  • number of steps;
  • random data, if any.

Then inspect the sequence of parameter values and losses.

A final loss can hide whether one optimizer approached smoothly, oscillated, or took a large detour.

Run the optimizer comparison​

The Lab minimizes the same simple quadratic with SGD and Adam.

  1. Click Run.
  2. Read the first steps before the final values. The minimum is at w = 3. SGD first steps: [-4.0, -1.2, 0.48, 1.488] walks steadily toward it, while Adam first steps: [-4.0, -3.5, -3.001, -2.505] moves about 0.5 per step, set by its learning rate.
  3. Find sgd_lr = 0.2 and change only 0.2 to 0.9.
  4. Before running, predict whether the first SGD updates will move farther and whether they may overshoot the minimum at 3.
  5. Click Run. Now SGD first steps: [-4.0, 8.6, -1.48, 6.584]: the very first step jumps past 3 to 8.6, then back below it. The final loss (0.231396) is worse than before even though each step was larger.
  6. Press Reset afterward.

Loading lab…

If loss begins bouncing or growing, that is evidence about update size. It is not evidence that SGD as an algorithm is universally bad.

What to check before switching optimizers​

When training behaves badly, inspect these boundaries first:

  1. Is the loss finite?
  2. Are the gradients finite?
  3. Do gradient signs make sense on a tiny case?
  4. Are parameter updates actually happening?
  5. Is the learning rate too large or too small?
  6. Is the data scale reasonable?

Only then is optimizer choice a well-grounded experimental variable.

A common misconception​

“Adam is smarter, so it should fix unstable or NaN gradients.”

Adam can rescale updates, but it cannot make an invalid forward computation or undefined gradient valid. Numerical failures should be diagnosed at their source.

A little more technical: Adam's state​

Adam tracks moving estimates often called first and second moments. Because those estimates start at zero, implementations also use bias correction in early steps.

You do not need to memorize Adam's complete equation yet. The important idea is that optimizer state affects future updates, so resetting or restoring an optimizer is different from restoring only model weights.

Quick Check

1. What state does plain SGD require beyond current parameters?
2. What can a too-large learning rate cause?
3. Should optimizer switching be the first response to NaN gradients?

0 of 3 questions answered.

Key Takeaways

  • Gradients describe local loss sensitivity; optimizers turn that evidence into updates.
  • Plain SGD uses the current gradient directly; Adam keeps adaptive state.
  • Compare optimizer trajectories under the same starting conditions.
  • Learning rate remains important regardless of optimizer.
  • Do not use a new optimizer to hide broken forward or gradient computations.

Next Lesson

Next, you will look at the state before the first optimizer step: how initial parameter values affect symmetry and signal scale.

References

Lesson actions

Completion is stored locally on this device.

View progress