본문으로 건너뛰기
L2.13

Regularization

Goal

By the end of this lesson, you can explain L2 regularization as a change to the training objective, compare its effect on weight size and validation behavior, and recognize when regularization has become too strong.

A flexible model can fit more than the useful pattern​

Suppose a model has enough freedom to match every wiggle in a small training set.

Some of those wiggles may reflect stable structure. Others may be sampling noise.

Regularization adds a preference that makes some fitted solutions less attractive, often encouraging a model not to use unnecessary complexity as aggressively.

L2 adds a cost for large weights​

A simplified L2-regularized objective looks like:

total_loss = data_loss + lambda * sum(weights²)

There are now two pressures:

  1. fit the training examples;
  2. avoid unnecessarily large parameter values.

As lambda grows, the penalty matters more strongly.

The model may accept slightly worse training fit in exchange for a parameter setting that performs better on validation data.

That tradeoff is the point. Regularization is not supposed to make training loss as small as possible at any cost.

Compare training fit, validation fit, and weight norm together​

Imagine three settings:

regularizationtraining errorvalidation errorweight norm
nonevery lowhighlarge
moderatea little higherlowersmaller
extremehighhightiny

The middle setting may generalize better even though its training error is not the smallest.

The extreme setting shows underfitting: the penalty has become strong enough to prevent the model from representing the useful pattern.

Run the regularization Lab​

The Lab fits a flexible polynomial-style linear model with and without an L2 penalty.

  1. Click Run. The starter uses lambda_l2 = 0.1.
  2. Read each line as “without penalty → with penalty.” weight norm: 1.3559 -> 0.8316 shows smaller weights. validation MSE: 0.010455 -> 0.006228 shows that this moderate penalty also improved held-out error.
  3. Find lambda_l2 = 0.1 and change only 0.1 to 5.0.
  4. Before running, predict: will the weight norm shrink further? Will validation error keep improving?
  5. Click Run. The weight norm drops to 0.3441, but validation MSE jumps to 0.211695, far worse than with no penalty. The model is now too constrained to follow the real pattern—the extreme row in the table above.
  6. Press Reset afterward.

Loading lab…

Do not judge the choice from weight norm alone. Small weights are not the final goal; useful held-out behavior is.

Choose regularization without peeking at the final test set​

Regularization strength is a model-selection choice.

Use training/validation evidence or cross-validation to choose it. If the final test set repeatedly guides lambda, that test set is no longer an untouched final evaluation.

This reuses the information-boundary reasoning from Level 1.

A common misconception​

“If some regularization helps, more regularization should help more.”

Too much pressure can erase useful structure. If both training and validation performance are poor, increasing regularization is unlikely to be the right fix.

Quick Check

1. What does L2 regularization penalize?
2. Can too much regularization hurt?
3. What evidence should guide regularization strength?

0 of 3 questions answered.

Key Takeaways

  • Regularization changes the training objective, not merely the reporting metric.
  • L2 discourages large parameter magnitudes.
  • Slightly worse training fit can accompany better validation behavior.
  • Too much regularization can cause underfitting.
  • Choose regularization strength using validation evidence, not the final test set.

Next Lesson

Next, you will learn to read the history of training and validation loss as diagnostic evidence rather than looking only at a final score.

References

Lesson actions

Completion is stored locally on this device.

View progress