Skip to main content
L1.16

Bias, Variance, and Overfitting

Goal

By the end of this lesson, you can recognize common underfitting and overfitting patterns from training and held-out evidence, and explain model flexibility as a tradeoff rather than a race toward maximum complexity.

A model can fail by being too stiff or too flexible​

Imagine trying to trace a curved path with two different rulers.

A very stiff ruler cannot bend enough to follow the curve. It misses important structure even on the examples you used to fit it.

A wildly flexible ruler can bend through every training point, including accidental noise. It can look perfect on those points but fail to follow the underlying path elsewhere.

These two ideas connect to bias and variance:

  • high bias often appears when a model is too limited for the important pattern;
  • high variance often appears when a model is too sensitive to details of the particular training sample.

The practical beginner terms are underfitting and overfitting.

Read training and held-out scores together​

Suppose Model A scores:

  • training: 60%
  • test: 58%

It performs poorly even on the training data. That is evidence of underfitting or another fundamental problem.

Now suppose Model B scores:

  • training: 100%
  • test: 72%

The large gap is a warning sign that the model may have fitted training-specific details that do not transfer well.

Neither pattern is a proof by itself. Data leakage, distribution shift, label noise, or a tiny test set can also create unusual scores.

Compare model capacity in the Lab​

The Lab creates a deterministic noisy classification dataset and compares a shallow decision tree with a very deep tree.

  1. Click Run.
  2. Read the scores as pairs: shallow_train: 0.837 with shallow_test: 0.75, and deep_train: 1.0 with deep_test: 0.821.
  3. Read deep train-test gap: 0.179. The deep tree memorized every training example (1.0) but scores much lower on held-out data. The shallow tree (max_depth=1, a single split) underfits: it is weak on both sets.
  4. Find shallow = DecisionTreeClassifier(max_depth=1, random_state=42). Change only max_depth=1 to max_depth=3.
  5. Before running, predict: with more capacity, should the shallow tree's training score rise? What about its test score?
  6. Click Run. The depth-3 tree reaches shallow_train: 0.952 and shallow_test: 0.804. Both scores improved, and its gap (about 0.15) is still a little smaller than the unlimited tree's.
  7. Press Reset afterward.

Loading lab…

A useful model is not necessarily the one with the highest training score. The goal is performance that transfers beyond the examples used for fitting.

Complexity is only one possible explanation​

Before blaming model capacity, check the experiment itself:

  • Did target or future information leak into features?
  • Does the split represent the real future task?
  • Is the test set large and varied enough to be informative?
  • Are labels noisy or inconsistent?
  • Does the same pattern appear across cross-validation folds?

Only after those checks does changing model complexity become a well-grounded debugging step.

Ways to respond to overfitting​

Depending on the evidence, possible responses include:

  • limiting model capacity;
  • using regularization;
  • collecting more representative data;
  • removing noisy or invalid features;
  • improving the split or evaluation procedure.

There is no universal “make the model simpler” button that solves every generalization problem.

A common misconception​

“A small train/test gap is always good.”

If both scores are poor, a small gap can simply mean the model underfits everywhere. Read the absolute performance and the gap together.

Quick Check

1. What often indicates underfitting?
2. What often indicates overfitting?
3. What should you check before blaming model complexity?

0 of 3 questions answered.

Key Takeaways

  • Underfitting often means the model cannot represent the pattern well enough.
  • Overfitting often means training-specific details do not generalize.
  • Training and held-out scores must be interpreted together.
  • A train/test gap is a clue, not a complete diagnosis.
  • Change model complexity only after checking data and evaluation boundaries.

Next Lesson

Next, you will use pipelines to keep preprocessing and model fitting in a reproducible order and make accidental leakage harder.

References

Lesson actions

Completion is stored locally on this device.

View progress