Skip to main content
L2.14

Training Curves and Debugging

Goal

By the end of this lesson, you can read training and validation curves, distinguish several common failure patterns, and choose a next diagnostic experiment from the evidence.

A final score hides the path​

Imagine two training runs that both end with validation loss 0.5.

Run A decreases smoothly from 2.0 to 0.5.

Run B drops to 0.2, explodes to 10, becomes unstable, then happens to end at 0.5.

The final number is identical. The training histories tell very different stories.

A training curve records a metric such as loss across epochs or steps so you can see when behavior changed.

Four patterns worth recognizing​

1. Training and validation both improve and stay reasonably close​

This often indicates useful learning, though you still need task-specific evaluation.

2. Training improves while validation gets worse​

A widening gap is a classic warning sign of overfitting: the model is fitting training-specific details that do not transfer.

3. Training and validation stay high and nearly flat​

Possible causes include:

  • model too limited;
  • bad features or labels;
  • gradients too small or wrong;
  • learning rate too small;
  • a broken forward path.

The curve suggests a symptom, not one unique cause.

4. Loss oscillates, explodes, or becomes non-finite​

Suspect update size, numerical instability, invalid loss inputs, or gradient problems.

Reveal the training history

Move through the epochs and compare the solid training-loss line with the dashed validation-loss line. The labels and values carry the same meaning without relying on color.

Epoch 8 · train 0.579 · validation 0.516
Training loss Validation loss
epochloss

The visual can show a symptom pattern, but it cannot tell you the root cause. Use the curve to choose the next check rather than treating the chart as an automatic diagnosis.

Practice diagnosis in the Lab​

The Lab contains deterministic synthetic curves and a small diagnostic rule.

  1. Before clicking Run, read the six lists at the top of the Lab. They are the curves.
  2. In overfit_train = [1.0, 0.70, 0.45, 0.25, 0.12] and overfit_val = [1.05, 0.78, 0.65, 0.72, 0.88], find the point where they start moving in different directions: after the third value, training keeps falling while validation rises from 0.65 to 0.72 and 0.88.
  3. Click Run. The output is {'healthy': 'healthy_progress', 'overfit': 'overfitting', 'unstable': 'unstable'}.
  4. Now change only the final validation value so it improves again: overfit_val = [1.05, 0.78, 0.65, 0.72, 0.60].
  5. Before running, predict whether the rule will still call this curve overfitting.
  6. Click Run. The label becomes healthy_progress, because the rule only compares the last validation point with earlier ones. A person reading the whole curve would still notice the rise in the middle. This is why the rule is only a helper, not an authority.
  7. Press Reset afterward.

Loading lab…

The rule in the Lab is intentionally simple. Real debugging should use the curves as evidence, not hand authority to a tiny automatic diagnosis.

Curves describe symptoms, not causes​

If validation loss rises, “overfitting” describes a pattern in measured behavior. It does not by itself explain whether the root cause is:

  • too much capacity;
  • too little representative data;
  • weak regularization;
  • distribution mismatch;
  • leakage elsewhere in the experiment.

A useful next step tests one plausible cause while keeping the rest of the setup fixed.

Keep enough metadata to reproduce the curve​

A screenshot of a curve without its experiment settings is hard to debug later.

Record at least:

  • dataset/version and split;
  • random seed;
  • architecture;
  • initialization;
  • optimizer and learning rate;
  • batch size;
  • regularization;
  • checkpoint or code revision.

The curve becomes much more useful when another run can recreate it.

Read curves as a timeline of evidence​

A final loss value hides the path taken to reach it.

Compare two runs:

Run A: 2.0 → 1.2 → 0.8 → 0.6
Run B: 2.0 → 5.0 → 20.0 → NaN

The second run tells you much more than “final loss is bad.” The rapid explosion suggests checking learning rate, gradient magnitude, numerical stability, or bad input scaling.

Another pattern is:

train loss: keeps decreasing
validation loss: decreases, then rises

That is evidence consistent with overfitting, assuming the validation pipeline is trustworthy.

Curves suggest hypotheses; they do not prove causes​

A flat curve might mean:

  • learning rate is too small;
  • gradients are zero;
  • parameters are not being updated;
  • the task/data provide little learnable signal;
  • a logging bug is showing stale values.

Use the curve to choose the next diagnostic check, then inspect the relevant intermediate evidence.

Debugging improves when each plot leads to a testable hypothesis instead of a random parameter change.

Quick Check

1. Training loss falls while validation loss rises. What pattern is most plausible?
2. Why record curves rather than only final loss?
3. What is a good debugging change policy?

0 of 3 questions answered.

Key Takeaways

  • Training curves expose dynamics hidden by a final score.
  • A widening training/validation gap can indicate overfitting.
  • Flat or unstable curves can have several causes; they are symptoms, not diagnoses.
  • Use a specific hypothesis and one controlled change for the next experiment.
  • Record enough configuration to reproduce the curve.

Next Lesson

Next, you will combine the whole Level 2 evidence chain into a structured neural-network debugging workshop.

References

Lesson actions

Completion is stored locally on this device.

View progress