Loss Functions
Goal
By the end of this lesson, you can compare absolute and squared error, explain why a loss function changes which mistakes matter most, and choose a sensible loss for a simple regression situation.
One set of errors can be summarized in different ways
Suppose two predictions miss their targets by 2 and 6 units.
With absolute error, the two contributions are simply:
|2| = 2|6| = 6
With squared error, they become:
2² = 46² = 36
Both methods agree that 6 is the larger mistake. They disagree about how much more strongly that large mistake should influence the summary.
That is why choosing a loss function is not a neutral formatting choice. It encodes how training should value different errors.
MAE and MSE answer slightly different questions
Mean absolute error (MAE) averages the sizes of the misses:
MAE = average(|target - prediction|)
Mean squared error (MSE) averages squared misses:
MSE = average((target - prediction)²)
Because MSE squares the error, one large miss can affect it much more strongly than several small misses.
MAE also has an intuitive unit: if the target is measured in minutes, MAE is measured in minutes. MSE is measured in squared minutes, so its raw unit is less intuitive even though it is useful for optimization.
Watch an outlier change both losses
The Lab compares MAE and MSE on the same four targets, then introduces one large prediction error.
- Click Run.
- Read the two output lines.
clean MAE/MSE: 1.0 1.0shows both losses when every prediction is off by 1.outlier MAE/MSE: 4.25 49.75shows them after the last prediction becomes30.0while its target is16.0, a miss of 14. - Compare the jumps: MAE grew from
1.0to4.25, but MSE grew from1.0to49.75. Squaring the miss (14 × 14 = 196) is what makes MSE jump so much. - Find
pred_outlier = np.array([9.0, 13.0, 13.0, 30.0]). Change only30.0to44.0, so the miss becomes 28 (twice as large). - Before running, predict: if the miss doubles, should MAE roughly double its extra error, and should MSE grow even faster?
- Click Run. You should see
outlier MAE/MSE: 7.75 196.75. The absolute error of the outlier doubled, but its squared error became four times larger. Press Reset afterward.
Loading lab…
The important observation is not that MSE is “better.” It is that MSE makes a deliberate tradeoff: large misses receive much more weight.
Which loss should you choose?
Suppose you predict delivery time.
If a 60-minute miss is dramatically more damaging than six 10-minute misses, a squared-error objective may match that concern better.
Now suppose your data occasionally contains extreme sensor glitches. If you do not want those few extreme values to dominate training, an absolute-error style objective may be more robust.
There is no universally best loss. The right question is:
How should different sizes or types of mistakes matter for this task?
The training loss and the reporting metric also do not have to be identical. You might train with one objective and report several metrics that communicate behavior more clearly.
A common misconception
“A lower numerical loss means the model is better in every sense.”
Loss values are only comparable when the loss definition, data, and evaluation conditions are aligned. A lower MSE does not automatically imply better fairness, better rare-case behavior, or even lower MAE.
Quick Check
Key Takeaways
- A loss turns many prediction errors into a training signal.
- MAE grows linearly with error size.
- MSE squares errors and emphasizes large misses.
- Loss choice encodes a tradeoff about which mistakes matter most.
Next Lesson
Next, you will see how gradient descent uses information about the loss to decide how model parameters should change.
References
- scikit-learn, sklearn.metrics.
Completion is stored locally on this device.