What Machine Learning Is Really Optimizing
Goal
By the end of this lesson, you can explain that training searches for model settings that make a chosen loss smaller, and you can distinguish that training objective from the larger real-world goal.
Training needs a scoreboard
Imagine that you are trying to improve three predictions. The real targets are:
2, 4, 6
Two candidate rules produce:
- Rule A:
3, 5, 7 - Rule B:
7, 9, 11
Rule A misses each target by 1. Rule B misses each target by 5. Even before using machine-learning vocabulary, you can see that Rule A is closer.
A training algorithm needs a way to turn that idea of “closer” into a number it can measure. That number is usually called a loss.
For this lesson, we use mean squared error (MSE). It squares each miss and then averages the squared values.
For Rule A, each squared error is 1² = 1.
For Rule B, each squared error is 5² = 25.
So Rule A receives the smaller loss.
That gives training a practical target:
change the adjustable model settings so the chosen loss becomes smaller.
Those adjustable settings are called parameters. Later, a model may have millions of them.
The training procedure that decides how to change parameters is called an optimizer. The repeated process of changing parameters to improve an objective is called optimization.
For now, keep this chain in mind:
parameter change
→ prediction changes
→ error changes
→ loss changes
→ optimizer uses that evidence for the next change
The objective is not the whole real-world goal
There is an important catch.
A model cannot directly optimize a vague instruction such as “be useful,” “be fair,” or “help every student.” Training needs something measurable.
Suppose a school prediction system minimizes average error across all students. The average can improve even while one small group of students is repeatedly predicted badly.
The training objective improved, but the product may still be failing an important requirement.
This is why we separate two ideas:
- training objective: the measurable quantity the optimizer directly tries to improve;
- real-world goal: the broader outcome people actually care about.
A useful objective should support the real-world goal, but the two are rarely identical.
Try the tiny loss experiment
The Lab below computes MSE for three candidate prediction rules.
- Click Run without editing anything.
- Find
MSE by rule:and note that Rule A starts with an MSE of1.0. - In
candidates, find Rule A:"A": np.array([2.0, 4.0, 6.0, 8.0]). - Change only its first prediction from
2.0to3.0. The matching first target is3.0, so this removes one squared error while leaving the other three predictions unchanged. - Before running, predict whether Rule A's MSE should become larger or smaller.
- Click Run. Look again at
MSE by rule:; Rule A should move from1.0to0.75and remain the lowest-loss rule. - Restore
3.0to2.0before moving on.
Loading lab…
After this guided pass, try one additional prediction change of your own. Keep every other prediction fixed and explain the direction of the resulting MSE change before you run it.
The important result is not merely which rule wins. It is the chain of reasoning:
prediction changes → error changes → loss changes → the objective gives the optimizer a preference.
A common misconception
“If training loss is low, the system must be good.”
Low loss is useful evidence about the objective you chose. It is not proof that the system is fair, safe, robust to new conditions, gives reliable confidence estimates, or is useful for every group.
Later lessons add held-out evaluation, multiple metrics, baselines, and failure analysis precisely because one training number cannot answer every question.
Transfer the idea
Suppose you are predicting bus arrival times. A loss based on average arrival-time error may be sensible for training.
But a transit team might also care about rare delays, reliability on less frequent routes, or whether errors are worse during accessibility-critical trips. Those are additional evaluation questions, not automatically captured by one average loss.
Quick Check
Key Takeaways
- Training needs a measurable objective.
- Loss summarizes prediction error according to a chosen rule.
- Parameters are adjustable model settings.
- An optimizer changes parameters; optimization is the repeated search for settings that improve the objective.
- A better objective value is evidence about that objective, not proof of complete real-world success.
Next Lesson
Next, in L1.2 — Datasets as Tables, you will examine the object that feeds every classical ML experiment: a dataset organized as examples and variables.
References
- scikit-learn, Getting Started.
Completion is stored locally on this device.