Loss for Neural Networks
Goal
By the end of this lesson, you can connect a network prediction to a scalar loss, compare numeric and binary-classification losses, and recognize dangerous logits/probability or numerical mismatches.
Training needs one number to improve
A neural network can produce many predictions in one batch. The optimizer still needs a measurable objective that says whether the current parameters are doing better or worse.
A loss function turns prediction error into a scalar number.
That scalar is where backpropagation starts.
The loss does not say everything about the real system. It answers the particular question encoded by the training objective.
Numeric targets: squared error
Suppose the target is 5 and the prediction is 3.
The error is -2. Squared error is:
(3 - 5)² = 4
If the prediction moves to 4, squared error becomes:
(4 - 5)² = 1
The lower loss reflects that the prediction moved closer under this objective.
Binary targets: reward probability on the true class
Suppose the true binary target is 1.
Compare two predicted positive-class probabilities:
- Model A:
0.9 - Model B:
0.55
Binary cross-entropy gives lower loss to 0.9 because the model assigned more probability to the true outcome.
If the true target were 0, the preference would reverse: lower positive-class probability would be better.
This is the core idea before the formula: assign high probability to what actually happened.
Logits and probabilities are not interchangeable
Many library losses expect logits, raw unbounded model scores, and internally apply a stable sigmoid or softmax calculation.
Other hand-written formulas expect probabilities already limited to a valid range.
Passing probabilities to a “with logits” loss, or passing raw logits to a formula that takes log(probability), can produce confusing results.
Always know what representation a loss expects.
Numerical safety near 0 and 1
A raw cross-entropy calculation may contain log(p) or log(1-p).
If p = 0 exactly where log(p) is needed, log(0) is undefined.
Production implementations use numerically stable formulations rather than trusting a naive raw logarithm at the boundaries.
Compare losses in the Lab
The Lab compares good and bad probability vectors and also computes a small MSE example.
- Click Run.
- Look at the first example.
targetsstarts with1.0, so the correct answer is “positive.”good_probsgives it0.90andbad_probsgives it0.55. - Compare the losses:
good BCE: 0.1643andbad BCE: 0.7571. Probabilities that lean toward the correct endpoint receive a much lower loss. - Find
good_probs = np.array([0.90, 0.10, 0.80, 0.20]). Change only the first value,0.90, to0.99, moving it closer to its true endpoint1. - Before running, predict whether
good BCEshould rise or fall. - Click Run.
good BCEfalls to0.1404.bad BCEandMSEstay the same because you did not change their inputs. - Press Reset afterward.
Loading lab…
Loss is an objective, not a complete quality certificate
A network can reach very low training loss and still fail on held-out data. It can also optimize a loss that does not capture an important real-world cost.
Later debugging therefore reads training loss together with validation behavior and task-relevant metrics.
Loss turns many wrong predictions into one training signal
For a batch, the model usually makes several prediction errors at once.
Suppose three examples have individual losses:
0.2, 0.7, 0.1
A training objective may average them:
mean loss = (0.2 + 0.7 + 0.1) / 3
= 0.333...
Gradient computation then asks how changing each parameter would change this aggregate objective.
That does not mean every example is equally easy or equally important. The average can hide one very bad case.
This is why training loss and diagnostic evaluation serve different roles: loss supplies an optimization target, while slice-level metrics and examples help you understand where the model is failing.
A lower loss is useful evidence under the same evaluation setup, but it is not a guarantee that every behavior you care about improved.
Quick Check
Key Takeaways
- Loss gives training a scalar objective.
- MSE is useful for numeric error; cross-entropy rewards probability assigned to the true class.
- Know whether a loss expects logits or probabilities.
- Numerically stable implementations matter near probability boundaries.
- Low training loss is not proof of held-out or real-world success.
Next Lesson
Next, you will ask how a change in an earlier parameter can change this final loss—the central question behind backpropagation.
References
- PyTorch, CrossEntropyLoss.
Completion is stored locally on this device.