Neural Network Debugging Workshop
Goal
By the end of this lesson, you can diagnose neural-network failures by moving through observable boundaries—data, shapes, forward values, loss, gradients, updates, and curves—and propose one evidence-driven repair experiment.
Debugging is easier when you inspect boundaries in order
A neural network is a numerical pipeline.
When the final result is wrong, jumping directly to “change the optimizer” or “add another layer” skips many places where the failure may actually begin.
Use an evidence ladder:
- Data and targets — Are the examples and labels what you think they are?
- Shapes — Do dimensions represent batch, features, and neurons correctly?
- Forward pass — Are weighted sums and activations finite and sensible?
- Loss — Is the objective finite and using the intended representation?
- Gradients — Do values exist, have plausible signs, and avoid NaN/Inf?
- Updates — Do parameters actually move by sensible amounts?
- Curves — Does learning improve, stall, diverge, or overfit?
This order moves from basic correctness toward higher-level modeling choices.
Follow the evidence, not a favorite fix
Consider three examples.
Case A: shape mismatch before the first layer
Do not tune the learning rate. The computation cannot even represent the intended operation yet. Name the dimensions and fix the data/layer contract.
Case B: finite loss and nonzero gradients, but parameters never change
The backward path exists. Inspect the optimizer/update boundary: Are parameters registered? Is step() called? Are updates overwritten or immediately restored?
Case C: training loss falls while validation loss rises
The forward and optimization paths are at least functioning. Now investigate generalization: capacity, data coverage, regularization, leakage, and split quality.
The same symptom word “model failed” leads to very different next checks.
Diagnose the four cases in the Lab
The Lab contains four named evidence records: shape, dead_relu, no_update, and overfit.
- Click Run and read all four
name -> diagnosislines. - Focus on the
no_updatecase. It starts withgradients_nonzero: Trueandparameters_changed: False, so the diagnosis should beinspect_optimizer_update. - In
cases["no_update"], change only"parameters_changed": Falseto"parameters_changed": True. - Before running, reason through the evidence ladder: shapes are valid, loss is finite, gradients exist, parameters now move, and validation is not rising. Predict that the diagnosis should move past the optimizer boundary.
- Click Run. Confirm
no_updatenow reportscontinue_training_or_collect_more_evidencewhile the other three cases keep their original diagnoses. - Restore
"parameters_changed": False.
Loading lab…
After this guided pass, choose a different case and change one evidence flag. Predict the first boundary that should become decisive before running again.
The diagnostic function is a practice aid. In a real system, your evidence—not the label returned by a helper—is what justifies the next action.
Keep a short experiment log
For every debugging attempt, record:
| Field | Example |
|---|---|
| Symptom | validation loss rises after epoch 8 |
| Evidence | train loss continues falling |
| Hypothesis | model capacity is fitting training noise |
| One change | reduce hidden width |
| Fixed | data split, seed, optimizer, metric |
| Result | validation curve improves/worsens/unchanged |
| Interpretation | hypothesis supported or not supported |
Revert failed hypotheses instead of stacking unexplained changes on top of one another.
A common mistake: applying several repairs at once
Suppose you change learning rate, batch size, initialization, and architecture together and the model improves.
You have an improved run but weak understanding. You do not know which change mattered or whether one change merely compensated for another mistake.
Controlled debugging produces slower-looking individual experiments but faster understanding.
The Level 2 mental model
At this point, you should be able to explain a neural network without treating it as a black box:
inputs -> weighted sums -> activations -> later layers -> prediction -> loss -> local derivatives -> gradients -> optimizer updates -> new parameters
Training curves then show what repeated updates do over time.
If you can trace that chain and identify where evidence first goes wrong, you have the main first-principles skill this level is designed to teach.
Quick Check
Key Takeaways
- Debug neural networks by observable boundaries from data to evaluation.
- Use the earliest wrong evidence to choose the next check.
- Keep symptom, evidence, hypothesis, controlled change, and result together.
- Reproducible failures are easier to fix than unexplained intermittent ones.
- Level 2 mastery means you can both build a small network and explain why it succeeds or fails.
Next Lesson
You are ready for the Level 2 project, Neural Network From Scratch. After that, Level 3 moves from tiny first-principles networks to learned representations across images, sequences, and pretrained models.
References
- PyTorch, Autograd.
- PyTorch, torch.optim.
- PyTorch, Reproducibility.
Completion is stored locally on this device.
Level project unlocked: Neural Network From Scratch