Forward Pass by Hand
Goal
By the end of this lesson, you can perform and verify a two-layer forward pass from inputs to prediction while keeping every important intermediate value visible.
A forward pass follows the arrows
A network does not learn during a forward pass. It simply computes an output using its current parameters.
For a two-layer network, the path can be written as:
z1 = X @ W1 + b1a1 = activation(z1)z2 = a1 @ W2 + b2- output from
z2using the task's final transformation
The names are useful because they separate two different things:
z1: the weighted sums before activation;a1: the transformed values passed to the next layer.
If you confuse them, downstream calculations can look reasonable while using the wrong quantity.
Trace one ReLU example
Suppose a hidden layer produces:
z1 = [-2, 3]
After ReLU:
a1 = [0, 3]
The first hidden unit contributes zero to the next layer because its pre-activation was negative. The second unit passes its positive value onward.
If you change one first-layer bias so -2 becomes +1, then that hidden unit wakes up:
a1 changes from [0, 3] to [1, 3]
and the output layer now receives a new contribution.
This lets you predict which downstream values should change before running code.
Run the forward pass Lab
The Lab uses one two-feature example, a two-neuron hidden layer, ReLU, and one output neuron.
- Click Run and read the outputs in order:
z1: [[0.1, -1.45]],a1: [[0.1, 0.0]], andprediction: 0.17. (predictionis the output layer'sz2.) - Verify the second hidden pre-activation by hand. With
x = [2.0, -1.0], the second column ofW1is[-0.25, 0.75]and its bias is-0.2: 2.0 × (-0.25) + (-1.0) × 0.75 + (-0.2) = -1.45. ReLU turns-1.45into0.0. - Find
b1 = np.array([0.1, -0.2]). Change only the second bias,-0.2, to1.5, so that hidden pre-activation becomes positive. - Before running, predict which values must change and which stay fixed. Does the first hidden neuron change?
- Click Run. Now
z1: [[0.1, 0.25]]anda1: [[0.1, 0.25]]: the second neuron is “on” and passes0.25forward. Because its output weight inW2is-0.7, the prediction falls from0.17to-0.005. The first neuron did not change at all. - Press Reset afterward.
Loading lab…
When debugging, stop at the first intermediate value that disagrees with your hand calculation. Every later mismatch may simply be a consequence of that earlier error.
A semantic failure can survive shape checks
Imagine applying W2 directly to the original input instead of a1.
In a specially chosen toy example, shapes might still happen to fit. The code would run, but the network would have skipped the hidden computation you intended.
Shape correctness is necessary. It is not sufficient. Also verify that each operation receives the right meaningful value.
Forward computation creates the dependency path for learning
Conceptually, each result depends on earlier values:
X -> z1 -> a1 -> z2 -> loss
Backpropagation will later traverse these dependencies in reverse to ask how a small change in an earlier parameter would affect the final loss.
A forward pass is a chain of named transformations
Consider a tiny network:
input x = [2, 1]
linear layer → z
ReLU → h
final linear layer → prediction
Do not jump directly from input to prediction. Record each intermediate value.
If the first layer produces z = [-1, 3], ReLU gives:
h = [0, 3]
The final layer therefore receives [0,3], not the original [2,1] and not the pre-activation [-1,3].
That distinction matters when debugging.
Stop at the first disagreement
Suppose code returns the wrong final prediction.
If your hand-calculated first-layer output already differs from the program, inspecting the final layer is wasted effort. The earliest mismatch is closer to the cause.
This gives a reusable debugging method:
- verify the input;
- verify each linear pre-activation;
- verify the activation output;
- continue layer by layer;
- only then inspect the final prediction.
Later Transformer debugging uses the same habit with much larger tensors.
Quick Check
Key Takeaways
- A forward pass computes predictions from current parameters; it does not update them.
- Keep pre-activations and activations distinct.
- Intermediate values make the computation auditable.
- Debug from the earliest mismatch.
- The forward dependency chain is the path that backpropagation will reverse.
Next Lesson
Next, you will turn prediction quality into a scalar loss that gives training a measurable objective.
References
Completion is stored locally on this device.