Activation Functions
Goal
By the end of this lesson, you can compare ReLU, sigmoid, and tanh on the same inputs and explain why neural networks need nonlinear activation functions between linear layers.
An activation transforms the weighted sum
A neuron first produces a pre-activation value z from its weighted sum.
An activation function decides what signal is passed onward.
Take the same inputs:
-5, -1, 0, 1, 5
Different activations transform them differently.
ReLU
ReLU(z) = max(0, z)
Negative values become 0. Positive values stay unchanged.
Sigmoid
Sigmoid squeezes any real number into the interval from 0 to 1.
Large negative inputs approach 0; large positive inputs approach 1.
tanh
tanh squeezes values into the interval from -1 to 1 and is centered around zero.
Why the bend matters
If a layer computes only a linear transformation and the next layer also computes only a linear transformation, the two can be combined into one linear transformation.
A nonlinear activation creates a bend that prevents this collapse.
This is the connection back to XOR: hidden layers become meaningfully more expressive when the transformations between them are nonlinear.
Compare the three functions in the Lab
The Lab sends the same five values through ReLU, sigmoid, and tanh.
- Run the starter once and observe that the ReLU check is incomplete.
- Complete the
reluTODO so negative values become 0 while positive values keep their magnitude. - Run again and confirm the ReLU checks pass.
- Compare the three output lines for the input
[-5.0, -1.0, 0.0, 1.0, 5.0]:relu: [0.0, 0.0, 0.0, 1.0, 5.0]— never negative, and no upper limit;sigmoid: [0.0067, 0.2689, 0.5, 0.7311, 0.9933]— always between 0 and 1;tanh: [-0.9999, -0.7616, 0.0, 0.7616, 0.9999]— between -1 and 1, and keeps the sign.
- Find
values = np.array([-5.0, -1.0, 0.0, 1.0, 5.0])and add-10.0at the start and10.0at the end:values = np.array([-10.0, -5.0, -1.0, 0.0, 1.0, 5.0, 10.0]). Before running, predict the sigmoid and tanh outputs for10.0. - Click Run. ReLU simply returns
10.0, but sigmoid shows1.0and tanh shows1.0(rounded to four decimals). Going from5to10barely changed them—this flattening is the saturation described below. - Press Reset afterward. (Reset also removes your ReLU code, so copy it first if you want to keep it.)
Loading lab…
For very large-magnitude inputs, sigmoid and tanh enter regions where their outputs change only a little. This is called saturation and later matters for gradients.
Debug the value before the activation first
If an activation output looks strange, do not immediately change the activation function.
Check the pre-activation z first.
If z is already wrong because the weighted sum or shape is wrong, the activation may be doing exactly what it should with a bad input.
This gives a general debugging rule for networks: inspect the earliest incorrect intermediate value.
No activation is universally best
ReLU, sigmoid, tanh, and many other activations have different shapes and gradient behavior.
The right choice depends on the architecture, task, initialization, and optimization setup.
For now, the important skill is to understand what numerical transformation the activation performs and what kind of gradient signal its shape can allow.
Why stacked linear layers need a nonlinearity
Imagine two layers that both perform only linear transformations:
h = W1 x + b1
y = W2 h + b2
Substitute the first equation into the second:
y = W2(W1 x + b1) + b2
The result can still be rewritten as one larger linear transformation of x.
So adding more purely linear layers does not automatically give the network a richer kind of boundary.
An activation function changes that. ReLU, for example, transforms each value with:
ReLU(z) = max(0, z)
Now the network can behave differently in different regions of its input space.
Compare values around the boundary
For inputs [-2, 0, 3], ReLU produces:
[0, 0, 3]
The positive value keeps its magnitude; the negative value is clipped to zero.
That simple bend is enough to break the “all layers collapse into one linear map” problem.
A common misconception is that activation functions are added only to force outputs into a convenient range. Some activations do control range, but their deeper role is to introduce nonlinear behavior so stacked layers can represent more complex functions.
Quick Check
Key Takeaways
- Activations transform pre-activation weighted sums.
- ReLU, sigmoid, and tanh have different ranges and shapes.
- Nonlinearity prevents stacked linear layers from collapsing into one linear map.
- Saturated regions can affect gradient flow.
- Debug the pre-activation before blaming the activation function.
Next Lesson
Next, you will scale from one neuron to a full dense layer and make every array shape explicit.
References
- PyTorch, torch.nn.
Completion is stored locally on this device.