본문으로 건너뛰기
L2.3

Activation Functions

Goal

By the end of this lesson, you can compare ReLU, sigmoid, and tanh on the same inputs and explain why neural networks need nonlinear activation functions between linear layers.

An activation transforms the weighted sum​

A neuron first produces a pre-activation value z from its weighted sum.

An activation function decides what signal is passed onward.

Take the same inputs:

-5, -1, 0, 1, 5

Different activations transform them differently.

ReLU​

ReLU(z) = max(0, z)

Negative values become 0. Positive values stay unchanged.

Sigmoid​

Sigmoid squeezes any real number into the interval from 0 to 1.

Large negative inputs approach 0; large positive inputs approach 1.

tanh​

tanh squeezes values into the interval from -1 to 1 and is centered around zero.

Why the bend matters​

If a layer computes only a linear transformation and the next layer also computes only a linear transformation, the two can be combined into one linear transformation.

A nonlinear activation creates a bend that prevents this collapse.

This is the connection back to XOR: hidden layers become meaningfully more expressive when the transformations between them are nonlinear.

Compare the three functions in the Lab​

The Lab sends the same five values through ReLU, sigmoid, and tanh.

  1. Run the starter once and observe that the ReLU check is incomplete.
  2. Complete the relu TODO so negative values become 0 while positive values keep their magnitude.
  3. Run again and confirm the ReLU checks pass.
  4. Compare the three output lines for the input [-5.0, -1.0, 0.0, 1.0, 5.0]:
    • relu: [0.0, 0.0, 0.0, 1.0, 5.0] — never negative, and no upper limit;
    • sigmoid: [0.0067, 0.2689, 0.5, 0.7311, 0.9933] — always between 0 and 1;
    • tanh: [-0.9999, -0.7616, 0.0, 0.7616, 0.9999] — between -1 and 1, and keeps the sign.
  5. Find values = np.array([-5.0, -1.0, 0.0, 1.0, 5.0]) and add -10.0 at the start and 10.0 at the end: values = np.array([-10.0, -5.0, -1.0, 0.0, 1.0, 5.0, 10.0]). Before running, predict the sigmoid and tanh outputs for 10.0.
  6. Click Run. ReLU simply returns 10.0, but sigmoid shows 1.0 and tanh shows 1.0 (rounded to four decimals). Going from 5 to 10 barely changed them—this flattening is the saturation described below.
  7. Press Reset afterward. (Reset also removes your ReLU code, so copy it first if you want to keep it.)

Loading lab…

For very large-magnitude inputs, sigmoid and tanh enter regions where their outputs change only a little. This is called saturation and later matters for gradients.

Debug the value before the activation first​

If an activation output looks strange, do not immediately change the activation function.

Check the pre-activation z first.

If z is already wrong because the weighted sum or shape is wrong, the activation may be doing exactly what it should with a bad input.

This gives a general debugging rule for networks: inspect the earliest incorrect intermediate value.

No activation is universally best​

ReLU, sigmoid, tanh, and many other activations have different shapes and gradient behavior.

The right choice depends on the architecture, task, initialization, and optimization setup.

For now, the important skill is to understand what numerical transformation the activation performs and what kind of gradient signal its shape can allow.

Why stacked linear layers need a nonlinearity​

Imagine two layers that both perform only linear transformations:

h = W1 x + b1
y = W2 h + b2

Substitute the first equation into the second:

y = W2(W1 x + b1) + b2

The result can still be rewritten as one larger linear transformation of x.

So adding more purely linear layers does not automatically give the network a richer kind of boundary.

An activation function changes that. ReLU, for example, transforms each value with:

ReLU(z) = max(0, z)

Now the network can behave differently in different regions of its input space.

Compare values around the boundary​

For inputs [-2, 0, 3], ReLU produces:

[0, 0, 3]

The positive value keeps its magnitude; the negative value is clipped to zero.

That simple bend is enough to break the “all layers collapse into one linear map” problem.

A common misconception is that activation functions are added only to force outputs into a convenient range. Some activations do control range, but their deeper role is to introduce nonlinear behavior so stacked layers can represent more complex functions.

Quick Check

1. What does ReLU return for a negative input?
2. Why not use only identity activations?
3. Which value should you check first to tell whether the activation or its input is wrong?

0 of 3 questions answered.

Key Takeaways

  • Activations transform pre-activation weighted sums.
  • ReLU, sigmoid, and tanh have different ranges and shapes.
  • Nonlinearity prevents stacked linear layers from collapsing into one linear map.
  • Saturated regions can affect gradient flow.
  • Debug the pre-activation before blaming the activation function.

Next Lesson

Next, you will scale from one neuron to a full dense layer and make every array shape explicit.

References

Lesson actions

Completion is stored locally on this device.

View progress