Recurrent Models and Memory
Goal
By the end of this lesson, you can trace a recurrent hidden state across time and explain how current input and previous state combine into a compressed memory.
A recurrent model carries a running note
Imagine reading a sentence one word at a time while keeping a small note about what has happened so far.
At each step you:
- read the new input;
- combine it with the previous note;
- write an updated note for the next step.
In a recurrent neural network, that carried note is the hidden state.
It is not a perfect recording of the entire sequence. It is a compressed representation shaped by the update rule and training objective.
Trace a tiny recurrence
The Lab uses:
h_t = tanh(0.7*h_(t-1) + x_t)
with h_0 = 0.
Suppose the input sequence is [1, 0, 0].
At the first step, the positive input makes h_1 positive.
At the second step, x_2 is zero, but 0.7*h_1 still carries some positive influence forward.
At the third step, that influence can remain again, though transformed by another recurrence and tanh.
This is memory in the simplest visible form: the current state depends partly on previous state.
What happens when recurrent weight becomes zero?
If the recurrence becomes:
h_t = tanh(0*h_(t-1) + x_t)
then:
h_t = tanh(x_t)
Previous hidden state no longer contributes at all. In this toy model, every step depends only on the current input.
That makes the role of the recurrent connection concrete.
Complete the recurrent update, then inspect memory strength
This Lab now asks you to write the one line that creates recurrent memory.
- Click Run unchanged. The TODO currently sets
h = 0.0, so the carried state is erased at every step and the exercise checks do not all pass. - Find the TODO inside
recurrent_trace. - Replace the placeholder with the recurrence introduced above:
h = tanh(recurrent_weight * previous_h + current_x)
Use the variables already present in the function: h, recurrent_weight, x, and math.tanh.
4. Run again. Confirm that the zero input after the first positive input still has a non-zero hidden state.
5. Compare the completed default trace with no_memory_states, where recurrent weight is 0.0.
6. Once the checks pass, change the default recurrent weight from 0.7 to 0.95 and predict whether early influence persists more strongly.
Loading lab…
If the final hidden state surprises you, do not debug only the last scalar. Find the first time step where the trace differs from your expectation.
Reusing one transition is powerful and limiting
An RNN uses the same transition parameters at every time step. This lets one learned update rule process sequences of different lengths.
But the repeated chain also means information and gradients must pass through many transformations to connect distant steps.
That leads directly to the long-context problem in the next lesson.
A common misconception
“The hidden state stores the whole past.”
A finite-dimensional hidden state is a learned summary. It can forget, compress, distort, or emphasize parts of history depending on the model and training.
LSTMs and GRUs add gates to control what is kept and forgotten, but they still do not create unlimited perfect memory.
Quick Check
Key Takeaways
- Recurrent models carry a hidden representation through sequence steps.
- Each update mixes current input with previous state.
- Hidden state is compressed memory, not a perfect recording.
- Shared transition parameters create a long dependency chain.
- Debug recurrent computation step by step from the first mismatch.
Next Lesson
You have reached the Level 3 mini checkpoint. After it, you will measure how the influence of early context changes as the dependency chain grows longer.
References
- PyTorch, nn.RNN.
Completion is stored locally on this device.