Why Long Context Is Hard
Goal
By the end of this lesson, you can calculate how repeated multipliers change distant influence and explain vanishing/exploding chains as a reason long dependencies can be hard to learn or use.
Repeated small losses of influence become large
Suppose a toy recurrent path keeps 80% of an earlier signal at each step.
After one step, influence is:
0.8
After two:
0.8² = 0.64
After five:
0.8⁵ ≈ 0.33
After ten:
0.8¹⁰ ≈ 0.11
The earlier input was processed, but its direct influence has become much smaller after many repeated transformations.
That makes an important distinction:
A model having seen an earlier item does not guarantee that item still strongly affects a distant prediction.
Multipliers above one create the opposite problem
With a repeated factor 1.5:
1.5² = 2.25
1.5⁵ ≈ 7.59
Repeated amplification can make signals or gradients grow rapidly.
These toy products illustrate two classic chain problems:
- factors below one can create vanishing influence/gradients;
- factors above one can create exploding influence/gradients.
Real recurrent networks contain nonlinearities and matrices, but the repeated-product intuition remains useful.
This affects both memory and learning
The forward hidden state may gradually lose evidence from an early input. During training, gradients also travel through a long chain of recurrent transformations. They can shrink or grow, making it difficult to learn which distant event mattered.
Gated recurrent networks such as LSTMs and GRUs improve this situation by creating more controlled paths for information.
Measure influence by distance
The Lab computes r^distance for several recurrent multipliers.
- Click Run.
- Compare influence at short and long distances for
r=0.5andr=0.8. - Change the multiplier to
1.05. - Predict after how many steps the influence first grows above its original magnitude.
- Run and inspect the trajectory.
- Reset the starter afterward.
Loading lab…
Aggregate accuracy can hide a distance problem
Suppose a sequence model performs:
- 95% when relevant context is 1–3 positions away;
- 90% at 4–8 positions;
- 55% at 20+ positions.
One overall average can hide the sharp long-distance weakness.
A useful failure slice therefore records performance by context distance.
This is a general evaluation habit: slice by the variable connected to your hypothesis.
Attention changes the information path, not the laws of evidence
Attention can create much shorter paths between distant positions than a purely recurrent chain. That is one reason it became important for long sequences.
But it does not give infinite memory or guarantee that every distant token is used well. Context length, computation, positional structure, training data, and retrieval behavior still matter.
Quick Check
Key Takeaways
- Distant influence passes through repeated transformations.
- Repeated factors below one can vanish; factors above one can explode.
- “Seen in context” is not the same as “strongly influences the output.”
- Evaluate sequence failures by context distance when distance is the suspected cause.
- Shorter information paths help but do not create unlimited context use.
Next Lesson
Next, you will turn discrete items into learned vectors and ask what geometric similarity means inside a representation.
References
- PyTorch, nn.RNN.
Completion is stored locally on this device.