Context Windows and Lost Information
Goal
Distinguish context-window capacity from reliable context use, explain how important information can be present but poorly used, and design tests that vary information position while holding the content fixed.
A large context window tells you how many tokens a model can accept. It does not guarantee that every token inside that window will influence the answer equally well. This distinction is easy to miss.
Present is not the same as used
Imagine a 20-page context containing one sentence that answers the question:
Project Cedar's launch month is May.
You can place that sentence:
- near the beginning;
- near the middle;
- near the end.
The total information is the same. Only its position changes. If answer quality changes substantially across those placements, the model is showing a context-utilization problem, not a context-capacity problem.
The text fits. The model simply does not use it equally reliably in every position.
Long context adds competition
Every extra document, example, and instruction consumes attention and context budget. Relevant evidence can compete with:
- repeated boilerplate;
- unrelated examples;
- old conversation history;
- duplicated documents;
- misleading but similar text.
More context can therefore make a task worse even when nothing is truncated. This is a key engineering lesson: adding information is not the same as adding useful signal.
Test position as one controlled variable
Suppose you want to know whether the model loses a fact in the middle. Use the same fact and same distractors in three versions:
A: launch fact near start
B: launch fact near middle
C: launch fact near end
Keep model, prompt wording, decoding, and total context as similar as possible. Then compare the answers. That experiment is much stronger than trying three completely different documents and concluding that “long context is unreliable.”
Context history can contain stale information
In a chat system, old turns may remain inside the context. Suppose an early message says:
Ship to Helsinki.
Later the user says:
Change the destination to Turku.
A workflow that keeps everything must make the newer instruction clearly relevant. Longer context means more opportunities for old, contradictory, or superseded information to remain present. The application may need rules for recency, versioning, or explicit state instead of asking the model to infer which old statement should be ignored.
Capacity, retrieval, and reasoning are separate questions
When a model fails on long input, ask:
- Did the relevant information fit inside the context window?
- Was it actually included in the request?
- Where was it placed?
- Was it surrounded by distractors?
- Did the prompt clearly connect the question to the evidence?
- Would retrieval or summarization create a cleaner evidence set?
Do not jump directly from “wrong answer” to “need a larger context model.”
Predict
Run the Lab
The Lab creates a synthetic context with one answer-bearing record.
The Lab places the same answer sentence, Project Cedar launch month is May., at the start, middle, and end of a small context, and uses a simple exact-match filter to find launch-month candidates.
- Click Run. The three lines show target index
0,2, and4. Only the position changed; the target text and the distractors stayed the same. That is what makes this a clean position experiment. - Notice that the simple filter finds the target at every position. A program that checks every line does not “lose” the middle. An LLM reading a long context can, which is why position is worth testing.
- Add a competing line. At the end of the
distractorslist, add"Draft note: Project Cedar launch month might move to June.",. - Before running, predict how many candidates each placement will return.
- Click Run. Every placement now returns two candidates, and in the
endplacement the draft note comes first. Choosing between them now needs more than keyword matching—you need to know which source is authoritative. - Press Reset afterward.
Loading lab…
Quick Check
Explain it back
Describe the difference between context capacity and context utilization. Then design a three-condition experiment that tests whether one fact is used differently at the beginning, middle, and end.
Key Takeaways
- Fitting inside the context window does not guarantee reliable use.
- More context can introduce competition, stale history, and distractors.
- Position sensitivity should be tested with controlled evidence placement.
- Old context can conflict with newer state.
- A larger context window is not the only response to long-input failures.
Next Lesson
Next, compare strategies for reducing long context to the evidence a task actually needs.
References
Completion is stored locally on this device.