Context Windows
Goal
Translate a token-position limit into actual text capacity, compare truncation with sliding windows, and explain why tokenizer efficiency changes how much source text reaches a model.
A context window counts token positions, not characters or words. If two tokenizers encode the same paragraph as 80 and 120 tokens, a 100-token model can receive the first representation in one window but not the second.
Think of the window as a fixed number of seats. Tokenizer choices determine how quickly those seats fill.
A context window is an information boundary
If a model supports a context length of 8 tokens, a single forward pass cannot directly use token position 9 as part of the same 8-token input.
That sounds obvious, but tokenization makes the boundary less intuitive. A “short” human sentence can use many tokens if it contains rare words, code, URLs, or text from a domain the tokenizer represents inefficiently.
So context limits are measured in tokens, not characters, words, or pages.
Tokenization and context length interact
Suppose two tokenizers encode the same paragraph:
Tokenizer A → 120 tokens
Tokenizer B → 190 tokens
With a 150-token context limit, A fits and B does not.
This is one reason tokenizer efficiency affects more than storage: it changes how much human-visible information fits inside the model's context budget.
Truncation creates a task-dependent risk
Imagine a support conversation where the latest message says:
Ignore my earlier request; I solved that problem. Help me with billing instead.
If a truncation policy keeps the wrong end of the conversation, the model may never receive that correction.
Long-context handling is therefore not merely “pick the maximum window.” You need a policy for what to keep when text exceeds the budget.
Later RAG and agent lessons will revisit this as a system-design problem: retrieve, summarize, compress, or prioritize information rather than assuming every relevant token can always stay in context.
Predict
Make coverage visible with indices
The Lab uses the ten token IDs 0..9, a window size of 4, and a stride of 2. The stride is how far each new window starts after the previous one.
- Before running, write the windows by hand: they start at positions
0,2,4, and6. - Click Run. Compare with
windows: [(0, [0, 1, 2, 3]), (2, [2, 3, 4, 5]), (4, [4, 5, 6, 7]), (6, [6, 7, 8, 9])]. Neighboring windows overlap by two positions, andcovered positionsincludes all ten. - Find
sliding_windows(ids, window_size=4, stride=2)and changestride=2tostride=5. Before running, predict which positions will never appear in any window. - Click Run. The windows start at
0and5, andcovered positionsis missing4and9. A stride larger than the window creates gaps. - Try
stride=4as well. Now there are no gaps and no overlap, but the last window is only[8, 9]. Overlap costs extra computation; gaps lose information. - Press Reset afterward.
Loading lab…
Do not confuse “the model supports a longer context” with “the model uses all distant information well.” At this level, the first question is simpler: which encoded content reaches the model at all?
Quick Check
Transfer the idea
Design a chunking policy for a long document where evidence near boundaries matters. State both the benefit and cost of overlap.
Key Takeaways
- Context capacity is measured in token positions after tokenization.
- Tokenizer efficiency directly affects how much source text fits.
- Truncation loses content; windows trade repeated work for broader coverage.
- Track source indices to reveal gaps and boundary mistakes.
Next Lesson
Next, evaluate tokenizer behavior across realistic text slices instead of trusting one average or one familiar sentence.
References
- Kudo and Richardson, SentencePiece.
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.