Skip to main content
L7.8

Long-Context Strategies

Goal

Choose among trimming, retrieval, chunking, summarization, and explicit state based on the information a task needs, and explain the failure risks of each strategy.

When input is too long or too noisy, “use a bigger context window” is only one option. A reliable workflow asks a more useful question:

Which information does this task actually need, and how can we preserve that information with the least unnecessary context?

Strategy 1: trim by a known rule​

If only the most recent conversation turns matter, keep the recent window and discard older turns. This is simple and deterministic. But it can fail when an old fact is still important. For example, the user's shipping address may have been given 20 turns ago and never repeated.

Trimming works best when the application knows which history can safely expire.

Strategy 2: keep explicit state​

Instead of relying on the whole conversation transcript, the application can maintain state:

{
"shipping_city": "Turku",
"plan": "standard",
"open_issue": "damaged item"
}

The state can be updated when the user changes a value. This makes important facts easy to find and reduces dependence on the model remembering which historical message is current. The risk is state-update correctness: if the application extracts the wrong value, the error becomes persistent.

Strategy 3: retrieve relevant chunks​

Split documents into chunks, score them against the current question, and send only the most relevant ones. A simple keyword-overlap retriever might choose:

question: "What is the launch code?"

chunk A: meeting schedule
chunk B: launch code is 4172
chunk C: team roster

Chunk B should rank highest. Real retrieval systems use stronger methods, which Level 9 will cover. The Level 7 lesson focuses on the workflow idea: select evidence before generation. Retrieval can fail if chunking separates needed evidence or if the query does not match the right document vocabulary.

Strategy 4: summarize or compress​

A long history can be summarized into a shorter record. This saves context, but summarization is lossy. If the summary omits one exception, the later model cannot recover it from the summary alone. Important facts may need structured state or source links rather than only prose compression.

Combine strategies intentionally​

A practical workflow might use:

recent turns
+ explicit user state
+ retrieved evidence
+ short system instruction

Each piece has a reason to be present. This is better than copying every available token into one prompt.

Measure information preservation​

For a long-context strategy, define cases where the answer depends on:

  • a recent fact;
  • an old but still active fact;
  • two facts from different chunks;
  • a corrected fact;
  • a fact absent from all sources.

Then test whether the strategy preserves the evidence needed for each case. The goal is not minimum token count. The goal is enough correct evidence with manageable noise and cost.

Predict

A user's current shipping city must survive a long conversation. Which approach gives the application the clearest control?

Complete the Lab retriever​

The Lab contains several small document chunks.

  1. Click Run. The output is retrieved: None None, and the checks fail, because retrieve is not written yet.
  2. Complete the TODO: use the provided word_overlap to score every chunk against the query, and return the ID and text of the highest-scoring chunk.
  3. Click Run again. The output should be retrieved: month Project Cedar launch month is May., and every check should pass.
  4. Now ask the same question in different words: change the query to "When does Cedar go live?" and run it. Look at the scores: the only shared word is cedar, which appears in both the month and the team chunks. The right chunk can only win by luck, because keyword overlap cannot see that “go live” means “launch.”
  5. Restore the original query. Add a distractor to chunks: "faq": "What is the Cedar launch month? Ask the project office.",.
  6. Click Run. The retriever now returns the faq chunk, because it repeats the question's words—but it contains no answer. Many shared words is not the same as useful evidence.

Loading lab…

Quick Check

1. What is the main risk of summarization?
2. Why can explicit state outperform raw history for current values?
3. What does retrieval try to do before generation?

0 of 3 questions answered.

Explain it back​

Choose one long-input task and compare trim, explicit state, retrieval, and summarization. State one benefit and one failure mode for each, then choose a combination and justify it.

Key Takeaways

  • Long-context problems are often information-selection problems.
  • Trimming is simple but can discard still-active facts.
  • Explicit state makes important current values easier to control.
  • Retrieval selects evidence before generation but can miss relevant chunks.
  • Summarization saves tokens but can omit critical details.
  • Evaluate strategies by whether required information survives, not by token count alone.

Next Lesson

Next, revisit decoding at the workflow level and separate reproducibility from randomness.

References

Lesson actions

Completion is stored locally on this device.

View progress