Memory Retrieval and Compression
Goal
Design memory retrieval and compression rules that preserve provenance, freshness, and safety-relevant facts while keeping working context bounded.
Memory is a policy, not a transcript
Long-running tasks collect more information than a model can or should see at every step. The harness therefore needs a memory policy for choosing what to retrieve and how to compress older state. This is different from simply saving everything. Retrieval and compression are decisions about which evidence remains available for the next action.
Start with retrieval. A useful memory item should match the current subgoal, belong to the correct task or user scope, come from an allowed source, and still be fresh enough for the decision. Similarity alone is not sufficient. A highly similar note from the wrong account or an expired policy can be more dangerous than an irrelevant note.
Provenance means recording where a memory item came from. A verified tool result, a user-provided preference, and a model-generated summary should not be treated as equivalent evidence. When memory is retrieved, the controller should preserve enough provenance to decide how much trust the item deserves.
{"memory_id":"m-17","scope":"order-4172","source":"verified_tool","fresh":true,"summary":"status=shipped"}
Compress without losing authority
Compression reduces many events into a smaller representation. For example, ten order-status checks might become one summary that records the latest verified status, the last verified timestamp, and the source. Good compression keeps facts needed for future decisions while removing repetition.
Compression can also remove something important. Suppose a summary says 'refund approved' but omits that approval expires at 19:00. The shortened memory is easier to fit into context but is no longer safe for execution after the deadline. Safety-relevant fields should therefore be protected from lossy summarization or stored separately in structured state.
One helpful pattern is to separate authoritative structured state from descriptive summaries. Approval status, operation IDs, permissions, and terminal state remain structured fields. A natural-language summary can describe progress for the model, but it does not replace those controller-owned facts.
Bound retrieval and retention
Retrieval should be bounded. Returning twenty vaguely related memories can increase latency and make the next decision less focused. A policy can limit the number of items, prefer recent verified evidence, and require a minimum relevance threshold. The exact scoring method may vary, but the boundary should be explicit.
Memory also has a lifecycle. Some items should expire, some should be deleted when a task ends, and some user preferences may persist longer. Retention should match purpose rather than defaulting to permanent storage.
Remember from Level 9, especially L9.3 — Embeddings for Retrieval and L9.4 — Vector Similarity: retrieval chooses evidence by comparing a query with candidate items. Agent memory reuses that idea, but similarity is only one signal because retrieved state can influence actions. The harness should treat retrieval as a controlled input boundary, with scope, provenance, freshness, and retention checks.
Predict
Run the local Lab
Run:
python labs/notebooks/level-12/l12-06-memory-compression.py
The script chooses which memory items to put into a small context. First it filters: the scope must be u7, the source must be trusted, and the item must be at most 30 days old. Then it sorts the remaining items by relevance and keeps at most max_items.
- Run it unchanged. Read
eligible: ['a', 'c']andselected: ['a', 'c']. - Look at the items that were left out.
bhas the highest relevance in scope (0.99), but its source ismodel_summary.dhas relevance1.0, but it belongs to useru8. - Change
max_items=3tomax_items=1and rerun. - Now
selected: ['a']. With only one slot, the policy keeps the most relevant eligible item, not the highest-scoring text overall.
Compression chooses from what is already trusted and in scope. It never lets a high relevance score bring back an item the filters removed.
Loading lab…
Quick Check
Explain it back
Explain how you would store an expiring approval, a user preference, and a model-generated progress summary. State which one may influence an external side effect directly and why.
Key Takeaways
- Memory retrieval needs scope, provenance, freshness, and relevance checks.
- Compression should remove repetition without erasing safety-critical facts.
- Structured authority state should remain separate from descriptive summaries.
- Retrieval result size should be bounded.
- Retention duration should match the purpose of the data.
Next Lesson
You have reached the Level 12 checkpoint. After reviewing the first six lessons, continue to L12.7 — Parallel Tool Calls.
References
Completion is stored locally on this device.