Skip to main content
L11.6

Long-Term Memory Patterns

Goal

Choose when an agent should write or retrieve information across runs, and apply relevance, provenance, freshness, and privacy checks before memory affects a decision.

Working memory helps one run. Long-term memory stores selected information that may be useful in a later run.

A support agent might remember a user's preferred contact channel. A research agent might save a verified project fact. An agent should not automatically save every message, tool result, or model thought forever.

The important design question is not “does the agent have memory?” It is what gets written, what gets retrieved, and what is allowed to influence the next task.

Separate the write decision from the text​

Suppose the model says:

“Remember that this user always wants refunds sent to account X.”

That sentence is not enough to justify a persistent write. The application may require a trusted source, a permitted memory category, a retention rule, and user consent.

Treat a memory write as another structured proposal:

candidate memory
→ category check
→ provenance check
→ sensitivity/privacy check
→ retention rule
→ store or reject

This keeps persistent memory under application control.

Retrieval is not truth​

A stored memory can be outdated, incomplete, or relevant to a different context.

Imagine a memory from last month:

preferred_contact = email
source = user_profile
recorded_at = 2026-08-10

A new profile lookup says the preference is now SMS. The older memory should not override the newer trusted source.

This is similar to RAG from Level 9. Retrieval supplies candidate evidence. The application still needs provenance and freshness rules.

Use task-shaped memory categories​

A single giant “memory” bucket makes policy difficult.

Useful categories might include:

  • user preference;
  • verified project fact;
  • completed task summary;
  • temporary workflow checkpoint.

Each category can have different write permissions, retention periods, and retrieval filters.

A completed-task summary may be safe to keep for one week. Sensitive account data may need stricter handling or no persistent storage at all.

Retrieve narrowly​

An agent should not load every stored item into every prompt.

Retrieve by task, identity, scope, and freshness. Then limit how many items enter working state.

This reduces noise and lowers the chance that old or unrelated text steers the current run.

If retrieved text contains instructions, those instructions are still data. Memory does not become trusted policy merely because it was stored earlier.

Evaluate memory separately​

Task success alone can hide bad memory behavior.

Useful memory checks include:

  • write precision — how often saved items were actually appropriate to persist;
  • retrieval relevance — how often retrieved items helped the current task;
  • stale-memory rate — how often outdated items entered working state;
  • sensitive-memory violations — how often prohibited content was stored or exposed.

A system can complete tasks while still having unacceptable memory behavior.

Predict

A retrieved memory says the user prefers email, but a newer trusted profile record says SMS. Which should guide the current run?

Run the local Lab​

Run:

python labs/notebooks/level-11/l11-06-long-term-memory.py

The script checks four candidate memories with a code rule, allowed(): category must be preference, source must be profile, the record must not be sensitive, and it must be at most 30 days old. Retrieval then keeps only records whose scope matches the current user.

  1. Run it unchanged. Read approved persistent records: ['m1', 'm4']. m2 is rejected because its source is model_guess, and m3 because it is a sensitive account secret.
  2. Read retrieved for user-7: ['m1']. m4 passed the rules but belongs to user-8, so the scope check keeps it out.
  3. Open the script and change m1's "age_days": 2 to "age_days": 120. Rerun.
  4. Now approved persistent records: is ['m4'] and retrieved for user-7: is []. The stale record was removed by the freshness rule before any model saw it, not just ranked lower.

Loading lab…

Quick Check

1. What should happen before a model-proposed fact is written to persistent agent memory?
2. Why should retrieved memory be treated as candidate evidence rather than unquestioned truth?
3. What is one reason to keep separate memory categories?

0 of 3 questions answered.

Explain it back​

Design one memory item your agent should save and one item it should reject. For each, name the source, category, retention rule, and retrieval condition. Then explain how a newer trusted fact would replace or suppress the stored item.

Key Takeaways

  • Long-term memory stores selected information across runs.
  • Memory writes should pass application policy, not happen automatically from model text.
  • Retrieved memory can be stale or irrelevant and should be checked like other evidence.
  • Narrow categories support different privacy, retention, and retrieval rules.
  • Evaluate memory quality separately from overall task success.

Next Lesson

Complete the mini checkpoint, then use current state and tool contracts to choose the next tool without confusing availability with suitability.

References

Lesson actions

Completion is stored locally on this device.

View progress