State and Working Memory
Goal
Represent the information an agent needs for the current run as explicit working state, and keep that state small, typed, and separate from the full conversation transcript.
An agent makes later decisions from what happened earlier. That means it needs state: information that survives from one cycle to the next.
State can contain trusted facts, progress markers, counters, the active subgoal, approval status, and recent observations. The subset needed for near-term decisions is often called working memory.
The term sounds human, but here it is an engineering structure. It can be a dictionary, object, or database row.
Start with explicit fields
For the running support example, a useful state might be:
{
"goal": "resolve order 4172",
"order_status": "delivered",
"damaged": true,
"active_subgoal": "check refund eligibility",
"approval_status": "not_requested",
"step_count": 2
}
Each field has a purpose. Application code can validate the type and update rules.
Compare that with storing only a long transcript and hoping the model remembers which order status is current. The transcript is still useful evidence, but explicit fields make critical state easier to inspect and test.
Working memory should be selective
Not every past sentence belongs in current working memory.
Suppose the agent has already looked up the order three times. Keeping all three raw responses in the active state may add noise. A normalized current fact such as order_status="delivered" may be enough, while the full raw responses remain in the audit trace.
This creates two useful layers:
- working state for current decisions;
- trajectory history for audit and debugging.
The first should be concise. The second can be more complete.
Track provenance for important facts
A fact is stronger when you know where it came from.
Instead of:
refund_eligible = true
you might store:
refund_eligible = true
source = refund_policy_tool
observed_at_step = 2
That helps distinguish a trusted tool result from a model inference or a user claim.
You do not need a provenance object for every trivial field. Use it where source identity affects correctness or authority.
State updates need rules
If a new observation contradicts an older value, do not silently keep both as if they are equally current.
Define update rules. A newer trusted order lookup may replace an older trusted order status. A model guess should not overwrite a trusted service result. An approval for one refund amount should not become a general approved=true flag.
State design is therefore part of safety, not merely bookkeeping.
Keep counters outside prose
Step count, retry count, repeated-action count, and cost budget should be numeric fields controlled by application code.
A prompt sentence saying “you have two steps left” can be useful context, but the real counter should live outside the model so the controller can enforce it.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-05-working-memory.py
The script feeds four observations into a small state object. Named fields (such as order_status) keep only the latest value, while recent keeps a short list of the last few events.
- Run it unchanged. In
working state:, find'order_status': 'delivered'. The earliershippingvalue was replaced, not added. - Count the events inside
recent. There are 3: the second order event, the policy event, and the approval event. The firstshippingevent has dropped out. - Read
audit event count: 4. The full event list still exists outside working memory. - Open
l11-05-working-memory.pyand changeMAX_RECENT = 3toMAX_RECENT = 2. Rerun. - Now
recentholds only 2 events (policy and approval), butorder_status,refund_eligible, andapproval_statusare unchanged, and the audit count is still4.
The lesson: shrinking the working list does not lose the facts the agent needs, because important facts live in named fields, and the full history lives in the audit log.
Loading lab…
Quick Check
Explain it back
List five fields your own bounded agent would need during a run. For each field, say whether it comes from the user, a trusted tool, application control logic, or model interpretation. Identify one field that should never be overwritten by model text.
Key Takeaways
- State carries important information across agent cycles.
- Working memory is a selective engineering structure, not the whole transcript.
- Important facts are easier to trust when their provenance is recorded.
- State updates need conflict and replacement rules.
- Budgets and counters belong under application control.
Next Lesson
Next, decide which information should survive beyond one run and how to keep long-term memory from becoming an uncontrolled pile of stale or sensitive text.
References
Completion is stored locally on this device.