본문으로 건너뛰기
L11.2

Observe, Decide, Act

Goal

Trace one agent cycle by separating observations, decisions, actions, and resulting state updates instead of treating the run as one opaque model conversation.

An agent loop becomes easier to reason about when each step has a job. The simplest useful cycle is observe → decide → act.

An observation is evidence available at the current step: a user goal, a tool result, an approval response, a timeout, or a value already stored in state. A decision chooses what should happen next. An action is the permitted operation that actually runs.

After the action, a new observation becomes available and the cycle can begin again.

Follow one cycle with concrete data​

Start with this state:

goal = "resolve order 4172"
order_status = unknown
approval = none
step = 0

The system observes that the order status is unknown. The decision is get_order. Application code validates and executes the tool. The tool returns {"status":"delivered","damaged":true}.

That result is not merely text to append somewhere. It is a new observation that should update state in a controlled form:

order_status = delivered
damaged = true
step = 1

Now the next decision has different evidence than the first decision.

Do not mix observation and action​

A common design mistake is to store something like “refund the order” as if it were an observation. That phrase is a proposed action, not external evidence.

Keeping categories separate helps answer debugging questions. Did the tool return the wrong fact? Did the model choose the wrong action from a correct fact? Did application code execute an action that policy should have denied?

If the trace blends everything into one transcript, those questions become much harder.

State is the bridge between cycles​

The complete conversation history can be useful, but an agent should not depend on an ever-growing text transcript as its only state representation. Important facts should have explicit fields when practical.

For the running example, useful state might include:

  • goal;
  • current subgoal;
  • trusted order facts;
  • requested and completed actions;
  • approval status;
  • step count;
  • terminal status.

This does not eliminate model context. It gives the application a stable record that does not depend on the model correctly re-reading every old sentence.

Decisions should name evidence​

When possible, record why an action was chosen using inspectable evidence.

For example:

decision: check_refund_policy
because:
order_status = delivered
damaged = true
policy_status = unknown

This is more useful than a vague note such as “the agent thought it was appropriate.” You can test whether the required evidence was present.

Actions still cross the Level 10 boundary​

The decision refund_order is not the refund. Before an external side effect, application code still checks the tool schema, principal permissions, required approval, and legal workflow transition.

The observe–decide–act loop extends the number of decision points. It does not remove the action boundary.

Predict

A tool returns order_status=delivered. Which record belongs in the observation part of the next cycle?

Run the local Lab​

Run:

python labs/notebooks/level-11/l11-02-observe-decide-act.py

The script starts with the action get_order. After each action it records the observation, updates state, and then decides the next action from that observation.

  1. Run it. You should see two cycles: get_order → decision check_refund_policy, then check_refund_policy → decision request_approval. The last line is stopped at: request_approval after 2 cycle(s).
  2. In TOOL_RESULTS, change the get_order result from "status": "delivered" to "status": "shipping" and rerun.
  3. Now the first decision is report_shipping, and the run stops after 1 cycle. The policy check never happens, because the new observation made it unnecessary. state still records order_status and step explicitly.
  4. Change the status back to delivered.

Loading lab…

Quick Check

1. What should an observation represent in an agent trace?
2. Why keep important state in explicit fields instead of relying only on the full text transcript?
3. What should happen after the decision proposes a side-effect action?

0 of 3 questions answered.

Explain it back​

For one two-step task, write four lines for each cycle: observation, decision, action, state update. If you cannot tell which line contains evidence and which contains a proposal, separate them more clearly.

Key Takeaways

  • Observe, decide, and act are different roles in the loop.
  • New tool results become observations that update controlled state.
  • Explicit state fields make important facts and counters inspectable.
  • Decision records are stronger when they name the evidence used.
  • Side effects still require application-level validation and authority checks.

Next Lesson

Next, wrap repeated cycles in an explicit agent loop and make the loop stop safely when progress ends or a budget is reached.

References

Lesson actions

Completion is stored locally on this device.

View progress