Skip to main content
L11.13

Bounded Agent Integration Workshop

Goal

Integrate explicit goals, subgoals, state, memory rules, tool selection, critique, approvals, stopping conditions, and trajectory evaluation into one reproducible bounded-agent run.

The final Level 11 system is not “a model that keeps calling tools.” It is a controlled loop whose decisions remain inspectable.

A reviewable run looks like:

trusted goal + principal
→ initial state
→ observe
→ choose or revise subgoal
→ choose next action
→ critique when policy calls for it
→ validate permission / approval / transition
→ execute or stop
→ normalize observation
→ update state and memory candidates
→ check terminal conditions
→ repeat within budget
→ evaluate the full trajectory

Each arrow is a place where a later failure can begin.

Keep four kinds of records separate​

A strong project trace separates:

  1. goal and policy — what the run is trying to do and what it is allowed to do;
  2. working state — current facts, subgoal, counters, and approval state;
  3. trajectory history — ordered observations, decisions, actions, and results;
  4. evaluation — recomputed metrics and release decision.

Do not let a generated summary silently become the authoritative policy or overwrite counters.

Reuse Level 10 action boundaries​

Every selected tool still passes schema, permission, approval, retry, and state-transition checks.

The agent loop adds continuation decisions around those boundaries. It does not weaken them.

For example, suppose the model proposes a refund on step 4. The application should be just as careful as a one-shot Level 10 workflow. Validate the arguments and check the principal. Verify the exact approval, execute once, normalize the result, and record the transition.

Make the stopping reason part of the deliverable​

Your project should show at least these outcomes:

  • successful completion;
  • blocked on missing information;
  • denied high-impact action;
  • repeated-action or no-progress stop;
  • budget stop;
  • intentional unsafe fixture rejected by validation.

The final text alone is not enough. Record the terminal category.

Keep memory writes narrow​

The project includes persistent-memory candidates, but not every observation should be saved.

A valid write should name its category, source, scope, and freshness or retention information. The validator checks the deterministic policy rules you implement.

Retrieved memory should enter working state only if it matches the current scope and freshness rule.

Evaluate the release rule from raw trajectories​

Do not trust precomputed metrics in a submitted JSON file.

The validator should recompute values such as:

task_success
action_selection_accuracy
unauthorized_executions
approval_bypasses
runaway_loops
memory_policy_violations
average_success_steps

Then recompute whether the release thresholds pass.

This prevents a project from declaring success by editing the summary numbers.

The required project path is provider-neutral​

The required Bounded Agent project uses recorded fixtures and the Python standard library. That lets another learner reproduce the same controller, policy, state, memory, and evaluation behavior without an API key.

You may add a live model as optional evidence. If you do, keep the deterministic baseline and record the model/provider version separately.

Review capability and control separately​

First ask whether the agent can make useful progress. Check its subgoals, tool choices, state updates, and task outcomes.

Then ask whether the run stayed inside its controls. Check legal transitions, bounded steps, scoped memory, exact approvals, and zero unauthorized side effects.

A capable but uncontrolled agent is not a passing project. A perfectly safe agent that always stops immediately is also not sufficient.

Bridge to Level 12​

Level 12 will add durability, tracing, recovery, sandboxing, testing, delegation, and operational controls. The bounded loop from this Level becomes the unit that the harness must run safely over time.

That next step only works if the basic trajectory is already explicit. Production infrastructure cannot rescue an agent whose goal, state, permissions, or stopping logic are undefined.

Predict

A submitted run JSON says release_passed=true, but recomputed traces contain one approval bypass. What should the validator do?

Run the integration Lab​

Run the passing fixture:

python labs/notebooks/level-11/l11-13-integration-check.py projects/reference/l11/sample-run/agent-run.json

Then run the intentional failure fixture:

python labs/notebooks/level-11/l11-13-integration-check.py projects/reference/l11/failure-run/agent-run.json

The first should print a passing integration marker. The second should return a non-zero status and identify at least one unsafe or unbounded condition.

Loading lab…

Quick Check

1. Why separate working state from trajectory history in the final project?
2. What must happen to submitted summary metrics such as task_success and unauthorized_executions?
3. What does Level 12 inherit from this project?

0 of 3 questions answered.

Explain it back​

Trace one project case from the initial goal through every state change to the terminal reason. For each action, identify the evidence that supported it, the deterministic controls that allowed or blocked it, and the metric that would reveal a failure at that boundary.

Key Takeaways

  • A bounded agent is a controlled decision loop, not simply repeated tool calls.
  • Keep goal/policy, working state, trajectory history, and evaluation separate.
  • Reuse Level 10 authorization and action boundaries inside every loop step.
  • Record terminal reasons and narrow memory behavior as project evidence.
  • Recompute release metrics from raw trajectories and keep safety constraints independent.

Next Lesson

Complete Bounded Agent. Level 12 will wrap this explicit bounded loop in a production-grade agent harness.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Bounded Agent

View progress