Delegation and Subagents
Goal
Delegate a subtask with explicit scope, tools, budget, evidence requirements, and return contract while keeping authority with the parent harness.
Delegate a bounded job
Imagine a group project leader asking a teammate to check three experiment logs. A good assignment says which logs to inspect, what question to answer, when to stop, and what evidence to bring back. It does not hand the teammate every password and permission the leader has.
A subagent is similar: it receives a bounded subtask, not the parent's entire authority. The parent should narrow the child's tools and budget and define what evidence must come back. The child's conclusion is then an input for the parent to verify, not automatic permission for the next high-impact action.
Narrow tools and budgets
A good delegation request defines the subgoal, available inputs, allowed tools, resource budget, stop conditions, and expected output shape. This is similar to giving a human teammate a clear assignment. 'Investigate this' is weaker than 'compare these three logs, use only read-only log tools, return the first failing request ID and supporting timestamps.'
{"subgoal":"check three logs","allowed_tools":["read_log"],"step_budget":4,"expected_fields":["request_id","timestamp","evidence"]}
The subagent should receive only the permissions it needs. If the parent can issue refunds, a research subtask usually should not inherit refund authority. Capability narrowing prevents a specialized helper from performing unrelated high-impact actions.
Budgets should also narrow. A parent with ten remaining steps should not spawn five subagents each with ten steps unless the system explicitly accounts for that expansion. The parent harness should reserve and track child budgets so delegation cannot multiply cost invisibly.
Require evidence back
The returned result should contain evidence, not only a conclusion. A subagent that says 'the inventory service is the problem' gives the parent little basis to verify the claim. A better return includes the affected request IDs, timestamps, tool observations, and a confidence or uncertainty note where appropriate.
The parent remains responsible for integrating the result. It should validate the child output shape, check evidence, and decide whether a high-impact next action is legal. A subagent recommendation is not approval.
Failure isolation is another reason to use explicit child tasks. If a delegated search hits its budget, the parent can record CHILD_BUDGET_EXCEEDED and choose another path. Without separate child identity and state, a failure can blur into the parent's own trajectory.
Coordinate child tasks safely
Delegation can be sequential or parallel. Independent research subtasks may run concurrently, but they still follow the dependency and join rules from L12.7. Shared state should be minimized so two children do not silently overwrite one another.
AgentBench and ReAct illustrate that agent behavior can be evaluated as trajectories of actions and observations. Delegated systems add hierarchy to that trajectory. The same principle holds: record observable decisions and effects so the parent-child boundary can be tested.
Predict
Run the local Lab
Run:
python labs/notebooks/level-12/l12-08-delegation.py
The Lab validates a delegation envelope containing subgoal, allowed_tools, step_budget, and expected_fields. Change the child tool list to include refund_order and rerun. Observe that validation rejects the capability because it is outside the delegated scope.
Loading lab…
Quick Check
Explain it back
Design a delegation for analyzing application logs. Specify the child's subgoal, two allowed tools, a budget, the evidence it must return, and one authority the child should not inherit.
Key Takeaways
- Delegation should narrow scope and capability.
- Child budgets must be accounted for by the parent.
- Subagents should return evidence, not only conclusions.
- The parent validates results and retains high-impact authority.
- Child task identity improves tracing and failure isolation.
Next Lesson
Next, L12.9 — Agent Test Fixtures turns known trajectories into repeatable regression tests for the harness.
References
- Liu et al., AgentBench.
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models.
Completion is stored locally on this device.