Code Execution as a Tool
Goal
Explain why code execution is a high-authority tool and design a bounded execution envelope with limits on code, files, network, time, processes, and returned output.
Code can calculate things that are awkward for a language model to do reliably. It can also read files, open network connections, start processes, consume resources, and change data.
That makes “run code” fundamentally different from a narrow calculator function.
Start with the minimum capability
If the task only needs arithmetic, expose arithmetic. If it needs dataframe operations on one uploaded file, give the executor only that file and the needed libraries.
A general-purpose interpreter should be the result of a deliberate requirement, not the default tool.
A sandbox limits blast radius
A code-execution boundary can restrict:
- filesystem paths;
- network access;
- environment variables and secrets;
- allowed packages or binaries;
- CPU time;
- memory;
- process count;
- output size;
- writable locations.
No sandbox is made safe by prompt wording. The restrictions must come from the runtime or surrounding application.
Inputs and outputs need boundaries too
Do not interpolate untrusted text into shell commands when a structured library call would work.
Likewise, large stdout or generated files should be capped and inspected before being placed back into model context. A program that prints a million lines can exhaust context even if it never escapes the sandbox.
Timeouts are not only convenience
An accidental infinite loop can consume resources. A timeout converts that failure into a bounded, observable result.
The workflow should distinguish:
completed
runtime_error
timeout
resource_limit
policy_blocked
These categories guide different next steps.
Generated code still needs review for high-impact actions
Even inside a sandbox, some tasks can have meaningful side effects—for example modifying an approved working file that will later be published.
Use human approval or deterministic checks where the consequence justifies it. “The model wrote valid Python” is not the same as “the action is acceptable.”
Distinguish code generation from code permission
A model can propose code that is logically correct for a calculation and still request capabilities the current job should not have. Review therefore happens on two axes: what the program computes and what the runtime lets it touch. Sandboxing answers the second question even when the first answer is uncertain.
For example, a CSV analysis may legitimately read /workspace/input.csv and write /workspace/result.csv. It does not need to read ~/.ssh, inspect environment secrets, or make outbound requests. A narrow mount and denied network make those boundaries real regardless of generated code content.
Reproducibility needs environment identity
If code output matters to a decision, record the interpreter/runtime version, allowed packages, input artifact hashes or IDs, resource limits, and generated source or source hash. “It ran successfully” is weak evidence if another person cannot recreate the same execution envelope.
Predict
Run the local Lab
python labs/notebooks/level-10/l10-08-code-policy.py
The Lab checks three code-execution requests against a policy with no network, a 3-second limit, a 2,000-byte output limit, and one writable folder.
- Run the command. You should see
calc -> allowed,web -> policy_blocked_network, andlarge-output -> resource_limit_output. - Notice the order of checks in
check: network permission first, then time, then output size, then write paths. The first failing rule is the one reported. - Lower the output limit: change
"max_output_bytes": 2000,to"max_output_bytes": 50,and rerun. - Now even
calc, which prints only 100 bytes, becomesresource_limit_output. A limit that is too tight blocks legitimate work, so limits must fit the task. The script then stops with anAssertionError, which is expected: its final checks were written for the original values. Change the value back afterward.
Loading lab…
Quick Check
Key Takeaways
- General code execution has much broader authority than narrow tools.
- Sandboxing must enforce filesystem, network, resource, and secret boundaries.
- Prefer structured library calls over shell interpolation.
- Time and output limits turn runaway behavior into bounded failures.
- High-impact generated actions may still need deterministic checks or approval.
Next Lesson
Next, move from text-only evidence to images and learn how visual inputs add another source of observations and ambiguity.
References
- Open Containers Initiative, Runtime Specification.
- OWASP GenAI Security Project, Top 10 for LLM Applications.
Completion is stored locally on this device.