Prompting as Interface Design
Goal
Turn a vague request into a clear model-facing contract with task, input, constraints, and output requirements that can be evaluated.
A prompt is often described as “what you ask the model.” That is too loose for reliable engineering. A useful prompt behaves more like an interface contract. It tells the model what job it is performing, which data belongs to that job, what constraints matter, and what output shape is expected.
Compare a vague request with a testable one
Vague:
Summarize this.
More explicit:
Task: summarize the supplied incident report for an engineering handoff.
Input:
<report>
...
</report>
Requirements:
- maximum 4 bullet points
- include observed impact
- include confirmed cause only if the report states one
- if cause is unknown, write "Cause: unknown"
The second prompt is not “magically better” because it is longer. It is better because more of the desired behavior is observable. You can test bullet count, required fields, and whether unsupported causes were invented.
Separate instructions from data
If a prompt mixes instructions and source text without clear boundaries, the model has a harder job distinguishing what to do from what to process. Use explicit delimiters or structured message fields:
Instruction:
Classify the review as positive, negative, or mixed.
Review text:
<<<
The battery is excellent, but the screen is dim.
>>>
Later security lessons add another reason for this separation: source text may be untrusted and can contain instruction-like strings.
Constraints should describe behavior, not wish for quality
Weak:
Be accurate and professional.
More useful:
Use only facts present in the supplied context.
If the context does not support a value, return null.
Return keys: item, quantity, unit.
“Be accurate” is a goal. “Use only supplied context” and “return null when unsupported” are behaviors that can be checked.
Do not overload one prompt with unrelated jobs
Suppose one call asks the model to:
- extract entities;
- summarize the document;
- decide a policy outcome;
- write an email;
- produce analytics.
Even if the model can sometimes do all of this, failures become harder to diagnose. A workflow is often easier to evaluate when it separates tasks with different success criteria. That is the same principle used earlier in ML debugging: isolate boundaries so evidence tells you where a failure began.
Define failure behavior as part of the interface
A useful interface says not only what success looks like but also what to do when the input is insufficient. For example:
If the report does not contain a confirmed cause:
- do not infer one;
- output "Cause: unknown".
That one rule removes an important ambiguity. Without it, a model may try to be “helpful” by filling the missing field with a plausible explanation. Negative requirements are especially useful around known failure modes:
- do not use facts outside the supplied context;
- do not invent a missing identifier;
- do not perform an action when required approval is absent;
- do not silently change the requested output format.
These rules should still be tested. Writing a prohibition in the prompt is not proof that the model will obey it.
Version the contract when behavior matters
If prompt wording changes in a production workflow, the effective interface changed. Record a prompt version or stable fingerprint together with evaluation results. Then “prompt B passed 19/20 cases” refers to a specific artifact rather than whatever text happens to be deployed today. Prompt versioning also makes rollback possible when a seemingly helpful rewrite causes a regression.
Predict
Complete the Lab contract builder
The starter contains one TODO.
- Run it unchanged and observe the failing checks.
- Complete
build_promptso it includes separate task, input, and requirements sections. - Preserve the supplied input exactly.
- Run again and inspect the rendered prompt.
- Change one requirement and explain which evaluation check should change with it.
Loading lab…
Quick Check
Explain it back
Rewrite one vague instruction you have seen into a contract with task, input, constraints, and output expectations. Then name at least two checks that could evaluate the result.
Key Takeaways
- Prompting is interface design, not a collection of magic phrases.
- Good prompts separate task instructions from input data.
- Observable constraints are more useful than vague quality wishes.
- Output requirements should connect directly to evaluation.
- Smaller stages can be easier to debug when one call contains unrelated jobs.
Next Lesson
Next, add role structure and learn why system, user, assistant, and untrusted data should not be treated as interchangeable text.
References
- Brown et al., Language Models are Few-Shot Learners.
Completion is stored locally on this device.