From Experiment to Explanation
Goal
By the end of this lesson, you can explain a small ML experiment by separating the question, setup, measured evidence, conclusion, limitation, and next step.
A result is not yet an explanation
Suppose a Lab prints:
before mistakes: 1
after mistakes: 0
Those numbers are useful, but someone reading them still has questions:
- What changed between
beforeandafter? - What stayed fixed?
- Which examples were tested?
- What does “0 mistakes” allow us to conclude?
- How much should we trust such a small experiment?
An explanation connects the measurement to the experiment that produced it.
A useful structure is:
question -> setup -> evidence -> conclusion -> limitation -> next step
You do not need to use those words as headings every time. You do need the ideas.
Observation and interpretation are different
Consider these two statements:
Observation: Threshold 4 made 1 mistake on six held-out examples. Threshold 3 made 0 mistakes on those same examples.
Interpretation: Lowering the threshold improved the measured result on this held-out set.
The first statement describes what was measured.
The second statement explains what that measurement means.
Now compare this statement:
“Threshold 3 is the perfect solution and will work for every future user.”
That claim is much larger than the evidence.
Six examples can show what happened on six examples. They cannot prove what will happen in every future situation.
Strong explanations stay close to the evidence
A careful conclusion often includes phrases such as:
- “on this held-out set”;
- “under this metric”;
- “in this experiment”;
- “for these examples.”
That may sound cautious, but caution is not weakness. It is precision.
Compare:
Too broad:
The new threshold is better.
Better:
On the same six held-out examples, lowering the threshold from 4 to 3 reduced the mistake count from 1 to 0.
The second sentence tells us exactly what evidence supports the claim.
Limitations are part of the result
A limitation tells the reader where the evidence may stop being reliable.
For a tiny Level 0 experiment, limitations might include:
- very few examples;
- only one feature;
- one simple threshold model;
- no check on a larger or more varied population;
- labels that may contain mistakes.
Mentioning a limitation does not cancel the experiment. It tells us what the next experiment should investigate.
For example:
“The lower threshold did better on these six examples, but six examples are too few to know whether it will generalize to a larger dataset.”
That sentence both reports the result and points toward the next test.
Try the browser Lab
The Lab below builds a small experiment card from a before/after threshold comparison.
- Click Run without changing anything.
- Read each printed field:
questionchangedkept_fixedevidenceconclusionlimitation
- In the
evidencefield, confirm that the before threshold makes 1 mistake and the after threshold makes 0. - Read the
conclusion. Notice that it says the lower threshold performs better on this held-out set. - Read the
limitation. It explicitly says that the small held-out set does not represent every future case.
Loading lab…
Now change the evidence and watch the explanation respond.
- Find:
after_threshold = 3
- Change only
3to5. - Click Run again.
- Read the new
evidencevalues. - Then read the generated
conclusion.
With threshold 5, the after setting makes more mistakes than before. The conclusion should therefore say that the change did not improve this held-out set.
The important point is that the conclusion follows the evidence. We do not decide the story first and then search for numbers that sound good.
A practical experiment explanation
When you finish a small experiment, try to answer these six questions in ordinary language:
- What were you trying to learn?
- What data or examples did you use?
- What changed?
- What stayed fixed?
- What did you measure?
- What does the result support, and what does it not support?
Here is an example:
I tested whether lowering a threshold from 4 to 3 reduced mistakes. I used the same six held-out examples for both settings and counted prediction mistakes. Threshold 4 made one mistake and threshold 3 made zero. So threshold 3 performed better on this small held-out set. The dataset is very small, so I would want a larger test before making a broad claim.
That paragraph is short, but it is much more informative than “My model got better.”
Avoid mixing evidence with hype
AI results are especially easy to overstate.
Words such as these often need more evidence than a tiny experiment provides:
- “perfect”;
- “intelligent”;
- “solved”;
- “safe”;
- “works for everyone”;
- “always better.”
A useful technical habit is to ask:
Which exact measurement supports this sentence?
If you cannot point to evidence, either gather more evidence or make the sentence narrower.
Reproducibility starts with explanation
A good explanation also helps another person repeat the experiment.
At Level 0, that may mean recording:
- the examples;
- the threshold values;
- the held-out set;
- the metric;
- the before/after result.
Later, reproducibility records will also include things such as code revision, package versions, model checkpoints, random seeds, and hardware settings.
The basic purpose stays the same: someone else should be able to understand what you did and why you believe the result.
Quick Check
Key Takeaways
- Measurements and interpretations are different.
- Explain what changed and what stayed fixed.
- Keep conclusions no broader than the evidence.
- Limitations are part of a trustworthy explanation.
- A good result record helps someone else reproduce the claim.
- Do not choose the story first and bend the evidence to fit it.
Next Lesson
You now have the pieces needed for a complete beginner ML workflow: data, predictions, errors, controlled comparisons, baselines, and evidence-based explanations. In the next Lesson you will use all of them to debug and improve one tiny predictor.
References
- Google for Developers, Introduction to Machine Learning.
Completion is stored locally on this device.