Rules, Examples, and Data
Goal
By the end of this lesson, you can tell the difference between a rule, an example, a dataset, and a label, and explain why labels matter in supervised learning.
Four words that are easy to mix up
Before we talk about learning, we need to separate a few basic ideas.
A rule tells a system what to do.
An example is one case we observed.
A dataset is a collection of examples.
A label is the target answer attached to an example when we want to teach or evaluate a supervised prediction task.
These are related, but they are not the same thing.
Start with a hand-written rule
Suppose we want to mark temperatures below 0°C as freezing.
A person could write this rule:
If temperature is below 0°C, output
freezing. Otherwise outputnot_freezing.
That rule already contains the decision logic. The computer does not need to discover the threshold. We gave it the threshold directly.
Now compare that with one observed example:
temperature = −3°C, label =
freezing
The example does not say, “all temperatures below 0°C are freezing.” It only tells us what happened, or what answer we assigned, for one case.
A dataset contains many such cases:
| Temperature | Label |
|---|---|
| −3°C | freezing |
| −1°C | freezing |
| 2°C | not_freezing |
| 5°C | not_freezing |
The rows are examples. The whole table is a dataset. The right-hand column contains labels.
Why do labels matter?
Imagine showing a learner only these temperatures:
−3, −1, 2, 5
Those numbers are inputs, but they do not tell the learner what answer we want.
Should −3 mean “freezing”? “dangerous”? “winter”? “turn on the heater”? The input alone does not define the task.
When we attach labels, we add information about the target answer:
−3 → freezing
−1 → freezing
2 → not_freezing
5 → not_freezing
In supervised learning, these target answers are part of the learning signal. A model can compare its prediction with the label and use the difference to decide whether its settings need to change.
Later we will often write one supervised example as (x, y):
xmeans the inputymeans the target answer or label
You do not need that notation yet, but it is simply a shorter way to describe the same pair.
Labels can be wrong or ambiguous
A label is an answer stored in a dataset, but that does not make it automatically correct or perfectly clear.
There are two different problems worth separating.
A wrong label: someone recorded the wrong answer
Suppose our labeling instruction is very clear:
Temperatures below 0°C should be labeled
freezing.
Then someone accidentally enters this row:
| Temperature | Label |
|---|---|
| −1°C | not_freezing |
Here the problem is straightforward: the stored label does not follow the labeling rule. It may be a typing or data-entry mistake.
An ambiguous label: reasonable people may disagree
Now imagine labeling email as spam or not_spam.
A shop sends this message:
“Weekend sale: 50% off.”
One person may have asked to receive shop promotions and call it not_spam. Another person may not want promotional email and call it spam.
Neither person necessarily made a typing mistake. The word spam is being used with a boundary that may not be perfectly clear unless the labeling instructions explain it.
That is what ambiguous means here: the example does not have one obvious label that every reasonable labeler would choose.
Why does this matter for a model?
A model sees the examples and labels we give it. It does not automatically know whether a surprising label is:
- a simple mistake;
- a real exception;
- a sign that useful information is missing;
- or a case where the target itself is uncertain.
So when labels look inconsistent, the first response should not be “make the model bigger.” We should inspect the examples, the labeling instructions, and the missing context.
The Lab below demonstrates the simpler case: one clearly changed label creates a disagreement with a fixed rule.
Guided Lab: change exactly one label
The Lab below contains the same four temperature examples in Python form. You do not need to understand Python dictionaries yet. Treat each {...} block as one row of the table.
First, click Run without changing anything.
Look at the final line beginning with:
mistakes:
With the original labels, it should show an empty list: []. That means the fixed rule agrees with every example in this tiny dataset.
Now make one controlled change:
-
In the Lab below, find the second example:
{"temperature": -1, "label": "freezing"} -
Change only the label text from
"freezing"to"not_freezing". -
Leave the temperature
-1unchanged. -
Click Run again.
-
Look at the
mistakes:output.
You should now see the edited row listed as a mistake because the rule still predicts freezing for −1°C, while the new label says not_freezing.
Loading lab…
What this experiment shows
We changed no input values and no rule. We changed only one target answer.
The disagreement appeared because labels affect what counts as correct.
That is a more useful lesson than simply saying “labels are important”: you can now point to exactly where the label enters the comparison.
Example versus dataset
One more distinction matters.
If one student solves one math problem correctly, that is one example. It does not prove the student can solve every problem of that type.
Likewise, one machine-learning example tells us about one case. A dataset gives us more evidence because it contains many cases.
But a large dataset can still be poor if it repeats the same narrow situation or contains unreliable labels. More rows and better evidence are not always the same thing.
A common misconception
It is easy to think:
“The data is the rule.”
Not quite.
Data provides examples. A hand-written rule may be chosen by a person. A learning algorithm may use examples to choose model settings. The examples are evidence the system can use; they are not automatically the final decision rule.
That distinction sets up the next lesson.
Quick Check
Key Takeaways
- A rule tells a system what to do; an example records one case.
- A dataset is a collection of examples.
- In supervised learning, labels are target answers attached to examples.
- A wrong label can be a recording mistake; an ambiguous label can come from a target whose boundary is not perfectly clear.
- Examples provide evidence; they are not automatically the final rule.
Next Lesson
Next, you will compare two ways to get a decision rule: a person can choose the setting directly, or a learning process can choose a setting from labeled examples.
References
- Google for Developers, Introduction to Machine Learning.
Completion is stored locally on this device.