Level 1 — Classical Machine Learning You Can Trust
Level 0 taught you to ask basic questions about an AI experiment: What information goes in? What answer are we trying to predict? What changed? What evidence says the result is useful?
Level 1 turns those questions into a complete machine-learning workflow.
Start with two prediction jobs
Imagine two small systems:
- One predicts how many minutes a bus will be late.
- One predicts whether an email is spam or not spam.
Both are machine-learning problems, but the type of answer is different.
- Predicting a number such as
7.5 minutesis called regression. - Predicting a category such as
spamornot spamis called classification.
In both cases, training means choosing model settings from examples. The hard part is not merely making the model produce an answer. We also need to know whether the way we trained and tested it was fair.
For example, a model can appear excellent if information from the test set accidentally leaks into training. A model can also have a high accuracy while repeatedly making the one kind of mistake that matters most for the real decision.
That is why this Level is called Classical Machine Learning You Can Trust: you will learn not only how to fit a model, but how to check the evidence around it.
A few words you will meet
You do not need to memorize these now. Each one will be taught again when it becomes useful.
| Term | Plain meaning in this Level |
|---|---|
| loss | a number that says how wrong the model is according to one chosen rule |
| parameter | an adjustable value inside a model |
| optimization | changing parameters in a direction that improves the chosen objective |
| metric | a measurement we use to evaluate behavior, such as accuracy or recall |
| data leakage | letting information reach training that should have been kept separate for fair evaluation |
| overfitting | doing very well on familiar training examples but not working as well on new examples |
A training objective is the quantity the training process directly tries to improve. A real-world goal is what people actually care about. Those two should be connected, but they are not automatically the same.
What you will learn
By the end of the level, you should be able to:
- organize a dataset into features and a target;
- keep training information separate from final evaluation information;
- fit and inspect regression and classification models;
- explain loss, gradient descent, learning rate, and feature scaling;
- use more than one feature without losing track of order or meaning;
- distinguish a model's score from the final decision made from that score;
- evaluate with accuracy, precision, recall, F1, and cross-validation;
- recognize leakage, underfitting, and overfitting;
- package preprocessing and models into a repeatable pipeline;
- write a model review that uses a baseline, multiple metrics, and failure analysis.
You do not need advanced algebra before starting. When an equation becomes useful, the lesson will first connect it to a small example you can inspect.
What “understanding” means here
A lesson is not finished just because a Lab ran or a Quick Check turned green.
For the important ideas in this Level, try to reach three levels of understanding:
- Explain it: describe the idea in ordinary language without copying the definition.
- Trace it: point to a small example and show where the idea appears in the numbers, table, or output.
- Transfer it: apply the same reasoning to a different small example.
Read Quick Check feedback even after a correct answer. The explanation is part of the lesson because it tells you why the answer follows from the idea.
The learning path
Part 1 — Build a fair regression experiment
L1.1–L1.3 connect training objectives to loss, show how a dataset is organized, and establish a fair train/test boundary.
L1.4 — Linear Regression introduces a simple model for predicting numbers from a straight-line relationship.
L1.5–L1.9 build the rest of the regression workflow: loss functions, gradient descent, learning rate, feature scaling, and multiple features.
After L1.9 — Multiple Features, complete Level 1 Mini Checkpoint — Regression Experiment Review. If you can explain the answers rather than only name the terms, you are ready to move from predicting numbers to predicting categories.
Part 2 — Build and review classifiers
L1.10–L1.14 introduce classification, logistic regression, decision boundaries, decision-relevant metrics, and cross-validation.
L1.15 — Data Leakage shows how an apparently strong result can become invalid when information crosses a boundary it should not cross.
L1.16–L1.18 finish the workflow with overfitting, reproducible pipelines, and a complete model review.
How to use the Labs
Every numbered lesson has a browser or notebook Lab matched to the lesson goal. When a Lab appears, the lesson should tell you:
- which value or line to inspect;
- exactly what to change, if anything;
- which output to read;
- what result to expect or reason about;
- what the result means for the concept.
When code is not the main concept, you do not need to understand every Python line. Focus first on the named input, change, and output. Then connect the result back to the explanation.
A rule that protects almost every experiment
When a score changes, ask more than “Did it get bigger?” Ask:
- What changed?
- What stayed fixed?
- Did any evaluation information influence training?
- Which examples or mistake types changed?
- Is the metric appropriate for the decision?
- Could another person repeat this comparison from the recorded settings?
That habit is the thread connecting the entire Level.
Level Project
After L1.18 — Model Review: Build a Trustworthy Classifier, complete Trustworthy ML: Compare Models Without Cheating.
You will compare a simple baseline with a learned classifier while preserving the train/test boundary, excluding a deliberately leaky feature, reporting more than one metric, and documenting at least one failure/debug path.
A strong project submission should let another learner answer two questions from your evidence:
- Why is this comparison fair?
- Why does the conclusion follow from the measurements?
Completion is stored locally on this device.