Linear Regression
Goal
By the end of this lesson, you can explain what linear regression does, read the meaning of a line's slope and intercept, fit a model in the browser, and debug a prediction that looks unreasonable.
A straight line as a prediction rule
Many prediction problems ask for a number: a price, a time, a demand estimate, or a temperature.
With one input feature, linear regression uses a rule of the form:
prediction = slope × input + intercept
Imagine placing a straight ruler through a cloud of points. The ruler will not usually touch every point. Fitting the model means choosing a line that keeps the prediction errors small overall.
The slope tells you how much the prediction changes when the input increases by one unit. The intercept is the model's predicted value when the input is zero.
Read one fitted example
The browser Lab uses six house examples:
| Size (m²) | Price (thousand) |
|---|---|
| 40 | 180 |
| 50 | 210 |
| 60 | 245 |
| 70 | 275 |
| 90 | 335 |
| 100 | 365 |
The fitted line is approximately:
price = 3.09 × square_meters + 57.39
So an 80 m² house is predicted at about 304.3 thousand.
That number is not a promise about a real house. It is what this small dataset and straight-line model imply.
Predict
Before running code, use the slope to make a rough prediction. Moving from 80 m² to 85 m² adds 5 m², so the fitted line should increase by about:
5 × 3.09 ≈ 15.45 thousand
That kind of estimate is useful debugging evidence: you know the expected direction and rough size before the computer answers.
Run one controlled change
- Click Run with the starter unchanged and find the prediction for
80.0m². - Find the prediction input
80.0in the Lab code. - Change only
80.0to85.0. - Before rerunning, predict that the output should increase by roughly 15 thousand.
- Click Run and compare the new prediction with your estimate.
- Restore
80.0afterward.
Loading lab…
The built-in checks verify the fitted model, the six-row one-feature dataset shape, and a sensible 80 m² reference prediction.
When a line gives a number you should not trust
Suppose you ask this model for a 300 m² house even though every training example lies between 40 and 100 m².
The formula can still produce a number. That does not mean the relationship continues that far outside the observed range. Predicting beyond the data range is called extrapolation.
A straight line can also fail when the true relationship is strongly curved or when important features are missing. House price, for example, may depend on location, condition, age, and many other variables besides size.
So a valid numerical output is not enough evidence that a prediction is useful.
Debug from assumptions inward
If a regression prediction looks unreasonable, check:
- units — is the input really square meters rather than another unit?
- shape — did the model receive the number of features it expects?
- fit state — was the model fitted before prediction?
- range — is the new input similar to the values used for training?
- model assumption — does a straight relationship make sense for this pattern?
A surprising output is often a data or assumption problem before it is a library problem.
Under the Hood
For each training example, linear regression compares the prediction with the target. A common objective squares those errors and averages them. Training chooses slope and intercept values that make that loss small. The next lesson studies loss functions directly.
Explain it back
If the fitted slope is about 3.09, explain that value using the units of this dataset. A strong answer says that, inside this fitted model, one additional square meter is associated with about 3.09 thousand more in predicted price while the model structure stays fixed. It does not claim that adding one square meter causes a real house to gain exactly that amount in value.
Quick Check
Key Takeaways
- Linear regression predicts a number with a fitted straight-line relationship.
- Slope describes how the model output changes as the input changes.
- Fitting chooses parameters that reduce prediction error on training data.
- Extrapolated outputs can be numerically valid without being trustworthy.
- Debug units, shapes, fit state, range, and assumptions before adding complexity.
Next Lesson
Next, turn individual prediction errors into a single training signal and compare how different loss functions treat small and large mistakes.
References
- scikit-learn, LinearRegression API.
- scikit-learn, Linear Models user guide.
Completion is stored locally on this device.