Data Leakage
Goal
By the end of this lesson, you can recognize data leakage, explain why it creates unrealistically good evaluation scores, and describe a safer train/evaluate workflow.
An excellent score can be evidence of a broken experiment
Imagine taking a quiz while the answer key is hidden inside your notes. Your score may be excellent, but the score no longer measures what you learned.
Data leakage is the machine-learning version of that problem: information reaches training or evaluation that would not really be available when the prediction must be made.
Suppose we want to predict whether a customer will cancel a subscription tomorrow.
| Feature | Available at prediction time? |
|---|---|
| Account age | Yes |
| Messages sent this week | Yes |
| Support tickets this month | Yes |
| Cancellation confirmation timestamp | No — it appears after cancellation |
If the model uses the cancellation timestamp, it can almost read the answer. A high score would estimate an easier, unrealistic task.
A useful question for every feature is:
Could I know this value at the exact moment the real prediction is made?
Predict
Two common ways leakage happens
Target or future leakage happens when an input directly or indirectly reveals the answer. Predicting loan default using a field created only after default is one example.
Train-test contamination happens when held-out information influences fitting or model selection. For example, computing a normalization mean from the full dataset before the split lets test-set information affect the representation used during training.
The metric calculation can be mathematically correct while the experiment answers the wrong question.
See the leak directly in the Lab
The Lab contains two candidate features:
signal— an imperfect feature that is allowed;leak— an answer-like feature that mirrors the label in this teaching fixture.
- Click Run without changing anything.
- Compare
leaky feature score:withclean feature score:. - Confirm that
features = ["signal"]excludes the answer-likeleakfield from the repaired feature list. - Find
train_meanand confirm it is calculated fromtrainrows only. - For one controlled failure, change
features = ["signal"]tofeatures = ["signal", "leak"]and explain why that feature list is invalid, even though the current printout does not automatically retrain a model from the list. - Restore
features = ["signal"]before moving on.
Loading lab…
The point is not that every suspicious feature is literally named leak. In a real dataset, the dangerous field may have a harmless-sounding name. You have to understand when and how the value was created.
A safer workflow
Keep evaluation data outside learning from the beginning:
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
A split alone is not enough. Transformations that learn from data must also learn from the training portion only.
A pipeline is one useful way to keep that boundary explicit:
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
During cross-validation, the same rule applies inside every fold: fit preprocessing on that fold's training portion, not once on the full dataset.
Audit a suspiciously good result
When a score jumps unexpectedly, check in this order:
- Prediction moment — when exactly must the real prediction exist?
- Feature availability — which values exist at that moment?
- Target relationship — was any feature created from the answer or from events after it?
- Split boundary — did held-out rows influence fitting or model selection?
- Preprocessing — were scalers, imputers, encoders, or selectors fitted on training data only?
- Duplicates/groups/time — does the split accidentally put nearly identical or future-related examples on both sides?
Treat an unexpectedly great score as something to explain, not just something to report.
Explain it back
Use the “answer key” analogy to explain leakage, then give one ML example where a field exists in the stored table but would not be available at the real prediction moment.
Quick Check
Key Takeaways
- Leakage gives a model information it should not have for the real prediction task.
- Great-looking metrics can result from an unfair experiment.
- Check feature availability at the exact prediction moment.
- Keep evaluation data outside fitting and data-dependent preprocessing.
- Cross-validation does not repair leakage automatically; fold boundaries still matter.
Next Lesson
Carry this rule forward: before asking whether a model is accurate, first ask whether the evaluation is fair. The next lessons continue model comparison under stricter validation evidence.
References
- scikit-learn, Common pitfalls and recommended practices — data leakage.
- scikit-learn, Pipeline.
- scikit-learn, Cross-validation: evaluating estimator performance.
Completion is stored locally on this device.