Skip to main content
L1.15

Data Leakage

Goal

By the end of this lesson, you can recognize data leakage, explain why it creates unrealistically good evaluation scores, and describe a safer train/evaluate workflow.

An excellent score can be evidence of a broken experiment​

Imagine taking a quiz while the answer key is hidden inside your notes. Your score may be excellent, but the score no longer measures what you learned.

Data leakage is the machine-learning version of that problem: information reaches training or evaluation that would not really be available when the prediction must be made.

Suppose we want to predict whether a customer will cancel a subscription tomorrow.

FeatureAvailable at prediction time?
Account ageYes
Messages sent this weekYes
Support tickets this monthYes
Cancellation confirmation timestampNo — it appears after cancellation

If the model uses the cancellation timestamp, it can almost read the answer. A high score would estimate an easier, unrealistic task.

A useful question for every feature is:

Could I know this value at the exact moment the real prediction is made?

Predict

A simple model suddenly reaches 99.9% accuracy after one new feature is added. What should you do before celebrating?

Two common ways leakage happens​

Target or future leakage happens when an input directly or indirectly reveals the answer. Predicting loan default using a field created only after default is one example.

Train-test contamination happens when held-out information influences fitting or model selection. For example, computing a normalization mean from the full dataset before the split lets test-set information affect the representation used during training.

The metric calculation can be mathematically correct while the experiment answers the wrong question.

See the leak directly in the Lab​

The Lab contains two candidate features:

  • signal — an imperfect feature that is allowed;
  • leak — an answer-like feature that mirrors the label in this teaching fixture.
  1. Click Run without changing anything.
  2. Compare leaky feature score: with clean feature score:.
  3. Confirm that features = ["signal"] excludes the answer-like leak field from the repaired feature list.
  4. Find train_mean and confirm it is calculated from train rows only.
  5. For one controlled failure, change features = ["signal"] to features = ["signal", "leak"] and explain why that feature list is invalid, even though the current printout does not automatically retrain a model from the list.
  6. Restore features = ["signal"] before moving on.

Loading lab…

The point is not that every suspicious feature is literally named leak. In a real dataset, the dangerous field may have a harmless-sounding name. You have to understand when and how the value was created.

A safer workflow​

Keep evaluation data outside learning from the beginning:

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)

model.fit(X_train, y_train)
score = model.score(X_test, y_test)

A split alone is not enough. Transformations that learn from data must also learn from the training portion only.

A pipeline is one useful way to keep that boundary explicit:

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
score = model.score(X_test, y_test)

During cross-validation, the same rule applies inside every fold: fit preprocessing on that fold's training portion, not once on the full dataset.

Audit a suspiciously good result​

When a score jumps unexpectedly, check in this order:

  1. Prediction moment — when exactly must the real prediction exist?
  2. Feature availability — which values exist at that moment?
  3. Target relationship — was any feature created from the answer or from events after it?
  4. Split boundary — did held-out rows influence fitting or model selection?
  5. Preprocessing — were scalers, imputers, encoders, or selectors fitted on training data only?
  6. Duplicates/groups/time — does the split accidentally put nearly identical or future-related examples on both sides?

Treat an unexpectedly great score as something to explain, not just something to report.

Explain it back​

Use the “answer key” analogy to explain leakage, then give one ML example where a field exists in the stored table but would not be available at the real prediction moment.

Quick Check

1. What makes a feature leaky in a future-prediction task?
2. Why can fitting a scaler on the full dataset before the train-test split be a problem?
3. Validation jumps from 72% to 99.9% after adding one feature. What is the best next step?

0 of 3 questions answered.

Key Takeaways

  • Leakage gives a model information it should not have for the real prediction task.
  • Great-looking metrics can result from an unfair experiment.
  • Check feature availability at the exact prediction moment.
  • Keep evaluation data outside fitting and data-dependent preprocessing.
  • Cross-validation does not repair leakage automatically; fold boundaries still matter.

Next Lesson

Carry this rule forward: before asking whether a model is accurate, first ask whether the evaluation is fair. The next lessons continue model comparison under stricter validation evidence.

References

Lesson actions

Completion is stored locally on this device.

View progress