본문으로 건너뛰기
L1.3

Splitting Data Without Cheating

Goal

By the end of this lesson, you can create a train/test split, explain why test examples must stay outside training, and use a fixed random seed to make a comparison reproducible.

Practice questions and a final check have different jobs​

Suppose you practice with one stack of questions and take a final exam with a different stack.

During practice, it is fine to inspect mistakes and change how you study. The practice questions are helping you learn.

The final exam has a different purpose: it checks whether what you learned works on questions that did not guide your studying.

Machine learning uses the same separation:

  • training data is allowed to influence the fitted model;
  • test data is held out so it can provide a more independent evaluation later.

If the same examples appear on both sides, the evaluation partly asks whether the model can handle cases it already used while learning.

A tiny split​

Imagine 12 numbered examples. A 75/25 split should put 9 examples in training and 3 in testing.

The specific IDs can vary when the split is random, but two rules should hold:

  1. training and test IDs should not overlap;
  2. the split should be reproducible when we are comparing models under the same conditions.

A random seed is a starting value used by a pseudo-random process. Using the same seed with the same data and split procedure makes the same random choices again.

The seed does not make the split automatically good. It makes the random part repeatable.

Run the split Lab​

The Lab creates 12 examples and calls train_test_split with test_size=0.25 and random_state=42.

  1. Before running, predict the train and test counts for the 75/25 split: 9 training examples and 3 test examples.
  2. Click Run.
  3. Read train ids:, test ids:, and overlap:. Confirm that overlap: is [].
  4. Find this argument in the split call:
test_size=0.25
  1. Change only 0.25 to 0.50.
  2. Before running, predict that 6 of the 12 examples will now be held out for testing and 6 will remain for training.
  3. Click Run and count the IDs on each side. Confirm that overlap: is still empty.
  4. Restore test_size=0.25 before moving on.

Loading lab…

The empty overlap is evidence that no row is literally present in both sets. It is necessary, but it is not the whole fairness story.

Random rows are not always the right split​

Suppose one patient has ten medical visits. A random row split might put some visits from that patient into training and other visits into testing.

The model may then benefit from recognizing patient-specific patterns rather than showing how well it handles a truly new patient.

Similar problems appear when:

  • several rows belong to the same user, machine, family, or document;
  • future records are mixed into training for a prediction that should happen earlier in time;
  • near-duplicate examples cross the split boundary.

In those cases, a group-based or time-based split may be more honest than a random row split.

Another common mistake: tuning on the test set​

Even without row overlap, a test set can lose its independence if you repeatedly look at its score and change the model because of what you saw.

Once test results guide decisions, those examples have become part of the development process.

Later you will use validation data and cross-validation for repeated model selection while protecting a final test set.

Quick Check

1. What is the main purpose of a test set?
2. Why use a fixed random seed in a comparison?
3. When might a random split be misleading?

0 of 3 questions answered.

Key Takeaways

  • Training data may influence learning; held-out test data is for a later check.
  • Training and test rows should not overlap.
  • A fixed seed makes random choices reproducible, not automatically fair.
  • Grouped or time-ordered data may need a structure-aware split.
  • Repeatedly tuning from test results weakens the test as independent evidence.

Next Lesson

Next, in L1.4 — Linear Regression, you will learn a simple model for predicting a number from a straight-line relationship.

References

Lesson actions

Completion is stored locally on this device.

View progress