Skip to main content
L1.14

Cross-Validation

Goal

By the end of this lesson, you can explain k-fold cross-validation, interpret both average performance and fold-to-fold variation, and keep learned preprocessing inside each training fold.

One split can give a lucky or unlucky impression​

Suppose a small dataset is divided into training and validation data once.

By chance, the validation set might contain unusually easy examples. Another split might contain several difficult cases.

If you choose a model from only one small split, your decision can depend too strongly on that accident.

Cross-validation repeats the train/validation process across several partitions so you can see whether performance is stable across different held-out subsets.

A fold is one of those subsets. In k-fold cross-validation, the dataset is divided into k folds, and each fold gets one turn as the validation data.

Five-fold cross-validation, step by step​

Imagine 10 examples divided into five folds of two examples each.

On the first round:

  • train on folds 2–5;
  • validate on fold 1.

On the second round:

  • train on folds 1, 3, 4, and 5;
  • validate on fold 2.

Continue until every fold has served as validation once.

The model is fitted from scratch five times. Cross-validation is not one fitted model being scored five times on different slices.

At the end, you have five validation scores.

The spread matters as well as the average​

Consider two candidates:

  • Model A: [0.82, 0.81, 0.83, 0.82, 0.81]
  • Model B: [0.95, 0.70, 0.94, 0.69, 0.95]

Model B has some excellent folds, but its performance changes dramatically across splits.

A mean score alone can hide that instability. The mean is the average of the fold scores. Their variation tells you how much the scores move from fold to fold.

When you review cross-validation results, inspect:

  • the individual fold scores;
  • their mean;
  • their variation;
  • whether one fold reveals a particular subgroup or difficult region.

Run reproducible folds in the Lab​

The Lab uses StratifiedKFold with five folds. Stratified means each fold tries to preserve roughly the same class proportions as the full dataset.

  1. Click Run without editing anything.
  2. Read fold scores: before the final mean: and std: values. Here std is the standard deviation, a compact measurement of how spread out the fold scores are.
  3. Find the weakest and strongest fold score.
  4. Find this line:
splitter = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
  1. Change only random_state=42 to random_state=17.
  2. Before running, predict what should stay the same and what may change. There should still be five folds and the same 80 examples, but shuffling can put different examples together, so the individual scores, mean, and standard deviation may change.
  3. Click Run and compare the new fold scores: with the first run.
  4. Restore random_state=42 before moving on.

Loading lab…

A fixed seed makes a randomized fold assignment reproducible. It does not make a poor splitting strategy appropriate for every dataset.

Preprocessing must happen inside each fold​

Suppose you standardize the full dataset once and then run cross-validation on the already-standardized values.

The scaler's mean and standard deviation were influenced by rows that later appear in validation folds. That gives the training process information about the validation data before evaluation.

This is a form of data leakage: information crosses a boundary that was supposed to keep evaluation independent.

The safer pattern is:

training fold
→ fit scaler using only that training fold
→ transform training fold
→ fit model
→ transform validation fold using the already-fitted scaler
→ evaluate

You will build this pattern explicitly in L1.17 — Pipelines and Reproducible Experiments.

Choose a splitter that matches the data structure​

Standard k-fold splitting may be inappropriate when:

  • several rows belong to the same person or device;
  • time order matters;
  • class balance needs to be preserved;
  • groups must stay together.

Cross-validation repeats a splitting rule. It cannot rescue a splitting rule that violates the real prediction boundary.

Quick Check

1. What changes between cross-validation folds?
2. Why inspect score variation across folds?
3. Where should learned preprocessing occur?

0 of 3 questions answered.

Key Takeaways

  • K-fold cross-validation refits the model once per fold.
  • Inspect individual fold scores and variation, not only the mean.
  • A fixed seed can make randomized folds reproducible.
  • Data-dependent preprocessing belongs inside each fold's training process.
  • The splitting method itself must match groups, time order, and other data structure.

Next Lesson

Next, in L1.15 — Data Leakage, you will see how an experiment can look excellent because information crossed a boundary it should not cross.

References

Lesson actions

Completion is stored locally on this device.

View progress