Pipelines and Reproducible Experiments
Goal
By the end of this lesson, you can build a preprocessing-and-model pipeline, explain how it protects the training boundary during validation, and record enough experiment settings for another person to repeat the comparison.
An ML experiment is more than a model object
Suppose someone tells you:
“I used logistic regression and got 84%.”
You still do not know enough to repeat or judge the experiment.
You need to know things such as:
- which dataset revision and features were used;
- how train and test data were split;
- which random seed controlled the split;
- whether features were standardized;
- where the scaler was fitted;
- the model hyperparameters;
- the evaluation metric.
If these steps live in scattered notebook cells, it is easy to run them in a different order the next time.
A pipeline is an ordered fitting procedure
Think of a pipeline as an assembly line.
For a scaled logistic-regression model, the intended order is:
raw features -> fit/transform scaler -> fit classifier
When the pipeline is fitted on training data, each learned step receives only that training data.
During prediction, the scaler reuses its learned training statistics and the classifier reuses its learned parameters. Neither should refit itself on the test example.
This makes the data boundary part of the code structure rather than a rule you must remember manually in every notebook cell.
Why pipelines matter during cross-validation
Without a pipeline, a common mistake is:
- fit a scaler on all rows;
- transform all rows;
- run cross-validation on the transformed data.
The validation folds influenced the scaler before cross-validation even began.
With a pipeline, each fold does this instead:
- take that fold's training portion;
- fit the scaler on that portion;
- transform the training portion;
- fit the classifier;
- transform and score the held-out fold using the training-fitted scaler.
That keeps learned preprocessing inside the correct boundary.
Run the same experiment twice
The Lab builds make_pipeline(StandardScaler(), LogisticRegression(...)) and uses a fixed split.
- Click Run and record the held-out result:
test accuracy: 0.8. - Click Run again without changing anything.
- Confirm that you get
0.8again. The split seed (random_state=42) and data seed are fixed, so the experiment repeats exactly. - Now change one model setting on purpose. Find
LogisticRegression(max_iter=500, random_state=42)and add a stronger regularization setting:LogisticRegression(max_iter=500, random_state=42, C=0.05). (SmallerCpushes the model toward smaller weights; you will study this idea later.) - Before running, write the change down in your own notes: “changed
Cfrom the default1.0to0.05.” - Click Run. The result becomes
test accuracy: 0.833on the same split and metric. - Now look at
experiment record:. It does not mentionC. If you saved only this record, nobody could reproduce the0.833result. Add"C": 0.05,to theexperiment_recorddictionary and run once more. - Press Reset afterward.
Loading lab…
Reproducibility makes a comparison auditable. It does not make a poor methodology correct.
What a minimal experiment record should contain
For this level, a useful record includes at least:
| Item | Example |
|---|---|
| Data | source/revision and feature definitions |
| Split | strategy, proportions, seed or grouping rule |
| Preprocessing | transformations and where they are fitted |
| Model | class and important hyperparameters |
| Metric | exact held-out or validation measure |
| Result | score plus important failure details |
Later projects add package versions, hardware, model checkpoints, and other reproducibility metadata.
A common misconception
“If a pipeline runs reproducibly, the experiment is trustworthy.”
A pipeline cannot detect that leaky_answer literally contains the target or that the dataset was collected after the event you claim to predict.
A reproducible broken experiment is still broken. You need both repeatable procedure and sound methodology.
Quick Check
Key Takeaways
- An ML result includes data, splitting, preprocessing, model settings, and evaluation choices.
- Pipelines keep learned preprocessing and model fitting in a fixed order.
- Cross-validation should refit the whole pipeline inside each training fold.
- Reproducibility makes evidence auditable but does not guarantee sound methodology.
Next Lesson
Next, you will use the entire Level 1 workflow in a model review: allowed features, fair split, baseline, pipeline, multiple metrics, and failure analysis.
References
- scikit-learn, Pipeline.
- scikit-learn, Common pitfalls and recommended practices.
Completion is stored locally on this device.