본문으로 건너뛰기
L1.8

Feature Scaling

Goal

By the end of this lesson, you can standardize numeric features using training data, explain why very different numerical scales can hinder some algorithms, and avoid fitting a scaler on evaluation data.

Different units can create very different numerical ranges​

Imagine two features for the same person:

  • age: about 20 to 60 years;
  • income: about 25,000 to 120,000 dollars per year.

The income numbers are thousands of times larger, but that does not mean income is automatically thousands of times more important.

Some algorithms care about numerical distance or use gradient-based updates. For them, very different feature scales can make learning awkward or make one coordinate dominate the geometry.

Feature scaling changes the numerical ruler used for a feature. It does not create new information about the example.

Standardization gives values a shared interpretation​

A common transformation is standardization:

z = (x - training_mean) / training_standard_deviation

After standardization, a value tells us roughly how many training-set standard deviations it lies above or below the training-set mean.

For example, z = 0 means “at the training mean,” while z = 2 means “two training standard deviations above it.”

The person or house represented by the row has not changed. Only the numerical representation has.

The most important boundary: fit the scaler on training data​

A scaler learns values such as the mean and standard deviation.

Those learned statistics are part of the fitting process.

So the correct order is:

  1. split the data;
  2. fit the scaler using training rows only;
  3. transform training rows with those learned statistics;
  4. transform test rows using the same training statistics;
  5. never refit the scaler on the test set for that evaluation.

If you compute the mean and standard deviation from the full dataset before splitting, information from the test rows has already influenced preprocessing.

That is a form of leakage.

Inspect scaling in the Lab​

The Lab contains age and income values and uses StandardScaler.

  1. Click Run.
  2. Read training means learned by scaler: [40.0, 68400.0]. These are the average age and average income of the five training rows only.
  3. Read scaled training column means: [0.0, 0.0]. After scaling, each training column is centered at zero.
  4. Read scaled test rows: [[-0.354, -0.544], [1.061, 1.082]]. The test rows were transformed with the training statistics. Do not expect them to average exactly zero; the scaler never looked at them.
  5. Find the last training row, [60.0, 120000.0],, and change only the income 120000.0 to 220000.0.
  6. Before running, predict which learned statistic should move: the age mean, the income mean, or both?
  7. Click Run. The income mean becomes 88400.0 while the age mean stays 40.0. The second test row's scaled income also changes, from 1.082 to 0.239, even though you never touched the test data. Test rows are transformed with whatever the training data taught the scaler.
  8. Press Reset afterward.

Loading lab…

The test transformation can change because the training mean or standard deviation changed. That is expected: the test rows are transformed with statistics learned from training.

Scaling is not useful in exactly the same way for every model​

Linear models, distance-based methods, and gradient-based optimization often benefit from comparable numerical scales.

Tree-based models usually do not need standardization to decide where to split a feature, because their split logic depends on ordering and thresholds rather than Euclidean scale in the same way.

This is a useful reminder that preprocessing should match the algorithm, not become a ritual applied to every dataset.

A common misconception​

“Scaling makes the data more informative.”

Scaling changes representation, not the underlying predictive signal. A useless feature remains useless after standardization, and a leaky feature remains leaky.

Quick Check

1. What does standardization use?
2. Why fit the scaler on training data only?
3. Does scaling add new predictive information?

0 of 3 questions answered.

Key Takeaways

  • Scaling changes numerical units, not the identity of an example.
  • Standardization uses a mean and standard deviation learned from training data.
  • Split first; fit data-dependent preprocessing only on training rows.
  • Different model families depend on scaling differently.

Next Lesson

Next, you will combine several features in one linear model and learn how to keep their shapes, order, and interpretations straight.

References

Lesson actions

Completion is stored locally on this device.

View progress