Skip to main content
L3.3

Build a Tiny Image Classifier

Goal

By the end of this lesson, you can trace a tiny image classifier from pixels through orientation-sensitive features to a label and use a controlled image change to identify what the decision depends on.

A classifier is a sequence of transformations​

Suppose a 5×5 image contains either a vertical bar or a horizontal bar.

A transparent classifier can work in two stages:

  1. compute a vertical-pattern score and a horizontal-pattern score;
  2. predict whichever orientation has the stronger score.

This is deliberately simple, but it makes the pipeline visible:

pixels -> local feature responses -> feature vector -> decision

A learned CNN makes the feature extraction trainable, but the same debugging question remains: which intermediate representation led to the final decision?

Why the intermediate feature vector matters​

Imagine a vertical image produces features:

[vertical_score=8, horizontal_score=2]

The final label vertical is easy to understand.

Now imagine a preprocessing bug swaps the feature order before classification:

[2, 8]

Shapes still match. The values are still valid numbers. The semantics are reversed.

This is why “the model ran without an error” is weak evidence.

Run the complete image pipeline​

The Lab creates deterministic vertical and horizontal bar images.

  1. Click Run.
  2. Follow the vertical-bar image in order: its pixels are a column of 1.0 values in column 2; its features are [10.0, 0.0] (lots of left-right change, no up-down change); its prediction is class 0, vertical. The horizontal bar gives [0.0, 10.0] and class 1.
  3. Add one bright noise pixel to the vertical image. Directly after the line horizontal[2, :] = 1.0, add a new line:
vertical[0, 0] = 1.0
  1. Before running, predict: which feature score should change, and is the change large enough to flip the label?
  2. Click Run. The vertical image's features become [11.0, 1.0]. Both scores moved by 1, but the vertical-change score is still far larger, so the prediction stays 0 and accuracy stays 1.0. One noisy pixel is small evidence compared with a whole bar.
  3. Press Reset afterward.

Loading lab…

Now think about rotating a vertical bar by 90 degrees while leaving the classifier unchanged. The image is still visually simple, but its orientation representation changes dramatically, so the predicted class should change.

Controlled perturbations tell you what the classifier relies on​

Useful tests include:

  • one-pixel noise;
  • small shifts;
  • rotation;
  • feature-order swaps in a deliberate debug copy.

Each perturbation asks a specific question about dependency.

If several changes are applied at once, you lose the ability to explain which property caused the new output.

A common misconception​

“If the final label is correct, the pipeline is correct.”

Two wrong transformations can occasionally cancel each other and produce the expected class. Inspecting intermediate features protects against that false confidence.

Follow the information path instead of treating the model as one box​

A tiny image classifier usually does more than “image in, label out.”

A simplified path is:

image pixels
→ convolution features
→ nonlinearity
→ pooling or downsampling
→ later features
→ classifier scores

Each stage changes what kind of information is available.

Early convolution filters can respond to small local patterns such as edges or corners. Later features can combine those local responses into patterns useful for the final classes.

When the classifier is wrong, this path gives you places to inspect. A failure may come from preprocessing, feature extraction, the classifier head, or the data itself.

Keep preprocessing part of the model contract​

If training images were scaled to [0,1] but evaluation images arrive as [0,255], the network receives a very different numerical range even though the image looks identical to a human.

Likewise, changing channel order or image size can silently invalidate what the filters learned.

So “same architecture” is not enough for reproduction. The input transformation pipeline is part of the experiment and must travel with the model.

Quick Check

1. What should you inspect before the final class when debugging?
2. Why can swapped feature order be dangerous?
3. What can a controlled corruption reveal?

0 of 3 questions answered.

Key Takeaways

  • An image classifier can be traced as pixels -> features -> decision.
  • Intermediate representations help localize failures.
  • Semantic bugs can survive shape checks.
  • Controlled perturbations reveal what a classifier depends on.
  • A correct final label does not prove every upstream transformation is correct.

Next Lesson

Next, you will deliberately compress spatial representations and examine exactly what max and average pooling keep or lose.

References

Lesson actions

Completion is stored locally on this device.

View progress