Failure Analysis Across Modalities
Goal
By the end of this lesson, you can design reproducible evaluation slices for image and sequence failures and explain why one overall metric can hide concentrated weaknesses.
An average can hide the condition that matters
Suppose an image classifier has:
- clean accuracy: 96%;
- noisy-image accuracy: 62%.
If you mix many clean examples with a few noisy examples, the combined average may still look high.
The model has not become robust. The aggregate has hidden a specific failure condition.
A slice is a meaningful subset or stress condition evaluated separately.
Different modalities suggest different hypotheses
For images, useful slices might include:
- controlled pixel noise;
- brightness change;
- rotation;
- small translation;
- particular classes or sources.
For sequences, useful slices might include:
- reversed order;
- longer sequence length;
- relevant context placed farther away;
- missing steps;
- a different time period.
A slice is useful when it tests a real hypothesis about model dependence.
Compare clean and stressed conditions in notebook
The notebook evaluates:
- transferred digit classifier on clean versus noisy inputs;
- a transparent sequence trend task on original versus reversed order.
- Open the notebook and run the cell.
- Record the four results as two separate pairs:
image_clean_accuracywithimage_noisy_accuracy, andsequence_clean_accuracywithsequence_reversed_accuracy. - Look at the sequence pair. The trend rule is perfect on the original sequences (
1.000) and completely wrong on the reversed ones (0.000). Reversal kept every value but flipped every trend, so the rule's single piece of evidence—last minus first—points the wrong way on every example. - Now change only the image-noise strength. Find
rng.normal(0,0.35,Xv.shape)and change0.35to0.6. - Before running, predict which result should change: the clean image score, the noisy image score, or the sequence scores?
- Run the cell again. Only
image_noisy_accuracycan change, and stronger noise usually lowers it. (If it does not drop, that is interesting evidence too: this classifier may rely on coarse strokes that survive noise.) The clean images, the validation split, and the seed stayed fixed, so the comparison isolates the noise strength. - Restore
0.35afterward.
Loading lab…
Reversal is a controlled test of order dependence
A reversed sequence contains the same values but different order.
If the task label depends on direction and you deliberately keep the original label, performance can collapse because the semantic condition changed.
This isolates a model/task assumption much more clearly than adding several sequence corruptions at once.
A slice must be reproducible too
“Hard examples” is not a sufficient slice definition.
Record:
- exact condition or transformation;
- sample count;
- metric;
- seed/procedure;
- clean control;
- representative examples or IDs.
Very tiny slices can also be noisy. Report their size so a dramatic percentage is not mistaken for strong evidence.
One average score can hide different failure populations
Imagine two image classifiers both score 90% accuracy.
Model A fails mostly on dark images. Model B fails mostly on one rare class.
The same average score leads to different engineering actions.
Sequence models have analogous slices: long sequences, rare tokens, early-versus-late events, or particular sequence patterns.
Choose slices connected to plausible mechanisms
Useful failure slices are not arbitrary columns from the data table. They should help test a hypothesis.
For images, inspect brightness, size, occlusion, or class.
For sequences, inspect length, position, missing events, or domain.
Then compare the same slices across candidate models under the same evaluation set.
Failure analysis is strongest when it turns “model B seems worse sometimes” into a measurable statement such as “error rate doubles for sequences longer than 50 steps.”
Quick Check
Key Takeaways
- Aggregate metrics can hide concentrated weaknesses.
- Evaluation slices should correspond to meaningful failure hypotheses.
- Image and sequence modalities require different stress conditions.
- Keep a clean control and preserve slice metadata.
- Report sample counts so slice uncertainty remains visible.
Next Lesson
Next, you will put scratch training, frozen transfer, and fine-tuning on one fair adaptation scorecard and make an engineering decision from the evidence.
References
Completion is stored locally on this device.