Representations: What Networks Learn
Goal
By the end of this lesson, you can explain a representation as an internal description of an input and reason about which information a transformation preserves, emphasizes, or discards.
The same example can be described in different coordinates
Imagine a tiny image. One description is the raw pixel grid. Another description might contain only two numbers:
- vertical-edge strength;
- horizontal-edge strength.
Both descriptions refer to the same image, but they make different information easy to use.
A representation is an internal description produced from an input for later computation.
In a neural network, a hidden layer computes something like:
h = f(x)
The next layer receives h, not the original x directly.
Why a new representation can make a task easier
Suppose four raw points are awkward to separate with one simple rule.
A transformation can move those points into new coordinates where examples from the same class are closer or a simple boundary becomes possible.
This does not mean the network discovers human-readable concepts automatically. A hidden coordinate may combine several properties in a way that is useful for the training objective without having a neat name.
The useful question is not “What English word does this neuron mean?” It is:
What information in this representation helps the next computation?
Compression can help and hurt at the same time
Imagine a representation that keeps only edge orientation and discards exact location.
That may help a task that asks “vertical or horizontal?”
It may hurt a task that asks “is the edge on the left or the right?”
Representation quality is therefore task-dependent. A compact feature is not automatically better merely because it uses fewer numbers.
Inspect a deterministic feature map
The Lab transforms four 2D points and prints both the coordinates and a distance between the first and last point.
- Click Run without editing anything.
- Read
raw:,representation:,pair distance raw:,pair distance representation:, andaccuracy:. - Find this line:
representation = transform(raw)
- Change it to:
representation = transform(raw, second_scale=0.25)
- Before running, predict what should stay fixed and what should change. The first representation coordinate
x1*x2is unchanged, so the simple classifier should keep the same accuracy. The second coordinate is four times smaller, so the printed representation distance between the first and last point should shrink. - Click Run and compare
pair distance representation:andaccuracy:with the first run. - Restore
representation = transform(raw).
Loading lab…
After the guided pass, try a different positive second_scale. Explain which part of the representation changes and why the classifier can remain unchanged even while a geometric distance changes.
Now consider a deliberately destructive change: set the transformed second coordinate to zero for every point. The transformation still produces valid numbers, but it erases one dimension of information.
That is the kind of failure you should learn to recognize: valid computation does not guarantee useful representation.
Debug representations by asking what collapsed
When an internal representation looks unhelpful, inspect:
- feature ranges;
- pairs of different inputs that became identical or nearly identical;
- which task-relevant distinctions disappeared;
- whether one scale dominates the distance or downstream score.
The first useful hypothesis often comes from identifying information that was unintentionally removed.
Quick Check
Key Takeaways
- A representation is an internal description produced from an input.
- Later layers operate on representations rather than raw inputs directly.
- A useful representation emphasizes information needed by the task.
- Compression and invariance always trade some information away.
- Representation quality must be judged relative to a task and evidence.
Next Lesson
Next, you will build image representations with a local operation that reuses the same detector across positions: convolution.
References
- PyTorch, Tensors.
Completion is stored locally on this device.