Layers and Shapes
Goal
By the end of this lesson, you can predict the input, weight, bias, and output shapes of a dense layer and diagnose a shape mismatch from dimension meaning rather than trial-and-error transposes.
Shapes tell you what an array means
Suppose a batch contains 3 examples and each example has 2 input features.
Then the input matrix has shape:
X: (3, 2)
Now suppose the next layer has 4 neurons.
Using the convention in this level:
X:(batch, input_features)=(3, 2)W:(input_features, output_features)=(2, 4)b:(output_features,)=(4,)X @ W + b:(batch, output_features)=(3, 4)
The result has one row per example and one column per output neuron.
Follow the dense-layer interfaces
Select each stage and read what every dimension means before checking whether the multiplication is legal.
B = 3- three examples in the batch
F_in = 2- two input features per example
Why the inner dimensions must match
In X @ W, each row of X contains 2 feature values. Each output neuron's weights must therefore contain 2 matching input weights.
That is why the inner dimensions are both 2:
(3, 2) @ (2, 4) -> (3, 4)
If you transpose W to (4, 2), you would try:
(3, 2) @ (4, 2)
The inner dimensions 2 and 4 disagree, so the multiplication does not make sense.
Changing batch size should not change learned parameter shapes
If the same model receives 5 examples instead of 3:
Xbecomes(5, 2);Wremains(2, 4);bremains(4,);- output becomes
(5, 4).
Batch size counts examples. It is not a learned feature dimension.
Verify the shapes in the Lab
The Lab sends three two-feature examples through a four-neuron dense layer.
- Before running, write the expected shapes of
X,W,b, and the outputZ. - Click Run and compare with the two output lines:
X (3, 2) W (2, 4) b (4,)andZ (3, 4) A (3, 4). - Now make the batch bigger. Find
X = np.array([[1.0, 2.0], [0.0, -1.0], [3.0, 0.5]])and add two more two-feature examples:
X = np.array([[1.0, 2.0], [0.0, -1.0], [3.0, 0.5], [2.0, 2.0], [-1.0, 0.0]])
- Before running, predict which shapes should change (only the batch size) and which should stay fixed (the layer's
Wandb). - Click Run. You should see
X (5, 2)andZ (5, 4) A (5, 4), whileW (2, 4)andb (4,)are unchanged. The layer's parameters do not depend on how many examples you send at once. - Press Reset afterward.
Loading lab…
When a framework later reports a mismatch such as (32, 128) @ (64, 10), translate the numbers into meanings: “the data currently has 128 features but the layer expects 64.”
A dangerous debugging habit
“Transpose or flatten until the error disappears.”
That can make the code run while mixing examples, features, or channels incorrectly.
Instead:
- write a semantic name beside each dimension;
- identify which dimensions the operation requires to match;
- fix the upstream representation or parameter shape that violates that meaning.
Trace dimensions before tracing numbers
Suppose a batch contains 5 examples, each with 3 input features.
X shape = (5, 3)
A linear layer with 4 output units needs a weight matrix compatible with those feature dimensions. Conceptually:
3 input features → 4 output features
After the layer:
output shape = (5, 4)
The batch count stays 5. The feature width changes from 3 to 4.
If the next layer expects 4 inputs and produces 2 outputs, the shape becomes (5,2).
This simple shape trace often catches bugs before any arithmetic is inspected.
Shape-valid does not always mean meaning-valid
A tensor can have the expected dimensions but the wrong axis meaning.
For example, (5,3) might mean “5 examples × 3 features,” while another operation accidentally treats it as “5 time steps × 3 channels.”
The dimensions still fit some calculations, but the semantics are wrong.
When writing shape notes, label axes with names such as B for batch and C for features instead of recording only raw numbers.
Quick Check
Key Takeaways
- Shape is part of a tensor's meaning.
- Dense-layer multiplication requires matching input-feature dimensions.
- Batch size can change without changing learned parameter shapes.
- Debug shape errors by naming dimensions before reshaping or transposing.
Next Lesson
Next, you will use those shapes to trace a complete two-layer forward pass from input to prediction.
References
Completion is stored locally on this device.