Embeddings as Learned Coordinates
Goal
By the end of this lesson, you can explain an embedding as a learned vector, compare Euclidean distance and dot-product similarity, and interpret nearest neighbors relative to the training objective rather than as universal semantic truth.
Discrete IDs do not tell a model how items relate
Suppose four items have IDs:
0, 1, 2, 3
The fact that item 3 has a larger ID than item 1 usually carries no useful semantic meaning.
An embedding maps each discrete item to a learned vector such as:
- item A ->
[0.8, 0.2] - item B ->
[0.7, 0.3] - item C ->
[-0.4, 0.9]
These coordinates can be adjusted by training so geometry becomes useful for the model's objective.
Distance and dot product ask different questions
Euclidean distance asks how far apart two vector endpoints are.
A dot product becomes larger when vectors have aligned directions and/or large magnitudes.
That “and/or” matters.
If you multiply one vector by 100, its raw dot products can become huge even though its direction did not change.
This is why similarity interpretation depends on the metric and normalization convention.
Now compare the same learned points with both measurements:
Embedding space: compare two learned vectors
Pick two points. Distance asks how far apart their endpoints are; dot product depends on both direction and magnitude.
cat vs dog: Euclidean distance = 0.141 · dot product = 1.800
| Item | Vector | What to notice |
|---|---|---|
| cat | [1.0, 0.9] | Close to dog in this toy geometry. |
| dog | [0.9, 1.0] | Close to cat in this toy geometry. |
| car | [-1.0, -0.8] | Close to bus in this toy geometry. |
| bus | [-0.9, -1.0] | Close to car in this toy geometry. |
Move one embedding and inspect geometry
The Lab starts with:
items = ["cat", "dog", "car", "bus"]
embeddings = np.array([
[1.0, 0.9],
[0.9, 1.0],
[-1.0, -0.8],
[-0.9, -1.0],
])
- Click Run and record
cat-dog dot,cat-dog distance,cat-car dot, andcat-car distance. - Change only the first coordinate of the
dogvector from0.9to1.0, sodogbecomes[1.0, 1.0]. - Before running, predict that the Euclidean distance from
cat = [1.0, 0.9]todogshould decrease because the first coordinates now match exactly. - Click Run and compare both
cat-dog distanceandcat-dog dotwith the first run. - Restore
dogto[0.9, 1.0].
Loading lab…
After the guided pass, try a magnitude experiment: multiply both coordinates of cat by 10 while leaving dog unchanged. Predict how the raw dot product changes and explain why a larger raw dot product alone does not prove a more meaningful relationship.
Embedding axes do not need human-readable names
A learned vector can be useful even if coordinate 1 does not mean “animalness” and coordinate 2 does not mean “size.”
The geometry is learned because it helps a training objective.
That also means nearest neighbors are contextual evidence:
“These items are nearby under this learned representation and this metric.”
It is too broad to conclude:
“These items have the same universal meaning.”
ID mapping is a simple but serious failure point
An embedding matrix is a trainable lookup table. Item ID i selects row i.
If preprocessing changes the ID-to-item mapping but the old embedding rows are reused, shapes can still match while every item receives the wrong vector.
Always keep vocabulary/category mapping as part of the model contract.
Quick Check
Key Takeaways
- Embeddings map discrete items to trainable vectors.
- Learned geometry reflects the objective and data.
- Euclidean distance and dot product encode different geometric relationships.
- Magnitude and normalization affect similarity scores.
- Item-to-row mapping is part of the embedding contract.
Next Lesson
Next, you will reuse a representation learned on one task and test whether it helps a related target task with limited labels.
References
- PyTorch, nn.Embedding.
Completion is stored locally on this device.