Vector Similarity
Goal
Compute dot product and cosine similarity, explain the role of vector magnitude, and use one consistent similarity contract for indexing and querying.
Before the formulas, picture arrows on graph paper. Two arrows can point in the same direction while one is much longer. A scoring rule can care about both direction and length, or it can compare mostly the direction after normalizing the lengths.
That is the intuition behind the two scores in this Lesson. Dot product is influenced by alignment and magnitude. Cosine similarity divides out vector length and compares direction more directly. The tiny two-dimensional vectors below let you see that difference with arithmetic before applying the same rule to high-dimensional retrieval embeddings.
Remember from L5.4 — Dot Products as Similarity: a dot product multiplies matching vector components and adds them. Here you reuse that same operation for retrieval, then compare it with cosine similarity.
After embedding a query and document chunks, a retriever needs a score. Two common choices are:
- dot product;
- cosine similarity.
They are related but not identical.
Dot product mixes alignment and magnitude
Take:
q = [1, 1]
d1 = [1, 1]
d2 = [10, 10]
Dot products:
q·d1 = 2
q·d2 = 20
Both document vectors point in the same direction, but d2 has larger magnitude. So raw dot product rewards both alignment and scale. That may be intended if the embedding model was trained for dot-product retrieval.
Cosine similarity normalizes magnitude
Cosine similarity is:
cos(q,d) = (q·d) / (||q|| ||d||)
For the two vectors above, both cosines are 1 because their directions match exactly. Cosine asks more directly about angle/direction. For:
q = [1,0]
d3 = [0,1]
the dot product is 0 and cosine is 0.
Normalized vectors make dot and cosine agree
If every vector has unit norm:
||q|| = 1
||d|| = 1
then:
q·d = cos(q,d)
Some systems normalize embeddings first and then use inner-product search. The important rule is consistency. Do not build an index assuming normalized cosine-style vectors and then query with unnormalized vectors under another score.
Zero vectors need handling
Cosine similarity divides by vector norms. A zero vector has norm 0. That makes cosine undefined. A robust implementation should reject or define a policy for zero vectors instead of silently dividing by zero.
If an embedding model unexpectedly emits many near-zero vectors, that is also useful diagnostic evidence.
Ranking depends on the score contract
Suppose:
query = [1,1]
A = [1,1]
B = [10,9]
Now compute both scores:
dot(query, A) = 2
dot(query, B) = 19
cos(query, A) = 1.000
cos(query, B) ≈ 0.999
Raw dot product ranks B first because B is much larger. Cosine ranks A first because A points in exactly the same direction as the query. The difference is small in cosine space but enough to reverse the order. Neither score is “mathematically wrong.” The retrieval model and index must use the score contract they were designed for.
Ranking is always relative to the candidate set
A cosine score of 0.82 is not universally “good.” Its meaning depends on the embedding model, corpus, query, and competing candidates. If every irrelevant chunk scores around 0.80, then 0.82 may be weak evidence. If most irrelevant chunks score below 0.20, it may be unusually strong. This is why retrieval systems should evaluate ranking metrics on labeled queries instead of relying on one global similarity threshold.
Normalization policy must match index and query
If stored vectors are normalized but query vectors are not, a system that expects normalized dot product can produce inconsistent scores. The failure may not raise an exception; all dimensions still match. Document whether normalization happens before storage, at query time, inside the index implementation, or not at all. Shape compatibility and score-contract compatibility are separate invariants.
Predict
Complete the Lab similarity function
- Click Run. The
dotlines already work (dot q,a: 2.0anddot q,b: 20.0), but everycosline shows0.0and checks fail. - Complete the TODO in
cosine: divide the dot product by the product of the two vector lengths (norms). If either length is zero, raise aValueErrorinstead of dividing by zero. - Click Run again. You should see
cos q,a: 1.0,cos q,b: 1.0, andcos q,c: 0.0. - Compare
a = [1.0, 1.0]withb = [10.0, 10.0]. They point in the same direction, butbis ten times longer. The dot product grows tenfold; cosine stays1.0. Cosine measures direction only. - Test your error policy. Add a line at the end:
print("cos zero:", cosine([0.0, 0.0], q)). Run it. The Lab should stop with yourValueError, a visible error instead of a silent0or a crash deep inside ranking.
Loading lab…
Quick Check
Explain it back
Compute dot product and cosine for two two-dimensional vectors. Explain one case where the rankings could differ and why a retrieval system must document its score contract.
Key Takeaways
- Dot product includes magnitude and alignment.
- Cosine normalizes vector magnitude.
- Unit-normalized dot product equals cosine.
- Zero vectors need an explicit policy.
- Index and query scoring must use the same contract.
Next Lesson
Next, store chunk vectors in a small index and return top-k chunk IDs with their scores.
References
- Johnson et al., Billion-scale similarity search with GPUs.
Completion is stored locally on this device.