Embeddings for Retrieval
Goal
Explain retrieval embeddings as vector representations used to compare queries with chunks, distinguish retrieval vectors from token IDs and generation hidden states, and diagnose representation mismatch.
Imagine a library that wants related questions and passages to be easy to compare even when they use different words. Instead of comparing the raw sentences directly, a retrieval model turns each text into a list of numbers. You can picture that list as a machine-made location in a many-dimensional map: texts the retriever considers useful for similar information needs tend to receive representations that a similarity rule can score as closer.
The coordinates themselves are not human-readable meanings, and "nearby" does not prove that a passage contains the answer. The vector is a retrieval representation—a tool for ranking candidates. This is why an embedding can help connect "How long does the battery last?" with "Typical runtime is ten hours" while still needing later checks for exact support, freshness, and authorization.
A semantic retriever needs a way to compare:
- a user query;
- a document chunk.
An embedding model maps each text into a fixed-width vector.
query text → query vector
chunk text → chunk vector
Then a similarity function compares those vectors.
Embeddings are not token IDs
Token IDs are discrete lookup addresses such as:
[17, 42, 9]
A retrieval embedding might be:
[0.21, -0.08, 0.44, ...]
The embedding summarizes a whole query or chunk into a vector intended for retrieval comparison. It is not simply “the token IDs converted to decimals.”
Query and document encoders must be compatible
Some retrieval models use the same encoder for queries and documents. Others use separate query/document encoders trained to produce compatible spaces. The invariant is:
The vectors must be comparable under the intended retrieval model contract.
If chunks are indexed with embedding model v1 and queries are embedded with incompatible model v2, dimensions may differ or, worse, dimensions may match while semantic geometry differs. Store embedding model/version with the index.
Semantic retrieval can bridge wording differences
Query:
How long does the battery last?
Chunk:
Typical runtime is approximately ten hours.
Keyword overlap may be small. A useful retrieval embedding can place these texts near each other because their meanings are related. But semantic embeddings are not perfect. They can struggle with:
- exact product codes;
- numbers;
- rare names;
- domain-specific terminology;
- negation or subtle distinctions.
That motivates hybrid search later.
Retrieval embeddings are task-dependent
An embedding model good for general web text may perform poorly on:
- source code;
- legal citations;
- medical abbreviations;
- multilingual support data.
Evaluate with your actual queries and chunks. Do not select an embedding model only from a generic benchmark name.
Inspect nearest neighbors
A useful diagnostic is not only recall@k. For a small query set, inspect:
- top retrieved chunk;
- score;
- next few alternatives;
- metadata;
- obvious false positives.
If a query for “battery duration” retrieves chunks about “battery replacement policy,” the representation may be collapsing nearby but distinct concepts.
Retrieval embeddings are an index contract
An index built with embedding model A should not be queried with vectors from an unrelated embedding model B. Even when both models produce vectors with the same dimension, their coordinate systems may represent text differently. The shapes can match while the semantics do not. Record at least:
embedding model/version
dimension
normalization policy
text preprocessing
index build version
When the embedding model changes, rebuild and reevaluate the index rather than assuming old stored vectors remain compatible.
Retrieval similarity is task evidence, not semantic truth
Two chunks can be close in embedding space because they share topic or phrasing, yet only one contains the required answer. An embedding score says the representation considers the texts related under that model. It does not prove support, authority, freshness, or authorization. Those checks belong to later stages.
Retrieval embeddings are useful because geometry is tied to a search objective
An embedding model maps text to vectors so that a similarity function can compare a query with candidate passages. The coordinates are not human-readable facts, and a nearby point is not automatically “true.” The useful property is empirical: texts that should be retrieved for similar information needs tend to receive compatible representations under the model and similarity rule.
That is why embedding quality must be evaluated on retrieval cases from the target domain. A general-purpose model may place common semantic paraphrases close together yet struggle with product codes, legal clauses, or specialized terminology. Keep the embedding model ID and preprocessing settings with the index because changing either can move every vector and invalidate old distance assumptions.
Predict
Run the Lab
The Lab uses small hand-authored semantic feature vectors.
Each vector has three hand-made features. You can read them roughly as “how much is this about battery duration, battery replacement, and warranty.”
- Click Run. With
query = [0.9, 0.1, 0.0](a battery-duration question), the scores arebattery_duration 0.74,battery_replacement 0.58, andwarranty 0.02. - Move the
battery_durationchunk toward replacement: change"battery_duration": [0.8, 0.2, 0.0],to"battery_duration": [0.5, 0.5, 0.0],. - Before running, predict the new ranking.
- Click Run.
battery_replacement(0.58) now ranks abovebattery_duration(0.5). The nearest neighbor changed because the chunk's representation changed, not the text you would read. - Press Reset. Now think about the query
XR-417 battery. Nothing in these three features represents the product codeXR-417, so every chunk about batteries would look equally close. Explain why keyword search, which matches the exact stringXR-417, can help for that case.
Loading lab…
Quick Check
Explain it back
Describe the difference among token IDs, retrieval embeddings, and generation hidden states. Then name one domain where you would insist on domain-specific retrieval evaluation.
Key Takeaways
- Retrieval embeddings map queries/chunks into a comparable vector space.
- They are not token IDs.
- Query and document vector contracts must be compatible.
- Semantic retrieval can bridge wording differences but may miss exact lexical signals.
- Evaluate embeddings on real task queries and chunks.
Next Lesson
Next, calculate vector similarity directly and see why normalization changes dot-product behavior.
References
- Lewis et al., Retrieval-Augmented Generation.
- Johnson et al., Billion-scale similarity search with GPUs.
Completion is stored locally on this device.