RAG Evaluation
Goal
Evaluate retrieval and answer generation separately, compute simple retrieval metrics, and diagnose whether a RAG regression begins before or after the needed source material reaches the generator.
Think of a relay team where one runner finds the correct document, another runner passes the relevant passage into the prompt, and a final runner produces the answer and citation. If the finish is wrong, one total score cannot tell you which exchange failed.
RAG evaluation therefore measures stages separately. Retrieval metrics ask whether the needed chunk appeared and how high it ranked. Grounding and answer checks ask whether the generator used the supplied context correctly. Citation checks ask whether claims point back to supplied sources that actually support them. Stage metrics turn one vague failure into a location you can investigate.
retrieval
→ grounded generation
→ citation/support
Retrieval recall@k
Suppose each evaluation query has one known relevant chunk. For five queries, top-3 retrieval contains the relevant chunk for four.
recall@3 = 4 / 5 = 0.8
This asks:
Did the candidate set include the needed chunk?
It does not care whether the generator later used that chunk correctly.
Reciprocal rank rewards earlier retrieval
If the relevant chunk appears at rank:
rank 1 → reciprocal rank 1.0
rank 2 → 0.5
rank 5 → 0.2
missing → 0
Mean reciprocal rank averages this value across queries. MRR distinguishes a relevant result consistently appearing first from one that appears near the bottom of the candidate list. It is most straightforward when the evaluation definition identifies a relevant item or first relevant rank clearly.
Answer quality needs separate checks
For grounded generation, possible checks include:
- exact task answer;
- answer supported by supplied sources;
- abstention when the allowed corpus has no answer;
- citation identity validity;
- citation support;
- format/schema validity.
A system can have high retrieval recall and low answer groundedness. Or retrieval recall can be low while the generator occasionally answers correctly from model memory. The latter should not be counted as retrieval success.
Use controlled ablations
Useful comparisons include:
generator without retrieval
generator with top-1
generator with top-5
hybrid retrieval
hybrid + reranker
Keep the evaluation set and generator settings fixed. Then differences are easier to attribute. If top-5 hurts grounded accuracy while recall improves, distractor context may be harming generation.
Evaluate no-answer-in-corpus cases
Include queries whose answer is absent from the authorized corpus. The desired behavior may be abstention. If the system answers confidently anyway, retrieval recall is not the right metric for that case. You need an unsupported-answer metric.
Report stage metrics together
A useful summary:
retrieval recall@5: 91%
MRR: 0.78
answer correctness: 84%
grounded-support pass: 80%
citation identity pass: 99%
no-answer abstention: 88%
Now the bottleneck is visible.
Build a gold set that diagnoses stages
A useful RAG evaluation case can record:
query
relevant source IDs
expected answer behavior
whether abstention is correct
required authorization context
That lets one case test more than final answer wording. If the relevant chunk is missing from top-k, retrieval recall failed. If the chunk is present but uncited, attribution failed. If the allowed corpus has no answer and the system abstains, that can be a successful result.
Protect the evaluation set from tuning leakage
If every prompt, chunking, and retrieval change is repeatedly chosen using the same small test set, the system can overfit that evaluation. Keep a development set for iteration and a held-out regression set for later confirmation. The principle is the same as earlier ML Levels: a test set becomes less independent when it repeatedly guides design choices.
Evaluate the pipeline in layers so one score does not hide the bottleneck
A useful RAG evaluation table has at least two kinds of labels: which chunks are relevant for the query and whether the final answer is supported and useful. Retrieval metrics such as recall@k can then be computed even if generation is temporarily disabled. Answer metrics can be computed on a fixed supplied-context set even if the retriever is being changed elsewhere.
This decomposition makes experiments interpretable. If recall@5 rises but supported-answer accuracy does not, the added candidates may be noisy or the generator may not use them well. If answer support improves with a fixed context set, the gain belongs to prompt/generation logic rather than retrieval. Keep case-level traces so aggregate averages cannot hide a critical subgroup such as no-answer abstention.
Predict
Run the local evaluator
Run the fixed evaluation set:
python labs/notebooks/level-09/l09-11-rag-evaluation.py
Record retrieval metrics and answer metrics separately. One case is intentionally correct even though its relevant chunk never reached the top-k context.
Now change only retrieval rank. Before running, predict which metric family should move:
python labs/notebooks/level-09/l09-11-rag-evaluation.py --move-q2-lower
Then change only answer support:
python labs/notebooks/level-09/l09-11-rag-evaluation.py --break-q1-support
The first experiment changes reciprocal rank without changing answer quality. The second changes grounded support without changing retrieval.
Loading lab…
Quick Check
Explain it back
Design four RAG metrics: two retrieval-side and two generation/support-side. Explain a failure pattern that each metric would expose.
Key Takeaways
- Evaluate retrieval separately from generation.
- Recall@k measures whether the relevant chunk reached the candidate set.
- Reciprocal rank measures how early the first relevant result appears.
- Correct answers from model memory are not retrieval successes.
- No-answer-in-corpus cases are required to test abstention.
- Stage metrics make bottlenecks visible.
Next Lesson
Next, turn common RAG failures into a boundary-by-boundary diagnostic procedure.
References
- Thakur et al., BEIR.
- Lewis et al., Retrieval-Augmented Generation.
Completion is stored locally on this device.