Reranking
Goal
Separate fast candidate retrieval from a more expensive reranking step, explain why reranking cannot recover missing candidates, and measure recall before reranking quality.
Imagine a librarian facing 100,000 books. A fast catalog search can quickly produce 20 plausible candidates. The librarian can then read those 20 descriptions more carefully and move the best matches to the top. It would be far too expensive to do the careful reading for every book in the library.
Retrieval plus reranking uses the same two-stage strategy. The first stage searches broadly and cheaply enough to keep the relevant item in the candidate set. The reranker spends more computation comparing the query with only those candidates. This also explains a hard limit: if the first stage never includes the relevant chunk, the reranker has nothing to rescue.
query
→ fast retriever returns candidates
→ reranker examines those candidates more carefully
→ final top results
Reranking cannot recover a missing document
Suppose the relevant chunk is R. First-stage top-5:
[A, B, C, D, E]
R is absent. No reranker operating only on those five candidates can return R. This gives a critical evaluation order:
- Was relevant evidence present in the candidate set?
- If yes, did reranking place it high enough?
Do not blame the reranker for candidate-generation misses.
A reranker can use richer query-document interaction
A vector retriever might score query and chunk from separately computed embeddings. A reranker can inspect the query and candidate together, using richer features or a cross-encoder-style model. That often costs more per pair. If 100,000 chunks exist, scoring every chunk with an expensive reranker may be impractical. So:
cheap/broad search
→ expensive/narrow rerank
is a common architecture.
Candidate k creates a recall/cost trade-off
If candidate k is too small:
- relevant chunks may be missing.
If candidate k is very large:
- reranking costs more;
- latency rises;
- more low-quality candidates are processed.
Evaluate recall@candidate-k separately from final top-n quality. Example:
candidate recall@20 = 95%
final relevant-in-top3 = 82%
Those two numbers point to different improvement opportunities.
Reranking scores need provenance too
Record:
- first-stage score/rank;
- reranker score/rank;
- candidate-set version;
- reranker version.
A final ranking alone hides which stage changed. If a new reranker improves final top-3 but candidate recall remains 70%, the system still has an upstream bottleneck.
Use deterministic tie rules in tests
Like vector search, reranking tests should have deterministic ordering when scores tie. Otherwise small test fixtures can flicker even when the underlying behavior is unchanged.
Candidate recall sets the reranker's ceiling
Suppose the correct chunk ranks 40th in the first-stage retriever. If the reranker receives only the top 20 candidates, the best possible reranker still cannot return that chunk. This creates a useful diagnostic inequality:
final relevant result
requires
relevant result inside candidate set
Measure candidate recall at the chosen k before evaluating reranker quality.
Reranking can improve precision and still hurt a slice
A reranker may move generally relevant documents upward while demoting exact identifiers, short fragments, or a domain it was not trained for. That is why final ranking evaluation should include slices tied to real query types. Do not accept “average ranking improved” as proof that every retrieval boundary improved. Store both pre-rerank and post-rerank rankings for failed cases so you can tell whether the reranker helped or caused the miss.
Candidate retrieval and reranking solve different budget problems
The first-stage retriever must search a large corpus cheaply enough to return a manageable candidate set. A reranker can then spend more computation comparing the query with only those candidates. This architecture makes sense only if the relevant evidence survives the first stage; a reranker cannot rescue a document that was never retrieved.
Evaluate both boundaries. Measure candidate recall before reranking, then measure whether reranking improves the order of those candidates. If recall is already poor, increase retrieval coverage or fix the representation first. If recall is good but relevant evidence is buried, the reranker is the more plausible place to investigate.
Predict
Complete the reranking Lab
The first stage returned three candidates, ranked by first_stage score: c1 (0.92), c2 (0.78), c3 (0.74). The query asks about the Cedar contractor policy, which is c2.
- Click Run once. The order is unchanged,
['c1', 'c2', 'c3'], and the checks fail. - Complete the TODO in
rerank: sort by how many query words each candidate contains (most first), then byfirst_stagescore (highest first), then byid. - Click Run again. You should see
['c2', 'c3', 'c1']: the relevant chunk moved from second to first. - Remove the relevant chunk from the candidates: delete the whole line for
c2. - Click Run. The result is
['c3', 'c1']. A reranker can only reorder what it receives, so it cannot recover a chunk the first stage missed. - Restore
c2. Explain the trade-off of asking the first stage for more candidates: a larger candidate k makes it more likely the relevant chunk is included, but the reranker then has to score more chunks, which costs time.
Loading lab…
Quick Check
Explain it back
Draw a two-stage retrieval pipeline. State one metric for candidate generation and one metric for final reranked output, and explain why each diagnoses a different boundary.
Key Takeaways
- Candidate retrieval and reranking are separate stages.
- Rerankers cannot recover absent candidates.
- Candidate k trades recall against cost.
- Record both first-stage and reranked scores/ranks.
- Evaluate upstream recall before optimizing downstream ranking.
Next Lesson
Complete the mini checkpoint, then study query rewriting as another upstream retrieval intervention.
References
- Thakur et al., BEIR.
Completion is stored locally on this device.