Skip to main content
L9.12

Failure Modes in Retrieval

Goal

Debug a wrong RAG answer by following the retrieval pipeline in order and fixing the first step where the needed source, chunk, permission, query, or ranking result goes missing.

When a RAG answer is wrong, start with the source the answer needed and follow its journey. Do not tune the final prompt until you know where that source disappeared.

Use one running question:

What is the XR-8 warranty?

Step 1: did the right source enter the library?​

If the current XR-8 policy was never ingested, no retriever can return it. Check the source inventory, ingestion status, version or update time, and deletion state.

Step 2: did chunking keep the useful sentence together?​

The document may exist while the warranty sentence is split away from the product name or another sentence it depends on. Inspect the chunk text, neighboring chunks, metadata, and source-to-chunk mapping.

If the answer needs two adjacent chunks but the system or evaluation expects one, chunking may be the first broken step.

Step 3: was the chunk allowed for this request?​

A permission or freshness rule may remove the chunk before ranking. That can be correct behavior. A missing result is not a retrieval bug when the user is not allowed to see the source.

Record the decision in plain terms:

eligible before scoring? yes/no
reason: ...

Step 4: did the search query keep the user's meaning?​

Compare the original and rewritten queries. Check exact product codes, important dates, filters, and synonyms. A rewrite that drops XR-8 can make every later stage look healthy while searching for the wrong thing.

Step 5: did the first search stage rank the chunk high enough?​

Now inspect the scores and versions used by keyword, embedding, or hybrid search. Compare an exact-search reference with the approximate index when useful. This tells you whether the problem is the representation or the index behavior.

Step 6: did the candidate cutoff remove it?​

Suppose the useful chunk is ranked 18th, but only the first 10 candidates are sent forward:

useful rank: 18
candidate limit: 10

The reranker never receives the useful chunk. Changing the reranker cannot fix a candidate it never saw.

Step 7: did reranking make the order worse?​

If the chunk enters the candidate set at rank 3 but leaves the final top 5 after reranking, compare the before/after ranks and inspect the reranker. At this point the earlier stages delivered the right candidate, so the problem really did begin later.

The important debugging habit is simple: find the first stage where the useful source stops being available. Fixing that stage is usually more informative than changing the final answer prompt.

Then inspect generation​

Only after verifying the correct evidence reached the model-facing context should you focus on:

  • prompt construction;
  • trust/injection handling;
  • grounded answer logic;
  • citation support.

The first broken step usually gives the most direct fix.

One symptom can have several upstream causes​

“No relevant source was retrieved” can mean:

  • the source was never ingested;
  • chunking separated the needed evidence;
  • authorization correctly filtered it;
  • query rewriting changed intent;
  • the embedding representation ranked it poorly;
  • candidate k cut it off.

Those causes require different fixes. Increasing top-k helps only one of them.

Write a retrieval trace before tuning​

For a failed query, record:

source snapshot
→ chunk IDs
→ eligible chunk IDs
→ original/re-written query
→ first-stage ranking
→ candidate cutoff
→ reranked order
→ supplied prompt evidence
→ cited answer

Then locate the first stage where expected and observed evidence diverge. Changing a downstream prompt when the correct source never reached the prompt is the RAG equivalent of debugging a final neural-network layer before checking a broken input.

Fix the first broken step before tuning the final answer​

Retrieval failures often propagate forward. A source can be absent from ingestion, split into an unhelpful chunk, filtered by metadata, represented poorly, ranked below k, removed by reranking, or dropped from the final prompt. By the time the answer is wrong, several downstream components may look suspicious even though only one boundary first lost the evidence.

Use a trace that asks the same ordered questions for every failure: did the correct source exist, was the correct chunk created, was it eligible, was it retrieved, was it reranked into the supplied set, and did the generator use it correctly? Stop at the first “no.” That habit prevents prompt changes from masking indexing bugs and prevents index changes from being blamed for generation mistakes.

Predict

A relevant chunk ranks 18th and candidate k is 10. Which subsystem should you inspect before the reranker?

Run the local failure trace​

The script contains seven fixed failure cases and checks the same ordered boundaries: source presence, chunk creation, eligibility, query intent, candidate retrieval, top-k inclusion, and reranking.

  1. Before running anything, open the script and inspect only the top-k-cutoff fixture. Its first five states are True and inside_top_k is False. Predict the earliest broken step.
  2. Run:
python labs/notebooks/level-09/l09-12-failure-trace.py
  1. Compare your prediction with the printed top-k-cutoff line.
  2. For a controlled change, change only the sixth state in that fixture from False to True, leaving the seventh state False and the stored expected value untouched. Before rerunning, predict which later boundary will now be observed and why the script should report a trace mismatch.
  3. Rerun, observe the mismatch, then restore the starter. The exercise shows that fixing an upstream boundary can reveal the next downstream failure.

Loading lab…

Quick Check

1. The source document was never ingested. Can better reranking solve the query?
2. A relevant chunk is excluded because the user lacks permission. Is that necessarily a retrieval bug?
3. Why inspect the earliest broken step?

0 of 3 questions answered.

Explain it back​

Create a seven-step RAG failure trace from source ingestion to reranking. For one hypothetical wrong answer, identify the earliest failed boundary and one focused diagnostic.

Key Takeaways

  • Retrieval failure can begin before embedding search.
  • Corpus presence, chunking, eligibility, rewriting, ranking, k cutoff, and reranking are separate boundaries.
  • Authorization exclusions may be correct behavior.
  • Candidate misses cannot be repaired by downstream reranking.
  • Debug the earliest mismatch before changing later stages.

Next Lesson

Next, make freshness, deletion, and access control explicit parts of the retrieval contract.

References

  • Thakur et al., BEIR.
Lesson actions

Completion is stored locally on this device.

View progress