Why Retrieval-Augmented Generation
Goal
Explain RAG as supplying selected external evidence to a fixed generator, distinguish retrieval failures from generation failures, and decide when retrieval is a better lever than weight adaptation.
A language model's parameters contain learned patterns from training, but they are not a live database. If a user asks about:
- a policy published this morning;
- a private company handbook;
- a customer's authorized account record;
- a document too specific to rely on model memory;
the workflow needs a way to supply that information. RAG adds an evidence-selection step before generation.
The core split is retrieval versus generation
A RAG request first searches for useful evidence, places the selected evidence into context, and then generates an answer. The same final symptom can come from different earlier mistakes.
RAG: see where a bad answer can start
Switch between three traces. A request can continue after an earlier failure, so follow both the first broken step and its downstream effects.
- 1QuestionThe user asks for information.
✓ Request received - 2RetrieveSearch eligible sources for relevant chunks.
✓ Needed evidence found - 3EvidenceKeep the selected chunks and their source IDs.
✓ Needed evidence selected - 4ContextPut the selected evidence into the model's request.
✓ Good evidence in context - 5GenerateCreate an answer using the supplied context.
✓ Answer uses the evidence - 6Answer + citationsReturn the answer and show which evidence supports it.
✓ Supported answer returned
Use the three trace buttons to compare a successful request with two failure cases. Retrieval failure means the needed evidence never reaches the model. Generation/grounding failure means the right evidence is present, but the answer ignores, distorts, or overstates it. If you only inspect the final answer, these failures can look similar.
If you log retrieved chunk IDs, they are much easier to separate.
RAG is useful for changing knowledge
Suppose a support policy changes every week. Fine-tuning every weekly revision would create repeated training, versioning, and rollback work. With RAG:
policy document updated
→ index updated
→ next request can retrieve new evidence
The generator checkpoint can stay fixed. That does not make RAG automatically fresh; the index/update pipeline can still be stale. But freshness becomes a data-system problem rather than a model-weight update problem.
RAG does not guarantee factual answers
Suppose the correct chunk says:
Warranty is 18 months.
The model can still produce:
Warranty is 24 months.
Retrieval made the evidence available. It did not force the generator to use it correctly. A strong RAG workflow therefore needs both:
- retrieval evaluation;
- grounded answer/citation evaluation.
RAG also does not replace fine-tuning
Fine-tuning may be useful for stable behavior or task conventions. RAG may be useful for changing/private evidence. They can be combined. For example:
fine-tuned model: follows the company's support-record format
RAG: supplies the current product policy
The correct intervention depends on what is missing: behavior, evidence, or a deterministic rule.
Define the retrieval contract
For one request, record:
- query;
- eligible document set;
- retrieval method/version;
- retrieved chunk IDs;
- scores;
- top-k or reranking policy;
- generation prompt version.
This makes a RAG answer traceable.
RAG changes where evidence comes from, not the model's training objective
A useful RAG trace separates three objects:
question
→ retrieved evidence
→ generated answer
The generator can still produce unsupported text even when retrieval found the right source. Retrieval can also fail while the generator writes a fluent answer from its learned parameters. That means a wrong answer does not tell you which stage failed.
A diagnostic run should record the retrieved source IDs and inspect them before judging the final answer. If the correct evidence never arrived, changing prompt wording is downstream of the real problem. Conversely, if the right chunk is present but the answer invents a different value, retrieval succeeded and grounding/generation needs attention. This stage separation is the central engineering reason to evaluate RAG as a system rather than as one opaque model call.
Predict
Run the Lab
The browser Lab contains several failure traces.
Each trace records two facts: whether the relevant source was retrieved, and whether the answer was supported by the context that was supplied.
- Click Run. Read
retrieval_miss -> retrieval,generation_miss -> generation_or_grounding, andsuccess -> none. - For each trace, point to the field that decided the answer.
earliest_failurechecks retrieval first, because generation cannot use evidence it never received. - Change the
retrieval_misstrace so the right source was retrieved: set its"relevant_source_retrieved"toTrue, but leave"answer_supported_by_supplied_context": False. - Before running, predict the new label.
- Click Run. The trace now reports
generation_or_grounding. The evidence arrived, so the investigation moves to how the answer used it. - Press Reset afterward.
Loading lab…
Quick Check
Explain it back
Describe one case where RAG is preferable to fine-tuning and one case where fine-tuning may still be useful. Then explain how you would tell retrieval failure from generation failure.
Key Takeaways
- RAG supplies external evidence without requiring a model-weight update.
- Retrieval and generation are separate failure boundaries.
- RAG is useful for changing, private, or specific knowledge.
- Correct retrieval does not guarantee a correct answer.
- Trace query, chunk IDs, scores, and prompt identity.
Next Lesson
Next, prepare source documents so retrieval can operate on chunks while preserving provenance and access metadata.
References
Completion is stored locally on this device.