Documents, Chunks, and Metadata
Goal
Split documents into retrieval chunks while preserving source identity, boundaries, ordering, freshness, and access metadata.
Retrieval systems rarely compare a query against an entire book or handbook as one indivisible object. They usually split sources into chunks. A chunk is a retrieval-sized unit that still points back to its source.
Chunking changes what can be retrieved
Suppose a document contains:
Paragraph 1: product dimensions
Paragraph 2: battery duration
Paragraph 3: warranty
Paragraph 4: return policy
If the entire document is one chunk, retrieval can return everything at once. If every sentence is a separate chunk, retrieval is more precise but related context may be separated. Chunk size trades off:
- precision;
- context completeness;
- number of index items;
- retrieval noise;
- prompt token cost.
There is no universally correct chunk size.
Boundaries can destroy meaning
Consider:
The exception applies only when...
[chunk boundary]
...the device was purchased before June.
A naive fixed-length split can separate a condition from the statement it qualifies. Possible strategies include:
- paragraph-aware splitting;
- heading-aware splitting;
- sentence grouping;
- overlapping windows.
Overlap can preserve boundary context, but it duplicates text and can cause near-duplicate search results.
Metadata is part of the retrieval object
A useful chunk record might contain:
{
"chunk_id": "policy-7#03",
"source_id": "policy-7",
"title": "Returns Policy",
"section": "Exceptions",
"updated_at": "2026-09-20",
"access_group": "support",
"text": "..."
}
The text supports semantic/keyword matching. Metadata supports:
- source attribution;
- freshness filtering;
- authorization;
- deduplication;
- ordering;
- debugging.
Do not force every control into the embedding vector.
IDs must survive re-indexing
If chunk IDs change every time the index is rebuilt, old citations and evaluation records become difficult to interpret. Use a stable chunk identity policy where practical. For example:
source_id + section_id + chunk_ordinal + content/version hash
The exact scheme varies, but it should let you answer:
Which source content produced this retrieved item?
Chunking belongs in evaluation
A retrieval miss can be caused by:
- a weak embedding;
- bad similarity;
- a query mismatch;
- or a bad chunk boundary.
If the answer spans two chunks, top-1 retrieval may return only half the needed evidence. Include evaluation cases that test:
- facts near boundaries;
- multi-chunk answers;
- short exact identifiers;
- long sections with one relevant sentence.
Chunk identity needs lineage
A chunk should remain traceable to its parent document and source version. For example:
document_id: policy-17
document_version: 2026-09-01
chunk_id: policy-17:v3:chunk-04
start_offset: 1200
end_offset: 1680
If the source changes and the index is rebuilt, old and new chunks should not silently share an identity unless they really represent the same source state. This matters for citations, deletion, freshness checks, and debugging. A citation to chunk-04 is weak if nobody can tell which document version produced that chunk.
Chunk size is a retrieval contract, not only preprocessing
Large chunks preserve context but can mix several topics into one retrieval unit. Small chunks sharpen topic focus but can split a fact from the condition that gives it meaning. Overlap can reduce boundary loss, but it also creates near-duplicate candidates.
So record the chunker version and parameters with evaluation results. If chunking changes, the retrieval corpus effectively changed too.
Predict
Build chunks in the Lab
This Lab contains a small sectioned document.
This Lab contains one small sectioned document, policy-7, with three sections.
- Click Run once. Nothing is printed, because
make_chunksreturns an empty list, and the checks fail. - Complete the TODO: return one chunk per section. Each chunk needs a stable
chunk_idbuilt from the source ID and a two-digit ordinal (policy-7#00,policy-7#01, ...), plus thesource_id,section,updated_at,access_group, andtext. - Click Run again. You should see three chunk records, and every check should pass.
- Add a fourth section at the end of
sections:("exceptions-2", "Gift cards are final sale."),. Before running, predict itschunk_id. - Click Run. The new chunk is
policy-7#03, and the first three IDs did not move. Appending keeps old IDs stable; inserting a section in the middle would shift every later ordinal. The Lab should reportResult: experiment ranbecause the original three-chunk baseline changed while the chunking and metadata invariants still pass. - Notice a boundary problem: the rule about safety equipment lives in
exceptions, far from the 30-day rule ineligibility. Write one evaluation question—such as “Can I return an opened helmet after 10 days?”—that needs both chunks. That kind of case reveals a bad chunk boundary.
Loading lab…
Quick Check
Explain it back
Design a chunk record for a policy document. Include at least five metadata fields and explain which are used for citation, freshness, and access control.
Key Takeaways
- Chunking determines the retrieval unit.
- Boundary choices can preserve or destroy useful context.
- Overlap trades context preservation for duplication.
- Metadata carries provenance, freshness, and access information.
- Chunking strategy should be evaluated, not assumed.
Next Lesson
Next, represent queries and chunks as vectors for semantic retrieval.
References
- Lewis et al., Retrieval-Augmented Generation.
Completion is stored locally on this device.