topic: RAG Chunking, Step by Step
Blog 1 — Chunking Decides What RAG Can Retrieve
In a RAG system, a language model does not depend only on what it learned during training. Before generating an answer, the system searches external documents and sends the relevant information to the model.
Table of contents
RAG stands for Retrieval-Augmented Generation.
In a RAG system, a language model does not depend only on what it learned during training. Before generating an answer, the system searches external documents and sends the relevant information to the model.
The component that performs this search is called a retriever.
A retriever usually does not search an entire PDF as one item. The document is first divided into smaller pieces called chunks. Each chunk can then be stored, searched, and retrieved independently.
This document preparation process is called indexing. It normally happens before a user asks a question.
Document
↓
Create chunks
↓
Create embeddings
↓
Store them in a searchable index
↓
User asks a question
↓
Retriever returns relevant chunks
An embedding is a list of numbers that represents the meaning of a piece of text. Texts with similar meanings should have similar embeddings.
Chunking is therefore an indexing step, but it is also an important retrieval design decision. It defines the units that the retriever will later search.
The split-boundary problem
Consider this text:
My name is Andy.
I like to eat beef noodle soup.
A useful chunk keeps the related information together:
Chunk 1:
My name is Andy.
I like to eat beef noodle soup.
The question below can be answered from one chunk:
What food does Andy like?
Now consider a poor split:
Chunk 1:
My name is Andy.
I like to eat
Chunk 2:
beef noodle soup.
The first chunk identifies Andy but does not contain the food. The second chunk contains the food but does not identify the person.
The language model can only use the chunks returned by the retriever. If the retriever returns only one of these chunks, the model does not receive the complete evidence.
This is a split-boundary problem. A boundary is the location where one chunk ends and the next chunk begins.
Four chunking design dimensions
Four design dimensions are useful when comparing chunking strategies.
1. Chunk size
Chunk size controls how much text is placed in one chunk.
A small chunk is usually more focused, but it may lose important context. A large chunk contains more context, but it may also contain unrelated information.
2. Overlap
Overlap means repeating some text in neighbouring chunks.
Chunk 1:
My name is Andy.
I like to eat beef noodle
Chunk 2:
I like to eat beef noodle soup.
I usually order it on weekends.
Overlap can protect information near a boundary. However, it also creates duplicated text, additional embeddings, and more storage.
3. Split boundaries
A splitter can create boundaries based on:
- character or token count
- sentences
- paragraphs
- headings
- changes in meaning
A token is a small unit of text processed by a language model. A token may be a whole word, part of a word, or punctuation.
4. Document structure
Different documents contain different structures.
Markdown contains headings. Code contains functions and classes. Legal documents contain sections and clauses. These structures often show which information belongs together.
A good chunking strategy should not ignore them.
Design decision
Compared with storing an entire document as one searchable unit, chunking creates smaller and more focused retrieval units. This can improve search speed and relevance, but every new boundary creates a risk of losing meaning.
The first chunking trade-off is therefore simple:
Smaller chunks
→ faster and more focused retrieval
→ greater risk of losing context
Larger chunks
→ more context
→ more noise and more tokens
The next article introduces the simplest implementation of this decision: fixed-size chunking.
References
- LangChain, Retrieval: definitions of text splitters, embeddings, vector stores, and retrievers in a RAG pipeline. (Docs by LangChain)
- freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course, created by Paulo Dichone. The document-processing and indexing section begins at approximately 28:27. Watch the YouTube course. (freeCodeCamp)