topic: RAG Chunking, Step by Step
Blog 6 — Choosing a Chunking Strategy for Production RAG
A production chunking strategy should begin with the document, not with the most advanced algorithm.
Table of contents
A production chunking strategy should begin with the document, not with the most advanced algorithm.
Different documents organise meaning in different ways. A plain article, a Python file, and a legal contract should not automatically use the same boundaries.
The general production process has three stages:
1. Preserve useful document structure
2. Start with a simple chunking strategy
3. Add complexity only when evaluation
shows a specific retrieval problem
Preserve document structure first
Documents often contain natural boundaries that already express meaning.
Markdown
Markdown uses headings:
# Main Topic
## Installation
## Configuration
These headings identify sections. A Markdown-aware splitter can create chunks under each heading and store the heading as metadata.
Metadata is additional information attached to a chunk, such as:
document name
page number
section title
source
Code
Source code contains functions and classes:
class Customer:
...
def calculate_total():
...
A code-aware splitter can prefer boundaries before functions or classes instead of cutting through the middle of them.
Legal and technical documents
Legal and technical documents often contain:
chapter
section
subsection
definition
clause
table
These elements should be preserved when they carry important meaning. If a section is still too large, a general splitter can divide it further.
Document-aware splitting and recursive splitting can therefore be used together:
Split by document structure
↓
Section is still too large
↓
Apply recursive splitting
A practical decision flow
The following process works as a general starting point.
START
|
v
Does the document have useful structure?
|
+-- Yes
| |
| v
| Preserve headings, functions,
| sections, or other structures
| |
| v
| Apply recursive splitting
| to oversized sections
|
+-- No
|
v
Start with recursive chunking
|
v
Evaluate retrieval results
|
v
Are unrelated topics mixed together?
|
+-- Yes --> Test semantic chunking
|
+-- No --> Keep the simpler strategy
A separate test should check for missing document context:
Do chunks contain words such as
"he," "it," "this result," or "the company"
without enough surrounding information?
When this problem appears frequently, late chunking may improve the chunk embeddings.
Relative speed and meaning
The following table shows general tendencies. It is not a universal performance benchmark.
| Strategy | Main priority | Relative indexing speed | Meaning preservation |
|---|---|---|---|
| Fixed-size | Simplicity and predictability | Fastest | Low when boundaries are poor |
| Recursive | Natural text boundaries | Fast | Good for general text |
| Semantic | Topic coherence | Slower | Better for mixed topics |
| Late chunking | Wider document context | Depends on model and document length | Better for cross-chunk references |
| Structure-aware | Document organisation | Usually fast to medium | Strong when structure is reliable |
These methods are not always competitors.
For example, one pipeline can use:
Markdown headings
+
recursive size control
+
late chunk embeddings
Each method handles a different design dimension.
Evaluate with real questions
A production decision should be based on representative user questions.
For each strategy, check:
Retrieval quality
Did the retriever return the text that contains the correct evidence?
Answer quality
Did the language model receive enough information to answer correctly?
Indexing time
How long did document processing and embedding take?
Storage
How many chunks and duplicated overlapping passages were stored?
Latency
Latency means the delay between a user request and the system response.
A more complex chunking strategy is useful only when the quality improvement justifies its additional processing or maintenance cost.
Production starting point
A practical starting policy is:
General prose
→ start with recursive chunking
Markdown, HTML, code, or structured documents
→ preserve structure first
Long text with mixed topics
→ test semantic chunking
Chunks that lose surrounding document context
→ test late chunking
For every strategy
→ evaluate speed and answer quality
Fixed-size chunking remains useful as a baseline. Recursive chunking is a strong starting point for general text because it provides better boundaries without requiring semantic analysis.
Semantic and late chunking should be introduced for specific, measured problems rather than applied automatically.
Design decision
The progression across this series can be summarised as follows:
Fixed-size
→ controls size
Recursive
→ improves structural boundaries
Semantic
→ improves topic boundaries
Late chunking
→ improves context inside chunk embeddings
Structure-aware splitting
→ respects the document's own organisation
The most advanced method is not automatically the best production method.
The final goal is to choose the simplest strategy that retrieves complete and meaningful evidence with acceptable speed, cost, and complexity.
References
- LangChain, Text Splitter Integrations: recursive splitting is recommended as a starting point for many general-text use cases, while document-aware strategies can use structures such as Markdown, HTML, JSON, and code. (Docs by LangChain)
- LangChain, Splitting Code: language-aware separators can prefer boundaries such as Python classes and functions. (Docs by LangChain)
- LangChain, Splitting Markdown: Markdown headings can be used as boundaries and retained as chunk metadata; recursive splitting can then control oversized sections. (Docs by LangChain)
- Günther et al., Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models. (arXiv)
- freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course. The course covers the indexing pipeline, embeddings, vector databases, retrieval optimisation, contextual retrieval, and late chunking. Watch the YouTube course. (freeCodeCamp)