topic: RAG Chunking, Step by Step
Blog 3 — Recursive Chunking: Preserve Natural Text Boundaries
Fixed-size chunking uses one main rule:
Table of contents
Fixed-size chunking uses one main rule:
Split the text when it reaches the size limit.
Recursive chunking keeps the size limit but adds a hierarchy of preferred boundaries.
The word recursive means that the splitting process is repeated using smaller text units until every piece fits within the required size.
A simplified hierarchy looks like this:
Try paragraph boundaries
↓
If a paragraph is too large,
try smaller boundaries
↓
Try line or sentence boundaries
↓
If the result is still too large,
try word boundaries
↓
Use characters as the final fallback
The exact hierarchy depends on the implementation. For example, LangChain’s default recursive character splitter uses blank lines, line breaks, spaces, and finally individual characters. Sentence punctuation can also be added to the separator list.
Preserving larger text units
Consider this paragraph:
My name is Andy. I work as a data engineer.
I build data pipelines with Python and SQL.
My favourite food is beef noodle soup.
A fixed-size splitter may stop in the middle of a sentence because the size limit has been reached.
A recursive splitter first checks whether the paragraph can remain together. If it is too large, the splitter moves to a smaller separator.
A possible result is:
Chunk 1:
My name is Andy.
I work as a data engineer.
Chunk 2:
I build data pipelines with Python and SQL.
My favourite food is beef noodle soup.
The chunks do not need to have exactly the same length. They only need to stay below the size limit while preserving natural text boundaries where possible.
Chunk size still matters
Recursive chunking does not remove the chunk-size setting.
A very small size limit will still force the splitter to create small pieces. If one sentence is longer than the limit, the splitter may eventually divide it by words or characters.
Recursive chunking therefore reduces boundary damage, but it cannot guarantee that every sentence or idea remains complete.
Overlap can also be added:
RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50
)
The size and overlap values still need to be tested against the real documents and questions.
Structure is not the same as meaning
Recursive chunking uses visible text structure, such as paragraphs and line breaks.
It does not fully understand the topic.
For example, one paragraph may contain these sentences:
The company opened a new office in Toronto.
The office will support customers across Canada.
Beef noodle soup is my favourite food.
I usually order it with spicy broth.
The paragraph contains two topics:
Topic 1: company office
Topic 2: food
A recursive splitter may keep them together because they appear inside the same paragraph and fit within the size limit.
This limitation leads to semantic chunking, which uses meaning rather than formatting to locate boundaries.
Design decision
Compared with fixed-size chunking, recursive chunking changes the boundary rule.
Fixed-size:
cut mainly by length
Recursive:
use length
+
prefer natural separators
Recursive chunking requires slightly more text processing, but it does not need an embedding model to choose boundaries. It therefore remains relatively fast while preserving more meaning than a hard fixed-size split.
Its design position is:
Speed: high
Meaning preservation: better than fixed-size
Complexity: low
The next strategy moves further toward meaning, but requires additional processing during indexing.
References
- LangChain, Splitting Recursively:
RecursiveCharacterTextSplitteruses a hierarchy of separators and supports configurable chunk size and overlap. (Docs by LangChain) - LangChain recommends recursive splitting as a starting point for many generic text use cases because it balances context preservation and size control. (Docs by LangChain)
- freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course. Watch the YouTube course. (freeCodeCamp)