topic: RAG Chunking, Step by Step
Recursive Chunking
How recursive chunking preserves document structure through progressively smaller separators, with size constraints, overlap, and a Gemini token-counting example.
Table of contents
Introduction
Recursive Chunking is one of the most commonly used chunking strategies in RAG systems.
Similar to Fixed Chunking, we still need to define:
- Chunk size: the maximum size of each chunk.
- Overlap size: the amount of content shared between consecutive chunks.
The size can be measured in characters or tokens, depending on the implementation.
The main difference from Fixed Chunking is how the chunk boundaries are selected.
Fixed Chunking mainly splits text according to size. This means a chunk boundary may occur in the middle of a paragraph or even a sentence.
Recursive Chunking tries to avoid this problem by splitting the document using progressively smaller structural boundaries.
How It Works
Suppose we set:
chunk_size = 500 tokens
chunk_overlap = 50 tokens
Instead of immediately cutting the document every 500 tokens, Recursive Chunking first tries to split the document using larger and more meaningful separators.
Conceptually, we can think about a document as having several levels of structure:
Document
↓
Page
↓
Paragraph
↓
Sentence
↓
Clause
↓
Word
These structures may be represented by separators such as:
Page → \f
Paragraph → \n\n
Line → \n
Sentence → ". "
Clause → ", "
Word → " "
The exact separators depend on the document format and splitter configuration.
The key idea is that the splitter tries larger boundaries first.
For example:
1. Try to split by page
2. If a page is still too large, split it by paragraph
3. If a paragraph is still too large, split it by lines or sentences
4. If a sentence is still too large, split it by smaller separators
5. Continue until the content fits within the chunk size
This is why the method is called Recursive Chunking.
For example, imagine the following document:
Paragraph A: 180 tokens
Paragraph B: 210 tokens
Paragraph C: 170 tokens
With:
chunk_size = 500 tokens
the splitter may combine Paragraph A and Paragraph B:
Chunk 1
├── Paragraph A: 180 tokens
└── Paragraph B: 210 tokens
Total ≈ 390 tokens