topic: RAG Chunking, Step by Step
Fixed Chunking
How fixed chunking uses chunk size and overlap to prepare documents for retrieval, with practical trade-offs and a LangChain example.
Table of contents
Introduction
Fixed chunking is a simple method for splitting a document into smaller chunks. It usually requires two main parameters:
- Chunk size: the maximum size of each chunk.
- Overlap size: the amount of content shared between consecutive chunks.
The size can be measured in characters or tokens, depending on the implementation.
How It Works
Suppose we set:
chunk_size = 500 tokens
overlap_size = 50 tokens
The document is divided into chunks of approximately 500 tokens. Each new chunk also includes the last 50 tokens from the previous chunk.
For example:
Chunk 1: Token 1 ───────────── Token 500
Chunk 2: Token 451 ───────────── Token 950
Chunk 3: Token 901 ───────────── Token 1400
The overlap helps preserve context that might otherwise be lost at the boundary between two chunks.
By splitting a large document into smaller pieces, the retrieval system can later search for and return only the chunks that are relevant to a user's question instead of sending the entire document to the LLM.
How to Choose the Chunk Size
There is no universally optimal chunk size. The best value depends on the document structure, embedding model, expected queries, and retrieval task.
As a practical starting point, chunks of around 256–1024 tokens are reasonable values to test, with approximately 500 tokens being a useful initial baseline.
Smaller chunks usually provide more granular retrieval because each chunk represents a narrower piece of information. However, if chunks are too small, they may lose important surrounding context.
Larger chunks preserve more context, but they may contain multiple topics or ideas. This can make the embedding less specific and potentially reduce retrieval precision.
Therefore, chunk size represents a trade-off:
Smaller chunks → more granular retrieval, less context
Larger chunks → more context, potentially less precise retrieval
In a production RAG system, the chunk size should ideally be selected through retrieval evaluation rather than relying only on a fixed rule of thumb.
How to Choose the Overlap
A reasonable starting point is around 10–20% of the chunk size, but there is no universally optimal overlap.
For example:
chunk_size = 500 tokens
overlap = 50–100 tokens