topic: RAG Chunking, Step by Step
Blog 2 — Fixed-Size Chunking: Fast but Blind to Meaning
Fixed-size chunking is the simplest way to divide a document.
Table of contents
Fixed-size chunking is the simplest way to divide a document.
The splitter receives a size limit, such as:
500 characters
or:
200 tokens
It then moves through the document and creates chunks around that size.
Document
↓
First 500 characters → Chunk 1
Next 500 characters → Chunk 2
Next 500 characters → Chunk 3
A character is a letter, number, punctuation mark, or space. A token is a unit used by a language model and may contain one word or only part of a word.
Strengths of fixed-size chunking
Fixed-size chunking requires very little processing.
It is usually:
- fast
- easy to implement
- predictable
- easy to debug
It is also deterministic, which means the same document and settings produce the same chunks every time.
These properties make fixed-size chunking a useful baseline. A baseline is a simple starting solution used to compare later improvements.
The boundary problem
Fixed-size chunking mainly pays attention to length. In its simplest form, it does not understand sentences, topics, or document structure.
Consider the following text:
My name is Andy.
I work as a data engineer.
My favourite food is beef noodle soup.
A fixed boundary may produce:
Chunk 1:
My name is Andy.
I work as a data engineer.
My favourite food is
Chunk 2:
beef noodle soup.
The phrase “beef noodle soup” remains complete, but its relationship with Andy has been separated.
The situation becomes worse when a hard character boundary cuts through a word:
Chunk 1:
My name is An
Chunk 2:
dy. I work as a data engineer.
Not every fixed-size implementation makes such a hard cut. Some libraries use separators or tokenizers to reduce broken words. However, a size limit alone cannot guarantee that a complete idea remains in one chunk.
Adding overlap
Overlap is often combined with fixed-size chunking.
For example, a 500-token chunk might repeat the final 50 tokens at the beginning of the next chunk.
Chunk 1: tokens 0–500
Chunk 2: tokens 450–950
Chunk 3: tokens 900–1400
This gives information near the boundary another chance to appear in a complete chunk.
Overlap reduces some boundary problems, but it does not remove them. An important explanation may still be longer than the overlapping area.
Large overlap also creates disadvantages:
More overlap
→ more duplicated text
→ more embeddings
→ more storage
→ more repeated search results
Suitable use cases
Fixed-size chunking can work well when:
- the documents have a uniform format
- each record already contains complete information
- indexing speed is important
- a simple comparison baseline is needed
It is less suitable when sentences, paragraphs, or sections must remain together.
Design decision
Compared with treating a whole document as one retrieval unit, fixed-size chunking improves speed, predictability, and retrieval focus.
Its main weakness is meaning preservation.
Fixed-size chunking
Main priority:
speed and simplicity
Main risk:
arbitrary boundaries
The next strategy keeps the size limit but changes how boundaries are selected. Instead of cutting immediately when the limit is reached, recursive chunking first looks for a more natural place to split.
References
- LangChain, Text Splitter Integrations: length-based chunking can use character counts or token counts to control chunk size. (Docs by LangChain)
- LangChain, Splitting by Character: character-based splitting uses a chosen separator and measures chunk length by character count. (Docs by LangChain)
- freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course. Watch the YouTube course. (YouTube)