topic: RAG Chunking, Step by Step
Blog 4 — Semantic Chunking: Split When Meaning Changes
Recursive chunking uses document structure to create better boundaries.
Table of contents
Recursive chunking uses document structure to create better boundaries.
Semantic chunking goes one step further. It attempts to place boundaries where the meaning changes.
The word semantic means related to meaning.
Consider these sentences:
I work as a data engineer.
I build data pipelines.
Most of my work uses Python and SQL.
My favourite food is beef noodle soup.
I usually order it with spicy broth.
The first three sentences discuss work. The last two discuss food.
Even without a heading or empty line, a person can recognise the topic change. Semantic chunking tries to detect the same change automatically.
Using embeddings to compare sentences
An embedding converts text into a numerical representation.
For a beginner, it is useful to imagine an embedding as a coordinate that represents the general meaning of a sentence.
Sentences about similar topics should have nearby coordinates:
"I build data pipelines."
"Most of my work uses Python and SQL."
Similarity: high
Sentences about different topics should be farther apart:
"Most of my work uses Python and SQL."
"My favourite food is beef noodle soup."
Similarity: lower
Semantic chunking uses this difference to locate possible boundaries.
A simplified process
A common semantic chunking pipeline contains four steps:
1. Split the document into sentences
2. Create embeddings for sentences
or small groups of sentences
3. Compare neighbouring embeddings
4. Create a boundary when
the semantic difference becomes large
The system normally uses a threshold. A threshold is a chosen cutoff value.
When the difference between neighbouring sentence groups passes the threshold, the splitter creates a new chunk.
Work sentences
↓
high similarity
↓
large similarity drop
↓
Food sentences
Advantages
Semantic chunking is useful when:
- paragraphs contain several topics
- natural formatting is missing
- sections are very long
- topic coherence is important for retrieval
A coherent chunk contains information that belongs to one clear topic.
This can help the embedding represent the chunk more accurately because unrelated topics are less likely to be mixed together.
Costs and limitations
Semantic chunking performs additional work during indexing.
It may need to:
- split text into sentences
- create many temporary embeddings
- compare neighbouring sentence groups
- select a threshold
The resulting chunks may also have very different sizes. A long section that stays on one topic can become one large chunk, while a document with frequent topic changes can produce many small chunks.
The quality also depends on the embedding model. A general embedding model may not recognise specialised differences in areas such as law, medicine, or engineering.
Semantic chunking therefore needs both semantic rules and size controls.
Design decision
Compared with recursive chunking, semantic chunking changes the source of the boundary signal.
Recursive chunking:
visible structure decides where to split
Semantic chunking:
changes in meaning help decide where to split
This moves the design further from speed and closer to meaning preservation.
Speed: lower during indexing
Meaning preservation: potentially higher
Complexity: medium
Semantic chunking is most valuable when evaluation shows that structural boundaries are mixing different topics. It should not be added only because it sounds more advanced.
The next method addresses a different problem. Instead of improving where the boundary is placed, late chunking improves how each chunk receives context from the surrounding document.
References
- LlamaIndex,
SemanticSplitterNodeParser: an official implementation that groups semantically related sentences and uses an embedding model to evaluate similarity between sentence groups. (GitHub) - freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course. Watch the YouTube course. (freeCodeCamp)