topic: RAG Chunking, Step by Step
Semantic Chunking
How semantic chunking uses embedding distances to detect topic changes, with variable-sized chunks, ingestion trade-offs, and a LangChain example.
Table of contents
Introduction
Semantic Chunking is a more advanced chunking strategy used in RAG systems.
Unlike Fixed Chunking and Recursive Chunking, semantic chunking does not mainly determine chunk boundaries using a predefined chunk size.
The main idea is simple: if we want the content inside a chunk to be semantically related, why not first embed the sentences into vectors and compare their semantic similarity?
For example:
Sentence 1: FAISS is used for vector similarity search.
Sentence 2: It can efficiently retrieve similar embedding vectors.
Sentence 3: Toronto has cold winters.
After embedding the sentences, we may find:
Sentence 1 → Sentence 2 similar
Sentence 2 → Sentence 3 large semantic distance
The large semantic change indicates that a new chunk should start before Sentence 3.
Therefore:
Chunk 1:
FAISS is used for vector similarity search.
It can efficiently retrieve similar embedding vectors.
Chunk 2:
Toronto has cold winters.
This can be useful for documents where topic boundaries do not always follow paragraphs, pages, or other structural separators.
Unlike fixed and recursive chunking, parameters such as:
chunk_size = 500
chunk_overlap = 50
are not the primary mechanism in semantic chunking. Instead, we usually define a threshold that determines how large a semantic change must be before creating a new chunk.
How It Works
Suppose we have:
S1: Machine learning models learn patterns from data.
S2: Supervised learning uses labeled training data.
S3: Classification is a supervised learning task.
S4: PostgreSQL is a relational database.
S5: PostgreSQL supports SQL queries.
Semantic chunking works approximately like this:
Document
↓
Split into sentences
↓
Generate embeddings
↓
Compare neighboring semantic representations
↓
Calculate semantic distance
↓
Detect a large distance
↓
Create a new chunk
For example:
S1 → S2 small distance
S2 → S3 small distance
S3 → S4 LARGE distance
S4 → S5 small distance
Therefore:
Chunk 1
├── S1
├── S2
└── S3
Chunk 2
├── S4
└── S5
In LangChain's SemanticChunker, we can use a breakpoint threshold such as:
breakpoint_threshold_type="percentile"
breakpoint_threshold_amount=95
This means unusually large semantic-distance changes are treated as potential chunk boundaries.
One important difference is that semantic chunking usually produces variable-sized chunks rather than guaranteeing every chunk is below a fixed number of tokens.
Downsides
As mentioned in the introduction, semantic chunking requires additional embedding work during ingestion.
First, embeddings are generated to determine semantic boundaries:
Document
↓
Sentence groups
↓
Embeddings
↓
Detect semantic boundaries
↓
Final chunks
After the final chunks are created, we normally need to embed those chunks again before storing them in the vector database:
Final chunks
↓
Embedding again
↓
FAISS / Vector Database
Therefore, compared with fixed or recursive chunking, semantic chunking can require more embedding cost and ingestion time.
For documents that already have clear paragraph or section structures, recursive chunking may be sufficient. Semantic chunking is more useful when topic changes do not align well with those structural boundaries.
Practical Code Quick View
If our RAG application uses a Google embedding model, we can use GoogleGenerativeAIEmbeddings with LangChain's SemanticChunker.
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_experimental.text_splitter import SemanticChunker
# Google Gemini embedding model
embeddings = GoogleGenerativeAIEmbeddings(
model="gemini-embedding-2"
)
# Semantic chunking
text_splitter = SemanticChunker(
embeddings=embeddings,
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=95,
)
chunks = text_splitter.split_text(document_text)
The important difference from recursive chunking is:
Recursive Chunking
→ separators + chunk size decide the boundaries
Semantic Chunking
→ embedding distance decides the boundaries
Here, the Google embedding model converts sentence groups into vectors:
Text
↓
Gemini Embedding Model
↓
Embedding vectors
↓
Compare semantic distance
↓
Detect breakpoint
↓
Create chunks
The embedding model is not directly generating or rewriting the text. It is only providing vector representations that allow the semantic chunker to measure how much the meaning changes.
If the input is already represented as LangChain Document objects:
chunks = text_splitter.split_documents(documents)
Each resulting chunk remains a Document, so its text and metadata can continue through the RAG pipeline.