topic: RAG Chunking, Step by Step
Blog 5 — Late Chunking: Give Each Chunk More Document Context
Fixed-size, recursive, and semantic chunking mainly differ in how they select boundaries.
Table of contents
Fixed-size, recursive, and semantic chunking mainly differ in how they select boundaries.
Late chunking changes a different part of the indexing process. It changes when the final chunk embedding is created.
The traditional process
A traditional chunking pipeline first divides the document and then embeds every chunk independently.
Document
↓
Split into chunks
↓
Embed Chunk 1 separately
Embed Chunk 2 separately
Embed Chunk 3 separately
This is sometimes called early chunking because the document is divided before the embedding model processes each chunk.
Consider this document:
Andy moved to Toronto in 2024.
He later opened a noodle restaurant there.
After early chunking, the second chunk may contain only:
He later opened a noodle restaurant there.
The words He and there refer to information in the previous chunk.
He → Andy
there → Toronto
When the second chunk is embedded independently, the embedding model does not receive those names as part of its input.
The late-chunking process
Late chunking changes the order:
Long document
↓
Embedding model reads the document
↓
Create context-aware representations
for the individual tokens
↓
Apply chunk boundaries
↓
Combine the token representations
inside each chunk
↓
Create the final chunk embeddings
The embedding model sees the larger document before the final chunk vectors are produced.
A context-aware representation means that the numerical representation of a word is influenced by the surrounding text.
The combining step is called pooling. Pooling takes multiple token-level representations and combines them into one vector for the chunk. The original late-chunking paper applies the chunk boundaries after the main model and before mean pooling, where mean pooling averages the token representations.
Because the model has already read the earlier sentence, the representation of He can contain information related to Andy. The representation of there can contain information related to Toronto.
Late chunking still needs boundaries
Late chunking does not remove the need for a chunking strategy.
The system still needs boundaries that define which tokens belong to each final chunk. Those boundaries may come from fixed-size, recursive, semantic, or document-structure rules.
Late chunking and semantic chunking therefore solve different problems:
Semantic chunking:
improves where the document is divided
Late chunking:
improves how much surrounding context
influences each chunk embedding
They can also be combined.
Infrastructure requirements
Late chunking requires a long-context embedding model. Long context means that the model can process a large number of tokens at one time.
The implementation also needs access to token-level representations before they are combined into the final vector.
A vector database does not need a special late-chunking feature. After the chunk vectors are created, they can be stored and searched like normal embeddings.
The practical processing cost depends on the model, document length, hardware, and implementation. Long documents may require more memory and more computation, so late chunking is a more specialised indexing design than recursive splitting.
Design decision
Compared with semantic chunking, late chunking does not mainly improve boundary detection. It improves context preservation inside the embedding.
Semantic chunking
→ better topic boundaries
Late chunking
→ more document context inside each chunk vector
The trade-off moves further toward meaning and context, with higher implementation complexity and potentially higher indexing cost.
Late chunking is most useful when retrieval failures involve references that cross chunk boundaries, such as pronouns, definitions, entity names, or earlier explanations.
References
- Günther et al., Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models. The method applies chunking after the transformer model and before mean pooling so chunk vectors can include wider document context. (arXiv)
- freeCodeCamp.org, Production RAG with LangChain & Vector Databases – Full Course. The chapter “Late Chunking vs Early Chunking” begins at approximately 6:24:26. Watch the YouTube course. (CHOOSE TO)