topic: RAG Chunking, Step by Step
Contextual Retrieval
How contextual retrieval restores document context before embedding, with prompt caching, retrieval failure-rate results, and a practical Gemini example.
Table of contents
Introduction
Contextual Retrieval is a technique used in RAG systems to reduce the context loss caused by chunking.
When we split a large document into smaller chunks, a chunk may still contain correct information, but it can lose the context needed to understand what that information refers to.
Anthropic gives a good example using a collection of SEC financial filings.
Suppose one chunk contains:
The company's revenue grew by 3% over the previous quarter.
The sentence makes sense by itself, but important information is missing:
Which company?
Which quarter?
Which financial filing?
This becomes a problem if our knowledge base contains thousands of financial reports with similar sentences.
A user may ask:
What was the revenue growth in Q2 2023?
but the relevant chunk only says:
The company's revenue grew by 3% over the previous quarter.
The embedding model may have difficulty connecting them.
Contextual Retrieval solves this by generating a short piece of context for each chunk before creating its embedding.
For example:
Context:
This chunk comes from a Q2 2023 SEC filing and discusses
the company's quarterly revenue performance.
Original Chunk:
The company's revenue grew by 3% over the previous quarter.
Now the chunk contains information about where it belongs in the original document.
Anthropic calls the embedding version of this approach Contextual Embeddings. The same idea can also be applied before building a BM25 index, which they call Contextual BM25.
How It Works
A normal RAG ingestion pipeline may look like:
Document
↓
Chunking
↓
Embedding
↓
Vector Database
With Contextual Retrieval, we add another step before embedding:
Document
↓
Chunking
↓
Whole Document + Current Chunk
↓
LLM generates short context
↓
Context + Original Chunk
↓
Embedding
↓
Vector Database
For example:
Original Chunk:
The company's revenue grew by 3%
over the previous quarter.
We provide both the whole document and this specific chunk to an LLM.
The LLM may generate:
This chunk discusses revenue performance
from the company's Q2 2023 financial filing.
Then we combine them:
This chunk discusses revenue performance
from the company's Q2 2023 financial filing.
The company's revenue grew by 3%
over the previous quarter.
This final contextualized text is what we embed.
At query time, the process is still normal retrieval:
User Question
↓
Query Embedding
↓
Vector Search
↓
Relevant Chunks
↓
LLM
↓
Answer
Therefore, Contextual Retrieval does not replace vector search.
It improves the chunks that vector search operates on.
When Is This Approach Useful?
A particularly good situation is the same type of problem shown in Anthropic's SEC filing example.
Imagine a RAG system containing:
Thousands of SEC filings
├── Different companies
├── Different years
├── Different quarters
└── Similar financial terminology
Many chunks may look very similar:
Revenue increased by 3%.
Revenue increased by 5%.
Operating income decreased by 2%.
Revenue increased compared with the previous quarter.
The problem is not that these chunks contain bad information.
The problem is that the information required to identify them may exist somewhere else in the document:
Document title
Company
Quarter
Year
Section
This is where Contextual Retrieval is especially worth trying.
In other words:
Chunk content is useful
+
Chunk loses identifying context
+
Many documents contain similar-looking chunks
↓
Contextual Retrieval can help
Simply increasing chunk_overlap may not fully solve this problem because the identifying information could be much earlier in the document rather than immediately beside the chunk.
Why Not Just Give the Whole Document to the LLM?
If the knowledge base is small enough, this may actually be the simpler solution.
Instead of:
Document
↓
Chunk
↓
Retrieve
↓
LLM
we could simply do:
Whole Document
+
User Question
↓
LLM
Anthropic specifically notes that if the entire knowledge base is below roughly 200,000 tokens, approximately 500 pages in their rough estimate, directly providing the knowledge base to the model can be a reasonable alternative to RAG.
For example, if our application only answers questions about one relatively small PDF, Contextual Retrieval may introduce unnecessary complexity.
The difference becomes important when we have:
1 small document
→ Sending the whole document may be reasonable.
Thousands of large documents
→ Retrieval becomes necessary.
Contextual Retrieval is mainly useful for the second situation.
What About Cost and Caching?
One obvious problem is that contextualization appears to repeatedly send the same document:
Full Document + Chunk 1
Full Document + Chunk 2
Full Document + Chunk 3
Full Document + Chunk 4
That looks expensive.
However, most of the prompt is identical:
Full Document ← repeated
Chunk ← changes
This makes prompt caching useful.
Conceptually:
Full Document
↓
Cache
↓
┌────┼────┐
↓ ↓ ↓
C1 C2 C3
Anthropic specifically recommends prompt caching when generating contextualized chunks because the document prefix can be reused across requests.
If we use Google models, Gemini also supports context caching. Gemini 2.5 and newer models provide implicit caching, while explicit caching can also be used through the appropriate Gemini API when we want more control over reuse and cache lifetime.
Therefore, in production we do not necessarily need to pay the full input cost of repeatedly processing the same document.
Anthropic's Experiment Results
Anthropic evaluated this approach across multiple knowledge domains and measured how often the relevant information failed to appear within the top 20 retrieved chunks.
Their experiments found:
Standard Embedding
Failure Rate: 5.7%
↓ Contextual Embedding
Failure Rate: 3.7%
35% relative reduction
When they combined Contextual Embeddings + Contextual BM25:
5.7% → 2.9%
49% relative reduction
After adding a reranking step:
5.7% → 1.9%
67% relative reduction
An important detail is that this does not mean:
Accuracy increased by 67%
It means that their measured top-20 retrieval failure rate was reduced by 67% relative to the baseline.
The main takeaway is that adding document-level context to chunks significantly improved the probability of retrieving the correct information in their experiments.
Downsides
Contextual Retrieval adds another LLM operation during ingestion.
Instead of:
Chunk
↓
Embedding
we now have:
Chunk + Document
↓
LLM
↓
Contextualized Chunk
↓
Embedding
This increases:
Ingestion time
LLM usage
Pipeline complexity
Prompt caching can reduce some of the repeated cost, but the contextualization step still exists.
Another concern is that the generated context must be accurate. If the LLM incorrectly describes a chunk, the extra context could make retrieval worse rather than better.
Therefore, Contextual Retrieval should solve a real retrieval problem rather than being automatically added to every RAG system.
Practical Code Quick View
If our RAG application uses Google models, we can use Gemini to generate the contextual description and a Gemini embedding model to generate the final vectors.
from langchain_google_genai import (
ChatGoogleGenerativeAI,
GoogleGenerativeAIEmbeddings,
)
from langchain_core.documents import Document
from langchain_community.vectorstores import FAISS
# Generate chunk context
llm = ChatGoogleGenerativeAI(
model="gemini-3.8-flash"
)
# Generate final embeddings
embeddings = GoogleGenerativeAIEmbeddings(
model="gemini-embedding-2"
)
def contextualize_chunk(full_document, chunk):
prompt = f"""
Here is the full document:
{full_document}
Here is one chunk from the document:
{chunk}
Give a short context explaining where this chunk
belongs in the document for retrieval purposes.
Return only the short context.
"""
context = llm.invoke(prompt).content
return f"{context}\n\n{chunk}"
contextualized_chunks = []
for chunk in chunks:
contextualized_text = contextualize_chunk(
document_text,
chunk.page_content,
)
contextualized_chunks.append(
Document(
page_content=contextualized_text,
metadata=chunk.metadata,
)
)
vector_store = FAISS.from_documents(
contextualized_chunks,
embeddings,
)
gemini-3.8-flash is currently a stable Gemini model, while gemini-embedding-2 is Google's current stable newer-generation embedding model.
The most important difference to remember is:
Traditional Retrieval
Chunk
↓
Embedding
↓
Vector Database
Contextual Retrieval
Document + Chunk
↓
LLM
↓
Short Chunk Context
↓
Context + Chunk
↓
Embedding
↓
Vector Database
The additional LLM is not answering the user's question.
Its job is to help each chunk remember where it came from before that chunk is stored for retrieval.