topic: RAG Retrieval, Step by Step
Rerank
Rerank retrieved candidates before generation, understand recall and latency trade-offs, and use Google's Ranking API in a RAG pipeline.
Table of contents
Introduction
Reranking is a retrieval optimization technique used after the initial search in a RAG system.
In Semantic Search, BM25, or Hybrid Search, we normally retrieve the top K chunks based on similarity or retrieval scores.
For example:
User Query
↓
Hybrid Search
↓
Top 20 candidate chunks
However, the initial retrieval score does not always mean that the most useful chunk is ranked first.
For example:
Rank 1 → somewhat relevant
Rank 2 → somewhat relevant
Rank 3 → correct answer
Rank 4 → irrelevant
Rank 5 → relevant
If we only send the top 2 chunks to the LLM, the correct information may never reach the generation stage.
Reranking adds another model after retrieval to re-evaluate the candidate chunks more carefully and reorder them.
How It Works
Suppose the user asks:
"What is the refund policy for enterprise customers
after cancelling an annual contract?"
The initial retrieval system may retrieve the top 20 chunks:
Query
↓
Semantic / BM25 / Hybrid Retrieval
↓
Top 20 candidate chunks
The first-stage retriever is designed to search a large knowledge base efficiently.
The reranker then receives the query together with each retrieved candidate:
Query + Chunk 1 → relevance score
Query + Chunk 2 → relevance score
Query + Chunk 3 → relevance score
...
Query + Chunk 20 → relevance score
For example, the initial retrieval may return:
Rank Chunk Score
1 General cancellation policy 0.89
2 Enterprise annual contract renewal 0.87
3 Enterprise cancellation refund policy 0.84
4 Monthly subscription refund policy 0.82
After reranking:
Rank Chunk Score
1 Enterprise cancellation refund policy 0.96
2 General cancellation policy 0.78
3 Enterprise annual contract renewal 0.65
4 Monthly subscription refund policy 0.41
We may then send only the top 3 reranked chunks to the LLM.
Knowledge Base
↓
Initial Retrieval
↓
Top 20 candidates
↓
Reranker
↓
Top 3 chunks
↓
LLM
↓
Answer
The important idea is that the first retriever focuses on finding candidates efficiently, while the reranker spends more computation deciding which candidates are actually the most relevant.
Downsides
The biggest downside is latency and cost.
A normal vector search can search through a large number of vectors very quickly.
A reranker needs to evaluate the relationship between the query and multiple retrieved chunks more carefully.
Therefore, we normally do not rerank the entire knowledge base.
Instead:
100,000 chunks
↓
Fast Retrieval
↓
Top 20–50 chunks
↓
Rerank
↓
Top 3–5 chunks
Another important limitation is that reranking cannot recover information that was never retrieved.
For example:
Correct chunk rank after retrieval = 87
Rerank input = Top 20
The reranker never sees the correct chunk, so it cannot fix the problem.
A useful diagnostic is:
Correct chunk is NOT in Top 20
→ Retrieval problem
Correct chunk is in Top 20 but ranked too low
→ Reranking may help
Practical
Google provides a dedicated Ranking API that can rerank documents retrieved by Semantic Search, BM25, or Hybrid Search.
Suppose our retriever already returned several candidate chunks:
retrieved_chunks = [
{
"id": "1",
"content": "Customers can cancel their subscription at any time."
},
{
"id": "2",
"content": (
"Enterprise annual contracts receive a prorated refund "
"when cancellation is approved."
)
},
{
"id": "3",
"content": "Monthly plans automatically renew every month."
},
]
We can rerank them using Google's semantic ranking model:
from google.cloud import discoveryengine_v1 as discoveryengine
PROJECT_ID = "YOUR_PROJECT_ID"
client = discoveryengine.RankServiceClient()
ranking_config = client.ranking_config_path(
project=PROJECT_ID,
location="global",
ranking_config="default_ranking_config",
)
query = (
"What is the refund policy for enterprise customers "
"after cancelling an annual contract?"
)
records = [
discoveryengine.RankingRecord(
id=chunk["id"],
content=chunk["content"],
)
for chunk in retrieved_chunks
]
request = discoveryengine.RankRequest(
ranking_config=ranking_config,
model="semantic-ranker-default@latest",
query=query,
records=records,
top_n=2,
)
response = client.rank(request=request)
for record in response.records:
print(record.id, record.score, record.content)
The complete RAG flow becomes:
Semantic / Hybrid Search
↓
Top 20 chunks
↓
Google Semantic Ranker
↓
Top 3 chunks
↓
Gemini
↓
Final Answer
For example:
retrieved_chunks = retriever.invoke(query) # retrieve Top 20
reranked_chunks = google_reranker(
query=query,
chunks=retrieved_chunks,
top_n=3,
)
answer = gemini.invoke(
question=query,
context=reranked_chunks,
)
The reranker therefore does not replace retrieval.
Instead, it sits between retrieval and generation:
Retrieval → Google Reranker → Gemini