topic: RAG Retrieval, Step by Step
Hybrid Search
Combine BM25 and semantic search with Reciprocal Rank Fusion, tune retriever weights, and implement hybrid retrieval with LangChain, FAISS, and Google embeddings.
Table of contents
Introduction
Hybrid Search combines multiple retrieval methods to improve the quality of retrieved results.
In the previous posts, we introduced:
- Semantic Search: retrieves based on meaning.
- BM25 Search: retrieves based on keyword matching.
Each method has different strengths.
For example, BM25 may work better when the query contains an exact term such as:
Section 15.2
ERROR-502
PRODUCT-123
Semantic Search may work better when the user describes the same idea using different words.
Instead of choosing only one retrieval method, Hybrid Search uses both and combines their results.
User Query
↓
┌────────┴────────┐
↓ ↓
BM25 Semantic Search
↓ ↓
└────────┬────────┘
↓
Rank Fusion
↓
Final Results
One common way to combine these rankings is Reciprocal Rank Fusion (RRF).
How It Works
When a query comes in, we send it to both BM25 and Semantic Search.
For example:
BM25 Ranking
1. Document A
2. Document B
3. Document C
Semantic Search may return:
Semantic Ranking
1. Document B
2. Document D
3. Document A
Now we need to combine these two rankings.
Reciprocal Rank Fusion
We usually should not directly add BM25 and semantic similarity scores because they use different scoring systems.
RRF avoids this problem by looking mainly at the rank position instead of the original score.
The simplified formula is:
RRF Score =
Σ 1 / (c + rank)
For example, if Document A is:
BM25 Rank = 1
Semantic Rank = 3
it receives a score from both rankings.
Documents that appear near the top of both retrieval methods will usually receive a higher final RRF score.
BM25 Ranking
↓
├───────┐
↓
RRF
↓
Final Ranking
↑
├───────┘
↑
Semantic Ranking
Weighted RRF
Sometimes we do not want both retrieval methods to have the same importance.
For example, a technical knowledge base may contain many exact identifiers such as:
ERROR-502
API-128
SERVICE-201
In this case, we may want BM25 to contribute more.
Weighted RRF adds a weight to each retriever:
Weighted RRF Score =
Σ weight / (c + rank)
For example:
BM25 Weight = 0.7
Semantic Weight = 0.3
This means BM25 has more influence on the final ranking.
For applications where users mostly ask natural-language questions, we may instead use:
BM25 Weight = 0.3
Semantic Weight = 0.7
There is no universal best weight. The weights should ideally be tested using retrieval evaluation data.
Downsides
Hybrid Search requires running multiple retrieval methods for the same query.
Instead of:
Query
↓
Retriever
we now have:
Query
↓
BM25 + Semantic Search
↓
Rank Fusion
This adds some retrieval complexity and computation.
Weighted RRF also introduces additional parameters that need to be tuned.
Another limitation is that RRF mainly considers ranking positions.
For example:
Document A → 0.95
Document B → 0.94
and:
Document A → 0.95
Document B → 0.50
both become:
Document A → Rank 1
Document B → Rank 2
RRF does not fully preserve the difference between the original scores.
However, this is also why RRF is useful: we can combine retrieval methods without having to normalize their different scoring systems.
Practical Code Quick View
We can implement Hybrid Search using LangChain, BM25, FAISS, and a Google embedding model.
from langchain_core.documents import Document
from langchain_community.retrievers import BM25Retriever
from langchain_community.vectorstores import FAISS
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_classic.retrievers import EnsembleRetriever
documents = [
Document(
page_content="Section 15.2 describes employee termination requirements."
),
Document(
page_content="Section 15.3 describes temporary employee termination."
),
Document(
page_content="Section 20.1 describes employee vacation policies."
),
]
# BM25
bm25_retriever = BM25Retriever.from_documents(documents)
bm25_retriever.k = 3
# Semantic Search
embeddings = GoogleGenerativeAIEmbeddings(
model="gemini-embedding-001"
)
vector_store = FAISS.from_documents(
documents,
embeddings
)
semantic_retriever = vector_store.as_retriever(
search_kwargs={"k": 3}
)
# Hybrid Search using Weighted RRF
hybrid_retriever = EnsembleRetriever(
retrievers=[
bm25_retriever,
semantic_retriever,
],
weights=[
0.3,
0.7,
],
)
results = hybrid_retriever.invoke(
"What are the rules for firing an employee?"
)
for document in results:
print(document.page_content)
Conceptually:
User Query
↓
┌────────┴────────┐
↓ ↓
BM25 Semantic Search
↓ ↓
Weight Weight
└────────┬────────┘
↓
Weighted RRF
↓
Final Ranking
Hybrid Search allows us to combine the strengths of keyword matching and semantic understanding, while RRF provides a simple way to merge their rankings.