topic: RAG Retrieval, Step by Step
BM25 Search
How BM25 uses term frequency, inverse document frequency, and document length for keyword retrieval, its strengths with exact identifiers, and a practical LangChain example.
Table of contents
Introduction
BM25 Search is a commonly used retrieval method for finding relevant information based on keywords.
The main idea is simple:
If important words from the user's query also appear in a document chunk, that chunk is probably relevant.
For example, suppose our knowledge base contains:
Section 15.2 describes the requirements for employee termination.
Section 15.3 describes the requirements for temporary employee termination.
A user asks:
What does Section 15.2 say?
BM25 can give the first chunk a high score because the exact term:
15.2
appears in both the query and the document.
This makes BM25 particularly useful when the user's query contains important exact words, names, numbers, or identifiers.
How It Works
BM25 does not convert text into vector embeddings.
Instead, it looks at the words in the user's query and checks how those words appear across the document chunks.
Conceptually:
User Query
↓
Tokenize Words
↓
Compare Query Terms with Each Chunk
↓
Calculate BM25 Score
↓
Rank Chunks
↓
Top-k Results
BM25 mainly considers three ideas.
Term Frequency
If words from the query appear inside a document, that document becomes more relevant.
For example:
Query:
machine learning engineer
Document A:
machine learning engineer responsibilities...
Document B:
marketing manager responsibilities...
Document A will receive a higher score because more of the query terms appear in the document.
However, repeating the same word many times should not make a document infinitely more relevant.
For example:
machine machine machine machine machine...
BM25 uses a saturation mechanism so that repeated occurrences provide less and less additional benefit.
Inverse Document Frequency
Some words are more useful than others.
Suppose a word appears in almost every document:
employee
company
system
These words are not very useful for identifying one specific document.
However, something like:
LAW-15.2
ERROR-502
POLICY-8231
may appear in only a few documents.
BM25 gives more importance to these relatively rare and distinctive terms.
Document Length
BM25 also considers the length of a document.
Imagine one chunk contains 100 words while another contains 2,000 words.
The 2,000-word chunk naturally has a higher chance of containing query keywords simply because it contains much more text.
BM25 applies document-length normalization so that long documents do not automatically receive higher scores.
After calculating the score for every chunk, the retriever can return the top-k highest-scoring chunks.
For example:
k = 3
means that we return the three chunks with the highest BM25 scores.
Downsides
The major limitation of BM25 is that it mainly matches words, not the meaning behind those words.
For example, suppose our knowledge base contains:
Employees are allowed to work remotely three days per week.
The user asks:
How many days can I work from home?
A human can easily understand that:
work remotely
and:
work from home
have almost the same meaning.
However, BM25 does not truly understand this relationship.
It mainly sees that the query contains:
work
home
while the document contains:
work
remotely
Because there is less keyword overlap, the document may receive a lower score even though it actually contains the correct answer.
This problem becomes more common when users describe the same concept using different words, such as:
car → automobile
work from home → remote work
terminate → fire
purchase → buy
BM25 is therefore especially useful when exact terms matter, such as:
- Law or regulation numbers
- Product names
- Product IDs
- Error codes
- Technical terminology
- Ticket numbers
But when users may express the same idea using different words, BM25 alone may not retrieve the best document.
Because of this, many RAG systems combine keyword-based retrieval such as BM25 with embedding-based retrieval. This approach is commonly called Hybrid Search.
Practical Code Quick View
A simple BM25 retriever can be implemented using LangChain.
pip install langchain-community rank-bm25
from langchain_core.documents import Document
from langchain_community.retrievers import BM25Retriever
documents = [
Document(
page_content="Section 15.2 describes the requirements for employee termination."
),
Document(
page_content="Section 15.3 describes the requirements for temporary employee termination."
),
Document(
page_content="Section 20.1 describes employee vacation policies."
),
]
# Build the BM25 retriever
retriever = BM25Retriever.from_documents(documents)
# Return the top 2 results
retriever.k = 2
query = "Section 15.2 termination"
results = retriever.invoke(query)
for document in results:
print(document.page_content)
Conceptually, the code is doing:
Documents
↓
Tokenize
↓
Build BM25 Index
User Query
↓
Tokenize
↓
Calculate BM25 Scores
↓
Rank Chunks
↓
Top-k Relevant Chunks
One important point is that BM25 does not require an embedding model.
There is no process like:
Text
↓
Embedding Model
↓
Vector
Instead, BM25 builds a keyword-based index and calculates how relevant each document is to the words in the user's query.
This makes BM25 a simple and effective retrieval method when exact keywords and identifiers contain important information.