topic: RAG Retrieval, Step by Step
Semantic Search
How semantic search uses embeddings and vector similarity to retrieve relevant chunks, its limitations with exact identifiers, and a practical LangChain, Google embeddings, and FAISS example.
Table of contents
Introduction
Semantic Search is one of the most common retrieval methods used in RAG systems.
Instead of searching for documents based only on exact keyword matches, semantic search focuses on the meaning of the query and the document chunks.
The main idea is simple:
If two pieces of text have similar meanings, their embeddings should also be close to each other in the vector space.
For example, suppose our knowledge base contains:
Employees are allowed to work remotely for up to three days per week.
A user asks:
How many days can I work from home?
Even though the query does not contain the exact words "work remotely", semantic search can still recognize that "work from home" and "work remotely" have similar meanings.
How It Works
During the ingestion stage, each document chunk is converted into a vector using an embedding model and stored in a vector store.
For example:
Document Chunk
↓
Embedding Model
↓
[0.12, -0.41, 0.73, ...]
When a user asks a question, we use the same embedding model to convert the query into an embedding.
User Query
↓
Same Embedding Model
↓
Query Vector
The vector store then compares the query vector against the stored document vectors.
Common similarity measurements include:
- Cosine similarity
- Euclidean distance
- Dot product
The system then returns the top-k most similar chunks.
For example, if:
k = 3
the retriever returns the three chunks considered most relevant to the user's question.
User Query
↓
Query Embedding
↓
Compare with Stored Embeddings
↓
Rank by Similarity
↓
Top 3 Chunks
These retrieved chunks can then be passed to the LLM as context for generating the final answer.
Downsides
One important limitation of semantic search is that semantic similarity is not always the same as exact matching.
Consider a legal knowledge base containing:
Section 15.2 - Requirements for employee termination
Section 15.3 - Requirements for temporary employee termination
Suppose the user asks:
What does Section 15.2 say about termination?
Semantic search understands the meaning of words such as "termination", but identifiers such as:
15.2
15.3
ABC-123
Policy-4721
may carry very little semantic meaning for an embedding model.
As a result, the retriever may return Section 15.3 because its surrounding text is semantically very similar to the query, even though the user explicitly asked for Section 15.2.
This problem commonly appears when documents contain exact identifiers such as:
- Law or regulation numbers
- Product IDs
- Error codes
- Ticket numbers
- Employee IDs
In these situations, pure semantic search may not be enough. Keyword-based retrieval such as BM25, or a hybrid search combining keyword and semantic retrieval, can help handle exact-match information.
Practical Code Quick View
A simple semantic search can be implemented with LangChain, a Google embedding model, and FAISS.
pip install langchain langchain-community langchain-google-genai faiss-cpu
from langchain_core.documents import Document
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_community.vectorstores import FAISS
# Example document chunks
documents = [
Document(
page_content="Section 15.2 describes the requirements for employee termination."
),
Document(
page_content="Section 15.3 describes the requirements for temporary employee termination."
),
Document(
page_content="Section 20.1 describes employee vacation policies."
),
]
# Create the embedding model
embeddings = GoogleGenerativeAIEmbeddings(
model="gemini-embedding-001"
)
# Embed the document chunks and store them in FAISS
vector_store = FAISS.from_documents(
documents,
embeddings
)
# User question
query = "What are the rules for terminating an employee?"
# Retrieve the top 2 semantically similar chunks
results = vector_store.similarity_search(
query,
k=2
)
for document in results:
print(document.page_content)
Conceptually, LangChain is doing:
Documents
↓
Embedding Model
↓
FAISS Vector Index
User Query
↓
Same Embedding Model
↓
Query Vector
↓
Similarity Search
↓
Top-k Relevant Chunks
The important idea is that we are not asking the LLM to search the documents directly. We first use embeddings and vector similarity to retrieve a small set of relevant chunks, which can then be provided to the LLM during the generation stage.