Naive RAG with Gemini
Document ingestion, vector retrieval, and context-based answers
A backend-focused RAG learning project that connects document ingestion, configurable chunking, Gemini embeddings, FAISS retrieval, and answer generation through a FastAPI backend and Streamlit demo.
Built with
- Python
- LangChain
- Google Gemini
- FAISS
- FastAPI
- Streamlit
- Docker Compose
The workflow
From documents to answers.
One pipeline, from document preparation to context-based generation.
- 01
Ingest
Text & documents
Paste text or upload documents. Keep the source metadata.
- 02
Split
LangChain
Create recursive chunks with configurable size and overlap.
- 03
Retrieve
Gemini embeddings · FAISS
Find the top-k similar chunks in the selected knowledge base.
- 04
Generate
Gemini
Prompt with retrieved context. Return answers, chunks, and sources.
01 / Intent
Build the whole loop.
Explore how to implement the full RAG workflow, from preparing user-provided documents to retrieving context and generating answers within a selected knowledge base.
Built an interactive RAG demo with responses that include retrieved chunks and source metadata. FastAPI exposes the backend API, Streamlit provides the demo interface, and Docker Compose starts both services.
02 / Implementation
Behind the interface.
Used LangChain recursive character splitting with configurable chunk size and overlap, Google Gemini embeddings, FAISS top-k vector similarity search, and Gemini generation prompted to use retrieved context. The backend was primarily implemented by me with some GitHub Copilot assistance; the Streamlit frontend was primarily generated with Codex assistance to demonstrate frontend-backend integration.
03 / Scope
A learning implementation.
User-provided pasted text and PDF, DOCX, TXT, and Markdown uploads, with source metadata retained. Knowledge bases and vector indexes are stored in memory and cleared when the backend restarts.
Current boundary
The project provides a learning environment for document ingestion, chunking, retrieval, and generation. Storage is in memory, so documents must be ingested again after a backend restart.
