Loading page…
THE ENGINEERING NOTEBOOK
Retrieval-Augmented Generation (RAG) is an architecture that combines information retrieval with a Large Language Model (LLM).
Traditional LLMs generate answers based on the knowledge learned during training. However, when users ask questions about private, proprietary, or newly created data that was not included in the training dataset, the model may provide inaccurate answers or hallucinate information.
One simple approach is to provide the entire document to the LLM and ask questions about it. However, this approach has several limitations. A document may contain many pages while the relevant information exists only in a small paragraph. Passing the entire document to the model increases token usage, cost, and processing time. More importantly, the LLM may have difficulty identifying the most relevant information when too much unrelated context is provided.
RAG addresses this problem by introducing a retrieval layer before generation. During the ingestion process, documents are divided into smaller chunks, converted into vector embeddings, and stored in a searchable database or vector index.
When a user asks a question, the retrieval system searches for the most relevant chunks from the knowledge base. Only those relevant chunks are then provided to the LLM as additional context.
By retrieving relevant information before generating an answer, RAG can improve answer accuracy, reduce hallucination, lower unnecessary token usage, and allow LLM applications to work with private or domain-specific knowledge without retraining the underlying model.
Read from top to bottom: earlier approaches to later developments.
Earlier
| Relative progression | 01 /Ingestion / Indexing | 02 /Retrieval | 03 /Generation | 04 /Evaluation |
|---|---|---|---|---|
| Evolution position 1 | ||||
| Evolution position 2 | ||||
| Evolution position 3 | ||||
| Evolution position 4 | ||||
| Evolution position 5 |
Later