RAG -- Retrieval-Augmented Generation
Knowledge
LLMs have a fundamental problem: their knowledge is limited to their training data. A model trained in January 2026 knows nothing about events after that date. It also doesn't know about internal company documents, private databases, or daily updates.
Retrieval-Augmented Generation (RAG) solves this problem. Instead of retraining the model, you provide relevant information directly in the prompt. RAG combines two strengths: the LLM's ability to generate natural language with the precision of a database.
iRAG in One Sentence
RAG = First find the right information, then let the LLM answer based on it.
Understand
The RAG Pipeline in Six Steps
Click a step to see its explanation. The animation shows the data flow through the RAG pipeline.
- Query: The user asks a question, e.g., "What is the return policy for online orders?"
- Embedding: The question is converted into an embedding vector -- using the same model that embedded the documents.
- VectorDB: The question vector is matched against the vector database holding all document embeddings.
- Retrieval: The top-K most similar documents are fetched and prepared as context for the LLM.
- LLM: The language model receives the question plus the retrieved documents and generates an answer grounded in both.
- Answer: The grounded response is returned -- backed by current, relevant sources rather than training data alone.
RAG vs. Pure Prompting
| Aspect | Pure Prompting | RAG |
|---|---|---|
| Knowledge Source | Training data only | Training data + external documents |
| Freshness | Training cutoff date | As current as the database |
| Hallucinations | Higher (no fact-checking) | Lower (answers grounded in sources) |
| Citations | Not possible | Possible (document references) |
| Effort | Minimal | Embedding pipeline + vector DB required |
Why does RAG reduce hallucinations compared to pure prompting?
Apply
When to use RAG and when not to
RAG makes sense for:
- Company-specific knowledge (internal documentation, policies)
- Frequently changing information (product catalogs, pricing)
- Areas where citations matter (compliance, legal)
RAG is overkill for:
- General knowledge that any LLM knows ("What is the capital of France?")
- Creative tasks without a factual basis (writing poems, brainstorming)
- Simple conversations without domain-specific content
The Indexing Phase
Before RAG can work, the documents need to be prepared:
- Chunking: Split documents into smaller sections (e.g., 500 tokens per chunk)
- Embedding: Convert each chunk into a vector
- Storage: Store vectors in a vector database (e.g., Pinecone, Weaviate, pgvector)
The quality of chunking has a major impact on results. Chunks that are too large dilute relevance; chunks that are too small lose context.
Reflect
Your company has 10,000 PDF documents with internal policies. Employees should be able to ask questions about them via chat. Which approach is most suitable?