Zum Inhalt springen

The Evolution of RAG

Knowledge

RAG is not a static concept. Since its introduction by Lewis et al. (2020), the architecture has evolved fundamentally. What started as simple "search and generate" is now a spectrum of architectures with different trade-offs.

Three generations have emerged:

  1. Naive RAG (2020--2023): Simple vector search + LLM generation
  2. Advanced RAG (2023--2025): Hybrid Search, Re-Ranking, Query Transformation
  3. Agentic RAG (2025+): LLM-driven, iterative retrieval loops

Stage 1: Naive RAG

Simple pipeline without post-processing

QueryEmbedTop-KLLM

Advantages

  • Easy to implement
  • Low latency
  • Minimal infrastructure required

Disadvantages

  • No relevance optimization
  • Keyword gaps with pure vector search
  • Quality depends heavily on embedding model

Common Pitfalls

  • Irrelevant chunks in Top-K results
  • Hallucinations from poor context
  • Synonyms and paraphrases are missed

When to use?

  • Prototypes, simple FAQ bots, small document collections

Click the stage tabs to explore the evolution of RAG architectures.

Understand

Naive RAG -- The First Generation

Naive RAG follows a rigid three-step process: embed the query, vector search with Top-K, generate the answer. No query preprocessing, no re-ranking of results, no feedback loop.

Why it works: For simple questions with a single document as the answer, this is perfectly sufficient. "What is the return policy?" finds the matching paragraph and the LLM generates a correct answer.

Why it fails:

  • Vague queries: "Tell me about returns" retrieves too many, overly unspecific chunks
  • Multi-hop questions: "Which customers had both returns and complaints?" requires information from multiple documents that must be linked together
  • Keyword mismatch: Semantic search finds "refund" for the query "money back", but misses exact terms like article numbers

A user asks: 'Which products with SKU number AB-1234 were returned in Q3?' -- Naive RAG returns irrelevant results. What is the most likely cause?

Advanced RAG -- Hybrid Search and Re-Ranking

Advanced RAG addresses the weaknesses of Naive RAG through three key extensions:

1. Query Transformation Before the search starts, the query is optimized:

  • Query Rewriting: The LLM rephrases the user's question to achieve better search results
  • HyDE (Hypothetical Document Embeddings): The LLM generates a hypothetical answer, and its embedding is used for the search -- often closer to the target documents than the question itself
  • Query Decomposition: Complex questions are broken down into sub-questions

2. Hybrid Search Combination of two search mechanisms:

  • Dense Retrieval: Vector search via embeddings (semantic similarity)
  • Sparse Retrieval: Keyword-based search like BM25 (exact matching)
  • Fusion: Results from both searches are combined using Reciprocal Rank Fusion (RRF)
Search TypeStrengthWeakness
Dense (Vector)Semantic similarity, synonymsExact terms, codes, IDs
Sparse (BM25)Exact matches, technical termsSynonyms, paraphrases
Hybrid (Fusion)Combines bothHigher effort, tuning required

3. Re-Ranking After retrieval, a cross-encoder model re-evaluates each result in the context of the query. Cross-encoders are slower than bi-encoders (thus not suitable for initial search) but significantly more precise in relevance assessment.

Agentic RAG -- The Iterative Loop

Agentic RAG is the paradigm shift: the LLM actively controls the retrieval process. Instead of a rigid pipeline, the LLM decides:

  • Whether to search at all (perhaps training knowledge is sufficient)
  • What to search for (query formulation and reformulation)
  • Where to search (which data source, which index)
  • Whether the results are sufficient (or if another search is needed)

The loop works like a researcher working iteratively:

  1. Analyze the question and plan a search strategy
  2. Perform the first search
  3. Evaluate results: Are they sufficient? Are they relevant?
  4. If not: reformulate the query or consult a different source
  5. Repeat steps 2--4 until sufficient information is gathered
  6. Synthesize the collected information into a final answer

iAgentic RAG in Practice

Agentic RAG is most powerful when the question is ambiguous or information from different sources needs to be synthesized. For simple lookup questions, it's overhead.

A company builds a technical support system. Customers ask questions like 'My device XZ-500 shows error code E42 -- what should I do?' -- Which RAG architecture is best suited?

Apply

Decision Matrix: Which RAG Generation Do I Need?

CriterionNaive RAGAdvanced RAGAgentic RAG
Question typeSimple lookup questionsStructured questions with technical termsComplex, ambiguous questions
Data sourcesOne homogeneous sourceOne or few sourcesMultiple heterogeneous sources
Latency budget< 1s1--3s3--10s+
Implementation effortLowMediumHigh
Typical use caseFAQ bot, document searchTechnical support, complianceResearch assistant, analysis

Upgrade Path

In practice, you start with Naive RAG and upgrade strategically:

  1. Start with Naive RAG: Build MVP, measure baseline
  2. Measure retrieval quality: How often are the top-5 results relevant?
  3. Identify weaknesses: Keyword mismatches? Irrelevant chunks? Missing connections?
  4. Upgrade strategically: Hybrid Search for keyword problems, Re-Ranking for relevance, Agentic for multi-source

Reflect

You're designing a RAG system for a company with 50,000 technical documents, internal wikis, and ticket systems. Users ask both simple questions ("Where do I find form X?") and complex analysis questions ("Which incidents in the last 6 months involve product Y and have priority 1?"). Consider which architecture you'd choose and why. The quiz questions above help you sharpen your decision criteria.