GraphRAG -- Knowledge Graphs Meet RAG
Knowledge
Standard RAG finds information that exists within a single chunk. But what if the answer is distributed across multiple documents and the connections between the information are what matters?
GraphRAG combines Knowledge Graphs with RAG. Instead of only storing text chunks, entities (people, products, concepts) and their relationships are modeled in a graph. The retrieval phase then uses both vector search and graph traversal.
Understand
Why Vector Search Fails for Multi-Hop
Consider this question: "Which customers who purchased Product A also filed complaints about Product B?"
This question requires multi-hop reasoning -- three information jumps:
- Find customers who purchased Product A
- Check if those customers filed complaints
- Filter for complaints about Product B
In a pure vector RAG system, the query is represented as a single vector and compared against all chunks. The search might find chunks about Product A and chunks about complaints regarding Product B -- but the connection between customers, purchases, and complaints is lost.
How GraphRAG Works
Phase 1: Knowledge Extraction (Indexing)
- An LLM extracts entities and relationships from the documents
- Entities: customers, products, events, concepts
- Relationships: "purchased", "filed complaint", "is part of"
- The resulting knowledge graph is stored in a graph database (e.g., Neo4j)
- In parallel, text chunks are stored in a vector database as usual
Phase 2: Hybrid Retrieval (Query)
- The question is analyzed: Which entities are mentioned? Which relationships are asked about?
- Graph traversal: Follow relationships in the knowledge graph to find connected entities
- Vector search: Find supplementary text chunks that provide context
- Fusion: Combine graph results and vector results
- Generation: The LLM synthesizes the answer from both sources
Microsoft's GraphRAG Approach
Microsoft published an influential GraphRAG approach in 2024 that uses community detection:
- Entity Extraction: LLM identifies entities and relationships in each chunk
- Community Detection: Closely connected entities are grouped into communities (e.g., Leiden algorithm)
- Community Summaries: For each community, the LLM generates a summary
- Query Routing: Questions are routed to the most relevant communities
This approach is particularly strong for global questions ("What are the main themes across these 10,000 documents?"), where local vector search fails because the answer is distributed across the entire corpus.
When GraphRAG and When Not
| Scenario | Vector RAG | GraphRAG | Rationale |
|---|---|---|---|
| "What does paragraph 3.2 say?" | Better | Overkill | Simple lookup question, one chunk suffices |
| "Which projects does Person X lead?" | Works | Better | Relationship between person and projects |
| "What dependencies exist between these 5 systems?" | Often fails | Significantly better | Multi-hop across multiple relationships |
| "What are the main themes across the entire document collection?" | Fails | Strong | Global question requires aggregation |
iGraphRAG Is Not a Replacement
GraphRAG doesn't replace Vector RAG -- it complements it. Most production-ready systems use both: vector search for local questions, graph traversal for relationship and aggregation questions.
A pharmaceutical company wants a system that answers questions like: 'Which studies show interactions between Drug A and Drug B, and which patient groups are affected?' -- Which approach is suitable?
Apply
Implementation Complexity
GraphRAG is significantly more complex than standard RAG:
- Entity extraction requires LLM calls per chunk (cost!)
- Graph database must be operated (Neo4j, Amazon Neptune)
- Schema design for the knowledge graph requires domain expertise
- Hybrid retrieval must meaningfully fuse graph results and vector results
- Maintenance: The graph must be updated when documents change
Rule of thumb: Start with Vector RAG. If multi-hop questions are a common pattern and answer quality is insufficient, invest in GraphRAG.
Reflect
Consider which data in your work environment is highly interconnected. Are there scenarios where the relationships between entities are just as important as the entities themselves? That's where GraphRAG delivers the most value.