Idris Carter 8 min readRetrieval-augmented generation is often discussed as though it were one architecture: split documents, create embeddings, retrieve passages, and ask a model to answer from them. That pattern is useful, but it conceals three materially different choices.
Vector RAG retrieves by semantic similarity. Hybrid search combines semantic retrieval with exact lexical matching. GraphRAG models entities and relationships, then retrieves through the resulting structure. Each can produce grounded answers. They differ in what they consider relevant, what they demand from your data, and how they fail.
The right question is not which approach is most advanced. It is which representation preserves the evidence your product must recover.
The architectural difference in one view
| Criterion | Vector RAG | Hybrid search | GraphRAG |
|---|---|---|---|
| Primary retrieval signal | Semantic similarity | Semantic similarity plus exact terms | Entities, relationships, paths, and graph communities |
| Best at | Conceptual questions over prose | Mixed natural-language and identifier-heavy queries | Questions spanning connected facts |
| Data preparation | Chunking and embedding | Chunking, embedding, and lexical indexing | Entity and relation extraction, resolution, and graph construction |
| Freshness burden | Relatively low | Relatively low | Higher because documents and graph structure must stay aligned |
| Typical failure | Semantically plausible but incomplete passages | Poor fusion or ranking between retrieval channels | Incorrect or missing graph structure |
| Operational complexity | Lowest | Moderate | Highest |
Vector RAG: the strongest default for meaning-rich documents
Vector RAG converts a query and document chunks into embeddings. A similarity search returns chunks whose positions are close to the query in embedding space. The generator receives those chunks as context.
This works well when users express an idea differently from the source material. A policy may say “employees must obtain written authorization before exporting customer records,” while a user asks, “Can I download client data without approval?” Lexical overlap is limited, but semantic proximity can still surface the correct passage.
The architecture is compact: ingest content, divide it into retrievable units, embed those units, store vectors, retrieve candidates, and optionally rerank them. That simplicity makes evaluation easier. Teams can inspect whether the correct chunk was indexed, retrieved, reranked, and cited.
Where vector retrieval breaks
Semantic similarity is not the same as evidential completeness. Suppose an analyst asks, “Which suppliers connected to Project Alder have unresolved security exceptions?” The answer may require one document naming Project Alder’s suppliers, another recording exception status, and a third mapping a subsidiary to its parent company. No single chunk resembles the entire question closely enough.
Vector search also struggles with exact tokens that embeddings may treat as incidental: product codes, ticket numbers, abbreviations, legal citations, version strings, and personal names with variant spellings. Chunking introduces another weakness. If a definition and its exception land in separate chunks, the retrieved evidence can preserve the rule while losing the qualification.
Choose vector RAG when documents are primarily prose, questions are concept-oriented, and most answers can be supported by a small number of locally coherent passages.
Hybrid search: the pragmatic choice for mixed vocabularies
Hybrid search runs semantic and lexical retrieval together. The lexical side rewards exact or near-exact term matches; the vector side catches paraphrases and conceptual similarity. Their candidate lists are then fused or reranked.
Consider a support assistant asked, “Does build XR-417 still trigger the OAuth redirect defect?” Vector retrieval may locate discussions of authentication loops. Lexical retrieval protects the exact build identifier and protocol term. A reranker can then prioritize passages containing both the right concept and the right artifact.
The central design decision is not merely to add two scores. Vector and lexical scores often live on incompatible scales. A robust system can use rank-based fusion, which combines each result’s position rather than assuming raw scores are directly comparable. Another option is to pool candidates and let a reranker evaluate their relevance to the query.
The hidden cost: tuning and diagnosis
Hybrid search creates more control surfaces: lexical analyzers, synonym rules, vector models, candidate counts, fusion logic, metadata filters, and reranking. That complexity is manageable, but only with a representative evaluation set.
A poor fusion strategy can erase the intended benefit. Overweight lexical retrieval and the system misses paraphrases. Overweight vectors and identifiers disappear beneath conceptually similar prose. Domain-specific tokenization also matters. A generic analyzer may split a component name, statute reference, or chemical notation in ways that destroy its retrieval value.
Hybrid search is usually the most defensible production baseline for enterprise knowledge: policies mixed with acronyms, documentation mixed with code symbols, or case records mixed with account numbers.
GraphRAG: when the answer lives between documents
GraphRAG represents knowledge as nodes and edges. Nodes may be people, organizations, systems, incidents, products, or claims. Edges encode relationships such as owns, depends on, reported by, or supersedes. Retrieval can begin with entities mentioned in a query, traverse relevant connections, and bring supporting source passages into the model’s context.
Return to the Project Alder question. A graph can connect Project Alder to Supplier North, connect Supplier North to its subsidiary North Systems, and connect that subsidiary to an unresolved exception. The useful evidence is a path, not merely a similar paragraph.
Graph structures also support corpus-level questions such as “What themes connect the incidents affecting our payment systems?” One implementation may cluster densely connected entities and generate summaries for those communities. The system can retrieve a high-level map first, then descend into supporting records.
Structure must be earned
GraphRAG is not simply vector RAG with a graph database attached. It requires a trustworthy ontology, entity extraction, entity resolution, relation extraction, provenance, and update logic. “North,” “North Systems,” and “North Systems Ltd.” may need to resolve to one entity—or remain distinct for legal reasons.
Incorrect edges are particularly dangerous because they can create persuasive chains of false association. Every extracted relationship should retain provenance to the source passage, and generated answers should cite that evidence rather than treating the graph as unquestionable truth.
Use GraphRAG when multi-hop relationships are central to the product and valuable enough to justify explicit modeling.
A worked comparison: investigating a service outage
Imagine an operations assistant answering: “Which recent deployments could have contributed to checkout failures, and which teams own the dependencies involved?” The corpus contains deployment notes, incident timelines, service ownership records, and dependency documentation.
- Vector RAG retrieves passages about checkout failures and recent releases. It may answer well if an incident review already narrates the relationship. It may miss a dependency whose documentation uses different language.
- Hybrid search preserves exact service names, release identifiers, and error codes while finding semantically related incident notes. It is stronger when the user supplies operational tokens such as CHK-219 or HTTP 502.
- GraphRAG can traverse from checkout to payment orchestration, from that service to a shared identity dependency, and from each component to its owning team and deployments. It is strongest when no document assembles the whole causal neighborhood.
GraphRAG still does not prove causality. A deployment connected to an affected service is a candidate, not a cause. The answer should distinguish recorded evidence, inferred relevance, and unresolved hypotheses. Retrieval architecture can improve discovery; it cannot replace epistemic discipline.
Quality, freshness, and governance
Vector and hybrid indexes are comparatively direct to refresh. When a document changes, its affected chunks can be re-embedded and reindexed. GraphRAG adds derived state: entities may merge or split, edges may expire, and community summaries may become stale.
Deletion is equally important. Removing a source document should remove or invalidate claims derived from it. In a graph, that may affect edges, paths, summaries, and cached answers. Provenance must therefore be a first-class field rather than an afterthought.
Access control also changes retrieval. A graph path can leak the existence of a restricted entity even when the source text is hidden. Permission filtering must apply to nodes, edges, evidence, and generated summaries. Vector and hybrid systems face similar concerns, but graph traversal enlarges the surface on which indirect disclosure can occur.
How to choose without overbuilding
- Start with answer shape. If answers usually come from one or two passages, begin with vector RAG. If exact strings determine relevance, begin with hybrid search. If answers require traversing relationships, test GraphRAG.
- Build an evaluation set from real work. Include paraphrases, identifiers, ambiguous entities, multi-hop questions, stale records, and questions that should not be answered.
- Measure retrieval before generation. Verify whether the necessary evidence appears in the candidate set. A stronger generator cannot reliably repair absent evidence.
- Add structure only where it pays. A graph may cover a relationship-rich subset of the corpus while hybrid search handles everything else.
Pick vector RAG for an early product, a meaning-rich document collection, or a workflow where local passages usually contain complete answers.
Pick hybrid search for technical, legal, operational, or enterprise repositories where exact identifiers and semantic intent matter together. For many production systems, this is the most resilient default.
Pick GraphRAG for investigations, dependency analysis, intelligence synthesis, and other products where the decisive knowledge exists in connections across records.
The deeper distinction is representational. Vector RAG preserves resemblance. Hybrid search preserves resemblance and exact language. GraphRAG preserves relationships. Choose the one that retains what your users cannot afford to lose.
This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.
From our own rounds
Measured on The Curator, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 169
- Questions per round
- 1.7
Rate this article
Discussion
Comments are moderated. Read our editorial policy.