Ask your assistant a real operational question: “Which of our services is exposed to CVE-2026-0142, and who’s on call for it right now?” It comes back with a tidy paragraph describing the vulnerability, and a vague line about the on-call rotation. It never connects the two. It never actually answers you.
Every fact you need is already in your docs. The security advisory behind that CVE number, which is just the public ID for a known security flaw, names the vulnerable library, libfoo 2.3. A service manifest lists libfoo as a dependency of billing-api. An ownership file maps billing-api to Team Payments. The on-call schedule says Dana is holding the pager this week. Four documents, four facts, and the answer only exists once you walk the chain from the CVE to Dana. Your vector search read all four and handed back the two that happened to look like your question.
Last time, we chased a different miss: the answer sat in one passage, and retrieval just failed to surface it. The fix was to patch the pipe. This is not that. This time the answer isn’t in any single passage at all, and no better search model, no smarter re-sorting of the results, no bigger model will conjure it. The problem isn’t retrieval quality. It’s that you’re asking a question your retrieval method structurally cannot answer.
First question: what shape is your question?
The instinct in RAG (retrieval-augmented generation, the standard setup where an assistant searches your documents and feeds the best hits to the model before it answers) is to treat every miss as a search-quality problem. Reach for a better model, slice the documents up differently, bolt on a step that re-sorts the results. But before any of that, ask a cheaper question: what shape is this question? Because different shapes need genuinely different retrieval, and no amount of tuning moves one shape into another.
There are three shapes, and only the first is vector search’s home turf. A single-fact question (“what’s the rate limit on /v2/upload?”) has its answer sitting in one passage; find the passage, you’re done. A multi-hop question (“who’s on call for a service hit by this CVE?”) has an answer that only exists as a chain of facts spread across documents. A global question (“what themes recur across this quarter’s postmortems?”) has an answer that lives in the whole corpus at once, in no particular passage. Same search box, three completely different jobs.
Why similarity can’t follow a chain
To see why the CVE question breaks, look at how vector search actually scores. Your documents were sliced into chunks, a few paragraphs each, and every chunk was turned into an embedding: a long list of numbers that places it at a point in space, positioned so that similar meanings land near each other. Your query gets turned into a point the same way. Then every chunk is ranked independently by how near it sits to the query. Independently is the load-bearing word. Nothing in the process represents a relationship between two chunks.
Now watch the chain fall apart. Your query mentions a CVE and on-call, so the chunks about the advisory and the rotation rank high; they resemble your words. But the bridge facts, the ones that say billing-api depends on libfoo and billing-api is owned by Team Payments, mention neither the CVE nor on-call. They’re dull dependency and ownership records. They don’t resemble your question, so they rank near the bottom and never make the candidate set. And without those bridges, there is no path from the CVE to Dana. Retrieval hands the model the two ends of a broken chain and the model, dutifully, can’t join them.
This is why multi-hop benchmarks like HotpotQA exist as a separate category at all: they deliberately require joining evidence across documents, precisely the thing single-passage retrieval can’t do. A sharper embedding might pull one more bridge into range, but it can’t create the missing idea, that these independently-ranked chunks are linked.
Why similarity can’t summarize a whole corpus
The global question fails for a different reason. Ask “what themes recur across this quarter’s postmortems?” and there is no set of twenty passages that answers it. Vector search pulls the top twenty of, say, ten thousand chunks, and those twenty are just the twenty nearest neighborhoods to your query, not a representative sample of the corpus. They might all come from the same loud incident. Retrieval throws away 99.8% of your documents by design, then you ask it a question about all of them.
That’s not really a retrieval problem, it’s a summarization problem over everything. Microsoft’s original GraphRAG work framed exactly this: for global sensemaking questions over corpora in the million-token range, it reports substantial gains over ordinary RAG in the comprehensiveness and diversity of the answers, because it stops pretending a handful of nearby chunks can speak for the whole.
The fix: store the connections, not just the text
If the problem is that plain retrieval has no notion of connections, the fix is to build them. That’s what GraphRAG does. At index time, a large language model (LLM) reads each chunk and extracts two things: the entities (the nouns that matter: CVE-2026-0142, libfoo, billing-api, Team Payments) and the relationships between them (the verbs: affects, depends on, owned by). Wire those together across your whole corpus and you have a knowledge graph, a web of facts that remembers how they connect.
Now the CVE question has a path to walk. A query does local search: it seeds on an entity, found by ordinary vector match, so the graph still leans on embeddings to get started, and then it traverses, following edges out to the neighbors. From CVE-2026-0142 it hops to libfoo, to billing-api, to Team Payments, to Dana. The exact bridges plain vector dropped are now first-class edges the search follows on purpose.
”Global search” is not what it sounds like
The whole-corpus question gets a cleverer treatment, and the name oversells what it actually does. GraphRAG groups the graph into communities, clusters of tightly-connected entities, using a standard algorithm (Leiden). Then, still at index time, it writes an LLM summary of each community. When a global query arrives, it map-reduces: it answers your question against each community summary in parallel (the map), then combines those partial answers into one (the reduce).
The code tells a plainer story. Global search does not traverse the graph at query time. It reads summaries that were written in advance and stitches them together. That’s a genuinely good design for “what are the themes,” but it means the corpus-wide view is only as fresh as your last indexing run, and every community summary was another LLM call you paid for up front. Keep that in mind before you assume “global” means the graph is being explored live.
The honest cost
Both stories keep landing on the same line: “at index time, an LLM call per chunk.” That’s the catch nobody puts on the slide. Building the graph means running a language model over every chunk of your corpus to pull out entities and edges, sometimes with follow-up passes to catch what it missed. Change your docs and you re-extract the affected parts. Want global search too? Add another LLM pass to summarize every community. Call it the indexing tax: real money and real time spent before a single question gets answered, on a scale a plain vector index never approaches. Microsoft puts a number on it in their own docs, estimating graph extraction at roughly 75% of indexing cost, and they ship a cheaper route (graphrag index --method fast) that swaps the LLM entity pass for plain noun-phrase extraction with NLTK or spaCy. You get the bill down; you also get a noisier graph with no descriptions on its entities and edges. That trade is fine if you only want global summaries, and wrong if you wanted the graph itself to be worth something.
The good news is the field is actively cutting that tax. LazyGraphRAG defers the heavy extraction to query time and, by Microsoft’s own numbers, drops indexing cost to 0.1% of full GraphRAG, the same order as a plain vector index, while matching its quality. HippoRAG reports up to 20% better multi-hop accuracy while being many times cheaper than iterative-retrieval approaches, LightRAG supports incremental updates so a changed doc doesn’t trigger a full rebuild, and newer agentic systems like GRASP push multi-hop accuracy up while spending 40 to 50% fewer tokens than the earlier graph pipelines. The trend line is clear: graphs are getting cheaper to run. That leaves the original in an odd spot. Microsoft’s graphrag repo now carries a maintenance-mode notice, bug fixes and dependency updates only, no new features. Read it to learn the mechanism, which is what we did here, but treat the successors as the live options.
But the honest default hasn’t changed. Build a graph only when the question shape demands it. If your users mostly ask single-passage questions, a documentation lookup, an API reference, an FAQ, then a tuned hybrid retriever is cheaper, faster, and far simpler. Hybrid means running two searches side by side, plain keyword matching (BM25, which just matches literal terms) and vector search, then letting a second model re-sort the combined list, which is the pipe we took apart in the retrieval piece. A graph on top of that just adds cost and staleness for no gain. Don’t put a knowledge graph under an FAQ.
Passages versus connections
Vector search was never broken. It answers “which passages resemble this?” beautifully, and for the huge share of questions whose answer sits in one passage, that’s the whole job. It just can’t answer “which facts connect?” or “what does the whole corpus say?”, because neither of those is about resemblance. One is about paths, the other is about coverage, and no embedding, however good, turns similarity into either.
So the job isn’t picking the best retriever. It’s reading the shape of the question first, then reaching for the machinery that shape needs: hybrid for the single passage, a graph for the chain, community summaries for the whole. That judgment, matching retrieval structure to question shape instead of forcing everything through one-size vectors, is the retrieval layer we build at Gracient, so the answer that was spread across five of your documents is the one your users actually get. Let’s talk.