RAG has been one of the most important ideas of the LLM era, but as we move towards agents, I think we’re running into a limitation in the way we think about retrieval. RAG, meaning Retrieval-Augmented Generation, was supposed to solve one of the biggest problems with LLMs: they don’t know everything and hallucinate a lot.
The idea behind RAG is very simple. We give the model access to documents, retrieve the relevant information, put it into context, and let it answer. It worked quite well. But as I started building and working with RAG systems, I became increasingly interested in a problem underneath it all:
How do we know we retrieved everything we needed?

To understand this better, think about these two questions: “What does error code E104 mean?” and “How many error codes exist for battery failure?” The first is a relevance problem. Find the right chunk and answer, but the second is different. One needs to discover the complete set before one can count it. So one could retrieve five highly relevant chunks and produce a perfectly grounded answer, and could still be wrong because another error code was somewhere else. This is what I think is one of the biggest challenges in RAG today:
relevance isn’t completeness.
We’ve responded to this in several ways, such as better embeddings, hybrid search, reranking, query expansion, agentic retrieval, and, more recently, approaches like GraphRAG. GraphRAG is interesting because it introduces structure like entities, relationships, communities, and summaries – rather than relying purely on the semantic similarity that traditional RAG uses. GraphRAG is great when done right and can make certain multi-hop reasoning tasks easier, but there is a nuance with it – where does the graph come from?
For unstructured data, you need to extract the entities and relationships. You can build domain-specific pipelines, which can require substantial domain knowledge, or you can use an LLM to construct the graph. The latter is powerful, but it adds cost, latency, and potentially another hallucination or incompleteness layer. If the graph misses an entity or relationship, your retrieval system may never find information that exists perfectly in the original document.
This doesn’t mean GraphRAG is bad. It means the representation itself becomes part of the retrieval problem. Also, a graph traversal can be perfectly exhaustive over the graph while being incomplete with respect to the source.
This led me to experiment with Seshat: instead of treating retrieval as a single top-k search operation, I wanted a system that makes it possible to choose among different acquisition strategies depending on the task. Semantic / Hybrid seach when relevance is enough, structural inspection when document organization matters, and systematic traversal when the task requires coverage and completeness is not negotiable.
There is an obvious objection:
Why not just put the entire document into the context window?
Yes, that could work. For a small enough document, it might even be the simplest solution. But not every knowledge base is one small document. Real systems have thousands of documents, revisions, permissions, and constantly changing information. And giving a model everything doesn’t mean it will necessarily use everything correctly.
Which brings me back to the question that started this whole experiment:
What does “retrieval” actually mean?
Maybe retrieval isn’t simply putting a few chunks into a prompt. Maybe it is the broader process of acquiring the evidence required to perform an information task. Sometimes that’s three paragraphs. Sometimes it’s an entire section. Sometimes it’s every occurrence of something across documents and the context window is a capacity.
Retrieval is an information-acquisition strategy.
And I think we still have a lot to figure out there.

AI-Generated Image




