all posts
retrieval · sep 2026

RAG is not memory: what the leaderboards teach about hybrid retrieval

Everyone bolted a vector database onto their agent and called it memory. The benchmark ablations tell a more specific story — about which signals actually matter, and why retrieving memories is a different job than retrieving documents.

The +9pp ablation that settled it

The single most replicated finding in memory retrieval: dense vectors alone lose to dense + sparse. Adding BM25 keyword search alongside vector similarity is worth roughly nine points of recall on memory benchmarks — not because keyword search is smarter, but because the two signals fail differently. Vectors miss exact names, dates, and rare terms; keyword search misses paraphrase. Memories are full of both: "my sister" in one session is "Maria" in another, and only one of your two retrievers will catch each phrasing.

This is why the serious systems all run multi-signal retrieval now. Mem0's 2026 pipeline fuses three parallel passes — semantic, BM25, and entity match — into one normalized score. Zep fuses cosine similarity, BM25, and graph traversal. The consensus first stage is hybrid dense + sparse, full stop.

Reranking: the cheapest +3.4pp in the field

The second robust finding: a cross-encoder rerank over the fused candidates is worth about three and a half points, and it's the rare improvement that's both large and boring. The first stage is optimized for recall — don't miss anything. The reranker is optimized for precision — put the right thing on top. Splitting those two jobs across two stages beats any single ranker trying to do both, and the reranker only ever sees a few dozen candidates, so it's cheap.

If you take one engineering lesson from this post: turn reranking on before you tune anything else. It outperforms most embedding model upgrades at a fraction of the effort.

marginal retrieval gains · percentage points of recall
Add BM25 to dense vectors+9.0pp
hybrid dense + sparse — the consensus first stage
Add cross-encoder rerank+3.4pp
cheapest gain in the field
Typical run-to-run noise±2pp
a +1.5pp “improvement” is luck

Magnitudes from reported ablations across memory benchmarks; bars scaled illustratively. Entity linking adds further headroom on top of dense+sparse — the signal most DIY systems skip.

Entities: the signal everyone underuses

The third signal — entity linking — is the one most DIY memory systems skip, and it's where the biggest headroom is. Resolving "my sister," "Maria," and "she" to one entity, then boosting memories attached to the entities in the query, consistently lifts recall beyond what dense+sparse achieves alone. It's the bridge between the vector world and the graph world: entities are what let retrieval follow relationships instead of just matching text.

takeaway
Memory retrieval is evidence assembly, not ranking. A document search returns the best page; a memory query must assemble a case — the fact, its provenance, its validity window, and the entities that connect it to the question. Optimize for assembly and the ranking takes care of itself.

Why documents and memories need different retrieval

RAG retrieves documents: self-contained, authored once, ranked by relevance to a standalone query. Memories are none of those things. They're fragments — half a sentence from March that only makes sense next to a correction from June. They contradict each other across time. They reference entities by nickname. A ranker trained on document relevance will happily return the March fragment without the June correction, because it doesn't know the correction exists.

That's the real argument of this post: the failure mode of treating memory as RAG isn't low recall, it's confident staleness. You retrieve the right-shaped memory from the wrong time, the LLM answers fluently, and nobody notices. Temporal validity filtering and contradiction edges aren't embellishments — they're what make the difference between a search engine and a memory.

figure · ranking vs. evidence assembly
document rag · rankingqueryrank documentstop hit: march fragmentno june correctionfluent, stale answerconfident staleness —the june correction was never retrievedmemory retrieval · evidence assemblyqueryhybrid fusionfts5 + vectors + entitiescross-encoder rerankevidence bundle:fact · provenance · validityentity linksgrounded answer

Our stack, concretely

Here's how MemOS instantiates the consensus, locally, with no API key:

  • FTS5 first. SQLite's full-text search is the keyword leg — fast, local, and free. It catches the names, dates, and exact terms that vectors fumble.
  • Embeddings as boost. Dense vectors (nomic-embed-text via Ollama, or BGE-family models) add the semantic leg for paraphrase. Optional and local — the system works without them, just less forgiving of rewording.
  • Entity-fused scoring. Query entities resolve against the memory graph, and memories attached to matched entities get boosted — the bridge from text matching to relationship following.
  • Trust-weighted. Every memory carries a trust score from provenance and corroboration, and it multiplies into the final ranking. A high-similarity memory from an unreliable source should lose to a medium-similarity one that's been corroborated twice.

And because retrieval quality is the thing we refuse to regress, it's gated in CI: a golden corpus with a committed baseline (recall@5 0.9583, MRR 0.9444), failing the build on drops. The leaderboard numbers are nice. The CI gate is the product.