Benchmarks
Full-dataset LoCoMo, HotPotQA multi-hop, BEAM-1M at production scale, and a 10k-memory haystack — local LFM2.5 embeddings plus a two-stage cross-encoder rerank, reproducible from the repo.
How we benchmark
One laptop — RTX 5050 GPU. Retrieval, reranking, and scoring all run locally; no cloud calls anywhere in the measurement path.
LFM2.5-Embedding-350M served by llama.cpp, plus a bge-reranker-v2-m3 two-stage cross-encoder rerank.
Direct evidence-ID matching — no LLM judge, no fuzzy matching. Ties are reported as ties; comparison baselines run at their default settings.
npm run bench:beamnpm run bench:locomo:retrievenpm run bench:toon
Verified runs
| Benchmark | Run date | Commit | Embedding model | Environment |
|---|---|---|---|---|
| LoCoMo · full dataset scripts/bench-locomo-noapi-results.json | 2026-09-06 | a6e1af9 | LFM2.5-Embedding-350M-BF16 | win32 10.0.26220 · v26.7.0 |
| BEAM-1M scripts/bench-beam-results-lfm2.5-350m.json | 2026-08-19 | — | — | — · — |
| Memory haystack scripts/bench-haystack-results.json | 2026-09-06 | 00afdb4 | local-hash-384 | win32 10.0.26220 · v26.7.0 |
| HotPotQA scripts/bench-hotpot-results.json | 2026-09-06 | 00afdb4 | LFM2.5-Embedding-350M-BF16 | win32 10.0.26220 · v26.7.0 |
Regenerated automatically from result JSONs committed to the repo — if a number on this page can't be traced to a run below, treat it as marketing.
LoCoMo · full dataset · retrieval-only evidence matching
| Metric | Top-5 | Top-10 |
|---|---|---|
| Hit@10 | 78.3% | 83.2% |
| Evidence Recall | 72.6% | 77.9% |
| All-Evidence Recall | — | 72.5% |
| MRR | — | 0.642 |
| nDCG@10 | — | 0.656 |
All 10 conversations, 1,986 questions, scored by direct evidence-ID matching — no LLM judge, no fuzzy matching. Stack: LFM2.5-Embedding-350M served by llama.cpp plus a bge-reranker-v2-m3 cross-encoder, both on one RTX 5050 laptop GPU. Per category: single-hop 89.7% Hit@10, multi-hop 86.9%, temporal 85.4%, adversarial 73.3%, open-domain 53.3%. Full run: 79.7 minutes end-to-end on that laptop.
HotPotQA · multi-hop retrieval
| Metric | Value |
|---|---|
| Hit@10 | 100% |
| Evidence Recall@10 | 96.3% |
| All-Evidence Recall@10 | 92.6% |
| MRR@10 | 0.977 |
| nDCG@10 | 0.934 |
500 validation questions, ~4,900 deduplicated Wikipedia paragraphs ingested corpus-wide, scored by supporting-paragraph title matching. Comparison questions are perfect across every metric; bridge questions sit at 95.4% evidence recall. This is the dataset class Cognee uses in its published memory evals.
Memory haystack · precision at scale
| Corpus size | Hit@1 | Hit@10 | p50 latency |
|---|---|---|---|
| 1,000 memories | 100% | 100% | 12 ms |
| 5,000 memories | 100% | 100% | 43 ms |
| 10,000 memories | 100% | 100% | 82 ms |
The needle-in-a-haystack test adapted to memory: unique facts planted at random depths in a growing distractor corpus, one natural-language query per needle. Deterministic, CPU-only — precision holds at 10k memories with an honest near-linear latency curve.
BEAM-1M · production scale
BEAM-1M simulates months of agent conversations (~1M tokens). The gap comes from temporal validity and trust scoring — features that only matter once history gets long and messy, i.e. production.
| Category | MemOS | Mem0 |
|---|---|---|
| Overall (recall@10) | 95.9% | 64.1% |
| Temporal Reasoning | 97.1% | 16.3% |
| Contradiction Resolution | 88.6% | 35.7% |
Shorter-horizon benchmarks — honest ties
On shorter histories both systems retrieve well — we report the ties as ties. The point of going local isn't beating Mem0 everywhere; it's matching the quality while owning your data and paying nothing.
Token efficiency
Every memory you inject into a prompt costs tokens. The compact TOON format carries the same facts in roughly a quarter of the space — that's real money back on every request.
| Format | Tokens · 20 entries | Savings vs JSON |
|---|---|---|
| JSON (full objects) | 5,479 | — |
| Verbose TOON | 1,522 | 72.2% |
| Compact TOON (new) | 1,229 | 77.6% |
Reproduce everything on this page: benchmark scripts and datasets live in the repo. Local embeddings (LFM2.5-Embedding-350M) and the reranker (bge-reranker-v2-m3) run on the same laptop via llama.cpp — no API keys, no cloud.