benchmarks

Benchmarks

Full-dataset LoCoMo, HotPotQA multi-hop, BEAM-1M at production scale, and a 10k-memory haystack — local LFM2.5 embeddings plus a two-stage cross-encoder rerank, reproducible from the repo.

How we benchmark

hardware

One laptop — RTX 5050 GPU. Retrieval, reranking, and scoring all run locally; no cloud calls anywhere in the measurement path.

retrieval stack

LFM2.5-Embedding-350M served by llama.cpp, plus a bge-reranker-v2-m3 two-stage cross-encoder rerank.

scoring

Direct evidence-ID matching — no LLM judge, no fuzzy matching. Ties are reported as ties; comparison baselines run at their default settings.

reproduce

npm run bench:beamnpm run bench:locomo:retrievenpm run bench:toon

Verified runs

BenchmarkRun dateCommitEmbedding modelEnvironment
LoCoMo · full dataset
scripts/bench-locomo-noapi-results.json
2026-09-06a6e1af9LFM2.5-Embedding-350M-BF16win32 10.0.26220 · v26.7.0
BEAM-1M
scripts/bench-beam-results-lfm2.5-350m.json
2026-08-19——— · —
Memory haystack
scripts/bench-haystack-results.json
2026-09-0600afdb4local-hash-384win32 10.0.26220 · v26.7.0
HotPotQA
scripts/bench-hotpot-results.json
2026-09-0600afdb4LFM2.5-Embedding-350M-BF16win32 10.0.26220 · v26.7.0

Regenerated automatically from result JSONs committed to the repo — if a number on this page can't be traced to a run below, treat it as marketing.

LoCoMo · full dataset · retrieval-only evidence matching

MetricTop-5Top-10
Hit@1078.3%83.2%
Evidence Recall72.6%77.9%
All-Evidence Recall—72.5%
MRR—0.642
nDCG@10—0.656

All 10 conversations, 1,986 questions, scored by direct evidence-ID matching — no LLM judge, no fuzzy matching. Stack: LFM2.5-Embedding-350M served by llama.cpp plus a bge-reranker-v2-m3 cross-encoder, both on one RTX 5050 laptop GPU. Per category: single-hop 89.7% Hit@10, multi-hop 86.9%, temporal 85.4%, adversarial 73.3%, open-domain 53.3%. Full run: 79.7 minutes end-to-end on that laptop.

HotPotQA · multi-hop retrieval

MetricValue
Hit@10100%
Evidence Recall@1096.3%
All-Evidence Recall@1092.6%
MRR@100.977
nDCG@100.934

500 validation questions, ~4,900 deduplicated Wikipedia paragraphs ingested corpus-wide, scored by supporting-paragraph title matching. Comparison questions are perfect across every metric; bridge questions sit at 95.4% evidence recall. This is the dataset class Cognee uses in its published memory evals.

Memory haystack · precision at scale

Corpus sizeHit@1Hit@10p50 latency
1,000 memories100%100%12 ms
5,000 memories100%100%43 ms
10,000 memories100%100%82 ms

The needle-in-a-haystack test adapted to memory: unique facts planted at random depths in a growing distractor corpus, one natural-language query per needle. Deterministic, CPU-only — precision holds at 10k memories with an honest near-linear latency curve.

BEAM-1M · production scale

recall @10 · million-token history · higher is better
MemOS95.9%
local SQLite + Gemma-300M embeddings
Mem064.1%
hosted platform, default settings

BEAM-1M simulates months of agent conversations (~1M tokens). The gap comes from temporal validity and trust scoring — features that only matter once history gets long and messy, i.e. production.

CategoryMemOSMem0
Overall (recall@10)95.9%64.1%
Temporal Reasoning97.1%16.3%
Contradiction Resolution88.6%35.7%

Shorter-horizon benchmarks — honest ties

recall · LoCoMo & LongMemEval
MemOS · LoCoMo92.5
Mem0 · LoCoMo92.5
MemOS · LongMemEval94.4
Mem0 · LongMemEval94.4

On shorter histories both systems retrieve well — we report the ties as ties. The point of going local isn't beating Mem0 everywhere; it's matching the quality while owning your data and paying nothing.

Token efficiency

tokens per 20 memory entries · lower is better
JSON (full objects)5,479 tok
Verbose TOON1,522 tok
Compact TOON1,229 tok
77.6% smaller than JSON

Every memory you inject into a prompt costs tokens. The compact TOON format carries the same facts in roughly a quarter of the space — that's real money back on every request.

FormatTokens · 20 entriesSavings vs JSON
JSON (full objects)5,479—
Verbose TOON1,52272.2%
Compact TOON (new)1,22977.6%

Reproduce everything on this page: benchmark scripts and datasets live in the repo. Local embeddings (LFM2.5-Embedding-350M) and the reranker (bge-reranker-v2-m3) run on the same laptop via llama.cpp — no API keys, no cloud.