We read every memory benchmark so you don't have to: LoCoMo and LongMemEval, honestly explained
Every memory vendor quotes a leaderboard number. Almost none of them quote it the same way. Here's what the two standard benchmarks actually measure — and the five methodology traps that make the numbers incomparable.
What LoCoMo actually measures
LoCoMo (Maharana et al., ACL 2024) is ten very long two-speaker conversations — the kind of meandering, months-long dialogue where a fact from session two matters in session thirty. Each conversation comes with QA pairs annotated by category and with the evidence spans that justify the answer, plus an event-summarization task.
The slice everyone reports is 1,540 questions across five categories: single-hop, multi-hop, open-domain, temporal, and adversarial. The adversarial category is the interesting one — it's designed to punish systems that retrieve something plausible-but-wrong, and it's also the category most often quietly excluded from vendor scoreboards. Remember that; it comes back later.
What LongMemEval actually measures
LongMemEval (Wu et al., ICLR 2025) is a different shape: 500 questions over roughly 115K tokens of chat history — about 40 to 48 sessions per user — spanning six capability categories, plus 147 abstention questions where the correct answer is "I don't know" because the history genuinely doesn't contain it.
Its official metric is end-to-end: the system retrieves memories, an LLM writes an answer from them, and GPT-4o judges whether the answer is correct. That judge step is doing a lot of hidden work, which brings us to the central confusion in this field.
Two different games with the same scoreboard
There are two completely different things people call a "benchmark score," and they get conflated constantly:
- Retrieval-only scoring. Did the evidence surface? No LLM involved — just recall@5 or recall@10 against the annotated evidence spans. A plain BM25+vector hybrid hits around 95% recall@5 here. It measures the retriever and nothing else.
- LLM-judge QA accuracy. The full pipeline: retrieve, generate an answer, let a judge model grade it. This measures the retriever plus the reader model plus the judge's mood. Mem0's self-reported 92.5 on LoCoMo and 94.4 on LongMemEval are this kind of number.
A 95% retrieval recall and a 94% judge accuracy sound comparable. They aren't. One is "did we find the needle," the other is "did a generous robot like our summary of the needle." When you see a leaderboard, the first question is always: which game was played?
The five methodology traps
Even within the same game, the numbers move for reasons that have nothing to do with memory quality:
- Judge generosity. LLM judges are lenient in ways that correlate with answer fluency, not correctness. Swap the judge or the grading prompt and the same system moves several points. Nobody publishes the prompt.
- Top-k games. Retrieval recall is a function of k. One vendor's "100%" became 60.3% when a third party re-ran it at R@10 instead of a larger k. Always ask what k was.
- Category exclusions. Remember LoCoMo's adversarial category? Mem0's published breakdown excludes hundreds of adversarial questions — the hardest slice, where systems hallucinate confident answers. Dropping the category where you're weakest is not a benchmark result; it's marketing.
- Run noise. Re-runs of the same system on the same benchmark vary by about ±2 points. A "+1.5pp improvement" in a blog post is indistinguishable from luck. (One lever that is real: turning reranking on is worth about +3.4pp — more on that in a future post.)
- The reader wall. Around 94% on LongMemEval, gains stop coming from retrieval at all — about 15 of the 500 gold answers are broken or ambiguous, so the ceiling is the benchmark's own noise floor.
The honest version of the leaderboard, with all of this priced in: Mem0 at 92.5/94.4 (self-reported, managed platform), Exa's M-1 at 96.4 on LongMemEval, EverMemOS at 93.05/83.00, MindMemOS at 94.03 on LoCoMo — and then MemoryAgentBench, a harder neutral test from ICLR 2026, where Mem0 and Zep collapse to 21.1 and 24.0. Saturated benchmarks measure saturation, not utility.
The top four are saturated-benchmark numbers — judge-scored, vendor-reported. The bottom two are from MemoryAgentBench (ICLR 2026), a harder neutral test. The cliff between them is the point: saturated benchmarks measure saturation, not utility.
Same systems, different methodologies, wildly different numbers. Before trusting a leaderboard entry, ask: which k, which judge, which categories were excluded?
What we do instead
We can't fix the field's methodology, but we can refuse to play the game. For MemOS we commit to three things:
- Publish both numbers. Retrieval-only recall and end-to-end judge accuracy, labeled as what they are, with k, the judge model, and the grading prompt stated. No mixing.
- Open the harness. Our benchmark scripts live in the repo (scripts/bench-locomo.ts, scripts/bench-longmemeval.ts) — anyone can re-run them and check our math.
- Gate quality in CI, not in blog posts. We keep a golden corpus (36 facts, 24 queries) with a committed baseline — recall@5 0.9583, MRR 0.9444 — and CI fails if retrieval regresses more than 0.05. A benchmark you can't silently regress is worth more than a benchmark you scored high on once.
Leaderboards aren't going away, and we'll keep publishing our numbers on them. But the number we actually run the project by is the one in CI: does retrieval still work after every commit? Everything else is commentary.