all posts
research · sep 2026

The 2026 agent-memory landscape: what we learned, and what MemOS will adopt

We spent a week reading the 2025–2026 memory wave end to end — the academic papers, the vendor reports, the new benchmarks, and the compression literature. This post is the full briefing: what the field converged on, where the benchmarks are breaking, and the concrete list of what's coming to MemOS because of it.

Five paradigms now define agent memory

The 2025–2026 research wave split into five identifiable camps. Every serious memory system you'll evaluate today descends from one of them, usually with the graph layer bolted on:

  • Memory-as-OS (Letta/MemGPT). The original insight — page memory in and out of context like an operating system — aged well. Letta's 2026 benchmarking even argues a plain filesystem is surprisingly competitive, which is a useful corrective against over-engineering. Their sleep-time compute paper is the more interesting contribution: agents that "think" offline about stored context before any query arrives cut test-time compute ~5× at equal accuracy.
  • Extract-and-retrieve (Mem0). An LLM extracts discrete facts from conversations, stores them, and retrieval fuses multiple signals — Mem0's 2026 state-of-the-industry report describes three parallel scoring passes (semantic, BM25, entity match) normalized into one fused score. Their published numbers: 92.5 on LoCoMo and 94.4 on LongMemEval at roughly ~6,900 tokens per query.
  • Temporal knowledge graphs (Zep/Graphiti). Model change over time as an evolving graph rather than replacement. Research-driven, and the reason Zep's temporal reasoning numbers look the way they do.
  • Neuroscience-inspired retrieval (HippoRAG 2). "From RAG to Memory" (ICML 2025) frames retrieval as non-parametric continual learning: a knowledge graph plus Personalized PageRank integrates factual, sense-making, and associative memory tasks — comprehensively beating standard RAG.
  • Self-organizing memory (A-Mem). A-MEM (NeurIPS 2025) applies the Zettelkasten method: memories generate their own contextual descriptions, link to related notes, and evolve as new memories arrive — the store restructures itself instead of sitting still.

Two more matter for completeness: MIRIX (six specialized memory types — core, episodic, semantic, procedural, resource, vault — with a meta-manager routing retrieval), and MemTensor's MemOS (unrelated to us despite the shared name — their MemCube abstraction and "LLM as kernel" scheduling led to EverMemOS at ACL 2026). The 2026 trend line: graph layers became table stakes (Mem0 added graph memory, Cognee is graph-centered, Graphiti went standalone), and offline consolidation is merging with memory formation.

takeaway
MemOS (ours) already ships the local-first versions of three of these ideas — a SQLite-native graph with temporal validity and trust scoring, hybrid retrieval, and TOON-packed context budgets. The gaps the research exposes are consolidation, richer pool structure, and retrieval upgrades — that's the roadmap below.

The platform vendors picked a side — and it favors tools

The consumer architectures diverged in a way Simon Willison documented nicely: ChatGPT preloads a server-side user profile into every conversation; Claude starts every session from a blank slate and consults memory on demand via a file-based memory tool — plain CRUD on files, combined with automatic context editing. OpenAI's own Agents SDK cookbook likewise treats memory as "managing what's stored, recalled, and injected into working memory" — an application concern, not a model feature. Neither vendor's memory is available as infrastructure.

That's the structural opening local-first memory layers sit in: the platforms gave agents memory tools, and somebody has to implement the store behind the tool. Ours answers with a SQLite file you own, an MCP server any harness can mount, and a Claude Code plugin that injects context at session start.

The benchmarks are breaking — and that's informative

Three things happened to memory evaluation in 2026, and all three matter to how you should read anyone's leaderboard claims (including ours):

  • LoCoMo is saturated and being gamed. With million-token context windows, naive "dump everything in context" solves most of it; at least one project claimed a 100% score, drawing justified community skepticism. A leaderboard number on LoCoMo no longer discriminates.
  • LongMemEval became the gold standard — the harder, six-category evaluation where the best published result sits around 94.4 (Mem0's April 2026 algorithm). And the scale cliff is real: Mem0's own BEAM results drop from 64.1 at 1M tokens to 48.6 at 10M — a ~25-point fall that they correctly attribute to temporal abstraction. Long histories still break everyone.
  • LongMemEval-V2 changed the subject. The new benchmark tests memory for web agents over 25M–115M-token environment trajectories — workflow knowledge and "environment gotchas" instead of chat recall. The results are humbling: naive RAG manages 38–51%, frontier parametric knowledge alone gets 14.1%, and the best system (a coding agent acting as memory controller over raw slices, event pools, and distilled runbook notes) reaches ~72.5%.

The V2 design lessons are the most actionable findings in the literature right now:

  • Store multiple granularities. Removing the raw-state pool crashed one system's accuracy from 0.661 to 0.286. Raw events, transition events, and distilled notes each earn their keep.
  • Distilled notes beat raw retrieval. Workflow documents (procedural notes) were the single most valuable addition in their ablations.
  • Multi-stream retrieval wins. Separate queries per memory pool, generated by an LLM controller, beat single-query RAG by a wide margin.
  • Evidence slicing helps reading. Radius-1 windows around key states substantially improve how well a model can use long traces.

Token compression split into hard, soft, and agentic

The compression literature now cleanly divides into four families, and the split matters for what a local-first stack can adopt:

  • Hard compression (token selection) is production-ready: the LLMLingua family scores tokens with a small model and drops the low-information ones; LLMLingua-2 is 3–6× faster than its predecessors. RECOMP takes the text-level version of the same idea — compress retrieved documents into extractive or abstractive summaries before injection.
  • Soft compression (latent/gist tokens) reaches up to 26× (gist tokens) and 2026 added latent-space variants — K-Token Merging, CoLaR, LCLMs — but requires trained adapters and model cooperation. Not viable for a drop-in layer serving arbitrary models.
  • KV-cache compression (RocketKV, ChunkKV) compresses the cache instead of the prompt — orthogonal, inference-engine territory.
  • Agentic context compression is the hot 2026 survey topic: compressing observations and agent trajectories, not single prompts — exactly the failure mode LongMemEval-V2 exposes.
takeaway
For a local-first memory layer, the practical recipe is hard compression + structured formats + budget-aware packing — no trained adapters required. That's why MemOS bets on compact TOON context packs (77.6% smaller than JSON at equal fidelity, verified in our token-efficiency benchmarks) plus token-budgeted retrieval. LLMLingua-style filtering slots in cleanly as an optional stage; soft compression doesn't.

Retrieval quality has known, cheap upgrades

The most cited engineering result of the period is Anthropic's contextual retrieval: prepend LLM-generated context to each chunk before embedding, and retrieval failures drop 35–49% — up to 67% when combined with reranking. Its mirror image is late chunking (embed the long context first, pool per chunk afterward). On the model side, Qwen3-Embedding-0.6B and Qwen3-Reranker-0.6B put 2025-MTEB-leading quality on a laptop, which matters to us specifically: the two-stage retrieve-then-rerank architecture everyone converges on no longer requires cloud APIs. And on the query side, the community consensus is that HyDE and multi-query expansion trade latency for recall with real but diminishing returns — worth having, cheap to toggle.

What's changing in MemOS

Everything below is scoped against our existing architecture (SQLite storage, hybrid keyword/semantic retrieval, graph edges with temporal validity and trust scores, TOON context packs, the MCP surface). Items 1–4 and 6 are already shipped — pools, consolidation with decay-forgetting, entity-fused scoring, write-time enrichment, and multi-stream packs landed together with this post; the rest is scoped in priority order.

1. Multi-granularity pools (LongMemEval-V2's top lesson) — shipped ✅

Raw event memories, transition events (state A → state B), and distilled notes as three queryable pools inside one store. Today MemOS keeps one node stream plus graph relations; the V2 ablations show the note pool and event pool each carry independent accuracy. Our consolidation pass (below) is what feeds the note pool.

2. Sleep-time consolidation — shipped ✅

An idle-time job — think memos consolidate --while-idle — that re-reads recent episodes, resolves contradictions (writing the superseded node's validTo instead of deleting — our temporal layer already models this), merges near-duplicates via the existing semantic-dedup path, distills procedural notes, and applies algorithmic forgetting with a decay curve. Letta's sleep-time compute shows this shift from query-time to idle-time is a Pareto improvement, and every ingredient already exists in the SDK.

3. Entity-linked fusion scoring — shipped ✅

Mem0's 2026 pivot is telling: they replaced a separate graph store with entities in a parallel collection that boost retrieval scores, at the cost of losing relation traversal. We don't have to make that trade — our graph already persists relations. The upgrade is fusing a third signal into scoring: entity-match (and eventually a bounded Personalized-PageRank expansion over graph neighbors, HippoRAG-style) alongside the existing keyword + semantic passes.

4. Contextual memory enrichment — shipped ✅

At write time, prepend a one-line LLM-generated context to each memory ("this was stated while debugging the deploy pipeline") so embeddings carry situational meaning. This is the single best-attested retrieval upgrade in the literature (−35–49% failures), and with a local 0.6B model it stays fully offline.

5. Local default embedder + reranker

Make Qwen3-Embedding-0.6B + Qwen3-Reranker-0.6B the batteries-included local defaults (LFM2.5 and bge-reranker-v2-m3 remain supported). Both run on CPU/laptop GPU via llama.cpp, and the two-stage architecture is now the consensus quality floor.

6. Multi-stream retrieval — shipped ✅

When a context pack is requested, issue separate pool-specific queries (events / notes / procedures) and fuse — the V2 result that beat single-query RAG by ~20 points. Our context-pack API already owns the budget arithmetic; this changes where candidates come from, not the contract.

7. LLMLingua-style filtering as an opt-in stage

For aggressive token budgets, an optional hard-compression pass over the assembled pack, downstream of TOON packing. TELeR-style tiered summaries for tool outputs are in the same bucket — this composes with LLM Guardian's pipeline rather than duplicating it.

8. Benchmarks: scale, provenance, and the agentic turn

Three commitments. First, extend BEAM beyond 1M toward the 10M regime where everyone (us included) falls off a cliff — temporal abstraction is the honest frontier. Second, keep publishing provenance: our benchmark page now auto-regenerates a "verified runs" table from committed result JSONs, and LoCoMo's gaming episode is exactly why we'll keep doing that. Third, treat LongMemEval-V2 as the target shape for an agentic-memory eval: environment trajectories, workflow knowledge, gotcha recall — chat-recall benchmarks alone no longer prove an agent memory works.

The local-first thesis got stronger

The quiet pattern across everything we read: every technique that mattered in 2025–2026 — graph memory, two-stage rerank, contextual enrichment, consolidation, hard compression — works on small local models and a SQLite file. The cloud memory platforms' structural advantages (hosted vector DBs, server-side profiles) turned out not to be where accuracy comes from. Accuracy comes from pipeline design, and the pipeline now fits on a laptop.

Sources