How MemOS works
A layered, local-first memory engine. Everything runs against a single SQLite file on your disk — no vector database, no cloud, no API keys.
Layered architecture
Retrieval pipeline — hybrid search
Every search() runs two legs in parallel against the same SQLite database, then fuses the rankings.
SQLite FTS5 with BM25 ranking — exact terms, names, and identifiers.
Cosine similarity over the embeddings table. Vectors come from the local embedding provider — no network calls.
Reciprocal Rank Fusion (K=60) merges both legs, weighted, then trust-weighted and filtered by temporal validity.
Context packs — feeding LLMs on a token budget
contextPack() builds a token-budgeted slice of memory for prompt injection. LLM Guardian consumes the envelope directly as a high-relevance shard.
The retrieval pipeline above returns ranked candidates for the query.
Filler is stripped from each candidate before it can spend tokens.
Candidates are ranked by relevance × trust and cut at the token budget.
JSON, TOON, or TOON-compact — compact TOON is ~77.6% smaller than JSON for a 20-entry pack. Envelope: ai-trio.memos.context-pack.v1.
Module map
| Module | Responsibility |
|---|---|
| memory.ts | Public API (MemOS), orchestration, hybrid search + RRF |
| storage/sqlite.ts | Persistence via better-sqlite3 — WAL, FTS5, embeddings table |
| graph.ts | Graph operations, text similarity, clustering |
| context-pack.ts | Context pack builder, TOON / TOON-compact serialization |
| confidence-machine.ts | Confidence score state machine (floor 0.3, cap 1.0) |
| embedding-queue.ts | Batched async embedding writes |
| importance.ts | Effective importance — recency decay + access reinforcement |
| mcp.ts | MCP server exposing MemOS tools (stdio + HTTP/SSE) |