architecture

How MemOS works

A layered, local-first memory engine. Everything runs against a single SQLite file on your disk — no vector database, no cloud, no API keys.

Layered architecture

APPLICATION LAYEROpenAI adapterAnthropic adapterCustom agents / LLMsTRANSPORT LAYERPython HTTP server (FastAPI)MCP server (stdio + HTTP/SSE)TypeScript SDK (direct)CORE ENGINE (TYPESCRIPT)MemOS — public API + orchestration (src/memory.ts)Retrieval pipeline(FTS5 + semantic + RRF)Graph engineContext packs(TOON)Confidence machineSTORAGE LAYERSQLite — WAL · FTS5 · embeddings tablefuture: Postgres, Redis, Qdrant…

Retrieval pipeline — hybrid search

Every search() runs two legs in parallel against the same SQLite database, then fuses the rankings.

01
Keyword leg

SQLite FTS5 with BM25 ranking — exact terms, names, and identifiers.

02
Semantic leg

Cosine similarity over the embeddings table. Vectors come from the local embedding provider — no network calls.

03
RRF fusion

Reciprocal Rank Fusion (K=60) merges both legs, weighted, then trust-weighted and filtered by temporal validity.

Context packs — feeding LLMs on a token budget

contextPack() builds a token-budgeted slice of memory for prompt injection. LLM Guardian consumes the envelope directly as a high-relevance shard.

01
Hybrid search

The retrieval pipeline above returns ranked candidates for the query.

02
Debloat

Filler is stripped from each candidate before it can spend tokens.

03
Budget cut

Candidates are ranked by relevance × trust and cut at the token budget.

04
Serialize

JSON, TOON, or TOON-compact — compact TOON is ~77.6% smaller than JSON for a 20-entry pack. Envelope: ai-trio.memos.context-pack.v1.

Module map

ModuleResponsibility
memory.tsPublic API (MemOS), orchestration, hybrid search + RRF
storage/sqlite.tsPersistence via better-sqlite3 — WAL, FTS5, embeddings table
graph.tsGraph operations, text similarity, clustering
context-pack.tsContext pack builder, TOON / TOON-compact serialization
confidence-machine.tsConfidence score state machine (floor 0.3, cap 1.0)
embedding-queue.tsBatched async embedding writes
importance.tsEffective importance — recency decay + access reinforcement
mcp.tsMCP server exposing MemOS tools (stdio + HTTP/SSE)