Cortex

MCP serverDocs & knowledge

Cortex gives your AI a memory that lasts between conversations. It stores information in four tiers, keeps track of the people you mention, and adjusts its beliefs as new information arrives. Everything stays on your device, encrypted, with recall in microseconds.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Add cortex, then start a normal conversation with your AI. What you share gets stored and can be recalled in later conversations.

What your AI can do with it

  • Remember facts and context across conversations
  • Track people and how they relate to each other
  • Adjust beliefs as new information comes in
  • Sort memories into four tiers
  • Keep memories encrypted on your own device
  • Recall stored information in microseconds

From the project's README

As published by gambletan/cortex in README.md.

🧠 Try Cortex in your browser — zero install, 124KB WASM, runs entirely client-side.

If Cortex helps your AI remember, give it a ⭐ — it takes 1 second and helps others discover the project.

中文 | 日本語 | 한국어

Memory for AI agents that never leaves your device.

Private. Free. Local. — a memory engine for personal AI agents.

Your AI's memory lives on your device — your data never leaves, never costs, never spies. Pure Rust. 3.8MB binary. No third-party servers in the data path, zero telemetry, zero cost. Syncs through your own cloud storage. (On-device semantic search downloads a ~30MB model once on first use, then runs fully offline — or go 100% offline with CORTEX_NO_EMBEDDINGS=1. See Security & Privacy.)

What you get

  • 🔒 Private by default — memories live in a local SQLite file, never leave your device, zero telemetry (CI-enforced).
  • 🧠 Real memory, not a text file — 4 tiers, multi-signal retrieval, self-correcting Bayesian beliefs, a cross-channel people graph.
  • Sub-millisecond — 156µs ingest, 568µs search. ~528× faster than cloud memory APIs, with no network round-trip.
  • 🔌 Drop-in for any agent — one MCP server gives Claude Code / Claude Desktop (or any MCP client) persistent cross-session memory.
  • ☁️ Yours across devices — optional end-to-end-encrypted sync through your own iCloud / Drive / Dropbox. No server of ours, ever.

See it remember across sessions — ~30 seconds:

brew install gambletan/tap/cortex-mcp-server          # or: cargo build --release -p cortex-mcp-server
claude mcp add cortex-memory -- cortex-mcp-server ~/.cortex/memory.db

Tell Claude "remember I deploy on Fly.io and always run tests before pushing." Open a brand-new session and ask "how do I deploy this project?" — it answers from memory, 100% on your machine.

⭐ If that's useful, give it a star — it helps others find a memory engine that respects their privacy.


LLMs start blank every session — they forget your name, your preferences, yesterday's conversation, last week's decision. The usual fixes are flat text files (no ranking, no decay), keyword grep, or cloud APIs that add 200–500ms, charge you, and ship your personal data to someone else's server. Cortex gives your AI structured, self-evolving long-term memory that persists across sessions and channels — all local, all yours. Your memories are not a cloud provider's training data, a startup's monetization asset, or a surveillance target.

Cortex vs Mem0 vs OpenAI Memory

CortexMem0OpenAI Memory
Privacy100% local, zero cloudCloud API (your data on their servers)OpenAI servers
Latency156µs ingest, 568µs search~200-500ms~300-800ms
CostFree, forever$99+/mo (Pro)ChatGPT Plus ($20/mo)
Memory tiers4 (Working/Episodic/Semantic/Procedural)1 (flat)1 (flat)
Bayesian beliefsSelf-correcting with evidenceNoNo
People graphCross-channel identity resolutionPaid tier onlyNo
Conversation compressionAutomatic session summarizationNoNo
Relationship inferencePattern-based (EN + CN)NoNo
Temporal retrievalIntent-aware ("recently" / "first time")NoNo
Contradiction detectionAutomatic with confidence scoresNoNo
ConsolidationEpisodic → Semantic auto-promotionNoNo
Context injectionToken-budgeted LLM-ready outputManualAutomatic but opaque
Import/ExportFull JSON backup & restoreAPI onlyNo export
Self-hostedNative binary, Docker, MCPCloud onlyCloud only
Binary size3.8 MBnpm packageN/A
Dependencies0 runtime services (single binary)Node.js + cloudN/A
Open sourceMITPartialNo
EncryptionAES-256-GCM encrypted sync (opt-in)NoNo
Key rotationVersioned envelopes, forward secrecyNoNo
Privacy levelsPrivate (default, never syncs) / Shared / Public — per-memory opt-in, demote retracts from other devicesNoNo
Tool authorizationDeny-by-default capability policy on the MCP surfaceNoNo
Zero telemetryNo analytics, no phone-home, verifiableUnknownNo
CostFree forever, unlimited$99+/mo (Pro)$20/mo (Plus)
Chinese NLPNative (inference, retrieval, relationships)NoLimited
Namespace isolationPer-user/context memory separationNoNo
Plugin systemCompile-time hooks for ingest/retrieve/consolidationNoNo
MCP tools30 tools for Claude/LLM integration3rd partyN/A

Performance Benchmarks

OperationCortexMem0 (cloud)File-based
Ingest156µs~200ms~1ms
Search (top-10)568µs~300ms~10ms
Context generation621µs~500msmanual
Belief update66µsN/AN/A
People graph51µspaid tierN/A
Structured facts45µsN/AN/A
1K memories search1.6ms~500ms~50ms

528x faster than Mem0 cloud. With features neither Mem0 nor OpenAI Memory offer.

Note: Benchmarks include proactive inference (auto-extracting facts, preferences, relationships) on every ingest. Raw ingest without inference is ~15µs. Numbers from cargo bench on M-series Mac.

LoCoMo Benchmark (ACL 2024)

Academic-grade long-term conversation memory evaluation — 10 conversations, 1540 QA pairs across 4 categories.

SystemSingle-hopMulti-hopOpen-domainTemporalOverall
Backboard89.4%75.0%91.2%91.9%90.0%
MemMachine v0.284.9%
Cortex72.5%59.5%88.8%74.1%73.7%
Mem0-Graph65.7%47.2%75.7%58.1%68.4%
Mem067.1%51.2%72.9%55.5%66.9%
OpenAI Memory52.9%

Key findings:

  • Open-domain 88.8% — leads Mem0 (72.9%) by +15.9%
  • Temporal 74.1% — leads Mem0 (55.5%) by +18.6%
  • Single-hop 72.5% — leads Mem0 (67.1%) by +5.4%
  • Multi-hop 59.5% — leads Mem0 (51.2%) by +8.3%
  • Overall 73.7% — beats Mem0 (66.9%) by +6.8%, beats OpenAI Memory (52.9%) by +20.8%

Cortex outperforms Mem0 on all 4 categories — while running 100% locally, end-to-end encrypted, at $0 cost.

Setup: Claude Sonnet 4 (QA + judge), nomic-embed-text (embeddings via Ollama), top-30 retrieval. Reproducible with that setup: python3 bench/locomo_bench.py (needs ANTHROPIC_API_KEY + a local Ollama with nomic-embed-text). Numbers measured on the v1.7 engine; the v2.2 retrieval beam fix (paraphrase recall 40%→90% at 5K, see docs/scale-test-2026-06-13.md) has not yet been re-run on LoCoMo, so these are reported as the last verified figures, not a v2.2 claim.

Architecture

Cortex implements a 4-tier memory model inspired by human cognition:

                    +---------------------+
                    |   Working Memory    |  Current session context
                    +---------------------+
                              |
                    +---------------------+
                    |   Episodic Memory   |  Raw experiences: conversations, events, observations
                    +---------------------+
                              |  consolidation (decay, promotion, pattern extraction)
                    +---------------------+
                    |   Semantic Memory   |  Distilled facts, preferences, relationships
                    +---------------------+
                              |
                    +---------------------+
                    | Procedural Memory   |  Learned routines, user-specific workflows
                    +---------------------+

Working holds the current session scratch pad. Episodic stores raw experiences with timestamps and source metadata. The Consolidation Engine periodically promotes recurring patterns into Semantic facts and decays stale episodes. Procedural captures learned workflows and routines.

Key Components

People Graph

Cross-channel identity resolution. The same person messaging you on Telegram, emailing you, and showing up in calendar events gets unified into a single identity node. Interactions, relationship strength, and communication patterns are tracked per-person.

Bayesian Belief System

Self-correcting understanding of the world. Beliefs are formed from evidence, updated with each new observation, and can be contradicted. Confidence scores reflect actual certainty rather than recency bias.

cortex.observe_belief("user_prefers_morning_meetings", true, 0.8)?;
cortex.observe_belief("user_prefers_morning_meetings", false, 0.6)?;
// Confidence adjusts automatically via Bayesian update

Consolidation Engine

Episodic-to-semantic promotion, decay of stale memories, and pattern extraction. Runs as a background cycle that keeps the memory store lean and queryable. Returns a report of what was promoted, decayed, and merged.

Multi-signal Retrieval

Queries combine five signals for relevance ranking:

  • Similarity -- vector cosine distance against query embedding
  • Temporal -- recency weighting with configurable decay
  • Salience -- importance scoring from access patterns and explicit hints
  • Social -- boost for memories involving specific people
  • Channel -- filter or boost by source channel

Context Injection Protocol

Generates LLM-ready context strings from memory state. Pass a token budget, optional channel/person filters, and get back a structured text block your LLM can consume directly.

Storage

SQLite for persistence, in-memory vector index for fast similarity search. Single-file database, no external services required. Designed for edge deployment -- runs on a laptop, a Raspberry Pi, or a server.

Cloud Sync

Sync memories across devices through your own cloud storage — no third-party server involved.

Device A (Mac)              Your Cloud Storage              Device B (iPhone)
┌──────────┐         ┌──────────────────────┐         ┌──────────┐
│ SQLite DB │ ──W──>  │ iCloud / GDrive /    │  <──R── │ SQLite DB│
│ (local)   │         │ OneDrive / Dropbox   │         │ (local)  │
│           │ <──R──  │                      │  ──W──> │          │
└──────────┘         └──────────────────────┘         └──────────┘
  • Changelog-based: Each device writes append-only operation logs to its own subfolder
  • No conflicts: Devices never write to the same file. Merge uses Last-Writer-Wins with Hybrid Logical Clocks
  • Encrypted: AES-256-GCM encryption (opt-in). Even if your cloud account is compromised, memories stay private
  • Tamper-evident: the sync manifest and every operation carry an HMAC; tampered or plaintext-injected oplog lines are rejected, and a manifest without integrity protection refuses to load (no key-rollback path)
  • Key rotation & forward secrecy: rotate to a new key version (ENC2 envelopes) without re-encrypting history; old versions stay readable, new writes are unreadable to a leaked old key
  • Privacy-aware, per-memory opt-in: Private memories (the default) never leave your device. Mark a memory shared to sync it; demote it back to private and a retraction deletes it from your other devices (local copy kept)
  • Survives restarts: sync settings persist in the database (passphrase never touches disk — macOS login Keychain or CORTEX_SYNC_PASSPHRASE); the server resumes sync and starts background pull (30s poll + fs watcher) automatically

Supported providers: iCloud Drive, Google Drive, OneDrive, Dropbox (auto-detected).

use cortex_core::sync::SyncConfig;
use cortex_core::types::PrivacyLevel;

// Enable sync with encryption (settings persist; passphrase goes to the OS keychain)
let config = SyncConfig::new(sync_dir, device_id, device_name)
    .with_encryption("my-strong-passphrase");
cortex.enable_sync(config)?;

// Opt a memory into sync — everything is Private unless you say otherwise
cortex.set_memory_privacy(mem_id, PrivacyLevel::Shared { scope: "all".into() })?;

// Pull changes from other devices (also happens automatically in the background)
let applied = cortex.sync_pull()?;
println!("Applied {} remote changes", applied);

Security & Privacy

FeatureDetail
EncryptionAES-256-GCM with Argon2id key derivation (per-line random nonce)
Key rotationVersioned ENC2 envelopes with per-version passphrase-derived keys — forward secrecy against AES-key exfiltration, no full re-encryption needed
IntegrityHMAC on the sync manifest and on every sync operation; plaintext lines in an encrypted oplog are rejected outright (injection defense)
Privacy levelsPrivate (default, never syncs), Shared, Public — set at ingest (privacy arg / --privacy) or later (memory_set_privacy); demoting to Private retracts the memory from other devices
Capability policyDeny-by-default tool authorization on the MCP surface: a capabilities.json grants tool groups (read/write/sync/plugins) or exact tools; ungranted tools are invisible and uncallable; malformed policy fails closed
Query budgetEvery retrieval is bounded (candidate cap + wall-clock cap) — query cost never scales with total store size; DoS guard and timing-side-channel bound in one
Secret handlingSync passphrase is never written to disk by Cortex — macOS login Keychain or env var only; missing passphrase fails safe (sync off, never plaintext)
Memory zeroizationSensitive data cleared from RAM on drop (zeroize crate)
Zero telemetryNo analytics, no phone-home, no user data ever leaves the device — enforced in CI (scripts/check-no-network-egress.sh): the build fails if any network/telemetry crate enters cortex-core's default tree, and the check also proves the --no-default-features binary is completely zero-network.
Embedding model fetch (one-time)The default cortex-mcp-server enables on-device semantic search, which downloads a ~30 MB model (all-MiniLM-L6-v2) from the Hugging Face CDN on first ingest, then runs fully offline and sends none of your data. For a 100%-offline setup: run with CORTEX_NO_EMBEDDINGS=1 (keyword/FTS recall, zero network) or build --no-default-features. A one-time stderr notice is printed before any download — nothing is ever fetched silently.
No accountsNo API key, no registration, no cloud dependency

See SECURITY.md for the full threat model.

Prerequisites

Install the Rust toolchain (provides cargo):

curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh

After installation, either restart your terminal or run:

source "$HOME/.cargo/env"

Verify:

cargo --version

Real-World Example: A Personal AI That Actually Remembers

Imagine your AI assistant across a week of real conversations:

# Day 1 — You chat on Telegram
You: "Sarah works at Stripe. She's interested in our API."

  Cortex auto-extracts:
  ├── episodic memory stored (156µs)
  ├── fact: Sarah → works_at → Stripe (confidence: 0.70)
  └── person resolved: sarah_telegram

# Day 2 — Sarah emails you
From: sarah@stripe.com
"Here's the technical spec we discussed."

  Cortex:
  ├── person resolved: sarah@stripe.com → merged with sarah_telegram
  │   (same person, different channel — automatic identity resolution)
  └── fact: Sarah → sent → technical spec

# Day 3 — You ask your AI
You: "What's the status with Stripe?"

  Cortex retrieves (568µs):
  ├── Sarah works at Stripe (semantic fact)
  ├── Meeting went well, interested in API (episodic, Day 1)
  ├── She sent technical spec (episodic, Day 2)
  └── Cross-channel context: Telegram + Email unified under one person

  Your AI responds with full context — no "sorry, I don't remember" 🎯

# Day 5 — New information arrives
You: "Sarah now works at Anthropic."

  Cortex:
  ├── contradiction detected: Sarah works_at Stripe vs Sarah works_at Anthropic
  ├── old fact superseded + decayed: Stripe (salience ×0.3, kept as history)
  ├── new fact stored: Sarah → works_at → Anthropic
  └── current employer now ranks first; self-correcting, no manual cleanup

  (Third-party relations are extracted from natural-language verbs —
   "works at / works for / joined / now works at", "runs on", "hosted in",
   "manages", "part of", … — between two proper-noun entities.)

# Day 7 — Consolidation runs
  Cortex auto-consolidation:
  ├── 3 episodic memories about Sarah → promoted to semantic summary
  ├── stale memories from other topics → decayed
  └── pattern detected: you have recurring Monday meetings

All of this happens locally in <1ms per operation. No cloud. No API calls. No one else sees your data.

Install

Homebrew (macOS / Linux)

brew tap gambletan/tap
brew install cortex-mcp-server

From source

cargo build --release -p cortex-mcp-server
cp target/release/cortex-mcp-server ~/.local/bin/

Official packages (avoid look-alikes)

Cortex is published under the cortex-ai-memory name. Several similarly-named packages on npm/PyPI are not affiliated with this project — use exactly these:

EcosystemOfficial packageUse for
Binary / MCP serverGitHub Releases, or brew install gambletan/tap/cortex-mcp-serverthe memory engine (primary)
PyPIcortex-ai-memoryPython bindings
npm@cortex-ai-memory/cortex-memory (scoped)OpenClaw memory plugin

⚠️ Not us: npm cortex-mcp, npm cortex-ai-memory (unscoped), PyPI cortex-memory. The source of truth is always this repo — github.com/gambletan/cortex. When in doubt, the binary from Releases is the canonical install.

Quick Start

use cortex_core::Cortex;

// Open (or create) a memory database
let cortex = Cortex::open("memory.db")?;

// Ingest a memory from a Telegram conversation
let embedding = your_embedding_fn("Met with Alice about the Q3 roadmap");
cortex.ingest(
    "Met with Alice about the Q3 roadmap",
    "telegram",               // source channel
    Some("alice_123"),         // user ID (triggers identity resolution)
    Some(0.8),                 // salience hint
    Some(embedding),           // vector embedding
)?;

// Add a semantic fact directly
cortex.add_fact(
    "Alice", "works_at", "Acme Corp",
    0.95, "telegram", None,
)?;

// Store a preference
cortex.add_preference("timezone", "America/Los_Angeles", 0.9)?;

// Retrieve relevant memories
let results = cortex.retrieve(
    "What do I know about Alice?",
    5,                         // top-k
    None,                      // any channel
    None,                      // any person
    Some(query_embedding),     // vector for similarity search
)?;

// Generate LLM-ready context (token-budgeted)
let context = cortex.get_context(
    2000,                      // max tokens
    Some("telegram"),          // channel filter
    None,                      // no person filter
)?;
// Pass `context` as system/user message prefix to your LLM

// Run consolidation (call periodically)
let report = cortex.run_consolidation()?;
println!("Promoted: {}, Decayed: {}", report.promoted, report.decayed);

Python Bindings

Coming soon via PyO3. The cortex-python crate will expose the full API as a native Python module:

from cortex import Cortex

cx = Cortex.open("memory.db")
cx.ingest("Had lunch with Bob at the Thai place", channel="imessage", user_id="bob")
results = cx.retrieve("Where does Bob like to eat?", limit=5)

Integration with unified-channel-hub

Cortex is designed as the memory layer for unified-channel-hub. Messages flow in from any channel adapter, Cortex ingests and indexes them, and the context injection protocol feeds relevant memory back to your LLM before each response.

Telegram ─┐                          ┌─ Context
Discord  ─┤  unified-channel-hub  →  │  Cortex  →  LLM
Email    ─┤  (ingest)                 │  (retrieve + inject)
Calendar ─┘                          └─ Response

Integration with LangGraph

Add persistent memory to any LangGraph agent via langchain-mcp-adapters — no custom code needed.

from langchain_mcp_adapters.client import MultiServerMCPClient
from langgraph.prebuilt import create_react_agent
from langchain_openai import ChatOpenAI

model = ChatOpenAI(model="gpt-4o")

async with MultiServerMCPClient({
    "cortex": {
        "command": "cortex-mcp-server",
        "args": ["~/.cortex/memory.db"]
    }
}) as client:
    agent = create_react_agent(model, client.get_tools())
    # Agent now has all 30 Cortex memory tools
    result = await agent.ainvoke({
        "messages": [{"role": "user", "content": "What do you remember about Alice?"}]
    })

Your LangGraph agent gets instant access to memory_search, memory_ingest, fact_add, belief_observe, person_resolve, and 25 more tools — all running locally.

Integration with DeerFlow (ByteDance)

Cortex works as a persistent memory layer for DeerFlow — ByteDance's open-source multi-agent orchestration platform. Zero code changes needed.

# Add to DeerFlow config.yaml
mcp_servers:
  cortex-memory:
    command: cortex-mcp-server
    args:
      - ~/.cortex/deerflow.db

All DeerFlow agents (Telegram, Slack, Feishu) get instant access to 30 memory tools — cross-session memory, fact storage, people graph, and belief tracking across all channels.

CLI

Cortex doubles as a standalone CLI tool — no MCP client required.

$ cortex-mcp-server --help
Cortex memory engine — MCP server & CLI tools

Usage: cortex-mcp-server [DB_PATH] [COMMAND]

Commands:
  ingest  Store a new memory
  search  Search memories
  stats   Show memory statistics
  sync    Show cloud sync status and detected providers
  export  Export all data as JSON
  import  Import data from JSON file
  info    Show version, DB path, and capabilities
  help    Print this message or the help of the given subcommand(s)

Arguments:
  [DB_PATH]  Path to the Cortex database file (default: ~/.cortex/memory.db)

Options:
  -h, --help     Print help
  -V, --version  Print version

Examples:

# Store a memory
cortex-mcp-server ~/.cortex/memory.db ingest "Met with Alice about Q3 roadmap"
cortex-mcp-server ~/.cortex/memory.db ingest -c telegram "Sarah now works at Anthropic"

# Search
cortex-mcp-server ~/.cortex/memory.db search "Alice"
cortex-mcp-server ~/.cortex/memory.db search -l 10 "Q3 roadmap"

# Stats
cortex-mcp-server ~/.cortex/memory.db stats

Shortened here. Read the whole README on GitHub.

Signals

GitHub stars
33
Forks
2
Last commit
Jun 2026
Advanced
Delivery
cortex MCP server → your ahel gateway (mcp.ahel.ai) → every connected AI client.
Catalog kind
mcp-server
Gateway key
io-github-gambletan-cortex
Source
github.com/gambletan/cortex